Transition frame generation method, device, equipment and storage medium

By acquiring action information and joint velocity prediction, and using the action prediction model to generate transition frames, the problem of unnatural transition frames in the prior art is solved, and high-quality transition frame generation is achieved.

CN115131475BActive Publication Date: 2025-08-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210461991.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-08-26
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

When generating character animations, the prior art, especially when the number of keyframes is small or the actions are complicated, the generation of transition frames is unnatural, unreasonable, and lacks general use.

Method used

By obtaining the action information of the starting frame and the target frame and the target frame number, predict the action change characteristics and joint speed of the target object, use the action prediction model to generate transition frames, including training sub-models of the action prediction model to improve prediction accuracy.

Benefits of technology

A smooth and natural transition frame sequence is generated, which improves the quality of transition frames and adapts to action predictions of different complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131475B_ABST
    Figure CN115131475B_ABST
Patent Text Reader

Abstract

The present application provides a transition frame generation method, device, equipment and storage medium, which are applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving. When generating a transition frame between a starting frame and a target frame, the target object's motion change characteristics and predicted target joint speed are obtained based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target number of frames, so as to predict the transition motion of the target object between the starting frame and the target frame to generate a transition frame. In this process, since different joints have different importance in motion generation, and the target joint speed largely determines the position of the target object in the next frame, more accurate motion information can be obtained by further predicting the transition motion of the target object by obtaining the predicted target joint speed, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a transition frame generation method, apparatus, device and storage medium. Background Art

[0002] With the rapid development of internet technology, the demand for high-quality character animation is growing in industries such as film and gaming. Currently, when creating character animation, the more intensive the character's movements, the smoother the resulting animation. Therefore, after generating keyframes based on the key actions during a character's motion or transformation, the transition movements between adjacent keyframes are often predicted, generating a specified number of transition frames.

[0003] Related technologies use linear interpolation between two adjacent keyframes to predict the target object's transition motion, thereby generating a transition frame between the two keyframes. However, this method is only suitable for motion prediction with short transition times, a small number of transition frames, and simple motions. It lacks versatility and is prone to generating unnatural and unreasonable transition frames when there are only a few keyframes or complex motions.

[0004] Therefore, there is an urgent need for a transition frame generation method that can effectively improve the quality of transition frames. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for generating transition frames, which can obtain a smooth and natural transition frame sequence and effectively improve the quality of transition frames. The technical solution is as follows:

[0006] In one aspect, a method for generating a transition frame is provided, the method comprising:

[0007] Acquire motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number, where the target frame number indicates the number of transition frames between the starting frame and the target frame;

[0008] obtaining a motion change feature of the target object and a predicted target joint velocity based on motion information of the target object in the starting frame, motion information of the target object in the target frame, and the target frame number, the motion change feature indicating a change in a transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame;

[0009] Based on the motion information of the target object in the starting frame, the motion change characteristics of the target object and the predicted target joint speed, the transition motion is predicted to generate a transition frame.

[0010] On the other hand, a transition frame generation device is provided, the device comprising:

[0011] A first acquisition module is configured to acquire motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number, where the target frame number indicates the number of transition frames between the starting frame and the target frame;

[0012] a second acquisition module, configured to acquire a motion change feature of the target object and predict a target joint velocity based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, wherein the motion change feature indicates a change in a transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame;

[0013] The transition frame generation module is used to predict the transition action based on the action information of the target object in the starting frame, the action change characteristics of the target object and the predicted target joint speed to generate a transition frame.

[0014] In some embodiments, the transition frame generation module is configured to:

[0015] Inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into a second sub-model of the motion prediction model to obtain a plurality of sub-motion information of the target object, wherein the sub-motion information is motion information predicted based on the motion stage of the target object;

[0016] Based on multiple target weights, the multiple sub-action information is weightedly summed to obtain the action information of the target object based on the transition action, so as to generate the transition frame.

[0017] In some embodiments, the transition frame generation module is configured to:

[0018] inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model; mapping the motion change characteristics into corresponding motion information based on a mapping relationship between the motion manifold space and the object motion information in the second sub-model, to obtain the predicted joint position and predicted joint velocity of the target object, wherein the motion manifold space indicates the motion change of the object in frames corresponding to two consecutive motions;

[0019] The plurality of sub-action information are obtained based on the action information of the target object in the starting frame, the predicted joint positions and predicted joint velocities of the target object, and the predicted target joint velocities.

[0020] In some embodiments, the apparatus further comprises:

[0021] a first training module, configured to train a second sub-model in the motion prediction model based on a sample data set and label information to obtain the trained second sub-model, wherein the sample data set includes sample frames of a sample object based on a plurality of continuous motions, and the label information indicates a sample target joint velocity of the sample object in the sample frame;

[0022] The second training module is used to train the first sub-model in the action prediction model based on the sample data set and the trained second sub-model to obtain the trained first sub-model.

[0023] In some embodiments, the first training module includes:

[0024] a first training unit, configured to update model parameters of the vector coding model and the second sub-model based on the sample data set, the label information, and the first loss function until a first training condition is satisfied, thereby obtaining an intermediate vector coding model and an intermediate second sub-model, wherein the vector coding model is configured to output a predicted motion change feature of the m+1th sample frame based on the mth sample frame and the m+1th sample frame, where m is a positive integer;

[0025] A second training unit is configured to update model parameters of the intermediate vector coding model and the intermediate second sub-model based on the sample data set, the label information, and a second loss function until a second training condition is satisfied, thereby obtaining the trained vector coding model and the trained second sub-model;

[0026] The first loss function indicates the motion reconstruction loss and information divergence of the sample frame, and the second loss function indicates the motion reconstruction loss, information divergence, footstep sliding loss and bone length loss of the sample frame.

[0027] In some embodiments, the first training unit is configured to:

[0028] Obtaining a motion reconstruction loss value of the (m+1)th sample frame based on the (m)th sample frame, the (m+1)th sample frame, the label information, the vector coding model, and the second sub-model;

[0029] Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame;

[0030] Based on the action reconstruction loss value and the information divergence, the model parameters of the vector coding model and the second sub-model are updated until the first training condition is met, thereby obtaining the intermediate vector coding model and the intermediate second sub-model.

[0031] In some embodiments, the second training unit is configured to:

[0032] Based on the mth sample frame, the (m+1th) sample frame, the label information, the intermediate vector coding model, and the intermediate second sub-model, obtaining a motion reconstruction loss value, a foot sliding loss value, and a bone length loss value of the (m+1th) sample frame;

[0033] Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame;

[0034] Based on the action reconstruction loss value, the foot sliding loss value, the bone length loss value and the information divergence, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

[0035] In some embodiments, the second training module is used to:

[0036] Based on the sample start frame, the sample target frame, the sample target frame number, the first sub-model and the trained second sub-model, obtaining a joint rotation loss value, a joint position loss value and a bone rotation loss value of a sample transition frame between the sample start frame and the sample target frame, wherein the sample target frame number indicates the number of sample transition frames between the sample start frame and the sample target frame;

[0037] Based on the joint rotation loss value, the joint position loss value and the bone rotation loss value, the model parameters of the first sub-model are updated until the training end condition is met, thereby obtaining the trained first sub-model.

[0038] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the transition frame generation method in the embodiment of the present application.

[0039] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the transition frame generation method in the embodiment of the present application.

[0040] In another aspect, a computer program product or computer program is provided, the computer program product or computer program including computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to implement the transition frame generation method of the embodiments of the present application.

[0041] In an embodiment of the present application, when generating a transition frame between a starting frame and a target frame, the target object's motion information in the starting frame, the target object's motion information in the target frame, and the target frame number are used to obtain the target object's motion change characteristics and predict the target joint velocity to predict the target object's transition motion between the starting frame and the target frame to generate a transition frame. In this process, since different joints have different importance in motion generation, and the target joint velocity largely determines the position of the target object in the next frame, by obtaining the predicted target joint velocity to further predict the target object's transition motion, more accurate motion information can be obtained, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0043] Figure 1 Schematic diagram of an implementation environment of a transition frame generation method provided in an embodiment of the present application;

[0044] Figure 2 is a flowchart of a method for generating a transition frame according to an embodiment of the present application;

[0045] Figure 3 is a flowchart of a method for generating a transition frame according to an embodiment of the present application;

[0046] Figure 4 is a schematic diagram of a first sub-model provided in an embodiment of the present application;

[0047] Figure 5 is a schematic diagram of a second sub-model provided in an embodiment of the present application;

[0048] Figure 6 This is a flowchart of a method for training an action prediction model provided in an embodiment of the present application;

[0049] Figure 7 This is a schematic diagram of a joint representation provided by an embodiment of the present application;

[0050] Figure 8 is a schematic diagram of a vector coding model and a second sub-model provided in an embodiment of the present application;

[0051] Figure 9 is a schematic diagram of a transition frame sequence provided in an embodiment of the present application;

[0052] Figure 10 1 is a structural diagram of a transition frame generating device provided according to an embodiment of the present application;

[0053] Figure 11 is a schematic structural diagram of a terminal provided according to an embodiment of the present application;

[0054] Figure 12 It is a structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0056] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0057] In this application, the terms "first," "second," and the like are used to distinguish identical or similar items having substantially the same role and function. It should be understood that "first," "second," and "nth" do not have a logical or temporal dependency, nor do they limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," and the like to describe various elements, these elements should not be limited by these terms.

[0058] These terms are simply used to distinguish one element from another. For example, without departing from the scope of various examples, a first action can be referred to as a second action, and similarly, a second action can also be referred to as a first action. Both the first action and the second action can be actions, and in some cases, can be separate and different actions.

[0059] Here, at least one refers to one or more than one. For example, at least one action can be one action, two actions, three actions, or any other action that is an integer greater than or equal to one. And multiple refers to two or more than two. For example, multiple actions can be two actions, three actions, or any other action that is an integer greater than or equal to two.

[0060] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the action information involved in this application is obtained with full authorization. In some embodiments, the embodiment of the present disclosure provides a permission inquiry page, which is used to inquire whether the permission to obtain the above information is granted. In the permission inquiry page, an authorization consent control and an authorization rejection control are displayed. When a trigger operation of the authorization consent control is detected, the transition frame generation method provided in the embodiment of the present application is used to obtain the above information, thereby realizing the prediction of the object action.

[0061] The following introduces the technologies that may be used in the transition frame generation solution provided in the embodiments of the present application.

[0062] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0063] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0064] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0065] The following introduces the key terms or abbreviations that may be used in the transition frame generation solution provided in the embodiments of the present application.

[0066] Frame is a unit of time used to represent a specific moment.

[0067] Variational Auto-Encoder (VAE) is an unsupervised / semi-supervised neural network architecture that compresses input information into a compact multivariate latent distribution through an encoder, and the decoder recovers the input information from the distribution as accurately as possible.

[0068] Conditional Variational Auto-Encoder (CVAE) is a VAE that introduces conditions and can generate different results under different given conditions.

[0069] Manifold learning is a machine learning method that assumes that the distribution of data in a high-dimensional space is similar to a low-dimensional manifold. The goal of manifold learning is to reduce the dimensionality of the data by finding this low-dimensional manifold.

[0070] Forward Kinematics (FK) refers to the process of determining the position and transformation of a child bone given the position and transformation of a parent bone. For example, when a person moves their arm, they can move their elbow, which in turn moves their palm.

[0071] Mixture of Experts (MOE) is a machine learning model that trains multiple neural networks (that is, multiple experts), each of which is suitable for different features of the data.

[0072] Before introducing the transition frame generation solution provided by the embodiment of the present application, for ease of understanding, the application scenario of the embodiment of the present application is first introduced below.

[0073] Illustratively, the embodiments of the present application can be applied to various scenarios, including but not limited to animation production, cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0074] For example, in an animation production scenario, the more intensive the target object's movements, the smoother the generated animation. Therefore, after generating some sparse key frames based on the key movements in the target object's movement or change process, the missing movements between the key frames are often generated through these key frames to obtain a natural, smooth, and fluid animation. This process can also be understood as, given the starting frame corresponding to the target object's starting movement, the target frame corresponding to the target movement of the target object, and the target number of frames (i.e., the number of transition frames between the starting frame and the target frame), predicting how the target object moves from the starting movement to the target movement according to the target number of frames, thereby generating corresponding transition frames (in-betweening, also called interpolated frames or intermediate frames, not limited here, for ease of description, collectively referred to as transition frames in the following embodiments), and supplementing these transition frames between the starting frame and the target frame to obtain a natural, smooth, and fluid animation.

[0075] It should be understood that the above animation production scenario is only for illustrative purposes. In other scenarios such as film production and video processing, the process of generating transition frames is similar to the above process and will not be repeated here.

[0076] Based on this, an embodiment of the present application provides a transition frame generation method, which can predict the transition action of the target object in real time given the starting frame corresponding to the starting action of the target object, the target frame corresponding to the target action of the target object, and the target number of frames to generate corresponding transition frames, thereby obtaining a natural, smooth, and smooth transition frame sequence and improving the quality of the transition frames.

[0077] The following introduces the implementation environment of the transition frame generation solution provided in the embodiment of the present application.

[0078] Figure 1 1 is a schematic diagram of an implementation environment of a transition frame generation method according to an embodiment of the present application. The implementation environment includes: a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via a wired network or a wireless network, which is not limited in this application.

[0079] Terminal 101 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Illustratively, terminal 101 can install and run an application that provides a transition frame generation function, such as an animation production application, but this is not limited to this. In some embodiments, during the transition frame generation process, a motion prediction model is used to predict the motion of the target object in the transition frame. Terminal 101 can provide server 102 with information required for the training method of the motion prediction model, such as training parameters, sample frames, and an initial AI model.

[0080] In some embodiments, terminal 101 generally refers to one of multiple terminals. This embodiment uses terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals 101 can be greater. For example, if there are dozens, hundreds, or even more terminals 101, the implementation environment of the transition frame generation method may also include other terminals. This embodiment of the present application does not limit the number of terminals or device types.

[0081] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. The number of the above-mentioned servers 102 can be more or less, and the embodiments of the present application are not limited to this. Of course, server 102 can also include other functional servers to provide more comprehensive and diversified services. In some embodiments, server 102 is used to execute the training method of the action prediction model provided in the embodiment of the present application, and the action prediction model is trained based on the information provided by terminal 101.

[0082] In some embodiments, during the transition frame generation process, the server 102 undertakes the main computing work and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work and the terminal 101 undertakes the main computing work; or, the server 102 or the terminal 101 can each undertake the computing work independently.

[0083] It should be noted that both the terminal 101 and the server 102 can generate transition frames using the method provided in the embodiments of the present application. In some embodiments, the terminal 101 sends the starting frame corresponding to the starting action of the target object, the target frame corresponding to the target action, and the target number of frames to the server 102, and the server 102 performs action prediction based on the action prediction model to generate a transition frame between the starting frame and the target frame, and sends the transition frame to the terminal 101. Of course, the terminal 101 can also directly use the action prediction model to perform action prediction to generate a transition frame. It should be noted that the above-mentioned action prediction model can be obtained by training through the terminal 101 or by training through the server 102, and the embodiments of the present application do not limit this.

[0084] In some embodiments, the wired or wireless network described above uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network, or any combination of virtual private networks. In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the above-mentioned data communication technologies.

[0085] The following describes several method embodiments to introduce the transition frame generation method provided in the embodiments of the present application and the training process of the motion prediction model involved in the transition frame generation method.

[0086] Figure 2 This is a flow chart of a method for generating a transition frame according to an embodiment of the present application. Figure 2 As shown, the method is executed by a computer device, which can be provided as the above Figure 1In the terminal or server shown, schematically, the method includes the following steps 201 to 203.

[0087] 201. A computer device obtains motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number.

[0088] In an embodiment of the present application, the starting frame refers to the frame corresponding to the starting action of the target object, the target frame refers to the frame corresponding to the target action of the target object, and the target frame number indicates the number of transition frames between the starting frame and the target frame. Schematically, the target frame number can also be understood as the number of transition actions that the computer device needs to predict during the process of the target object moving from the starting action to the target action. The target object is a person or animal (including a virtual person or virtual animal) with joints, which is not limited to this. The action information of the target object indicates the action posture of the target object in the frame. In some embodiments, the action information includes the joint position, joint rotation, and joint speed of the target object.

[0089] Schematically, taking the target object as a person as an example, the action information S of the target object in any frame i Expressed as ; Where i is a positive integer, p represents joint position, r represents joint rotation, v represents joint velocity, L represents lower body joint, h represents hip joint, and U represents upper body joint.

[0090] 202. The computer device obtains the motion change characteristics of the target object and predicts the target joint velocity based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number.

[0091] In an embodiment of the present application, the motion change feature indicates a change in the transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame. In some embodiments, the transition motion refers to the next motion of the starting motion during the process of the target object moving from the starting motion to the target motion. In some embodiments, the target joint is a hip joint, and schematically, the predicted target joint velocity of the target object refers to the predicted velocity of the hip joint during the process of the target object moving from the starting motion to the transition motion.

[0092] It should be understood that since different joints of the target object have different importance in motion generation, and the hip joint velocity largely determines the position of the next frame, predicting the hip joint velocity can improve the accuracy of the motion information prediction results, obtain more accurate motion information, and improve the quality of transition frames.

[0093] 203. The computer device predicts the transition action based on the action information of the target object in the starting frame, the action change characteristics of the target object and the predicted target joint speed to generate a transition frame.

[0094] In an embodiment of the present application, a computer device predicts the transition action, obtains action information of the target object based on the transition action, and thereby generates a corresponding transition frame.

[0095] It should be noted that in steps 201 to 203 above, the computer device generates a transition frame based on the starting frame, the target frame, and the target frame number. The transition frame is the next frame after the starting frame. In some embodiments, after the computer device generates the transition frame, the transition frame is used as the new starting frame in a similar process to steps 201 to 203 above. Based on the new starting frame, the target frame, and the new target frame number (i.e., the original target frame number - 1), the transition motion of the target object is predicted to generate corresponding transition frames, until the number of transition frames required to supplement between the new starting frame and the target frame is zero, thereby obtaining a set of natural and smooth transition frame sequences. For example, taking the target object as a person, the action of the target object in the starting frame is half-squatting, the action of the target object in the target frame is standing, and the target number of frames is 10 frames, then the computer device repeatedly performs the above steps 201 to 203, predicts the transition action of the target object from "half-squatting" to "standing", and generates 10 transition frames, thereby obtaining a set of smooth and natural transition frame sequences, which can indicate the movement process of the target object from half-squatting to standing. That is, the above process can be understood as a process of generating transition frames in real time. In other embodiments, the computer device generates transition frames corresponding to the target number of frames based on the action information of the target object in the starting frame, the action information of the target object in the target frame, and the target number of frames. This process can be understood as a process of generating transition frames offline. It should be understood that the principle of this process is the same as that of the above steps 201 to 203, so it is not repeated here.

[0096] In summary, the embodiment of the present application provides a transition frame generation method. When generating a transition frame between a starting frame and a target frame, the target object's motion change characteristics and predicted target joint velocity are obtained based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, so as to predict the transition motion of the target object between the starting frame and the target frame to generate a transition frame. In this process, since different joints have different importance in motion generation, and the target joint velocity largely determines the position of the target object in the next frame, by obtaining the predicted target joint velocity to further predict the transition motion of the target object, more accurate motion information can be obtained, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frame.

[0097] According to the above Figure 2 The embodiment shown briefly describes the transition frame generation method provided by this application. Figure 3 The transition frame generation method provided in this application is described in detail.

[0098] Figure 3 This is a flow chart of a method for generating a transition frame according to an embodiment of the present application. Figure 3 As shown, the method is executed by a computer device, which can be provided as the above Figure 1 In the terminal or server shown, schematically, the method includes the following steps 301 to 306.

[0099] 301. A computer device obtains motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number.

[0100] In an embodiment of the present application, the target frame number indicates the number of transition frames between the starting frame and the target frame. Schematically, the computer device is a terminal, and an animation production application is running on the terminal. The computer device responds to the selection operation for the starting frame and the target frame to obtain the action information of the target object in the starting frame and the action information of the target object in the target frame. This embodiment of the present application does not limit this. In some embodiments, the target frame number is set by default, or is set according to actual needs, and this is not limited. In addition, the specific content of the target object's action information is the same as the above Figure 2 The same is true for step 201 in the illustrated embodiment, which will not be described again here.

[0101] After obtaining the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, the computer device predicts the transition motion of the target object based on the following steps 302 to 306 to generate a corresponding transition frame.

[0102] 302. The computer device inputs the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number into the first submodel of the motion prediction model to obtain the offset embedding vector between the starting frame and the target frame, the embedding vector of the starting frame, and the embedding vector of the target frame.

[0103] In an embodiment of the present application, the motion prediction model includes a first sub-model and a second sub-model, wherein the first sub-model is used to predict the target joint speed and motion change characteristics of the object in the next frame (the next frame of the current frame) based on the current frame, the target frame, and the number of transition frames between the current frame and the target frame; the second sub-model is used to predict the motion information of the object in the next frame based on the current frame and the predicted target joint speed and motion change characteristics of the object in the next frame.

[0104] In some embodiments, the computer device inputs the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number into the first sub-model, encodes the motion information of the starting frame through the first sub-model to obtain the embedding vector of the starting frame; encodes the motion information of the target frame to obtain the embedding vector of the target frame; and encodes the offset between the motion information of the starting frame and the target frame to obtain the offset embedding vector.

[0105] In some embodiments, the computer device inputs partial motion information of the target object in the starting frame, partial motion information of the target object in the target frame, and the target frame number into the first sub-model to obtain a corresponding embedding vector. Wherein, the partial motion information indicates the lower body motion information and hip joint motion information of the object. For example, the partial motion information includes the lower body joint position, hip joint position, upper body rotation, lower body joint speed, and hip joint speed. It should be understood that since different joints have different importance in motion generation, the hip joint speed largely determines the position of the target object in the next frame (for example, distinguishing high-speed motion from low-speed motion), and the lower body joints have a greater impact on visual quality. In contrast, the upper body joints have a smaller impact on visual quality. Therefore, by inputting partial motion information into the first sub-model to obtain the corresponding prediction results, it is possible to reduce the amount of data processing and improve the efficiency of transition frame generation on the basis of ensuring the accuracy of the model prediction results.

[0106] 303. The computer device predicts the change of the transition action of the target object relative to the action in the starting frame and the target joint speed of the target object based on the offset embedding vector, the embedding vector of the starting frame, the embedding vector of the target frame, and the target frame number, and obtains the action change feature and the predicted target joint speed.

[0107] In an embodiment of the present application, the motion change feature indicates the change of the transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame. Schematically, the motion change feature is an N-dimensional feature vector (N is a positive integer). In some embodiments, the first sub-model is also used to predict the rotation information of the object in the next frame based on the current frame, the target frame, and the number of transition frames between the current frame and the target frame. Schematically, the computer device predicts the joint rotation of the target object through the first sub-model to obtain the rotation information of the target object, such as upper body joint rotation and hip joint rotation, etc., which is not limited to this.

[0108] In some embodiments, the first sub-model is a model built based on a recurrent neural network (RNN). Figure 4Taking the target joint as the hip joint as an example, the process of the computer device obtaining the motion change characteristics of the target object and predicting the target joint speed through the first sub-model (that is, the above-mentioned steps 302 and 303) is introduced.

[0109] Figure 4 This is a schematic diagram of a first sub-model provided in an embodiment of the present application. Figure 4 As shown, the first sub-model includes a coding layer 401 , a prediction layer 402 and a decoding layer 403 .

[0110] The coding layer 401 is used to encode the target object's motion information in the starting frame, the target object's motion information in the target frame, and the target frame number to obtain a target embedding vector, and input the target embedding vector into the prediction layer 402. Schematically, the coding layer 401 includes a state encoder, a target encoder, and an offset encoder. For example, each of the three encoders includes two hidden layers, the first layer includes 512 units, the second layer includes 256 units, and the activation function uses a PLU activation function, which is not limited. The state encoder is used to encode the target object's motion information in the starting frame (such as ) is encoded; the target encoder is used to encode the action information of the target object in the target frame (such as t represents the target frame); the offset encoder is used to encode the offset between the action information of the starting frame and the target frame (such as ) for encoding.

[0111] The prediction layer 402 is used to predict the hip joint velocity and motion change characteristics of the target object based on the target embedding vector. For example, the prediction layer 402 is a network layer based on a long short-term memory (LSTM) network.

[0112] The decoding layer 403 is used to decode the data output by the prediction layer 402 to obtain the final prediction result. For example, the decoding layer 403 includes a Parse decoder, which includes three hidden layers and uses an ELU activation function, which is not limited to this.

[0113] In some embodiments, the computer device encodes the target frame number into a time embedding vector and adds the time embedding vector to the target embedding vector. Schematically, the time embedding vector is similar to the position encoding, and its principle is referred to the following formula (1) and formula (2):

[0114]

[0115]

[0116] Where z dt represents the time embedding vector, dt represents the target frame number, D represents the vector dimension, and k represents the vector dimension index.

[0117] In some embodiments, the computer device further adds time-varying Gaussian noise to the target embedding vector. Schematically, the time-varying Gaussian noise is represented as z target , whose variance is equal to 0.5. The amplitude of the time-varying Gaussian noise decreases according to the following formula (3). By adding the time-varying Gaussian noise to the target embedding vector, the attention of the first sub-model is focused on the target frame only when it is close to the target frame, which can improve the robustness of the first sub-model.

[0118]

[0119] Where dt represents the target frame number, t zero Indicates the number of noise-free interval frames close to the target frame, t period Indicates the number of frames at which the noise decreases linearly. For example, t zero Set it to 5 frames, and set t period It is set to 30 frames, which can be set according to needs in actual applications, and is not limited in this embodiment of the present application.

[0120] In some embodiments, the computer device uses the hyperbolic tangent Tanh activation function through the decoding layer 203 in the first sub-model to output the target object's motion change characteristics. In this way, the output results can be converted into sample values ​​of a standard normal distribution (for example, the output results are expanded by 4.5 times through the Tanh function), thereby ensuring that the motion change characteristics can cover the range of the normal distribution and improving the accuracy of the object motion information prediction results.

[0121] In some embodiments, the computer device predicts the joint rotation of the target object through the first sub-model to obtain the upper body joint rotation offset of the target object. and hip rotation deviation Etc., so that the computer device can combine the upper body joint rotation and hip joint rotation of the target object in the starting frame to obtain the upper body joint rotation and hip joint rotation of the target object based on the transition action, which is not limited in this embodiment of the present application.

[0122] It should be noted that the above Figure 4 The structure of the first sub-model shown is only for illustration. In some embodiments, the first sub-model may also be a model with other structures. The embodiments of the present application do not limit the structure of the first sub-model.

[0123] After the above steps 302 and 303, the computer device obtains the target object's motion change characteristics and predicted target joint speed by inputting the target object's motion information in the starting frame, the target object's motion information in the target frame, and the target frame number into the first sub-model.

[0124] The following describes the process of the computer device further predicting the motion information of the transition motion through the second sub-model through steps 304 and 305.

[0125] 304. The computer device inputs the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model of the motion prediction model to obtain multiple sub-motion information of the target object.

[0126] In an embodiment of the present application, the sub-action information is action information predicted based on the motion stage of the target object. Schematically, the motion stage refers to the motion stage corresponding to the object's motion posture. For example, taking a walking action of the object as an example, the motion stage corresponding to the walking action includes moving the right leg forward, pushing off the ground with the left leg, and bending the right knee. In some embodiments, the target object is in multiple motion stages. The computer device uses the second sub-model to predict the first sub-action information of the target object based on the first motion stage of the target object, and predicts the second sub-action information of the target object based on the second motion stage of the target object, and so on. The embodiment of the present application does not limit the specific prediction process of the sub-action information.

[0127] In some embodiments, this step 304 includes the following steps A and B:

[0128] Step A: Input the motion information of the target object in the starting frame, the motion change characteristics of the target object and the predicted target joint velocity into the second sub-model; based on the mapping relationship between the motion manifold space and the object motion information in the second sub-model, map the motion change characteristics into corresponding motion information to obtain the predicted joint position and predicted joint velocity of the target object.

[0129] Among them, the action manifold space indicates the action changes of the object in the frames corresponding to two consecutive actions. The action change feature is also a data point in the action manifold space. The second sub-model can map the action change feature to the corresponding action information based on the mapping relationship between the action manifold space and the object action information to obtain the predicted joint position and predicted joint speed of the target object. In other words, the second sub-model can map the data points in the action manifold space to the action posture of the object. Schematically, the computer device trains the second sub-model based on the manifold learning method, so that it learns the above-mentioned action manifold space (this training process will be discussed later). Figure 6Detailed description is given in the embodiment shown and will not be repeated here).

[0130] In some embodiments, the second sub-model is a model constructed based on the decoder of CVAE. It should be understood that in the related art, CVAE includes an encoder and a decoder, wherein the encoder is used to encode the sample X, output the log value of the mean and variance of the normal distribution, and then sample the variable z from the normal distribution, and the decoder decodes it according to the control condition and the variable z to reconstruct the sample X. In an embodiment of the present application, the computer device trains CVAE based on the manifold learning method so that it learns the above-mentioned action popular space. After the training is completed, the decoder in CVAE is used as the above-mentioned second sub-model to participate in the transition frame generation process. Through the second sub-model, the transition action of the target object is predicted based on the action change characteristics of the target object (i.e., variable z) and the predicted target joint speed (i.e., control condition).

[0131] The following introduces the principle of the above-mentioned action manifold space based on formulas (4) to (7).

[0132] Schematically, taking any frame sequence of an object as an example, the starting frame is represented as S 0 , the target frame is represented as S t , starting frame S 0 and target frame S t The transition frame set between 1 ,...,S t-1}, the target joint is the hip joint. For the i-th frame in the frame sequence (i is a positive integer), the action information of the i-th frame is expressed as Based on this, the joint probability of M is shown in the following formula (4):

[0133] P(M)=∫∫∫P(M|S 0 , S t , z dt )P(S 0 , S t , z dt )dS 0 dS t dz dt (4)

[0134] Where z dt Represents the temporal embedding vector obtained after encoding the target number of frames. Based on the Markov hypothesis, the above formula (4) is decomposed to obtain the following formula (5):

[0135]

[0136] By introducing the variable z to encode the common embedding vector of two consecutive frames, and the hip joint velocity of the next frame As a control condition, the probability of the object's motion change transition in two consecutive frames is expressed as follows:

[0137]

[0138] Furthermore, since different joints have different importance in motion generation, the hip joint velocity largely determines the position of the target object in the next frame, and the lower body joints have a greater impact on visual quality. In contrast, the upper body joints have a smaller impact on visual quality. Therefore, by focusing on learning the motion information corresponding to the lower body joints and hip joints in the motion information, the above formula (6) is transformed to obtain the following formula (7):

[0139]

[0140] By training the CVAE, the CVAE decoder learns the above formula (7), that is, learns the above action manifold space. When predicting the object's action, the computer device uses the trained CVAE decoder as the second sub-model to obtain the predicted joint positions and predicted joint velocities of the target object.

[0141] Step B: obtaining the plurality of sub-action information based on the action information of the target object in the starting frame, the predicted joint positions and predicted joint velocities of the target object, and the predicted target joint velocities.

[0142] Among them, the second sub-model includes a hybrid expert network, and the computer device predicts the transition action of the target object based on the action information of the target object in the starting frame, the predicted joint position of the target object, the predicted joint speed, the predicted target joint speed and the motion stage of the target object through the hybrid expert network, thereby obtaining multiple sub-action information. This process can also be understood as using multiple expert networks in the hybrid expert network to predict the transition action of the target object based on multiple motion stages corresponding to the action of the target object (such as a certain expert network focusing on learning the motion stage with the left foot in front, and another expert network focusing on learning the motion stage with the right foot in the back, etc., which are not limited to this), thereby obtaining multiple sub-action information. In this way, the action changes of the target object in two consecutive frames are modeled as multi-mode mapping between frames, thereby improving the accuracy of the action information prediction results.

[0143] In some embodiments, the computer device predicts the joint rotation of the target object through the second sub-model to obtain rotation information of the target object, such as the joint rotation of the lower body, which is not limited to this.

[0144] 305. The computer device performs weighted summation on the multiple sub-action information based on multiple target weights to obtain action information of the target object based on the transition action.

[0145] In an embodiment of the present application, the second sub-model further includes multiple gating networks configured to output target weights indicating the importance of different motion phases to motion prediction. The computer device, through the hybrid expert network in the second sub-model, performs a weighted summation of the multiple sub-action information based on the multiple target weights to obtain motion information of the target object based on the transition motion. Illustratively, the motion information obtained by the computer device through the second sub-model includes the target object's lower body joint velocities and lower body joint rotations.

[0146] Reference below Figure 5 Taking the target joint as the hip joint as an example, the process of the computer device obtaining the motion information of the target object through the second sub-model (that is, the above-mentioned steps 304 and 305) is introduced.

[0147] Figure 5 This is a schematic diagram of a second sub-model provided in an embodiment of the present application. Figure 5 As shown, the second sub-model includes an input layer 501, a gate layer 502, a hybrid expert layer 503 and an output layer 504. The output layer 501 is used to input the action information of the target object in the starting frame (such as ), the target object's motion change characteristics (i.e., variable z) and the predicted hip joint velocity (i.e., ). The gating layer 502 is used to output the target weight corresponding to each expert network. For example, the gating layer adopts the Softmax activation function, which is not limited to this. The hybrid expert layer 503 is used to predict the action of the target object based on the motion stage of the target object, obtain multiple sub-action information, and perform weighted summation on the multiple sub-action information based on multiple target weights to obtain the action information of the target object. For example, each expert network is a three-layer 256-unit feedforward network, which adopts the ELU activation function, which is not limited to this. The output layer 504 is used to output the action information of the target object (such as ).

[0148] It should be noted that the above Figure 5 The structure of the second sub-model shown is only for illustration. In some embodiments, the second sub-model may also be a model with other structures. The embodiments of the present application do not limit the structure of the second sub-model.

[0149] After the above steps 304 and 305, the computer device obtains the motion information of the target object based on the transition motion by inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object and the predicted target joint velocity into the second sub-model.

[0150] 306. The computer device generates a transition frame based on the motion information of the target object.

[0151] In the embodiment of the present application, the process of the computer device generating the transition frame is the same as that described above. Figure 2 The same is true for step 203 in the illustrated embodiment, so it will not be described again here.

[0152] Through the above steps 301 to 306, taking the starting frame as an example, the process of a computer device generating a transition frame (the next frame of the starting frame) is introduced. After the computer device generates the transition frame, the transition frame is used as the new starting frame according to the same process as the above steps 301 to 306, and the above steps 301 to 306 are repeated until the number of transition frames required to be supplemented between the starting frame and the target frame is zero, thereby obtaining a set of natural and smooth transition frame sequences.

[0153] In summary, the embodiment of the present application provides a transition frame generation method. When generating a transition frame between a starting frame and a target frame, the target object's motion change characteristics and predicted target joint velocity are obtained based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, so as to predict the transition motion of the target object between the starting frame and the target frame to generate a transition frame. In this process, since different joints have different importance in motion generation, and the target joint velocity largely determines the position of the target object in the next frame, by obtaining the predicted target joint velocity to further predict the transition motion of the target object, more accurate motion information can be obtained, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frame.

[0154] Based on the above Figures 2 to 5 The embodiment shown in the figure introduces the method for generating a transition frame provided by the embodiment of the present application. Figure 6 , the training process of the action prediction model involved in the above method is introduced.

[0155] Figure 6 This is a flow chart of a method for training an action prediction model provided in an embodiment of the present application. Figure 6 As shown, the method is executed by a computer device, which can be provided as the above Figure 1 In the terminal or server shown, schematically, the method includes the following steps 601 to 604.

[0156] 601. The computer device obtains a sample data set and label information.

[0157] In an embodiment of the present application, the sample data set includes sample frames of a sample object based on multiple continuous actions, and the label information indicates the sample target joint velocity of the sample object in the sample frame. In some embodiments, the sample data set includes multiple continuous sample frames corresponding to an action. In other embodiments, the sample data set includes multiple actions, each action includes multiple continuous sample frames, which is not limited to this. For example, the sample data set includes the open source dataset Lafan1 dataset and / or the Human3.6M dataset.

[0158] In some embodiments, the computer device uses different representations for different joints based on the sample data set. Figure 7 , Figure 7 This is a schematic diagram of a joint representation provided by an embodiment of the present application. Figure 7 As shown in Figure (a), the sample objects in the Human3.6M dataset have 21 joints; Figure 7 As shown in Figure (b), the sample objects in the Lafan1 dataset have 22 joints. Among them, a position-based representation is used for the 8 lower body joints, and a rotation-based representation is used for the upper body joints. All lower body joints are connected to less than two other joints to determine their direction. In addition, a 2-axis rotation matrix representation is used to represent the joint rotation, which includes a 3-dimensional up vector and a forward vector. In order to facilitate the conversion of joint position to joint rotation, in addition to using a 3-dimensional vector to represent the joint velocity, a vector is also used to represent the up direction of the joint. In some embodiments, a 2-axis rotation matrix representation is used instead of the up vector to achieve unified representation.

[0159] It should be noted that in an embodiment of the present application, the action prediction model includes a first sub-model and a second sub-model. In the process of training the action prediction model, the computer device first trains the second sub-model, and then trains the first sub-model based on the trained second sub-model. In the process of training the second sub-model, a two-stage training method is adopted. In the first training stage, based on the sample data set, the label information and the first loss function, the model parameters of the vector coding model and the second sub-model (that is, the encoder and decoder of CVAE) are updated to obtain an intermediate vector coding model and an intermediate second sub-model; in the second training stage, based on the sample data set, the label information and the second loss function, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated to obtain the trained vector coding model and the trained second sub-model. Through this two-stage training method, the accuracy of the action information prediction results can be further improved while ensuring that the model predicts approximate actions.

[0160] The following introduces the above two-stage training method based on step 602 and step 603.

[0161] 602. The computer device updates the model parameters of the vector coding model and the second sub-model based on the sample data set, the label information and the first loss function until the first training condition is met, thereby obtaining an intermediate vector coding model and an intermediate second sub-model.

[0162] In an embodiment of the present application, the vector coding model is used to output a predicted motion change feature of the m+1th sample frame based on the mth sample frame and the m+1th sample frame, where m is a positive integer. The first loss function indicates the motion reconstruction loss and information divergence of the sample frame. Schematically, the vector coding model is a model constructed based on the CVAE encoder.

[0163] For example, reference Figure 8 , Figure 8 Schematic diagram of a vector coding model and a second sub-model provided in an embodiment of the present application, such as Figure 8 As shown, taking the target joint as the hip joint as an example, the vector coding model 801 is a model constructed by the encoder based on CVAE, and the second sub-model 802 is a model constructed by the decoder based on CVAE. Among them, the vector coding model 801 is used to encode according to two consecutive sample frames (such as the m-th sample frame and the m+1-th sample frame, where m is a positive integer) to obtain the mean and variance of the normal distribution, and sample the variable z from the normal distribution (that is, the predicted motion change feature of the m+1-th sample frame); the second sub-model 802 is used to decode according to the motion information of the m-th sample frame, the variable z and the sample hip joint velocity of the m+1-th sample frame to obtain the motion information of the m+1-th sample frame.

[0164] The following describes the above training process using the e-th iteration (e is a positive integer) of the training process as an example. Schematically, the training process includes the following steps A to C:

[0165] Step A: Based on the m-th sample frame, the m+1-th sample frame, the label information, the vector coding model and the second sub-model, obtain the motion reconstruction loss value of the m+1-th sample frame.

[0166] The computer device calculates the motion reconstruction loss value of the m+1th sample frame based on the mean square error between the predicted motion information and the true motion information of the m+1th sample frame. Schematically, the motion reconstruction loss value is calculated using the following formula (8):

[0167]

[0168] Where, represents the predicted joint positions of the lower body joints, p L represents the true joint positions of the lower body joints, represents the predicted joint rotation of the lower body joints, r L Represents the true joint rotation.

[0169] Step B: Based on the mth sample frame, the (m+1)th sample frame and the vector coding model, obtain the information divergence of the (m+1)th sample frame.

[0170] The information divergence is also known as the KL divergence (Kullback-Leibler divergence), which is used to constrain the predicted motion change characteristics of the m+1th sample frame output by the vector coding model so that its distribution is close to the Gaussian distribution. Schematically, the information divergence is calculated using the following formula (9):

[0171] L kl =-0.5(1+σ-μ 2 -e σ ) (9)

[0172] Where μ and σ are the logarithms of the mean and variance.

[0173] Step C: Based on the action reconstruction loss value and the information divergence, the model parameters of the vector coding model and the second sub-model are updated until the first training condition is met, thereby obtaining the intermediate vector coding model and the intermediate second sub-model.

[0174] The computer device calculates the loss value based on the action reconstruction loss value and the information divergence according to the first loss function. When the loss value satisfies the first training condition, the intermediate vector coding model and the intermediate second sub-model are output. When the loss value does not meet the first training condition, the model parameters of the vector coding model and the second sub-model are adjusted, and the e+1th iteration is performed based on the adjusted vector coding model and the second sub-model until the first training condition is met. In some embodiments, the first training condition refers to the number of iterations reaching the target number or the loss value being less than a set threshold, etc., which is not limited to this. Schematically, the first loss function is shown in the following formula (10):

[0175] L loss1 =L rec +L kl (10)

[0176] Where, L rec is the action reconstruction loss value, L kl is the information divergence of the model.

[0177] In some embodiments, during any iteration process, the computer device takes multiple sample frames as input, obtains the total loss value corresponding to the multiple sample frames, and updates the model parameters of the vector coding model and the second sub-model based on the total loss value until the first training condition is met. This embodiment of the present application is not limited to this.

[0178] After the above step 602, the first stage of training of the vector coding model and the second sub-model is achieved. This process can also be understood as training the model to predict approximate actions, that is, ensuring that the similarity between the model's predicted results and the actual results meets certain conditions.

[0179] 603. The computer device updates the model parameters of the intermediate vector coding model and the intermediate second sub-model based on the sample data set, the label information and the second loss function until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

[0180] In an embodiment of the present application, the second loss function indicates motion reconstruction loss, information divergence, foot sliding loss, and bone length loss of the sample frame.

[0181] The following describes the training process using the fth iteration (f is a positive integer) of the training process as an example. Schematically, the training process includes the following steps A to C:

[0182] Step A: Based on the mth sample frame, the m+1th sample frame, the label information, the intermediate vector coding model and the intermediate second sub-model, obtain the motion reconstruction loss value, foot sliding loss value and bone length loss value of the m+1th sample frame.

[0183] The calculation method of the motion reconstruction loss value is the same as that of the above step 602.

[0184] The computer device calculates the foot sliding loss value of the m+1th sample frame based on the predicted relative hip joint velocity and the actual hip joint velocity of the m+1th sample frame. Schematically, the foot sliding loss value is calculated using the following formula (11):

[0185]

[0186] Where, represents the predicted relative hip joint velocity of the subject’s foot contacting the ground relative to the hip joint, v h Indicates the true hip velocity.

[0187] The computer device calculates the bone length loss value of the m+1th sample frame based on the predicted joint position and the actual joint position of the m+1th sample frame. Schematically, the bone length loss value is calculated using the following formula (12):

[0188]

[0189] Where p j , p k Represents any two adjacent joints in the lower body of the sample object.

[0190] Step B: Based on the mth sample frame, the (m+1)th sample frame and the vector coding model, obtain the information divergence of the (m+1)th sample frame.

[0191] The process of the computer device obtaining the information divergence is the same as that of the above step 602, so it will not be repeated here.

[0192] Step C: Based on the action reconstruction loss value, the foot sliding loss value, the bone length loss value and the information divergence, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

[0193] The computer device calculates the loss value based on the second loss function, the action reconstruction loss value, the foot sliding loss value, the bone length loss value, and the information divergence. If the loss value meets the second training condition, the trained vector coding model and the second sub-model are output. If not, the model parameters of the intermediate vector coding model and the intermediate second sub-model are adjusted, and the f+1th iteration is performed based on the adjusted intermediate vector coding model and the intermediate second sub-model until the second training condition is met. In some embodiments, the second training condition refers to the number of iterations reaching the target number or the loss value being less than a set threshold, etc., which is not limited to this. Schematically, the second loss function is shown in the following formula (13):

[0194] L loss2 =L rec +L kl +L foot +L bone (13)

[0195] Where, L foot is the foot sliding loss, L bone is the bone length loss, L rec is the motion reconstruction loss, the mean square error between the predicted pose and the actual pose, L kl is the information divergence.

[0196] After step 603, the second stage of training for the vector encoding model and the second sub-model is completed. Because bone length loss can limit joint velocities to zero, resulting in strange motion sequences, the second stage of training can avoid this situation, thereby improving the accuracy of the model's prediction results.

[0197] In addition, during the training process shown in the above steps 602 and 603, the computer device divides the sample data set into multiple 50-frame windows, and evenly divides each 50-frame window into two 25-frame windows, and trains in groups of 25 sample frames, thereby improving the convergence efficiency of the model.

[0198] In some embodiments, during the above-mentioned training process, the computer device uses the motion information of the m+1 sample frames obtained by model prediction as input to predict the motion information of the m+2th sample frame. In some embodiments, the computer device performs training based on the sample data set during the above-mentioned training process, that is, each input is the real motion information of the sample frame. In other embodiments, the computer device adopts a sampling strategy and performs training based on the sample data set to increase the robustness of the model. For example, sampling is performed from the sample data set based on the target probability. When training starts, the target probability is set to 1, indicating that all is provided by the sample data set. After y (y is a positive integer) epochs (Epochs), the target probability is adjusted based on a linear function until the target probability is 0. For example, for the Lafan1 data set, y is set to 5, and for the Human3.6M data set, y is set to 20. The embodiments of the present application are not limited to this.

[0199] In some embodiments, during the above training process, the computer device uses the AMSgrad optimizer. In the first training phase, the learning rate is initialized to 1e-4 and linearly reduced to 1e-5 through 50,000 iterations. In the second training phase, the learning rate is initialized to 0 and increased from 0. After 10 epochs, it is increased to 1e-4 to ensure that the newly added loss does not significantly change the model parameters. The learning rate is then reduced at the same rate as in the first training phase. For example, in the AMSgrad optimizer, β1 = 0.5 and β2 = 0.9, which is not limited in the embodiments of the present application.

[0200] In some embodiments, during the training process, the computer device adjusts the weight ratio of each loss value to improve the accuracy of the model prediction results. For example, the weight ratio of each loss value is adjusted to 1 (or approximately equal to 1), which is not limited in this embodiment of the present application.

[0201] After the above steps 602 and 603, the computer device trains the trained vector encoding model and the second sub-model based on manifold learning, that is, the computer device trains CVAE to enable it to learn the action manifold space. It should be understood that although CVAE learns the action manifold space, it can only perform an uncontrolled generation process. Therefore, by training the following first sub-model, it learns how to sample frames from the action manifold space, so that the first sub-model can generate the i+1 frame under the constraints of the target frame and the target number of frames when the i-th frame is given. That is, the first sub-model is equivalent to a sampler. In order to generate the i+1 frame under the constraints of the target frame and the target number of frames given the i-th frame, the first sub-model is required to sample the action change feature (variable z) from the action manifold space and predict the target joint velocity (such as hip joint velocity) of the object in the i+1-th frame. ).

[0202] The following describes the training process of the first sub-model based on step 604.

[0203] 604. The computer device trains the first sub-model in the action prediction model based on the sample data set and the trained second sub-model to obtain the trained first sub-model.

[0204] In an embodiment of the present application, a computer device connects the output layer of the first sub-model to the input layer of the trained second sub-model, and trains the first sub-model based on the sample data set. During this process, only the model parameters of the first sub-model are updated until the training end conditions are met, thereby obtaining the trained first sub-model.

[0205] The following describes the training process using the g-th iteration (g is a positive integer) of the training process as an example. Schematically, the training process includes the following steps A and B:

[0206] Step A: Based on the sample start frame, the sample target frame, the sample target frame number, the first sub-model and the trained second sub-model, obtain the joint rotation loss value, the joint position loss value and the bone rotation loss value of the sample transition frame between the sample start frame and the sample target frame.

[0207] The sample target frame number indicates the number of sample transition frames between the sample start frame and the sample target frame. The computer device calculates the joint rotation loss value of the sample transition frame based on the predicted joint rotations of the lower body joints, the predicted joint rotations of the upper body joints, the actual joint rotations of the lower body joints, and the actual rotations of the upper body joints of the sample object in the sample transition frame. For example, the joint rotation loss value is defined as the L1 norm of all joint losses in the rotation space, and is calculated by the following formula (14):

[0208]

[0209] Where, represents the predicted joint rotation of the lower body joints, r L represents the true joint rotation of the lower body joints, represents the predicted joint rotation of the upper body joint, r U Represents the true joint rotations of the upper body joints.

[0210] In some embodiments, the joint position loss value is calculated based on all joint positions of the sample object in the sample transition frame. In other embodiments, the joint position loss value is calculated based on the joint positions of the lower body of the sample object in the sample transition frame, and the embodiments of the present application are not limited to this. Schematically, taking the calculation of the joint position loss value based on the joint positions of the lower body as an example, the computer device calculates the joint position loss value of the sample transition frame based on the predicted joint positions and the actual joint positions of the lower body joints of the sample object in the sample transition frame. For example, the joint position loss value is calculated by the following formula (15):

[0211]

[0212] Where, represents the predicted joint position, p L represents the true joint position.

[0213] The computer device converts the rotation space of all joints into the position space based on the FK algorithm, and calculates the bone rotation loss value of the sample transition frame using the following formula (16).

[0214]

[0215] Where, represents the joint position determined based on the predicted joint rotation, and p represents the true joint position.

[0216] Step B: Based on the joint rotation loss value, the joint position loss value and the bone rotation loss value, the model parameters of the first sub-model are updated until the training end condition is met, thereby obtaining the trained first sub-model.

[0217] Wherein, the computer device calculates the loss value based on the joint rotation loss value, the joint position loss value and the bone rotation loss value according to the third loss function, and outputs the trained first sub-model when the loss value meets the training end condition. If it does not meet the condition, the model parameters of the first sub-model are adjusted, and the g+1th iteration is performed based on the adjusted first sub-model until the training end condition is met. In some embodiments, the training end condition refers to the number of iterations reaching the target number, such as 3 million times, which is not limited to this. Schematically, taking the joint position loss value calculated based on the lower body joint position as an example, the above-mentioned third loss function is shown in the following formula (17):

[0218] L loss3 =L rot +L leg +L pos,rot (17)

[0219] Where, L rot is the joint rotation loss value, L leg is the joint position loss value, L pos,rot is the bone rotation loss value.

[0220] In some embodiments, the third loss function further includes step slip loss and bone length loss, thereby further improving the accuracy of the model prediction results. It should be noted that the calculation method of step slip loss and bone length loss is similar to that of step 603 above, so it will not be repeated here.

[0221] In addition, during the training process shown in the above step 604, the computer device divides the sample data set into multiple 50-frame windows, and samples from the windows based on different sample target frame numbers (such as 5 frames to 30 frames), so that the model can be trained based on different transition lengths and target frames, thereby improving the accuracy of the motion information prediction results.

[0222] In some embodiments, during the above training process, the computer device uses the AMSgrad optimizer to initialize the learning rate to 1e-3, which is not limited.

[0223] Through the training method shown in steps 601 to 604 above, the computer device trains the vector encoding model and the second sub-model of the motion prediction model using a two-stage training method based on manifold learning, enabling them to learn the motion manifold space. This improves the accuracy of subsequent motion information prediction results while ensuring that the second sub-model can predict approximate motions. Furthermore, based on the trained second sub-model, the first sub-model of the motion prediction model is trained to learn how to sample frames from the motion manifold space, thereby improving the accuracy of motion information prediction results and providing technical support for generating natural and smooth transition frame sequences.

[0224] The following is the above Figures 2 to 5 The object action method shown above Figure 6 Taking the training method shown as an example, based on the experimental results of the embodiment of the present application and related solutions, the beneficial effects brought by the embodiment of the present application are explained.

[0225] Schematically, the relevant solutions compared with this application include: a standard linear interpolation solution, RTN (Robust Motion In-betweening, a transition frame generation solution), RTN with added foot sliding loss (+skatingloss, referred to as SL), a solution obtained by replacing the CVAE involved in the embodiment of this application with AE (autoencoder), and a solution obtained by replacing the CVAE involved in the embodiment of this application with VAE. The effect of each solution is reflected by the following three indicators: the L2 norm of the global position of the joint, the normalized power spectrum similarity (NPSS), and the foot sliding index.

[0226] This application tests the solutions provided in the embodiments of this application and related solutions based on the Lafan1 dataset. Schematically, the Lafan1 dataset contains 5 sample objects, 77 action sequences, and a total of 496,672 action frames. The fifth sample object is used as the test set for testing.

[0227] The present application is tested based on different target frame numbers (i.e., the number of transition frames). Schematically, the target frame numbers are set to 5 frames, 15 frames, and 30 frames, respectively. Tables 1 to 3 are the test results of the schemes and related schemes provided in the embodiments of the present application based on the above three indicators. Among them, Table 1 is the L2 norm of the global position of the joint under different target frame numbers, Table 2 is the NPSS index under different target frame numbers, and Table 3 is the footstep sliding index under different target frame numbers. Based on the following Tables 1 to 3, it is concluded that the transition frame generation method provided in the embodiments of the present application can generate high-quality transition frame sequences under different target frame numbers.

[0228] Table 1

[0229] Target frame rate 5 15 30 Interpolation 0.37 1.24 2.31 RTN 0.22 0.59 1.16 RTN(+SL) 0.28 0.68 1.27 AE 0.28 0.63 1.16 VAE 0.20 0.56 1.11 This application 0.20 0.56 1.12

[0230] Table 2

[0231]

[0232]

[0233] Table 3

[0234] Target frame rate 5 15 30 Interpolation 1.708 2.081 2.144 RTN 0.483 0.698 0.930 RTN(+SL) 0.249 0.349 0.455 AE 0.294 0.485 0.649 VAE 0.255 0.353 0.502 This application 0.244 0.343 0.469

[0235] In addition, the present application is tested based on different types of action sequences. Schematically, the running and walking actions in the Lafan1 data set are classified as walking sequences, the dance actions, fighting actions and sports actions are classified as dancing sequences, all jumping actions are classified as jumping sequences, and the actions of all characters avoiding obstacles are classified as obstacle sequences, etc. Taking the target number of frames as 30 frames as an example, the various indicators are shown in Tables 4 to 6 below, wherein Table 4 is the L2 norm of the global position of the joints under different action sequences, Table 5 is the NPSS index under different action sequences, and Table 6 is the foot sliding index under different action sequences. Based on Tables 4 to 6 below, it is concluded that the transition frame generation method provided in the embodiments of the present application can generate high-quality transition frame sequences under different action sequences.

[0236] Table 4

[0237] action walk Dance jump Interpolation 2.76 2.40 1.89 RTN 0.99 1.51 1.21 This application 0.95 1.48 1.18

[0238] Table 5

[0239] action walk Dance jump Interpolation 0.6430 0.6405 0.4000 RTN 0.3380 0.5197 0.3123 This application 0.3306 0.5141 0.3205

[0240] Table 6

[0241] action walk Dance jump Interpolation 2.743 1.844 1.381 RTN 1.187 1.103 0.640 This application 0.589 0.571 0.326

[0242] Furthermore, the present application was tested based on different time lengths. Schematically, by discarding the 30-frame sequence of each sample, the network was required to recover it using 8 frames, 15 frames, 60 frames and 100 frames respectively, and the footstep sliding indicators of different methods were measured. The test results are shown in Table 7 below. It can be seen that under the condition of 100 frames, since the interval between the two frames is very short, the footstep difference between the two frames is reduced by 100 times for the interpolation method, so the indicator is the best. For all other time lengths, the present application is significantly better than RTN. When the time is very short, it is difficult to generate reasonable actions, but the present application can still obtain high-quality transition frame sequences. Visually, when the starting frame and the target frame are in the same motion phase of a walking cycle (for example, both are left feet in front), RTN is likely to generate a sliding action, while the present application can quickly generate a small step to make it walk to the position corresponding to the target frame.

[0243] Table 7

[0244] time 8 15 60 100 Interpolation 7.302 3.917 1.075 0.004 RTN 4.050 2.087 0.814 0.522 This application 3.363 1.350 0.438 0.110

[0245] The present application also conducted tests based on different target positions. Schematically, the present application conducted tests based on the following two situations: one is to move the target position forward to a position twice as far as the initial target position, and the other is to move the target position forward in the opposite direction of the initial position.

[0246] In the forward test, the present application tends to generate several large steps or more steps to compensate for the extra distance, while the RTN always generates the same step size as the dataset, compensating the distance by sliding. In the backward test, the RTN always generates recognizable sliding, while the present application does not. This is because the action manifold space involved in the embodiments of the present application can successfully capture the backward-moving speed in the dataset.

[0247] In addition, this application also tested an extreme example, setting the target frame to 10m away from the starting frame and letting the sample object move to it in 60 frames. In the dataset, the farthest distance is only 5.79m and the longest time length is only 30 frames. The specific test results are as follows Figure 9 As shown, Figure 9 Schematic diagram of a transition frame sequence provided by an embodiment of the present application. Figure 9 As shown, due to the long distance, RTN (first row) generates a sliding action, and the result generated by VAE (second row) is a mixture of crawling and running actions. However, since the hip joint speed is used as a control condition in this application (third row), it is able to generate a transition frame sequence of fast running, and the transition frame sequence has a higher quality.

[0248] As can be seen, the transition frame generation method provided by the embodiments of the present application can generate reasonable and natural results even when input conditions change, or even become unreasonable (for example, when the time interval is very long or very short, or when the target is in the opposite direction of the initial position or is very far away from the initial position). Moreover, compared to other methods, the generated transition frames are essentially free of errors such as foot slipping.

[0249] In summary, the embodiment of the present application provides a transition frame generation method. When generating a transition frame between a starting frame and a target frame, the target object's motion change characteristics and predicted target joint velocity are obtained based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, so as to predict the transition motion of the target object between the starting frame and the target frame to generate a transition frame. In this process, since different joints have different importance in motion generation, and the target joint velocity largely determines the position of the target object in the next frame, by obtaining the predicted target joint velocity to further predict the transition motion of the target object, more accurate motion information can be obtained, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frame.

[0250] Figure 10 This is a schematic diagram of the structure of a transition frame generation device provided according to an embodiment of the present application. The device is used to execute the steps of the above-mentioned transition frame generation method. Figure 10The transition frame generation device includes: a first acquisition module 1001, a second acquisition module 1002 and a transition frame generation module 1003.

[0251] A first acquisition module 1001 is configured to acquire motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number, where the target frame number indicates the number of transition frames between the starting frame and the target frame;

[0252] A second acquisition module 1002 is configured to acquire a motion change feature of the target object and predict a target joint velocity based on the motion information of the target object in the start frame, the motion information of the target object in the target frame, and the target frame number, wherein the motion change feature indicates a change in a transition motion of the target object between the start frame and the target frame relative to the motion in the start frame;

[0253] The transition frame generating module 1003 is configured to predict the transition motion based on the motion information of the target object in the starting frame, the motion change characteristics of the target object and the predicted target joint velocity, so as to generate a transition frame.

[0254] In some embodiments, the second acquisition module 1002 is configured to:

[0255] Inputting the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number into a first sub-model of the motion prediction model to obtain an offset embedding vector between the starting frame and the target frame, an embedding vector of the starting frame, and an embedding vector of the target frame;

[0256] Based on the offset embedding vector, the embedding vector of the starting frame, the embedding vector of the target frame, and the target frame number, the change of the transition movement of the target object relative to the movement in the starting frame and the hip joint velocity of the target object are predicted to obtain the movement change characteristics and the predicted target joint velocity.

[0257] In some embodiments, the transition frame generation module 1003 is configured to:

[0258] Inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into a second sub-model of the motion prediction model to obtain a plurality of sub-motion information of the target object, wherein the sub-motion information is motion information predicted based on the motion stage of the target object;

[0259] Based on multiple target weights, the multiple sub-action information is weightedly summed to obtain the action information of the target object based on the transition action, so as to generate the transition frame.

[0260] In some embodiments, the transition frame generation module 1003 is configured to:

[0261] inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model; mapping the motion change characteristics into corresponding motion information based on a mapping relationship between the motion manifold space and the object motion information in the second sub-model, to obtain the predicted joint position and predicted joint velocity of the target object, wherein the motion manifold space indicates the motion change of the object in frames corresponding to two consecutive motions;

[0262] The plurality of sub-action information are obtained based on the action information of the target object in the starting frame, the predicted joint positions and predicted joint velocities of the target object, and the predicted target joint velocities.

[0263] In some embodiments, the apparatus further comprises:

[0264] a first training module, configured to train a second sub-model in the motion prediction model based on a sample data set and label information to obtain the trained second sub-model, wherein the sample data set includes sample frames of a sample object based on a plurality of continuous motions, and the label information indicates a sample target joint velocity of the sample object in the sample frame;

[0265] The second training module is used to train the first sub-model in the action prediction model based on the sample data set and the trained second sub-model to obtain the trained first sub-model.

[0266] In some embodiments, the first training module includes:

[0267] a first training unit, configured to update model parameters of the vector coding model and the second sub-model based on the sample data set, the label information, and the first loss function until a first training condition is satisfied, thereby obtaining an intermediate vector coding model and an intermediate second sub-model, wherein the vector coding model is configured to output a predicted motion change feature of the m+1th sample frame based on the mth sample frame and the m+1th sample frame, where m is a positive integer;

[0268] A second training unit is configured to update model parameters of the intermediate vector coding model and the intermediate second sub-model based on the sample data set, the label information, and a second loss function until a second training condition is satisfied, thereby obtaining the trained vector coding model and the trained second sub-model;

[0269] The first loss function indicates the motion reconstruction loss and information divergence of the sample frame, and the second loss function indicates the motion reconstruction loss, information divergence, footstep sliding loss and bone length loss of the sample frame.

[0270] In some embodiments, the first training unit is configured to:

[0271] Obtaining a motion reconstruction loss value of the (m+1)th sample frame based on the (m)th sample frame, the (m+1)th sample frame, the label information, the vector coding model, and the second sub-model;

[0272] Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame;

[0273] Based on the action reconstruction loss value and the information divergence, the model parameters of the vector coding model and the second sub-model are updated until the first training condition is met, thereby obtaining the intermediate vector coding model and the intermediate second sub-model.

[0274] In some embodiments, the second training unit is configured to:

[0275] Based on the mth sample frame, the (m+1th) sample frame, the label information, the intermediate vector coding model, and the intermediate second sub-model, obtaining a motion reconstruction loss value, a foot sliding loss value, and a bone length loss value of the (m+1th) sample frame;

[0276] Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame;

[0277] Based on the action reconstruction loss value, the foot sliding loss value, the bone length loss value and the information divergence, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

[0278] In some embodiments, the second training module is used to:

[0279] Based on the sample start frame, the sample target frame, the sample target frame number, the first sub-model and the trained second sub-model, obtaining a joint rotation loss value, a joint position loss value and a bone rotation loss value of a sample transition frame between the sample start frame and the sample target frame, wherein the sample target frame number indicates the number of sample transition frames between the sample start frame and the sample target frame;

[0280] Based on the joint rotation loss value, the joint position loss value and the bone rotation loss value, the model parameters of the first sub-model are updated until the training end condition is met, thereby obtaining the trained first sub-model.

[0281] In an embodiment of the present application, a transition frame generation device is provided. When generating a transition frame between a starting frame and a target frame, the device obtains the target object's motion change characteristics and predicts the target joint velocity based on the target object's motion information in the starting frame, the target object's motion information in the target frame, and the target frame number, so as to predict the target object's transition motion between the starting frame and the target frame to generate the transition frame. In this process, since different joints have different importance in motion generation, and the target joint velocity largely determines the position of the target object in the next frame, more accurate motion information can be obtained by further predicting the target object's transition motion by obtaining the predicted target joint velocity, thereby generating a smooth and natural transition frame sequence, effectively improving the quality of the transition frame.

[0282] It should be noted that the above-described embodiments of the transition frame generation device, when generating transition frames, illustrate the division of the aforementioned functional modules only as an example. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, i.e., the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the transition frame generation device and the transition frame generation method embodiments provided in the above-described embodiments share the same concept. The specific implementation process is detailed in the method embodiments and will not be further elaborated here.

[0283] In an exemplary embodiment, a computer device is also provided, which includes a processor and a memory, wherein the memory is used to store at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the transition frame generation method in the embodiment of the present application.

[0284] Taking computer equipment as the terminal as an example, Figure 11 The figure is a schematic diagram of the structure of a terminal provided according to an embodiment of the present application. Terminal 1100 may be a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. Terminal 1100 may also be referred to as user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.

[0285] Typically, the terminal 1100 includes a processor 1101 and a memory 1102 .

[0286] The processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0287] Memory 1102 may include one or more computer-readable storage media, which may be non-transitory. Memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1102 is used to store at least one program code, which is executed by processor 1101 to implement the transition frame generation method provided in the method embodiment of the present application.

[0288] In some embodiments, terminal 1100 may optionally include a peripheral device interface 1103 and at least one peripheral device. Processor 1101, memory 1102, and peripheral device interface 1103 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1103 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.

[0289] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0290] The RF circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 1104 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0291] The display screen 1105 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch screen display, the display screen 1105 also has the ability to collect touch signals on the surface or above the surface of the display screen 1105. The touch signal can be input as a control signal to the processor 1101 for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1105, which is set on the front panel of the terminal 1100; in other embodiments, there can be at least two display screens 1105, which are respectively set on different surfaces of the terminal 1100 or in a folding design; in other embodiments, the display screen 1105 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1100. Even more, the display screen 1105 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0292] The camera assembly 1106 is used to capture images or videos. Optionally, the camera assembly 1106 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0293] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1101 for processing, or input into the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0294] The positioning component 1108 is used to locate the current geographical location of the terminal 1100 to implement navigation or LBS (Location Based Service).

[0295] Power supply 1109 is used to power various components in terminal 1100. Power supply 1109 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1109 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0296] In some embodiments, the terminal 1100 further includes one or more sensors 1110 , including but not limited to: an acceleration sensor 1111 , a gyroscope sensor 1112 , a pressure sensor 1113 , an optical sensor 1114 , and a proximity sensor 1115 .

[0297] The accelerometer 1111 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal 1100. For example, the accelerometer 1111 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 1111. The accelerometer 1111 can also be used to collect game or user motion data.

[0298] The gyroscope sensor 1112 can detect the orientation and rotation angle of the terminal 1100. The gyroscope sensor 1112 can work with the acceleration sensor 1111 to collect the user's 3D movements on the terminal 1100. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can implement the following functions: motion sensing (such as changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0299] The pressure sensor 1113 can be set on the side frame of the terminal 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is set on the side frame of the terminal 1100, it can detect the user's grip signal of the terminal 1100, and the processor 1101 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 1113. When the pressure sensor 1113 is set on the lower layer of the display screen 1105, the processor 1101 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0300] Optical sensor 1114 is used to collect ambient light intensity. In one embodiment, processor 1101 can control the display brightness of display screen 1105 based on the ambient light intensity collected by optical sensor 1114. Specifically, when the ambient light intensity is high, the display brightness of display screen 1105 is increased; when the ambient light intensity is low, the display brightness of display screen 1105 is decreased. In another embodiment, processor 1101 can also dynamically adjust the shooting parameters of camera assembly 1106 based on the ambient light intensity collected by optical sensor 1114.

[0301] Proximity sensor 1115, also known as a distance sensor, is typically located on the front panel of terminal 1100. Proximity sensor 1115 is used to detect the distance between the user and the front of terminal 1100. In one embodiment, when proximity sensor 1115 detects that the distance between the user and the front of terminal 1100 is gradually decreasing, processor 1101 controls display screen 1105 to switch from the screen-on state to the screen-off state. When proximity sensor 1115 detects that the distance between the user and the front of terminal 1100 is gradually increasing, processor 1101 controls display screen 1105 to switch from the screen-off state to the screen-on state.

[0302] Those skilled in the art will understand that Figure 11 The structure shown in the figure does not constitute a limitation on the terminal 1100, and the terminal 1100 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0303] Taking the computer device as a server as an example, Figure 12 1 is a schematic diagram of the structure of a server provided according to an embodiment of the present application. The server 1200 may vary significantly due to different configurations or performance, and may include one or more processors (Central Processing Units, CPUs) 1201 and one or more memories 1202, wherein the memories 1202 store at least one computer program, which is loaded and executed by the processor 1201 to implement the transition frame generation method provided by each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.

[0304] An embodiment of the present application also provides a computer-readable storage medium, which is applied to a computer device. The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the transition frame generation method in the above embodiment.

[0305] The present application also provides a computer program product or computer program, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to implement the transition frame generation method of the above-described embodiment.

[0306] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0307] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for generating a transition frame, characterized in that: The method comprises: Acquire motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number, where the target frame number indicates the number of transition frames between the starting frame and the target frame; acquiring a motion change feature of the target object and a predicted target joint velocity based on motion information of the target object in the starting frame, motion information of the target object in the target frame, and the target frame number, wherein the motion change feature indicates a change in a transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame; Inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into a second sub-model of the motion prediction model to obtain a plurality of sub-motion information of the target object, wherein the sub-motion information is motion information predicted based on the motion stage of the target object; Based on multiple target weights, the multiple sub-action information is weightedly summed to obtain the action information of the target object based on the transition action, so as to generate a transition frame.

2. The method according to claim 1, characterized in that The acquiring the motion change characteristics of the target object and predicting the target joint velocity based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number includes: Inputting the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number into a first submodel of the motion prediction model to obtain an offset embedding vector between the starting frame and the target frame, an embedding vector of the starting frame, and an embedding vector of the target frame; Based on the offset embedding vector, the embedding vector of the starting frame, the embedding vector of the target frame, and the target frame number, the change of the transition movement of the target object relative to the movement in the starting frame and the hip joint speed of the target object are predicted to obtain the movement change feature and the predicted target joint speed.

3. The method according to claim 1, characterized in that The step of inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model of the motion prediction model to obtain a plurality of sub-motion information of the target object includes: inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model; mapping the motion change characteristics into corresponding motion information based on a mapping relationship between the motion manifold space and the object motion information in the second sub-model, to obtain the predicted joint position and predicted joint velocity of the target object, wherein the motion manifold space indicates the motion change of the object in frames corresponding to two consecutive motions; The plurality of sub-action information are obtained based on the action information of the target object in the starting frame, the predicted joint positions and predicted joint velocities of the target object, and the predicted target joint velocities.

4. The method according to claim 1, wherein The method further comprises: training a second sub-model in the motion prediction model based on a sample data set and label information to obtain a trained second sub-model, wherein the sample data set includes sample frames of a sample object based on a plurality of continuous motions, and the label information indicates a sample target joint velocity of the sample object in the sample frame; Based on the sample data set and the trained second sub-model, the first sub-model in the action prediction model is trained to obtain the trained first sub-model.

5. The method according to claim 4, characterized in that The step of training the second sub-model in the action prediction model based on the sample data set and the label information to obtain the trained second sub-model includes: Based on the sample data set, the label information, and the first loss function, updating model parameters of the vector coding model and the second sub-model until a first training condition is met, thereby obtaining an intermediate vector coding model and an intermediate second sub-model, wherein the vector coding model is used to output a predicted motion change feature of the m+1th sample frame based on the mth sample frame and the m+1th sample frame, where m is a positive integer; Based on the sample data set, the label information, and the second loss function, updating model parameters of the intermediate vector coding model and the intermediate second sub-model until a second training condition is satisfied, thereby obtaining the trained vector coding model and the trained second sub-model; The first loss function indicates the motion reconstruction loss and information divergence of the sample frame, and the second loss function indicates the motion reconstruction loss, information divergence, footstep sliding loss and bone length loss of the sample frame.

6. The method according to claim 5, characterized in that The updating of model parameters of the vector coding model and the second sub-model based on the sample data set, the label information, and the first loss function until a first training condition is satisfied to obtain an intermediate vector coding model and an intermediate second sub-model includes: Obtaining a motion reconstruction loss value of the (m+1)th sample frame based on the (m)th sample frame, the (m+1)th sample frame, the label information, the vector coding model, and the second sub-model; Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame; Based on the action reconstruction loss value and the information divergence, the model parameters of the vector coding model and the second sub-model are updated until the first training condition is met, thereby obtaining the intermediate vector coding model and the intermediate second sub-model.

7. The method according to claim 5, characterized in that The updating of model parameters of the intermediate vector coding model and the intermediate second sub-model based on the sample data set, the label information, and the second loss function until a second training condition is satisfied to obtain the trained vector coding model and the trained second sub-model includes: Based on the mth sample frame, the (m+1th) sample frame, the label information, the intermediate vector coding model, and the intermediate second sub-model, obtaining a motion reconstruction loss value, a foot sliding loss value, and a bone length loss value of the (m+1th) sample frame; Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame; Based on the action reconstruction loss value, the footstep sliding loss value, the bone length loss value and the information divergence, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

8. The method according to claim 4, characterized in that The step of training the first sub-model in the action prediction model based on the sample data set and the trained second sub-model to obtain the trained first sub-model includes: Based on a sample start frame, a sample target frame, a sample target frame number, the first sub-model, and the trained second sub-model, obtaining a joint rotation loss value, a joint position loss value, and a bone rotation loss value of a sample transition frame between the sample start frame and the sample target frame, wherein the sample target frame number indicates the number of sample transition frames between the sample start frame and the sample target frame; Based on the joint rotation loss value, the joint position loss value, and the bone rotation loss value, the model parameters of the first sub-model are updated until a training end condition is met, thereby obtaining the trained first sub-model.

9. A transition frame generating device, characterized in that: The device comprises: A first acquisition module is configured to acquire motion information of a target object in a starting frame, motion information of a target object in a target frame, and a target frame number, where the target frame number indicates the number of transition frames between the starting frame and the target frame; a second acquisition module, configured to acquire a motion change feature of the target object and predict a target joint velocity based on the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number, wherein the motion change feature indicates a change in a transition motion of the target object between the starting frame and the target frame relative to the motion in the starting frame; a transition frame generation module, configured to input the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into a second sub-model of the motion prediction model to obtain a plurality of sub-motion information of the target object, wherein the sub-motion information is motion information predicted based on the motion stage of the target object; Based on multiple target weights, the multiple sub-action information is weightedly summed to obtain the action information of the target object based on the transition action, so as to generate a transition frame.

10. The device according to claim 9, characterized in that The second acquisition module is used to: Inputting the motion information of the target object in the starting frame, the motion information of the target object in the target frame, and the target frame number into a first submodel of the motion prediction model to obtain an offset embedding vector between the starting frame and the target frame, an embedding vector of the starting frame, and an embedding vector of the target frame; Based on the offset embedding vector, the embedding vector of the starting frame, the embedding vector of the target frame, and the target frame number, the change of the transition movement of the target object relative to the movement in the starting frame and the hip joint speed of the target object are predicted to obtain the movement change feature and the predicted target joint speed.

11. The device according to claim 9, characterized in that The transition frame generation module is used to: inputting the motion information of the target object in the starting frame, the motion change characteristics of the target object, and the predicted target joint velocity into the second sub-model; mapping the motion change characteristics into corresponding motion information based on a mapping relationship between the motion manifold space and the object motion information in the second sub-model, to obtain the predicted joint position and predicted joint velocity of the target object, wherein the motion manifold space indicates the motion change of the object in frames corresponding to two consecutive motions; The plurality of sub-action information are obtained based on the action information of the target object in the starting frame, the predicted joint positions and predicted joint velocities of the target object, and the predicted target joint velocities.

12. The device according to claim 9, characterized in that The device further comprises: a first training module, configured to train a second sub-model in the motion prediction model based on a sample data set and label information to obtain a trained second sub-model, wherein the sample data set includes sample frames of a sample object based on a plurality of continuous motions, and the label information indicates a sample target joint velocity of the sample object in the sample frame; The second training module is used to train the first sub-model in the action prediction model based on the sample data set and the trained second sub-model to obtain the trained first sub-model.

13. The device according to claim 12, characterized in that The first training module includes: a first training unit, configured to update model parameters of the vector coding model and the second sub-model based on the sample data set, the label information, and the first loss function until a first training condition is satisfied, thereby obtaining an intermediate vector coding model and an intermediate second sub-model, wherein the vector coding model is configured to output a predicted motion change feature of the m+1th sample frame based on the mth sample frame and the m+1th sample frame, where m is a positive integer; a second training unit, configured to update model parameters of the intermediate vector coding model and the intermediate second sub-model based on the sample data set, the label information, and a second loss function until a second training condition is satisfied, thereby obtaining the trained vector coding model and the trained second sub-model; The first loss function indicates the motion reconstruction loss and information divergence of the sample frame, and the second loss function indicates the motion reconstruction loss, information divergence, footstep sliding loss and bone length loss of the sample frame.

14. The device according to claim 13, characterized in that The first training unit is used to: Obtaining a motion reconstruction loss value of the (m+1)th sample frame based on the (m)th sample frame, the (m+1)th sample frame, the label information, the vector coding model, and the second sub-model; Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame; Based on the action reconstruction loss value and the information divergence, the model parameters of the vector coding model and the second sub-model are updated until the first training condition is met, thereby obtaining the intermediate vector coding model and the intermediate second sub-model.

15. The device according to claim 13, characterized in that The second training unit is used to: Based on the mth sample frame, the (m+1th) sample frame, the label information, the intermediate vector coding model, and the intermediate second sub-model, obtaining a motion reconstruction loss value, a foot sliding loss value, and a bone length loss value of the (m+1th) sample frame; Based on the m-th sample frame, the (m+1)-th sample frame and the vector coding model, obtaining the information divergence of the (m+1)-th sample frame; Based on the action reconstruction loss value, the footstep sliding loss value, the bone length loss value and the information divergence, the model parameters of the intermediate vector coding model and the intermediate second sub-model are updated until the second training condition is met, thereby obtaining the trained vector coding model and the trained second sub-model.

16. The device according to claim 12, characterized in that The second training module is used to: Based on a sample start frame, a sample target frame, a sample target frame number, the first sub-model, and the trained second sub-model, obtaining a joint rotation loss value, a joint position loss value, and a bone rotation loss value of a sample transition frame between the sample start frame and the sample target frame, wherein the sample target frame number indicates the number of sample transition frames between the sample start frame and the sample target frame; Based on the joint rotation loss value, the joint position loss value, and the bone rotation loss value, the model parameters of the first sub-model are updated until a training end condition is met, thereby obtaining the trained first sub-model.

17. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory is used to store at least one computer program, and the at least one computer program is loaded by the processor and executes the transition frame generation method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by the processor to implement the transition frame generation method according to any one of claims 1 to 8.

19. A computer program product, characterized in that The computer program product includes at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the transition frame generation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Training method and device of action completion model, completion method, equipment and medium

    CN113345061A