A transition action generation method, device, apparatus and storage medium
By acquiring joint information and using a pre-trained model to generate transition joint trajectories, the problem of insufficient realism in transition motion synthesis is solved, and more realistic and natural transition motion generation is achieved.
Patent Information
- Application Number
- CN202211430976.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2026-07-10
- Estimated Expiration
- 2042-11-15
AI Technical Summary
Existing deep learning-based methods produce poor realism in transition action synthesis.
By acquiring the key point information from the initial input data, and utilizing the pre-trained transition frame prediction model, transition trajectory generation model, and transition action synthesis model, transition key point trajectories are generated to guide the transition action synthesis process and improve the realism of transition actions.
It improves the realism and naturalness of transition actions, and enhances the controllability and efficiency of transition action synthesis.
Smart Images

Figure CN115761081B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animation production technology, and more specifically, to a method, apparatus, device, and storage medium for generating transitional actions. Background Technology
[0002] Transition animation synthesis technology can automatically synthesize a reasonable intermediate transition animation from two action segments of a given virtual character, and it has wide applications in game animation and film and television animation production.
[0003] Currently, transitional actions can be synthesized from given source and target actions using traditional art production methods, linear interpolation-based methods, and deep learning-based methods. While existing deep learning-based methods can produce transitional actions, they simply synthesize transitional actions based on the source and target actions, resulting in transitional actions with poor realism.
[0004] Therefore, how to improve the realism of transition actions is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The purpose of this application is to address the shortcomings of the prior art by providing a method, apparatus, device, and storage medium for synthesizing transition actions, which can improve the realism of transition actions.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:
[0007] In a first aspect, embodiments of this application provide a method for generating transition actions, the method comprising:
[0008] Acquire initial input data, which includes: global position information and local rotation information of each joint in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint in at least one frame of animation corresponding to the target action.
[0009] The number of transition frames is obtained based on the initial input data and the pre-trained transition frame prediction model;
[0010] The transition joint trajectory is obtained based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model;
[0011] Based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model, a transition action sequence between the source action and the target action is obtained.
[0012] Secondly, embodiments of this application also provide a transition action generation device, the device comprising:
[0013] The acquisition module is used to acquire initial input data, which includes: global position information and local rotation information of each joint point in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint point in at least one frame of animation corresponding to the target action.
[0014] The first determining module is used to obtain the number of transition frames based on the initial input data and the pre-trained transition frame prediction model;
[0015] The second determining module is used to obtain the transition joint trajectory based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model;
[0016] The third determining module is used to obtain the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model.
[0017] Thirdly, embodiments of this application provide an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the transition action generation method described in the first aspect.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the transition action generation method described in the first aspect.
[0019] The beneficial effects of this application are:
[0020] This application provides a method, apparatus, device, and storage medium for generating transition actions. The method includes: acquiring initial input data, which includes global position information and local rotation information of each joint in at least one frame of animation corresponding to a source action, and global position information and local rotation information of each joint in at least one frame of animation corresponding to a target action; obtaining a number of transition frames based on the initial input data and a pre-trained transition frame prediction model; obtaining a transition joint trajectory based on the number of transition frames, the initial input data, and a pre-trained transition trajectory generation model; and obtaining a sequence of transition actions between the source action and the target action based on the number of transition frames, the initial input data, the transition joint trajectory, and a pre-trained transition action synthesis model.
[0021] The transition action generation method provided in this application allows for the prediction of transition frame numbers using a transition frame number prediction model after initial input data is obtained. Then, based on the transition frame numbers and the initial input data, the input to the transition trajectory generation model is obtained. The transition trajectory generation model, based on the transition frame numbers and the initial input data, generates transition joint trajectory data. Finally, during the transition action synthesis process, the transition action synthesis model uses the transition joint trajectory generated by the transition trajectory generation model to constrain the transition action synthesis process. In other words, this application not only synthesizes transition actions based on the transition frame numbers and initial input data but also introduces the transition joint trajectory as a parameter. The transition joint trajectory guides the transition action synthesis model to synthesize transition actions based on the source and target actions included in the initial input data, thus improving the realism of the transition actions. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A schematic diagram of the structure of a target deep learning model provided in an embodiment of this application;
[0024] Figure 2 A flowchart illustrating a transition action generation method provided in an embodiment of this application;
[0025] Figure 3 A flowchart illustrating another transition action generation method provided in this application embodiment;
[0026] Figure 4 This is a schematic diagram of the structure of a transition trajectory generation model provided in an embodiment of this application;
[0027] Figure 5 A flowchart illustrating another method for generating transition actions provided in this application embodiment;
[0028] Figure 6 A flowchart illustrating another method for generating transition actions provided in an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of the structure of a transition frame number prediction model provided in an embodiment of this application;
[0030] Figure 8A flowchart illustrating another method for generating transition actions provided in this application embodiment;
[0031] Figure 9 This is a schematic diagram of the structure of a transition action generation device provided in an embodiment of this application;
[0032] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0034] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0035] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0036] Before providing a detailed explanation of the embodiments of this application, the application scenario of this application will first be introduced. Specifically, this application scenario can be based on deep learning-based transition action synthesis. For example, in the field of game animation production, a transition action can be automatically synthesized based on two preceding and following action segments (source action and target action) of a given virtual character, according to the method described in the following example of this application. It should be noted that the application fields of this application may include game animation production, film and television animation production, etc., and this application does not limit it to these fields.
[0037] Understandably, one exemplary approach involves pre-training an initial deep learning network based on constructed training samples. Upon meeting a training stopping condition, a target deep learning network is obtained, comprising a transition frame count prediction model, a transition trajectory generation model, and a transition action synthesis model. Another exemplary approach involves training the initial transition frame count prediction model, the initial transition trajectory generation model, and the initial transition action synthesis model separately to obtain a transition frame count prediction model, a transition trajectory generation model, and a transition action synthesis model that meet the training stopping condition. These three models are then combined to form the target deep learning network. It should be noted that this application does not limit the scope of the proposed approach.
[0038] This section introduces the transition frame prediction model, transition trajectory generation model, and transition action synthesis model in target deep learning networks, focusing on the application of target deep learning networks.
[0039] Figure 1 This is a schematic diagram of the structure of a target deep learning model provided in an embodiment of this application, such as... Figure 1 As shown, the initial input data can first be input into the transition frame prediction model 101 to predict the number of transition frames; then, the transition frame number and the initial input data are combined to form transition input data and input into the transition trajectory generation model 102 to generate the transition joint trajectory; finally, the transition action synthesis model 103 uses the transition joint trajectory output by the transition trajectory generation model 102 to constrain the transition action synthesis process during the transition action synthesis process based on the transition input data, so that the final transition action sequence has strong realism and naturalness.
[0040] It should be noted that the specific details of the model application stage can be found in the example description below. Furthermore, the model training stage is similar to the model application stage, so the specific details of the model training stage can be found in the relevant section of the model application stage description, and will not be explained here.
[0041] The following is an example illustration of the transition action generation method mentioned in this application, with reference to the accompanying drawings. Figure 2 This is a flowchart illustrating a transition action generation method provided in an embodiment of this application. Figure 2 The method may include:
[0042] S201. Obtain initial input data.
[0043] The initial input data includes: global position information and local rotation information of each joint in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint in at least one frame of animation corresponding to the target action.
[0044] It should be noted that the number of frames corresponding to the source action and the number of frames corresponding to the target action may be the same or different, and this application does not limit them. For example, the source action contains 10 frames of animation and the target action contains 8 frames of animation.
[0045] For example, suppose we need to synthesize transitional motions between two action segments (i.e., the source action and the target action) of a given target virtual character. An actor can then perform actions according to pre-designed movements, such as dancing, running, or walking. During the actor's performance, the actor's skeletal animation can be captured using pre-set motion capture equipment. After acquiring the actor's skeletal animation, it can be redirected to the target character's skeletal animation. For instance, the redirection function built into 3D creation software can be used to redirect the actor's skeletal animation to the target character's skeletal animation, where the target character's skeleton is the aforementioned target virtual character skeleton.
[0046] After obtaining the target character's skeletal animation, artists can repair it to obtain a repaired version. This avoids issues such as clipping, distortion, and flickering. After obtaining the repaired skeletal animation, the original input data can be extracted. This original input data includes the global position information of the root joints in each frame of the target character's skeletal animation, as well as the local rotation information of each joint relative to its parent joint. After obtaining the original input data, forward dynamics processing can be performed to obtain the global position information of each joint. Simultaneously, the local rotation information of each joint relative to its parent joint can be converted into 6D local rotation information, which is the content included in the initial input data mentioned above.
[0047] S202. Based on the initial input data and the pre-trained transition frame prediction model, obtain the number of transition frames.
[0048] The initial input data, including the global position information and local rotation information of the joints corresponding to the source action and the target action, is input into the transition frame prediction model to predict the number of transition frames from the source action to the target action. For example, the number of transition frames is 20, 30, or 40.
[0049] S203. Based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model, the transition joint trajectory is obtained.
[0050] For example, after predicting the number of transition frames using the transition frame number prediction model, the transition frame number and the initial input data can be combined to form transition input data, which can then be used as input to the transition trajectory generation model. By encoding and decoding the transition input data using the transition trajectory generation model, the trajectory of the transition joint can be obtained.
[0051] It is worth noting that the transition joint trajectory mentioned here can be understood as the transition joint trajectory corresponding to at least one joint included in the initial input data. In other words, the transition trajectory generation model can output at least one transition joint trajectory. Furthermore, it is understandable that the number of transition frames affects the length of the transition joint trajectory.
[0052] S204. Based on the number of transition frames, initial input data, transition joint trajectories, and the pre-trained transition action synthesis model, obtain the transition action sequence between the source action and the target action.
[0053] For example, after obtaining the transition joint trajectory, the input data for the transition motion synthesis model can be obtained based on the aforementioned transition input data (transition frame number and initial input data) and the transition joint trajectory. The transition motion synthesis model encodes and decodes this input data, and synthesizes transition motions for the source motion and target motion in the initial input data, including the transition joint trajectory, based on the constraints of the transition joint trajectory. This results in a transition motion sequence between the source motion and the target motion. The transition motion sequence includes the global position information and local rotation information of each joint in each transition frame corresponding to the transition motion. Finally, the transition motion animation can be obtained by combining the aforementioned target character skeleton and the transition motion sequence between the source motion and the target motion.
[0054] In summary, the transition action generation method provided in this application can predict the number of transition frames through a transition frame prediction model after obtaining the initial input data. Then, based on the number of transition frames and the initial input data, the input to the transition trajectory generation model is obtained. The transition trajectory generation model can obtain the transition joint trajectory based on the number of transition frames and the initial input data. Finally, during the transition action synthesis process based on the number of transition frames and the initial input data, the transition joint trajectory generated by the transition trajectory generation model constrains the transition action synthesis process. In other words, this application not only synthesizes transition actions based on the number of transition frames and the initial input data, but also introduces the parameter of transition joint trajectory. The transition joint trajectory can guide the transition action synthesis model to synthesize transition actions based on the source action and target action included in the initial input data, thus improving the realism of the transition actions.
[0055] Figure 3This is a flowchart illustrating another transition action generation method provided in an embodiment of this application. Optionally, as... Figure 3 As shown, the transition joint trajectory is obtained based on the number of transition frames, initial input data, and a pre-trained transition trajectory generation model, including:
[0056] S301. Based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model, the first embedded feature is obtained.
[0057] As described above, the initial input data includes the global position information of each joint in each frame of the animation. For example, the input data of the input trajectory encoding module in the encoder of the transition trajectory generation model can be determined based on the global position information of joints (such as end joints) that meet preset requirements. The input trajectory encoding module in the encoder of the transition trajectory generation model encodes the input data to obtain the first embedded feature.
[0058] S302. Construct the initial transition trajectory based on the number of transition frames.
[0059] S303. Based on the first embedded feature, the initial transition trajectory, and the transition trajectory encoding module in the encoder of the transition trajectory generation model, the transition joint trajectory is obtained.
[0060] For example, the initial transition trajectory constructed based on the number of transition frames is a tensor of all zeros, and the time dimension is the number of transition frames. For instance, assuming the number of transition frames is 6, then the tensor corresponding to the initial transition trajectory is [000000].
[0061] After obtaining the first embedded feature and the initial transition trajectory, the initial transition trajectory can be first input into the input layer of the transition trajectory generation model. The input layer sequentially performs position concatenation encoding and linear projection on the initial transition trajectory to obtain the projected features. The projected features and the first embedded feature are simultaneously input into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition joint trajectory.
[0062] Optionally, obtaining the first embedded feature based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model includes: obtaining the input trajectory based on the initial input data; inputting the input trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the second projected feature; and inputting the second projected feature into the input trajectory encoding module in the encoder of the transition trajectory generation model to obtain the first embedded feature.
[0063] For example, the initial input data is filtered to identify key points that meet preset requirements. The input trajectory is then obtained based on the global position information of these key points. The transition trajectory generation model includes an input layer, an encoder, and a linear projection layer, which are connected sequentially. The input layer of the transition trajectory generation model includes a position concatenation coding layer and a linear projection layer. The position concatenation coding layer is connected to the linear projection layer. Position concatenation coding can be understood as concatenating the vector corresponding to the received input trajectory with the position coding vector to obtain a concatenated vector. After processing by the linear projection layer, the concatenated vector yields the second projected feature.
[0064] For example, the transition trajectory generation model may include multiple cascaded encoders, such as six. Each encoder includes an input trajectory encoding module and a transition trajectory encoding module. The input trajectory encoding module and the transition trajectory encoding module have the same structure, and the structure of the input trajectory encoding module in each encoder is similar to that of the transition trajectory encoding module. Here, we take the input trajectory encoding module of one encoder in the transition trajectory generation model as an example. This input trajectory encoding module includes a multi-head attention layer, a normalization layer, an activation function layer, and a residual perception layer, which are connected sequentially. The residual perception layer may include multiple fully connected layers (such as three), each followed by an activation layer. The output of each activation function layer in the residual perception layer, plus the input of the connected fully connected layer, is the final output. Based on this, after encoding the aforementioned second projected feature through the multi-head attention layer, normalization layer, activation function layer, and residual perception layer in this input trajectory encoding module, the first embedded feature is obtained. It should be noted that if an encoder is followed by another encoder, the first embedded feature output by the input trajectory encoding module of that encoder is input into the input trajectory encoding module of the next encoder for encoding. Simultaneously, the first embedded feature output by the input trajectory encoding module of that encoder is also input into the transition trajectory generation module of the current encoder for encoding. Similarly, the transition trajectory embedding feature output by the transition trajectory encoding module is input into the transition trajectory encoding module of the next encoder for encoding, and so on. If no encoder is connected after that encoder, the first embedded feature output by the input trajectory encoding module of that encoder is only input into the transition trajectory generation module of the current encoder for encoding. After linear projection of the transition trajectory embedding feature output by the transition trajectory encoding module of that encoder, the transition joint trajectory can be obtained. The encoding process of one encoder in the transition trajectory generation model can be found in [reference needed]. Figure 4 , Figure 4 This is a schematic diagram of a transition trajectory generation model provided in an embodiment of this application. It should be noted that... Figure 4 This is merely an example and is not intended to limit the structure of transition trajectory generation models.
[0065] Optionally, obtaining the input trajectory based on the initial input data includes: determining multiple target global position information that satisfy the selection information and belong to key joints from the initial input data based on the user's selection information; and obtaining the input trajectory based on the global position information of each target.
[0066] The user-input selection information may include at least one key joint, such as the identifier of the end joint. This identifier can be understood as the number of the end joint. It should be noted that this application does not limit the number of end joints. As described above, the initial input data includes the global position information of each joint in each frame of the animation. Therefore, end joints that meet the selection information can be selected from the joints in each frame of the animation, thereby obtaining the global position information of the end joints. The global position information of the end joints can be referred to as the target global position information.
[0067] For example, suppose the user input includes identifiers for 12 end joints, such as 2 for each wrist, elbow, shoulder, thigh, knee, and ankle. It should be understood that the identifiers for joints at the same position are the same across different animation frames. Therefore, these 12 end joints can be selected from the joints in each frame corresponding to the initial input data. Then, based on the global position information of each joint in each frame, the target global position information corresponding to these 12 end joints can be obtained. In other words, each end joint can correspond to multiple target global position information. Therefore, based on the target global position information corresponding to each end joint, the input trajectory, i.e., the vector corresponding to the input trajectory, can be constructed.
[0068] As can be seen, users can control the input trajectory according to actual needs, and thus control the final transition joint trajectory. In other words, users can indirectly control and adjust the synthesis of transition actions, improving controllability.
[0069] Optionally, obtaining the input trajectory based on the global position information of each target includes: obtaining a coordinate system with the root joint as the origin corresponding to the global position information of each target; transforming the global position information of each target based on the coordinate system with the root joint as the origin to obtain multiple transformed position information; and obtaining the input trajectory based on the transformed position information.
[0070] Continuing with the example above, we can obtain the target global position information of each end joint in each frame of the animation. Taking one frame as an example, this frame corresponds to a root joint. Using this root joint as the origin coordinate system, we can transform the target global information of each end joint in this frame based on the global position information of the root joint, thus obtaining the transformed position information of each end joint in this frame. Referring to the above description, we can finally obtain the transformed position information of each end joint in each frame of the animation, and thus obtain the input trajectory.
[0071] Figure 5 This is a flowchart illustrating another transition action generation method provided in an embodiment of this application. Optionally, as... Figure 5 As shown, the transition trajectory encoding module in the encoder of the transition trajectory generation model, based on the first embedded feature, the initial transition trajectory, and the transition trajectory generation model, obtains the transition joint point trajectory, including:
[0072] S501. Input the initial transition trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the third projected features.
[0073] S502. Input the first embedded feature and the third projected feature into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition trajectory embedded feature, and perform linear projection on the transition trajectory embedded feature to obtain the transition joint trajectory.
[0074] As described above, the transition trajectory generation model includes an input layer, an encoder, and a linear projection layer. For example, the initial transition trajectory and the input trajectory can share a single input layer. Based on this, the position concatenation encoding layer in the input layer of the transition trajectory generation model concatenates the vector corresponding to the initial transition trajectory with the position encoding vector to obtain a concatenated vector. This concatenated vector is then processed by the linear projection layer in the input layer to obtain the third projected feature. The process by which the transition trajectory encoding module in the encoder of the transition trajectory generation model encodes the first embedded feature and the third projected feature can be referenced to the process by which the input trajectory encoding module in the encoder of the transition trajectory generation model encodes the second projected feature. A further explanation is not provided here; please refer to [link to relevant documentation]. Figure 4 After obtaining the transition trajectory embedding features, the transition trajectory encoding module in the last encoder of the transition trajectory generation model can input these features into the linear projection layer included in the transition trajectory generation model for linearization to obtain the transition joint trajectory.
[0075] Figure 6 This is a flowchart illustrating another transition action generation method provided in an embodiment of this application. Optionally, as... Figure 6As shown, the transition frame number prediction model obtained above, based on the initial input data and the pre-trained transition frame number prediction model, yields the following:
[0076] S601. Input the initial input data into the transition frame prediction model for position concatenation encoding and linear projection to obtain the first projected features.
[0077] For example, the transition frame number prediction model includes an input layer, an encoder, and a decoder, which are connected sequentially. The input layer includes a position concatenation coding layer and a linear projection layer. The position concatenation coding layer can be understood as concatenating the received initial input data with the position coding vector to obtain a concatenated vector. After the concatenated vector is processed by the linear projection layer, the first projected feature is obtained.
[0078] S602. Input the first projected features into the encoder of the transition frame number prediction model, perform multi-head attention processing, normalization, activation, and residual perception processing to obtain the encoded features.
[0079] For example, the transition frame prediction model may include multiple cascaded encoders, such as six. Each encoder encodes the first projected feature in a similar process; here, we take one encoder as an example. This encoder includes a multi-head attention layer, a normalization layer, an activation function layer, and a residual perception layer, which are connected sequentially. The residual perception layer may include multiple fully connected layers (such as three), each followed by an activation layer. The output of each activation function layer in the residual perception layer, plus the input of its connected fully connected layer, becomes the final output. Based on this, after the multi-head attention layer, normalization layer, activation function layer, and residual perception layer in this encoder encode the first projected feature, intermediate features are obtained. These intermediate features are then input into the encoders connected in series for encoding. After iterative encoding by each encoder, the encoded features are obtained.
[0080] S603. Input the encoded features into the decoder of the transition frame number prediction model and perform decoding multi-layer perception processing to obtain the decoded features.
[0081] S604. Perform linear projection on the decoded features to obtain the number of transition frames.
[0082] For example, the decoder of the transition frame prediction model may include two fully connected layers, each followed by an activation function layer, and finally a linear projection layer. The fully connected layers and activation function layer in the decoder decode the encoded features to obtain decoded features, which are then input into the linear projection layer. Assuming this linear projection layer is a 40-class linear projection layer, the 40-class linear projection layer performs linear projection on the decoded features to obtain the probabilities of belonging to frames 1-40. Finally, the frame number with the highest probability can be determined as the final transition frame number using the argmax function.
[0083] The process of predicting the transition frame number using the aforementioned transition frame number prediction model can be found in [reference needed]. Figure 7 , Figure 7 This is a schematic diagram illustrating the structure of a transition frame number prediction model provided in an embodiment of this application. It should be noted that... Figure 7 This is merely an example and is not intended to limit the structure of transition frame prediction models.
[0084] As can be seen, this application can automatically predict the number of transition frames by using initial input data and a transition frame prediction model, which can improve the efficiency of transition action synthesis and also increase the versatility of the model network (target deep learning network).
[0085] Figure 8 This is a flowchart illustrating another transition action generation method provided in an embodiment of this application. Optionally, as... Figure 8 As shown, based on the number of transition frames, initial input data, transition joint trajectories, and a pre-trained transition motion synthesis model, the transition motion sequence between the source motion and the target motion is obtained, including:
[0086] S801. Based on the number of transition frames, the initial input data, and the first encoding module in the encoder of the transition action synthesis model, the second embedded feature is obtained.
[0087] Optionally, the number of transition frames and the initial input data are input into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fourth projected feature; the fourth projected feature is input into the first encoding module in the encoder of the transition action synthesis model to obtain the second embedded feature.
[0088] For example, the transition action synthesis model includes an input layer, an encoder, and a decoder, which are connected sequentially. The transition action synthesis model may include multiple cascaded encoders, such as six. Each encoder includes a first encoding module and a second encoding module. The first encoding modules have the same structure, and the structures of the first and second encoding modules in each encoder are similar. The structures of the input layer, the first encoding module, and the second encoding module in the encoder of the transition action synthesis model can be referenced from the structures of the input layer and the input trajectory encoding module and transition trajectory encoding module in the encoder of the aforementioned transition frame number prediction model. The position concatenation encoding layer in the input layer of the transition action synthesis model concatenates the vector corresponding to the transition frame number and the initial input data with the position encoding vector to obtain a concatenated vector. After processing by the linear projection layer in the input layer, the concatenated vector yields the fourth projected feature.
[0089] This section uses the first encoding module of an encoder in a transition action synthesis model as an example. This first encoding module includes a multi-head attention layer, a normalization layer, an activation function layer, and a residual perception layer, which are connected sequentially. The residual perception layer may include multiple fully connected layers (e.g., three), each followed by an activation layer. The output of each activation function layer in the residual perception layer, plus the input of its connected fully connected layer, becomes the final output. Based on this, the fourth projected feature mentioned above is encoded using the multi-head attention layer, normalization layer, activation function layer, and residual perception layer in this first encoding layer to obtain the second embedded feature.
[0090] It should be noted that if an encoder is followed by another encoder, the second embedding feature output by the first encoding module of that encoder is input into the first encoding module of the next encoder for encoding. Simultaneously, the second embedding feature output by the first encoding module of that encoder is also input into the second encoding module of the current encoder for encoding. Similarly, the residual embedding feature output by the second encoding module (described below) is input into the second encoding module of the next encoder for encoding, and so on. If no encoder is followed by that encoder, the second embedding feature output by the first encoding module of that encoder is only input into the second encoding module of the current encoder for encoding, and the residual embedding feature output by the second encoding module of that encoder is directly input into the decoder of the transition motion synthesis model for decoding.
[0091] S802. Based on the number of transition frames and the trajectory of the transition joints, the initial residual motion is obtained.
[0092] S803. Based on the number of transition frames, initial input data, second embedded features, initial residual actions, and the second decoding module and decoder in the encoder of the transition action synthesis model, the transition action sequence is obtained.
[0093] For example, a tensor consisting entirely of zeros is constructed based on the number of transition frames, with the time dimension being the number of transition frames. For instance, assuming the number of transition frames is 6, the tensor would be [000000]. This tensor is then concatenated (i.e., spliced) with the vectors corresponding to the transition joint trajectories to obtain the initial residual action. This initial residual action can be understood as an initial residual action guided by the trajectory.
[0094] For example, the initial residual action and the initial input data can share a single input layer. Based on this, the position concatenation encoding layer in the input layer of the transition action synthesis model concatenates the vector corresponding to the initial residual action with the position encoding vector to obtain a concatenated vector. After the concatenated vector is processed by the linear projection layer in the input layer, the projected features are obtained. Based on the results of encoding and decoding the projected features and the second embedded features in the encoder and decoder of the transition action synthesis model, the number of transition frames, and the initial input data, the transition action sequence is obtained.
[0095] Optionally, the process of obtaining the transition action sequence based on the number of transition frames, initial input data, second embedded features, initial residual actions, and the second encoding module and decoder in the encoder of the transition action synthesis model includes: inputting the initial residual actions into the transition action synthesis model for position concatenation encoding and linear projection to obtain a fifth projected feature; inputting the second embedded features and the fifth projected feature into the encoder of the transition action synthesis model to obtain residual embedded features; inputting the residual embedded features into the second encoding module in the decoder of the transition action synthesis model to obtain residual action information; inputting the number of transition frames and initial input data into the transition action synthesis model for linear interpolation to obtain initial linear action information; and obtaining the transition action sequence based on the initial linear action information and the residual action information.
[0096] The transition motion synthesis model includes a linear interpolator. The number of transition frames and initial input data are input into this linear interpolator. Based on the number of transition frames, the linear interpolator generates initial linear motion information between the last frame of the source motion and the first frame of the target motion included in the initial input data. It is understood that this initial linear motion has linear characteristics.
[0097] In the transition action synthesis model, the input layer performs position concatenation encoding and linear projection on the initial residual action containing the transition joint trajectory to obtain the fifth projected feature. The second embedded feature and the fifth projected feature are then input into the second encoding module of the encoder in the transition action synthesis model. These modules undergo multi-head attention processing, normalization, activation, and residual perception to obtain the residual embedded feature. This residual embedded feature is then input into the decoder of the transition action synthesis model, which includes a position decoder and a rotation decoder. The residual embedded feature is decoded by both the position decoder and the rotation decoder to obtain the residual action information. Finally, the initial linear action output by the linear interpolator in the transition action synthesis model is added to the residual action information output by the decoder to obtain the transition action sequence.
[0098] It can be seen that, since the residual motion information contains the constraint of the transition joint trajectory, the linear constraint of the linear difference (i.e., the initial linear motion information) can be weakened by the constraint of the transition joint trajectory, so that the obtained transition motion sequence can have a sense of realism and naturalness.
[0099] Figure 9 This is a schematic diagram of a transition action generation device provided in an embodiment of this application. Figure 9 As shown, the device includes:
[0100] The acquisition module 901 is used to acquire initial input data, which includes: global position information and local rotation information of each joint point in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint point in at least one frame of animation corresponding to the target action.
[0101] The first determining module 902 is used to obtain the number of transition frames based on the initial input data and the pre-trained transition frame prediction model;
[0102] The second determining module 903 is used to generate a model based on the number of transition frames, initial input data and pre-trained transition trajectory to obtain the trajectory of the transition joint point;
[0103] The third determining module 904 is used to obtain the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model.
[0104] Optionally, the second determining module 903 is specifically used to obtain a first embedded feature based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model; construct an initial transition trajectory based on the number of transition frames; and obtain the transition joint trajectory based on the first embedded feature, the initial transition trajectory, and the transition trajectory encoding module in the encoder of the transition trajectory generation model.
[0105] Optionally, the second determining module 903 is further specifically used to obtain the input trajectory based on the initial input data; input the input trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the second projected features; and input the second projected features into the input trajectory encoding module in the encoder of the transition trajectory generation model to obtain the first embedded features.
[0106] Optionally, the second determining module 903 is further specifically used to determine, based on the selection information input by the user, multiple target global position information that satisfy the selection information and belong to key joint points from the initial input data; and to obtain the input trajectory based on the global position information of each target.
[0107] Optionally, the second determining module 903 is further specifically used to obtain a coordinate system with the root joint as the origin corresponding to the global position information of each target; to transform the global position information of each target according to the coordinate system with the root joint as the origin, to obtain multiple transformed position information; and to obtain the input trajectory according to the transformed position information.
[0108] Optionally, the second determining module 903 is further specifically used to input the initial transition trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the third projected feature; input the first embedded feature and the third projected feature into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition trajectory embedded feature, and perform linear projection on the transition trajectory embedded feature to obtain the transition joint trajectory.
[0109] Optionally, the first determining module 902 is specifically used to input the initial input data into the transition frame number prediction model for position concatenation encoding and linear projection to obtain the first projected features; input the first projected features into the encoder of the transition frame number prediction model for multi-head attention processing, normalization, activation, and residual perception processing to obtain the encoded features; input the encoded features into the decoder of the transition frame number prediction model for decoding multi-layer perception processing to obtain the decoded features; and perform linear projection on the decoded features to obtain the transition frame number.
[0110] Optionally, the third determining module 904 is specifically used to obtain the second embedded feature based on the number of transition frames, the initial input data, and the first encoding module in the encoder of the transition action synthesis model; to obtain the initial residual action based on the number of transition frames and the trajectory of the transition joints; and to obtain the transition action sequence based on the number of transition frames, the initial input data, the second embedded feature, the initial residual action, and the second encoding module and decoder in the encoder of the transition action synthesis model.
[0111] Optionally, the third determining module 904 is further specifically used to input the number of transition frames and the initial input data into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fourth projected feature; and to input the fourth projected feature into the first encoding module in the encoder of the transition action synthesis model to obtain the second embedded feature.
[0112] Optionally, the third determining module 904 is further specifically used to input the initial residual action into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fifth projected feature; input the second embedded feature and the fifth projected feature into the second encoding module in the encoder of the transition action synthesis model to obtain the residual embedded feature; input the residual embedded feature into the decoder of the transition action synthesis model to obtain the residual action information; input the transition frame number and the initial input data into the transition action synthesis model for linear interpolation processing to obtain the initial linear action information; and obtain the transition action sequence based on the initial linear action information and the residual action information.
[0113] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0114] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SoC).
[0115] Figure 10This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 10 As shown, the electronic device may include: a processor 1001, a storage medium 1002, and a bus 1003. The storage medium 1002 stores machine-readable instructions executable by the processor 1001. When the electronic device is running, the processor 1001 communicates with the storage medium 1002 via the bus 1003. The processor 1001 executes the machine-readable instructions to perform the following steps:
[0116] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: acquire initial input data, the initial input data including: global position information and local rotation information of each joint point in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint point in at least one frame of animation corresponding to the target action; obtain the number of transition frames based on the initial input data and a pre-trained transition frame prediction model; obtain the trajectory of the transition joint points based on the number of transition frames, the initial input data, and a pre-trained transition trajectory generation model; and obtain the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joint points, and a pre-trained transition action synthesis model.
[0117] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: obtain a first embedded feature based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model; construct an initial transition trajectory based on the number of transition frames; and obtain the transition joint trajectory based on the first embedded feature, the initial transition trajectory, and the transition trajectory encoding module in the encoder of the transition trajectory generation model.
[0118] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: obtain an input trajectory based on initial input data; input the input trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain a second projected feature; and input the second projected feature into the input trajectory encoding module in the encoder of the transition trajectory generation model to obtain a first embedded feature.
[0119] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: determine multiple target global position information that satisfy the selection information and belong to key joint points from the initial input data based on the selection information input by the user; and obtain the input trajectory based on the global position information of each target.
[0120] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: obtain a coordinate system with the root joint point as the origin corresponding to the global position information of each target; transform the global position information of each target according to the coordinate system with the root joint point as the origin to obtain multiple transformed position information; and obtain the input trajectory according to the transformed position information.
[0121] In one feasible implementation, when the processor 1001 executes the transition action generation method, it is specifically used to: input the initial transition trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the third projected feature; input the first embedded feature and the third projected feature into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition trajectory embedded feature, and perform linear projection on the transition trajectory embedded feature to obtain the transition joint trajectory.
[0122] In one feasible implementation, when the processor 1001 executes the transition action generation method, it specifically performs the following steps: inputs the initial input data into the transition frame number prediction model for position concatenation encoding and linear projection to obtain the first projected features; inputs the first projected features into the encoder of the transition frame number prediction model for multi-head attention processing, normalization, activation, and residual perception processing to obtain the encoded features; inputs the encoded features into the decoder of the transition frame number prediction model for decoding multi-layer perception processing to obtain the decoded features; and performs linear projection on the decoded features to obtain the transition frame number.
[0123] In one feasible implementation, when the processor 1001 executes the transition action generation method, it specifically performs the following steps: obtaining a second embedded feature based on the number of transition frames, initial input data, and a first encoding module in the encoder of the transition action synthesis model; obtaining an initial residual action based on the number of transition frames and the trajectory of the transition joints; and obtaining a transition action sequence based on the number of transition frames, initial input data, the second embedded feature, the initial residual action, and a second encoding module and a decoder in the encoder of the transition action synthesis model.
[0124] In one feasible implementation, when the processor 1001 executes the transition action generation method, it specifically performs the following: inputting the number of transition frames and the initial input data into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fourth projected feature; inputting the fourth projected feature into the first encoding module in the encoder of the transition action synthesis model to obtain the second embedded feature.
[0125] In one feasible implementation, when the processor 1001 executes the transition action generation method, it specifically performs the following steps: inputting the initial residual action into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fifth projected feature; inputting the second embedded feature and the fifth projected feature into the second encoding module of the encoder of the transition action synthesis model to obtain the residual embedded feature; inputting the residual embedded feature into the decoder of the transition action synthesis model to obtain the residual action information; inputting the transition frame number and the initial input data into the transition action synthesis model for linear interpolation processing to obtain the initial linear action information; and obtaining the transition action sequence based on the initial linear action information and the residual action information.
[0126] Optionally, this application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, the processor performs the following steps:
[0127] In one feasible implementation, when the processor executes the transition action generation method, it is specifically used to: acquire initial input data, the initial input data including: global position information and local rotation information of each joint point in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint point in at least one frame of animation corresponding to the target action; obtain the number of transition frames based on the initial input data and a pre-trained transition frame prediction model; obtain the trajectory of the transition joint points based on the number of transition frames, the initial input data, and a pre-trained transition trajectory generation model; and obtain the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joint points, and a pre-trained transition action synthesis model.
[0128] In one feasible implementation, when the processor executes the transition action generation method, it is specifically used to: obtain a first embedded feature based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model; construct an initial transition trajectory based on the number of transition frames; and obtain the transition joint trajectory based on the first embedded feature, the initial transition trajectory, and the transition trajectory encoding module in the encoder of the transition trajectory generation model.
[0129] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following steps: obtaining an input trajectory based on initial input data; inputting the input trajectory into a transition trajectory generation model for position concatenation encoding and linear projection to obtain a second projected feature; and inputting the second projected feature into the input trajectory encoding module in the encoder of the transition trajectory generation model to obtain a first embedded feature.
[0130] In one feasible implementation, when the processor executes the transition action generation method, it is specifically used to: determine multiple target global position information that satisfy the selection information and belong to key joint points from the initial input data based on the selection information input by the user; and obtain the input trajectory based on the global position information of each target.
[0131] In one feasible implementation, when the processor executes the transition action generation method, it is specifically used to: obtain a coordinate system with the root joint as the origin corresponding to the global position information of each target; transform the global position information of each target according to the coordinate system with the root joint as the origin to obtain multiple transformed position information; and obtain the input trajectory according to the transformed position information.
[0132] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following: inputs the initial transition trajectory into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the third projected feature; inputs the first embedded feature and the third projected feature into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition trajectory embedded feature, and performs linear projection on the transition trajectory embedded feature to obtain the transition joint trajectory.
[0133] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following steps: inputs the initial input data into the transition frame number prediction model for position concatenation encoding and linear projection to obtain the first projected features; inputs the first projected features into the encoder of the transition frame number prediction model for multi-head attention processing, normalization, activation, and residual perception processing to obtain the encoded features; inputs the encoded features into the decoder of the transition frame number prediction model for decoding multi-layer perception processing to obtain the decoded features; and performs linear projection on the decoded features to obtain the transition frame number.
[0134] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following steps: obtaining a second embedded feature based on the number of transition frames, initial input data, and a first encoding module in the encoder of the transition action synthesis model; obtaining an initial residual action based on the number of transition frames and the trajectory of the transition joints; and obtaining a transition action sequence based on the number of transition frames, initial input data, the second embedded feature, the initial residual action, and a second encoding module and a decoder in the encoder of the transition action synthesis model.
[0135] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following: inputting the number of transition frames and the initial input data into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fourth projected feature; inputting the fourth projected feature into the first encoding module in the encoder of the transition action synthesis model to obtain the second embedded feature.
[0136] In one feasible implementation, when the processor executes the transition action generation method, it specifically performs the following steps: inputting the initial residual action into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fifth projected feature; inputting the second embedded feature and the fifth projected feature into the second encoding module of the encoder of the transition action synthesis model to obtain the residual embedded feature; inputting the residual embedded feature into the decoder of the transition action synthesis model to obtain the residual action information; inputting the transition frame number and the initial input data into the transition action synthesis model for linear interpolation processing to obtain the initial linear action information; and obtaining the transition action sequence based on the initial linear action information and the residual action information.
[0137] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.
[0140] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0141] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0142] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need further definition and explanation in subsequent figures. The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for generating transitional actions, characterized in that, The method includes: Acquire initial input data, which includes: global position information and local rotation information of each joint in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint in at least one frame of animation corresponding to the target action. The number of transition frames is obtained based on the initial input data and the pre-trained transition frame prediction model; The transition joint trajectory is obtained based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model; Based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model, a transition action sequence between the source action and the target action is obtained. The transition action synthesis model includes multiple cascaded encoders. Each encoder includes an input trajectory encoding module and a transition trajectory encoding module. The input trajectory encoding module includes a multi-head attention layer, a normalization layer, an activation function layer, and a residual perception layer.
2. The method according to claim 1, characterized in that, The step of obtaining the transition joint trajectory based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model includes: Based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model, the first embedded feature is obtained; Based on the number of transition frames, an initial transition trajectory is constructed; The transition joint trajectory is obtained based on the first embedded feature, the initial transition trajectory, and the transition trajectory encoding module in the encoder of the transition trajectory generation model.
3. The method according to claim 2, characterized in that, The step of obtaining the first embedded feature based on the initial input data and the input trajectory encoding module in the encoder of the transition trajectory generation model includes: Based on the initial input data, the input trajectory is obtained; The input trajectory is input into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the second projected features; The second projected feature is input into the input trajectory encoding module in the encoder of the transition trajectory generation model to obtain the first embedded feature.
4. The method according to claim 3, characterized in that, The step of obtaining the input trajectory based on the initial input data includes: Based on the selection information input by the user, determine the global position information of multiple targets that satisfy the selection information and belong to key nodes from the initial input data; The input trajectory is obtained based on the global position information of each target.
5. The method according to claim 4, characterized in that, The step of obtaining the input trajectory based on the global location information of each target includes: Obtain a coordinate system with the root joint point as the origin corresponding to the global position information of each target; Based on the coordinate system with each root joint as the origin, the global position information of each target is transformed to obtain multiple transformed position information; The input trajectory is obtained based on the transformed position information.
6. The method according to claim 2, characterized in that, The step of obtaining the transition joint trajectory based on the first embedded feature, the initial transition trajectory, and the transition trajectory generation model's encoder, via a transition trajectory encoding module, includes: The initial transition trajectory is input into the transition trajectory generation model for position concatenation encoding and linear projection to obtain the third projected feature. The first embedded feature and the third projected feature are input into the transition trajectory encoding module in the encoder of the transition trajectory generation model to obtain the transition trajectory embedded feature, and the transition trajectory embedded feature is linearly projected to obtain the transition joint trajectory.
7. The method according to claim 1, characterized in that, The step of obtaining the number of transition frames based on the initial input data and the pre-trained transition frame prediction model includes: The initial input data is input into the transition frame prediction model for position concatenation encoding and linear projection to obtain the first projected features. The first projected features are input into the encoder of the transition frame number prediction model, where multi-head attention processing, normalization, activation, and residual perception processing are performed to obtain the encoded features. The encoded features are input into the decoder of the transition frame number prediction model for decoding multilayer perceptual processing to obtain the decoded features; The number of transition frames is obtained by linearly projecting the decoded features.
8. The method according to any one of claims 1-7, characterized in that, The step of obtaining the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model includes: The second embedding feature is obtained based on the number of transition frames, the initial input data, and the first encoding module in the encoder of the transition action synthesis model; The initial residual motion is obtained based on the number of transition frames and the trajectory of the transition joints; The transition action sequence is obtained based on the number of transition frames, the initial input data, the second embedded feature, the initial residual action, and the second encoding module and decoder in the encoder of the transition action synthesis model.
9. The method according to claim 8, characterized in that, The step of obtaining the second embedded feature based on the number of transition frames, the initial input data, and the first encoding module in the encoder of the transition action synthesis model includes: The number of transition frames and the initial input data are input into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fourth projected feature. The fourth projected feature is input into the first encoding module of the encoder of the transition action synthesis model to obtain the second embedded feature.
10. The method according to claim 8, characterized in that, The step of obtaining the transition action sequence based on the number of transition frames, the initial input data, the second embedded features, the initial residual action, and the second encoding module and decoder in the encoder of the transition action synthesis model includes: The initial residual action is input into the transition action synthesis model for position concatenation encoding and linear projection to obtain the fifth projected feature. The second embedded feature and the fifth projected feature are input into the second encoding module of the encoder of the transition action synthesis model to obtain the residual embedded feature; The residual embedding features are input into the decoder of the transition action synthesis model to obtain residual action information; The number of transition frames and the initial input data are input into the transition motion synthesis model for linear interpolation to obtain initial linear motion information. The transition action sequence is obtained based on the initial linear action information and the residual action information.
11. A transitional motion generation device, characterized in that, The device includes: The acquisition module is used to acquire initial input data, which includes: global position information and local rotation information of each joint point in at least one frame of animation corresponding to the source action, and global position information and local rotation information of each joint point in at least one frame of animation corresponding to the target action. The first determining module is used to obtain the number of transition frames based on the initial input data and the pre-trained transition frame prediction model; The second determining module is used to obtain the transition joint trajectory based on the number of transition frames, the initial input data, and the pre-trained transition trajectory generation model; The third determining module is used to obtain the transition action sequence between the source action and the target action based on the number of transition frames, the initial input data, the trajectory of the transition joints, and the pre-trained transition action synthesis model. The transition action synthesis model includes multiple encoders in series. Each encoder includes an input trajectory encoding module and a transition trajectory encoding module. The input trajectory encoding module includes a multi-head attention layer, a normalization layer, an activation function layer, and a residual perception layer.
12. An electronic device, characterized in that, include: The electronic device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the transition action generation method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the transition action generation method as described in any one of claims 1-10.
Citation Information
Patent Citations
Training method and device of action completion model, completion method, equipment and medium
CN113345061A
Training method of animation generation model and animation generation method and device
CN114972591A