A method and device for generating an animation of a character interacting with a dynamic scene

By constructing a dynamic scene interactive animation generation model and utilizing feature extraction and the Transformer structure, the problem that existing technologies can only handle static scenes is solved. This enables the real-time generation of reasonable character movements in dynamic scenes, making it suitable for applications with dynamic scenes and high real-time requirements.

CN120747317BActive Publication Date: 2025-11-28SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511274752.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-11-28
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing methods for generating character animations can only handle static scenes and cannot adapt to real-time changes in dynamic scenes, resulting in the inability to generate reasonable dynamic character movements.

Method used

By constructing a model for generating interactive animations between characters and dynamic scenes, features are extracted using a near-human scene spatial feature extractor and a character motion feature extractor. Combined with a Transformer structure and an autoregressive approach, dynamic scene information is processed in real time to generate reasonable character motions.

Benefits of technology

It enables real-time generation of reasonable character actions in dynamic scenes, breaking through the limitations of static scenes. It is applicable to dynamic scenes and supports continuous generation without being limited by the generation time, making it suitable for scenes with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747317B_ABST
    Figure CN120747317B_ABST
Patent Text Reader

Abstract

The application provides a kind of generation method and device of character and dynamic scene interactive animation, it is related to character animation generation technical field, the method includes first obtaining character and scene interaction dataset, after dividing training set and test set, down-sampling training set data;Again, the near-human scene space feature of down-sampling data and the character action feature with momentum information are extracted, and stored in database according to time sequence;Then, retrieve the next frame data corresponding to K most similar features for each frame data of training set as future reference data;Subsequently, train the model with training data and future reference data, constrain joint rotation, displacement and the like through loss function, and train until convergence;Finally, in a dynamic scene, the reference data is cyclically retrieved in a self-recurrent manner, the next frame of action is generated and the pose is updated until the specified number of frames of animation is generated.The application can process dynamic scene changes in real time, generate reasonable and diverse character actions, continuously generate unlimited length animations, and is suitable for scenes with high real-time requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of character animation generation technology, and in particular to a method and apparatus for generating interactive animations between characters and dynamic scenes. Background Technology

[0002] Character-scene interaction animation generation aims to create smooth and natural character animations within a given scene, enabling interaction between the character and the scene. It has broad application value in animation and embodied intelligence, assisting in character animation design and production for film, television, and games. It can also guide robot behavior in real-world scenarios. In situations requiring real-time responses to environmental changes, such as game characters and real-world robots, real-time processing of dynamic scene changes and adjustments to subsequent character animation behavior are particularly important.

[0003] Currently, the field of character and scene interaction animation generation mainly focuses on the interaction between characters and static scenes, and there are mainly two-stage methods and one-stage methods.

[0004] The two-stage approach, which first plans the character's movement path or keyframes and then generates the specific character actions, first uses a path planning algorithm to plan the character's movement path on flat ground, and then uses reinforcement learning to train a reasonable character action generation strategy model so that the character can move along the planned path. However, this method is limited by existing path planning algorithms and can only be used on flat ground, making it unable to adapt to more complex terrains in virtual or real-world scenes; at the same time, because the movement path needs to be determined in advance, this type of method cannot adapt to dynamic scenes that change in real time.

[0005] One-stage methods that directly generate complete character animations use an end-to-end approach, taking the entire scene as input and directly outputting the character's motion sequence. However, this method can only generate character animations of a limited length and also cannot handle dynamic, changing scenes.

[0006] In existing technologies, when generating character animations, only a single static scene can be input, and it is impossible to generate corresponding character actions based on real-time changes in the scene. There is an urgent need for a method and device for generating character and scene interaction animations that can be applied to dynamic scenes. Summary of the Invention

[0007] To address this issue, this invention provides a method and apparatus for generating interactive animations between characters and dynamic scenes, which solves the problem in the prior art where only a single static scene can be input when generating character animations, and the corresponding character actions cannot be generated according to real-time changes in the scene.

[0008] To address the aforementioned problems, embodiments of the present invention provide a method for generating interactive animations between a character and a dynamic scene, the method comprising:

[0009] S1: Obtain the dataset of character and scene interaction and divide it into training set and test set. Downsample the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals.

[0010] S2: Input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; input the downsampled data into the human motion feature extractor to extract the human motion features with motion information of each frame; store the human motion features and the near-human scene spatial features in the database in chronological order to obtain the human motion feature-near-human scene spatial feature pair;

[0011] S3: For each frame of each data in the training set: extract the spatial features of the current frame's near human scene, calculate its similarity with all spatial features of the near human scene in the database, select the top K most similar features other than the current frame itself, and retrieve the next frame's human action feature-near human scene spatial feature pair corresponding to the K most similar features from the database as future reference data;

[0012] S4: Construct a character-scene interaction animation generation model and train the model to obtain a trained character-scene interaction animation generation model, specifically including:

[0013] S41: Train the model using the training set data and its future reference data as input;

[0014] S42: Based on the task action data of the next frame in the time series and the model output, construct a loss function constraining the following terms:

[0015] The SMPL model represents the rotation of each joint relative to its parent joint in a character's joint tree.

[0016] The displacement and rotation of the character's position relative to the current position in the next frame;

[0017] The distance L2 between each joint position and the distance L2 between the velocity vectors;

[0018] S43: Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other.

[0019] S5: Given a dynamic scene and the initial pose of the character, execute the following steps in an autoregressive manner until a specified number of frames are generated:

[0020] (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data;

[0021] (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame.

[0022] (c) Add the displacement and rotation to the current character pose and update the character animation state.

[0023] Preferably, the near-human scene spatial feature extractor performs:

[0024] Based on the character's position in the scene, m feature vectors are sampled radially and uniformly from the character's position as the center of the sphere. The length of each feature vector is the distance from the center of the sphere to the first spatial obstacle entity encountered in that direction. The direction unit vector and length of each feature vector are concatenated to obtain a spatial feature vector containing its direction and length information. The m spatial feature vectors are then concatenated horizontally to obtain a spatial feature matrix describing the spatial feature extraction time of the near-human scene at the current frame.

[0025] Preferably, the process of extracting the character motion features of each frame's motion information includes:

[0026] The 22 joint velocities represented by the SMPL model relative root joint position vector The data are concatenated into a vector, and then this vector is input into a character motion feature extractor through character motion-scene data pairs to extract character motion features of each frame, where the character motion feature extractor is represented as follows:

[0027] ;

[0028] In the formula, To drive the movement characteristics of the character, It is a multilayer perceptron. This indicates that operands, including numerical values ​​and one-dimensional vectors, are concatenated into a one-dimensional vector.

[0029] Preferably, the similarity calculation method for the spatial features of the near-human scene is as follows:

[0030] ;

[0031] In the formula, The L1 distance between two spatial characteristic matrices. and All are spatial feature matrices.

[0032] Preferably, the character-dynamic scene interaction animation generation model is a Transformer structure:

[0033] The input includes the character's motion features in the current frame. Spatial feature vectors and K future references , ,in This means expanding the matrix row by row into a one-dimensional vector. It is a multilayer perceptron. It is the spatial characteristic matrix;

[0034] The output is represented as:

[0035] ;

[0036] In the formula, The output of the model contains predicted values ​​for the next frame's motion, displacement, and rotation. and These are the character motion data and scene space data for the next frame corresponding to the K most similar features, respectively.

[0037] Preferably, the loss function includes:

[0038] The loss of rotation of each joint relative to its parent joint in a character joint tree represented by the SMPL model. :

[0039] ;

[0040] In the formula, For the first Rotation of a joint relative to its parent joint The joint rotation values ​​output by the SMPL model. This represents the total number of joints in the SMPL model. It is an L1 norm;

[0041] Loss of displacement and rotation of the character's position relative to the current position in the next frame :

[0042] ;

[0043] in, It is the character's displacement vector. It is the character displacement vector output by the model. It represents the change in the character's orientation. It is the change in the character's orientation output by the model. It is an L2 norm;

[0044] Loss of L2 distance at each joint position and L2 distance of velocity vector :

[0045] ;

[0046] In the formula, This indicates the character output by the model. The position vectors of each joint This indicates the character's position in the next frame of the actual data. The position vectors of each joint This indicates the character output by the model. The velocity vector of each joint This indicates the character's position in the next frame of the actual data. The velocity vector of each joint;

[0047] Total loss function during model training Represented as:

[0048] .

[0049] Preferably, the character pose update method is as follows:

[0050] ;

[0051] ;

[0052] In the formula, It is the position vector of the character's root joint in the world coordinate system. It is the character displacement vector output by the model. This is the updated character position vector. It is relative to the world coordinate system The direction the figure on the axis is facing. It is the change in the character's orientation output by the model. This refers to the direction the character is facing after the update.

[0053] Preferably, the world coordinate system is a three-dimensional Cartesian coordinate system determined by the right-hand rule, with the z-axis pointing in the opposite direction of gravity.

[0054] Preferably, the method for determining the direction the person is facing is as follows:

[0055] Take the direction vector from the left knee joint to the right knee joint of the figure, multiply it by the negative z-axis, and the projection of the resulting vector direction onto the xy plane is the counterclockwise rotation angle relative to the positive y-axis direction. This is the facing direction of the figure, expressed in radians.

[0056] This invention also provides an apparatus for generating animations of characters interacting with dynamic scenes. This apparatus is used to implement the aforementioned method for generating animations of characters interacting with dynamic scenes, specifically including:

[0057] The data sampling module is used to acquire the dataset of character and scene interaction and divide it into training set and test set. It downsamples the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals.

[0058] The feature extraction module is used to input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; input the downsampled data into the human motion feature extractor to extract the human motion features with motion information of each frame; and store the human motion features and the near-human scene spatial features in the database in chronological order to obtain human motion feature-near-human scene spatial feature pairs.

[0059] The future reference data retrieval module is used for each frame of each data in the training set: extracting the spatial features of the near human scene in the current frame, calculating its similarity with all spatial features of the near human scene in the database, selecting the top K most similar features other than the current frame itself, and retrieving the next frame's human action feature-near human scene spatial feature pair corresponding to the K most similar features from the database as future reference data;

[0060] The model training module is used to build and train a model for generating animations of characters interacting with dynamic scenes, resulting in a trained model for generating animations of characters interacting with dynamic scenes. This includes:

[0061] The model is trained using the training set data and its future reference data as input.

[0062] Based on the task action data of the next frame in the time series and the model output, the loss function is constructed with the following constraints:

[0063] The SMPL model represents the rotation of each joint relative to its parent joint in the character's joint tree; the displacement and rotation of the character's position in the next frame relative to the current position; the L2 distance of each joint position and the L2 distance of its velocity vector.

[0064] Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other.

[0065] The animation generation module, given a dynamic scene and the initial pose of a character, executes the following steps in an autoregressive manner until a specified number of frames are generated:

[0066] (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data;

[0067] (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame.

[0068] (c) Add the displacement and rotation to the current character pose and update the character animation state.

[0069] As can be seen from the above technical solutions, this invention application has the following beneficial effects:

[0070] (1) The present invention uses an autoregressive method to process dynamic scene information around the character in real time, and can generate corresponding character actions according to the real-time changes of the scene. This breaks through the limitation of existing technologies that can only be adapted to static scenes, and can be applied to dynamic scenes as well as static scenes.

[0071] (2) The present invention constructs a character action-scene space database and retrieves multiple future reference data similar to the current scene in real time to input the model. Combined with loss function training with multi-dimensional constraints (joint rotation, position, speed, etc.), the rationality and diversity of generated character actions are improved. Moreover, the autoregressive characteristics of the model support continuous generation of character actions without being limited by the generation time, and are suitable for scenarios with high real-time requirements (such as games and real-world robot interaction). Attached Figure Description

[0072] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly described below. Referring to the accompanying drawings will provide a clearer understanding of the features and advantages of the present invention. The drawings are illustrative and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. Wherein:

[0073] Figure 1 A flowchart illustrating a method for generating interactive animations between a character and a dynamic scene, provided by this invention;

[0074] Figure 2 A block diagram of a device for generating interactive animations between a character and a dynamic scene, provided by the present invention. Detailed Implementation

[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0076] Example 1: To address the problem in existing technologies where only a single static scene can be input when generating character animations, and the ability to generate corresponding character actions based on real-time scene changes is not possible. For example... Figure 1 As shown, this invention proposes a method for generating interactive animations between characters and dynamic scenes, the method comprising:

[0077] S1: Obtain the dataset of character and scene interaction and divide it into training set and test set. Downsample the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals.

[0078] S2: Input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; input the downsampled data into the human motion feature extractor to extract the human motion features with motion information of each frame; store the human motion features and near-human scene spatial features in chronological order into the database;

[0079] S3: For each frame of each data in the training set: extract the spatial features of the current frame's near human scene, calculate its similarity with all spatial features of the near human scene in the database, select the top K most similar features other than the current frame itself, and retrieve the next frame's human action features-near human scene spatial feature pairs corresponding to the K most similar features from the database as future reference data;

[0080] S4: Construct a character-scene interaction animation generation model and train the model to obtain a trained character-scene interaction animation generation model, specifically including:

[0081] S41: Train the model using the training set data and its future reference data as input;

[0082] S42: Based on the task action data of the next frame in the time series and the model output, construct a loss function constraining the following terms:

[0083] The SMPL model represents the rotation of each joint relative to its parent joint in a character's joint tree.

[0084] The displacement and rotation of the character's position relative to the current position in the next frame;

[0085] The distance L2 between each joint position and the distance L2 between the velocity vectors;

[0086] S43: Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other.

[0087] S5: Given a dynamic scene and the initial pose of the character, execute the following steps in an autoregressive manner until a specified number of frames are generated:

[0088] (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data;

[0089] (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame.

[0090] (c) Add the displacement and rotation to the current character pose and update the character animation state.

[0091] As can be seen from the above technical solution, this invention proposes a method for generating interactive animations between characters and dynamic scenes. First, it acquires a dataset of character-scene interaction and divides it into training and testing sets. The training set data is downsampled to simplify the data and improve processing efficiency. Then, it extracts near-human scene spatial features from the downsampled data, capturing key information about the character's surrounding environment and the character's motion features that drive movement, preserving motion continuity, and storing it in a database to provide a foundation for subsequent reference. Next, it retrieves the next frame data corresponding to the K most similar features for each frame of the training set as future reference data, enhancing the rationality and diversity of motion generation. Subsequently, it constructs a character-dynamic scene interaction animation generation model, using training data and future reference data as input, and trains it to convergence through a loss function that constrains joint rotation, displacement, joint position, and velocity, ensuring the physical consistency and coherence of the motion. Finally, in the dynamic scene, it retrieves reference data in real time using an autoregressive approach and generates the next frame action, overlaying and updating the pose to achieve real-time, continuous interactive animation generation in dynamic scenes, solving the problem that existing technologies are only suitable for static scenes.

[0092] In step S1, the character and scene interaction dataset is obtained and divided into training set and test set. For each data point in the training set, the multi-frame character action-scene data pairs are downsampled at fixed time intervals.

[0093] Specifically, this embodiment uses the CIRCLE dataset as the basis for training and testing. This dataset contains several common home scenarios (such as kitchen, bedroom, study, etc.) and has been pre-divided into training and testing sets, which are used for model training and performance verification, respectively.

[0094] Meanwhile, to reduce the amount of data and the computational complexity of subsequent processing, the character action-scene data pairs in the dataset were downsampled: one frame was selected every two frames in chronological order, retaining key temporal information and removing redundant data. This provides more efficient input data for subsequent feature extraction (such as extracting spatial features of near-human scenes and character action features) and model training. This step belongs to the data preprocessing stage, aiming to improve the efficiency and stability of model training by reasonably reducing the data scale.

[0095] In step S2, the downsampled data is input into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; the downsampled data is input into the human motion feature extractor to extract the human motion features with motion information of each frame; the human motion features and near-human scene spatial features are stored in the database in chronological order to obtain the human motion feature-near-human scene spatial feature pair.

[0096] Specifically, the working principle of the near-human scene spatial feature extractor in this embodiment is as follows:

[0097] Using the character's position in the scene as the center of a sphere, m feature vectors are sampled radially and uniformly within the sphere centered on this center (m=1024 in this embodiment). The length of each feature vector is defined as the distance from the center of the sphere to the first spatial obstacle entity encountered in the direction of that vector; the direction of the vector is represented by a unit vector.

[0098] Next, the unit vector (direction information) and length (distance information) of each feature vector are concatenated to form a spatial feature vector containing the direction and length of that vector. Finally, these 1024 spatial feature vectors are concatenated horizontally to obtain a spatial feature matrix describing the spatial space of the near-human scene at the current frame's near-human scene at the time of feature extraction. In this embodiment, the size of this matrix is ​​4×1024.

[0099] In this way, the near-human scene spatial feature extractor can accurately capture the distribution of spatial obstacles around the character, providing key scene feature support for the subsequent model to understand the scene and generate reasonable character actions.

[0100] Furthermore, the logic for extracting the character's action features based on the volume information is as follows:

[0101] The core of this feature is the fusion of the motion state (momentum) and position information of human joints, which is achieved through the following steps:

[0102] 1. Feature Source Selection: Based on the human body structure represented by the Skinned Multi-Person Linear Model (SMPL model), focusing on 22 joints, two types of key information are extracted—the motion velocity of each joint (denoted as...). This reflects momentum information, i.e., the trend and speed of joint movement, as well as the position vector of each joint relative to the root joint (denoted as...). This reflects the relative position of the joint in space.

[0103] It should be noted that the SMPL model is a parametric 3D human body model that controls the movement of human mesh vertices and joints through shape and pose parameters. This invention uses the SMPL model as the basis for human motion representation. This model defines a human motion tree structure containing N joints (e.g., N=22). The rotation of each joint... All coordinates are referenced to their parent joint coordinate system, conforming to biological kinematic constraints. The model output includes joint positions. and speed It is used to construct character movement features.

[0104] 2. Vector concatenation: combining the velocity vectors of the same joint. relative root joint position vector splicing (through) The operation merges the two into a single one-dimensional vector, forming a comprehensive vector that simultaneously contains motion state and position information.

[0105] 3. Feature Extraction: The concatenated composite vector is input into a character motion feature extractor (this extractor is a multilayer perceptron MLP). Through the nonlinear transformation of the MLP, the final feature of each frame is obtained (using...). Character motion features that represent the amount of motion information (indicated by the frame number), denoted as ,Right now .

[0106] This process integrates velocity (momentum) and position information, enabling the extracted character movement features to reflect the spatial position of joints while preserving the continuity and trend of movement, thus providing a foundation for subsequent model generation of coherent and reasonable dynamic movements.

[0107] In step S3, for each frame of each data in the training set: extract the spatial features of the current frame's near-human scene, calculate its similarity with all spatial features of the near-human scene in the database, select the top K most similar features (K=3 in this embodiment) excluding the current frame itself, and retrieve the next frame's character action features-near-human scene spatial feature pairs corresponding to the K most similar features from the database as future reference data.

[0108] Specifically, in this embodiment, the similarity calculation method for spatial features of near-human scenes is based on the L1 distance of the spatial feature matrix of near-human scenes, as explained below:

[0109] In this embodiment, the method for calculating the similarity between two spatial features is based on the L1 distance of the spatial feature matrix of a near-human scene, as detailed below:

[0110] Computational Object: This method measures the similarity of near-scene spatial features between two person action-scene data pairs. Specifically, and These represent the spatial features (i.e., spatial feature matrices, such as a 4×1024 matrix in the example) of two character action-scene data pairs obtained by the near-human scene spatial feature extractor.

[0111] L1 distance definition: Formula In this context, L1 distance (Manhattan distance) refers to the sum of the absolute values ​​of the differences between corresponding elements of two spatial feature matrices. That is, for a matrix... and The similarity measure between the two pairs is obtained by subtracting the elements at each position, taking the absolute values, and summing them. .

[0112] Application purpose: The smaller the value, the more similar the spatial features of two near-human scenes. This calculation method can retrieve the top K data points from the database that are closest to the current scene features, providing future reference data for the model and improving the rationality of generated actions.

[0113] In step S4, a character-dynamic scene interaction animation generation model is constructed and trained to obtain a trained character-dynamic scene interaction animation generation model. This specifically includes the following steps:

[0114] S41: Train the model using the training set data and its future reference data as input.

[0115] S42: Based on the task action data of the next frame in the time series and the model output, construct a loss function constraining the following terms:

[0116] The SMPL model represents the rotation of each joint relative to its parent joint in a character's joint tree.

[0117] The displacement and rotation of the character's position relative to the current position in the next frame;

[0118] The distance L2 between each joint position and the distance L2 between the velocity vectors.

[0119] S43: Train the model until the loss function converges to obtain a trained model for generating interactive animations between characters and dynamic scenes.

[0120] Specifically, in this embodiment, the character-dynamic scene interaction animation generation model is a Transformer structure:

[0121] The input includes the character's motion features in the current frame. Spatial feature vectors and K future references , ,in This means expanding the matrix row by row into a one-dimensional vector. It is a multilayer perceptron. It is the spatial characteristic matrix;

[0122] The output is represented as:

[0123] ;

[0124] In the formula, The output of the model contains predicted values ​​for the next frame's motion, displacement, and rotation. and These are the character motion data and scene space data for the next frame corresponding to the K most similar features, respectively.

[0125] For example, when K=3, the model's output is represented as:

[0126] ;

[0127] In the formula, and The model outputs the character motion data and scene space data for the next frame, corresponding to the three most similar features. It contains 8 vectors, which are concatenated into a one-dimensional vector and then passed through a multilayer perceptron to obtain the character's motion and displacement relative to the current character position in the next frame, represented as:

[0128] ;

[0129] in, It represents the rotation of each joint relative to its parent joint in the character's joint tree as depicted in the SMPL model. It is the character displacement vector output by the model. It is the change in the character's orientation output by the model. It is the position vector of each joint of the character. It is the velocity vector of each joint of the character.

[0130] Furthermore, the loss function in this embodiment includes:

[0131] The loss of rotation of each joint relative to its parent joint in a character joint tree represented by the SMPL model. :

[0132] ;

[0133] In the formula, For the first The rotation of each joint relative to its parent joint is represented using a continuous 6-dimensional rotation representation. The joint rotation values ​​output by the SMPL model. This represents the total number of joints in the SMPL model. It is an L1 norm.

[0134] Loss of displacement and rotation of the character's position relative to the current position in the next frame :

[0135] ;

[0136] in, It is the character's displacement vector. It is the character displacement vector output by the model. It represents the change in the character's orientation. It is the change in the character's orientation output by the model. It is an L2 norm.

[0137] Loss of L2 distance at each joint position and L2 distance of velocity vector :

[0138] ;

[0139] In the formula, This indicates the character output by the model. The position vectors of each joint This indicates the character's position in the next frame of the actual data. The position vectors of each joint This indicates the character output by the model. The velocity vector of each joint This indicates the character's position in the next frame of the actual data. The velocity vector of each joint.

[0140] Total loss function during model training Represented as:

[0141] .

[0142] In step S5, given the dynamic scene and the initial pose of the character, the following steps are executed in an autoregressive manner until a specified number of frames are generated: (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data; (b) Using the current character action, scene information and future reference data as input, generate the character action, displacement and rotation of the next frame through the trained character and dynamic scene interaction animation generation model; (c) Superimpose the displacement and rotation onto the current character pose and update the character animation state.

[0143] Specifically, the method for updating character pose is as follows:

[0144] ;

[0145] ;

[0146] In the formula, It is the position vector of the character's root joint in the world coordinate system. It is the character displacement vector output by the model. This is the updated character position vector. It is relative to the world coordinate system The direction the figure on the axis is facing. It is the change in the character's orientation output by the model. This refers to the direction the character is facing after the update.

[0147] Furthermore, in this embodiment, the world coordinate system adopts a three-dimensional Cartesian coordinate system based on the right-hand rule, wherein the coordinate axes are defined as follows: The axis points in the opposite direction of gravity (i.e., perpendicular to the ground and upward). shaft and The axes form a horizontal plane (following the right-hand rule, with the right thumb pointing towards...). Positive axis direction, index finger pointing Positive axis direction, middle finger pointing (Positive direction of the axis).

[0148] Furthermore, in this embodiment, the direction the character is facing is calculated through the following steps and ultimately expressed in radians:

[0149] Take the direction vector from the left knee joint to the right knee joint of the character (denoted as vector A);

[0150] Multiply vector A by the negative z-axis (the direction perpendicular to the downward direction, denoted as vector B) to obtain a new vector C (the direction of which is perpendicular to both A and B).

[0151] Projecting vector C onto the xy plane (horizontal plane) yields the projected vector D;

[0152] The counterclockwise rotation angle of the projection vector D relative to the positive y-axis of the world coordinate system is the facing direction of the character (measured in radians).

[0153] In summary, this method updates the character's position and orientation through simple vector overlay, and combined with a clear coordinate system definition and facing direction calculation rules, ensures the spatial consistency and logical rationality of the character's movement in dynamic scenes.

[0154] Example 2: Figure 2 As shown, the present invention provides a device for generating animations of interaction between a character and a dynamic scene. This device is used to implement the method for generating animations of interaction between a character and a dynamic scene as described in Embodiment 1 above, and specifically includes:

[0155] The data sampling module 100 is used to acquire the dataset of character and scene interaction and divide it into training set and test set. It downsamples the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals.

[0156] The feature extraction module 200 is used to input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; input the downsampled data into the human motion feature extractor to extract the human motion features with motion information of each frame; and store the human motion features and near-human scene spatial features in the database in chronological order to obtain human motion feature-near-human scene spatial feature pairs.

[0157] The future reference data retrieval module 300 is used for each frame of each data in the training set: extracting the spatial features of the near human scene in the current frame, calculating its similarity with all spatial features of the near human scene in the database, selecting the top K most similar features other than the current frame itself, and retrieving the next frame's human action features-near human scene spatial feature pairs corresponding to the K most similar features from the database as future reference data;

[0158] The model training module 400 is used to construct and train a character-dynamic scene interaction animation generation model, resulting in a trained model that generates such an animation. This model includes:

[0159] The model is trained using the training set data and its future reference data as input.

[0160] Based on the task action data of the next frame in the time series and the model output, the loss function is constructed with the following constraints:

[0161] The SMPL model represents the rotation of each joint relative to its parent joint in the character's joint tree; the displacement and rotation of the character's position in the next frame relative to the current position; the L2 distance of each joint position and the L2 distance of its velocity vector.

[0162] Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other.

[0163] Animation generation module 500, given a dynamic scene and the initial pose of a character, executes the following steps in an autoregressive manner until a specified number of frames are generated:

[0164] (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data;

[0165] (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame.

[0166] (c) Add the displacement and rotation to the current character pose and update the character animation state.

[0167] This embodiment provides a device for generating animations of characters interacting with dynamic scenes, used to implement the aforementioned method for generating animations of characters interacting with dynamic scenes. Therefore, the specific implementation of the device for generating animations of characters interacting with dynamic scenes can be found in the previous embodiment section of the method for generating animations of characters interacting with dynamic scenes. For example, the data sampling module 100, the feature extraction module 200, the future reference data retrieval module 300, the model training module 400, and the animation generation module 500 are respectively used to implement steps S1, S2, S3, S4, and S5 in the above-mentioned method for generating animations of characters interacting with dynamic scenes. Therefore, its specific implementation can be referred to the description of the corresponding embodiments. To avoid redundancy, it will not be repeated here.

[0168] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0170] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0171] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for generating interactive animations between characters and dynamic scenes, characterized in that, include: S1: Obtain the dataset of character and scene interaction and divide it into training set and test set. Downsample the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals. S2: Input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; The downsampled data is input into the character motion feature extractor to extract the character motion features with motion information in each frame; the character motion features and the near-human scene spatial features are stored in the database in chronological order to obtain the character motion feature-near-human scene spatial feature pair. S3: For each frame of each data in the training set: extract the spatial features of the current frame's near human scene, calculate its similarity with all spatial features of the near human scene in the database, select the top K most similar features other than the current frame itself, and retrieve the next frame's human action feature-near human scene spatial feature pair corresponding to the K most similar features from the database as future reference data; S4: Construct a character-scene interaction animation generation model and train the model to obtain a trained character-scene interaction animation generation model, specifically including: S41: Train the model using the training set data and its future reference data as input; S42: Based on the task action data of the next frame in the time series and the model output, construct a loss function constraining the following terms: The SMPL model represents the rotation of each joint relative to its parent joint in a character's joint tree. The displacement and rotation of the character's position relative to the current position in the next frame; The distance L2 between each joint position and the distance L2 between the velocity vectors; S43: Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other. S5: Given a dynamic scene and the initial pose of the character, execute the following steps in an autoregressive manner until a specified number of frames are generated: (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data; (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame. (c) Add the displacement and rotation to the current character pose and update the character animation state.

2. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The near-human scene spatial feature extractor performs the following: Based on the character's position in the scene, m feature vectors are sampled radially and uniformly from the character's position as the center of the sphere. The length of each feature vector is the distance from the center of the sphere to the first spatial obstacle entity encountered in that direction. The direction unit vector and length of each feature vector are concatenated to obtain a spatial feature vector containing its direction and length information. The m spatial feature vectors are then concatenated horizontally to obtain a spatial feature matrix describing the spatial feature extraction time of the near-human scene at the current frame.

3. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The process of extracting character motion features from each frame includes: The 22 joint velocities represented by the SMPL model relative root joint position vector The data are concatenated into a vector, and then this vector is input into a character motion feature extractor through character motion-scene data pairs to extract character motion features of each frame, where the character motion feature extractor is represented as follows: ; In the formula, To drive the movement characteristics of the character, It is a multilayer perceptron. This indicates that operands, including numerical values ​​and one-dimensional vectors, are concatenated into a one-dimensional vector.

4. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The similarity calculation method for the spatial features of the near-human scene is as follows: ; In the formula, The L1 distance between two spatial characteristic matrices. and All are spatial feature matrices.

5. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The character-dynamic scene interaction animation generation model is a Transformer structure: The input includes the character's motion features in the current frame. Spatial feature vectors and K future references , ,in This means expanding the matrix row by row into a one-dimensional vector. It is a multilayer perceptron. It is the spatial characteristic matrix; The output is represented as: ; In the formula, The output of the model contains predicted values ​​for the next frame's motion, displacement, and rotation. and These are the character motion data and scene space data for the next frame corresponding to the K most similar features, respectively.

6. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The loss function includes: The loss of rotation of each joint relative to its parent joint in a character joint tree represented by the SMPL model. : ; In the formula, For the first Rotation of a joint relative to its parent joint The joint rotation values ​​output by the SMPL model. This represents the total number of joints in the SMPL model. It is an L1 norm; Loss of displacement and rotation of the character's position relative to the current position in the next frame : ; in, It is the character's displacement vector. It is the character displacement vector output by the model. It represents the change in the character's orientation. It is the change in the character's orientation output by the model. It is an L2 norm; Loss of L2 distance at each joint position and L2 distance of velocity vector : ; In the formula, This indicates the character output by the model. The position vectors of each joint This indicates the character's position in the next frame of the actual data. The position vectors of each joint This indicates the character output by the model. The velocity vector of each joint This indicates the character's position in the next frame of the actual data. The velocity vector of each joint; Total loss function during model training Represented as: 。 7. The method for generating interactive animations between characters and dynamic scenes according to claim 1, characterized in that, The method for updating the character's pose is as follows: ; ; In the formula, It is the position vector of the character's root joint in the world coordinate system. It is the character displacement vector output by the model. This is the updated character position vector. It is relative to the world coordinate system The direction the figure on the axis is facing. It is the change in the character's orientation output by the model. This refers to the direction the character is facing after the update.

8. The method for generating interactive animations between characters and dynamic scenes according to claim 7, characterized in that, The world coordinate system is a three-dimensional Cartesian coordinate system defined by the right-hand rule, with the z-axis pointing in the opposite direction of gravity.

9. The method for generating interactive animations between characters and dynamic scenes according to claim 7, characterized in that, The method for determining the direction the character is facing is as follows: Take the direction vector from the left knee joint to the right knee joint of the figure, multiply it by the negative z-axis, and the projection of the resulting vector direction onto the xy-plane is relative to... The counterclockwise rotation angle along the positive axis is the direction the person is facing, expressed in radians.

10. A device for generating interactive animation between a character and a dynamic scene, characterized in that, The apparatus is used to implement the method for generating interactive animations between characters and dynamic scenes as described in any one of claims 1 to 9, specifically including: The data sampling module is used to acquire the dataset of character and scene interaction and divide it into training set and test set. It downsamples the multi-frame character action-scene data pairs of each data in the training set at fixed time intervals. The feature extraction module is used to input the downsampled data into the near-human scene spatial feature extractor to extract the near-human scene spatial features of each frame; input the downsampled data into the human motion feature extractor to extract the human motion features with motion information of each frame; and store the human motion features and the near-human scene spatial features in the database in chronological order to obtain human motion feature-near-human scene spatial feature pairs. The future reference data retrieval module is used for each frame of each data in the training set: extracting the spatial features of the near human scene in the current frame, calculating its similarity with all spatial features of the near human scene in the database, selecting the top K most similar features other than the current frame itself, and retrieving the next frame's human action feature-near human scene spatial feature pair corresponding to the K most similar features from the database as future reference data; The model training module is used to build and train a model for generating animations of characters interacting with dynamic scenes, resulting in a trained model for generating animations of characters interacting with dynamic scenes. This includes: The model is trained using the training set data and its future reference data as input. Based on the task action data of the next frame in the time series and the model output, the loss function is constructed with the following constraints: The SMPL model represents the rotation of each joint relative to its parent joint in the character's joint tree; the displacement and rotation of the character's position in the next frame relative to the current position; the L2 distance of each joint position and the L2 distance of its velocity vector. Train the model until the loss function converges to obtain a trained model for generating animations of characters and dynamic scenes interacting with each other. The animation generation module, given a dynamic scene and the initial pose of a character, executes the following steps in an autoregressive manner until a specified number of frames are generated: (a) Extract the spatial features of the near-human scene in the current frame, retrieve the top K most similar features other than the current frame itself, and obtain the data of the next frame as future reference data; (b) Using the current character's actions, scene information, and future reference data as input, the trained character and dynamic scene interaction animation generation model generates the character's actions, displacement, and rotation for the next frame. (c) Add the displacement and rotation to the current character pose and update the character animation state.

Citation Information

Patent Citations

  • Virtual scene interaction method and system

    CN118377384A

  • Multi-condition human body action generation method and system based on scene and text

    CN119741408A