Action redirection method, device, electronic device and storage medium based on dense geometric interaction perception

By extracting semantically consistent feature points in motion redirection and establishing a dense mesh interaction field, the conflict between skeletal motion redirection and geometric correction is resolved, and the accuracy of motion redirection and contact preservation are achieved.

CN119762634BActive Publication Date: 2025-09-23TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411647500.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-09-23
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing motion redirection methods ignore the body's geometry, resulting in conflicts between skeletal motion redirection and geometric correction, and causing problems such as jitter, penetration, and contact mismatch.

Method used

By extracting semantically consistent feature points from the resting postures of the source and target characters respectively, calculating the tangent space coordinate system of the feature points, establishing dense grid correspondences, and using the dense grid interaction field (DMI) for action sequence prediction, the geometric interactions are directly modeled to avoid the geometric correction stage.

Benefits of technology

It promotes geometric interaction awareness between different mesh topologies, preserves motion semantics, prevents self-penetration, and ensures contact preservation in a single pass.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762634B_ABST
    Figure CN119762634B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, electronic device and storage medium for action redirection based on dense geometric interaction perception, which relates to the technical field of skinned animation. The method comprises: extracting semantically consistent feature points from the resting postures of a source character and a target character respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, performing spatiotemporal representation of the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, performing action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid and the second feature point grid to obtain a target action sequence of the target character. The method can promote geometric interaction perception action redirection between different grid topologies in a single process, which not only preserves the semantics of motion, but also prevents self-penetration and ensures contact maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of skinned animation, and in particular to a motion redirection method, device, electronic device and storage medium based on dense geometric interactive perception. Background Art

[0002] Skinned character animation is prevalent in virtual reality, game development, and various other fields. However, driving a target character often presents significant challenges due to differences in body proportions between the source and target characters. Motion retargeting is crucial for adjusting for these differences in body proportions and maintaining the integrity of the source motion signature in the target character's animation.

[0003] Existing motion retargeting methods often ignore the body geometry or add a geometry correction stage after skeletal motion retargeting, which leads to conflicts between skeletal motion retargeting and geometry correction, resulting in problems such as jitter, clipping, and contact mismatch. Summary of the Invention

[0004] The present invention provides a motion redirection method, device, electronic device and storage medium based on dense geometric interaction perception, which is used to solve the defects of motion redirection in the existing technology that causes jitter, self-penetration and contact mismatch. It can promote the motion redirection of geometric interaction perception between different grid topologies in a single processing, which not only retains the motion semantics, but also prevents self-penetration and ensures contact maintenance.

[0005] The present invention provides an action redirection method based on dense geometric interaction perception, comprising the following steps:

[0006] Extracting semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid each include positions and feature vectors of a plurality of feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship;

[0007] Performing a spatiotemporal representation on the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of a plurality of interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between body geometric shapes of the source character;

[0008] An action sequence is predicted based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, wherein the target action sequence is aligned with a geometric shape of the target character and the source grid interaction field.

[0009] According to the present invention, a method for action redirection based on dense geometric interaction perception is provided, wherein semantically consistent feature points are extracted from the resting postures of a source character and a target character, and a tangent space coordinate system of the feature points is calculated to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, including:

[0010] Acquire a resting posture of the source character and the target character, wherein the resting posture includes a plurality of skeletons;

[0011] For each bone of the source character and the target character, a target transformation is used to obtain semantic coordinates. The target transformation includes projecting a ray from the bone axis of each bone to the vertical plane corresponding to the bone axis. The origin parameter, direction parameter, and bone index of the ray constitute the semantic coordinates of the feature point. The semantic coordinates of the feature point are used to describe the connection relationship between the feature point and the bone.

[0012] Valid feature points are screened based on whether the ray of the feature point in the semantic coordinates of the feature point intersects with the mesh connected to the skeleton, to obtain a first feature point mesh of the source character and a second feature point mesh of the target character, wherein if the ray of the feature point intersects with the mesh connected to the skeleton, the feature point is valid.

[0013] According to an action redirection method based on dense geometric interaction perception provided by the present invention, performing spatiotemporal representation on the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field includes:

[0014] Performing a forward dynamics operation on the first feature point grid to obtain feature point features, the feature point features including a position of the feature point and a tangent matrix;

[0015] According to the feature point features and the geometric interaction model, paired interaction features between feature points are obtained to obtain the source grid interaction field. The geometric interaction model is obtained based on feature modeling between paired interaction feature points.

[0016] According to an action redirection method based on dense geometric interaction perception provided by the present invention, obtaining paired interaction features between feature points based on the feature point features and the geometric interaction model to obtain the source grid interaction field includes:

[0017] Obtaining initial pairwise interaction features based on the feature point features and the geometric interaction model;

[0018] The initial paired interaction features are subjected to feature screening using target screening conditions to obtain the source grid interaction field, wherein the target screening conditions include limiting the interaction between feature points to feature points of preset body parts, and / or limiting the number of feature points of each body part to no more than a preset number.

[0019] According to the present invention, a method for action redirection based on dense geometric interaction perception is provided, wherein the action sequence is predicted based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain the target action sequence of the target character, including:

[0020] Extracting features of the source grid interaction field to obtain grid interaction features;

[0021] Extracting features from the source motion sequence to obtain source joint rotation features;

[0022] Performing geometric feature extraction on the first feature point grid and the second feature point grid respectively to obtain a first geometric feature vector and a second geometric feature vector;

[0023] The first geometric feature vector and the second geometric feature vector are input into an encoder of a trained redirection network, and the mesh interaction feature and the source joint rotation feature are input into a decoder of the trained redirection network to obtain the target action sequence.

[0024] According to a method for action redirection based on dense geometric interaction perception provided by the present invention, before extracting semantically consistent feature points from the resting postures of a source character and a target character, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, the method further includes:

[0025] The initial redirection network is trained according to the target loss function to obtain the trained redirection network, wherein the target loss function includes reconstruction loss, mesh interaction field consistency loss, adversarial loss and end effector loss, wherein the reconstruction loss is used to minimize the motion change during the redirection process, the mesh interaction field consistency loss is used to maintain the geometric interaction between the source mesh interaction field and the target mesh interaction field, the adversarial loss is used to promote real action redirection, and the end effector loss is used to promote the consistent direction of the end effector in the redirection motion.

[0026] The present invention also provides an action redirection device based on dense geometric interaction perception, comprising the following modules:

[0027] a grid acquisition module, configured to extract semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculate a tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid each include positions and feature vectors of a plurality of feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship;

[0028] an interaction acquisition module, configured to perform spatiotemporal representation of the first feature point grid according to a source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of a plurality of interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between body geometric shapes of the source character;

[0029] An action prediction module is used to predict an action sequence based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, wherein the target action sequence is aligned with the geometric shape of the target character and the source grid interaction field.

[0030] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for action redirection of dense geometric interaction perception as described above is implemented.

[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described dense geometric interaction perception action redirection methods.

[0032] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described dense geometric interaction perception action redirection methods.

[0033] The present invention provides a method, device, electronic device and storage medium for action redirection based on dense geometric interaction perception. The method extracts semantically consistent feature points from the resting postures of the source character and the target character respectively, and calculates the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character. The first feature point grid is spatiotemporally represented according to the source action sequence of the source character to obtain a source grid interaction field. The action sequence is predicted based on the source grid interaction field, the source action sequence, the first feature point grid and the second feature point grid to obtain a target action sequence of the target character. The method can promote action redirection based on geometric interaction perception between different grid topologies in one processing, which not only retains the motion semantics but also prevents self-penetration and ensures contact maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 It is a flowchart of the action redirection method based on dense geometric interaction perception provided by the present invention.

[0036] Figure 2 It is a schematic diagram comparing the method provided by the present invention with the existing method.

[0037] Figure 3 It is a flow chart of the method for obtaining a feature point grid provided by the present invention.

[0038] Figure 4 Schematic diagram of the process of deriving semantically consistent feature points provided by the present invention.

[0039] Figure 5 It is a flow chart of the method for obtaining the source grid interaction field provided by the present invention.

[0040] Figure 6 It is a flowchart of the method for predicting an action sequence provided by the present invention.

[0041] Figure 7 It is a schematic diagram of the overall flow of the action redirection method provided by the present invention.

[0042] Figure 8 It is a structural diagram of the dense geometric interactive perception action redirection device provided by the present invention.

[0043] Figure 9It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0045] Skinned character animation is prevalent in virtual reality, game development, and various other fields. However, driving a target character often presents significant challenges due to differences in body proportions between the source and target characters. Motion retargeting is crucial for adjusting for these differences in body proportions and maintaining the integrity of the source motion signature in the target character's animation.

[0046] Existing motion retargeting methods often ignore the body geometry or add a geometry correction stage after skeletal motion retargeting, which leads to conflicts between skeletal motion retargeting and geometry correction, resulting in problems such as jitter, clipping, and contact mismatch.

[0047] In view of this, an embodiment of the present invention provides a method for motion redirection based on dense geometric interaction perception. By extracting semantically consistent feature points from the resting postures of the source character and the target character, and calculating the tangent space coordinate system of the feature points, a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character are obtained. The first feature point grid is spatiotemporally represented according to the source action sequence of the source character to obtain a source grid interaction field. The action sequence is predicted based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence for the target character. This method can facilitate geometric interaction-aware motion redirection between different mesh topologies in a single process, not only preserving motion semantics but also preventing self-penetration and ensuring contact maintenance.

[0048] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0049] Figure 1: is a flow chart of the motion redirection method based on dense geometric interaction perception provided by the present invention. The motion redirection method based on dense geometric interaction perception can be applied to electronic devices, which can be various types of devices with information processing capabilities during implementation. For example, the electronic device can include a personal computer, a laptop, a PDA or a server, etc.; the electronic device can also be a mobile terminal, for example, the mobile terminal can include a mobile phone, a car computer, a tablet computer or a projector, etc. Figure 1 As shown, the method may include the following steps 101 to 103:

[0050] Step 101: Semantically consistent feature points are extracted from the resting postures of the source character and the target character respectively, and the tangent space coordinate system of the feature points is calculated to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character. The first feature point grid and the second feature point grid both include positions and feature vectors of multiple feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence.

[0051] It should be noted that the resting poses of the source and target characters can include 3D model data. "Semantic consistency" here refers to feature points that have the same or similar meaning or function between the two characters. In other words, these feature points not only correspond in physical location (e.g., at the knees, wrists, etc.), but also play the same role or function in the character animation or behavior. For example, in animation, the knee is a key joint that controls leg flexion and extension. Knees as feature points should be semantically consistent in both the source and target characters, as they both play a key role in controlling leg movement. Extracting semantically consistent feature points from the resting poses of the source and target characters can be accomplished by automatically extracting feature points from 3D models using existing algorithms or tools (such as OpenPose and DeepLabCut), or by employing other extraction methods.

[0052] In addition, the tangent space coordinate system for calculating the feature points can be to define a local two-dimensional plane coordinate system tangent to the surface for each feature point. This coordinate system is crucial for understanding the geometric shapes around the feature points, texture mapping, lighting effects, and migration of animation data. For example, it can be achieved through methods such as calculation of normal vectors, calculation of tangent vectors, orthogonalization of the coordinate system, or standardization of the coordinate system. The present invention extracts semantically consistent feature points from the resting postures of the source character and the target character respectively, and calculates the tangent space coordinate system of the feature points, and does not limit the method of obtaining the first feature point grid corresponding to the source character and the second feature point grid corresponding to the target character.

[0053] Previous methods usually deduce correspondences from vertex coordinates, virtual feature points or through boundary meshes, however these methods are limited to template meshes that share the same topology, such as MANO or SMPL. There are also methods that suggest using nearest neighbor search on predefined feature vectors to determine vertex correspondences. However, this method often lacks precision and simplicity, resulting in inaccurate contact representation and a large optimization burden. The method provided by the present invention extracts semantically consistent feature points from the resting postures of the source character and the target character respectively, and calculates the tangent space coordinate system of the feature points to obtain a first feature point mesh corresponding to the source character and a second feature point mesh corresponding to the target character. This can promote dense geometric interaction, establish dense mesh correspondences between the source character and the target character, and is effective in different mesh topologies.

[0054] Step 102: Perform spatiotemporal representation of the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of multiple interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between the body geometry of the source character.

[0055] It should be noted that in order to effectively represent the interaction between the limbs and torso of the character, the present invention provides a dense mesh interaction (DMI) field. The first feature point grid is represented in time and space according to the source action sequence of the source character to obtain a source mesh interaction field. "Time and space representation" is used to describe the dynamic changes of objects or characters in time and space. Specifically, the time and space representation of the first feature point grid by the source action sequence of the source character means that the present invention needs to map the action sequence of the source character (that is, a series of postures or positions that change over time) to its first feature point grid and capture the changes of these feature points in three-dimensional space and time. Feature point extraction, feature encoding and other methods can be used. The present invention does not limit the method of performing time and space representation on the first feature point grid according to the source action sequence of the source character to obtain the source mesh interaction field.

[0056] Based on the feature point grid obtained in the above steps, the DMI field can comprehensively capture the contact and non-contact interactions of different body part geometries. Leveraging the DMI field, dense, geometric interaction-aware action redirection can be achieved, eliminating the need for a geometry correction stage.

[0057] Step 103: Perform action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, wherein the target action sequence is aligned with the geometric shape of the target character and the source grid interaction field.

[0058] It should be noted that, when performing action sequence prediction based on the source mesh interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain the target action sequence for the target character, appropriate algorithms (such as ICP or Procrustes analysis) may be used to align the first feature point grid of the source character with the second feature point grid of the target character to ensure that the aligned feature points are semantically consistent, i.e., they represent the same body part or function. The present invention does not limit the method for performing action sequence prediction based on the source mesh interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain the target action sequence for the target character.

[0059] The motion redirection method provided by the present invention first obtains a feature point grid to establish a dense grid correspondence between characters. This method is applicable to various grid topologies. Subsequently, a new spatiotemporal representation method is provided, called a dense mesh interaction (DMI) field. The DMI field is a collection of feature vectors of interacting feature points, which cleverly captures the contact and non-contact interactions between body geometries. During the redirection process, by aligning the DMI field, not only can the motion semantics be preserved, but self-penetration is also prevented and contact maintenance is ensured.

[0060] The motion retargeting method proposed in this paper focuses solely on dense geometric interactions for motion retargeting. Character animation videos rendered from skinned meshes rely on geometric interactions to shape user perception. In contrast, skeletal interactions represent only a simplified, sparse form of geometric interaction. Therefore, maintaining the correct interactions between the geometries of different body parts not only preserves motion semantics but also prevents mesh interpenetration and ensures contact preservation.

[0061] Figure 2 Schematic diagram comparing the method provided by the present invention with the existing method. Figure 2 As shown, unlike earlier redirection-correction methods, which suffer from internal inconsistencies and lead to problems such as interpenetration, jitter, and contact mismatch, this paper considers the importance of geometric interactions and proposes a motion redirection method based on dense geometric interaction awareness for skinned motion redirection. This motion redirection method utilizes DMI fields to accurately simulate the complex interactions between character meshes without relying on predefined vertex correspondences. This not only preserves motion semantics, but also prevents self-penetration and ensures contact preservation.

[0062] In some embodiments, the present invention uses semantically consistent sensors (SCS) to establish dense grid correspondences between characters. This method is applicable to various grid topologies.

[0063] Figure 3 FIG. 1 is a flow chart of the method for obtaining a feature point grid provided by the present invention. Figure 3 As shown, extracting semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character may include:

[0064] Step 201: Acquire resting postures of the source character and the target character, wherein the resting postures include multiple skeletons.

[0065] Here, it is assumed that the action sequence has T frames and the character has N skeletal joints. The action sequence is translated by the global root and local joint rotations The character's resting pose geometry G is composed of the resting pose mesh O and the resting pose joint positions express.

[0066] Step 202: For each bone of the source character and the target character, a target transformation is used to obtain semantic coordinates. The target transformation includes projecting a ray from the bone axis of each bone to the vertical plane corresponding to the bone axis. The origin parameter, direction parameter and bone index of the ray constitute the semantic coordinates of the feature point. The semantic coordinates of the feature point are used to describe the connection relationship between the feature point and the bone.

[0067] Step 203: Filter valid feature points based on whether the ray of the feature point in the semantic coordinates of the feature point intersects with the mesh connected to the skeleton, and obtain a first feature point mesh of the source character and a second feature point mesh of the target character, wherein if the ray of the feature point intersects with the mesh connected to the skeleton, the feature point is valid.

[0068] It should be noted that the present invention introduces semantically consistent feature points (SCS), which can work effectively in various mesh topologies while ensuring accurate semantic correspondence. The method of the present invention is inspired by the Medial Axis Inverse Transform (MAIT). The skeleton of each character is regarded as the approximate medial axis of its limbs and torso. For each bone, the present invention applies a MAIT-like transformation to generate a corresponding feature point grid. Starting from the bone axis, the present invention projects a ray on the perpendicular plane of the bone. The origin parameter l and direction parameter φ of the ray, combined with the bone index b, together constitute the semantic coordinates of the feature point. These semantic coordinates describe the connection relationship between the feature point and the bone. If the ray of the feature point intersects with the mesh connected to the bone, the feature point is considered valid; otherwise, it is considered invalid. In this way, the present invention establishes a dense geometric correspondence based on the sparse bone correspondence.

[0069] Figure 4 This is a schematic diagram of the process of deriving semantically consistent feature points provided by the present invention. Figure 4 Given a set of unified SCS semantic coordinates , the present invention can derive SCS features for each role .

[0070] in, Figure 4 The left picture shows the semantic coordinates of different roles. The method for deriving the feature point feature s. The lines in the circle represent the projected rays. Feature s contains the location of the feature point and its tangent space matrix. In the right figure: The DMI field (Dynamic Multi-Interaction Field) effectively captures contact and non-contact interactions. The circle and lines in the DMI field represent In the second example, the body landmarks (points in the image) lie on the tangent plane of the hand landmarks (points in the image), indicating a contact interaction.

[0071] The present invention uses semantically consistent feature points to establish semantic correspondences between character meshes of different topological structures. Specifically, the coordinates of a given feature point in the semantic coordinate system are , we can identify semantically consistent feature point locations on the meshes of different characters and obtain the feature vectors of these feature points. The detailed steps of this process are shown in the following algorithm.

[0072] Algorithm: Identify corresponding semantically consistent feature points from semantic coordinates.

[0073] Input: Grid , joint position , bone number , origin parameters , direction parameter ;

[0074] Output: Feature point feature vector ;

[0075] ;

[0076] ;

[0077] : Get the position of the parent node and child node;

[0078] : Calculate the ray origin, which is the l position on the line connecting the parent joint and the child joint;

[0079] ;

[0080] : Calculate the unit direction vector of the bone;

[0081] : Calculate the direction vector perpendicular to both the forward direction and the bone direction;

[0082] : Calculate the direction of the ray, which is determined by the parameter Decision is the forward direction and A linear combination of

[0083] : Get the mesh part associated with bone b;

[0084] : Construct a ray starting from the origin o and with direction n;

[0085] : Calculate the intersection of ray r and grid B;

[0086] if (that is, there is an intersection), then , calculate the tangent plane matrix of the intersection point xp on the grid B, or, s ← concat(p, t): concatenate the intersection point p and the tangent plane matrix t into the feature point feature s; otherwise (that is, there is no intersection), The feature point feature s is set to the zero vector. In some embodiments, the present invention provides a new spatiotemporal representation method called a dense mesh interaction (DMI) field. The DMI field is a collection of feature vectors of interacting SCS feature points, cleverly capturing the contact and non-contact interactions between body geometries.

[0087] Figure 5 FIG. 1 is a flow chart of the method for obtaining the source grid interaction field provided by the present invention. Figure 5 As shown, performing spatiotemporal representation on the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field may include:

[0088] Step 301: performing a forward dynamics operation on the first feature point grid to obtain feature point features, wherein the feature point features include the position of the feature point and a tangent matrix;

[0089] Step 302: According to the feature point features and the geometric interaction model, paired interaction features between feature points are obtained to obtain the source grid interaction field. The geometric interaction model is obtained based on feature modeling between paired interaction feature points.

[0090] It should be noted that, for a given action sequence m, the present invention first performs a forward kinematics operation on S to obtain the feature point features . Each Contains the positions and cut matrices of S feature points at frame t. The FK transformation of a single feature point is expressed as: in, is the global transformation matrix of bone n, derived from its local rotation matrix, Representing feature points The linear blend skinning (LBS) weights of a mesh are determined by interpolating the centroids of its neighboring mesh vertices.

[0091] Next, we model the geometric interactions as pairwise interaction features between feature points. Ideally, for each frame, we obtain a comprehensive DMI field, denoted as , which represents Paired vectors between pairs of feature points: Features between feature points:

[0092] ,in, is a feature point i The tangent matrix of Indicates that the observed feature points i In the tangent space, the target feature point j The relative position of the DMI field It consists of two parts: the relative position of feature point pairs, and the semantic coordinates of the observed and target feature points. The use of semantic coordinates rather than spatial coordinates is crucial because it avoids dependence on the actual feature point positions, making the DMI field suitable for motion redirection applications.

[0093] Furthermore, obtaining paired interaction features between feature points based on the feature point features and the geometric interaction model to obtain the source mesh interaction field may include: obtaining initial paired interaction features based on the feature point features and the geometric interaction model; performing feature screening on the initial paired interaction features using target screening conditions to obtain the source mesh interaction field, wherein the target screening conditions include limiting the interaction between feature points to feature points of preset body parts, and / or limiting the number of feature points of each body part to no more than a preset number.

[0094] It should be noted that shows quadratic growth relative to S because it contains S square feature point pairs, which makes it impractical to manage thousands of feature points. Two sparsification strategies are implemented. Initially, we restrict the interactions only to key body parts such as arm-torso, arm-head, arm-arm, and leg-leg, rather than interactions between all feature point pairs, thereby limiting the focus of the present invention to K observed feature points. Subsequently, for each observed feature point, we select L target feature points from each relevant body part, where L is a predetermined hyperparameter. Specifically, we empirically select the L / 2 nearest target feature points and the L / 2 farthest target feature points. We find that close feature point pairs are crucial to minimizing interpenetration and maintaining contact, while far feature point pairs depict the overall spatial relationship between body parts, such as Figure 4 These strategies lead to the final DMI field The selected feature point pairs are formed by Figure 3 The sparse DMI mask shown in express.

[0095] In some embodiments, during the redirection process, by aligning the DMI fields, not only can the motion semantics be preserved, but self-penetration is also prevented and contact preservation is ensured.

[0096] Figure 6 : is a flow chart of the method for predicting an action sequence provided by the present invention. Figure 6 As shown, performing action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain the target action sequence of the target character may include:

[0097] Step 401: extracting features of the source grid interaction field to obtain grid interaction features;

[0098] Step 402: extracting features from the source motion sequence to obtain source joint rotation features;

[0099] Step 403: performing geometric feature extraction on the first feature point grid and the second feature point grid respectively to obtain a first geometric feature vector and a second geometric feature vector;

[0100] Step 404: Input the first geometric feature vector and the second geometric feature vector into the encoder of the trained redirection network, and input the mesh interaction feature and the source joint rotation feature into the decoder of the trained redirection network to obtain the target action sequence.

[0101] It should be noted that the feature extraction of the source mesh interaction field can be performed to obtain the mesh interaction feature through the DMI encoder, the feature extraction of the source action sequence can be performed to obtain the source joint rotation feature through the action encoder, and the geometric feature extraction of the first feature point grid and the second feature point grid can be performed respectively to obtain the first geometric feature vector and the second geometric feature vector through the geometric encoder. The network structures of these encoders, namely the action encoder, DMI encoder and geometric encoder, are similar to the structure of PointNet. However, since all the data of the present invention are essentially in the canonical space, the present invention has eliminated T-Net from PointNet to reduce the network complexity. Before being input into the encoder, the feature point features pass through the feature point group embedding layer, which converts the bone index b into an 8-dimensional embedding vector. This embedding vector is updated during the training process. The geometric encoder consists of six PointNet layers, It is set to 256, and there is a different geometry encoder for the body, head, arms, and legs. The DMI encoder consists of a per-feature point encoder and a per-frame encoder, each of which is built from six PointNet layers, with each interaction pair having its own encoder. Specific interaction pairs include: [(left arm), (right arm, head, torso)], [(right arm), (left arm, head, torso)], [(left leg), (right leg, torso)], and [(right leg), (left leg, torso)]. The action encoder is a multi-layer perceptron (MLP). Both the transformer encoder and transformer decoder have eight layers, the number of heads is set to four, and the feedforward size is 256. An alignment mask is used between the transformer encoder and transformer decoder to ensure that each frame feature in the decoder only focuses on the corresponding DMI frame and initial token, so that the network's output action sequence is aligned with the input features.

[0102] In order to avoid the conflict between skeleton interaction and geometric correction, the motion redirection method proposed in this paper uses DMI field to directly model geometric interaction. Extract the DMI field ,field The interaction between various body parts in the source motion is encapsulated, including both contact and non-contact interactions. The DMI field consists of feature point pairs and feature vectors, which has the disordered characteristics of point clouds. Therefore, the present invention implements a PointNet-like network architecture for the DMI encoder of the present invention, which is divided into two components: a feature point encoder and a frame encoder. Given , each feature point encoder initially processes it into T * K separate point clouds, generating a representation for each observed feature point Model, where Represents the feature dimension. Subsequently, the frame-by-frame encoder generates a frame-by-frame representation by encoding these T point clouds Model.

[0103] Since the DMI field DA lacks geometric information about the characters, this paper introduces a geometric encoder Fg to extract geometric features from their SCS. For each feature point, this paper connects its rest-pose feature si with its semantic coordinates To form a feature vector. The resulting geometric features are represented as character A and character b's The semantic coordinates of the feature points act as an intermediary connecting the DMI field and the character geometry. The geometry encoder uses a point-net architecture to convert the geometric features C into a geometric latent code. .

[0104] Transformer-based redirection network processes input features, including source DMI features , source joint rotation Q A , source geometry is faint and target geometry is vague Specifically, the encoder processes and , while the decoder processes Q A and Potential and As the initial token in the sequence, both the encoder and decoder can operate on sequences of length T + 1. The final T frames of the output sequence are represented as .

[0105] The motion redirection method provided by the present invention directly models dense geometric interactions in motion, which not only preserves motion semantics but also prevents self-penetration and ensures contact preservation.

[0106] In some embodiments, when a trained redirection network is obtained for training, the initial redirection network may be trained using four loss functions.

[0107] In an embodiment of the present invention, before extracting semantically consistent feature points from the resting postures of the source character and the target character respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, the method may further include: training the initial redirection network according to a target loss function to obtain the trained redirection network, the target loss function including reconstruction loss, mesh interaction field consistency loss, adversarial loss and end effector loss, the reconstruction loss is used to minimize the motion change during the redirection process, the mesh interaction field consistency loss is used to maintain the geometric interaction between the source mesh interaction field and the target mesh interaction field, the adversarial loss is used to promote real action redirection, and the end effector loss is used to promote the consistent direction of the end effector in the redirection motion.

[0108] It should be noted that due to the lack of paired true value data, the present invention adopts an unsupervised method. The network of the present invention uses four loss functions for training: reconstruction loss, grid interaction field consistency loss, adversarial loss and end effector loss. The supervision signal comes from the source action sequence. The present invention combines the source DMI field DA with the target DMI field Align to maintain geometric interactions. First, the feature point forward kinematics is applied to , and then use the target sparse DMI mask Select feature point pairs to generate the target DMI field This mask It is through It is obtained by excluding the invalid feature points of the target character.

[0109] The grid interaction field consistency loss is quantified as and Cosine similarity loss between pairwise relative positions in :

[0110] ,

[0111] Among them, if the feature point pair (k, l) is and If all are valid, the value of c(k, l) is 1, otherwise it is 0.

[0112] The reconstruction loss serves as a regularization mechanism to minimize the motion change during the redirection process and is defined as follows:

[0113] .

[0114] To achieve realistic motion redirection, we use a discriminator, denoted as δ(·). Subsequently, the adversarial loss is defined as:

[0115] .

[0116] We observe that the global orientation of the end-effector has a significant impact on user experience. Therefore, we introduce an end-effector loss to promote consistency of the end-effector orientation during redirection motion.

[0117] .

[0118] where R(·) converts the local rotation of joint i along the kinematic chain into a global rotation, and X represents the set of end effectors. Our MeshRet network is trained as follows:

[0119]

[0120] Among them, the training details of the present invention use PyTorch to implement the network of the present invention, running on a machine equipped with NVIDIARTX A6000 GPU and AMD EPYC 9654 CPU. The data set is uniformly processed at a frame rate of 30 fps. During the training process, the present invention randomly clips a 30-frame sequence from the data set. The target character is set to be the same as the source character with a probability of 50% and different from the source character with a probability of 50%, and is randomly selected from the data set. On the system of the present invention, 36 epochs of training take approximately 40 hours. During the inference process, the action redirection method model of the present invention can achieve a performance of more than 30 fps.

[0121] In this paper, we propose SCS and a new DMI field to guide the training of motion redirection methods, effectively encapsulating contact and contactless interaction semantics.

[0122] The motion redirection method provided by the present invention includes multiple technical innovations. First, it requires the establishment of dense mesh correspondences between different characters. Inspired by the Medial Axis Inverse Transform (MAIT), the present invention develops a technique called Semantically Consistent Points (SCS), which can automatically derive dense mesh correspondences from sparse skeletal correspondences. This technique enables the present invention to sample a set of feature point clouds on the mesh to represent each character. Next, to account for dense mesh interactions between body parts while maintaining generality, the present invention adopts interacting mesh feature point pairs. These pairwise interactions are encoded in a new spatiotemporal representation called dense mesh interaction (DMI) field. The DMI field cleverly incorporates the semantic information of contact and non-contact interactions. Finally, the present invention learns a motion manifold that is consistent with the geometry of the target character and the DMI field of the source action.

[0123] The following describes an exemplary application of an embodiment of the present invention in a practical application scenario.

[0124] Figure 7 This is a schematic diagram of the overall flow of the action redirection method provided by the present invention. Figure 7 As shown in , the redirection process first extracts the DMI field using feature point forward kinematics (denoted as Fk) and pairwise interactive feature selection (denoted as Fc). This DMI field, combined with the geometric features derived from Fg, is fed into the encoder-decoder network. The network predicts a target action sequence that is consistent with the geometry of the target character and the original DMI field. Figure 7 As shown, the method includes the following steps 1 to 6:

[0125] Step 1: Extract semantically consistent feature points from the resting postures of the source and target characters, and calculate the tangent space coordinate system of the feature points, corresponding to Figure 7 SA and SB in;

[0126] Step 2: Input SA and SB into Fg to obtain the static geometric feature vector;

[0127] Step 3: Use the source action sequence QA to perform forward kinematics on SA, extract interactive features between feature points, and filter interactive features (corresponding to Fk and Fc) to obtain the source DMI field.

[0128] Step 4: Encode the source DMI field using the DMI Encoder and the source action QA using the Motion Encoder, input them into the Transformer network, and predict the target action sequence QB;

[0129] Step 5: Use the predicted QB to drive SB through forward kinematics operation, interactive feature extraction between feature points, and interactive feature screening (corresponding to Fk and Fc) to obtain the target DMI field;

[0130] Step 6: Calculate the loss and back-propagate the gradient to train the network.

[0131] Among them, steps 1-6 are the training process of the action redirection network, and steps 1-4 are the inference process of the action redirection method.

[0132] It should be noted that the task definition is first performed: given a source action sequence m A , and the geometry G of the source and target characters in their t poses A and G B The goal of this invention is to generate motion m for the target character B The process aims to preserve essential aspects of the source motion, including its semantics, contact preservation, and avoidance of interpenetration.

[0133] According to the definition of the task, the action redirection method model of the present invention initially derives semantically consistent feature points (SCS) , which provides the necessary dense geometric correspondence for the reorientation process, where S captures the feature point positions and feature point tangent space matrix, which helps to enhance the perception of geometric surfaces. Subsequently, the present invention performs feature point forward kinematics (FK) and pairwise interaction extraction to generate the source DMI field ,in Among them, K is the number of SCS in the DMI domain, L is the hyperparameter of feature selection, and P is the feature dimension of DMI. Finally, the transformer-based network ingests m A 、D A 、S A and S B , and predict the target action sequence m B , the sequence is consistent with the target character's geometry and the source DMI field. The entire pipeline is represented as follows:

[0134] .

[0135] Extensive experiments on the public Mixamo dataset and the newly collected ScanRet dataset demonstrate that our dense geometric interaction-aware action redirection method achieves state-of-the-art performance.

[0136] In order to more closely integrate the evaluation process of the present invention with real animation production, the present invention collected a field motion dataset called ScanRet, which is characterized by rich contact semantics and minimal mesh interpenetration. ScanRet consists of 100 human actors, ranging from fat to thin, each performing 83 action clips carefully reviewed by human animators. The action redirection method model is trained on the ScanRet dataset and the widely used Mixamo dataset. The present invention evaluates the method of the present invention on a variety of actions and a variety of target character arrays. Qualitative and quantitative analysis shows that the action redirection method model of the present invention significantly outperforms existing methods.

[0137] We use the Mixamo dataset and the newly curated ScanRet dataset to train and evaluate our method. We downloaded 3,675 action clips performed by 13 cartoon characters from the Mixamo dataset, while the ScanRet dataset consists of 8,298 clips performed by 100 human actors. Notably, the Mixamo dataset often has corrupted data due to interpenetration and contact mismatch. To overcome these issues, we created the ScanRet dataset, which provides detailed contact semantics and improved mesh interactions, with each clip carefully inspected by human animators. The training set includes 90% of the action clips from both datasets, including 9 characters from Mixamo and 90 characters from ScanRet. Our experiments tested the motion redirection capabilities between cartoon characters and real people, closely aligned with a typical redirection workflow. During the inference process, the present invention adopts four data segmentations based on feature and motion visibility: unseen features and unseen motion (UC+UM), unseen features and seen motion (UC+SM), unseen features and unseen motion (SC+UM), and seen features and seen motion (SC+SM). The present invention gives the average results of these segmentations.

[0138] Implementation details hyperparameters and L are empirically set to 1.0, 5.0, 1.0, 1.0 and 20 respectively. body -1}×{0,0.25, 0.5, 0.75}×{0,0.5π,π,1.5π} as the SCS semantic coordinate set, where N body = 18 is the number of human bones, and × represents is the Cartesian product. The present invention uses the Adam optimizer with a learning rate of 10-4 to optimize the network of the present invention. The training process requires 36 epochs.

[0139] Evaluation Metrics The effectiveness of the proposed method is evaluated through three metrics: joint accuracy, contact preservation, and geometric interpenetration. Joint accuracy is quantified by calculating the mean squared error (MSE) between the relocated joint positions and the real data provided by the animator in ScanRet. The analysis takes into account global and local joint positions and is normalized by the character height. Contact preservation is evaluated by measuring the contact error (Contact Error), which is defined as the mean squared distance between the feature points that were originally in contact in the source motion clip. Geometric interpenetration is determined by the ratio of penetrating limb vertices to the total limb vertices per frame.

[0140] We introduce a novel geometric interaction-aware action redirection framework, referred to as the MeshRet framework for action redirection. This framework explicitly models the dense geometric interactions between various body parts by first establishing dense mesh correspondences between characters using semantically consistent feature points. We then develop a unique spatiotemporal representation, referred to as the DMI field, that expertly captures both contact and non-contact interactions between body geometries. By aligning the DMI domains, our action redirection method achieves detailed contact preservation and seamless geometric interactions. Performance evaluation using the Mixamo dataset and our newly compiled ScanRet dataset confirms that our action redirection method provides state-of-the-art results.

[0141] The main limitation of the motion redirection method is its reliance on input with clean contact; motion clips exhibiting severe interpenetration produce poor results. Therefore, it cannot effectively handle noisy input. Future efforts will focus on enhancing its robustness to noisy data.

[0142] Based on the foregoing embodiments, an embodiment of the present invention provides an action redirection device for dense geometric interaction perception. The modules included in the device and the units included in each module can be implemented by a processor; of course, they can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0143] The following describes the motion redirection device for dense geometric interaction perception provided by the present invention. The motion redirection device for dense geometric interaction perception described below and the motion redirection method for dense geometric interaction perception described above can refer to each other.

[0144] Figure 8 Schematic diagram of the structure of the dense geometric interaction perception action redirection device provided by the present invention. Figure 8 As shown, the apparatus 500 includes a grid acquisition module 501, an interaction acquisition module 502, and an action prediction module 503, wherein:

[0145] A grid acquisition module 501 is configured to extract semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculate the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid each include positions and feature vectors of multiple feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship;

[0146] an interaction acquisition module 502 for performing spatiotemporal representation of the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of a plurality of interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between body geometric shapes of the source character;

[0147] The action prediction module 503 is used to predict the action sequence based on the source grid interaction field, the source action sequence, the first feature point grid and the second feature point grid to obtain the target action sequence of the target character, and the target action sequence is aligned with the geometric shape of the target character and the source grid interaction field.

[0148] In some embodiments, the grid acquisition module 501 includes a posture acquisition unit, a coordinate acquisition unit and a grid acquisition unit, wherein:

[0149] The posture acquisition unit is used to acquire the resting postures of the source character and the target character, wherein the resting postures include multiple skeletons;

[0150] The coordinate acquisition unit is configured to acquire semantic coordinates for each bone of the source character and the target character using a target transformation, wherein the target transformation includes projecting a ray from the bone axis of each bone to a vertical plane corresponding to the bone axis, wherein the origin parameter, direction parameter, and bone index of the ray constitute the semantic coordinates of the feature point, and the semantic coordinates of the feature point are used to describe the connection relationship between the feature point and the bone;

[0151] The grid acquisition unit is used to screen valid feature points based on whether the ray of the feature point in the semantic coordinates of the feature point intersects with the grid connected to the skeleton, to obtain the first feature point grid of the source character and the second feature point grid of the target character, wherein if the ray of the feature point intersects with the grid connected to the skeleton, the feature point is valid.

[0152] In some embodiments, the interaction acquisition module 502 includes a feature operation unit and a feature modeling unit, wherein:

[0153] The feature operation unit is configured to perform a forward dynamics operation on the first feature point grid to obtain feature point features, wherein the feature point features include a position of the feature point and a tangent matrix;

[0154] The feature modeling unit is used to obtain paired interaction features between feature points based on the feature point features and the geometric interaction model to obtain the source grid interaction field. The geometric interaction model is obtained based on feature modeling between paired interaction feature points.

[0155] In some embodiments, the feature modeling unit is further specifically used to: obtain initial paired interaction features based on the feature point features and the geometric interaction model; perform feature screening on the initial paired interaction features using target screening conditions to obtain the source grid interaction field, and the target screening conditions include limiting the interaction between feature points to feature points of preset body parts, and / or limiting the number of feature points of each body part to no more than a preset number.

[0156] In some embodiments, the action prediction module 503 is further specifically used to: perform feature extraction on the source mesh interaction field to obtain mesh interaction features; perform feature extraction on the source action sequence to obtain source joint rotation features; perform geometric feature extraction on the first feature point mesh and the second feature point mesh respectively to obtain a first geometric feature vector and a second geometric feature vector; input the first geometric feature vector and the second geometric feature vector into the encoder of the trained redirection network, and input the mesh interaction features and the source joint rotation features into the decoder of the trained redirection network to obtain the target action sequence.

[0157] In some embodiments, the device also includes a model training module, which is used to: train the initial redirection network according to the target loss function to obtain the trained redirection network, the target loss function includes reconstruction loss, grid interaction field consistency loss, adversarial loss and end effector loss, the reconstruction loss is used to minimize the motion change during the redirection process, the grid interaction field consistency loss is used to maintain the geometric interaction between the source grid interaction field and the target grid interaction field, the adversarial loss is used to promote real action redirection, and the end effector loss is used to promote the consistent direction of the end effector in the redirection motion.

[0158] In an embodiment of the present invention, geometric interaction-aware motion redirection between different mesh topologies can be promoted in one process, which not only preserves motion semantics but also prevents self-penetration and ensures contact preservation.

[0159] Figure 9 Schematic diagram of the physical structure of the electronic device provided by the present invention. Figure 9As shown, the electronic device 600 may include: a processor 610 , a communications interface 620 , a memory 630 and a communication bus 640 , wherein the processor 610 , the communications interface 620 , and the memory 630 communicate with each other via the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute the action redirection method based on dense geometric interaction perception, which includes: extracting semantically consistent feature points from the resting postures of the source character and the target character respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, the first feature point grid and the second feature point grid both including the positions and feature vectors of multiple feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship; performing spatiotemporal representation of the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, the source grid interaction field including a set of feature vectors of multiple interacting feature points, the source grid interaction field being used to describe contact and non-contact interactions between the body geometry of the source character; performing action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, the target action sequence being aligned with the geometry of the target character and the source grid interaction field.

[0160] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0161] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the action redirection method based on dense geometric interaction perception provided by the above methods. The method includes: extracting semantically consistent feature points from the resting postures of the source character and the target character, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid both include the positions and feature vectors of multiple feature points. The first feature point grid and the second feature point grid have a dense grid correspondence; the first feature point grid is represented in time and space according to the source action sequence of the source character to obtain a source grid interaction field, the source grid interaction field includes a set of feature vectors of multiple interacting feature points, and the source grid interaction field is used to describe the contact and non-contact interactions between the body geometry of the source character; action sequence prediction is performed based on the source grid interaction field, the source action sequence, the first feature point grid and the second feature point grid to obtain the target action sequence of the target character, and the target action sequence is aligned with the geometry of the target character and the source grid interaction field.

[0162] The computer program product includes one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially perform the processes or functions described in accordance with the embodiments of the present invention. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium capable of computer storage or a data storage device such as a server or data center that integrates one or more available media. The available medium may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0163] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the action redirection method based on dense geometric interaction perception provided by the above-mentioned methods, the method comprising: extracting semantically consistent feature points from the resting postures of a source character and a target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, the first feature point grid and the second feature point grid both including the positions and feature vectors of a plurality of feature points, and the first feature point grid and the second feature point grid having a dense grid correspondence relationship; performing spatiotemporal representation of the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, the source grid interaction field including a set of feature vectors of a plurality of interacting feature points, the source grid interaction field being used to describe contact and non-contact interactions between the body geometric shapes of the source character; performing action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, the target action sequence being aligned with the geometric shape of the target character and the source grid interaction field.

[0164] The computer-readable storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0165] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0166] Program code embodied on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the foregoing.

[0167] Computer program code for performing the operations of this specification may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0169] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An action redirection method based on dense geometric interaction perception, characterized in that: include: Extracting semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid each include positions and feature vectors of a plurality of feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship; Performing a spatiotemporal representation on the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of a plurality of interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between body geometric shapes of the source character; An action sequence is predicted based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, wherein the target action sequence is aligned with a geometric shape of the target character and the source grid interaction field.

2. The action redirection method based on dense geometric interaction perception according to claim 1 is characterized in that: Extracting semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, including: Acquire a resting posture of the source character and the target character, wherein the resting posture includes a plurality of skeletons; For each bone of the source character and the target character, a target transformation is used to obtain semantic coordinates. The target transformation includes projecting a ray from the bone axis of each bone to the vertical plane corresponding to the bone axis. The origin parameter, direction parameter, and bone index of the ray constitute the semantic coordinates of the feature point. The semantic coordinates of the feature point are used to describe the connection relationship between the feature point and the bone. Valid feature points are screened based on whether the ray of the feature point in the semantic coordinates of the feature point intersects with the mesh connected to the skeleton, to obtain a first feature point mesh of the source character and a second feature point mesh of the target character, wherein if the ray of the feature point intersects with the mesh connected to the skeleton, the feature point is valid.

3. The action redirection method based on dense geometric interaction perception according to claim 1, characterized in that: The performing spatiotemporal representation on the first feature point grid according to the source action sequence of the source character to obtain a source grid interaction field includes: Performing a forward dynamics operation on the first feature point grid to obtain feature point features, the feature point features including a position of the feature point and a tangent matrix; According to the feature point features and the geometric interaction model, paired interaction features between feature points are obtained to obtain the source grid interaction field. The geometric interaction model is obtained based on feature modeling between paired interaction feature points.

4. The action redirection method based on dense geometric interaction perception according to claim 3 is characterized in that: The step of acquiring paired interaction features between feature points based on the feature point features and the geometric interaction model to obtain the source grid interaction field includes: Obtaining initial pairwise interaction features based on the feature point features and the geometric interaction model; The initial paired interaction features are subjected to feature screening using target screening conditions to obtain the source grid interaction field, wherein the target screening conditions include limiting the interaction between feature points to feature points of preset body parts, and / or limiting the number of feature points of each body part to no more than a preset number.

5. The action redirection method based on dense geometric interaction perception according to claim 1, characterized in that: The performing action sequence prediction based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain the target action sequence of the target character includes: Extracting features of the source grid interaction field to obtain grid interaction features; Extracting features from the source motion sequence to obtain source joint rotation features; Performing geometric feature extraction on the first feature point grid and the second feature point grid respectively to obtain a first geometric feature vector and a second geometric feature vector; The first geometric feature vector and the second geometric feature vector are input into an encoder of a trained redirection network, and the mesh interaction feature and the source joint rotation feature are input into a decoder of the trained redirection network to obtain the target action sequence.

6. The action redirection method based on dense geometric interaction perception according to claim 5, characterized in that: Before extracting semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculating the tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, the method further includes: The initial redirection network is trained according to the target loss function to obtain the trained redirection network, wherein the target loss function includes reconstruction loss, mesh interaction field consistency loss, adversarial loss and end effector loss, wherein the reconstruction loss is used to minimize the motion change during the redirection process, the mesh interaction field consistency loss is used to maintain the geometric interaction between the source mesh interaction field and the target mesh interaction field, the adversarial loss is used to promote real action redirection, and the end effector loss is used to promote the consistent direction of the end effector in the redirection motion.

7. An action redirection device based on dense geometric interaction perception, characterized in that: include: a grid acquisition module, configured to extract semantically consistent feature points from the resting postures of the source character and the target character, respectively, and calculate a tangent space coordinate system of the feature points to obtain a first feature point grid corresponding to the source character and a second feature point grid corresponding to the target character, wherein the first feature point grid and the second feature point grid each include positions and feature vectors of a plurality of feature points, and the first feature point grid and the second feature point grid have a dense grid correspondence relationship; an interaction acquisition module, configured to perform spatiotemporal representation of the first feature point grid according to a source action sequence of the source character to obtain a source grid interaction field, wherein the source grid interaction field includes a set of feature vectors of a plurality of interacting feature points, and the source grid interaction field is used to describe contact and non-contact interactions between body geometric shapes of the source character; An action prediction module is used to predict an action sequence based on the source grid interaction field, the source action sequence, the first feature point grid, and the second feature point grid to obtain a target action sequence of the target character, wherein the target action sequence is aligned with the geometric shape of the target character and the source grid interaction field.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the action redirection method based on dense geometric interaction perception is implemented as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the action redirection method based on dense geometric interaction perception is implemented as claimed in any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the action redirection method based on dense geometric interaction perception is implemented as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Virtual character processing method and device and storage medium

    CN113920229A

  • Multi-style lip shape synthesis method, device and equipment and storage medium

    CN114022597A