Grid sequence driven two-dimensional face animation generation method and device, equipment and medium
By using Gaussian splattering representation and mesh sequence driving methods in two-dimensional face animation generation, we bind 3D mesh topology and 2D face identity information to optimize Gaussian avatars, solving the problems of high animation generation cost and poor time stability in the existing technology, and achieving efficient and flexible animation generation and time stability.
Patent Information
- Application Number
- CN202411811049.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The prior art has problems such as strong data dependence, poor time stability, poor migration, high cost and resource requirements when generating two-dimensional face animations, which are difficult to meet the market's demand for personalized and dynamic content.
The grid sequence driving method based on Gaussian splattering characterization is adopted to bind the topology of the 3D mesh and the identity information of the 2D face to the Gaussian model, optimize the Gaussian avatar through the deformation field, and generate animated video sequences.
Reduces rendering cost and complexity, improves the flexibility and efficiency of animation generation, ensures the time stability of long video sequences, and avoids inter-frame jitter and inconsistency problems.
Smart Images

Figure CN119941944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a method, device, equipment and medium for generating two-dimensional facial animation driven by a grid sequence. Background Art
[0002] In the field of modern digital content creation, especially in the film, television, game and virtual reality industries, the production and rendering of 3D assets has always been a time-consuming and expensive process. Traditional content creation processes usually require the creation of detailed 3D assets for each static image, which not only increases production costs, but also prolongs the production cycle, making it difficult to meet the market's urgent demand for personalized and dynamic content.
[0003] The current technical solutions include: (1) Reenactment network in 2D scene - static image animation generation. This solution has the following defects: a) Strong data dependence: Since the reenactment network has never seen data in the form of 3D gray model mesh rendering, the driving effect under the 2D reenactment network is not good. b) Poor temporal stability: When generating long video sequences, temporal inconsistency and inter-frame jitter are prone to occur, affecting user experience. c) Poor migration: The model is usually trained for a specific data set and is difficult to migrate to static images of different identities, limiting the breadth of application. (2) Dynamic portrait generation based on 3D models. This solution has the following defects: a) High cost and resource requirements: Building a high-precision 3D face model requires expensive scanning equipment and high-performance computing resources, which increases the production cost. b) Complex rendering process: Relying on a complex 3D rendering process, resulting in long rendering time and difficulty in achieving real-time generation. c) High manual adjustment requirements: A large amount of manual adjustment and professional knowledge are required, which increases the difficulty and time cost of content creation. d) Identity restrictions: Each static image needs to be specially built with a corresponding 3D model, which lacks flexibility and is difficult to apply on a large scale. Summary of the invention
[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a method, device, equipment and medium for generating two-dimensional facial animation driven by a grid sequence based on Gaussian splash representation.
[0005] The first technical solution adopted by the present invention is:
[0006] A method for generating a two-dimensional face animation driven by a grid sequence comprises the following steps:
[0007] Get the 2D face image that needs to be driven;
[0008] The initial frame of the mesh sequence motion and the 2D face image are bound using Gaussian splash expression, and the topology of the 3D mesh and the identity information of the 2D face are bound to the Gaussian model to generate the initial Gaussian avatar;
[0009] Combining deformation fields to characterize the motion of mesh sequences on Gaussian representations, Gaussian avatars are optimized and animated video sequences are generated through rendering;
[0010] Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
[0011] Furthermore, a Gaussian splash representation method is used as an intermediary between 2D faces and 3D meshes of different identities to achieve the transformation from mesh sequence M to animation sequence video Transition; where T is the number of frames of the animation.
[0012] Furthermore, the initial frame of the mesh sequence motion and the 2D face image are bound using Gaussian splash expression, the topology of the 3D mesh and the identity information of the 2D face are bound to the Gaussian model, and the initial Gaussian avatar is generated, including:
[0013] Using 3D Gaussian splashing as an intermediary, the connection between 3D meshes and 2D face images under different identities is established;
[0014] For each unit triangle in the mesh, a 3D Gaussian splash unit is bound to it, and the initial center position μ, rotation matrix r, and scaling factor s of the bound Gaussian are given according to the properties of the unit triangle.
[0015] Furthermore, for each unit triangle in the mesh, a 3D Gaussian splash unit is bound to it, and the initial center μ, rotation matrix r, and scaling factor s of the bound Gaussian are given according to the properties of the unit triangle, including:
[0016] For the center position μ of the bound Gaussian, the expression based on the barycentric coordinates is:
[0017] For any point V of the spatial triangle ΔV1V2V3, there must be unique coefficients λ1,λ2 such that:
[0018]
[0019] in, They are the three vertices of the triangle;
[0020] For each triangle, the normal vector n is obtained by computing the difference between the two edge vectors:
[0021]
[0022] For the direction vector d, the edge composed of vertex 1 and vertex 2 is selected to calculate the direction vector d, that is:
[0023]
[0024] In order to describe the direction of the triangle in the global space, the direction vector d of an edge, the normal vector n of the triangle and their cross product d×n are used as column vectors to form the rotation matrix R:
[0025] R = [d,n,d×n]]
[0026] The scalar k describing the scaling of the triangle is calculated by taking the length of the shortest side l and its perpendicular h:
[0027]
[0028] For the paired 3D Gaussian, define its position μ, rotation r, and anisotropic scale s in local space; initialize the position to the local origin, the rotation r to the unit rotation matrix, and the scale s to the unit vector; at rendering time, transform these attributes to the global space as follows:
[0029] r'=Rr
[0030] μ'=kRμ+T
[0031] s'=ks
[0032] Where r' is the rotation of the unit Gaussian splash in the global space, μ' is the center position of the unit Gaussian splash in the global space, s' is the scale of the unit Gaussian splash in the global space, and T is the center position of the grid where the Gaussian splash is located in the global space.
[0033] Furthermore, generating the animation video sequence by rendering includes:
[0034] For any given frame i, the deformation of each vertex is defined as:
[0035]
[0036] In the formula, is the position of vertex j in the deformed mesh in the i-th frame, v j,can is the position of the same vertex in the canonical grid;
[0037] For every Gaussian x in the canonical space i,can , the MLP network is used to calculate its characterization of its deformation, and then the Gaussian is transformed from the standard space to the deformation space:
[0038] x i,def =x i,can +Def(v j,can ,δv j )
[0039] By directly using the deformed vertices of the mesh and the vertex positions in the canonical space, complex deformations in dynamic scenes can be captured more accurately to achieve high-quality dynamic scene rendering.
[0040] Furthermore, the deformation field is combined to characterize the movement of the grid sequence on the Gaussian expression to optimize the Gaussian avatar, including:
[0041] The 2D reference face image is input into a pre-trained 2D reenactment network, and the grid state driven by the current frame is used as a conditional drive. The 2D reenactment network can give a reference prior of the 2D reference face image after being driven in the current frame based on the current driving information. These reference priors provide guidance for maintaining the consistency between the original image features and the reenactment motion.
[0042] These reference priors are combined with gradient-based optimization techniques to adjust the Gaussian deformation field to ensure accurate alignment of unique facial features and poses between motion and static images.
[0043] Furthermore, during the training optimization process, a combined L1 term and a D-SSIM term are used to supervise the rendered images:
[0044] L rgb =(1-λ)L1+λL D-SSIM
[0045]
[0046] In the formula, and Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo The average value of and Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo The variance of Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo covariance, C1 and C2 are small constants used to stabilize the denominator; L1 is the loss calculated by the absolute difference between each pixel of the rendered image and the image generated as a condition by the pre-trained 2D reenactment network; L D-SSIM is the structural similarity loss function, which represents the difference between images; λ is L D-SSIM The weight of is 0.2. The smaller the value of D-SSIM is, the more similar the two images are.
[0047] An L1 term is also used to supervise the transparency produced by the rendering and the loss between the image masks, hoping to ensure pixel alignment:
[0048]
[0049] In the formula, The mask image of the face part generated as a condition for the pre-trained 2D reenactment network, I alpha A mask image for the face part of the rendered image;
[0050] The final loss function is:
[0051] L=L rgb +λ mask L mask
[0052] In the formula, λ mask For L mask The weight of the total loss function, L mask The difference between the absolute values of the pixels of the two mask images is used as the loss value.
[0053] The second technical solution adopted by the present invention is:
[0054] A grid sequence driven two-dimensional face animation generation device, comprising:
[0055] 2D image acquisition module, used to acquire the 2D face image that needs to be driven;
[0056] The Gaussian avatar generation module is used to bind the initial frame of the grid sequence motion and the 2D face image using Gaussian splash expression, bind the topology of the 3D grid and the identity information of the 2D face to the Gaussian model, and generate the initial Gaussian avatar;
[0057] The video sequence generation module is used to combine the deformation field to characterize the motion of the grid sequence on the Gaussian expression, optimize the Gaussian head image, and generate an animated video sequence through rendering;
[0058] Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
[0059] The third technical solution adopted by the present invention is:
[0060] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement a grid sequence driven two-dimensional facial animation generation method as described above.
[0061] The fourth technical solution adopted by the present invention is:
[0062] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a grid sequence driven two-dimensional facial animation generation method as described above.
[0063] The fifth technical solution adopted by the present invention is:
[0064] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0065] The beneficial effects of the present invention are as follows: by introducing an intermediate representation method based on Gaussian Splatting and a Gaussian deformation field optimization method combined with two-dimensional prior and gradient guidance, the present invention can effectively reduce rendering cost and complexity, improve the flexibility and efficiency of animation generation, and improve the temporal stability of long video sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0067] Figure 1 It is a schematic flow chart of a method for generating a two-dimensional facial animation driven by a grid sequence in an embodiment of the present invention;
[0068] Figure 2 is a 3D Gaussian binding grid topology flow chart in an embodiment of the present invention;
[0069] Figure 3is a schematic diagram of 3D Gaussian modeling of a corresponding character ID constructed according to a given 2D identity image and 3D mesh topology in an initial state in an embodiment of the present invention;
[0070] Figure 4 is a schematic diagram of visualization results of a 2D human portrait model driven by a grid using a 3D Gaussian as an intermediary and a reenactment network as a priori in an embodiment of the present invention;
[0071] Figure 5 It is a schematic diagram of an experiment comparison between the method provided by an embodiment of the present invention and the current 2D replay method. DETAILED DESCRIPTION
[0072] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0073] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0074] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0075] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0076] Technical explanation:
[0077] (1) Mesh-Driven Animation: A method that uses a mesh sequence to drive animation generation, achieving dynamic performance of objects through continuous mesh deformation.
[0078] (2) Gaussian Splatting: A rendering method that represents a 3D scene as a point cloud defined by a Gaussian distribution. Efficient image generation and rendering are achieved by projecting these Gaussian points onto a 2D image plane.
[0079] (3) Deformation Field: Deformation field refers to a mathematical model used in Gaussian splattering technology to describe and control the deformation of the surface or volume of a three-dimensional object in space. It represents the local displacement and shape change of an object during the deformation process by representing the rotation offset, center point offset and other parameters of each point or voxel in the 3D Gaussian, thereby achieving smooth and continuous deformation of the object shape.
[0080] (4) 2D reenactment: 2D reenactment is a computer vision task that aims to generate a reenactment image that combines information from two input images: it preserves the pose and expression of the driving image while retaining the identity features of the source image. The core task is to transfer facial movements and expressions from one image or video to another.
[0081] (5) Face Reenactment
[0082] The reenacted image is generated by transferring the facial movements and expressions in a given driven image to a given source image. This scheme retains the posture information in the driven image while retaining the identity information in the source image. The specific steps include: first, rendering the 3D mesh into a 2D animation, extracting the facial posture information from the driven image, and extracting the identity features from the source image; then, using a deep learning model to map the expression features of the driven image to the source image to ensure that the identity features of the source face are preserved; finally, the generated reenacted image shows what the source face looks like under the expression and posture of the driven image, and necessary post-processing is performed to improve the image quality. This scheme achieves high-fidelity facial reenactment and is suitable for a variety of application scenarios.
[0083] (6) Dynamic portrait generation based on 3D models
[0084] Retargeting is the process of transferring the motion of a 3D mesh to another mesh and finally rendering it as a 2D video. First, motion data, including joint positions and animation features, is extracted from the source mesh, and then the motion retargeting algorithm is used to map this data to the target mesh to ensure the naturalness and smoothness of the motion. Next, an animation sequence of the target mesh is generated based on the retargeted motion, and the rendering engine is used to set the camera view and lighting conditions to generate high-quality 2D video output. Finally, the rendering result is post-processed to improve the visual effect.
[0085] (7) 3D Gaussian Splatting
[0086] The 3D Gaussian splash method provides a solution for reconstructing static scenes based on images and camera parameters. This method represents the scene as a set of 3D anisotropic Gaussian splash points, where each splash point is defined by a covariance matrix centered at the point (mean) μ:
[0087]
[0088] It is worth noting that the covariance matrix only has physical meaning when it is semi-positive, which is not guaranteed when using gradient descent for optimization. Therefore, Kerbl et al. define a parameter ellipse with a scaling matrix S and a rotation matrix R, and then construct the covariance matrix in the following way:
[0089] Σ=RSS T R T .
[0090] Thus, an ellipse can be constructed by the position vector Scaling Vectors And quaternions In this paper, we use the rotation matrix To represent the corresponding rotation.
[0091] During rendering, the color C of a pixel is calculated by blending all 3D Gaussian splatter points that overlap the pixel:
[0092]
[0093] where c i represents the color of each point, which is modeled by a third-order spherical harmonic function. The blending weight α′ is determined by the 2D projection of the 3D Gaussian multiplied by the transparency α of each point. To maintain the visible order, the Gaussian splash points are sorted by depth before blending.
[0094] In response to the existing technical problems, the present invention proposes an innovative method based on Gaussian splash representation, focusing on the field of mesh motion sequence driven 2D facial animation. This method simplifies the mesh model and rendering process, and enhances the portability of motion data, so that art and content creation related work no longer need to create detailed 3D assets for each static image. Motion data can be flexibly applied to any head avatar, achieving high efficiency and flexibility in animation generation.
[0095] Specifically, this method acts as an intermediary between a 2D portrait and a 3D mesh, and combines a 2D prior gradient-guided optimization technique to achieve efficient and flexible animation generation of arbitrary static portraits. By significantly simplifying the traditional 3D driver and rendering process, the present invention reduces the resource requirements for 3D face asset production and simplifies the rendering process from 3D to 2D. At the same time, it provides excellent temporal stability in the generation of long video sequences.
[0096] Example 1
[0097] like Figure 1 As shown, this embodiment provides a method for generating a two-dimensional face animation driven by a grid sequence based on Gaussian splash characterization, comprising the following steps:
[0098] S1, obtain the 2D face image to be driven;
[0099] S2, the initial frame of the mesh sequence motion and the 2D face image are bound using Gaussian splash expression, the topology of the 3D mesh and the identity information of the 2D face are bound to the Gaussian model, and an initial Gaussian avatar is generated;
[0100] S3, combining the deformation field to characterize the movement of the mesh sequence on the Gaussian expression, optimizing the Gaussian avatar, and generating an animated video sequence through rendering;
[0101] Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
[0102] The process of the method in this embodiment is as follows Figure 1 Specifically, we use mesh sequence driving technology to realize the animation generation of human portraits. First, given a 2D human portrait image to be driven, we can use a face mesh sequence of arbitrary topology to make the 2D portrait move according to the specified mesh sequence.
[0103] Specifically, first, the initial frame of the mesh sequence motion and the 2D initial image are bound using Gaussian splash expression, and the topology of the 3D mesh and the identity information of the 2D portrait are bound to the Gaussian model to generate the initial Gaussian avatar. Then, the deformation field is combined to characterize the motion of the mesh sequence on the Gaussian expression, the Gaussian splash avatar is optimized, and an animated video sequence is generated by rendering.
[0104] At the same time, the 2D reenactment pre-trained network processes the input reference image and the mesh current rendering frame as conditions, generates gradients and optimizes the avatar's posture and features, making the generated animation highly consistent with the features of the input image. The entire process ensures that the generated 3D dynamic avatar has high visual fidelity and accurate facial movements.
[0105] in, Figure 1 , Figure 4 and Figure 5 The facial images in the video are not generated through AI processing, so there is no infringement of other people’s portrait rights.
[0106] The above method is supplemented below with reference to the accompanying drawings and specific implementation methods.
[0107] (1) Network framework
[0108] The goal of this invention is to generate a corresponding animation video sequence by driving the movement of a static 2D portrait image I and a 3D head mesh M with arbitrary topology. Where T is the number of frames of the animation. To this end, the embodiment of the present invention proposes a method based on Gaussian Splatting representation, which serves as an intermediate expression between 2D portraits and 3D grids of different identities, and can realize the conversion from the grid sequence M to the animation sequence video. transition.
[0109] (2) Gaussian Binding
[0110] like Figure 2 As shown, the core of the method of the present invention is to use 3D Gaussian splash as an intermediary to establish a connection between the 3D mesh and the 2D character portrait under different identities. For each unit triangle in the mesh, we bind a 3D Gaussian splash unit to it, and give the initial center μ, rotation matrix r, and scaling factor s of the bound Gaussian according to the properties of the unit triangle.
[0111] Specifically, for the center position μ of the bound Gaussian, we express it based on the barycentric coordinates, namely:
[0112] For any point V of the spatial triangle ΔV1V2V3, there must be a unique λ1,λ2 such that:
[0113]
[0114] in, They are the three vertices of the triangle.
[0115] For each triangle, we obtain the normal vector n by computing the difference between the two edge vectors:
[0116]
[0117] Then, we normalize the normal vector:
[0118]
[0119] For the direction vector d, we use the edge composed of vertex 1 and vertex 2 to calculate the direction vector d, that is:
[0120]
[0121] To describe the orientation of the triangle in global space, we form a rotation matrix R using the direction vector d of an edge, the normal vector n of the triangle, and their cross product d×n as column vectors.
[0122] R = [d,n,d×n]]
[0123] We also compute a scalar k describing the scaling of the triangle by taking the length of the shortest side l and its perpendicular h.
[0124]
[0125] For the paired 3D Gaussian, we define its position μ, rotation r, and anisotropic scale s in local space. We initialize the position μ to the local origin, the rotation r to the unit rotation matrix, and the scale s to the unit vector. At rendering time, we transform these attributes to global space as follows:
[0126] r'=Rr
[0127] μ'=kRμ+T
[0128] s'=ks
[0129] The initial state of the constructed 3D mesh topology is similar to the Gaussian topology of the 2D character. Figure 3 shown.
[0130] (3) Learnable prior deformable Gaussian field
[0131] The standard 3D Gaussian splatting technique does not support dynamic scene rendering because the initial point cloud is assumed to be static. Our method needs to handle dynamic scenes, where the subject constantly presents different facial expressions and head poses, and we also need to model the inconsistencies between the mesh and the 2D image identity. We model this dynamic behavior as a deformation from a predefined canonical space to a deformed space. In the canonical space, the 3D Gaussian is located at position x i,can , in the deformation space, its position is x i,def .
[0132] For any given frame i, the deformation of each vertex is defined as:
[0133]
[0134] in, is the position of vertex j in the deformed mesh in the i-th frame, v j,can are the positions of the same vertices in the canonical mesh.
[0135] Next, for every Gaussian x in the canonical space i,can , we compute it through the MLP network to represent its deformation, and then we transform the Gaussian from the standard space to the deformation space:
[0136] x i,def =x i,can +Def(v j,can ,δv j )
[0137] By directly using the deformed vertices of the mesh and the vertex positions in the canonical space, our model is able to more accurately capture the complex deformations in dynamic scenes and achieve high-quality dynamic scene rendering.
[0138] (4) Optimization using 2D priors and gradient guidance
[0139] To further optimize the actuation effect, we introduce 2D priors derived from static images. We feed the 2D reference portrait into a pre-trained 2D reenactment network, and use the mesh state driven by the current frame as a conditional drive. The 2D reenactment network can give reference priors for the 2D reference image after being driven in the current frame based on the current actuation information. These priors provide guidance for maintaining consistency between the original image features and the reenacted motion, especially in areas prone to distortion, such as the eyes and mouth. We combine these priors with gradient-based optimization techniques to adjust the Gaussian deformation field to ensure accurate alignment of the unique facial features and poses of the motion and static images.
[0140] (5) Optimization and Regularization
[0141] a) Color loss
[0142] We use a combination of the L1 term and the D-SSIM term to supervise the rendered images:
[0143] L rgb =(1-λ)L1+λL D-SSIM
[0144] In this embodiment, λ=0.2.
[0145] b) Mask loss
[0146] At the same time, we also use an L1 term to supervise the transparency generated by the rendering and the loss between the image masks, hoping to ensure pixel alignment:
[0147]
[0148] c) Total loss function
[0149] Our final loss function is:
[0150] L=L rgb +λ mask L mask
[0151] Among them, λ mask = 1. Note that we only apply L to the visible Gaussian. mask , so only when there is a color loss L rgb Regularize the points.
[0152] (6) Experimental details
[0153] The Adam optimizer is used for parameter optimization. We set the position learning rate of the 3D Gaussian to 5e -3 , the scaled learning rate is 1.7e -2 , the learning rates of the remaining parameters are kept the same as those of the 3D Gaussian splash. We train for 60,000 iterations and exponentially decay the Gaussian position learning rate until it reaches 0.01 times the initial value at the final iteration.
[0154] (7) Advantages and beneficial effects
[0155] The present invention brings the following beneficial effects by introducing an intermediate representation method based on Gaussian splatting and a Gaussian deformation field optimization method combined with two-dimensional prior and gradient guidance:
[0156] a) Reduce rendering cost and complexity:
[0157] 1. No need for fine 3D models: Traditional methods usually rely on high-precision 3D models and complex rendering processes, resulting in large resource consumption and high costs. The present invention uses coarse mesh models and motion data with unrestricted identity to achieve high-quality animation generation, significantly reducing the reliance on fine 3D assets.
[0158] 2. Simplify the rendering process: Gaussian splashing is used as an intermediary between the two-dimensional portrait and the three-dimensional mesh, which avoids the complexity of traditional three-dimensional rendering and reduces rendering time and computing resource consumption.
[0159] b) Improve the flexibility and efficiency of animation generation:
[0160] 1. Enhanced mobility of motion data: The method of the present invention can transfer motion trajectories from any grid model to any static image, breaking through identity restrictions. This allows motion data to be flexibly applied to different head avatars to meet personalized and diversified needs.
[0161] 2. Reduce manual adjustments: Using two-dimensional priors and gradient-guided optimization methods, the optimization of the Gaussian deformation field is automatically completed, reducing the need for a large number of manual adjustments and professional knowledge, and accelerating the content production process.
[0162] c) Improving the temporal stability of long video sequences: Natural temporal consistency: The grid sequence driven method can maintain temporal continuity and stability when processing long video sequences, avoiding the problems of jitter and inconsistency between frames.
[0163] Compared with the prior art, the present invention has the following advantages:
[0164] a) Reduced resource usage: High-precision 3D models and expensive rendering equipment are no longer required, and hardware requirements are reduced.
[0165] b) Improved time consistency.
[0166] c) Ability to maintain stability when processing long sequences.
[0167] The results of the method of the present invention are as follows Figure 4 As shown in the figure, it can be seen that the 3D Gaussian Splatting modeling can provide good temporal consistency and spatial consistency for the driving video, significantly improving the fidelity under large rotation angles.
[0168] Comparative test: At the same time, we also compared with the current 2D replay method, and the results are as follows Figure 5As shown in the figure, based on the mesh motion sequence driving this character, the visualization results comparing our method with the representative work in the reenactment direction show that the 2D reenactment method performs poorly at extreme viewing angles and has ghosting phenomena. Our method provides higher temporal consistency and spatial consistency, significantly improving the fidelity of the driving video. It can be seen that through 3D Gaussian Splatting modeling, it can provide good temporal consistency and spatial consistency for the driving video, significantly improving the fidelity under large rotation viewing angles.
[0169] In general, the method of the present invention uses motion data with a coarse grid model, allowing the motion trajectory to be transferred to any static image, thereby generating a corresponding video sequence of the speaker's head portrait. This method effectively reduces the rendering cost and complexity, so that art and content creation no longer need to create detailed three-dimensional assets for each static image, improving the efficiency and flexibility of animation generation. In addition, for long video sequences, the grid sequence driven method can provide temporal stability to ensure high continuity and consistency of animation effects. The grid sequence driven method still has obvious advantages in long sequence processing.
[0170] The method of the present invention mainly focuses on migrating a mesh sequence to different character models to drive the same action. Specifically, the technology can be applied to multiple products and scenarios:
[0171] First, in the field of animation production, animators can use this technology to quickly transfer the action sequence of one character to other characters, saving a lot of time and energy while maintaining the consistency and naturalness of the action. This is especially important for industries such as film and television production and game development.
[0172] Secondly, in virtual reality (VR) and augmented reality (AR) applications, this technology can enable users to experience the same action performance when switching between different virtual characters, enhancing immersion and interactivity. For example, users can choose different characters in a VR game and still enjoy the same action experience.
[0173] In addition, in the field of education and training, this technology can be used to create a variety of teaching roles, so that the same explanation content can be presented by different roles, which can enhance the fun and participation of learning. Teachers can transfer the explanation actions of one role to other roles to maintain the consistency of teaching content.
[0174] Finally, on social media and content creation platforms, users can use this technology to transfer their own action performances to different virtual images, create personalized content, and attract more viewers and fans.
[0175] In summary, the technical solution of the present invention has broad application potential in multiple fields such as animation production, virtual reality, education and training, and social media, and can achieve efficient motion migration and consistent performance.
[0176] Example 2
[0177] This embodiment provides a grid sequence driven two-dimensional face animation generation device, comprising:
[0178] 2D image acquisition module, used to acquire the 2D face image that needs to be driven;
[0179] The Gaussian avatar generation module is used to bind the initial frame of the grid sequence motion and the 2D face image using Gaussian splash expression, bind the topology of the 3D grid and the identity information of the 2D face to the Gaussian model, and generate the initial Gaussian avatar;
[0180] The video sequence generation module is used to combine the deformation field to characterize the motion of the grid sequence on the Gaussian expression, optimize the Gaussian head image, and generate an animated video sequence through rendering;
[0181] Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
[0182] Since the device is a grid sequence driven two-dimensional facial animation generation device of an embodiment of the present invention, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation process of the above method embodiment, and the repeated parts will not be repeated.
[0183] Example 3
[0184] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A mesh sequence driven two-dimensional facial animation generation method is shown.
[0185] It is understandable that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0186] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0187] Since the electronic device is an electronic device corresponding to a grid sequence driven two-dimensional facial animation generation method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0188] Example 4
[0189] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A mesh sequence driven two-dimensional facial animation generation method is shown.
[0190] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0191] Since the storage medium is a storage medium corresponding to a grid sequence driven two-dimensional facial animation generation method of an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0192] Example 5
[0193] In some possible implementations, various aspects of the method of the embodiments of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a grid sequence driven two-dimensional face animation generation method according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0194] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0195] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0196] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating two-dimensional facial animation driven by a grid sequence, characterized in that: The following steps are involved: Get the 2D face image that needs to be driven; The initial frame of the mesh sequence motion and the 2D face image are bound using Gaussian splash expression, and the topology of the 3D mesh and the identity information of the 2D face are bound to the Gaussian model to generate the initial Gaussian avatar; Combining deformation fields to characterize the motion of mesh sequences on Gaussian representations, Gaussian avatars are optimized and animated video sequences are generated through rendering; Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
2. A method for generating a two-dimensional facial animation driven by a grid sequence according to claim 1, characterized in that: A Gaussian splash representation method is used as an intermediary between 2D faces and 3D meshes of different identities to achieve the transformation from mesh sequence M to animation sequence video. Transition; where T is the number of frames of the animation.
3. The method for generating a two-dimensional facial animation driven by a grid sequence according to claim 1, characterized in that: The initial frame of the grid sequence motion and the 2D face image are bound using Gaussian splash expression, the topology of the 3D grid and the identity information of the 2D face are bound to the Gaussian model, and the initial Gaussian head portrait is generated, including: Using 3D Gaussian splashing as an intermediary, the connection between 3D meshes and 2D face images under different identities is established; For each unit triangle in the grid, a 3D Gaussian splash unit is bound to it, and the initial center position μ, rotation matrix r, and scaling factor s of the bound Gaussian are given according to the properties of the unit triangle.
4. A method for generating a two-dimensional facial animation driven by a grid sequence according to claim 3, characterized in that: For each unit triangle in the mesh, a 3D Gaussian splash unit is bound to it, and the initial center μ, rotation matrix r, and scaling factor s of the bound Gaussian are given according to the properties of the unit triangle, including: For the center position μ of the bound Gaussian, the expression based on the barycentric coordinates is: For any point V of the spatial triangle ΔV1V2V3, there must be unique coefficients λ1,λ2 such that: Among them, V1 fid 、V2 fid 、V3 fid They are the three vertices of the triangle; For each triangle, the normal vector n is obtained by computing the difference between the two edge vectors: n=(V2 fid -V1 fid )×(V3 fid -V1 fid ) For the direction vector d, the edge composed of vertex 1 and vertex 2 is selected to calculate the direction vector d, that is: In order to describe the direction of the triangle in the global space, the direction vector d of an edge, the normal vector n of the triangle and their cross product d×n are used as column vectors to form the rotation matrix R: R = [d,n,d×n]] The scalar k describing the scaling of the triangle is calculated by taking the length of the shortest side l and its perpendicular h: For the paired 3D Gaussian, define its position μ, rotation r, and anisotropic scale s in local space; initialize the position to the local origin, the rotation r to the unit rotation matrix, and the scale s to the unit vector; at rendering time, transform these attributes to the global space as follows: r’=Rr μ'=kRμ+T s'=ks Where r' is the rotation of the unit Gaussian splash in the global space, μ' is the center position of the unit Gaussian splash in the global space, s' is the scale of the unit Gaussian splash in the global space, and T is the center position of the grid where the Gaussian splash is located in the global space.
5. The method for generating a two-dimensional facial animation driven by a grid sequence according to claim 1, characterized in that: The step of generating an animation video sequence by rendering includes: For any given frame i, the deformation of each vertex is defined as: In the formula, is the position of vertex j in the deformed mesh in the i-th frame, v j,can is the position of the same vertex in the canonical grid; For every Gaussian x in the canonical space i,can , the MLP network is used to calculate its characterization of its deformation, and then the Gaussian is transformed from the standard space to the deformation space: x i,def =x i,can +Def(v j,can ,δv j ) By directly using the deformed vertices of the mesh and the vertex positions in the canonical space, complex deformations in dynamic scenes can be captured more accurately to achieve high-quality dynamic scene rendering.
6. The method for generating a two-dimensional facial animation driven by a grid sequence according to claim 1, characterized in that: The method of combining the deformation field to characterize the movement of the grid sequence on the Gaussian expression and optimizing the Gaussian avatar includes: The 2D reference face image is input into a pre-trained 2D reenactment network, and the grid state driven by the current frame is used as a conditional drive. The 2D reenactment network can give a reference prior of the 2D reference face image after being driven in the current frame based on the current driving information. These reference priors provide guidance for maintaining the consistency between the original image features and the reenactment motion. These reference priors are combined with gradient-based optimization techniques to adjust the Gaussian deformation field to ensure accurate alignment of unique facial features and poses between motion and static images.
7. A method for generating a two-dimensional facial animation driven by a grid sequence according to claim 6, characterized in that: During the training optimization process, a combination of L1 and D-SSIM terms is used to supervise the rendered images: L rgb =(1-λ)L1+λL D-SSIM In the formula, and Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo The average value of and Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo The variance of Is the rendered image I render and the pre-trained 2D reenactment network as the conditional generated image I persudo covariance; C1 and C2 are small constants used to stabilize the denominator; L1 is the loss calculated by the absolute difference between each pixel of the rendered image and the image generated as a condition by the pre-trained 2D reenactment network; L D-SSIM is the structural similarity loss function; λ is L D-SSIM The weight of An L1 term is also used to supervise the transparency produced by the rendering and the loss between the image masks, hoping to ensure pixel alignment: In the formula, The mask image of the face part is generated as a condition for the pre-trained 2D reenactment network. I alpha A mask image for the face part of the rendered image; The final loss function is: L=L rgb +λ mask L mask In the formula, λ mask For L mask The weight of the total loss function.
8. A grid sequence driven two-dimensional face animation generation device, characterized in that: include: 2D image acquisition module, used to acquire the 2D face image that needs to be driven; The Gaussian avatar generation module is used to bind the initial frame of the grid sequence motion and the 2D face image using Gaussian splash expression, bind the topology of the 3D grid and the identity information of the 2D face to the Gaussian model, and generate the initial Gaussian avatar; The video sequence generation module is used to combine the deformation field to characterize the motion of the grid sequence on the Gaussian expression, optimize the Gaussian head image, and generate an animated video sequence through rendering; Among them, the 2D reenactment network processes the input reference face image and the current rendering frame of the mesh as conditions, generates gradients and optimizes the posture and features of the avatar so that the generated animation is highly consistent with the features of the reference face image.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion
CN118351252A
Three-dimensional figure model generation method based on Gaussian splashing
CN118429505A
Sparse visual angle three-dimensional reconstruction method based on depth prior information
CN118657888A
Face high-fidelity and drivable reconstruction method based on three-dimensional Gaussian splashing
CN118736108A
Three dimensional gaussian splatting initialization based on trained neural radiance field representations
US20240355047A1
Cited By
Modeling method and electronic equipment
CN120599148A
Modeling methods and electronic devices
CN120599148B
Digital human grid creating method capable of ensuring continuity of sequence frames
CN121147364A