Video synthesis method and system based on 3D deformation model

Through a video synthesis method based on 3D deformation model, combined with dense motion estimation and target conditional memory compensation network, the problem of low synthetic video quality in the prior art is solved, accurate motion representation and detailed presentation are achieved, and the synthesis effect of speaking portrait video is significantly improved.

CN120034703APending Publication Date: 2025-05-23GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411900435.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to achieve precise motion representation and detail presentation in talking head synthesis, resulting in low quality of synthetic videos, especially when cross-identity synthesis, identity appearance information cannot be well preserved.

Method used

Using a video synthesis method based on 3D deformation model, by obtaining the identity coefficient and head attitude coefficient of the source image and driving the video sequence, combining the dense motion estimation network and the target conditional memory compensation network, more accurate motion flow field and detail compensation characteristics are generated, and high-quality synthetic video is finally achieved.

Benefits of technology

Accurate motion representation and detailed presentation, significantly improving the synthesis effect of speaking portrait videos, ensuring the preservation of identity appearance information and improving video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034703A_ABST
    Figure CN120034703A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video generation, and discloses a video synthesis method and system based on a 3D deformation model, and the method comprises the steps: obtaining a first motion flow field according to a source image and a drive video sequence; inputting the source image and the driving video sequence into the 3D deformation model to obtain an identity coefficient and a head posture coefficient; inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and a shielding graph; inputting the source image into the U-Net, using the second motion flow field to carry out a distortion operation on the coding features, and using the shielding image to adjust an up-sampling mode and a jump connection mode of the features after the distortion operation to obtain aligned features; inputting the aligned features, the identity coefficient and the head posture coefficient into a target condition memory compensation network to obtain a compensated first face feature; further compensating the compensated first face feature by taking the identity coefficient and the head posture coefficient as constraint conditions to obtain a compensated second face feature; and performing up-sampling on the compensated second face feature to obtain a synthetic video. According to the method, accurate motion representation and detail presentation can be realized, so that the quality of the synthesized video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video generation, and in particular to a video synthesis method and system based on a 3D deformable model. Background Art

[0002] One-shot talking head synthesis is the task of synthesizing a talking video from a driving video for a given source person's portrait image by correctly transferring head motion while preserving the source identity. However, the human visual perception system is highly sensitive to the authenticity of portraits, and synthesizing faithful videos with correct facial geometry and appearance remains challenging due to subtle facial motion variations.

[0003] In the existing technology, the latest progress of talking head synthesis can be roughly divided into two main methods: facial landmark-based methods and differentiable neural rendering-based methods. Although the method shows better performance in preserving identity appearance through motion flow distortion, the estimated motion flow can capture the overall movement, but its lack of subtle correspondence leads to artifacts when using sparse facial key points to restore detailed facial expressions. In cross-identity synthesis, due to inaccurate flow fields and occlusion maps, identity appearance information may not be well preserved, resulting in low quality of synthesized videos; and the technology based on 3D deformable model - 3DMM (3D Morphable Model), which pioneered the solution in computer vision, estimating the 3D shape of the head (face) in a single image, and restoring high-quality 3D representations based on images alone. The features estimated by this model are also called face basis vectors. By building a series of independent basis vectors, in theory, a 3D face with any shape and appearance can be represented as a combination of target face basis vectors. However, its estimated parameters can only represent the deformable features of facial geometry, and thus cannot capture areas outside the 3DMM facial contour, such as hair, neck, accessories, background, etc. Therefore, it reduces the quality of the generated areas not covered by the deformation parameters. Summary of the invention

[0004] The primary purpose of the present invention is to overcome the problems existing in the prior art and provide a video synthesis method based on a 3D deformable model. The present invention can achieve accurate motion representation and detail presentation, thereby improving the quality of the synthesized video.

[0005] As another object of the present invention, a system adapted to the method based on the aforementioned object is provided.

[0006] As another object of the present invention, a non-volatile storage medium suitable for storing a computer program implemented according to the method described is provided.

[0007] In order to achieve the above object, the present invention provides a video synthesis method based on a 3D deformable model, the method comprising the following steps: Acquire a source image and a driving video sequence, and obtain a first motion flow field according to the source image and the driving video sequence; Inputting the source image and the driving video sequence into a 3D deformable model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; Inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; Input the source image into U-Net, use the second motion flow field to perform a distortion operation on the encoded features, and use the occlusion map to adjust the upsampling mode and the skip connection mode of the features after the distortion operation to obtain aligned features; Inputting the aligned features, the identity coefficients and the head posture coefficients into a target conditional memory compensation network to obtain a compensated first face feature; Taking the identity coefficient and the head posture coefficient as constraint conditions, further compensating the compensated first facial feature to obtain a compensated second facial feature; The compensated second facial features are up-sampled to obtain a synthetic video.

[0008] Further, obtaining a first motion flow field according to the source image and the driving video sequence includes: Input the source image and the driving video sequence into a key point detector based on the FOMM model to obtain a first key point corresponding to the source image and a second key point corresponding to the driving video sequence; A 2D dense motion estimation network is used to calculate the coordinate difference between the first key point and the second key point, and motion and appearance information are decoupled to obtain a first motion flow field.

[0009] Furthermore, the dense motion estimation network includes several multi-layer perceptrons, affine transformation layers, adaptive instance normalization layers and MotionNet connected in sequence, wherein the several multi-layer perceptrons connected in sequence are used to map the head posture coefficients to a latent space to obtain a latent vector; the affine transformation layer is used to perform an affine transformation operation on the latent vector to obtain an operation result; the adaptive instance normalization layer is used to obtain a feature tensor based on the operation result and the first motion flow field; and the MotionNet is used to obtain a second motion flow field and an occlusion map based on the feature tensor.

[0010] Furthermore, the adaptive instance normalization layer is used to obtain a feature tensor according to the operation result and the first motion flow field, and the calculation method is as follows:

[0011] in, is the result of affine transformation operation, is the first motion flow field, i is the i-th frame of the driving video, and Represent the mean and variance calculation respectively.

[0012] Further, the source image is input into U-Net, the encoded features are distorted using the second motion flow field, and the upsampling mode and the jump connection mode of the features after the distorted operation are adjusted using the occlusion map to obtain the aligned features, specifically including: Input the source image into the encoder of U-Net and output multi-level encoding features; Using the second motion flow field to perform a distortion operation on the multi-level coding features to obtain a distorted feature of each level of coding features; The occlusion map is used to adjust the upsampling mode and the skip connection mode of the distorted features to obtain the aligned features, as follows:

[0013] in, is the characteristic deformation function, represents the upsampling operation, is the first Level features, in particular, for =1, .

[0014] Furthermore, the target conditional memory compensation network includes an implicit identity conditional memory module and a memory compensation module, the implicit identity conditional memory module is used to process according to the aligned features, the identity coefficients and the head posture coefficients to obtain a feature tensor related to the identity, and the memory compensation module obtains the compensated first facial features according to the feature tensor related to the identity and the aligned features.

[0015] Furthermore, the identity coefficient and the head posture coefficient are used as constraints to further compensate the compensated first facial feature to obtain the compensated second facial feature, specifically including: inputting the identity coefficient and the head posture coefficient into a multi-layer perceptron to obtain a mapping matrix, and performing adaptive instance normalization processing on the mapping matrix and the compensated first facial feature to obtain the compensated second facial feature.

[0016] Furthermore, upsampling the compensated second facial feature to obtain a composite video specifically includes upsampling the compensated second facial feature at least 3 times to obtain a composite video.

[0017] In order to achieve another object of the present invention, the present invention provides a video synthesis system based on a 3D deformable model. The system is based on the video synthesis method based on a 3D deformable model, and includes: A first acquisition module: used for acquiring a source image and a driving video sequence, and obtaining a first motion flow field according to the source image and the driving video sequence; A second acquisition module: inputting the source image and the driving video sequence into a 3D deformation model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; A dense motion estimation module: used for inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; Alignment module: used for inputting the source image into U-Net, using the second motion flow field to perform a distortion operation on the encoded features, and using the occlusion map to adjust the upsampling mode and the jump connection mode of the features after the distortion operation to obtain aligned features; A first compensation module: used for inputting the aligned features, the identity coefficients and the head posture coefficients into a target condition memory compensation network to obtain a compensated first face feature; A second compensation module: using the identity coefficient and the head posture coefficient as constraints, further compensating the compensated first facial feature to obtain a compensated second facial feature; Synthesis module: up-sample the compensated second facial features to obtain a synthesized video.

[0018] To achieve another purpose of the present invention, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer-stored program is executed by a processor, the video synthesis method based on a 3D deformable model is implemented.

[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention firstly obtains the identity coefficient of the source image and the head posture coefficient of the driving video sequence as compensation information through the 3DMM model, then compensates the first motion flow field based on the detected key points through the head posture coefficient to generate a more accurate second motion flow field, and then uses the accurate second motion flow field to perform a distortion operation on the source image, and then performs detail compensation through a jump connection based on an occlusion map, uses the learned rich facial feature memory library, and then uses the coefficients of the 3DMM as additional constraint guidance to restore facial details, and finally synthesizes a speaking face video. The present invention establishes a unified model of efficient and accurate motion migration and effectively restores the features of the occluded part through the compensation information, which can achieve accurate motion representation and detail presentation, and significantly improve the synthesis effect of the speaking portrait video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of a video synthesis method based on a 3D deformable model according to embodiment 1 of the present invention; Figure 2 is a block diagram of a video synthesis system based on a 3D deformable model according to Embodiment 3 of the present invention; Figure 3 is a flowchart of a video synthesis method based on a 3D deformable model according to Embodiment 1 of the present invention; Figure 4 4 is a structural diagram of a target conditional memory compensation network (TCMC) according to embodiment 1 of the present invention. DETAILED DESCRIPTION

[0021] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0022] In the description of the present invention, it should be noted that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.

[0023] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0024] Furthermore, in the description of the present invention, unless otherwise specified, “plurality” means two or more.

[0025] Example 1 like Figure 1 , 3 As shown, a video synthesis method based on a 3D deformable model in a preferred embodiment of the present invention includes the following steps: S1: Acquire a source image and a driving video sequence, and obtain a first motion flow field according to the source image and the driving video sequence; In a feasible embodiment, S1 specifically includes: S1.1: Obtain a source image and a driving video sequence, input the source image and the driving video sequence into a key point detector (KPDetector) based on the FOMM model, and obtain a first key point corresponding to the source image and a second key point corresponding to the driving video sequence; S1.2: Use a 2D dense motion estimation network to calculate the coordinate difference between the first key point and the second key point, and decouple motion and appearance information to obtain a first motion flow field, where the first motion flow field is a rough motion flow field.

[0026] S2: inputting the source image and the driving video sequence into a 3D deformation model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; In one possible embodiment, the identity coefficient is expressed as , the expression coefficient is expressed as , and the head rotation coefficient and the head translation coefficient is expressed as ,These coefficients are introduced to encourage the model to predict precise and subtle motions based on the conditional 3DMM parameters.

[0027] S3: inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; In a feasible embodiment, the dense motion estimation network includes a plurality of sequentially connected multi-layer perceptrons, an affine transformation layer, an adaptive instance normalization layer and a MotionNet, wherein the plurality of sequentially connected multi-layer perceptrons (this embodiment uses a cascade of three multi-layer perceptrons MLP) are used to map the head posture coefficients to a latent space to obtain a latent vector, as follows:

[0028]

[0029] in, is the target motion coefficient, including the expression coefficient, head rotation coefficient and head translation coefficient. is a three-layer MLP; The affine transformation layer is used to perform an affine transformation operation on the potential vector to obtain an operation result, which is as follows:

[0030] The adaptive instance normalization layer is used to obtain a feature tensor according to the operation result and the first motion flow field, and the specific calculation is as follows:

[0031] in, is the result of affine transformation operation, is the first motion flow field, i is the i-th frame of the driving video, and Represents mean and variance calculation respectively The MotionNet is used to obtain a second motion flow field and an occlusion map according to the feature tensor.

[0032] S4: inputting the source image into U-Net, using the second motion flow field to perform a distortion operation on the encoded features, and using the occlusion map to adjust the upsampling mode and the skip connection mode of the features after the distortion operation to obtain aligned features; In a feasible embodiment, S4 specifically includes: S4.1: Input the source image into the encoder of U-Net and output multi-level encoding features, as follows:

[0033] Among them, the source portrait image passes through the encoder The four convolution samplings of get multi-level appearance features, where Indicates that the encoder The i-th level features generated are S4.2: Using the second motion flow field to perform a distortion operation on the multi-level coding features to obtain a distorted feature of each level of coding features; S4.3: Using the occlusion map, adjust the upsampling mode and the skip connection mode of the distorted features to obtain aligned features, as follows:

[0034] in, is the characteristic deformation function, represents the upsampling operation, is the first Level features, in particular, for =1, . is a weight vector of the same size as the feature tensor, with values ​​in the interval [0-1], used to select between aligned features and upsampled features, and to transfer features generated by the occlusion map based on the skip connection of feature warping. Select the aligned source features to preserve the rich source portrait information.

[0035] S5: inputting the aligned features, the identity coefficients and the head posture coefficients into a target condition memory compensation network to obtain a compensated first face feature; In a feasible embodiment, the target conditional memory compensation network includes an implicit identity conditional memory module and a memory compensation module, see Figure 4 , the implicit identity condition memory module is used to process according to the aligned features, the identity coefficients and the head posture coefficients to obtain a feature tensor related to the identity, Specifically, since there is an inconsistency problem in maintaining identity when synthesizing a talking head, this embodiment introduces a combined 3DMM parameter set (i.e., identity coefficient and head pose coefficient), which are used to guide the memory network to generate speaking portraits with consistent identities. In particular, the aligned features are processed through projection layers and convolution layers (CONV), global average pooling (GAP) operations, and flattening operations (Flatten). For simplicity, the superscripts are ignored in the following description of this embodiment. The process can be formulated as:

[0036] On this basis, this embodiment introduces the 3DMM parameter set As an additional guidance signal, the accurate features are queried from the feature memory. Specifically, this parameter set contains the target head pose parameters and the identity coefficients of the source portrait, including the source identity coefficients. , driving head posture and expressions The parameters of , whose combination is expressed as:

[0037] in are the identity coefficients of the source image I, represents the combined 3DMM parameter set. It is further injected into the TCMC module to encourage the model to extract identity-related features from the learned memory bank. Specifically, each 3DMM coefficient component is first standardized (Norm) and then compared with the flattened feature Combination (CANCAT):

[0038] The memory compensation module obtains the compensated first face feature according to the identity-related feature tensor and the aligned feature, specifically: The integrated feature tensor is further input into the implicit identity conditional memory (IICM) submodule to learn the meta-memory feature library. Compensated face features Conditioned on the source identity and driving motion information, and queried from the memory compensation module (MCM). Intuitively, this can be formulated as:

[0039] Through these steps, the present embodiment can effectively query and compensate for facial details using the combined 3DMM parameter set, ensure consistent identity features during the speaking head generation process, and improve the accuracy of facial feature recovery and detail reproduction capabilities.

[0040] S6: using the identity coefficient and the head posture coefficient as constraint conditions, further compensating the compensated first facial feature to obtain a compensated second facial feature; In a feasible embodiment, due to the facial features after memory compensation contains rich appearance information and can thus be directly used to render talking head videos. However, the results show artifacts and errors with incorrect facial shapes, as well as mixed identities with appearance leakage. Figure 4 As shown in the right half, this embodiment also introduces a 3DMM coefficient constraint (CC) module to Further modulation compensation characteristics:

[0041] S7: upsampling the compensated second facial features to obtain a synthetic video, upsampling the compensated second facial features at least 3 times to obtain a synthetic video, and during the upsampling stage of the U-Net generator, ensuring that the decoded features do not deviate from expectations and that non-relevant features are not introduced from the memory library to optimize the generation effect: .

[0042] This embodiment first obtains the identity coefficient of the source image and the head posture coefficient of the driving video sequence as compensation information through the 3DMM model, then compensates the first motion flow field based on the detected key points through the head posture coefficient to generate a more accurate second motion flow field, and then uses the accurate second motion flow field to distort the source image, and then performs detail compensation through jump connections based on the occlusion map, uses the learned rich facial feature memory library, and then uses the coefficients of 3DMM as additional constraint guidance to restore facial details, and finally synthesizes the speaking face video. The present invention establishes a unified model for efficient and accurate motion migration and effectively restores the features of the occluded part through compensation information, which can achieve accurate motion representation and detail presentation, and significantly improve the synthesis effect of the speaking portrait video.

[0043] Example 2 In order to verify the superiority of the video synthesis method proposed in Example 1, this embodiment compares the method proposed in Example 1 with five state-of-the-art speaking portrait synthesis methods, including FOMM, PIRendererren and StyleHEAT, DaGAN and MCNet. Specifically, FOMM is a representative key point-based method for face generation and animation. DaGAN combines estimated depth maps (3D information) learned in a self-supervised manner to guide facial generation. In addition, PIRenderer and StyleHEAT are representative methods for speaking head synthesis using 3DMM parameters. MCNet pioneered the application of key point-guided memory networks to speaking portrait synthesis tasks.

[0044] This embodiment adopts a set of widely used indicators to evaluate the performance of image generation, identity preservation, and motion transfer. Specifically, we used four metrics, including structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and learned perceptual image block similarity. In addition, for motion transfer, the average expression distance (AED) and average pose distance (APD) ren2021pirenderer are used to measure the differences in facial expressions and poses between the source image and the generated image. The following experimental results are all completed on the Voxceleb1 public dataset, as shown in Table 1: Table 1

[0045] Same and Cross respectively represent that the source frame and the driving frame are from the same video character and different video characters. ↑ means the lower the better. Conversely ↓ means the higher the better. According to Table 1, the method proposed in Example 1 is better in most indicators.

[0046] Example 3 like Figure 2 As shown, a video synthesis system based on a 3D deformable model according to an embodiment of the present invention includes: A first acquisition module: used for acquiring a source image and a driving video sequence, and obtaining a first motion flow field according to the source image and the driving video sequence; A second acquisition module: inputting the source image and the driving video sequence into a 3D deformation model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; A dense motion estimation module: used for inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; Alignment module: used for inputting the source image into U-Net, using the second motion flow field to perform a distortion operation on the encoded features, and using the occlusion map to adjust the upsampling mode and the jump connection mode of the features after the distortion operation to obtain aligned features; A first compensation module: used for inputting the aligned features, the identity coefficients and the head posture coefficients into a target condition memory compensation network to obtain a compensated first face feature; A second compensation module: using the identity coefficient and the head posture coefficient as constraints, further compensating the compensated first facial feature to obtain a compensated second facial feature; Synthesis module: up-sample the compensated second facial features to obtain a synthesized video.

[0047] The video synthesis system based on a 3D deformable model proposed in this embodiment is based on the method proposed in Example 1. It can be understood that the options proposed in Example 1 are also applicable to this embodiment, so they are not repeated here.

[0048] Example 4 An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer-stored program is executed by a processor, the video synthesis method based on a 3D deformable model is implemented.

[0049] In summary, the embodiments of the present invention provide a video synthesis method, system and storage medium based on a 3D deformable model, which obtains the identity coefficient of the source image and the head posture coefficient of the driving video sequence as compensation information through the 3DMM model, and then compensates the first motion flow field based on the detected key points through the head posture coefficient to generate a more accurate second motion flow field, and then uses the accurate second motion flow field to distort the source image, and then compensates for the details through the jump connection based on the occlusion map, and uses the learned rich facial feature memory library, and then uses the coefficients of the 3DMM as additional constraint guidance to restore the facial details, and finally synthesizes the speaking face video. The present invention establishes a unified model for efficient and accurate motion migration through compensation information, and effectively restores the features of the occluded part, which can achieve accurate motion representation and detail presentation, and significantly improve the synthesis effect of the speaking portrait video.

[0050] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention. These improvements and substitutions should also be regarded as the scope of protection of the present invention.

Claims

1. A video synthesis method based on a 3D deformable model, characterized in that: The method comprises the following steps: Acquire a source image and a driving video sequence, and obtain a first motion flow field according to the source image and the driving video sequence; Inputting the source image and the driving video sequence into a 3D deformable model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; Inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; Input the source image into U-Net, use the second motion flow field to perform a distortion operation on the encoded features, and use the occlusion map to adjust the upsampling mode and the skip connection mode of the features after the distortion operation to obtain aligned features; Inputting the aligned features, the identity coefficients and the head posture coefficients into a target conditional memory compensation network to obtain a compensated first face feature; Taking the identity coefficient and the head posture coefficient as constraint conditions, further compensating the compensated first facial feature to obtain a compensated second facial feature; The compensated second facial features are up-sampled to obtain a synthetic video.

2. The video synthesis method based on 3D deformable model according to claim 1, characterized in that: The step of obtaining a first motion flow field according to the source image and the driving video sequence comprises: Input the source image and the driving video sequence into a key point detector based on the FOMM model to obtain a first key point corresponding to the source image and a second key point corresponding to the driving video sequence; A 2D dense motion estimation network is used to calculate the coordinate difference between the first key point and the second key point, and motion and appearance information are decoupled to obtain a first motion flow field.

3. The video synthesis method based on 3D deformable model according to claim 1, characterized in that: The dense motion estimation network includes a plurality of multi-layer perceptrons, affine transformation layers, adaptive instance normalization layers and MotionNet connected in sequence, wherein the plurality of multi-layer perceptrons connected in sequence are used to map the head posture coefficients to a latent space to obtain a latent vector; the affine transformation layer is used to perform an affine transformation operation on the latent vector to obtain an operation result; the adaptive instance normalization layer is used to obtain a feature tensor based on the operation result and the first motion flow field; and the MotionNet is used to obtain a second motion flow field and an occlusion map based on the feature tensor.

4. The video synthesis method based on 3D deformable model according to claim 3, characterized in that: The adaptive instance normalization layer is used to obtain a feature tensor according to the operation result and the first motion flow field, and the calculation method is as follows: in, is the result of affine transformation operation, is the first motion flow field, i is the i-th frame of the driving video, and Represent the mean and variance calculation respectively.

5. The video synthesis method based on 3D deformable model according to claim 1, characterized in that: Inputting the source image into U-Net, using the second motion flow field to perform a distortion operation on the encoded features, and using the occlusion map to adjust the upsampling mode and the jump connection mode of the features after the distortion operation to obtain the aligned features, specifically including: Input the source image into the encoder of U-Net and output multi-level encoding features; Using the second motion flow field to perform a distortion operation on the multi-level coding features to obtain a distorted feature of each level of coding features; The occlusion map is used to adjust the upsampling mode and the skip connection mode of the distorted features to obtain the aligned features, as follows: in, is the characteristic deformation function, represents the upsampling operation, is the first Level features, in particular, for =1, .

6. The video synthesis method based on 3D deformable model according to claim 1, characterized in that: The target conditional memory compensation network includes an implicit identity conditional memory module and a memory compensation module. The implicit identity conditional memory module is used to process according to the aligned features, the identity coefficients and the head posture coefficients to obtain a feature tensor related to the identity. The memory compensation module obtains the compensated first facial features according to the feature tensor related to the identity and the aligned features.

7. The video synthesis method based on 3D deformable model according to claim 1, characterized in that: The identity coefficient and the head posture coefficient are used as constraints to further compensate the compensated first facial feature to obtain the compensated second facial feature, specifically including: inputting the identity coefficient and the head posture coefficient into a multi-layer perceptron to obtain a mapping matrix, and performing adaptive instance normalization processing on the mapping matrix and the compensated first facial feature to obtain the compensated second facial feature.

8. A video synthesis method based on a 3D deformable model according to any one of claims 1 to 7, characterized in that: The compensated second facial feature is up-sampled to obtain a composite video, specifically comprising up-sampling the compensated second facial feature at least 3 times to obtain a composite video.

9. A video synthesis system based on a 3D deformable model, the system being based on a video synthesis method based on a 3D deformable model according to any one of claims 1 to 8, comprising: A first acquisition module: used for acquiring a source image and a driving video sequence, and obtaining a first motion flow field according to the source image and the driving video sequence; A second acquisition module: inputting the source image and the driving video sequence into a 3D deformation model to obtain an identity coefficient and a head posture coefficient, wherein the head posture coefficient includes at least one of an expression coefficient, a head rotation coefficient and a head translation coefficient; A dense motion estimation module: used for inputting the first motion flow field and the head posture coefficient into a dense motion estimation network, and outputting a second motion flow field and an occlusion map; Alignment module: used for inputting the source image into U-Net, using the second motion flow field to perform a distortion operation on the encoded features, and using the occlusion map to adjust the upsampling mode and the jump connection mode of the features after the distortion operation to obtain aligned features; A first compensation module: used for inputting the aligned features, the identity coefficients and the head posture coefficients into a target condition memory compensation network to obtain a compensated first face feature; A second compensation module: using the identity coefficient and the head posture coefficient as constraints, further compensating the compensated first facial feature to obtain a compensated second facial feature; Synthesis module: up-sample the compensated second facial features to obtain a synthesized video.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer stored program is executed by a processor, a video synthesis method based on a 3D deformable model as described in any one of claims 1 to 8 is implemented.