Video synthesis method and device and electronic equipment

By extracting and fusing the features of the driver video and reference images, high-quality dynamic digital physical videos are generated, which solves the problems of appearance consistency and motion control in the prior art and achieves efficient video synthesis effect.

CN120472052APending Publication Date: 2025-08-12NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510370704.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality dynamic digital physical videos in the fields of virtual reality and augmented reality, and there are problems such as difficulty in maintaining appearance consistency, limited motion control accuracy, poor scene fusion effect and lack of flexibility in the generation process.

Method used

By obtaining reference images and driving videos, the spatio-temporal features, segmentation masks and pose information of the replacement object are extracted, and fused with the appearance features of the target object to generate synthetic videos to realize object replacement and action migration.

Benefits of technology

Achieve a high level of appearance consistency and precise action migration, improve the quality of video synthesis and flexibility in the generation process, and reduce production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472052A_ABST
    Figure CN120472052A_ABST
Patent Text Reader

Abstract

The invention provides a video synthesis method and device and electronic equipment, and the method comprises the steps: extracting spatial-temporal features corresponding to a replacement object from an obtained driving video, and extracting a segmentation mask and attitude information corresponding to each frame of image, the spatial-temporal features being used for indicating motion features and scene features corresponding to the replacement object; extracting appearance features of the target object from the obtained reference image; and fusing the appearance feature, the spatial-temporal feature, the segmentation mask and the attitude information to obtain a synthetic video in which the replacement object in the driving video is replaced by the target object. According to the mode, the object in the driving video can be replaced by the target object in the reference image through the reference image, and meanwhile, high-level appearance consistency is kept; meanwhile, by extracting the attitude information, the segmentation mask and the spatial-temporal characteristics in the driving video, the action in the driving video can be accurately migrated to the target object, and the target object is naturally fused into the scene of the driving video, so that the video synthesis quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video synthesis method, device and electronic device. Background Art

[0002] With the rapid development of emerging technologies such as virtual reality, augmented reality, and the metaverse, the demand for dynamic digital entities in entertainment, education, social interaction, and other fields is growing. High-quality dynamic digital entity video generation technology has become the key to supporting these applications.

[0003] In related technologies, a 3D model of a digital entity is first constructed, then animated using animation techniques, and rendered to generate a video. While 3D models offer good control over geometric structure and perspective consistency, they are expensive to construct and animate, and the resulting rendering lacks realism, making it difficult to capture fine details and dynamic textures. Summary of the Invention

[0004] The present invention aims to provide a video synthesis method, apparatus and electronic device to improve the quality of digital entity video generation.

[0005] In a first aspect, the present disclosure provides a video synthesis method, which includes: obtaining a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame image, and the posture information is used to indicate the action of the replacement object in each frame image; extracting the appearance features of the target object from the reference image; fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame image to obtain a synthesized video; wherein the synthesized video is a video in which the replacement object in the driving video is replaced with the target object.

[0006] In a second aspect, the present disclosure provides a video synthesis device, which includes: a data acquisition module for acquiring a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; a feature extraction module for extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame image, and the posture information is used to indicate the action of the replacement object in each frame image; an appearance extraction module for extracting the appearance features of the target object from the reference image; a feature fusion module for fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame image to obtain a synthesized video; wherein the synthesized video is a video in which the replacement object in the driving video is replaced with the target object.

[0007] In a third aspect, the present disclosure provides an electronic device, which includes a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned video synthesis method.

[0008] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned video synthesis method.

[0009] The embodiments of the present disclosure bring the following beneficial effects:

[0010] The present disclosure provides a video synthesis method, apparatus, and electronic device. The method first obtains a reference image and a driving video, wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and a background scene. The method then extracts spatiotemporal features corresponding to the replacement object from the driving video, as well as segmentation masks and posture information corresponding to each frame. The spatiotemporal features indicate the action and scene features corresponding to the replacement object, the segmentation mask indicates the position of the replacement object in each frame, and the posture information indicates the action of the replacement object in each frame. The method then extracts appearance features of the target object from the reference image, and fuses the appearance features, spatiotemporal features, and the segmentation masks and posture information corresponding to each frame to obtain a synthesized video. The synthesized video is a video in which the replacement object in the driving video is replaced with the target object. This method only requires a single reference image and can replace the object in the driving video with the target object in the reference image while maintaining a high level of appearance consistency. Furthermore, by extracting posture information, segmentation masks, and spatiotemporal features from the driving video, the action in the driving video can be accurately transferred to the target object, and the target object can be naturally integrated into the scene of the driving video, thereby improving the quality of video synthesis.

[0011] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by practicing the above-mentioned technology of the present disclosure.

[0012] In order to make the above-mentioned objects, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0014] Figure 1 A flowchart of a video synthesis method provided by an embodiment of the present disclosure;

[0015] Figure 2 A schematic diagram of a video synthesis system provided by an embodiment of the present disclosure;

[0016] Figure 3 A schematic structural diagram of a video synthesis device provided in an embodiment of the present disclosure;

[0017] Figure 4A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all of them. Generally, the components of the embodiments of the present disclosure described and shown in the drawings herein can be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure as claimed, but merely represents selected embodiments of the present disclosure. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative effort shall fall within the scope of protection of the present disclosure.

[0020] With the rapid development of emerging technologies such as virtual reality, augmented reality, and the metaverse, the demand for dynamic digital entities in entertainment, education, social interaction, and other fields is growing. High-quality dynamic digital entity video generation technology has become the key to supporting these applications.

[0021] Among related technologies, existing dynamic digital entity video generation technologies mainly face the following challenges:

[0022] 1. Difficulty in maintaining appearance consistency: Existing methods find it difficult to ensure that the synthesized digital entity always remains highly consistent with the reference appearance in the video sequence, especially in the case of complex movements and perspective changes, which easily leads to identity drift and blurred details.

[0023] 2. Limited motion control accuracy: Users often struggle to precisely control the movements of synthetic digital entities, ensuring they conform to their intended trajectory and timing. This is especially true when interacting with scene elements or other characters, making natural and coordinated movements difficult to ensure.

[0024] 3. Poor scene fusion effect: When integrating the synthesized digital entity into the target scene, problems such as unnatural edges, lighting mismatch, and incorrect occlusion relationships are likely to occur, resulting in distorted overall visual effects.

[0025] 4. Lack of flexibility in the generation process: Existing technologies usually require retraining the model or performing complex parameter adjustments when changing the appearance of digital entities or target scenes, which lacks flexibility and ease of use.

[0026] To address the above challenges, existing technologies mainly adopt the following solutions, but they all have certain limitations:

[0027] 3D model-based methods first construct a 3D model of the digital entity, then use animation techniques to drive the model's motion and render it into a video. While 3D models offer good control over geometric structure and perspective consistency, building and animating 3D models is expensive, and the rendered results often lack realism, making it difficult to capture fine appearance details and dynamic textures.

[0028] Methods based on 2D image manipulation directly edit and synthesize digital entities in the 2D image space, for example, through techniques such as image deformation and texture replacement. However, 2D manipulation struggles with complex occlusions and perspective changes, and is prone to temporal discontinuities and jitter, making it difficult to generate long, high-quality videos.

[0029] Based on the generative adversarial network (GAN) method, GANs have achieved remarkable results in the field of image generation and have also been applied to video generation. However, GANs training is unstable, the mode collapse problem is prominent, and it is difficult to achieve precise control of the generated video content.

[0030] Based on the variational autoencoder (VAE) method, VAEs can learn the latent space of data distribution and generate new samples. However, the videos generated by VAEs are usually blurry, lack details, and have limited control capabilities.

[0031] Diffusion model-based methods, as emerging generative models, have shown great potential in the field of image and video generation. While the samples generated by diffusion models are high-quality and diverse, effectively guiding the generation process of diffusion models to achieve precise control over the appearance, movement, and scene of digital entities remains a challenge.

[0032] Based on the above problems, the embodiments of the present disclosure provide a video synthesis method, device and electronic device. This technology can be applied to virtual image customization, film and television special effects production, game character generation, online education, virtual / augmented reality and other fields. It can significantly reduce video production costs, improve production efficiency, and provide users with more creative and personalized content creation tools.

[0033] In order to facilitate understanding of the embodiments of the present disclosure, a video synthesis method provided by the embodiments of the present disclosure is first described in detail. Figure 1 As shown, the method includes the following specific steps:

[0034] Step S102 , obtaining a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image including the action of the replacement object and the background scene.

[0035] The reference image is a static image provided by the user. This image defines the appearance of the target synthetic digital entity, which is also the target object in the reference image. The target object can be a person, cartoon character, game character, etc. In one embodiment, the reference image can be a full-body frontal photo of the target object, with a minimally invasive background and uniform lighting.

[0036] The driving video is a user-provided video consisting of a multi-frame video sequence containing the target object, the background scene, and the interaction between the selectable objects. The replacement object in the driving video is replaced with the digital entity represented by the reference image, which is also the target object. Specifically, the driving video can be a game video, a movie clip, or a filmed video.

[0037] Step S104, extracting the spatiotemporal features corresponding to the replacement object from the driving video, as well as extracting the segmentation mask and posture information corresponding to each frame image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame image, and the posture information is used to indicate the action of the replacement object in each frame image.

[0038] In a specific implementation, after acquiring the driving video, it is necessary to extract the spatiotemporal features of the replacement object from the driving video. These spatiotemporal features include the replacement object's motion features and scene features, and may also include temporal features. Motion features may include, but are not limited to, the replacement object's posture, action, and trajectory; scene features include the replacement object's surrounding background, lighting, and occlusion relationships; and temporal features may include the changes between different image frames in the driving video. Specifically, neural networks or deep learning models can be used to extract the corresponding spatiotemporal features of the replacement object from the driving video.

[0039] After acquiring the driving video, it's necessary to extract the segmentation mask and pose information corresponding to each frame in the driving video. The segmentation mask precisely indicates the pixel area within each frame of the driving video's replacement object, or, more accurately, the pixel-by-pixel position of the replacement object within each frame. The pose information indicates the movement of the replacement object within each frame. Specifically, the pose information can indicate a specific action or a sequence of pose keypoint positions corresponding to the replacement movement, including the two-dimensional coordinates of multiple pose keypoints.

[0040] In practice, the specific method for extracting the segmentation mask and posture information corresponding to each frame of the driving video can be determined based on R&D requirements. For example, information can be extracted using a pre-trained neural network model or deep learning model.

[0041] Step S106: extracting appearance features of the target object from the reference image.

[0042] The method of extracting the appearance features of the target object from the reference image can be determined according to R&D requirements. For example, the appearance features can be extracted through a pre-trained neural network model or a deep learning model. Appearance features are used to describe the appearance details of the target object in the reference image. For example, appearance details include the color information, texture information, and shape information of the target object. Color information includes but is not limited to the color of the target object's skin, hair, and clothing; texture information includes but is not limited to the target object's skin texture, clothing material, and facial details (such as wrinkles, freckles, etc.); shape information includes but is not limited to the target object's facial contours, reminders, and clothing style.

[0043] Step S108 , fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain a synthesized video; wherein the synthesized video is a video in which the replacement object in the driving video is replaced with the target object.

[0044] In the specific implementation, the appearance features of the target object are fused with the spatiotemporal features, segmentation mask and posture information provided by the driving video to obtain a synthetic video. The synthetic video is the driving video in which the replacement object is replaced by the target object. The object contained in the synthetic video is the target object, and the target object has the action features of the replacement object, and the environmental background of the target object is the environmental background of the replacement object.

[0045] A video synthesis method provided by the embodiments of the present disclosure can replace the object in the driving video with the target object in the reference image with only one reference image while maintaining a high level of appearance consistency; at the same time, by extracting posture information, segmentation masks and spatiotemporal features in the driving video, the action in the driving video can be accurately transferred to the target object, and the target object can be naturally integrated into the scene of the driving video, thereby improving the quality of video synthesis.

[0046] The following embodiments are used to describe how to extract features from a driving video.

[0047] Specifically, the above-mentioned process of extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image, can be implemented by the following steps 10-11:

[0048] Step 10: Extracting spatiotemporal features corresponding to the replacement object from the driving video through a first neural network.

[0049] In specific implementations, the network structure corresponding to the first neural network can be determined based on R&D requirements. For example, the first neural network can adopt a three-dimensional convolutional neural network, or a structure that combines a two-dimensional convolutional neural network and a recurrent neural network. The first neural network is used to encode the driving video, extract the spatiotemporal features of the driving video, and store the spatiotemporal features in the form of feature maps. The function of the first neural network is to convert the driving video into a compact feature representation for subsequent video synthesis.

[0050] Step 11: For each frame image in the driving video, perform mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image; extract multiple posture key point coordinates of the replacement object from the current frame image, and determine the multiple posture key point coordinates as the posture information corresponding to the current frame image.

[0051] In a specific implementation, mask segmentation processing is performed frame by frame from the driving video to extract the position of the replacement object in the current frame image, and the feature map obtained after extraction is determined as the segmentation mask corresponding to the current image. In addition, it is also necessary to extract multiple posture key point positions of the replacement object from the driving video frame by frame. The posture key point positions can be automatically generated using a pre-trained posture estimation model (for example, OpenPose, DWPose, etc.); wherein the multiple posture key points may include at least one of the following: key points such as the neck, shoulder, elbow, wrist, hip, knee, ankle, etc. of the replacement object. The multiple posture key point positions of the replacement object are used to guide the learning of the action information in the driving video during video synthesis.

[0052] In an optional embodiment, the above-mentioned process of performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image may include: inputting the current frame image into a pre-trained instance segmentation model; identifying the outline of the replacement object in the current frame image through the instance segmentation model, and outputting a segmentation mask based on the outline; wherein the pixel value of the image area containing the replacement object in the segmentation mask is a first value, and the pixel value of the image area not containing the replacement object is a second value.

[0053] In its implementation, an instance segmentation model identifies the outlines of different objects in an image and outputs pixel-level segmentation masks. The corresponding model structure can be determined based on R&D requirements. For example, the instance segmentation model can use the Mask R-CNN model, which performs well in general object segmentation tasks. Specifically, to ensure segmentation quality, a Mask R-CNN model pre-trained on the COCO dataset can be used and fine-tuned according to the actual application scenario. The COCO dataset is a large image dataset primarily used for object detection, segmentation, and captioning tasks.

[0054] Specifically, the instance segmentation model takes each frame of the driving video as input and outputs a binary image containing a segmentation mask of the replacement object. The binary image is also the above-mentioned segmentation mask. The pixel value of each pixel in the binary image can only have two possibilities: a first value and a second value. The specific values corresponding to the first value and the second value can be determined according to research and development needs. In a specific embodiment, the first value is set to 0 and the second value is set to 1. Pixels with a pixel value of 1 (usually displayed as white) represent image areas belonging to the replacement object in the driving video, and pixels with a pixel value of 0 (usually displayed as black) represent background areas. The segmentation masks corresponding to each frame of the driving video can form a binary image sequence that can accurately indicate the position of the replacement object at the pixel level.

[0055] To enhance the robustness of the model, data augmentation is performed on the segmentation mask to update the segmentation mask to simulate the imperfections of segmentation results in real scenes and reduce the model's over-reliance on the precise mask shape. Data augmentation can include at least one of the following: random image dilation, image erosion, and Gaussian blurring.

[0056] It should be noted that the above-mentioned models for enhancing robustness are the first model and the second model described in the following embodiments.

[0057] The following embodiment is used to describe a method for extracting features from a reference image.

[0058] Specifically, the specific process of extracting the appearance features of the target object from the reference image may include: extracting the appearance features corresponding to the target object from the reference image through the second neural network.

[0059] In a specific implementation, the network structure corresponding to the above-mentioned second neural network can be determined according to R&D requirements. The second neural network can adopt a convolutional neural network, such as the ResNet series model. In order to better capture the appearance details of the target object in the reference image, the second neural network can select a model pre-trained on a large image dataset (e.g., ImageNet), remove its fully connected layers, and retain the convolutional layers as feature extractors. The extracted appearance features can be the feature maps output by the intermediate layer of the second network model of the reference image.

[0060] The purpose of removing the fully connected layers here is to preserve the feature maps extracted by the convolutional layers. Fully connected layers are typically used for classification tasks, flattening the feature maps into one-dimensional vectors, thus losing spatial information. However, the task of this disclosure is video generation, which requires preserving spatial information. Therefore, the fully connected layers are removed, and only the convolutional layers are used as feature extractors.

[0061] The following embodiments are used to describe the method of feature fusion.

[0062] Specifically, the above-mentioned process of fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain the synthesized video can be achieved through the following steps 20-22:

[0063] Step 20: Encode the segmentation mask and posture information corresponding to each frame of image to obtain position features and posture features corresponding to the replacement object.

[0064] In the specific implementation, in order to facilitate subsequent video synthesis, the segmentation mask and posture information corresponding to each frame image in the driving video need to be converted into guiding features respectively. The guiding features include: the position features obtained after the segmentation mask corresponding to each frame image is converted and the posture features obtained after the posture information is converted.

[0065] In one specific embodiment, a mask feature converter can be used to convert the segmentation mask corresponding to each frame into a position feature. The position feature indicates the pixel-level position of the replacement object in each frame of the driving video. The position feature is represented in the form of a feature map, where each pixel value in the feature map corresponds to the probability that the corresponding position in the image belongs to the replacement object. The mask feature converter can employ a neural network or a deep learning model.

[0066] In an optional embodiment, a posture feature converter can be used to convert the posture information corresponding to each frame of the driving video into posture features. Since posture features are sequences of posture keypoint coordinates, the posture feature converter can be processed using a multi-layer perceptron (MLP) or a small CNN. To better utilize the temporal information of the posture, the posture feature converter can be combined with a recurrent neural network (RNN), such as an LSTM or GRU.

[0067] Among them, the posture feature is used to indicate the posture and movement information of the replacement object in each frame image of the driving video. It is expressed in the form of a vector and contains the position and motion information of each posture key point of the replacement object. This information can help to control the movement of the generated target object more accurately in the subsequent process, so that the target object is consistent with the movement of the replacement object in the driving video.

[0068] Step 21: Preliminarily fuse the appearance features, spatiotemporal features, position features, and posture features to obtain an initial synthesized video.

[0069] In step 22, the position features, appearance features, spatiotemporal features, and posture features of the target object determined based on the initial synthesized video are subjected to secondary fusion to obtain a final synthesized video.

[0070] The present disclosure achieves effective control and optimization of the target object generation process by adopting feature information of different precisions in different synthesis stages. Specifically, the present disclosure adopts a two-stage synthesis process: first, fine spatial guidance information (e.g., segmentation mask) and motion guidance information (e.g., spatiotemporal features and posture features) are used to preliminarily fuse the appearance features of the target object in the reference image with the action and scene information provided by the driving video. This stage focuses on accurate spatial positioning and action matching; then, based on the basic synthesis stage, a bootstrap learning mechanism is introduced, and coarse spatial guidance information (e.g., bounding box) is used to further synthesize the video. This stage focuses on improving the appearance detail preservation and overall generation quality, and reducing the negative impact that may be introduced by fine guidance information.

[0071] In an optional embodiment, the above-mentioned specific process of preliminarily fusing appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video may include: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a first fusion result; fusing the position features with the first fusion result to obtain a second fusion result; fusing the posture features with the second fusion result to obtain a third fusion result; and denoising the third fusion result to obtain an initial synthetic video.

[0072] In a specific implementation, an intermediate representation is obtained by encoding the spatiotemporal features and gradually adding noise. This intermediate representation is the aforementioned noisy spatiotemporal features. The appearance features corresponding to the target object are then fused with the noisy spatiotemporal features, positional features, and posture features to obtain a fusion result. This fusion result is then denoised to produce an initial composite video. This initial composite video is a video in which the replacement object in the driving video is replaced with the appearance of the target object. However, the appearance of the target object in the initial composite video may not exist, requiring further video synthesis.

[0073] In an optional embodiment, the above-mentioned process of performing noise processing on the spatiotemporal features to obtain the noisy spatiotemporal features may include: at each noise step, determining the noise variance according to a preset noise scheduling strategy, generating a Gaussian noise tensor with the same size as the spatiotemporal features based on the noise variance, and adding the Gaussian noise tensor to the spatiotemporal features to obtain the noisy spatiotemporal features.

[0074] In the specific implementation, noise is gradually added to the spatiotemporal features according to a predefined noise scheduling strategy. At each noise step, the corresponding noise variance is calculated according to the noise scheduling strategy, and then a Gaussian noise tensor with the same size as the spatiotemporal feature is generated. The Gaussian noise tensor is added to the spatiotemporal feature to obtain the noisy spatiotemporal feature.

[0075] The present disclosure can use a diffusion model to add noise. To better understand the noise addition process, the diffusion model, noise injection, and noise scheduling strategy are further explained here:

[0076] Diffusion models are a type of generative model whose core idea is to transform a data distribution into a pure noise distribution by gradually adding noise, and then learn the inverse denoising process to gradually recover the original data from the pure noise. Diffusion models have achieved breakthroughs in image, audio, and video generation, generating high-quality and diverse samples. In this disclosure, the Denoising Diffusion Probabilistic Model (DDPM) can be used as the basic generative model.

[0077] Noise injection is the forward pass of the diffusion model, which involves gradually noise-addressing the spatiotemporal features extracted from the driving video. Specifically, for each frame of the driving video, Gaussian noise is added to its pixel values according to a predefined noise scheduling strategy. As the number of noise steps increases, the amount of added noise gradually increases, ultimately completely destroying the original driving video image into random noise. The noised spatiotemporal features are an intermediate product of the forward pass of the diffusion model, not its output; the output of the diffusion model is the predicted noise.

[0078] The noise scheduling strategy controls the amount of noise added at different noise step numbers. A reasonable noise scheduling strategy is crucial for training diffusion models. In this disclosure, a linear noise schedule can be used, where the noise variance increases linearly with the number of noise steps. Linear noise scheduling is simple to implement and works well in practice. In addition to linear scheduling, other commonly used noise scheduling strategies include cosine scheduling.

[0079] In one embodiment, the parameters included in the noise injection process are: the initial noise level can be set to 0.01, the linear noise schedule can be set from β_start = 0.0001 to β_end = 0.02, and the number of noise steps T = 1000. The specific value or formula for the noise increase at each step can be clearly defined as the linear interpolation formula: β_t = β_start + (β_end - β_start) * (t / T).

[0080] In view of the characteristics of video data, the present disclosure introduces a temporal attention module to enhance the modeling capability of video temporal information during video synthesis. The temporal attention module can adopt a self-attention mechanism or a cross-attention mechanism to capture the temporal dependency between image frames of the video. Specifically, the specific process of fusing the position feature with the first fusion result to obtain the second fusion result can include: obtaining the temporal dependency between different image frames in the driving video through the temporal attention module; based on the temporal dependency, fusing the position feature with the first fusion result to obtain the second fusion result.

[0081] In an optional embodiment, the above-mentioned specific process of performing secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video may include: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a fourth fusion result; fusing the position features of the target object determined based on the initial synthetic video with the fourth fusion result to obtain a fifth fusion result; fusing the posture features with the fifth fusion result to obtain a sixth fusion result; and denoising the sixth fusion result to obtain the final synthetic video.

[0082] In specific implementation, the above-mentioned secondary fusion and initial fusion are performed in the same way, but the position features used in the secondary fusion and the initial fusion are different. The position features used in the initial fusion are determined based on the pixel-level segmentation mask extracted from each frame image of the driving video; the position features used in the secondary fusion are determined based on the position of the target object in the initial synthetic video.

[0083] In practical applications, before performing a secondary fusion of the position features, appearance features, spatiotemporal features, and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video, it is necessary to determine the bounding box corresponding to the target object in each frame of the initial synthetic video, encode the bounding box, and obtain the target position features; wherein the bounding box is used to indicate the position of the target object in each frame.

[0084] In a specific implementation, a coarse bounding box sequence is extracted frame by frame from the initial composite video. This bounding box sequence includes the bounding box corresponding to the target object in each frame of the initial composite video. This bounding box is also the minimum bounding rectangle that frames the target object in the image. The bounding box can be automatically generated using an object detection model or a simple heuristic algorithm (for example, based on the bounding rectangle of a segmentation mask). The coarse bounding box only needs to roughly frame the location of the target object in the image, and does not need to be accurate to the pixel level like a fine segmentation mask.

[0085] In another optional embodiment, the specific process of performing preliminary fusion of appearance features, spatiotemporal features, position features, and posture features to obtain an initial composite video may further include: performing preliminary fusion of appearance features, spatiotemporal features, position features, and posture features using a first model to obtain an initial composite video. The specific process of performing secondary fusion of position features, appearance features, spatiotemporal features, and posture features of a target object determined based on the initial composite video to obtain a final composite video may further include: performing secondary fusion of position features, appearance features, spatiotemporal features, and posture features of a target object determined based on the initial composite video using a second model to obtain a final composite video.

[0086] In a specific implementation, the above-mentioned first model is pre-trained and can preliminarily fuse the appearance features of the target object with the action and scene information provided by the driving video. The first model uses fine segmentation masks and posture key points for model training.

[0087] In a specific embodiment, the first model can adopt a three-dimensional convolutional neural network based on the U-Net architecture. U-Net is a commonly used image segmentation and generation network. Its jump connection structure can effectively fuse shallow and deep features, which is conducive to generating high-quality images and videos. The input of the first model is appearance features, spatiotemporal features, position features and posture features, and the output is predicted noise. The initial synthetic video is gradually recovered from the noisy spatiotemporal features through iterative denoising in the first model. Specifically, the first model first performs noise processing on the spatiotemporal features to obtain noisy spatiotemporal features, and then fuses the noisy spatiotemporal features with the appearance features to obtain a first fusion result. The position features are then fused with the first fusion result to obtain a second fusion result. The posture features are then fused with the second fusion result to obtain a third fusion result. Based on the third fusion result, a preset noise is output, and the third fusion result is denoised to obtain the initial synthetic video.

[0088] In specific implementations, the first model adds and removes noise using a diffusion model. The training goal of the diffusion model is to enable the first model to accurately predict noise under given conditions. The training goal of the diffusion model is to enable the first model to learn the denoising process, that is, to predict the noise components contained in the noisy spatiotemporal feature view under a given noise level (also known as the number of noise steps) and conditional information (such as appearance features, position features, and posture features).

[0089] In the specific implementation, the various input features of the first model are fused at different levels of the U-Net structure, and the fusion methods are mainly splicing and adaptive feature fusion. In a specific embodiment, the U-Net structure includes 5 layers, and at each layer, the appearance features are fused with the feature map of the current layer through the Adaptive Feature Fusion Module (AFFM). The AFFM module can adaptively learn the fusion weights and effectively fuse features from different sources. Then, in the middle layer (between the 3rd and 4th layers), the position features are spliced and fused with the feature map of the current layer. In the deepest layer (bottleneck layer), the posture features are spliced and fused with the feature map of the current layer to obtain the third fusion result.

[0090] In a specific embodiment, the U-Net encoder includes a total of 5 layers. The structure of each layer is as follows:

[0091] Layer 1: Two 3D convolutional layers (kernel 3x3x3, stride 1x1x1, padding 'same') followed by ReLU activation and normalization (num_groups=32).

[0092] Layer 2: Two 3D convolutional layers (kernel size 3x3x3, stride 1x2x2, padding 'same'), followed by a ReLU activation function and a group normalization layer (num_groups=32). Note that the stride here is 1x2x2, which is downsampling.

[0093] Layer 3: Two 3D convolutional layers (kernel 3x3x3, stride 1x2x2, padding 'same') followed by ReLU activation and group normalization (num_groups=32).

[0094] Layer 4: Two 3D convolutional layers (kernel 3x3x3, stride 1x2x2, padding 'same') followed by ReLU activation and group normalization (num_groups=32).

[0095] Layer 5 (bottleneck layer): two 3D convolutional layers (kernel 3x3x3, stride 1x1x1, padding 'same'), followed by a ReLU activation function and a group normalization layer (num_groups=32).

[0096] The convolutional layer extracts local features using convolution kernels; the activation function introduces nonlinearity to enhance the model's expressiveness; the group normalization layer accelerates model training and improves model stability; and a convolution operation with a stride of 1x2x2 reduces the spatial resolution of the feature map. Furthermore, the feature maps of each layer contain feature information at different levels and scales. Shallow feature maps contain more detailed information, while deep feature maps contain more semantic information.

[0097] The specific structure of the above adaptive feature fusion module (AFFM) is as follows:

[0098] The AFFM module consists of a convolutional layer and a Sigmoid activation function. The input is two feature maps, and AFFM first converts the feature Figure 2 The convolution layer is used to reduce the dimension, and then the fusion weight (between 0 and 1) is obtained through the Sigmoid activation function. Finally, the weight is combined with the feature Figure 1 Perform element-by-element multiplication to obtain the fused output features. AFFM can adaptively learn features Figure 1 and features Figure 2 The fusion ratio is increased to achieve more effective feature fusion.

[0099] Optionally, the loss function of the first model in model training can be determined according to R&D requirements. For example, the loss function L uses mean square error loss to measure the difference between the predicted noise ε^ and the actual noise ε. By minimizing the loss function L, the backpropagation algorithm and optimizer (for example, AdamW) are used to iteratively update the parameters of the first model, the first neural network (optional, if the first neural network also participates in the training) and the feature transformer. During the model training process, the parameters of the second neural network are usually kept fixed to utilize the pre-trained spatiotemporal feature extraction capabilities.

[0100] In an optional embodiment, the first model and the second model have different model structures, and the first model and the second model process input data in the same manner, but the input data of the first model and the second model are different.

[0101] In another optional embodiment, the first model and the second model have the same model structure; wherein the second model is trained based on the video obtained by the first model.

[0102] The goal of the second model is to further improve the quality of the synthesized video based on the first model, especially the ability to preserve the appearance details of the target object. This stage introduces a bootstrap learning mechanism and uses coarse spatial guidance information (bounding boxes) for model training.

[0103] In the bootstrap learning mechanism, bootstrap data must first be generated. This involves using the first model to generate bootstrap training data. The specific steps are as follows: A reference image Ir and a driving video Vd are randomly selected from the training dataset. Then, using the first model, an initial composite video is generated based on the reference image Ir and the driving video. This initial composite video serves as the "target video" in the bootstrap training data. A sequence of coarse bounding boxes is then extracted frame by frame from the initial composite video. These coarse bounding boxes only roughly define the target object's position in the image. Finally, pose information is extracted from the driving video.

[0104] To better understand bootstrap learning, a more detailed explanation is provided here: Bootstrap learning is an iterative learning strategy whose core idea is to use the model's own predictions as new training data, which in turn improves the model itself. In the field of machine learning, bootstrap learning is commonly used for tasks such as semi-supervised learning, weakly supervised learning, and model self-improvement. In this disclosure, a bootstrap learning mechanism is introduced, aiming to use the "target video" generated by the first model as new training data to train the second model, thereby iteratively improving the quality of the synthesized video.

[0105] In addition, although the first model can initially synthesize digital entity videos, it may be insufficient in preserving details. For example, the synthesized digital entity (equivalent to the target object) may appear slightly blurry or slightly different from the reference image. To address this issue, the present disclosure introduces a bootstrap learning mechanism. By using the videos generated by the first model as training data, the second model can learn how to better preserve the detailed features of the digital entity.

[0106] The iteration rules for the bootstrap learning described above can be determined based on R&D needs. For example, it can be set to two iterations. The specific steps of each iteration include: Step 1, using the first model to generate N initial synthetic videos to construct a bootstrap training dataset; Step 2, using the bootstrap training dataset to train the second model for a certain number of epochs (for example, 10 epochs). The performance of the second model is evaluated. If the performance improvement reaches a preset threshold or the number of iterations reaches the upper limit, the iteration is stopped; otherwise, steps 1 and 2 are repeated.

[0107] It should be noted that the data update method explicitly regenerates bootstrapped training data for each iteration, rather than accumulating data based on previous iterations. To prevent error accumulation, the following measures can be taken: limit the number of bootstrapped learning iterations to avoid over-reliance on the model's own predictions; add a small amount of real data to the bootstrapped training dataset to maintain the connection between the model and the real data distribution; and use appropriate loss functions and regularization methods to enhance the model's robustness and generalization capabilities.

[0108] The second model reuses the network structure and initial parameters of the first model, and only fine-tunes its parameters during the bootstrap learning process. The extraction methods for the appearance, spatiotemporal, and posture features from the input features of the second model are exactly the same as those used for the first model, and will not be further described.

[0109] like Figure 2 The diagram shows a schematic diagram of a video synthesis system provided by an embodiment of the present disclosure. First, a reference image and a driving video are obtained. Then, the appearance features of the target object are extracted from the reference image. The segmentation mask, posture information, and spatiotemporal features of the replacement object are extracted frame by frame from the driving video. Then, the segmentation mask is converted into position features, and the posture information is converted into posture features. Then, based on the appearance features, position features, posture features, and spatiotemporal features, a basic video synthesis is performed to obtain an initial synthesized video. That is, the appearance features of the target object in the reference image are fused to the action information and scene information provided by the driving video through a first model to obtain an initial synthesized video. Then, the position of the target object is located again through the initial synthesized video, and refined synthesis is performed based on the position. That is, the appearance features of the target object in the reference image are fused to the initial synthesized video through a second model to obtain a final synthesized video. The final synthesized video is a video in which the replacement object in the driving video is replaced with the appearance of the target object.

[0110] The above-mentioned video synthesis method, through a "fine to coarse" multi-stage guidance strategy, can more effectively utilize different types of guidance information, gradually improve the quality and controllability of the synthesized video, and ultimately achieve high-quality, highly flexible and highly controllable dynamic digital entity video generation.

[0111] In addition, this method only requires one reference image to replace the characters in the driving video with those in the reference image, while maintaining a high level of appearance consistency, including details, lighting, and texture, avoiding the "stiffness" and "unnaturalness" of traditional action replacement technology. Moreover, by extracting posture information (such as key points of the human skeleton) and fine segmentation masks from the driving video, this technology can accurately transfer the actions in the target video to the synthesized digital human, achieving "copy and paste" of the actions without the user having to perform any performance. At the same time, this technology uses advanced deep learning models and fine-to-coarse guidance strategies to naturally integrate the synthesized digital human into the scene of the target video, avoiding problems such as edge artifacts and inconsistent lighting, and achieving "seamless fusion."

[0112] In addition, this disclosure adopts a two-stage generation process, combined with a bootstrap learning mechanism, to achieve high-quality video generation without the need for large amounts of labeled data. Users only need to provide reference images and target videos to quickly generate the required digital human video.

[0113] To facilitate understanding of the embodiments of the present disclosure, the following embodiment provides a specific implementation method for video synthesis, which details how to use the video synthesis method proposed in the present disclosure to replace the character in the game video with the user's own image and achieve action migration.

[0114] First, prepare the input data, which includes a reference image and a driving video. The reference image is a full-body, frontal photo provided by the user, with a clean background and even lighting. The image size is adjusted to 768x768 pixels, with pixel values normalized to the range [-1, 1] and a data type of float32. The driving video is a game video, for example, a clip of the target game character performing a special skill. The video should be approximately 5 seconds long, with a frame rate of 30 FPS, a resolution of 768x768 pixels, pixel values normalized to the range [-1, 1], and a data type of float32.

[0115] Then, using the Mask R-CNN model pre-trained on the COCO dataset, we perform instance segmentation on each frame of the driving video and extract the segmentation mask of the target game character. To enhance the robustness of the model, the segmentation mask undergoes data augmentation: random dilation or erosion with a probability of 50%, with the kernel size randomly selected in the range [3, 7]; and Gaussian blurring with a probability of 50%, with the sigma randomly selected in the range [0.5, 1.5]. The augmented segmentation mask sequence is a binary image sequence with dimensions (batch = 1, time_length = 150, height = 768, width = 768) and data type int64 (0 or 1).

[0116] Using the OpenPose model, we extract the coordinates of the target game character's 17 pose keypoints from each frame of the driving video. These keypoints include the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles. The pose keypoint sequence contains the coordinates of all pose keypoints for all image frames. The sequence size is (batch = 1, time_length = 150, num_keypoints = 17 * 2 = 34), and the data type is float32.

[0117] The second neural network is then a ResNet50 model pre-trained on ImageNet (with the fully connected layers removed and the convolutional layers up to stage 4 retained). The input of the second neural network is the reference image, and the output is appearance features with a size of (batch = 1, channel = 2048, height = 24, width = 24) and a float32 data type. A 3D CNN is used as the first neural network, consisting of five 3D convolutional layers, each with a kernel size of 3x3x3 and a stride of 1x2x2. The output channels are 64, 128, 256, 512, and 512, respectively. Each convolution layer is followed by a ReLU activation function and a 3D BatchNorm. The input of the first neural network is the driving video, and the output is spatiotemporal features with a size of (batch = 1, channel = 512, time_length = 5, height = 24, width = 24) and a float32 data type.

[0118] The segmentation mask sequence is then converted into positional features using a mask feature transformer. This transformer consists of three 2D convolutional layers, each with a 3x3 kernel size, a 2x2 stride, 64 output channels, a True bias, 'same' padding, a ReLU activation function, and BatchNorm. The input is the enhanced fine-grained segmentation mask sequence, and the output is positional features of size (batch=1, channel=64, time_length=150, height=24, width=24) and data type float32.

[0119] The pose information is converted into pose features through the pose feature transformer. The pose feature transformer consists of three MLP layers, each containing a linear layer and a Reluctant Unit (ReLU) activation function. The linear layer has an input dimension of 34, the intermediate layer has a dimension of 128, and the output has a dimension of 64. The input is a sequence of pose keypoints, and the output is pose features with a size of (batch = 1, feature_dim = 64) and a float32 data type.

[0120] The first model uses a U-Net 3D architecture, consisting of a 5-layer encoder and a 5-layer decoder. Each encoder and decoder layer consists of two 3D convolutional layers (3x3x3 kernel, 1x1x1 stride, and 'same' padding), followed by ReLU activation and group normalization (num_groups = 32). A temporal attention module is added between the bottleneck layer (layer 3) of the encoder and the corresponding layer (layer 3) of the decoder. This module uses a multi-head self-attention mechanism with 8 heads. Appearance features are fused with the feature maps of each layer of the U-Net encoder via an AFFM module. The convolutional layer parameters of the AFFM module are: kernel size 3x3, stride 1x1, input channels equal to the number of channels in the current layer feature map, output channels 1, bias set to True, and 'same' padding. Position features are concatenated and fused with the feature map of the third layer of the U-Net encoder; pose features are concatenated and fused with the feature map of the bottleneck layer (layer 5) of the U-Net encoder. The noise steps are converted into vector representation through an Embedding layer (embedding_dim=256) and added to the feature maps of each layer of the U-Net encoder.

[0121] In the first model, noise injection uses a linear noise schedule with β_start = 0.0001, β_end = 0.02, and T = 1000. For a given number of noise steps t (uniformly sampled from [1, 1000]), the noise variance β_t = β_start + (β_end - β_start) * (t / T). Gaussian noise is added to the spatiotemporal features to obtain noisy spatiotemporal features. The loss function uses mean squared error (MSE) to measure the difference between the predicted noise ε^ and the true noise ε. The AdamW optimizer was used for training the first model, with a learning rate of 1e-4, β1 = 0.9, β2 = 0.999, a weight decay coefficient of 0.01, and a batch size of 1.

[0122] The first model uses a user-provided reference image and game video as input to generate an initial synthetic video. In this initial synthetic video, the target game character's appearance is replaced with the user's image, while the action and scene remain unchanged. A sequence of rough bounding boxes is extracted from the initial synthetic video using a simple heuristic algorithm: for each frame of the initial synthetic video, the region bounded by the union of the region and the segmentation mask generated by the first model is calculated. The bounding rectangle of this region is then calculated as the rough bounding box for that frame.

[0123] The input data for the second model is identical to the first model except for the positional features, which are replaced with rough bounding boxes. After passing through the second model, the final composite video is generated. In this composite video, the target game character's appearance is successfully replaced with the user's image, while preserving the target character's movements and the game scene.

[0124] The present invention can effectively overcome the shortcomings of existing technologies and achieve high-quality, highly controllable dynamic digital entity video generation. Compared with existing technologies, the present invention has the following significant advantages:

[0125] Better preservation of appearance details: Through bootstrap learning and coarse mask guidance, the appearance details of the reference digital entity (the digital entity here is also the object) can be better preserved, generating more realistic and delicate digital entity videos.

[0126] Higher motion control accuracy: The joint guidance of fine segmentation masks and pose key point sequences enables precise control of digital entity movements and spatial positions, generating videos with more natural and smoother movements.

[0127] More natural scene fusion effects: Through fine mask guidance and self-supervised training, more natural scene fusion effects can be achieved, reducing artifacts and inconsistencies.

[0128] The generation process is more flexible: the introduction of a two-stage training framework and multimodal guidance information makes the present disclosure more flexible and adaptable when replacing the reference digital entity appearance or driving video, without the need to retrain the entire model.

[0129] The disclosed technology can be widely applied in the following scenarios:

[0130] Avatar customization: Users can provide their own photos or cartoon images as reference images, select their favorite driving videos, and quickly generate personalized dynamic digital avatar videos for use in scenarios such as social media and virtual live broadcasts.

[0131] Movie special effects production: The present disclosure can be used to quickly replace characters in movies or advertising videos, achieving low-cost and efficient character replacement and special effects production.

[0132] Game character generation: Game developers can use this disclosure to quickly generate game character videos of various styles and actions for game promotion, character display, etc.

[0133] Education and training: Static photos of teachers or trainers can be combined into vivid teaching videos to enhance the interactivity and fun of online education and training.

[0134] Virtual Reality / Augmented Reality Applications: Provide high-quality virtual character content for VR / AR applications to enhance user immersion and interactive experience.

[0135] Corresponding to the above method embodiment, the present disclosure provides a video synthesis device, such as Figure 3 As shown, the device includes:

[0136] The data acquisition module 30 is used to acquire a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene.

[0137] The feature extraction module 31 is used to extract the spatiotemporal features corresponding to the replacement object from the driving video, as well as the segmentation mask and posture information corresponding to each frame of the image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame of the image, and the posture information is used to indicate the action of the replacement object in each frame of the image.

[0138] The appearance extraction module 32 is used to extract the appearance features of the target object from the reference image.

[0139] The feature fusion module 33 is used to fuse the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame image to obtain a synthetic video; wherein the synthetic video is a video in which the replacement object in the driving video is replaced by the target object.

[0140] The above-mentioned video synthesis device only needs one reference image to replace the object in the driving video with the target object in the reference image while maintaining a high level of appearance consistency; at the same time, by extracting posture information, segmentation masks and spatiotemporal features in the driving video, it can accurately transfer the action in the driving video to the target object, and naturally integrate the target object into the scene of the driving video, thereby improving the quality of video synthesis.

[0141] Furthermore, the feature extraction module 31 is used to: extract the spatiotemporal features corresponding to the replacement object from the driving video through a first neural network; perform mask segmentation processing on the replacement object in the current frame image for each frame image in the driving video to obtain a segmentation mask corresponding to the current frame image; extract multiple posture key point coordinates of the replacement object from the current frame image, and determine the multiple posture key point coordinates as the posture information corresponding to the current frame image.

[0142] Furthermore, the feature extraction module 31 is also used to: input the current frame image into a pre-trained instance segmentation model; identify the outline of the replacement object in the current frame image through the instance segmentation model, and output a segmentation mask based on the outline; wherein the pixel value of the image area containing the replacement object in the segmentation mask is a first value, and the pixel value of the image area not containing the replacement object is a second value.

[0143] Furthermore, the above-mentioned device also includes a data enhancement module, which is used to: perform data enhancement processing on the segmentation mask and update the segmentation mask; wherein the data enhancement processing includes at least one of the following: image random dilation processing, image corrosion processing and image Gaussian blur processing, etc.

[0144] Furthermore, the appearance extraction module 32 is configured to extract appearance features corresponding to the target object from the reference image through a second neural network.

[0145] Furthermore, the feature fusion module 33 is used to: encode the segmentation mask and posture information corresponding to each frame of the image respectively to obtain the position features and posture features corresponding to the replacement object; preliminarily fuse the appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video; and perform secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video.

[0146] Furthermore, the feature fusion module 33 is also used to: perform noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fuse the noisy spatiotemporal features with the appearance features to obtain a first fusion result; fuse the position features with the first fusion result to obtain a second fusion result; fuse the posture features with the second fusion result to obtain a third fusion result; and denoise the third fusion result to obtain an initial synthesized video.

[0147] Furthermore, the feature fusion module 33 is also used to: determine the noise variance according to a preset noise scheduling strategy at each noise step, generate a Gaussian noise tensor with the same size as the spatiotemporal feature based on the noise variance, add the Gaussian noise tensor to the spatiotemporal feature, and obtain the noisy spatiotemporal feature.

[0148] Furthermore, the feature fusion module 33 is also used to: obtain the temporal dependency between different image frames in the driving video through the temporal attention module; and fuse the position feature with the first fusion result based on the temporal dependency to obtain a second fusion result.

[0149] Furthermore, the feature fusion module 33 is also used to: perform noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fuse the noisy spatiotemporal features with the appearance features to obtain a fourth fusion result; fuse the position features of the target object determined based on the initial synthetic video with the fourth fusion result to obtain a fifth fusion result; fuse the posture features with the fifth fusion result to obtain a sixth fusion result; and denoise the sixth fusion result to obtain a final synthetic video.

[0150] Furthermore, the above-mentioned device also includes a position determination module, which is used to: determine the bounding box corresponding to the target object in each frame image in the initial synthetic video before performing a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video, encode the bounding box, and obtain the target position features; wherein the bounding box is used to indicate the position of the target object in each frame image.

[0151] Furthermore, the feature fusion module 33 is also used to: perform a preliminary fusion of the appearance features, spatiotemporal features, position features and posture features through a first model to obtain an initial synthetic video; and perform a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video through a second model to obtain a final synthetic video.

[0152] Furthermore, the above-mentioned first model and second model have the same model structure; wherein, the second model is obtained by training the video obtained by the first model.

[0153] The video synthesis device provided in the embodiment of the present disclosure has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding content in the aforementioned method embodiment.

[0154] The present disclosure also provides an electronic device, such as Figure 4 As shown, the electronic device includes a processor and a memory, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned video synthesis method.

[0155] Specifically, the above-mentioned video synthesis method includes: obtaining a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame image, and the posture information is used to indicate the action of the replacement object in each frame image; extracting the appearance features of the target object from the reference image; fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame image to obtain a synthesized video; wherein the synthesized video is a video in which the replacement object in the driving video is replaced with the target object.

[0156] The above-mentioned video synthesis method only requires a reference image to replace the object in the driving video with the target object in the reference image while maintaining a high level of appearance consistency. At the same time, by extracting posture information, segmentation masks and spatiotemporal features from the driving video, it can accurately transfer the action in the driving video to the target object and naturally integrate the target object into the scene of the driving video, thereby improving the quality of video synthesis.

[0157] In an optional embodiment, the above-mentioned steps of extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image, include: extracting the spatiotemporal features corresponding to the replacement object from the driving video through a first neural network; for each frame image in the driving video, performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image; extracting multiple posture key point coordinates of the replacement object from the current frame image, and determining the multiple posture key point coordinates as the posture information corresponding to the current frame image.

[0158] In an optional embodiment, the above-mentioned step of performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image includes: inputting the current frame image into a pre-trained instance segmentation model; identifying the outline of the replacement object in the current frame image through the instance segmentation model, and outputting a segmentation mask based on the outline; wherein the pixel value of the image area containing the replacement object in the segmentation mask is a first value, and the pixel value of the image area not containing the replacement object is a second value.

[0159] In an optional embodiment, the above method further includes: performing data enhancement processing on the segmentation mask to update the segmentation mask; wherein the data enhancement processing includes at least one of the following: image random dilation processing, image erosion processing, and image Gaussian blur processing, etc.

[0160] In an optional embodiment, the step of extracting appearance features of the target object from the reference image includes: extracting appearance features corresponding to the target object from the reference image through a second neural network.

[0161] In an optional embodiment, the above-mentioned step of fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain a synthetic video includes: encoding the segmentation mask and posture information corresponding to each frame of the image respectively to obtain the position features and posture features corresponding to the replacement object; preliminarily fusing the appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video; and performing a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain a final synthetic video.

[0162] In an optional embodiment, the above-mentioned step of preliminarily fusing appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a first fusion result; fusing the position features with the first fusion result to obtain a second fusion result; fusing the posture features with the second fusion result to obtain a third fusion result; and denoising the third fusion result to obtain an initial synthetic video.

[0163] In an optional embodiment, the above-mentioned step of performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features includes: at each noise step, determining the noise variance according to a preset noise scheduling strategy, generating a Gaussian noise tensor with the same size as the spatiotemporal features based on the noise variance, and adding the Gaussian noise tensor to the spatiotemporal features to obtain the noisy spatiotemporal features.

[0164] In an optional embodiment, the above-mentioned step of fusing the position feature with the first fusion result to obtain the second fusion result includes: obtaining the time dependency between different image frames in the driving video through a time attention module; based on the time dependency, fusing the position feature with the first fusion result to obtain the second fusion result.

[0165] In an optional embodiment, the above-mentioned step of performing a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a fourth fusion result; fusing the position features of the target object determined based on the initial synthetic video with the fourth fusion result to obtain a fifth fusion result; fusing the posture features with the fifth fusion result to obtain a sixth fusion result; and denoising the sixth fusion result to obtain the final synthetic video.

[0166] In an optional embodiment, before the step of performing a secondary fusion of the position features, appearance features, spatiotemporal features, and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video, the above method further includes: determining a bounding box corresponding to the target object in each frame image in the initial synthetic video, encoding the bounding box, and obtaining target position features; wherein the bounding box is used to indicate the position of the target object in each frame image.

[0167] In an optional embodiment, the above-mentioned step of preliminarily fusing the appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video includes: preliminarily fusing the appearance features, spatiotemporal features, position features and posture features through a first model to obtain an initial synthetic video; and performing a second fusing of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain a final synthetic video includes: performing a second fusing of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video through a second model to obtain a final synthetic video.

[0168] In an optional embodiment, the first model and the second model have the same model structure; wherein the second model is obtained by training based on the video obtained by the first model.

[0169] Furthermore, Figure 4 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 101 , the communication interface 103 and the memory 100 are connected via the bus 102 .

[0170] The memory 100 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0171] The processor 101 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 101 or by instructions in the form of software. The above-mentioned processor 101 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register. The storage medium is located in the memory 100, and the processor 101 reads the information in the memory 100 and, in conjunction with its hardware, completes the steps of the method of the aforementioned embodiment.

[0172] The embodiment of the present disclosure also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the above-mentioned video synthesis method. The specific implementation can be found in the method embodiment, which will not be repeated here.

[0173] Specifically, the above-mentioned video synthesis method includes: obtaining a reference image and a driving video; wherein the reference image includes a target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image; wherein the spatiotemporal features are used to indicate the action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame image, and the posture information is used to indicate the action of the replacement object in each frame image; extracting the appearance features of the target object from the reference image; fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame image to obtain a synthesized video; wherein the synthesized video is a video in which the replacement object in the driving video is replaced with the target object.

[0174] The above-mentioned video synthesis method only requires a reference image to replace the object in the driving video with the target object in the reference image while maintaining a high level of appearance consistency. At the same time, by extracting posture information, segmentation masks and spatiotemporal features from the driving video, it can accurately transfer the action in the driving video to the target object and naturally integrate the target object into the scene of the driving video, thereby improving the quality of video synthesis.

[0175] In an optional embodiment, the above-mentioned steps of extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame image, include: extracting the spatiotemporal features corresponding to the replacement object from the driving video through a first neural network; for each frame image in the driving video, performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image; extracting multiple posture key point coordinates of the replacement object from the current frame image, and determining the multiple posture key point coordinates as the posture information corresponding to the current frame image.

[0176] In an optional embodiment, the above-mentioned step of performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image includes: inputting the current frame image into a pre-trained instance segmentation model; identifying the outline of the replacement object in the current frame image through the instance segmentation model, and outputting a segmentation mask based on the outline; wherein the pixel value of the image area containing the replacement object in the segmentation mask is a first value, and the pixel value of the image area not containing the replacement object is a second value.

[0177] In an optional embodiment, the above method further includes: performing data enhancement processing on the segmentation mask to update the segmentation mask; wherein the data enhancement processing includes at least one of the following: image random dilation processing, image erosion processing, and image Gaussian blur processing, etc.

[0178] In an optional embodiment, the step of extracting appearance features of the target object from the reference image includes: extracting appearance features corresponding to the target object from the reference image through a second neural network.

[0179] In an optional embodiment, the above-mentioned step of fusing the appearance features, spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain a synthetic video includes: encoding the segmentation mask and posture information corresponding to each frame of the image respectively to obtain the position features and posture features corresponding to the replacement object; preliminarily fusing the appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video; and performing a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain a final synthetic video.

[0180] In an optional embodiment, the above-mentioned step of preliminarily fusing appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a first fusion result; fusing the position features with the first fusion result to obtain a second fusion result; fusing the posture features with the second fusion result to obtain a third fusion result; and denoising the third fusion result to obtain an initial synthetic video.

[0181] In an optional embodiment, the above-mentioned step of performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features includes: at each noise step, determining the noise variance according to a preset noise scheduling strategy, generating a Gaussian noise tensor with the same size as the spatiotemporal features based on the noise variance, and adding the Gaussian noise tensor to the spatiotemporal features to obtain the noisy spatiotemporal features.

[0182] In an optional embodiment, the above-mentioned step of fusing the position feature with the first fusion result to obtain the second fusion result includes: obtaining the time dependency between different image frames in the driving video through a time attention module; based on the time dependency, fusing the position feature with the first fusion result to obtain the second fusion result.

[0183] In an optional embodiment, the above-mentioned step of performing a secondary fusion of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a fourth fusion result; fusing the position features of the target object determined based on the initial synthetic video with the fourth fusion result to obtain a fifth fusion result; fusing the posture features with the fifth fusion result to obtain a sixth fusion result; and denoising the sixth fusion result to obtain the final synthetic video.

[0184] In an optional embodiment, before the step of performing a secondary fusion of the position features, appearance features, spatiotemporal features, and posture features of the target object determined based on the initial synthetic video to obtain the final synthetic video, the above method further includes: determining a bounding box corresponding to the target object in each frame image in the initial synthetic video, encoding the bounding box, and obtaining target position features; wherein the bounding box is used to indicate the position of the target object in each frame image.

[0185] In an optional embodiment, the above-mentioned step of preliminarily fusing the appearance features, spatiotemporal features, position features and posture features to obtain an initial synthetic video includes: preliminarily fusing the appearance features, spatiotemporal features, position features and posture features through a first model to obtain an initial synthetic video; and performing a second fusing of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video to obtain a final synthetic video includes: performing a second fusing of the position features, appearance features, spatiotemporal features and posture features of the target object determined based on the initial synthetic video through a second model to obtain a final synthetic video.

[0186] In an optional embodiment, the first model and the second model have the same model structure; wherein the second model is obtained by training based on the video obtained by the first model.

[0187] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0188] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate the description of this disclosure and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0189] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

1. A video synthesis method, characterized in that: The method comprises: Acquire a reference image and a driving video; wherein the reference image includes the target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; Extracting spatiotemporal features corresponding to the replacement object from the driving video, and extracting a segmentation mask and posture information corresponding to each frame of the image; wherein the spatiotemporal features are used to indicate action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame of the image, and the posture information is used to indicate the action of the replacement object in each frame of the image; extracting appearance features of the target object from the reference image; The appearance features, the spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image are fused to obtain a synthetic video; wherein the synthetic video is a video in which the replacement object in the driving video is replaced by the target object.

2. The method according to claim 1, characterized in that The step of extracting the spatiotemporal features corresponding to the replacement object from the driving video, and extracting the segmentation mask and posture information corresponding to each frame of the image, comprises: extracting spatiotemporal features corresponding to the replacement object from the driving video using a first neural network; For each frame image in the driving video, mask segmentation processing is performed on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image; multiple posture key point coordinates of the replacement object are extracted from the current frame image, and the multiple posture key point coordinates are determined as the posture information corresponding to the current frame image.

3. The method according to claim 2, characterized in that The step of performing mask segmentation processing on the replacement object in the current frame image to obtain a segmentation mask corresponding to the current frame image includes: Inputting the current frame image into a pre-trained instance segmentation model; The contour of the replacement object in the current frame image is identified by the instance segmentation model, and a segmentation mask is output based on the contour; wherein the pixel value of the image area containing the replacement object in the segmentation mask is a first value, and the pixel value of the image area not containing the replacement object is a second value.

4. The method according to claim 3, characterized in that The method further comprises: Performing data enhancement processing on the segmentation mask to update the segmentation mask; wherein the data enhancement processing includes at least one of the following: image random dilation processing, image erosion processing, image Gaussian blur processing, etc.

5. The method according to claim 1, wherein The step of extracting the appearance features of the target object from the reference image comprises: The appearance features corresponding to the target object are extracted from the reference image through a second neural network.

6. The method according to claim 1, characterized in that The step of fusing the appearance features, the spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain a synthesized video includes: Encoding the segmentation mask and posture information corresponding to each frame of image to obtain position features and posture features corresponding to the replacement object; Preliminarily fusing the appearance features, the spatiotemporal features, the position features, and the posture features to obtain an initial synthesized video; The position features, the appearance features, the spatiotemporal features and the posture features of the target object determined based on the initial synthesized video are secondary fused to obtain a final synthesized video.

7. The method according to claim 6, characterized in that The step of preliminarily fusing the appearance features, the spatiotemporal features, the position features, and the posture features to obtain an initial synthesized video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a first fusion result; fusing the position feature with the first fusion result to obtain a second fusion result; fusing the posture feature with the second fusion result to obtain a third fusion result; The third fusion result is subjected to denoising processing to obtain an initial synthesized video.

8. The method according to claim 7, characterized in that The step of performing noise processing on the spatiotemporal features to obtain the noisy spatiotemporal features includes: At each noise step, the noise variance is determined according to a preset noise scheduling strategy, a Gaussian noise tensor with the same size as the spatiotemporal feature is generated based on the noise variance, and the Gaussian noise tensor is added to the spatiotemporal feature to obtain a noisy spatiotemporal feature.

9. The method according to claim 7, characterized in that The step of fusing the position feature with the first fusion result to obtain a second fusion result includes: Acquire the temporal dependency between different image frames in the driving video through a temporal attention module; Based on the time dependency, the position feature is fused with the first fusion result to obtain a second fusion result.

10. The method according to claim 6, characterized in that The step of performing secondary fusion of the position features, the appearance features, the spatiotemporal features, and the posture features of the target object determined based on the initial synthesized video to obtain a final synthesized video includes: performing noise processing on the spatiotemporal features to obtain noisy spatiotemporal features; fusing the noisy spatiotemporal features with the appearance features to obtain a fourth fusion result; fusing the position feature of the target object determined based on the initial synthesized video with the fourth fusion result to obtain a fifth fusion result; fusing the posture feature with the fifth fusion result to obtain a sixth fusion result; The sixth fusion result is subjected to denoising processing to obtain a final composite video.

11. The method according to claim 10, characterized in that Before the step of performing secondary fusion of the position features, the appearance features, the spatiotemporal features, and the posture features of the target object determined based on the initial synthesized video to obtain a final synthesized video, the method further includes: Determine a bounding box corresponding to the target object in each frame of the initial synthesized video, encode the bounding box, and obtain a target position feature; wherein the bounding box is used to indicate the position of the target object in each frame.

12. The method according to claim 6, characterized in that The step of preliminarily fusing the appearance features, the spatiotemporal features, the position features, and the posture features to obtain an initial synthesized video includes: Preliminarily fusing the appearance features, the spatiotemporal features, the position features, and the posture features through a first model to obtain an initial synthesized video; The step of performing secondary fusion of the position features, the appearance features, the spatiotemporal features, and the posture features of the target object determined based on the initial synthesized video to obtain a final synthesized video includes: The position features, the appearance features, the spatiotemporal features and the posture features of the target object determined based on the initial synthesized video are secondary fused through a second model to obtain a final synthesized video.

13. The method according to claim 12, characterized in that The first model and the second model have the same model structure; wherein, the second model is trained based on the video obtained by the first model.

14. A video synthesis device, characterized in that: The device comprises: A data acquisition module is used to acquire a reference image and a driving video; wherein the reference image includes the target object, and the driving video is a multi-frame image containing the action of the replacement object and the background scene; a feature extraction module, configured to extract spatiotemporal features corresponding to the replacement object from the driving video, and extract a segmentation mask and posture information corresponding to each frame of the image; wherein the spatiotemporal features are used to indicate action features and scene features corresponding to the replacement object, the segmentation mask is used to indicate the position of the replacement object in each frame of the image, and the posture information is used to indicate the action of the replacement object in each frame of the image; An appearance extraction module, configured to extract appearance features of the target object from the reference image; A feature fusion module is used to fuse the appearance features, the spatiotemporal features, and the segmentation mask and posture information corresponding to each frame of the image to obtain a synthetic video; wherein the synthetic video is a video in which the replacement object in the driving video is replaced by the target object.

15. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the video synthesis method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the video synthesis method described in any one of items 1 to 13.

Citation Information

Cited By

  • Action migration method and device, electronic equipment, computer readable storage medium and program product

    CN121354212A