Video head changing method and device, electronic equipment and storage medium
By preprocessing the source character image and using the fine-tuned head generation model to generate a head exchange data set, combined with the statistics-based extended human body three-dimensional model to generate 4D neural Gaussian feature field and video image front-to-back scene fusion re-rendering technology, the target video is replaced, which solves the problem of unnatural video head change effect and insufficient time consistency in the existing technology, and achieves a high-quality, natural and consistent video head change effect.
Patent Information
- Application Number
- CN202510042579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing video head change technology is difficult to achieve high-quality, natural and consistent head change in dynamic videos, especially in terms of head-to-background matching and time consistency.
By preprocessing the source character image, the fine-tuned pre-trained head generation model is used to generate the head exchange data set, and a 4D neural Gaussian feature field is generated based on the statistically extended human body three-dimensional model. Combined with the front and back scene fusion re-rendering technology of the video image, the head change operation is performed on the target video.
It improves the adaptability and accuracy of the video head change technology, achieves a more natural head-body relationship and seamless front and back scene splicing, and improves the overall smoothness and naturalness of the video.
Smart Images

Figure CN119991462A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision, graphics and deep learning technology, and in particular to a video head-changing method and device, an electronic device and a storage medium. Background Art
[0002] Head-changing technology has very broad application prospects in many fields such as filmmaking, artistic creation, augmented reality (AR) and virtual reality (VR). With the continuous advancement of technology, especially the innovation in the fields of computer vision and deep learning, head-changing technology is no longer limited to traditional static image processing, and more and more research has begun to explore how to achieve head-changing effects in dynamic videos. However, how to seamlessly transfer the identity information of the source person to the target video without changing other attributes of the target video (such as posture, expression, background, etc.) is still a challenging research problem. This is not only a technical problem, but also a complex subject involving multiple disciplines such as deep learning, computer graphics and human-computer interaction.
[0003] Early video head-changing technology solutions mainly used the generative adversarial network (GAN) model. Although some good results were achieved, it was still difficult to ensure the high fidelity of identity features when faced with difficult head-changing tasks. Some other head-changing technology solutions, such as those based on diffusion models, often cannot guarantee time consistency, resulting in insufficient smoothness in facial switching in the video. In addition, other video head-changing solutions in the prior art also have problems such as poor facial imaging, unnatural head-changing effects, and easy discontinuity. Summary of the invention
[0004] In view of the above problems, the present invention provides a method and device for improving the video head-changing effect, an electronic device and a storage medium.
[0005] According to a first aspect of the present invention, there is provided a video head-changing method, comprising:
[0006] The background and clothing in the source person image are separated to obtain the preprocessed source person image, and the head area of the target video is redrawn using the fine-tuned pre-trained head generation model to obtain the head swap dataset;
[0007] A 4D neural Gaussian feature field is generated by using a statistically based extended human 3D model, and the 4D neural Gaussian feature field is trained using a head swap dataset to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation;
[0008] The foreground head area of the target video is processed to obtain a repaired target video, and based on the foreground and background fusion re-rendering technology of the video image, the trained 4D neural Gaussian feature field is used to perform a head-changing operation on the repaired target video to obtain the head-changing result of the repaired target video.
[0009] According to an embodiment of the present invention, the foreground header area of the target video is processed to obtain a restored target video including:
[0010] Removing the foreground header area of the target video to obtain the target video with the foreground header area removed;
[0011] The target video with the foreground head area removed is repaired using a predefined image repair tool to obtain a repaired target video, wherein the repaired target video includes a background area and a torso area of the target person.
[0012] According to an embodiment of the present invention, the above-mentioned predefined image restoration tool includes an image segmentation model, an image repair model and an artificial intelligence-based content generation model.
[0013] According to an embodiment of the present invention, the foreground and background fusion re-rendering technology based on video images uses the trained 4D neural Gaussian feature field to perform a head-changing operation on the repaired target video, and the head-changing result of the repaired target video includes:
[0014] The repaired target video is converted into a feature domain using a 2D background encoder to obtain a target video feature domain;
[0015] Perform two-dimensional splashing on the trained 4D neural Gaussian feature field to obtain the foreground head region;
[0016] Process the trained 4D neural Gaussian feature field to obtain the rendered alpha channel;
[0017] In the process of head reconstruction using the trained 4D neural Gaussian feature field, multiple key points of the character's face are deformed to obtain a head region mask;
[0018] Based on the overlapping area between the rendered alpha channel and the head area mask, the foreground head area and the background area are fused in the target video feature domain to obtain a fused feature map;
[0019] The fused feature map is converted into RGB domain using a 2D neural re-renderer to obtain the head-changing result of the repaired target video.
[0020] According to an embodiment of the present invention, the above-mentioned separation process of the background and clothing in the source person image to obtain the pre-processed source person image, and the head region of the target video is redrawn using the fine-tuned pre-trained head generation model to obtain the head exchange data set, which includes:
[0021] The background and clothing in the source person image are separated by cropping the foreground person region in the source person image to obtain a cropped source person image;
[0022] The cropped source person image is repaired using the pre-trained head generation model to obtain the source person image without clothing information;
[0023] Based on a predefined model fine-tuning method, the parameters of the pre-trained head generation model are fine-tuned using the source person image without clothing information to obtain a fine-tuned pre-trained head generation model;
[0024] The key point information injection method of the neural network is controlled based on predefined conditions, and the generation of head posture and expression is controlled by using the fine-tuned pre-trained head generation model. The head area of the target video is repaired to obtain the head swap dataset.
[0025] According to an embodiment of the present invention, the above-mentioned generation of a 4D neural Gaussian feature field using a statistically-based extended human body three-dimensional model includes:
[0026] A unified 3D human body surface is obtained by using a statistically extended 3D human body model, and a 3D neural Gaussian field is constructed in the unified 2D texture coordinate space of the 3D human body surface.
[0027] The learnable vertex offsets are introduced into the 3D neural Gaussian field to obtain the 4D neural Gaussian feature field.
[0028] According to an embodiment of the present invention, the 4D neural Gaussian feature field is trained using the head exchange data set to obtain the trained 4D neural Gaussian feature field, including:
[0029] The head swap dataset is used to track and adjust the parameters of the 4D neural Gaussian feature field. Based on the dynamic scene modeling technology, the temporal consistency of the 4D neural Gaussian feature field in the image generation process is enhanced by introducing temporal conditional features to obtain the trained 4D neural Gaussian feature field.
[0030] According to a second aspect of the present invention, there is provided a video head-changing device, comprising:
[0031] The image preprocessing and data generation module is used to obtain a preprocessed source person image by separating the background and clothing in the source person image, and redraw the head area of the target video using the fine-tuned pre-trained head generation model to obtain a head swap dataset;
[0032] A feature field generation and model training module is used to generate a 4D neural Gaussian feature field using a statistically based extended human 3D model, and train the 4D neural Gaussian feature field using a head swap dataset to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation;
[0033] The video restoration and video head-changing module is used to process the foreground head area of the target video to obtain the restored target video, and based on the foreground and background fusion re-rendering technology of the video image, use the trained 4D neural Gaussian feature field to perform a head-changing operation on the restored target video to obtain the head-changing result of the restored target video.
[0034] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0035] The fourth aspect of the present invention further provides a computer-readable storage medium on which a computer program or instruction is stored, and the steps of the above method are implemented when the above computer program or instruction is executed by a processor.
[0036] The video head-changing method provided by the present invention utilizes a fine-tuned pre-trained head generation model so that the generated exchange data can have better identity similarity and skin color consistency, greatly improving the adaptability and accuracy of the video head-changing technology and being able to more accurately adapt to the characteristics of the source head domain; at the same time, the present invention can make full use of geometric relationships through the generated 4D neural Gaussian feature field, and upgrade the traditional 2D portrait video to a 4D neural Gaussian feature field, so that the generated video head-changing result has better geometric consistency and inter-frame consistency, and can more naturally handle the connection between the head and neck, thereby making the head-changing effect more realistic and stable; in addition, based on the foreground and background fusion re-rendering technology, the present invention enables the video head-changing structure to have a more natural head-to-body relationship and a seamless foreground and background splicing relationship, thereby solving the problem of foreground and background incoordination that may occur during the video head-changing process, thereby improving the overall fluency and naturalness of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0038] Figure 1 is an application scenario diagram of a video head-changing method according to an embodiment of the present invention;
[0039] Figure 2 is a flow chart of a video head-changing method according to an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of a fine-tuning process of a pre-trained head generation model and a process of generating an exchange data set according to an embodiment of the present invention;
[0041] Figure 4 is a framework diagram of a video head-changing method according to an embodiment of the present invention;
[0042] Figure 5 is a structural block diagram of a video head-changing device according to an embodiment of the present invention;
[0043] Figure 6 4 is a block diagram of an electronic device suitable for implementing a video head-changing method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.
[0045] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0046] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0047] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0048] In the head-swapping task, the transfer of facial features between the source video and the target video is often not a simple matter. With the deepening of research in the field of head-swapping, face-swapping technology has gradually become a hot topic in computer vision. Early methods mainly used the generative adversarial network (GAN) model to fuse the identity features of the source face with other features of the target face through adversarial training to ensure the visual authenticity of the generated face. Although these methods have achieved good results in many tasks, they also have some inherent limitations. For example, the representation ability of the GAN model is limited, especially in complex facial details and expression changes, and it often fails to achieve the desired effect. This makes early face-swapping methods often unable to cope with the difficult head-swapping tasks and it is difficult to ensure the high fidelity of identity features.
[0049] In recent years, with the emergence of diffusion models, face-swapping technology has been further improved. Some new studies have begun to use diffusion models for face-swapping. Compared with traditional GAN models, diffusion models can capture the details in the image more delicately during the generation process and have better results. However, although diffusion models perform very well on single images, they still face great challenges when processing dynamic videos. Due to the randomness of the diffusion process, these methods often cannot maintain temporal consistency in video sequences, resulting in unclear face switching in the video, and even abrupt changes. Therefore, how to achieve high-quality head-swapping in video sequences is still a major problem facing diffusion models.
[0050] In addition, some studies have also tried to use 3D deformable models (3DMM) to perform face swapping, aiming to enhance the expressiveness of the model by modeling the three-dimensional geometric shape of the face. Some methods use 3DMM to generate more realistic face swapping effects, especially when faced with complex facial expressions and angle changes. These methods can better handle changes in facial geometry. However, 3DMM is limited to modeling the facial area, so when the task requires a full-head face swap, the application of 3DMM is inadequate. Changes in other areas of the head, such as hair, neck and other structures, are still problems that current 3DMM technology cannot effectively solve.
[0051] Compared with face swapping, head swapping is more difficult because it requires not only visually maintaining a high degree of consistency in facial features and expressions in the target video, but also capturing structural details of the source character's head, hair, neck and other areas. In particular, traditional methods often have difficulty in dealing with the problem of regional matching between the head and the background. This is because there are often certain spatial structural differences between the head and the background, and how to achieve seamless connection has become a core issue in the head swapping task. In response to these challenges, the latest method proposed a method of fusing the segmentation results of the reenacted source and the target video to achieve head switching between the source and target videos. Although this method solves the matching problem of the head and the background to a certain extent, it still has the problem of poor effect when the posture changes greatly, and due to the lack of 3D prior knowledge, it causes problems in temporal consistency. This makes the head swapping effect unsatisfactory in dynamic videos, and it is easy to have a sense of discontinuity and unnatural effects.
[0052] In addition, due to the limited capacity of the GAN model, although this method achieves a preliminary head-changing effect, the generated results often have certain limitations. Specifically, due to the complexity of the facial and head structure, as well as the high variability in the target video, the existing GAN model often has difficulty maintaining the consistency of the head with the target video, resulting in the lack of sufficient naturalness and details in the generated results. This makes the high-quality head-changing task still a huge challenge in the field of computer vision.
[0053] Therefore, although existing methods have promoted the progress of head-changing technology to varying degrees, how to achieve high-quality, natural and consistent head-changing in dynamic videos is still an important issue in this field. Future research may need to combine more advanced 3D modeling technology, extended spatiotemporal consistency models, and more efficient generation models to cope with various complex situations in head-changing tasks, thereby promoting the widespread application of this technology in film, VR, AR and other fields.
[0054] It should be particularly noted that in the technical solution of the present invention, the user information (including but not limited to user personal information, user image information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0055] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided by the embodiments of the present invention provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs, and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.
[0056] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned, which should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0057] An embodiment of the present invention provides a video head-changing method, which is used to solve the technical problems of the existing video head-changing method, such as poor image quality, inconsistent head-changing results, and poor 3D consistency and identity consistency.
[0058] Figure 1 2 is an application scenario diagram of a video head-changing method according to an embodiment of the present invention.
[0059] like Figure 1 As shown, the application scenario 100 according to this embodiment may include technical fields such as computer vision, graphics, artificial intelligence, and deep fake. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0060] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0061] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0062] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0063] It should be noted that the video head-changing method provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the video head-changing device provided in the embodiment of the present invention can generally be set in the server 105. The video head-changing method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the video head-changing device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0064] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.
[0065] The following will be based on Figure 1 The scene described by Figure 2~Figure 4 The video head-changing method of the disclosed embodiment is described in detail.
[0066] Figure 2 4 is a flowchart of a video head-changing method according to an embodiment of the present invention.
[0067] like Figure 2 As shown, the video head swapping of this embodiment includes operations S210 to S230.
[0068] In operation S210, the background and clothing in the source person image are separated to obtain a preprocessed source person image, and the head region of the target video is redrawn using the fine-tuned pre-trained head generation model to obtain a head swap dataset.
[0069] The source character image includes a source character portrait used to replace the head of the target character in the target video. After the source character authorizes, the source character image is processed, and the processing of the source character image is only used for the video head replacement method provided by the present invention. The relevant processing process strictly abides by the provisions of laws and regulations, and takes strict confidentiality measures and anti-abuse measures.
[0070] In addition, before performing operation S210, an obvious and clear prompt message needs to be given to inform the owner or source person of the source person image that this operation S210 needs to perform a video head-changing operation and related image processing operations, and after obtaining the authorization or permission of the owner or source person of the source person image, operation S210 is performed. If the owner or source person of the source person image refuses to provide authorization or permission, operation S210 is not performed.
[0071] Before using the head generation model, it is necessary to pre-train the head generation model and fine-tune the parameters of the pre-trained head generation model to obtain a fine-tuned pre-trained head generation model, and use the fine-tuned model to generate a head swap dataset.
[0072] The above-mentioned head exchange dataset refers to a head sample dataset used for training, verification, and inference.
[0073] By redrawing the head region of the target video using the fine-tuned pre-trained head generation model, the resulting head swap dataset can focus on the head and facial regions.
[0074] The above operation S210 adapts the pre-trained 2D head generation model to the source image domain by using a fine-tuning strategy. The above fine-tuning strategy enables the pre-trained 2D head generation model to better capture the characteristics of the source person's head through careful tuning. Subsequently, the present invention generates a set of high-quality exchange data sets using diffusion repair technology. These exchange data sets provide rich source images for model training and ensure that the present invention can perform video head swap tasks with high accuracy.
[0075] In operation S220, a 4D neural Gaussian feature field is generated using a statistically based extended human body three-dimensional model, and the 4D neural Gaussian feature field is trained using a head exchange dataset to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation.
[0076] The above-mentioned statistically based extended human 3D models include SMPL-X (Skinned Multi-Person LinearModel-X), which can be used to reconstruct 3D posture and shape from 2D images.
[0077] In order to ensure 3D consistency and temporal consistency during the head replacement process and to model the relationship between the head and neck more naturally, the present invention introduces an intrinsic 3D Gaussian feature field and embeds it into the complete human SMPL-X surface model. Through this innovative technical means, the present invention can accurately capture the spatial structure of the source character's head and maintain the coordination of the head with the body and background.
[0078] In operation S230, the foreground head area of the target video is processed to obtain a repaired target video, and based on the foreground and background fusion re-rendering technology of the video image, a head replacement operation is performed on the repaired target video using the trained 4D neural Gaussian feature field to obtain a head replacement result of the repaired target video.
[0079] The target video includes a portrait of the target person whose head is replaced. After the target person, the target video is processed, and the processing of the target video is only used for the video head replacement method provided by the present invention. The relevant processing process strictly abides by the provisions of laws and regulations, and takes strict confidentiality measures and anti-abuse measures.
[0080] In addition, before performing operation S230, an obvious and clear prompt message needs to be given to inform the owner of the target video or the target person in the target video that this operation S230 requires a video head-changing operation and related image processing, and operation S230 is performed after obtaining authorization or permission from the owner of the target video or the target person in the target video. If the owner of the target video or the target person in the target video refuses to provide authorization or permission, operation S230 is not performed.
[0081] The foreground feature map generated by Gaussian scattering is fused with the background feature map. This process can effectively enhance the visual consistency between the head and the background. On this basis, the present invention uses a neural re-rendering technology to seamlessly fuse the two feature maps, thereby achieving a more natural and smooth head-changing effect.
[0082] From the above operations S210 to S230, it can be seen that the present invention can maintain a high identity similarity with the source image by utilizing the expressive power of the diffusion model (i.e., the pre-trained head generation model, the same below); at the same time, the portrait representation is promoted to the 4D space, so that the result has better 3D consistency and time consistency; in addition, by utilizing the re-rendering module on the latent space feature map, it can have a natural head and torso relationship and can naturally fill the mismatched areas.
[0083] The video head-changing method provided by the present invention utilizes a fine-tuned pre-trained head generation model so that the generated exchange data can have better identity similarity and skin color consistency, greatly improving the adaptability and accuracy of the video head-changing technology and being able to more accurately adapt to the characteristics of the source head domain; at the same time, the present invention can make full use of geometric relationships through the generated 4D neural Gaussian feature field, and upgrade the traditional 2D portrait video to a 4D neural Gaussian feature field, so that the generated video head-changing result has better geometric consistency and inter-frame consistency, and can more naturally handle the connection between the head and neck, thereby making the head-changing effect more realistic and stable; in addition, based on the foreground and background fusion re-rendering technology, the present invention enables the video head-changing structure to have a more natural head-to-body relationship and a seamless foreground and background splicing relationship, thereby solving the problem of foreground and background incoordination that may occur during the video head-changing process, thereby improving the overall fluency and naturalness of the video.
[0084] The above operations S210 to S230 are further described in detail below through specific embodiments and in conjunction with the accompanying drawings.
[0085] The present invention provides a real video head-changing technical solution, which mainly solves the various defects of previous video head-changing algorithms. For example, previous studies on head-changing through 3D deformable models (3DMM) can handle facial geometric deformation, but are limited to the facial area and cannot effectively handle changes in other head areas such as hair and neck. Compared with facial head-changing, the head-changing task is more challenging. In addition to maintaining the consistency of facial features, it is also necessary to capture details such as the head, hair, and neck, and solve the problem of matching the head with the background. Other existing technologies, such as using GAN to fuse the segmentation results of the reenacted source and the target video, solve the background matching to a certain extent, but still have shortcomings in large posture changes and time consistency. In addition, the representation ability of the existing GAN model is limited, and the generated results often lack naturalness and details, making it difficult to cope with highly changing dynamic videos. To this end, the present invention provides a video head-changing method that has obvious advantages in terms of the quality, naturalness, 3D consistency, and identity retention of the generated results.
[0086] According to an embodiment of the present invention, the above-mentioned process of separating the background and clothing in the source person image to obtain a preprocessed source person image, and using the fine-tuned pre-trained head generation model to repair the head area of the preprocessed source person image to obtain a head exchange data set includes: separating the background and clothing in the source person image by cropping the foreground person area in the source person image to obtain a cropped source person image; using the pre-trained head generation model to repair the cropped source person image to obtain a source person image with clothing information removed; based on a predefined model fine-tuning method, using the source person image with clothing information removed to fine-tune the parameters of the pre-trained head generation model to obtain a fine-tuned pre-trained head generation model; based on predefined conditions, controlling the key point information injection method of the neural network, using the fine-tuned pre-trained head generation model to control the generation of head posture and expression, and repairing the head area of the target video to obtain a head exchange data set.
[0087] The above-mentioned pre-trained head generation model includes the Arc2Face model; the predefined model fine-tuning method includes the Dreambooth model adjustment method based on the generative adversarial network; and the predefined conditional control neural network includes ControlNet.
[0088] The above Arc2Face is an identity-conditioned face model that can generate diverse, photo-realistic images based on a person's ArcFace embedding, and these images far exceed existing models in terms of face similarity. The core advantage of Arc2Face is that it can generate face images that are highly consistent with the input ArcFace embedding, which is particularly important for application scenarios that require high identity consistency.
[0089] The above-mentioned Dreambooth is a personalized generative model technology that is mainly used to fine-tune the pre-trained generative model through a small number of pictures to generate images with specific attributes or styles. The core idea of Dreambooth is to retrain part of the generative model with a small amount of personalized data so that it can generate images that are highly similar to the input data.
[0090] The above-mentioned ControlNet is a plug-in based on the Stable Diffusion function, which aims to improve the controllability and accuracy of AI image generation by adding additional control conditions to guide image generation. It can accurately guide the image generation process by using edge features, depth features, or skeletal features of human posture in the input image.
[0091] Figure 3 2 is a schematic diagram of a fine-tuning process of a pre-trained head generation model and a process of generating an exchange data set according to an embodiment of the present invention.
[0092] like Figure 3 As shown in the figure, the present invention uses a 2D head generation model (i.e., head generation model, the same below) to repair the head area in the target video, thereby constructing a dataset for head replacement. Although existing human head generation models perform well in head generation tasks that maintain identity consistency, they often have deficiencies in retaining the details and personalized features of the source person's head. In order to solve this problem, the present invention fine-tunes the pre-trained head generation model: Arc2Face, and uses a small number of source images to enhance identity similarity.
[0093] Before fine-tuning, the present invention first needs to separate the background and clothing information from the source image. Specifically, the present invention crops the foreground character area in the source image and superimposes it on the background of the target video. At the same time, the present invention uses a diffusion model to repair the cropped image to remove the clothing information. These steps are very important because they can effectively avoid the leakage of irrelevant information (such as background and clothing details), thereby ensuring that the fine-tuning process focuses on the face and head areas. The present invention fine-tunes the portrait generation model with reference to the Dreambooth method, and uses ControlNet to inject key point information to accurately control the posture and expression of the head. Finally, the present invention uses the fine-tuned model to repair the head area in the target video to obtain a head swap dataset. The mask of the head area is based on the head mask in the rough reproduction result, and is obtained by deforming five facial key points.
[0094] According to an embodiment of the present invention, the above-mentioned generation of a 4D neural Gaussian feature field using a statistically based extended human body three-dimensional model includes: obtaining a unified human body three-dimensional surface using a statistically based extended human body three-dimensional model, and constructing a 3D neural Gaussian field in the unified human body three-dimensional surface two-dimensional texture coordinate space; introducing a learnable vertex offset into the 3D neural Gaussian field to obtain a 4D neural Gaussian feature field.
[0095] According to an embodiment of the present invention, the above-mentioned training of the 4D neural Gaussian feature field using the head exchange dataset to obtain the trained 4D neural Gaussian feature field includes: using the head exchange dataset to track and adjust the parameters of the 4D neural Gaussian feature field, and based on the dynamic scene modeling technology, by introducing time conditional features to enhance the time consistency of the 4D neural Gaussian feature field in the image generation process, so as to obtain the trained 4D neural Gaussian feature field.
[0096] Since the postures and expressions in the training data generated by the fine-tuned diffusion model (i.e., the fine-tuned pre-trained head generation model) are slightly different from the target video in some cases, showing a certain randomness, this may lead to minor inconsistencies in temporal consistency. To solve this problem, the present invention re-extracts the SMPL-X parameters in the training data and introduces time-based conditional features.
[0097] Figure 4 4 is a framework diagram of a video head-changing method according to an embodiment of the present invention.
[0098] The present invention uses the generated data frames to supervise the 4D neural Gaussian portrait model (i.e., 4D neural Gaussian feature field, the same below). In order to deal with the inconsistency of generating training data frames, the present invention adopts SMPL-X re-tracking technology. In addition, the present invention adds temporal conditional features to the neural Gaussian texture mechanism to effectively deal with the inconsistency between different frames.
[0099] Due to the randomness of the diffusion model itself, the generated head-changing frames are usually inconsistent, especially in terms of facial expressions and head postures. This inconsistency may cause large differences between the generated head-changing images and the target video. Especially in the process of dynamic video generation, each generated frame is affected by the randomness of the model, which may lead to unstable performance between different frames, thereby affecting the coherence and naturalness of the entire video sequence. In order to effectively solve this problem, the present invention draws on the previous 3D Gaussian portrait modeling method and proposes an innovative solution, which converts the knowledge of the fine-tuned diffusion model into an efficient 3D representation method, thereby significantly enhancing the consistency and stability of the generation process.
[0100] Specifically, if Figure 4 As shown, the present invention constructs a 3D Gaussian field (i.e., 3D neural Gaussian field, the same below) in the UV space of the SMPL-X surface, and introduces a learnable vertex offset so as to dynamically adjust and deform the Gaussian body according to the deformation of the underlying grid in the input video. With this method, the present invention can more accurately align the morphology of the Gaussian body with the posture and expression of the character in the target video, thereby achieving a more natural and realistic head-changing effect. In addition, by embedding the 3D Gaussian field onto the surface, the present invention can use the shape, posture, and expression parameters of the SMPL-X model to efficiently transform the Gaussian field, ensuring that the generated image can be consistent with the dynamics of the target video in various dimensions such as shape, expression, and posture.
[0101] In this process, the present invention not only realizes the deformation of the Gaussian body, but also stores learnable features for each Gaussian body so that these features can be continuously optimized during the generation process, thereby improving the quality and naturalness of the generated image. In order to further refine this process, the present invention establishes a Neural Gaussian Field in UV space, in which each pixel is characterized by four key attributes: neural features, opacity, scale, and rotation. These attributes enable each pixel to maintain consistency under different transformations while precisely controlling the details. Through UV mapping technology, the present invention can accurately convert the neural Gaussian body from two-dimensional UV space to three-dimensional space, thereby completing the spatial transformation in the entire head-changing process and ensuring that the generated image can be naturally embedded in the target scene in three-dimensional space.
[0102] Although the present invention has used ControlNet to inject key point information into the generation process to help better adjust the posture and expression in the head-changing image, in some cases, there are still certain expression and posture misalignment problems between the generated image and the target video. This problem is mainly manifested in the fact that there are subtle differences in facial expressions, head postures, etc. between the generated head-changing image and the characters in the target video, resulting in imperfect generation effects. In order to effectively alleviate this problem, the present invention re-tracks and adjusts the SMPL-X parameters of the generated image during the training phase. By accurately tracking and optimizing the SMPL-X parameters of the generated image, the present invention can better control the posture and expression in the generated image, thereby reducing the inconsistency problems caused by differences in posture or expression. In the reasoning stage, the present invention uses the SMPL-X parameters tracked from the original target video to further optimize the posture and expression of the head during the generation process, so that the generated head-changing image can match the dynamics of the target video more closely.
[0103] In order to further solve the inconsistency problem between different frames in the generated data set, the present invention draws on the latest dynamic scene modeling technology and enhances the temporal consistency of the generated images by introducing temporal conditional features. Specifically, in the training stage, the present invention stores a set of learnable temporal features and broadcasts these features to each Gaussian in all frames, thereby ensuring that the generation of each frame can maintain temporal consistency. This method effectively avoids mutations or inconsistencies between different frames, so that the generated head-changing sequence remains consistent in time, improving the smoothness and naturalness of the video. In the reasoning stage, the present invention uses features to maintain the consistency of all frames, further enhancing temporal coherence, and ensuring that the generated video shows consistent head posture and expression in multiple consecutive frames, thereby significantly improving the stability and naturalness of the generated sequence.
[0104] According to an embodiment of the present invention, the above-mentioned processing of the foreground head area of the target video to obtain a repaired target video includes: removing the foreground head area of the target video to obtain a target video with the foreground head area removed; using a predefined image repair tool to repair the target video with the foreground head area removed to obtain a repaired target video, wherein the repaired target video includes a background area and a torso area of the target person.
[0105] According to an embodiment of the present invention, the above-mentioned predefined image restoration tool includes an image segmentation model, an image repair model and an artificial intelligence-based content generation model.
[0106] According to an embodiment of the present invention, the foreground and background fusion re-rendering technology based on video images uses the trained 4D neural Gaussian feature field to perform a head-changing operation on the repaired target video to obtain the head-changing result of the repaired target video, including: using a 2D background encoder to convert the repaired target video into a feature domain to obtain a target video feature domain; performing two-dimensional splashing on the trained 4D neural Gaussian feature field to obtain a foreground head area; processing the trained 4D neural Gaussian feature field to obtain a rendered alpha channel; deforming multiple key points of the character's face during head reproduction with the trained 4D neural Gaussian feature field to obtain a head area mask; based on the overlapping area between the rendered alpha channel and the head area mask, fusing the foreground head area and the background area in the target video feature domain to obtain a fused feature map; using a 2D neural re-renderer to convert the fused feature map into an RGB domain to obtain the head-changing result of the repaired target video.
[0107] After obtaining the Gaussian rendered feature map, the neural re-renderer of the present invention is used to seamlessly merge the scattered portrait features of the foreground with the background features, thereby completing a natural head-changing effect.
[0108] The 4D portrait representation method of the present invention can effectively model the head after the head swap, but reintegrating this 3D head into the original video still faces challenges. In particular, due to the differences in head shape and hairstyle, there is often a significant mismatch in the head region between the foreground and target images. Therefore, the present invention designs a neural re-rendering module to solve this problem.
[0109] The present invention first removes the foreground head region in the target image and inpaints it using Inpaint-Anything. Then, the inpainted image (containing only the background and torso) is processed through a 2D background encoder to convert it into the feature domain. Based on the overlapping area between the rendered alpha channel and the head mask, the present invention fuses the foreground head and background in the feature domain. Then, the fused feature map is processed through a 2D neural re-renderer to convert it into the RGB domain, and finally the natural fusion of the head is completed.
[0110] Compared with various video head-changing / face-changing technical solutions in the prior art, the present invention has the following advantages: The present invention proposes to use the source image to personalize and fine-tune the pre-trained head diffusion model, so that the generated training data can have better identity similarity and skin color consistency. The present invention proposes to use the re-tracking method for the generated data and the learnable time conditional features to achieve any reconstruction of consistent three-dimensional representations from inconsistent training data, which is reflected in the expression consistency with the target image in the video head-changing task. The feature map-based foreground and background fusion re-rendering technology proposed in the present invention makes the head-changing result have a more natural head-body relationship and a seamless foreground and background splicing relationship. The method of solving the head-changing task with a four-dimensional Gaussian model proposed in the present invention makes full use of geometric relationships, so that the present invention has better geometric consistency and inter-frame consistency. Due to the network structure design proposed in the present invention, due to the addition of learnable features of the foreground and background and the convolutional neural network, the high-frequency information of the training data is fully extracted, so that the present invention has high clarity.
[0111] Based on the above video head-changing method, the present invention also provides a video head-changing device. Figure 5 The device is described in detail.
[0112] Figure 5 4 is a structural block diagram of a video head-changing device according to an embodiment of the present invention.
[0113] like Figure 5 As shown, the video head-changing device 500 of this embodiment includes an image preprocessing and data generation module 510, a feature field generation and model training module 520, and a video restoration and video head-changing module 530.
[0114] The image preprocessing and data generation module 510 is used to obtain a preprocessed source character image by separating the background and clothing in the source character image, and redraw the head area of the target video using the fine-tuned pre-trained head generation model to obtain a head exchange data set; in one embodiment, the image preprocessing and data generation module 510 can be used to perform the operation S210 described above, which will not be repeated here.
[0115] The feature field generation and model training module 520 is used to generate a 4D neural Gaussian feature field using a statistically based extended human body three-dimensional model, and train the 4D neural Gaussian feature field using a head exchange data set to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation. In one embodiment, the feature field generation and model training module 520 can be used to perform the operation S220 described above, which will not be repeated here.
[0116] The video repair and video head-changing module 530 is used to process the foreground head area of the target video to obtain a repaired target video, and based on the foreground and background fusion re-rendering technology of the video image, use the trained 4D neural Gaussian feature field to perform a head-changing operation on the repaired target video to obtain a head-changing result of the repaired target video. In one embodiment, the video repair and video head-changing module 530 can be used to perform the operation S230 described above, which will not be repeated here.
[0117] According to an embodiment of the present invention, any multiple modules of the image preprocessing and data generation module 510, the feature field generation and model training module 520, and the video restoration and video head swap module 530 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the image preprocessing and data generation module 510, the feature field generation and model training module 520, and the video restoration and video head swap module 530 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of them. Alternatively, at least one of the image preprocessing and data generation module 510, the feature field generation and model training module 520, and the video restoration and video head replacement module 530 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is executed.
[0118] The present invention proposes a coherent and realistic video head-changing device based on 4D neural Gaussian portrait prior. The present invention innovatively embeds the intrinsic 3D Gaussian feature field into the whole-body SMPL-X surface, thereby upgrading the traditional 2D portrait video to a 4D neural Gaussian field. This effectively ensures the high consistency of the 3D structure and can naturally handle the connection between the head and the neck, making the head-changing effect more realistic and stable. When creating the training data set required for reconstruction, the present invention uses a small number of source images and fine-tunes the pre-trained 2D portrait generation model to accurately adapt to the characteristics of the source domain. This process greatly improves the adaptability and accuracy of the model. At the same time, the present invention introduces an advanced neural re-rendering strategy that can achieve seamless fusion of foreground and background. This strategy effectively eliminates the problem of foreground and background inconsistency that may occur during the head-changing process, thereby improving the overall fluency and naturalness of the video. Experimental results show that compared with the existing head-changing methods, the device of the present invention shows significant advantages in image quality, naturalness, 3D consistency and identity preservation, and can generate a more realistic, stable and highly consistent head-changing effect.
[0119] Figure 6 4 is a block diagram of an electronic device suitable for implementing a video head-changing method according to an embodiment of the present invention.
[0120] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 to a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0121] In RAM 603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM 602 and RAM 603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.
[0122] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.
[0123] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0124] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.
[0125] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.
[0127] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A video head-changing method, characterized in that: The method comprises: The background and clothing in the source person image are separated to obtain the preprocessed source person image, and the head area of the target video is redrawn using the fine-tuned pre-trained head generation model to obtain the head swap dataset; Generate a 4D neural Gaussian feature field using a statistically based extended human three-dimensional model, and train the 4D neural Gaussian feature field using the head exchange data set to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation; The foreground head area of the target video is processed to obtain a repaired target video, and based on the foreground and background fusion re-rendering technology of the video image, the trained 4D neural Gaussian feature field is used to perform a head replacement operation on the repaired target video to obtain a head replacement result of the repaired target video.
2. The method according to claim 1, characterized in that The foreground head area of the target video is processed, and the repaired target video includes: Removing the foreground header region of the target video to obtain the target video with the foreground header region removed; The target video with the foreground head area removed is repaired using a predefined image repair tool to obtain a repaired target video, wherein the repaired target video includes a background area and a torso area of the target person.
3. The method according to claim 2, characterized in that The predefined image restoration tools include an image segmentation model, an image repair model, and an artificial intelligence-based content generation model.
4. The method according to claim 2, characterized in that: Based on the foreground and background fusion re-rendering technology of the video image, the head-changing operation is performed on the repaired target video using the trained 4D neural Gaussian feature field, and the head-changing result of the repaired target video is obtained, including: Using a 2D background encoder to convert the restored target video into a feature domain to obtain a target video feature domain; Performing two-dimensional splashing on the trained 4D neural Gaussian feature field to obtain a foreground head region; Processing the trained 4D neural Gaussian feature field to obtain a rendered alpha channel; In the process of reproducing the head using the trained 4D neural Gaussian feature field, multiple key points on the face of the person are deformed to obtain a head region mask; Based on the overlapping area between the rendered alpha channel and the head area mask, fusing the foreground head area and the background area in the target video feature domain to obtain a fused feature map; The fused feature map is converted into an RGB domain using a 2D neural re-renderer to obtain a head-changing result of the repaired target video.
5. The method according to claim 1, characterized in that By separating the background and clothing in the source person image, the preprocessed source person image is obtained, and the head area of the target video is redrawn using the fine-tuned pre-trained head generation model. The head swap dataset includes: The background and clothing in the source person image are separated by cropping the foreground person region in the source person image to obtain a cropped source person image; Restoring the cropped source person image using the pre-trained head generation model to obtain the source person image without clothing information; Based on a predefined model fine-tuning method, fine-tuning parameters of the pre-trained head generation model using the source person image with clothing information removed, to obtain a fine-tuned pre-trained head generation model; The key point information injection method of the neural network is controlled based on predefined conditions, the fine-tuned pre-trained head generation model is used to control the generation of head posture and expression, and the head area of the target video is repaired to obtain the head exchange data set.
6. The method according to claim 1, characterized in that The 4D neural Gaussian feature field is generated using a statistically based extended human 3D model including: A unified three-dimensional human body surface is obtained by using a statistically extended three-dimensional human body model, and a 3D neural Gaussian field is constructed in the two-dimensional texture coordinate space of the unified three-dimensional human body surface; The learnable vertex offset is introduced into the 3D neural Gaussian field to obtain the 4D neural Gaussian feature field.
7. The method according to claim 6, characterized in that The 4D neural Gaussian feature field is trained using the head exchange data set, and the trained 4D neural Gaussian feature field is obtained, including: The head exchange dataset is used to track and adjust the parameters of the 4D neural Gaussian feature field, and based on the dynamic scene modeling technology, the time conditional features are introduced to enhance the time consistency of the 4D neural Gaussian feature field in the image generation process, so as to obtain the trained 4D neural Gaussian feature field.
8. A video head-changing device, characterized in that: The device comprises: The image preprocessing and data generation module is used to obtain a preprocessed source person image by separating the background and clothing in the source person image, and redraw the head area of the target video using the fine-tuned pre-trained head generation model to obtain a head swap dataset; A feature field generation and model training module, used to generate a 4D neural Gaussian feature field using a statistically based extended human three-dimensional model, and train the 4D neural Gaussian feature field using the head exchange data set to obtain a trained 4D neural Gaussian feature field, wherein the attributes of each pixel in the 4D neural Gaussian feature field include neural features, opacity, scale, and rotation; The video repair and video head-changing module is used to process the foreground head area of the target video to obtain a repaired target video, and based on the foreground and background fusion re-rendering technology of the video image, use the trained 4D neural Gaussian feature field to perform a head-changing operation on the repaired target video to obtain the head-changing result of the repaired target video.
9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Construction method and device of deformable neural radiation field network
CN115909015A
Virtual anchor whole-body video generation method and system based on diffusion model
CN117979115A
Head image generation method and system based on 3D Gaussian field
CN119006676A
Person replacement utilizing deferred neural rendering
US11582519B1