Character background editing method, device, storage medium and computer equipment
By training the target video to generate a large model and utilizing sample video data and portrait mask synthesis technology, the problems of high threshold and low fusion quality in existing technologies are solved, and a low-threshold, high-quality video character background replacement effect is achieved.
Patent Information
- Application Number
- CN202510855605.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing video character background replacement technology has a high threshold for use and is difficult to achieve high-quality fusion, especially when the proportions of the character and background are inconsistent or there is a difference in lighting.
By training the target video to generate a large model, using sample video data with people as sample labels, and combining portrait masks with sample background data without people, we can generate real and unmodified video data to replace the background data in the original video.
It lowers the threshold for video creation, improves the effect and authenticity of character and background editing, and generates more natural and realistic video data.
Smart Images

Figure CN120378699B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, storage medium and computer equipment for editing character backgrounds. Background Art
[0002] With the rapid development of artificial intelligence (AI), its application in video editing is becoming increasingly widespread, with more and more self-media creators and ordinary users experimenting with AI for video creation. Currently, replacing backgrounds with characters has become a common and frequent requirement in video creation.
[0003] The current mainstream video character and background replacement technologies can be roughly divided into two categories: traditional color domain cutouts and deep learning model cutouts. Traditional color domain cutouts have a high threshold for use, requiring users to have certain image processing capabilities, and there are many software parameters, which requires a lot of manpower and material resources to achieve cutouts. The character and background editing technology based on deep learning usually requires the use of a semantic segmentation model to identify and extract the character mask, and then use a generative adversarial network to merge the character with the new background, or directly merge the character with the new background at the pixel level without the help of a model. This type of technology often finds it difficult to achieve high-quality fusion between the character and the new background, especially when there are problems such as disharmony in proportion and lighting differences between the character and the target background. Existing algorithms often perform poorly and cannot meet the needs of complex scenes. Summary of the Invention
[0004] The purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defect that the video character background replacement technology in the existing technology cannot have both a low usage threshold and a good fusion effect.
[0005] This application provides a method for editing a character background, the method comprising:
[0006] Obtain original video data with people and target background data without people;
[0007] Determining a target video generation large model, wherein the target video generation large model is obtained by training using sample video data with people as sample labels, a sample reference background after removing the people from the sample video data, and sample reference data obtained by synthesizing the people mask extracted from the sample video data and the sample background data without people as training samples;
[0008] The target video generates a large model to replace the background data in the original video data with the target background data to obtain the target video data.
[0009] Optionally, the training process of generating a large model from the target video includes:
[0010] Obtain sample video data with people and sample background data without people;
[0011] Extracting a portrait mask from the sample video data, and eliminating the portrait mask from the sample video data to obtain a sample reference background after eliminating the portrait;
[0012] Synthesizing the portrait mask with the sample background data to obtain sample reference data;
[0013] Replacing the background data in the sample reference data with the sample reference background through a preset initial video generation model to obtain predicted video data;
[0014] With the goal of making the predicted video data approach the sample video data, training the initial video generation model until a preset end condition is reached;
[0015] The trained initial video generation model is used as the target video generation model.
[0016] Optionally, extracting the portrait mask from the sample video data includes:
[0017] Get the portrait cutout model;
[0018] The portrait mask in the sample video data is extracted using the portrait cutout model.
[0019] Optionally, removing the portrait mask from the sample video data to obtain a sample reference background after removing the portrait includes:
[0020] Get the image restoration model;
[0021] The image restoration model is used to eliminate the portrait corresponding to the portrait mask in the sample video data to obtain a sample reference background after the portrait is eliminated.
[0022] Optionally, the initial video generation model includes an initial encoding module, an initial feature fusion module and an initial decoding module;
[0023] The method of generating a large model by using a preset initial video to replace background data in the sample reference data with the sample reference background to obtain predicted video data includes:
[0024] Inputting the sample reference data and the sample reference background into the initial encoding module, obtaining the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module;
[0025] The predicted spatiotemporal semantic features are fused with the predicted background features by the initial feature fusion module to obtain predicted fused features;
[0026] The predicted fusion features are decoded into predicted video data by the initial decoding module.
[0027] Optionally, the initial coding module includes an initial background coding module and an initial spatiotemporal semantic coding module;
[0028] Inputting the sample reference data and the sample reference background into the initial encoding module, and obtaining the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module, includes:
[0029] Inputting the sample reference data into the initial spatiotemporal semantic encoding module to obtain the predicted spatiotemporal semantic features output by the initial spatiotemporal semantic encoding module;
[0030] The sample reference background is input into the initial background encoding module to obtain the predicted background features output by the initial background encoding module.
[0031] Optionally, the target video generation model includes a target encoding module, a target feature fusion module and a target decoding module;
[0032] The step of generating a large model by using the target video to replace background data in the original video data with the target background data to obtain target video data includes:
[0033] Inputting the original video data and the target background data into the target encoding module, obtaining target spatiotemporal semantic features corresponding to the original video data and target background features corresponding to the target background data output by the target encoding module;
[0034] The target spatiotemporal semantic feature is fused with the target background feature by the target feature fusion module to obtain a target fusion feature;
[0035] The target fusion feature is decoded into target video data by the target decoding module.
[0036] This application also provides a character background editing device, comprising:
[0037] A data acquisition module is used to acquire original video data with people and target background data without people;
[0038] A model determination module is configured to determine a target video generation large model, wherein the target video generation large model is obtained by training using sample video data with people as sample labels, a sample reference background after removing the people from the sample video data, and sample reference data synthesized by synthesizing the people mask extracted from the sample video data and the sample background data without people as training samples;
[0039] The video generation module is used to replace the background data in the original video data with the target background data through the target video generation large model to obtain the target video data.
[0040] The present application also provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the character background editing method as described in any of the above embodiments.
[0041] The present application also provides a computer device, comprising: one or more processors, and a memory;
[0042] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the character background editing method described in any one of the above embodiments are performed.
[0043] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0044] The character background editing method, device, storage medium and computer equipment provided by the present application, after obtaining the original video data with characters and the target background data without characters, the present application can first determine the target video generation model. Since the target video generation model uses the sample video data with characters as the sample label, the sample reference background after eliminating the portrait in the sample video data, and the sample reference data synthesized by the portrait mask extracted from the sample video data and the sample background data without characters as the training samples, the output obtained after training is equivalent to the expected output when training the target video generation model, which is real and unmodified data, so that the target video generation model can simulate the real video data. Therefore, after the background data in the original video data is replaced with the target background data through the target video generation model, the present application can obtain more realistic target video data, thereby effectively improving the character background editing effect in the video creation scene, and the process only requires the user to input the original video data with characters and the target background data without characters to obtain the target video data, thereby greatly reducing the usage threshold. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0046] Figure 1 A flowchart of a method for editing a character background provided in an embodiment of the present application;
[0047] Figure 2 A diagram showing the process of removing human figures from sample video data provided in an embodiment of the present application;
[0048] Figure 3 A diagram showing the process of synthesizing a portrait mask and sample background data provided in an embodiment of the present application;
[0049] Figure 4 This is a diagram of the model architecture when using the Wan2.1 model for training provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of the structure of a character background editing device provided in an embodiment of the present application;
[0051] Figure 6 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0053] In one embodiment, Figure 1 As shown, Figure 1 A flowchart of a method for editing a character background provided in an embodiment of the present application is provided. The present application provides a method for editing a character background, which may include:
[0054] S110: Acquire original video data with people and target background data without people.
[0055] In this step, when replacing the background of the person in the video data, the original video data with the person and the target background data without the person can be obtained first, so as to replace the background in the original video data with the target background data.
[0056] Among them, the present application can collect original video data and target background data through the Internet and upload them, or can shoot original video data and target background data through a shooting device and upload them. The specific method of obtaining video data and background data can be selected according to actual conditions and is not limited here. The texture of the characters in the original video data with characters obtained by the present application is clear and continuous, and the characters in the original video data can be one or more, which can be selected according to actual conditions and is not limited here. The target background data without characters obtained by the present application refers to a video or image that is different from the background in the original video data, which can be selected according to actual conditions and is not limited here.
[0057] S120: Determine a target video generation large model, wherein the target video generation large model is obtained through training using sample video data with people as sample labels, sample reference background after removing the portrait from the sample video data, and sample reference data synthesized by synthesizing the portrait mask extracted from the sample video data and the sample background data without people as training samples.
[0058] In this step, after obtaining the original video data with people and the target background data without people through S110, the target video generation large model can also be determined. The target video generation large model is a model that has been pre-trained and saved locally. The target video data can be obtained by using this model to replace the background with people.
[0059] In a specific implementation, during the training phase of generating a large model from a target video, the present application may collect sample video data with people as sample labels. The sample video data may be collected via the Internet or captured by a camera, and the specific selection may depend on the actual situation and is not limited here. The texture of the people in the sample video data with people obtained by the present application is clear and continuous, and the number of people in the sample video data may be one or more, and the specific selection may depend on the actual situation and is not limited here.
[0060] The present application can also collect background videos or images that do not contain people through the Internet, or use a camera to shoot background videos or images that do not contain people, and use them as sample background data without people to synthesize them with the portrait mask extracted from the sample video data to obtain sample reference data, and the video data after removing the portrait from the sample video data can be used as the sample reference background. During training, the present application can use the sample reference data and the sample reference background as training samples, and use the training samples and sample labels to train the initial video generation model to obtain the target video generation model.
[0061] Furthermore, in order to improve the quality of video generation, this application can use large open source video generation models such as Wan2.1 and HunyanVideo as the basic model, that is, the initial video generation model, and fine-tune the basic model in combination with pre-acquired training samples and sample labels, so as to train a target video generation model that can be used for video character background editing. It is understandable that large open source video generation models such as Wan2.1 and HunyanVideo all use diffusion models or their variants as the core framework, generate high-quality video content through gradual denoising, and support multimodal inputs such as text and images, thereby realizing the text-to-video or image-to-video generation process. In addition, in the process of generating videos, the temporal coherence and spatial details of the video can also be jointly modeled through mechanisms such as 3D convolution and spatiotemporal attention to ensure the dynamic naturalness of the generated video. Therefore, this application uses this type of video generation model as the basic model, which can significantly improve the quality and efficiency of video generation and lower the threshold for video creation.
[0062] S130: Generate a large model through the target video to replace the background data in the original video data with the target background data to obtain the target video data.
[0063] In this step, after determining the target video generation large model through S120, the present application can replace the background data in the original video data with the target background data through the target video generation large model, so as to obtain the target video data.
[0064] Among them, since the target video generation large model of this application is obtained after training with sample video data with people as sample labels, sample reference background after eliminating the portrait in the sample video data, and sample reference data synthesized by synthesizing the portrait mask extracted from the sample video data and the sample background data without people as training samples, therefore, this application can not only make use of the powerful generation ability of the target video generation large model to realize the video character background editing ability, but also make the generated target video data more realistic and natural, thereby effectively improving the editing effect of the character background in the video creation scene.
[0065] In the above embodiment, after obtaining the original video data with people and the target background data without people, the present application can first determine the target video generation model. Since the target video generation model uses the sample video data with people as the sample label, the sample reference background after eliminating the portrait in the sample video data, and the sample reference data synthesized by the portrait mask extracted from the sample video data and the sample background data without people as the training samples, the output obtained after training is equivalent to the expected output when training the target video generation model, which is real and unmodified data, so that the target video generation model can simulate the real video data. Therefore, after the target video generation model is used to replace the background data in the original video data with the target background data, the present application can obtain more realistic target video data, thereby effectively improving the character background editing effect in the video creation scene, and the process only requires the user to input the original video data with people and the target background data without people to obtain the target video data, thereby greatly reducing the usage threshold.
[0066] In one embodiment, the training process of generating a large model from a target video may include:
[0067] S210: Obtain sample video data with people and sample background data without people.
[0068] S211: extracting a portrait mask from the sample video data, and eliminating the portrait mask from the sample video data to obtain a sample reference background after eliminating the portrait.
[0069] S212: Synthesize the portrait mask and the sample background data to obtain sample reference data.
[0070] S213: Using a preset initial video to generate a large model, the background data in the sample reference data is replaced with the sample reference background to obtain predicted video data.
[0071] S214: With the goal of making the predicted video data approach the sample video data, the initial video generation model is trained until a preset end condition is reached.
[0072] S215: Generate a large model using the trained initial video as a target video.
[0073] In this embodiment, when training the target video generation model, a batch of sample video data with people and sample background data without people can be collected through the Internet or by shooting with a shooting device, and then the portrait mask in the sample video data is extracted, and the portrait corresponding to the portrait mask is eliminated from the sample video data to obtain the sample reference background after the portrait is eliminated. Then, the present application can synthesize the portrait mask with the sample background data to obtain the sample reference data. The present application can input the sample video data, sample background data and sample reference data into the initial video generation model together, and replace the background data in the sample reference data with the sample reference background through the initial video generation model to obtain the predicted video data. Then, with the goal of the predicted video data approaching the sample video data, the initial video generation model is trained until the preset end condition is reached, so as to obtain the trained initial video generation model. The present application can use the trained initial video generation model as the target video generation model, so that the background in the original video data input by the user can be replaced by the target video generation model.
[0074] Schematically, as Figure 2 、 3 As shown, Figure 2 This is a diagram showing the process of removing human figures from sample video data provided by an embodiment of the present application. Figure 3 A diagram showing the process of synthesizing a portrait mask and sample background data provided in an embodiment of the present application; Figure 2 、 Figure 3 It can be seen that this application can be used to sample video data (i.e. Figure 3 The original person video in the image is extracted and the portrait mask is obtained and compared with the sample background data (i.e. Figure 3 The background image / video in the image is synthesized to obtain the sample reference data (i.e. Figure 3 Then, the present application can input the sample video data, sample background data and sample reference data into the initial video generation model for training, thereby obtaining the target video generation model.
[0075] Among them, the termination conditions preset in this application may include the number of training times reaching a preset threshold, the training loss converging to a preset range, etc., which can be set according to actual conditions and are not limited here. Through the above training process, this application can ensure that the target video generation model can learn how to accurately separate and synthesize people and backgrounds, thereby generating high-quality target video data.
[0076] In addition, it is worth noting that in the process of editing the character background, it is also necessary to consider the consistency of lighting, shadows, perspective, etc. between the character and the background. In order to solve these problems, the present application can introduce corresponding loss functions in the training stage of the target video generation large model to constrain the generated video data to be consistent with the real video data in these aspects. For example, the present application can introduce a lighting consistency loss function to constrain the continuity of lighting between the character and the background in the generated video, thereby avoiding obvious lighting inconsistencies in the generated video, thereby further improving the accuracy and naturalness of character background editing, and bringing more convenient and efficient solutions to the field of video creation.
[0077] In one embodiment, extracting the portrait mask from the sample video data in S211 may include:
[0078] S2111: Obtain a portrait cutout model.
[0079] S2112: Extracting a portrait mask from the sample video data using the portrait cutout model.
[0080] In this embodiment, when extracting the portrait mask from the sample video data, a mainstream or latest portrait cutout model may be obtained first, and then the portrait cutout model is used to extract the portrait mask from the sample video data, and the portrait mask is saved separately.
[0081] Specifically, the portrait cutout model of the present application can be constructed using deep learning technology, such as architectures based on convolutional neural networks (CNN) or generative adversarial networks (GAN). By training a large amount of image data containing people and backgrounds, the portrait cutout model can learn the characteristic differences between people and backgrounds, thereby accurately extracting portrait masks. Moreover, during the extraction process, the portrait cutout model processes each frame of the sample video data and extracts the portrait mask frame by frame, thereby ensuring the continuity and accuracy of the extraction results. The extracted portrait mask can be used in the subsequent background replacement process and synthesized with new background data to generate high-quality video data.
[0082] In one embodiment, removing the portrait mask from the sample video data in S211 to obtain a sample reference background after removing the portrait may include:
[0083] S2113: Obtain an image restoration model.
[0084] S2114: Eliminate the portrait corresponding to the portrait mask in the sample video data using the image restoration model to obtain a sample reference background after the portrait is eliminated.
[0085] In this embodiment, when removing the portrait in the sample video data, the mainstream or latest image restoration model can be obtained first, and then the image restoration model is used to remove the portrait corresponding to the portrait mask in the sample video data, thereby obtaining the sample reference background after removing the portrait.
[0086] Specifically, the image restoration model of the present application can also be constructed using deep learning technology, such as architectures based on convolutional neural networks (CNN) or generative adversarial networks (GAN). The image restoration model learns how to repair missing areas in images by training on a large amount of image data containing missing areas, so that it can accurately eliminate the portrait corresponding to the portrait mask in the sample video data and generate a coherent and natural background. The generated sample reference background after removing the portrait can be used in the subsequent background replacement process and synthesized with the new background data to generate high-quality video data. In this way, the present application can efficiently implement character background editing and provide users with a more convenient and efficient video creation experience.
[0087] In one embodiment, the initial video generation model may include an initial encoding module, an initial feature fusion module and an initial decoding module.
[0088] In step S213, the background data in the sample reference data is replaced with the sample reference background by using a preset initial video generation model to obtain predicted video data, which may include:
[0089] S2131: Input the sample reference data and the sample reference background into the initial encoding module to obtain the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module.
[0090] S2132: Fusing the predicted spatiotemporal semantic features with the predicted background features through the initial feature fusion module to obtain predicted fusion features.
[0091] S2133: Decoding the predicted fusion features into predicted video data through the initial decoding module.
[0092] In this embodiment, the initial video generation model may include multiple modules, specifically an initial encoding module, an initial feature fusion module and an initial decoding module. This application trains the initial video generation model so that the initial video generation model learns the characteristics of the sample labels, that is, learns the characteristics of the sample video data with people, and then uses the learned characteristics to adjust the parameters in the initial encoding module, the initial feature fusion module and the initial decoding module to improve the video generation quality.
[0093] Specifically, the present application can first input the sample reference data and the sample reference background into the initial encoding module. The initial encoding module will encode the input sample reference data and the sample reference background to obtain the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background. Then, the present application inputs the obtained predicted spatiotemporal semantic features and predicted background features into the initial feature fusion module. The initial feature fusion module will fuse these two features to obtain predicted fused features. Finally, the present application inputs the predicted fused features into the initial decoding module. The initial decoding module will decode the predicted fused features to obtain predicted video data.
[0094] The predicted video data is generated by the initial video generation model based on the input sample reference data and sample reference background. The background has been replaced with the sample reference background, while the characters retain the character features in the sample reference data. Therefore, this application continuously trains and adjusts the initial video generation model through the above process, ultimately obtaining a target video generation model that can generate high-quality target video data.
[0095] In one embodiment, the initial encoding module may include an initial background encoding module and an initial spatiotemporal semantic encoding module.
[0096] Inputting the sample reference data and the sample reference background into the initial encoding module in S2131 to obtain the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module may include:
[0097] S21311: Input the sample reference data into the initial spatiotemporal semantic coding module to obtain the predicted spatiotemporal semantic features output by the initial spatiotemporal semantic coding module.
[0098] S21312: Input the sample reference background into the initial background encoding module to obtain the predicted background features output by the initial background encoding module.
[0099] In this embodiment, after the sample reference data and sample reference background are input into the initial encoding module, the initial encoding module will encode the input sample reference data and sample reference background to obtain the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background, respectively.
[0100] Specifically, the initial encoding module of the present application may include an initial background encoding module and an initial spatiotemporal semantic encoding module, wherein the initial spatiotemporal semantic encoding module is responsible for encoding the sample reference data to extract the spatiotemporal semantic information of the characters therein, which includes the characters' movements, positions, and relative relationships with the background, etc., thereby obtaining predicted spatiotemporal semantic features. The initial background encoding module is responsible for encoding the sample reference background to extract the spatial structure and texture features of the background, thereby obtaining predicted background features. In this way, the present application can ensure that in the subsequent feature fusion process, the features of the characters and the background can be effectively combined, thereby generating high-quality target video data.
[0101] In a specific implementation, the initial video generation model of this application can be the Wan2.1 model. When this application uses the Wan2.1 model as the basic model, its specific model architecture is as follows: Figure 4 As shown, Figure 4 This is a diagram of the model architecture when using the Wan2.1 model for training provided in an embodiment of the present application; Figure 4 The reference background image / video is the sample reference background of this application, the reference image is the sample reference data in this application, and the expected generation result is the sample video data of this application.
[0102] Figure 4 In the
[15] , the initial background encoding module extracts static features (such as scene layout and hue) from the reference background image and outputs a low-dimensional latent variable bg_latent. The initial spatiotemporal semantic encoding module compresses the input video frame (or single image frame) into a spatiotemporal latent variable video_latent while preserving temporal dynamic information. The initial feature fusion module uses the Transformer architecture to denoise the spatiotemporal latent variable (noise prediction) and supports long sequence modeling and multi-condition fusion. Specifically, the initial feature fusion module receives bg_latent (the background latent variable) from the initial background encoding module and video_latent (the noisy spatiotemporal latent variable) from the initial spatiotemporal semantic encoding module and implements multi-condition control through cross-attention or feature concatenation. The initial decoding module decodes the low-dimensional spatiotemporal latent variable denoised by the initial feature fusion module into a video frame in pixel space, ultimately generating the predicted video data.
[0103] Furthermore, the initial background encoding module is typically designed to process static background images. This application can be expanded through technology to support video background input. For example, this application can extract background features of dynamic backgrounds through keyframe extraction + static encoding, or can achieve dynamic background feature extraction through a time-series background encoder or static and dynamic background separation training. The specific settings can be made according to the actual situation and are not limited here.
[0104] In one embodiment, the target video generation model may include a target encoding module, a target feature fusion module and a target decoding module.
[0105] In step S130, the target video is used to generate a large model to replace the background data in the original video data with the target background data to obtain the target video data, which may include:
[0106] S131: Inputting the original video data and the target background data into the target encoding module, obtaining the target spatiotemporal semantic features corresponding to the original video data and the target background features corresponding to the target background data output by the target encoding module.
[0107] S132: The target spatiotemporal semantic feature is fused with the target background feature by the target feature fusion module to obtain a target fusion feature.
[0108] S133: Decoding the target fusion features into target video data through the target decoding module.
[0109] In this embodiment, the target video generation model can also include multiple modules, specifically a target encoding module, a target feature fusion module, and a target decoding module. After training the target video generation model, the application can use the target video generation model to replace the background of the original video data, thereby generating the target video data.
[0110] Specifically, the present application can first input the original video data and target background data into the target encoding module, and the target encoding module will encode the input original video data and target background data to obtain the target spatiotemporal semantic features corresponding to the original video data, and the target background features corresponding to the target background data. Then, the present application can input the obtained target spatiotemporal semantic features and target background features into the target feature fusion module, and the target feature fusion module will fuse these two features to obtain the target fusion features. Finally, the present application inputs the target fusion features into the target decoding module, and the target decoding module will decode the target fusion features to obtain the target video data.
[0111] It is understandable that the target video data of this application is generated by the target video generation large model based on the input original video data and target background data. The background has been replaced with the target background data, while the characters retain the character features in the original video data. In this way, this application can efficiently implement character background editing and provide users with a more convenient and efficient video creation experience. At the same time, since the target video generation large model is trained on a large amount of sample data, it can learn the complex relationship between the character and the background, thereby generating more natural and realistic target video data.
[0112] The following describes a device for editing a character background provided in an embodiment of the present application. The device for editing a character background described below and the method for editing a character background described above can refer to each other.
[0113] In one embodiment, Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of a character background editing device provided in an embodiment of the present application. The present application also provides a character background editing device, which may include a data acquisition module 210, a model determination module 220, and a video generation module 230, specifically including the following:
[0114] The data acquisition module 210 is used to acquire original video data with people and target background data without people.
[0115] The model determination module 220 is used to determine the target video generation large model, wherein the target video generation large model is obtained after training with sample video data with people as sample labels, sample reference background after eliminating the portrait in the sample video data, and sample reference data synthesized by synthesizing the portrait mask extracted from the sample video data and the sample background data without people as training samples.
[0116] The video generation module 230 is configured to generate a large model using the target video to replace the background data in the original video data with the target background data to obtain the target video data.
[0117] In the above embodiment, after obtaining the original video data with people and the target background data without people, the present application can first determine the target video generation model. Since the target video generation model uses the sample video data with people as the sample label, the sample reference background after eliminating the portrait in the sample video data, and the sample reference data synthesized by the portrait mask extracted from the sample video data and the sample background data without people as the training samples, the output obtained after training is equivalent to the expected output when training the target video generation model, which is real and unmodified data, so that the target video generation model can simulate the real video data. Therefore, after the target video generation model is used to replace the background data in the original video data with the target background data, the present application can obtain more realistic target video data, thereby effectively improving the character background editing effect in the video creation scene, and the process only requires the user to input the original video data with people and the target background data without people to obtain the target video data, thereby greatly reducing the usage threshold.
[0118] In one embodiment, the present application also provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the character background editing method as described in any of the above embodiments.
[0119] In one embodiment, the present application further provides a computer device, including: one or more processors, and a memory.
[0120] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the character background editing method described in any one of the above embodiments are performed.
[0121] Schematically, as Figure 6 As shown, Figure 6 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. The computer device 300 can be provided as a server. Figure 6 Computer device 300 includes a processing component 302, which further includes one or more processors, and a memory resource represented by memory 301 for storing instructions executable by processing component 302, such as an application. The application stored in memory 301 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 302 is configured to execute the instructions to perform the character background editing method of any of the above-mentioned embodiments.
[0122] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or the like.
[0123] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0124] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0125] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
[0126] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A character background editing method, characterized in that: The method comprises: Obtain original video data with people and target background data without people; Determining a target video generation large model, wherein the target video generation large model is obtained by training using sample video data with people as sample labels, a sample reference background after removing the people from the sample video data, and sample reference data obtained by synthesizing the people mask extracted from the sample video data and the sample background data without people as training samples; Generating a large model by using the target video to replace the background data in the original video data with the target background data to obtain target video data; The training process of generating a large model from the target video includes: Obtain sample video data with people and sample background data without people; Extracting a portrait mask from the sample video data, and eliminating the portrait mask from the sample video data to obtain a sample reference background after eliminating the portrait; Synthesizing the portrait mask with the sample background data to obtain sample reference data; Replacing the background data in the sample reference data with the sample reference background through a preset initial video generation model to obtain predicted video data; With the goal of making the predicted video data approach the sample video data, training the initial video generation model until a preset end condition is reached; The trained initial video generation model is used as the target video generation model.
2. The character background editing method according to claim 1, characterized in that: The extracting of the portrait mask from the sample video data includes: Get the portrait cutout model; The portrait mask in the sample video data is extracted using the portrait cutout model.
3. The character background editing method according to claim 1, characterized in that: The removing the portrait mask from the sample video data to obtain a sample reference background after the portrait is removed includes: Get the image restoration model; The image restoration model is used to eliminate the portrait corresponding to the portrait mask in the sample video data to obtain a sample reference background after the portrait is eliminated.
4. The character background editing method according to claim 1, characterized in that: The initial video generation model includes an initial encoding module, an initial feature fusion module and an initial decoding module; The method of generating a large model by using a preset initial video to replace background data in the sample reference data with the sample reference background to obtain predicted video data includes: Inputting the sample reference data and the sample reference background into the initial encoding module, obtaining the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module; The predicted spatiotemporal semantic features are fused with the predicted background features by the initial feature fusion module to obtain predicted fused features; The predicted fusion features are decoded into predicted video data by the initial decoding module.
5. The character background editing method according to claim 4, characterized in that: The initial coding module includes an initial background coding module and an initial spatiotemporal semantic coding module; Inputting the sample reference data and the sample reference background into the initial encoding module, and obtaining the predicted spatiotemporal semantic features corresponding to the sample reference data and the predicted background features corresponding to the sample reference background output by the initial encoding module, includes: Inputting the sample reference data into the initial spatiotemporal semantic encoding module to obtain the predicted spatiotemporal semantic features output by the initial spatiotemporal semantic encoding module; The sample reference background is input into the initial background encoding module to obtain the predicted background features output by the initial background encoding module.
6. The method for editing character background according to any one of claims 1 to 5, characterized in that: The target video generation model includes a target encoding module, a target feature fusion module and a target decoding module; The step of generating a large model by using the target video to replace background data in the original video data with the target background data to obtain target video data includes: Inputting the original video data and the target background data into the target encoding module, obtaining target spatiotemporal semantic features corresponding to the original video data and target background features corresponding to the target background data output by the target encoding module; The target spatiotemporal semantic feature is fused with the target background feature by the target feature fusion module to obtain a target fusion feature; The target fusion feature is decoded into target video data by the target decoding module.
7. A character background editing device, characterized in that: include: A data acquisition module is used to acquire original video data with people and target background data without people; A model determination module is configured to determine a target video generation large model, wherein the target video generation large model is obtained by training using sample video data with people as sample labels, a sample reference background after removing the people from the sample video data, and sample reference data synthesized by synthesizing the people mask extracted from the sample video data and the sample background data without people as training samples; A video generation module, configured to replace background data in the original video data with the target background data by generating a large model of the target video to obtain target video data; The model determination module includes: Obtain sample video data with people and sample background data without people; Extracting a portrait mask from the sample video data, and eliminating the portrait mask from the sample video data to obtain a sample reference background after eliminating the portrait; Synthesizing the portrait mask with the sample background data to obtain sample reference data; Replacing the background data in the sample reference data with the sample reference background through a preset initial video generation model to obtain predicted video data; With the goal of making the predicted video data approach the sample video data, training the initial video generation model until a preset end condition is reached; The trained initial video generation model is used as the target video generation model.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the character background editing method according to any one of claims 1 to 6.
9. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the character background editing method according to any one of claims 1 to 6 are performed.
Citation Information
Patent Citations
Model training method, a method and device for replacing image background, and an electronic system
CN109377445A