Video editing method and device, equipment, medium and product
By jointly encoding video features, mask features and noise features, background feature representations are generated, and used to edit the foreground area in the video as a noise reduction condition, solving the problem of weak control signals in the prior art, resulting in poor video effects, and achieving higher quality video editing.
Patent Information
- Application Number
- CN202510208659.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the existing video repair methods, the control signal condition control is weak, resulting in poor video effects after repair.
By obtaining the video feature representation of the video, the mask feature representation of the mask image sequence, and the noise feature representation of the sampling noise, the joint encoding obtains the background feature representation, which is used to edit the foreground area in the video as a noise reduction condition, and generates a new video.
The background feature represents the guidance of the reservation of the background area when generating a new video, improving the edited video quality.
Smart Images

Figure CN120050482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a video editing method, apparatus, device, medium and product. Background Art
[0002] Video restoration plays a key role in the media industry, and its goal is to restore damaged video content.
[0003] In related technologies, the development of generative models has spawned many video restoration methods, which use additional modules or training strategies to expand the video restoration capabilities of the backbone network. For example, VideoComposer is based on a pre-trained video generation backbone network, and integrates various control signals (text, depth, mask, motion vector) through a shared spatio-temporal conditional fusion module, thereby assisting the generative model to achieve the video restoration function.
[0004] In the video restoration achieved by the above method, the conditional control that the control signal can provide is weak, resulting in poor effects of the restored video. Summary of the Invention
[0005] Embodiments of this application provide a video editing method, apparatus, device, medium and product. The technical solution is as follows:
[0006] On the one hand, a video editing method is provided, and the method includes:
[0007] Obtain a video feature representation of a first video, where the video image frames of the first video include a background region and a foreground region;
[0008] Obtain a mask feature representation of a mask image sequence, where the mask image sequence is used to indicate the position of the foreground region in the video image frames of the first video; and obtain a noise feature representation of sampling noise;
[0009] Jointly encode the video feature representation, the mask feature representation, and the noise feature representation to obtain a background feature representation corresponding to the background region, where the background feature representation is used to indicate the background information in the first video;
[0010] Taking editing the foreground region in the first video as a noise reduction condition, perform noise reduction on the noise feature representation based on the background feature representation to obtain a second video as the editing result of the first video.
[0011] On the other hand, a video editing apparatus is provided, and the apparatus includes:
[0012] An obtaining module, configured to obtain a video feature representation of a first video, where the video image frames of the first video include a background region and a foreground region;
[0013] The obtaining module is further configured to obtain a mask feature representation of a mask image sequence, where the mask image sequence is used to indicate the position of the foreground region in the video image frames in the first video; and obtain a noise feature representation of sampling noise.
[0014] The encoding module is configured to jointly encode the video feature representation, the mask feature representation, and the noise feature representation to obtain a background feature representation corresponding to the background region, where the background feature representation is used to indicate background information in the first video.
[0015] The prediction module is configured to use editing the foreground region in the first video as a noise reduction condition, and perform noise reduction on the noise feature representation based on the background feature representation to obtain a second video as an editing result of the first video.
[0016] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the video editing method according to any one of the foregoing embodiments of the present application.
[0017] On the other hand, a computer-readable storage medium is provided, where at least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the video editing method according to any one of the foregoing embodiments of the present application.
[0018] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video editing method according to any one of the foregoing embodiments.
[0019] The technical solutions provided in the present application at least include the following beneficial effects:
[0020] When editing the first video, a background feature representation is generated based on the video feature representation corresponding to the first video, the mask feature representation corresponding to the foreground region, and the noise feature representation. This background feature representation can indicate the background information of the background region to be retained in the first video. During the process of generating the edited second video from the noise feature representation, the generation of the content of the foreground region and the retention of the content of the background region are decoupled in the process of generating a new video from the noise feature representation. Thus, the retention of the background region is guided by the background feature representation when generating a new video from the noise feature representation, thereby improving the video quality of the edited second video. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a block diagram of the structure of a computer system provided by an exemplary embodiment of the present application;
[0023] Figure 2 is a flowchart of a video editing method provided by an exemplary embodiment of the present application;
[0024] Figure 3 is a schematic diagram of a video editing model provided by an exemplary embodiment of the present application;
[0025] Figure 4 is a flowchart of a video editing method provided by an exemplary embodiment of the present application;
[0026] Figure 5 is a schematic diagram of a feature screening process provided by an exemplary embodiment of the present application;
[0027] Figure 6 is a flowchart of a video editing method provided by an exemplary embodiment of the present application;
[0028] Figure 7 is a schematic diagram of a video editing model provided by an exemplary embodiment of the present application;
[0029] Figure 8 is a flowchart of a model training method provided by an exemplary embodiment of the present application;
[0030] Figure 9 is a flowchart of the construction of a data set provided by an exemplary embodiment of the present application;
[0031] Figure 10It is a schematic diagram of feature generation when the adaptation network performs ID resampling in the diffusion network provided by an exemplary embodiment of the present application;
[0032] Figure 11 It is a flowchart of a video restoration method provided by an exemplary embodiment of the present application;
[0033] Figure 12 It is a schematic diagram for comparing the model generation results provided by an exemplary embodiment of the present application;
[0034] Figure 13 It is a schematic diagram for comparing the model generation results provided by an exemplary embodiment of the present application;
[0035] Figure 14 It is a structural block diagram of an editing device for a video provided by an exemplary embodiment of the present application;
[0036] Figure 15 It is a structural block diagram of an editing device for a video provided by an exemplary embodiment of the present application;
[0037] Figure 16 It is a schematic diagram of the structure of a server provided by an exemplary embodiment of the present application. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0039] In the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first" and "second", nor are the quantity and execution order limited.
[0040] First, a brief introduction to the nouns involved in the embodiments of the present application will be given.
[0041] Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0042] Machine Learning (ML): It is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0043] Artificial Intelligence Generated Content (AIGC): It refers to the content generated by artificial intelligence (AI) technology. This content can include various forms such as text, images, audio, and video. AIGC technology uses machine learning and deep learning algorithms to automatically generate high-quality content, reducing the time and cost of manual creation.
[0044] Diffusion Model: It is a generative model. Its core principle is to simulate a forward diffusion process that gradually transitions from the data distribution to a simple distribution (usually Gaussian noise), and a reverse denoising process that recovers from the simple distribution to the original data distribution. Structurally, a diffusion model usually consists of a forward diffusion process and a trainable reverse denoising process. The forward process gradually adds Gaussian noise to the image until pure noise is finally obtained; the reverse process is to train a neural network to gradually denoise from the pure noise until a real image is obtained. Diffusion models have a wide range of application scenarios, including but not limited to image super-resolution, image restoration, object detection, video generation, etc.
[0045] Diffusion Transformer (DiT): It is a diffusion model based on the Transformer architecture for image and video generation tasks. It precisely simulates the diffusion process by converting the spatial input into a token sequence and then processing these tokens using Transformer blocks. The core innovations of DiT include the Adaptive Layer Normalization (AdaLN) and the Adaptive Layer Normalization with Zero Initialization (AdaLN-Zero) mechanisms, which enhance the scalability and generation performance of the model. In addition, DiT also supports multimodal generation.
[0046] Identifier (ID) resampling technique: It is a technique for resampling data entries with the same identifier (ID). Its purpose is to address data imbalance problems or enhance the model's learning ability for specific classes by upsampling or downsampling the data.
[0047] Figure 1 The structural block diagram of a computer system provided by an exemplary embodiment of the present application is shown. This computer system 100 includes: a terminal 120 and a server 140.
[0048] The terminal 120 installs and runs an application program that supports video editing functions. This application program can be any one of an e-commerce application program, a payment application program, a video application program, a social application program, a game application program, etc.
[0049] The device types of the terminal 120 include at least one of a smart phone, a laptop computer, a desktop computer, a tablet computer, a smart speaker, a smart robot, an extended reality device, etc.
[0050] The terminal 120 is connected to the server 140 via a wireless network or a wired network.
[0051] Those skilled in the art can understand that the number of the above devices can be more or less. For example, there can be only one of the above devices, or there can be dozens or hundreds of the above devices, or even more. The embodiments of the present application do not limit the number and type of the devices.
[0052] The server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 140 is used to provide background services for an application program that supports video editing functions. Optionally, a video editing model is installed in the server 140, and the video editing model can implement the AI editing function of the video. Optionally, the server 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, the server 140 and the terminal 120 adopt a distributed computing architecture for collaborative computing.
[0053] It should be noted that the above server 140 can be implemented as a physical server or as a cloud server in the cloud. In some embodiments, the above server 140 can also be implemented as a node in a blockchain system.
[0054] Illustratively, the terminal 120 uploads the first video 121 to be edited to the server 140. After receiving the first video 121, the server 140 executes an editing process for the first video 121 by invoking a video editing service. During the editing process for the first video 121, the server 140 generates a mask image sequence 141 corresponding to the first video 121. The mask image sequence 141 is used to indicate the position of the foreground region in the video image frames of the first video 121, and the foreground region is the region to be edited. The server 140 extracts the features of the first video 121 to obtain a video feature representation; extracts the features of the mask image sequence 141 to obtain a mask feature representation; and obtains a noise feature representation corresponding to the sampled sampling noise. The server 140 encodes the above video feature representation, mask feature representation, and noise feature representation into a background feature representation through a context encoder 142 in the video editing model, inputs the background feature representation and the noise feature representation into a diffusion network 143 in the video editing model, and uses the diffusion network 143 to denoise the noise feature representation based on the background feature representation with the condition of editing the foreground region in the first video 121, so as to obtain a second video 144 as the editing result of the first video 121. The server 140 sends the generated second video 144 to the terminal 120, and the terminal 120 displays the second video 144 to show the user the AI editing result of the first video 121.
[0055] It should be noted that the video editing method provided by the embodiments of the present application can also be executed independently by the terminal 120 or the server 140, which is not limited herein. In one example, the terminal 120 includes a functional device for implementing video editing. The terminal 120 directly inputs the first video 121 into the above functional device to obtain the second video 144 as the editing result of the first video 121. In another example, the server 140 obtains the first video 121 from the database, inputs the first video 121 into the video editing service, and obtains the second video 144 as the editing result of the first video 121.
[0056] Combined with the above noun introduction and application scenarios, the video editing method provided by the present application will be described. Taking the application of this method to the terminal as an example, as Figure 2 shown, this method includes the following steps 210 to 250.
[0057] Step 210, obtain the video feature representation of the first video.
[0058] Schematically, the first video is a video with a video editing requirement. Optionally, the first video can be implemented as a video collected by an image acquisition device; or, the first video can be implemented as a video produced by a computer device.
[0059] The video image frames of the first video include a background area and a foreground area to be edited. Among them, the foreground area is the area with an editing requirement in the first video, and the background area is the area where the original image content needs to be retained during the editing of the first video. That is, during the editing of the first video, the content of the background area is retained and the content of the foreground area is edited.
[0060] Optionally, the above editing requirement can be implemented as a content repair requirement, a content modification requirement, a content addition requirement, a content deletion requirement, etc. for the first video, which is not limited herein.
[0061] Optionally, the foreground area with an editing requirement can be an area marked by the user. Schematically, the terminal provides a selection tool to the user, and the user draws a custom graphic on at least one video image frame in the first video through the selection tool, and takes the area indicated by the custom graphic as the foreground area.
[0062] Optionally, the foreground area to be edited can be identified by a pre-trained foreground recognition model. Schematically, the first video and an editing requirement prompt are input into the foreground recognition model, and the foreground recognition model identifies a foreground area corresponding to the editing requirement prompt from the first video according to the editing requirement prompt. Optionally, the editing requirement prompt can be used to indicate an object (e.g., a person, an item, a building, a prop, etc.) and / or a location (e.g., the lower right corner, the lower left corner, the middle, the upper half, etc. of the video frame) that appears in the first video.
[0063] Optionally, the above foreground recognition model can be implemented by neural network models such as Convolutional Neural Networks (CNN), Feedforward Neural Network (FNN), Residual Network (ResNet), Transformer, etc., and no specific limitation is imposed here.
[0064] In some embodiments, the video feature representation of the first video is used to indicate the texture features of the video image frames of the first video. Among them, the texture features are used to describe the regularity shown by the local spatial arrangement pattern of pixel values in the video image frames of the first video, and reflect the texture, structure, and details of the object surface existing in the video image frames.
[0065] In some embodiments, the video image frames of the first video are encoded by a pre-trained first encoder to obtain the video feature representation of the first video. Schematically, the first video is input into the pre-trained first encoder to obtain the video feature representation.
[0066] Optionally, the pre-trained first encoder can be implemented as an Autoencoder, a Variational Autoencoder (VAE), a deep autoencoder, a Recurrent Neural Network Encoder, a Convolutional Neural Network Encoder, an Attention-based Encoder, etc., and no limitation is imposed here.
[0067] In some embodiments, before encoding the first video, a preprocessing operation is first performed on the first video. Optionally, the preprocessing operation includes at least one of video format conversion, frame splitting, noise reduction processing, foreground masking processing, etc.
[0068] Among them, the video format conversion instruction converts the data format of the first video into a data format that can be input into a downstream network (e.g., a pre-trained first encoder). Optionally, the above data formats include MP4 format, Audio Video Interleave (AVI) format, Windows Media Video (WMV) format, Flash Video (FLV) format, etc.
[0069] The frame splitting process instruction decomposes the first video into multiple video image frames. Schematically, the first video is frame-by-frame extracted according to the frame rate of the first video to obtain multiple video image frames. In some embodiments, in order to reduce the amount of data processing, key frame recognition is performed on the first video to determine multiple key frames in the first video, and the multiple key frames are extracted from the first video for downstream processing. For example, the video image frames with foreground regions are extracted from the first video as key frames, so that the downstream only performs video content editing processing on the key frames. In other embodiments, in order to reduce the amount of data processing, video image frames are extracted from the first video according to a preset frame extraction interval to obtain multiple video image frames for downstream processing.
[0070] The noise reduction process instruction improves the video quality of the first video by removing the noise in the video image frames of the first video. Optionally, the noise reduction process for the first video can be implemented as spatial domain noise reduction (directly processing on the video frame, using the statistical characteristics of the local area to estimate and remove noise), temporal threshold noise reduction (using the similarity between multiple frames of video to smooth the noise), frequency domain noise reduction (converting the video signal to the frequency domain and removing high-frequency noise components), and deep learning noise reduction (using a pre-trained noise reduction model to analyze the noise pattern in the video frame and intelligently remove the noise).
[0071] The foreground masking process instruction masks the foreground region in the first video so that the extracted video feature representation focuses on the features of the background region.
[0072] Schematically, annotation information for indicating the foreground region in the first video is obtained; the foreground region in the first video is masked based on the annotation information to obtain a masked video; and the masked video is encoded to obtain a video feature representation. Optionally, the annotation information for the foreground region can be implemented as marker points, marker boxes, position description texts, masked image sequences, etc. for the foreground region, which is not limited here. That is, before extracting the video feature representation, the foreground region is masked first, so as to ensure that the output background feature representation can carry more background region features to realize the guidance of background region content restoration.
[0073] Step 220, obtain the mask feature representation of the mask image sequence.
[0074] Schematically, the mask image sequence is used to indicate the position of the foreground region in the video image frames of the first video, and the mask image sequence is an image sequence composed of multiple mask images.
[0075] In some embodiments, the mask images in the mask image sequence are images that correspond one-to-one with the video image frames in the first video, that is, the number of mask images in the mask image sequence is the same as the number of video image frames in the first video. In other embodiments, the mask images in the mask image sequence correspond to the video image frames in the first video where the foreground region changes.
[0076] Among them, the mask image is an image having the same size as the video image frame of the first video.
[0077] Optionally, the mask image can be implemented as a binary image with only two states of pixel values (such as 0 and 1), which is used to represent the masked area and the non-masked area. The masked area in the mask image indicates that the mapped area in the corresponding video image frame is blocked, that is, this area is not processed, and the non-masked area in the mask image indicates that the mapped area in the corresponding video image frame is not blocked, that is, this area can be processed.
[0078] Optionally, the mask image can be implemented as a grayscale image with pixel values between 0 and 255. Among them, different grayscale values represent different masking degrees of the pixels in the video image frame. That is, the mask image is implemented as a ChangeMap, and the pixel values in the mask image can indicate the change intensity of each pixel in the video image frame during editing.
[0079] In some embodiments, the mask image sequence can be generated according to the foreground region where the user indicates an editing requirement. Schematically, according to the user drawing a custom graphic on at least one video image frame in the first video through a selection tool, the mask image corresponding to the video image frame is generated.
[0080] In other embodiments, the mask image sequence can be data output by a foreground recognition model. Schematically, the first video and an editing requirement prompt word are input into the foreground recognition model. The foreground recognition model identifies and determines the foreground region corresponding to the editing requirement prompt word from the first video according to the editing requirement prompt word, and outputs the mask images of the foreground regions corresponding to each video image frame to obtain the mask image sequence.
[0081] In some embodiments, the mask feature representation corresponding to the mask image sequence is obtained by downsampling. Schematically, the mask image sequence of the first video is acquired; the mask image sequence is subjected to downsampling to obtain the mask feature representation. In one example, each of the multiple mask images in the mask image sequence is separately subjected to downsampling to obtain the mask image feature representation, and the mask image feature representations are concatenated in the sequence order of the multiple mask images to obtain the mask feature representation.
[0082] Among them, downsampling refers to reducing the sampling rate of an image or a signal, thereby reducing the amount of data. In image processing, downsampling is usually achieved by reducing the width and height of the image. In the embodiments of the present application, the mask feature representation obtained by downsampling and the video feature representation obtained by encoding are feature representations of the same dimension.
[0083] Optionally, downsampling can be implemented as at least one of the following:
[0084] · Nearest Neighbor: Select the nearest pixel value as the value after downsampling;
[0085] · Bilinear Interpolation: Calculate the weighted average of the surrounding four pixels to obtain the pixel value after downsampling;
[0086] · Bicubic Interpolation: Perform cubic polynomial interpolation based on the surrounding 16 pixels;
[0087] · Pooling: Implement downsampling through the pooling layer in a convolutional neural network. Optionally, the pooling methods include Max Pooling and Average Pooling. Among them, Max Pooling takes the maximum value of the local area, and Average Pooling takes the average value of the local area.
[0088] Step 230, obtain the noise feature representation of the sampling noise.
[0089] Optionally, sampling noise is a kind of signal interference applied to the original video image frame, which can cause changes in the image information or pixel brightness in the original video image frame. Optionally, the sampling noise can be Gaussian noise, salt-and-pepper noise, etc. In the case where the sampling noise is Gaussian noise, each pixel point in the video image frame after adding noise is a pixel point to which noise is applied.
[0090] In some embodiments, noise is added to the video image frames of the first video through sampling noise to obtain a noisy video, and feature extraction is performed on the video image frames of the noisy video to obtain a noise feature representation. Schematically, sampling noise is acquired; the sampling noise is added to each video image frame of the first video to obtain a noisy video; the noisy video is encoded to obtain a noise feature representation.
[0091] Among them, the noisy video is a video subjected to signal interference. Compared with the original video, the image information or pixel brightness in the video image frames of the noisy video has changed.
[0092] In one example, taking the sampling noise as Gaussian noise, schematically, standard Gaussian noise is sampled to obtain sampling noise; noise is added to the video image frames of the first video according to the sampling noise and the diffusion duration to obtain a noisy video; the noisy video is encoded to obtain a noise feature representation. Among them, the diffusion duration is used to indicate the number of steps required to generate the second video from the noise feature representation in the iterative denoising process, that is, in the process of generating the second video from the noise feature representation, the number of steps required to gradually recover (generate) from the initial state (noise feature representation) to the second video. The diffusion duration is a continuous time variable t ∈ [0, T], and the data distribution corresponding to the 0 moment can be expressed as x 0 ~q(x 0 ), and the data distribution corresponding to the T moment can be expressed as x T ~q(x T ).
[0093] In some embodiments, for the noise addition process of the first video, it can be implemented through the forward diffusion sub-network in the diffusion network. That is, the process of adding noise to the original video through the forward diffusion sub-network is the process of applying signal interference to each pixel point in the video image frames of the original video using the sampling noise. Among them, the video image frames of the original video can be expressed as x 0 , the sampling noise can be expressed as ∈, and the diffusion duration can be expressed as T. Thus, the computing device uses the forward diffusion sub-network to apply noise to each pixel point in the video image frames of the original video based on the sampling noise within the continuous diffusion duration T, and the video image frames x T of the noisy video can be obtained, and the video image frames x T of the noisy video can be pure noise images.
[0094] In the embodiments of the present application, after the forward diffusion sub-network in the diffusion network adds noise to the first video, in order to align the data dimensions of the noise representation with the video feature representation in step 210 and the mask feature representation in step 220, therefore, it is also necessary to perform feature extraction on the noisy video to obtain a noise feature representation with the same dimension as the video feature representation and the mask feature representation.
[0095] In some embodiments, the feature extraction of the noisy video can be implemented by a pre-trained second encoder. Optionally, the pre-trained second encoder can be implemented as an autoencoder, variational autoencoder, deep autoencoder, recurrent neural network encoder, convolutional neural network encoder, attention mechanism-based encoder, etc., which are not limited herein.
[0096] Step 240: Jointly encode the video feature representation, mask feature representation, and noise feature representation to obtain a background feature representation corresponding to the background region.
[0097] In the embodiments of the present application, the background feature representation is used to indicate the background information in the first video. When guiding the recovery of the second video from the noise feature representation, the background feature representation is used to guide the content of the background region that does not need to be edited, thereby improving the editing quality of the generated second video.
[0098] In some embodiments, a pre-trained third encoder is used to jointly encode the video feature representation, mask feature representation, and noise feature representation to obtain a background feature representation corresponding to the background region. Schematically, the video feature representation, mask feature representation, and noise feature representation are input into the pre-trained third encoder to obtain the background feature representation.
[0099] Optionally, the pre-trained third encoder can be implemented as an autoencoder, variational autoencoder, deep autoencoder, recurrent neural network encoder, convolutional neural network encoder, attention mechanism-based encoder, etc., which are not limited herein.
[0100] In some embodiments, before jointly encoding the video feature representation, mask feature representation, and noise feature representation, feature fusion is performed on the video feature representation, mask feature representation, and noise feature representation to obtain a fused feature representation, and the fused feature representation is encoded to obtain the background feature representation.
[0101] Optionally, the feature fusion method for the video feature representation, mask feature representation, and noise feature representation can be implemented as at least one of the following:
[0102] · Feature concatenation: Concatenate the video feature representation, mask feature representation, and noise feature representation to obtain a fused feature representation;
[0103] · Feature addition: Element-wise add the video feature representation, mask feature representation, and noise feature representation to obtain a fused feature representation;
[0104] · Feature averaging: Element-wise average the video feature representation, mask feature representation, and noise feature representation to obtain a fused feature representation;
[0105] · Feature stacking: Stack the video feature representation, mask feature representation, and noise feature representation to obtain a fused feature representation.
[0106] · Deep learning-based fusion: Input the video feature representation, mask feature representation, and noise feature representation into a pre-trained convolutional neural network (or convolutional layer) for feature fusion to obtain a fused feature representation.
[0107] Step 250: Using the foreground region in the first video as the noise reduction condition, reduce the noise in the noise feature representation based on the background feature representation to obtain the second video as the editing result of the first video.
[0108] In the embodiments of the present application, the noise reduction process of the noise feature representation is a process of removing the signal interference applied to each pixel point in the video of the noisy video. In this process, the image content of the background region is restored according to the background feature representation, and new image content of the foreground region is generated from the noise feature representation.
[0109] In some embodiments, the noise reduction process of the noise feature representation is implemented by the reverse diffusion sub-network in the diffusion network. The reverse diffusion sub-network needs to learn the features of the video image frames and measure the noise applied to the noisy video to achieve denoising of the noisy image. Among them, the diffusion duration in the denoising process is also T, that is, the reverse diffusion process from x T to x 0 of.
[0110] In some embodiments, in order to improve the editing effect of the foreground region, the description text for the foreground region input by the user is also combined in the denoising process. Schematically, obtain the description text, where the description text is used to indicate the editing requirements for the foreground region; extract the text feature representation of the description text, and the text feature representation is used to indicate the semantic features of the description text; using the foreground region in the first video as the noise reduction condition, reduce the noise in the noise feature representation based on the background feature representation and the text feature representation to obtain the second video as the editing result of the first video. That is, in the noise reduction process, the image content of the background region is restored according to the background feature representation, and new image content of the foreground region is generated from the noise feature representation according to the text feature representation.
[0111] In summary, when editing the first video, a background feature representation is generated through the video feature representation corresponding to the first video, the mask feature representation corresponding to the foreground region, and the noise feature representation. This background feature representation can indicate the background information of the background region to be retained in the first video. In the process of generating the edited second video from the noise feature representation, the generation of the content of the foreground region and the retention of the content of the background region are decoupled in the process of generating a new video from the noise feature representation. Thus, the retention of the background region when generating a new video from the noise feature representation is guided by the background feature representation, thereby improving the video quality of the edited second video.
[0112] In some alternative embodiments, the video editing method provided by the embodiments of the present application is implemented through a video editing model. Schematically, as Figure 3 shown, it shows a schematic diagram of a video editing model 300 provided by an exemplary embodiment of the present application. The video editing model 300 includes a context encoder 320 and a diffusion network 330.
[0113] Among them, the input of the video editing model 300 includes a video feature representation 301, a mask feature representation 302, and a noise feature representation 303. The context encoder 320 integrates the video feature representation 301, the mask feature representation 302, and the noise feature representation 303 into the downstream diffusion network 330, and generates a second video 304 through the diffusion network 330.
[0114] Schematically, for the implementation of the context encoder 320, please refer to Figure 4 , which shows a video editing method provided by an exemplary embodiment of the present application. This method further includes step 241, where step 241 is a subordinate step of step 240.
[0115] Step 241: Input the video feature representation, the mask feature representation, and the noise feature representation into a pre-trained context encoder to obtain a background feature representation corresponding to the background region.
[0116] Among them, the context encoder is used to extract the semantic features of video image frames. In the embodiments of the present application, the video feature representation, the mask feature representation, and the noise feature representation are integrated into the downstream noise reduction process through an efficient context encoder.
[0117] In some embodiments, before jointly encoding the video feature representation, the mask feature representation, and the noise feature representation by the context encoder, the video feature representation, the mask feature representation, and the noise feature representation are fused for feature fusion. Schematically, the video feature representation, the mask feature representation, and the noise feature representation are fused to obtain a fused feature representation; the fused feature representation is input into the context encoder to obtain a background feature representation corresponding to the background region.
[0118] Optionally, the feature fusion method for video feature representation, mask feature representation, and noise feature representation can be implemented as at least one of the following: feature concatenation; feature addition; feature averaging; feature stacking; deep learning-based fusion.
[0119] In one example, the feature fusion method between video feature representation, mask feature representation, and noise feature representation can be implemented as three types of feature concatenations in units of pixels. Schematically, the video feature representation includes first sub-feature representations corresponding to each pixel point in the video image frames of the first video, the mask feature representation includes second sub-feature representations corresponding to each pixel point in the video image frames of the first video, and the noise feature representation includes third sub-feature representations corresponding to each pixel point in the video image frames of the first video. The fusion of video feature representation, mask feature representation, and noise feature representation is implemented as follows: taking a single pixel point in the video image frame of the first video as a unit, concatenating the first sub-feature representation, the second sub-feature representation, and the third sub-feature representation corresponding to each pixel point to obtain an image frame feature representation corresponding to the video image frame; according to the temporal relationship between the video image frames in the first video, concatenating the image frame feature representations to obtain a fusion feature representation.
[0120] That is, for each video image frame, the video feature representation, mask feature representation, and noise feature representation are concatenated pixel by pixel, so that the obtained fusion feature representation can carry the features of the background region of the first video, the features of the masking situation of the first video, and the features of the noisy video, improving the diversity of the semantic information carried by the background feature representation generated downstream.
[0121] In some embodiments, since after the context encoder encodes the fusion feature representation, the obtained feature representation includes the relevant features of all pixel positions in the video image frames of the first video, and the features provided by the background feature representation in the downstream denoising process are mainly used to indicate the image content of the background region to be retained, therefore, the output of the context encoder is filtered so that the background feature representation input to the next stage only retains the features corresponding to the background region.
[0122] Schematically, the fusion feature representation is input into the context encoder to obtain an intermediate encoded representation; the background feature representation corresponding to the background region is screened out from the intermediate encoded representation.
[0123] In some embodiments, the feature screening for the intermediate encoded representation is implemented based on the identifier (ID) of the feature, where the identifier of the feature is used to indicate the label or number for identifying the feature. Exemplarily, the feature identifiers for the background region and the foreground region are different. Through the ID, it can be determined whether the feature element in the intermediate encoded representation belongs to the background region or the foreground region, so as to screen out the feature elements corresponding to the foreground region from the intermediate encoded representation.
[0124] As shown in Figure 5 , it shows a schematic diagram of the feature screening process provided by an exemplary embodiment of the present application. The intermediate encoding representation 510 output by the context encoder includes tokens 511 corresponding to the foreground region and tokens 512 corresponding to the background region. The tokens 511 corresponding to the foreground region in the intermediate encoding representation 510 are screened out to obtain the background feature representation 520, and the background feature representation 520 is input into the downstream diffusion network.
[0125] In one example, when screening out the feature elements corresponding to the foreground region, in order to ensure that the data dimension of the obtained background feature representation can be the same as that of the noise feature representation downstream, therefore, when screening out the feature elements corresponding to the foreground region, the feature elements corresponding to the foreground region are set to 0.
[0126] That is, before inputting the output of the context encoder into the downstream diffusion network serving as the backbone network, the output of the context encoder is screened to ensure that the output to the downstream only involves background features, preventing potential ambiguities during the generation process of the backbone network, and thereby improving the generation quality of the video.
[0127] In summary, when editing the first video, a background feature representation is generated through the video feature representation corresponding to the first video, the mask feature representation corresponding to the foreground region, and the noise feature representation. This background feature representation can indicate the background information of the background region to be retained in the first video. During the process of generating the edited second video from the noise feature representation, the context encoder decouples the generation of the content in the foreground region and the retention of the content in the background region during the process of generating a new video from the noise feature representation, so as to guide the retention of the background region when generating a new video from the noise feature representation through the background feature representation, thereby improving the video quality of the edited second video.
[0128] The context encoder provided by the embodiments of the present application can also be used as a plug-and-play plugin to guide any pre-trained diffusion network downstream, realizing plug-and-play control and zero-shot adaptation across various stylized backbone networks.
[0129] As shown in Figure 3 , after obtaining the background feature representation through the context encoder 320, the diffusion network 330 realizes the generation process of the second video.
[0130] Schematically, for the implementation of the diffusion network 330, please refer to Figure 6 , which shows a video editing method provided by an exemplary embodiment of the present application. The method further includes step 251, where step 251 is a subordinate step of step 250.
[0131] Step 251: Input the background feature representation and the noise feature representation into a pre-trained diffusion network to obtain a second video.
[0132] Among them, the diffusion network is used to retain the background region in the first video according to the background feature representation and generate the content of the foreground region from the noise feature representation.
[0133] Optionally, the above diffusion network can be implemented as a Conditional Diffusion Models, Stable Diffusion Models, Latent Diffusion Models (LDM), Autoregressive Diffusion Models, Diffusion Transformer, etc., which are not limited here. In the embodiments of this application, is used as an example for illustrative purposes.
[0134] Schematically, when the diffusion network is implemented as a Diffusion Transformer, the process of the diffusion network generating the second video is as follows: in the j-th round of noise reduction process, fuse the j-th noise feature representation and the background feature representation to obtain the j-th conditional noise feature representation, where j is a positive integer; denoise the j-th conditional noise feature representation based on the self-attention mechanism to obtain the (j + 1)-th noise feature representation; use the n-th noise feature representation as the denoised feature representation to perform decoding to obtain the second video.
[0135] Among them, the feature processing process implemented by the self-attention mechanism is: perform a linear transformation on the j-th conditional noise feature representation to obtain a first query vector, a first key vector, and a first value vector; obtain an attention score through the first query vector and the first key vector, where the attention score is used to indicate the correlation between pixels in the image frame; weight and sum the first value vector based on the attention score to obtain the (j + 1)-th noise feature representation.
[0136] Specifically, in each DiT block of the Diffusion Transformer, the self-attention mechanism is used to capture the global dependencies of the features, and at the same time, conditional information (background feature representation, text feature representation) is combined for denoising. The adaptive layer normalization or cross-attention mechanism is used to integrate the conditional information into the feature processing process. At each time step of the diffusion duration, the mean and variance of the noise are predicted, and a new feature state is generated through the reparameterization trick.
[0137] In some embodiments, when the first video has a long video duration, that is, when the first video is a long video, in order to maintain the consistency of feature identity during the new video generation process, a resampling method for the foreground region is introduced. Schematically, during the j-th denoising process, the j-th noise feature representation and the background feature representation are fused to obtain the j-th conditional noise feature representation; resampling is performed on the region feature representation corresponding to the foreground region in the j-th conditional noise feature representation to obtain a sampled feature representation; the sampled feature representation is fused with the first query vector to obtain a second query vector; the sampled feature representation is fused with the first key vector to obtain a second key vector; attention scores are obtained through the second query vector and the second key vector; the first value vector is weighted and summed based on the attention scores to obtain the (j + 1)-th noise feature representation; the n-th noise feature representation is used as the denoised feature representation for decoding to obtain the second video.
[0138] In some embodiments, the decoding process of the denoised feature representation is implemented by a pre-trained decoder, and the network structure of the pre-trained decoder can be symmetric to the network structure of the encoder used to extract video feature representations.
[0139] Among them, resampling refers to recalculating image features to adapt to a new resolution or size. In some embodiments, the above-mentioned resampling for the foreground region is implemented as ID resampling, that is, the feature vectors of the foreground region have the same ID, and then the feature vectors corresponding to the ID of the foreground region are resampled. The specific implementation is to concatenate the feature vectors from the ID corresponding to the foreground region with the key vector and value vector in the implementation process of the self-attention mechanism, and enhance the ID retention in the foreground region through resampling with additional key vectors and value vectors, so that in the long video processing scenario, the smooth transition between video image frames and the long-term identity consistency of features in the long video can be guaranteed, thereby improving the video quality of the second video.
[0140] In some embodiments, the above-mentioned ID resampling for the foreground region can be achieved by adding a trainable ID resampling adapter in the diffusion network. In one example, the above-mentioned ID resampling adapter can be implemented as a Low-Rank Adaptation of Large Language Models (LoRA model).
[0141] In some embodiments, when the background feature representation input into the diffusion network is generated by a pre-trained context encoder, the pre-trained context encoder includes at least two encoding layers, and the pre-trained diffusion network is divided into at least two network sub-parts, where the number of at least two encoding layers is the same as the number of at least two sub-network parts, and the at least two sub-network parts include the i-th sub-network part, where i is a positive integer. Schematically, in the process of generating the content of the foreground region from the noise feature representation through the i-th sub-network part, the background region in the first video is retained based on the background feature representation output by the i-th encoding layer, and the noise feature representation is iteratively denoised to obtain a second video.
[0142] In one example, when the diffusion network is a network composed of n Transformer blocks, the context encoder for feature integration can be implemented as being achieved by two layers of Transformer blocks, that is, the context encoder includes two encoding layers implemented by Transformer blocks, and the n Transformer blocks in the diffusion network are divided into two groups to obtain two sub-network parts, where each sub-network part includes n / 2 Transformer blocks, and n is an integer greater than or equal to 2. It can be seen that the context encoder only needs to clone the first two layers of the diffusion network, so the network parameters it contains only account for 6% of the backbone network. On the premise of providing strong prior features for the backbone network, it avoids a large increase in the number of model parameters of the overall model, and realizes an efficient improvement in video quality through a lightweight context encoder.
[0143] In some embodiments, in order to reduce the size of tokens during the processing of long videos, the video editing model divides the first video into at least two video segments for data processing, where, among the at least two video segments, the first video image frame of the (i + 1)-th video segment is the last video image frame of the i-th video segment, so as to ensure the visual continuity of the video during grouped processing.
[0144] In summary, when editing the first video, a background feature representation is generated through the video feature representation corresponding to the first video, the mask feature representation corresponding to the foreground region, and the noise feature representation. This background feature representation can indicate the background information of the background region to be retained in the first video. In the process of generating the edited second video from the noise feature representation, the generation of the content of the foreground region and the retention of the content of the background region in the process of generating a new video from the noise feature representation are decoupled, so as to guide the retention of the background region when generating a new video from the noise feature representation through the background feature representation, thereby improving the video quality of the edited second video.
[0145] Exemplarily, such as Figure 7As shown, it shows a schematic diagram of a video editing model provided by an exemplary embodiment of the present application. The model architecture 700 of the video editing model includes a context encoder 710, a DiT network 720, and an ID resampling adapter 730.
[0146] Among them, the context encoder 710 includes an encoding layer 711 and an encoding layer 712. The DiT network 720 includes a sub-network part 721 and a sub-network part 722. Each Transformer block in the DiT network 720 is mounted with an ID resampling adapter 730.
[0147] Schematically, the background feature representation output by the encoding layer 711 in the context encoder 710 is input to each Transformer block in the sub-network part 721 of the DiT network 720 after passing through the screening operation 740. The background feature representation output by the encoding layer 712 in the context encoder 710 is input to each Transformer block in the sub-network part 722 of the DiT network 720 after passing through the screening operation 740. Among them, the above grouping and integration process is shown in Formula 1.
[0148] Formula 1:
[0149]
[0150] Among them, ∈ θ (z t , t, C) i represents the feature of the i-th layer in the DiT network 720, i ∼ [1, n], n is the number of layers, is the feature output by the context encoder 710, z t is the noise feature representation, is the video feature representation with a mask, m resized is the mask feature representation, Z() is a zero linear operation, and t is the time step in the diffusion duration.
[0151] It can be seen that the feature of the encoding layer 711 is added back to the first half of the DiT network 720 as the backbone, and the feature of the DiT network 720 is integrated into the second half, thereby achieving lightweight and efficient context control.
[0152] In addition to the background feature representation output by the context encoder 710, the input of the DiT network 720 also includes a noise feature representation. In some embodiments, the input of the DiT network 720 also includes the text feature representation of the description text indicating the editing requirements.
[0153] In some alternative embodiments, the training of the above video editing model 700 is divided into two stages. One stage is the training stage for the context encoder 710, and the other stage is the training stage for the ID resampling adapter 730.
[0154] Please refer to Figure 8 , which shows a flowchart of a model training method provided by an exemplary embodiment of the present application. The method includes steps 801 to 812.
[0155] Step 801: Obtain a first model to be trained. The first model includes a context encoder to be trained and a pre-trained diffusion network.
[0156] In the embodiments of the present application, in order to train a video editing model, a first model composed of a context encoder to be trained and a pre-trained diffusion network is first obtained. Among them, the above pre-trained diffusion network can be implemented as a publicly available model that can be directly obtained. In one example, the above pre-trained diffusion network is implemented as the Cogvideox model.
[0157] In some embodiments, the network structure of the context encoder to be trained can be obtained by copying at least one processing block in the pre-trained diffusion network, and the network parameters are initialized to obtain the context encoder to be trained.
[0158] In one example, when the pre-trained diffusion network is implemented as the Cogvideox model, the pre-trained diffusion network includes n Transformer blocks, and the network structure of the context encoder to be trained can be implemented as a network structure composed of two Transformer blocks.
[0159] Step 802: Freeze the network parameters of the pre-trained diffusion network in the first model.
[0160] In the embodiments of the present application, since the diffusion network in the first model has been trained, during the training process, there is no need to adjust the network parameters of this part. By freezing the network parameters of the pre-trained diffusion network in the first model, the network parameters of this part will not change during the training process.
[0161] Step 803: Obtain a first sample video and a first sample mask sequence.
[0162] The first sample video is sample data for training the first model. Optionally, the first sample video can be implemented as a video collected by an image acquisition device; or, the first sample video can be implemented as a video produced by a computer device.
[0163] The first sample mask sequence is used to indicate a first sample area in the first sample video where there is an editing requirement. The sample mask images in the first sample mask sequence are images that correspond one-to-one with the video image frames in the first sample video, that is, the number of sample mask images in the first sample mask sequence is the same as the number of video image frames in the first sample video. In some other embodiments, the sample mask images in the first sample mask sequence correspond to the video image frames in which the first sample area in the first sample video changes. The above sample mask images are images having the same size as the video image frames of the first sample video.
[0164] Optionally, the sample mask image can be implemented as a binary image with only two states for pixel values (such as 0 and 1), which is used to represent the masked area and the non-masked area. The masked area in the sample mask image indicates that the mapped area in the corresponding video image frame is blocked, that is, this area is not processed. The non-masked area in the sample mask image indicates that the mapped area in the corresponding video image frame is not blocked, that is, this area can be processed.
[0165] Optionally, the sample mask image can be implemented as a grayscale image with pixel values between 0 and 255. Among them, different grayscale values represent different masking degrees for the pixels in the video image frame. That is to say, the sample mask image is implemented as a change map. The pixel values in the sample mask image can indicate the change intensity of each pixel in the video image frame during editing.
[0166] In some embodiments, the first sample mask sequence is generated according to the annotation information for the first sample area. Schematically, the first sample mask sequence is generated according to the annotation information of the first sample area in the first sample video.
[0167] In some embodiments, the construction process of the dataset corresponding to the first sample video and the first sample mask sequence is as Figure 9 shown, and this process includes five preprocessing steps:
[0168] 910, Collection: Exemplarily, select a publicly available dataset as the data source. For example, Videvo and Pexels;
[0169] 920, Annotation: Exemplarily, for each collected video, a cascaded workflow was implemented for automatic annotation: The Recognize Anything Model (RAM) was used for open-set video tagging to identify the main objects; Based on the detected object labels, the visual-language fusion model for open-set object detection (Marrying DINO with Grounded Pre-Training for Open-Set Object Detection, Grounding DINO) was used to detect the bounding boxes of the objects at fixed intervals; These bounding boxes were used as prompts for the Segment Anything Model 2 (SAM2) to generate high-quality sample mask images;
[0170] 930, Segmentation: Exemplarily, scene transitions may occur when detecting different angles of the same object, resulting in interrupted perspective changes. Therefore, PySceneDetect was used to identify scene transitions and then segment the masks. Subsequently, the sequence was segmented into 10-second intervals, and shorter segments (e.g., segments < 6 seconds) were discarded;
[0171] 940, Screening: Exemplarily, screening was implemented using at least one of the following criteria: 1. Aesthetic quality, evaluated using the Laion-Aesthetic Score Predictor (LAP); 2. Motion intensity, predicted by measuring the optical flow using the deep learning optical flow estimation model (Recurrent All-Pairs Field Transforms, RAFT); 3. Content safety, evaluated using the Stable Diffusion Safety Checker;
[0172] 950, Description Generation: Exemplarily, key frames in the video were uniformly sampled using a vision-language model, and dense video descriptions and detailed descriptions of the masked objects were generated to obtain sample description texts. Optionally, the above vision-language model can be implemented as a large language model, a generative transformer, etc.
[0173] Step 804, Mask the first sample area in the first sample video to obtain a second sample video.
[0174] In the embodiments of the present application, the sample video input to the model during training is the video after masking the first sample area to be generated. Schematically, the first sample video is masked by the first sample mask sequence to obtain a second sample video.
[0175] Step 805: Input the second sample video and the first sample mask sequence into the first model to obtain the first predicted video.
[0176] The first predicted video is the video obtained by the first model after editing the first sample area.
[0177] In some embodiments, the first model further includes a pre-trained encoding network. The encoding network is used to encode the second sample video to obtain a sample video feature representation, and to downsample the first sample mask sequence to obtain a sample mask feature representation. In some embodiments, the first model is also used to generate a sample noise feature representation in Gaussian noise.
[0178] Schematically, the sample video feature representation, the sample mask feature representation, and the sample noise feature representation are input into a context encoder to be trained, and encoded to obtain a sample background feature representation. The sample background feature representation is input into a pre-trained diffusion network to output the first predicted video.
[0179] Step 806: Based on the difference between the first predicted video and the first sample video, iteratively train the network parameters of the context encoder to be trained to obtain the second model.
[0180] In some embodiments, based on the difference between the first predicted video and the first sample video, obtain a first loss value; iteratively train the first model based on the first loss value to obtain the second model.
[0181] In some embodiments, the obtaining manner of the first loss value is implemented as: input the first predicted video and the first sample video into a first preset loss function to obtain the first loss value. Optionally, the above first preset loss function can be implemented as at least one of a cross-entropy loss function, a binary cross-entropy loss (BCE Loss) function, a mean squared error loss (MSE) function, a logarithmic loss function, a least absolute deviations loss (L1 Loss) function, etc., which is not limited herein.
[0182] In the embodiments of the present application, when iteratively adjusting the network parameters of the first model according to the first loss value, only the network parameters of the context encoder are adjusted.
[0183] Schematically, the second model includes a trained context encoder and a pre-trained diffusion network.
[0184] Step 807: Add an adaptation network to be trained to the pre-trained diffusion network in the second model.
[0185] Schematically, after the training of the context encoder is completed, an adaptation network to be trained is added to the diffusion network in the second model. The adaptation network to be trained is used to resample the features of the foreground region during the denoising process of the pre-trained diffusion network.
[0186] Step 808: Freeze the network parameters of the pre-trained diffusion network in the second model and the network parameters of the trained context encoder.
[0187] During the training process of this stage, the network parameters of the diffusion network and the context encoder are fixed to train the additional adaptation network added to the diffusion network in this stage.
[0188] Step 809: Obtain a third sample video and a second sample mask sequence.
[0189] The third sample video is sample data for training the second model. Optionally, the second sample video can be implemented as a video collected by an image acquisition device; or, the third sample video can be implemented as a video produced by a computer device.
[0190] The second sample mask sequence is used to indicate the second sample region with an editing requirement in the third sample video. The sample mask images in the second sample mask sequence are images corresponding one-to-one to the video image frames in the third sample video, that is, the number of sample mask images in the second sample mask sequence is the same as the number of video image frames in the third sample video. In some other embodiments, the sample mask images in the second sample mask sequence correspond to the video image frames in which the first sample region in the third sample video changes. The above sample mask images are images having the same size as the video image frames of the third sample video.
[0191] In some embodiments, the second sample mask sequence is generated according to the annotation information for the second sample region. Schematically, the second sample mask sequence is generated according to the annotation information of the second sample region in the third sample video.
[0192] In some embodiments, the data set formed by the third sample video and the second sample mask sequence can be the same data set as the data set formed by the first sample video and the first sample mask sequence, that is, the first model and the second model are trained on the same data set; it can also be implemented as a different data set, that is, the first model and the second model are trained on different data sets.
[0193] Step 810: Mask the second sample region in the third sample video to obtain a fourth sample video.
[0194] In an embodiment of the present application, during the training process, the sample video input into the model is the video after masking the second sample area to be generated. Schematically, the third sample video is masked by the second sample mask sequence to obtain the fourth sample video.
[0195] Step 811, input the fourth sample video into the second model to obtain the second predicted video.
[0196] Among them, the second predicted video is the video obtained by the second model after editing the second sample area.
[0197] In some embodiments, the second model further includes a pre-trained encoding network, which is used to encode the fourth sample video to obtain the sample video feature representation, and downsample the second sample mask sequence to obtain the sample mask feature representation. In some embodiments, the second model is also used to generate the sample noise feature representation in Gaussian noise.
[0198] Schematically, the sample video feature representation, the sample mask feature representation, and the sample noise feature representation are input into the context encoder that has completed training, encoded to obtain the sample background feature representation, and the sample background feature representation is input into the diffusion network with an adaptation network added to output the second predicted video.
[0199] Schematically, when the adaptation network resamples the second sample area, ID resampling is used, that is, if the feature vectors of the second sample area have the same ID, then the feature vectors corresponding to the ID of the second sample area are resampled.
[0200] In one example, as Figure 10 shown, it shows a schematic diagram of feature generation when the adaptation network in the diffusion network provided by an exemplary embodiment of the present application performs ID resampling, taking this process combined with the grouping operation of the input video as an example.
[0201] In the training stage 1010, the feature representation 1011 corresponding to the video segment can obtain the Q vector 1012, the K vector 1013, and the V vector 1014 in the diffusion network based on the self-attention mechanism. The feature representation 1015 corresponding to the foreground area (sample area) obtains the K vector 1016 and the V vector 1017 through resampling. Among them, the K vector 1013 and the K vector 1016 are concatenated, and the V vector 1014 and the V vector 1017 are concatenated.
[0202] In the application stage 1020, the feature representation 1021 corresponding to the video segment can obtain a Q vector 1022, a K vector 1023, and a V vector 1024 in the diffusion network based on the self-attention mechanism. The feature representation 1025 corresponding to the foreground region (sample region) obtains a K vector 1026 and a V vector 1027 through resampling. Among them, the K vector 1023 and the K vector 1026 are concatenated, and the V vector 1024 and the V vector 1027 are concatenated. Among them, in the application stage 1020, for the resampling process of the foreground region, it is aimed at the feature representation 1025 of the foreground region corresponding to the last video image frame in the previous video segment.
[0203] Step 812: Based on the difference between the second predicted video and the third sample video, iteratively train the network parameters of the to-be-trained adaptation network to obtain a video editing model.
[0204] In some embodiments, based on the difference between the second predicted video and the third sample video, obtain a second loss value; iteratively train the second model based on the second loss value to obtain a video editing model.
[0205] In some embodiments, the obtaining manner of the second loss value is implemented as: inputting the second predicted video and the third sample video into a second preset loss function to obtain the second loss value. Optionally, the above-mentioned second preset loss function can be implemented as at least one of a cross-entropy loss function, a binary cross-entropy loss function, a mean square error loss function, a logarithmic loss function, a least absolute deviation loss function, etc., which is not limited herein.
[0206] In the embodiments of the present application, when iteratively adjusting the network parameters of the second model according to the second loss value, only the network parameters of the adaptation network are adjusted.
[0207] In summary, when editing the first video, a background feature representation is generated through the video feature representation corresponding to the first video, the mask feature representation corresponding to the foreground region, and the noise feature representation. This background feature representation can indicate the background information of the background region to be retained in the first video. In the process of generating the edited second video from the noise feature representation, the context encoder decouples the generation of the content of the foreground region and the retention of the content of the background region in the process of generating a new video from the noise feature representation. Thus, the retention of the background region is guided by the background feature representation when generating a new video from the noise feature representation, thereby improving the video quality of the edited second video.
[0208] Optionally, the video editing model provided by the embodiments of the present application can be applied to diverse video editing scenarios, such as video content repair, video content addition, video content elimination, video content modification, and other scenarios.
[0209] In one example, taking the above video editing model applied to the video restoration scenario as an example, as Figure 11 shown, it shows a flowchart of a video restoration method provided by an exemplary embodiment of the present application. The method includes steps 1101 to 1108.
[0210] Step 1101, obtain the video to be restored and the corresponding mask image sequence of the video to be restored.
[0211] Schematically, the video to be restored is a video with a video restoration requirement. Optionally, the video to be restored can be implemented as a video collected by an image acquisition device; or, the video to be restored can be implemented as a video produced by a computer device. For example, restoring an old low-quality video, or, for example, restoring an interference video with human interference, etc.
[0212] The video image frames of the video to be restored include a background region and a foreground region to be edited. Among them, the foreground region is the region with a restoration requirement in the video to be restored, and the background region is the region where the original image content needs to be retained during the editing of the video to be restored. That is, during the editing of the video to be restored, the content of the background region is retained and the content of the foreground region is restored.
[0213] Schematically, the mask image sequence is used to indicate the position of the foreground region in the video image frames of the video to be restored. The mask image sequence is an image sequence composed of multiple mask images.
[0214] In some embodiments, the mask images in the mask image sequence are images that are in one-to-one correspondence with the video image frames in the video to be restored, that is, the number of mask images in the mask image sequence is the same as the number of video image frames in the video to be restored. In other embodiments, the mask images in the mask image sequence correspond to the video image frames in the video to be restored where the foreground region changes. Among them, the mask image is an image having the same size as the video image frame of the video to be restored.
[0215] Step 1102, mask the video to be restored through the mask image sequence to obtain a masked video.
[0216] Schematically, the mask image sequence and the video to be restored are superimposed frame by frame to obtain a masked video.
[0217] In some embodiments, after obtaining the masked video, a preprocessing operation is performed on the masked video. Optionally, the above preprocessing operation includes video format conversion, frame splitting, noise reduction processing, etc.
[0218] Step 1103, encode the masked video to obtain a video feature representation.
[0219] In some embodiments, the video image frames of the first video are encoded by a pre-trained first encoder to obtain a video feature representation of the first video. Schematically, the first video is input into the pre-trained first encoder to obtain the video feature representation.
[0220] Optionally, the pre-trained first encoder can be implemented as an autoencoder, variational autoencoder, deep autoencoder, recurrent neural network encoder, convolutional neural network encoder, attention mechanism-based encoder, etc., which is not limited herein.
[0221] Step 1104: Downsample the mask image sequence to obtain a mask feature representation.
[0222] The mask feature representation corresponding to the mask image sequence is obtained by downsampling. Schematically, the mask image sequence of the video to be repaired is acquired; the mask image sequence is downsampled to obtain the mask feature representation. In one example, each of the multiple mask images in the mask image sequence is downsampled to obtain a mask image feature representation, and the mask image feature representations are concatenated in the sequence order of the multiple mask images to obtain the mask feature representation.
[0223] Optionally, the downsampling can be implemented as at least one of the following: nearest neighbor interpolation; bilinear interpolation; bicubic interpolation; pooling, etc.
[0224] Step 1105: Sample from Gaussian noise to generate a noise feature representation.
[0225] Schematically, sample noise is sampled from Gaussian noise, and the noise feature representation is generated through the sample noise.
[0226] Optionally, the sample noise is a signal interference applied to the original video image frames, which can cause changes in the image information or pixel brightness in the original video image frames. Optionally, the sample noise can be Gaussian noise, salt-and-pepper noise, etc. When the sample noise is Gaussian noise, each pixel point in the video image frame after adding noise is a pixel point to which the noise is applied.
[0227] In some embodiments, noise is added to the video image frames of the video to be repaired through the sample noise to obtain a noisy video, and the video image frames of the noisy video are feature-extracted to obtain the noise feature representation. Schematically, the sample noise is acquired; the sample noise is added to each video image frame of the video to be repaired to obtain a noisy video; the noisy video is encoded to obtain the noise feature representation.
[0228] Step 1106: Concatenate the video feature representation, the mask feature representation, and the noise feature representation and input them into a context encoder to generate an intermediate encoded representation through the context encoder.
[0229] In some embodiments, before jointly encoding the video feature representation, the mask feature representation, and the noise feature representation through the context encoder, the video feature representation, the mask feature representation, and the noise feature representation are fused for feature fusion. Schematically, the video feature representation, the mask feature representation, and the noise feature representation are fused to obtain a fused feature representation; the fused feature representation is input into the context encoder to obtain a background feature representation corresponding to the background region.
[0230] Optionally, the feature fusion method for the video feature representation, the mask feature representation, and the noise feature representation can be implemented as at least one of the following: feature concatenation; feature addition; feature averaging; feature stacking; deep learning-based fusion.
[0231] In one example, the feature fusion method between the video feature representation, the mask feature representation, and the noise feature representation can be implemented as three-way feature concatenation on a pixel-by-pixel basis. Schematically, the video feature representation includes first sub-feature representations corresponding to each pixel point in the video image frames of the video to be repaired, the mask feature representation includes second sub-feature representations corresponding to each pixel point in the video image frames of the video to be repaired, and the noise feature representation includes third sub-feature representations corresponding to each pixel point in the video image frames of the video to be repaired. Fusing the video feature representation, the mask feature representation, and the noise feature representation is implemented as: taking a single pixel point in the video image frames of the video to be repaired as a unit, concatenating the first sub-feature representation, the second sub-feature representation, and the third sub-feature representation corresponding to each pixel point to obtain an image frame feature representation corresponding to the video image frame; according to the temporal relationship between the video image frames in the video to be repaired, concatenating the image frame feature representations to obtain a fused feature representation.
[0232] Schematically, the fused feature representation is input into the context encoder to obtain an intermediate encoded representation.
[0233] Step 1107, screening out the background feature representation corresponding to the background region from the intermediate encoded representation.
[0234] In some embodiments, since after the context encoder encodes the fused feature representation, the obtained feature representation includes the relevant features of all pixel positions in the video image frames of the video to be repaired, and the features provided by the background feature representation in the downstream noise reduction process are mainly used to indicate the image content of the background region to be retained, therefore, the output of the context encoder is filtered so that the background feature representation input to the next stage only retains the features corresponding to the background region.
[0235] Schematically, the background feature representation corresponding to the background region is screened out from the intermediate encoded representation.
[0236] In some embodiments, feature screening for the intermediate coding representation is implemented based on the feature identifier (ID), where the identifier of the feature is used to indicate the label or number for identifying the feature. Exemplarily, features of the background region and the foreground region have different IDs, and through the ID, it can be determined whether the feature elements in the intermediate coding representation belong to the background region or the foreground region, so as to screen out the feature elements corresponding to the foreground region from the intermediate coding representation.
[0237] Step 1108: Input the background feature representation into the backbone network. Taking the restoration of the foreground region in the video to be restored as the noise reduction condition, perform noise reduction on the noise feature representation based on the background feature representation to obtain the restored video as the restoration result of the video to be restored.
[0238] In the embodiments of the present application, the backbone network includes a diffusion network added with an ID resampling adapter. Schematically, in the j-th round of noise reduction process, the j-th noise feature representation and the background feature representation are fused to obtain the j-th conditional noise feature representation; perform resampling on the region feature representation corresponding to the foreground region in the j-th conditional noise feature representation to obtain the sampled feature representation; fuse the sampled feature representation with the first query vector to obtain the second query vector; fuse the sampled feature representation with the first key vector to obtain the second key vector; obtain the attention score through the second query vector and the second key vector; perform weighted summation on the first value vector based on the attention score to obtain the (j + 1)-th noise feature representation; use the n-th noise feature representation as the denoised feature representation to perform decoding to obtain the restored video.
[0239] As shown in Table 1, it shows the comparison of the effects of the model of the present application with other models in the video restoration scenario. Among them, there are three indicator categories: Masked Region Preservation, TextAlignment, and Video Quality. Among them, the indicators used for Masked Region Preservation include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), Mean Squared Error (MSE), Mean Absolute Error (MAE). The indicators used for TextAlignment include Contrastive Language-Image Pre-trainingSim (CLIP Sim), Multimodal CLIP Sim (CLIP Sim(M)). The indicators used for Video Quality include Fréchet Video Distance (FVID). This evaluation process is implemented on three datasets, including VPBench-S, VPBench-L, and Davis. The models compared with the model of the present application include Model 1, Model 2, and Model 3.
[0240] Table 1
[0241]
[0242] Please refer to Figure 12 , which shows a comparison diagram of the generated results 1200 of each model in Table 1. Combining the data shown in Table 1, it can be seen that: Combining the quantitative comparison results on the datasets VPBench and Davis, comparing the restoration results of the non-generative Model 1, the generative Model 2, and the strong benchmark Model 3, Model 3 uses an image restoration model to repair the first frame and uses an image-to-video backbone network to propagate the results through a latent mixing operation. In the segmentation-based dataset VPBench, Model 1 and Model 2 perform the worst in most indicators, mainly because they are unable to repair fully occluded objects and there are difficulties in balancing the two competing tasks of background retention and foreground generation in a single backbone architecture. In the randomly masked benchmark Davis, Model 1 shows improvement by using some background information. However, the model of the present application achieves the optimal performance in both segmentation (standard and long-term) and randomly masked tests through its dual-branch architecture that effectively decouples background retention and foreground generation.
[0243] In another example, the embodiments of the present application provide a video editing model, which combines a vision-language model to generate a modified description based on user editing instructions and a source description, and can implement the process of video content editing based on the modified description, for example, modifying the content in the video. In this scenario, as shown in Table II, it shows the effect comparison between the model of the present application and other models in the video restoration scenario. Among them, there are three index categories: mask area retention, text alignment, and video quality. The indexes corresponding to each category are the same as those in Table I and will not be elaborated here. This evaluation process is implemented on two datasets, including Standard and Long. The models compared with the model of the present application include Model 4, Model 5, and Model 6.
[0244] Table II
[0245]
[0246] Please refer to Figure 13 , which shows a comparison schematic diagram of the generation results 1300 of each model in Table I. Combining the data shown in Table I, it can be known that: in the quantitative comparison on the dataset VPBench, the editing results of the inversion-based Model 4, the DiT-based Model 5, and the end-to-end Model 6 are compared. For the standard and long videos in VPBench, it can be seen from the data in the table that the model of the present application has achieved superior performance, even exceeding the end-to-end Model 6. This success can be attributed to its dual-branch architecture, which ensures excellent background retention and foreground generation capabilities, ensures tight alignment of the edited area with the editing instructions while maintaining high fidelity in the non-edited area, and maintains ID consistency in long videos through resampling of the repair area ID. Figure 13 shows a qualitative comparison with previous video restoration methods. The model of the present application shows superior performance in maintaining visual fidelity and text prompt consistency. For example, the model of the present application successfully generated a seamless animation of a future spaceship flying through the sky, maintaining a smooth temporal transition and precise background boundaries throughout the removal process, avoiding artifacts observed in Model 6.
[0247] It should be noted that, before collecting relevant data of the user and during the process of collecting relevant data of the user, this application can display a prompt interface, a pop-up window or output a voice prompt message. The prompt interface, pop-up window or voice prompt message is used to prompt the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation of the user on the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation of the user on the prompt interface or the pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by this application is collected with the consent and authorization of the user, and the collection, use and processing of relevant user data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0248] Please refer to Figure 14 , which shows a structural block diagram of an editing device for a video provided by an exemplary embodiment of this application. The device includes the following modules:
[0249] An acquisition module 1410, configured to acquire a video feature representation of a first video, where the video image frames of the first video include a background region and a foreground region;
[0250] The acquisition module 1410 is further configured to acquire a mask feature representation of a mask image sequence, where the mask image sequence is used to indicate the position of the foreground region in the video image frames of the first video; and acquire a noise feature representation of sampled noise;
[0251] An encoding module 1420, configured to jointly encode the video feature representation, the mask feature representation, and the noise feature representation to obtain a background feature representation corresponding to the background region, where the background feature representation is used to indicate background information in the first video;
[0252] A prediction module 1430, configured to use editing the foreground region in the first video as a noise reduction condition, and denoise the noise feature representation based on the background feature representation to obtain a second video as an editing result of the first video.
[0253] In some optional embodiments, the prediction module 1430 is further configured to input the video feature representation, the mask feature representation, and the noise feature representation into a pre-trained context encoder to obtain the background feature representation corresponding to the background region, where the context encoder is used to extract semantic features of the video image frames.
[0254] In some optional embodiments, as Figure 15 shown, the encoding module 1420 further includes:
[0255] A fusion unit 1421, configured to fuse the video feature representation, the mask feature representation, and the noise feature representation to obtain a fused feature representation;
[0256] An encoding unit 1422, configured to input the fused feature representation into the context encoder to obtain the background feature representation corresponding to the background region.
[0257] In some alternative embodiments, the video feature representation includes first sub-feature representations respectively corresponding to each pixel point in the video image frames of the first video, the mask feature representation includes second sub-feature representations respectively corresponding to each pixel point in the video image frames of the first video, and the noise feature representation includes third sub-feature representations respectively corresponding to each pixel point in the video image frames of the first video;
[0258] The fusion unit 1421 is further configured to concatenate the first sub-feature representation, the second sub-feature representation, and the third sub-feature representation corresponding to each pixel point in a single pixel point of the video image frames in the first video to obtain an image frame feature representation corresponding to the video image frames; and splice the image frame feature representations according to the temporal relationship between the video image frames in the first video to obtain the fused feature representation.
[0259] In some alternative embodiments, the encoding unit 1422 is further configured to input the fused feature representation into the context encoder to obtain an intermediate encoding representation;
[0260] The encoding module 1420 further includes:
[0261] A screening unit 1423, configured to screen out the background feature representation corresponding to the background region from the intermediate encoding representation.
[0262] In some alternative embodiments, the prediction module 1430 is further configured to input the background feature representation and the noise feature representation into a pre-trained diffusion network to obtain the second video, and the diffusion network is configured to retain the background region in the first video according to the background feature representation and generate the content of the foreground region from the noise feature representation.
[0263] In some alternative embodiments, the background feature representation is generated by a pre-trained context encoder, the pre-trained context encoder includes at least two encoding layers, the pre-trained diffusion network includes at least two sub-network parts, the number of the at least two encoding layers is the same as the number of the at least two sub-network parts, and the at least two sub-network parts include the i-th sub-network part;
[0264] The prediction module 1430 is further configured to retain the background region in the first video based on the background feature representation output by the i-th encoding layer during the process of generating the content of the foreground region from the noise feature representation through the i-th sub-network part, and iteratively denoise the noise feature representation to obtain the second video, where i is a positive integer.
[0265] In some alternative embodiments, the prediction module 1430 is further configured to, during the j-th denoising process, fuse the j-th noise feature representation and the background feature representation to obtain the j-th conditional noise feature representation, where j is a positive integer; denoise the j-th conditional noise feature representation based on the self-attention mechanism to obtain the (j + 1)-th noise feature representation; and use the n-th noise feature representation as the denoised feature representation to perform decoding to obtain the second video.
[0266] In some alternative embodiments, the prediction module 1430 is further configured to perform a linear transformation on the j-th conditional noise feature representation to obtain a first query vector, a first key vector, and a first value vector; obtain an attention score through the first query vector and the first key vector, where the attention score is used to indicate the correlation between pixels in the image frame; and perform weighted summation on the first value vector based on the attention score to obtain the (j + 1)-th noise feature representation.
[0267] In some alternative embodiments, the prediction module 1430 is further configured to perform resampling on the region feature representation corresponding to the foreground region in the j-th conditional noise feature representation to obtain a sampled feature representation; fuse the sampled feature representation with the first query vector to obtain a second query vector; fuse the sampled feature representation with the first key vector to obtain a second key vector; and obtain an attention score through the second query vector and the second key vector.
[0268] In some alternative embodiments, the acquisition module 1410 is further configured to acquire annotation information for indicating the foreground region in the first video.
[0269] The apparatus further includes:
[0270] A preprocessing module 1440, configured to mask the foreground region in the first video based on the annotation information to obtain a masked video.
[0271] The preprocessing module 1440 is further configured to encode the masked video to obtain the video feature representation.
[0272] In some alternative embodiments, the acquisition module 1410 is further configured to acquire the masked image sequence of the first video.
[0273] The preprocessing module 1440 is further configured to perform downsampling on the mask image sequence to obtain the mask feature representation.
[0274] In some alternative embodiments, the preprocessing module 1440 is further configured to sample standard Gaussian noise to obtain the sampled noise; add noise to the video image frames of the first video according to the sampled noise and the diffusion duration, where the diffusion duration is used to indicate the number of steps required to generate the second video from the noise feature representation during the iterative noise reduction process; and encode the noise video to obtain the noise feature representation.
[0275] In some alternative embodiments, the obtaining module 1410 is further configured to obtain a description text, where the description text is used to indicate the editing requirements for the foreground region;
[0276] The preprocessing module 1440 is further configured to extract a text feature representation of the description text, where the text feature representation is used to indicate the semantic features of the description text;
[0277] The prediction module 1430 is further configured to use editing the foreground region in the first video as a noise reduction condition, and perform noise reduction on the noise feature representation based on the background feature representation and the text feature representation to obtain a second video as the editing result of the first video.
[0278] In some alternative embodiments, the apparatus further includes: a training module 1450, including:
[0279] A freezing unit 1451, configured to freeze the network parameters of the pre-trained diffusion network in the first model;
[0280] An obtaining unit 1452, configured to obtain a first sample video and a first sample mask sequence, where the first sample mask sequence is used to indicate a first sample region in the first sample video where there are editing requirements;
[0281] A preprocessing unit 1453, configured to mask the first sample region in the first sample video to obtain a second sample video;
[0282] A training unit 1454, configured to input the second sample video and the first sample mask sequence into the first model to obtain a first predicted video, where the first predicted video is a video obtained by the first model after editing the first sample region;
[0283] The training unit 1454 is further configured to iteratively train the network parameters of the context encoder to be trained based on the difference between the first predicted video and the first sample video, so as to obtain a second model, where the second model includes the trained context encoder and the pre-trained diffusion network;
[0284] The obtaining unit 1452 is further configured to add an adaptation network to be trained to the pre-trained diffusion network in the second model, where the adaptation network to be trained is configured to resample the features of the foreground region during the denoising process of the pre-trained diffusion network;
[0285] The freezing unit 1451 is further configured to freeze the network parameters of the pre-trained diffusion network in the second model and the network parameters of the trained context encoder;
[0286] The obtaining unit 1452 is further configured to obtain a third sample video and a second sample mask sequence, where the second sample mask sequence is used to indicate a second sample region with an editing requirement in the third sample video;
[0287] The preprocessing unit 1453 is further configured to mask the second sample region in the third sample video to obtain a fourth sample video;
[0288] The training unit 1454 is further configured to input the fourth sample video into the second model to obtain a second predicted video, where the second predicted video is a video obtained by editing the second sample region by the second model;
[0289] The training unit 1454 is further configured to iteratively train the network parameters of the adaptation network to be trained based on the difference between the second predicted video and the third sample video to obtain the video editing model.
[0290] It should be noted that: for the video editing device provided in the above embodiment, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the video editing device provided in the above embodiment and the video editing method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.
[0291] Figure 16 The structural schematic diagram of a server provided by an exemplary embodiment of the present application is shown. Specifically, it includes the following structure.
[0292] The server 1600 includes a Central Processing Unit (CPU) 1601, a system memory 1604 including a Random Access Memory (RAM) 1602 and a Read Only Memory (ROM) 1603, and a system bus 1605 connecting the system memory 1604 and the central processing unit 1601. The server 1600 also includes a mass storage device 1606 for storing an operating system 1613, application programs 1614, and other program modules 1615.
[0293] The mass storage device 1606 is connected to the central processing unit 1601 through a mass storage controller (not shown) connected to the system bus 1605. The mass storage device 1606 and its associated computer-readable medium provide non-volatile storage for the server 1600. That is to say, the mass storage device 1606 can include computer-readable media (not shown) such as a hard disk or a Compact Disc Read Only Memory (CD-ROM) drive.
[0294] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, Erasable Programmable Read Only Memory (EPROM), Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory, or other solid-state memory technologies, CD-ROM, Digital Versatile Disc (DVD), or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage, or other magnetic storage devices. Of course, those skilled in the art know that computer storage media is not limited to the above several. The above-mentioned system memory 1604 and mass storage device 1606 can be collectively referred to as memory.
[0295] According to various embodiments of the present application, the server 1600 may also run on a remote computer on the network through a network such as the Internet. That is, the server 1600 may be connected to the network 1612 through the network interface unit 1611 connected to the system bus 1605. Or rather, the network interface unit 1611 may also be used to connect to other types of networks or remote computer systems (not shown).
[0296] The above-mentioned memory further includes one or more programs, and the one or more programs are stored in the memory and configured to be executed by the CPU.
[0297] Embodiments of the present application further provide a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the video editing method provided by the above-mentioned method embodiments. Optionally, the computer device may be a terminal or a server.
[0298] Embodiments of the present application further provide a computer-readable storage medium, on which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the video editing method provided by the above-mentioned method embodiments.
[0299] Embodiments of the present application further provide a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video editing method described in any one of the above embodiments.
[0300] Optionally, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid state drive (SSD, Solid State Drives), or optical disc, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance RandomAccess Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory). The above serial numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0301] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0302] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A video editing method, characterized in that: The method comprises: Obtaining a video feature representation of a first video, wherein a video image frame of the first video includes a background area and a foreground area; Acquire a mask feature representation of a mask image sequence, wherein the mask image sequence is used to indicate a position of the foreground area in the video image frame in the first video; and acquire a noise feature representation of sampled noise; The video feature representation, the mask feature representation and the noise feature representation are jointly encoded to obtain a background feature representation corresponding to the background area, wherein the background feature representation is used to indicate background information in the first video; Taking editing the foreground area in the first video as a noise reduction condition, the noise feature representation is denoised based on the background feature representation, and a second video is obtained as an editing result of the first video.
2. The method according to claim 1, characterized in that The step of jointly encoding the video feature representation, the mask feature representation, and the noise feature representation to obtain a background feature representation corresponding to the background area includes: The video feature representation, the mask feature representation and the noise feature representation are input into a pre-trained context encoder to obtain the background feature representation corresponding to the background area. The context encoder is used to extract semantic features of the video image frame.
3. The method according to claim 2, characterized in that The step of inputting the video feature representation, the mask feature representation, and the noise feature representation into a pre-trained context encoder to obtain the background feature representation corresponding to the background area includes: fusing the video feature representation, the mask feature representation and the noise feature representation to obtain a fused feature representation; The fused feature representation is input into the context encoder to obtain the background feature representation corresponding to the background area.
4. The method according to claim 3, characterized in that: The video feature representation includes a first sub-feature representation corresponding to each pixel point in the video image frame of the first video, the mask feature representation includes a second sub-feature representation corresponding to each pixel point in the video image frame of the first video, and the noise feature representation includes a third sub-feature representation corresponding to each pixel point in the video image frame of the first video; The fusing the video feature representation, the mask feature representation and the noise feature representation to obtain a fused feature representation comprises: Taking a single pixel of the video image frame in the first video as a unit, the first sub-feature representation, the second sub-feature representation, and the third sub-feature representation corresponding to each pixel are connected in series to obtain an image frame feature representation corresponding to the video image frame; The image frame feature representations are spliced according to the temporal relationship between the video image frames in the first video to obtain the fused feature representation.
5. The method according to claim 3 or 4, characterized in that: The step of inputting the fused feature representation into the context encoder to obtain the background feature representation corresponding to the background area includes: Inputting the fused feature representation into the context encoder to obtain an intermediate encoding representation; The background feature representation corresponding to the background area is obtained by screening from the intermediate coding representation.
6. The method according to any one of claims 1 to 4, characterized in that: The step of taking editing the foreground area in the first video as a noise reduction condition, reducing the noise feature representation based on the background feature representation, and obtaining a second video as an editing result of the first video includes: The background feature representation and the noise feature representation are input into a pre-trained diffusion network to obtain the second video, wherein the diffusion network is used to retain the background area in the first video according to the background feature representation and generate the content of the foreground area from the noise feature representation.
7. The method according to claim 6, characterized in that The background feature representation is generated by a pre-trained context encoder, the pre-trained context encoder includes at least two encoding layers, the pre-trained diffusion network includes at least two sub-network parts, the number of the at least two encoding layers is the same as the number of the at least two sub-network parts, and the at least two sub-network parts include the i-th sub-network part; The step of inputting the background feature representation and the noise feature representation into a pre-trained diffusion network to obtain the second video includes: In the process of generating the content of the foreground area from the noise feature representation through the i-th sub-network part, the background area in the first video is retained based on the background feature representation output by the i-th encoding layer, and the noise feature representation is iteratively denoised to obtain the second video, where i is a positive integer.
8. The method according to any one of claims 1 to 4, characterized in that: The step of taking editing the foreground area in the first video as a noise reduction condition, reducing the noise feature representation based on the background feature representation, and obtaining a second video as an editing result of the first video includes: In the j-th round of denoising, the j-th noise feature representation is fused with the background feature representation to obtain the j-th conditional noise feature representation, where j is a positive integer; De-noising the j-th conditional noise feature representation based on the self-attention mechanism to obtain the j+1-th noise feature representation; The nth noise feature representation is decoded as the denoising feature representation to obtain the second video.
9. The method according to claim 8, characterized in that The denoising the j-th conditional noise feature representation based on the self-attention mechanism to obtain the j+1-th noise feature representation includes: Performing a linear transformation on the j-th conditional noise feature representation to obtain a first query vector, a first key vector, and a first value vector; Obtain an attention score through the first query vector and the first key vector, where the attention score is used to indicate the correlation between pixels in the image frame; The first value vector is weightedly summed based on the attention score to obtain the j+1th noise feature representation.
10. The method according to claim 9, characterized in that The acquiring the attention score by using the first query vector and the first key vector includes: Resampling the region feature representation corresponding to the foreground region in the j-th conditional noise feature representation to obtain a sampled feature representation; Fusion the sampled feature representation with the first query vector to obtain a second query vector; Fusion the sampled feature representation with the first key vector to obtain a second key vector; An attention score is obtained by using the second query vector and the second key vector.
11. The method according to any one of claims 1 to 4, characterized in that: The obtaining of the video feature representation of the first video includes: Acquire annotation information for indicating the foreground area in the first video; Masking the foreground area in the first video based on the annotation information to obtain a masked video; The mask video is encoded to obtain the video feature representation.
12. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a mask feature representation of a mask image sequence includes: Acquire the mask image sequence of the first video; Downsampling is performed on the mask image sequence to obtain the mask feature representation.
13. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a noise characteristic representation of the sampling noise includes: Sampling standard Gaussian noise to obtain the sampled noise; Adding noise to the video image frame of the first video according to the sampling noise and the diffusion duration to obtain a noisy video, wherein the diffusion duration is used to indicate the number of steps required to generate the second video from the noise feature representation in an iterative noise reduction process; The noisy video is encoded to obtain the noise feature representation.
14. The method according to any one of claims 1 to 4, characterized in that: The step of taking editing the foreground area in the first video as a noise reduction condition, reducing the noise feature representation based on the background feature representation, and obtaining a second video as an editing result of the first video includes: Acquire a description text, where the description text is used to indicate an editing requirement for the foreground area; Extracting and obtaining a text feature representation of the description text, wherein the text feature representation is used to indicate a semantic feature of the description text; Taking editing the foreground area in the first video as a noise reduction condition, the noise feature representation is denoised based on the background feature representation and the text feature representation, and a second video is obtained as an editing result of the first video.
15. The method according to any one of claims 1 to 4, characterized in that: The method is implemented by a video editing model, and the video editing model is obtained by training a first model, wherein the first model includes a context encoder to be trained and a pre-trained diffusion network, and the training process of the first model includes: Freezing network parameters of the pre-trained diffusion network in the first model; Acquire a first sample video and a first sample mask sequence, where the first sample mask sequence is used to indicate a first sample area in the first sample video where editing is required; Masking the first sample area in the first sample video to obtain a second sample video; Inputting the second sample video and the first sample mask sequence into the first model to obtain a first predicted video, where the first predicted video is a video obtained after the first model edits the first sample area; Iteratively training network parameters of the context encoder to be trained based on the difference between the first predicted video and the first sample video to obtain a second model, wherein the second model includes the trained context encoder and the pre-trained diffusion network; Adding an adaptation network to be trained in the pre-trained diffusion network in the second model, wherein the adaptation network to be trained is used to resample the features of the foreground area during the process of performing noise reduction on the pre-trained diffusion network; Freezing the network parameters of the pre-trained diffusion network and the network parameters of the trained context encoder in the second model; Acquire a third sample video and a second sample mask sequence, where the second sample mask sequence is used to indicate a second sample area in the third sample video where editing is required; Masking the second sample area in the third sample video to obtain a fourth sample video; Inputting the fourth sample video into the second model to obtain a second predicted video, where the second predicted video is a video obtained after the second sample area is edited by the second model; Based on the difference between the second predicted video and the third sample video, the network parameters of the adaptation network to be trained are iteratively trained to obtain the video editing model.
16. A video editing device, characterized in that: The device comprises: An acquisition module, configured to acquire a video feature representation of a first video, wherein a video image frame of the first video includes a background area and a foreground area; The acquisition module is further used to acquire a mask feature representation of a mask image sequence, wherein the mask image sequence is used to indicate a position of the foreground area in the video image frame in the first video; and to acquire a noise feature representation of sampled noise; an encoding module, used for jointly encoding the video feature representation, the mask feature representation and the noise feature representation to obtain a background feature representation corresponding to the background area, wherein the background feature representation is used to indicate background information in the first video; A prediction module is used to take the foreground area in the first video as a noise reduction condition, reduce the noise feature representation based on the background feature representation, and obtain a second video as the editing result of the first video.
17. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the video editing method as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that: At least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by the processor to implement the video editing method as described in any one of claims 1 to 15.
19. A computer program product, characterized in that It comprises a computer program or an instruction, which, when executed by a processor, implements the video editing method as described in any one of claims 1 to 15.
Citation Information
Cited By
Video processing method and device, readable storage medium and program product
CN121619457A