A video processing method, device, medium and program product

Through optical flow network processing and diffusion model editing, the problem of limited efficiency and accuracy of existing video editing models when editing real videos is solved, and efficient and accurate video editing is achieved.

CN119906869BActive Publication Date: 2025-06-17LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510387243.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-17
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing video editing model has limited efficiency and accuracy when editing real videos, and the fine-tuning process consumes a lot of computing resources, extending the model inference time.

Method used

By obtaining the initial video and editing information, the preset optical flow network is used to perform forward and reverse optical flow processing, and the video is edited based on the editing information and optical flow motion information in combination with the diffusion model to realize video editing.

Benefits of technology

This method does not require adjusting model parameters, saves computing resources and model inference time, improves video editing efficiency and accuracy, and enhances the editing ability of the diffusion model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119906869B_ABST
    Figure CN119906869B_ABST
Patent Text Reader

Abstract

The present application discloses a video processing method, device, medium and program product, which relates to the field of computer technology. In the present application, the diffusion model can perceive the adjacent frame optical flow motion information in the real initial video, and the adjacent frame optical flow motion information corresponds to the temporal motion related information in the video. Then, the diffusion model edits the initial video into a new video according to the editing information and the adjacent frame optical flow motion information, so that the diffusion model not only has the editing ability of real videos, but also improves the video editing accuracy by virtue of the temporal motion related information in the video. This solution does not require adjusting model parameters, saves computing resources and model inference time, and thus also improves the video editing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a video processing method, device, medium, and program product. Background Art

[0002] Currently, the editing ability of existing video editing models for real videos is limited. Although the model parameters of the video editing model can be fine-tuned using real samples, each fine-tuning consumes a certain amount of computing resources, which not only easily causes waste of computing power but also prolongs the model inference time. Among them, fine-tuning refers to further training on a dataset for a specific task on the basis of using a pre-trained model to adjust the model parameters so that it can better adapt to the target task. During the fine-tuning process, most layers of the pre-trained model are usually frozen, and only the newly added layers or a small number of key layers are trained. This can not only retain the features learned by the pre-trained model but also quickly adapt to the specific requirements of the new task. In addition, selecting appropriate learning rates and the number of training epochs is also the key to successful fine-tuning. It can be seen that fine-tuning refers to small-scale training on the basis of a pre-trained model for specific task objectives and task data to achieve minor adjustments to the parameters of the pre-trained model, and finally obtain a model adapted to specific tasks and data.

[0003] Therefore, how to improve the editing efficiency and accuracy of video editing models is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a video processing method, device, medium, and program product to improve the editing efficiency and accuracy of video editing models.

[0005] In a first aspect, this application provides a video processing method, including:

[0006] Obtaining an initial video and editing information;

[0007] Performing forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video;

[0008] Performing backward optical flow processing on the initial video using the optical flow network to obtain a backward optical flow video;

[0009] Using a preset diffusion model, editing the initial video into a new video according to the editing information and the adjacent frame optical flow motion information carried in the forward optical flow video and the backward optical flow video.

[0010] In a second aspect, this application provides a video processing device, including:

[0011] An obtaining module, configured to obtain an initial video and editing information;

[0012] A forward optical flow processing module, configured to perform forward optical flow processing on the initial video by using a preset optical flow network to obtain a forward optical flow video;

[0013] A backward optical flow processing module, configured to perform backward optical flow processing on the initial video by using the optical flow network to obtain a backward optical flow video;

[0014] An editing module, configured to edit the initial video into a new video by using a preset diffusion model according to the editing information and the adjacent-frame optical flow motion information carried in the forward optical flow video and the backward optical flow video.

[0015] In a third aspect, the present application provides an electronic device, including:

[0016] A memory, configured to store a computer program;

[0017] A processor, configured to execute the computer program to implement the video processing method disclosed above.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein when the computer program is executed by a processor, the video processing method disclosed above is implemented.

[0019] In a fifth aspect, the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the video processing method disclosed above are implemented.

[0020] Through the present application, the diffusion model (i.e., the video editing model) can perceive the adjacent-frame optical flow motion information in the real initial video, and the adjacent-frame optical flow motion information corresponds to the temporal motion-related information in the video. Then, the diffusion model edits the initial video into a new video according to the editing information and the adjacent-frame optical flow motion information, so that the diffusion model not only has the editing ability of real videos, but also improves the video editing accuracy by virtue of the temporal motion-related information in the video. This solution does not require adjusting model parameters, saves computing resources and model inference time, and thus also improves the editing efficiency of the diffusion model for videos.

[0021] Correspondingly, a video processing device, medium, and program product provided by the present application also have the above technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 Flow chart of a video processing method disclosed in this application;

[0024] Figure 2 Comparison diagram of an optical flow map and an original video frame disclosed in this application;

[0025] Figure 3 Schematic diagram of the processing flow of a video description model disclosed in this application;

[0026] Figure 4 Schematic diagram of the processing flow of another video description model disclosed in this application;

[0027] Figure 5 Schematic diagram of the training process of a video reconstruction model disclosed in this application;

[0028] Figure 6 Schematic diagram of the processing flow of a trained diffusion model disclosed in this application;

[0029] Figure 7 Server structure diagram provided by this application;

[0030] Figure 8 Terminal structure diagram provided by this application. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.

[0032] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0033] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be made in conjunction with the accompanying drawings and specific implementation manners.

[0034] See Figure 1 As shown, an embodiment of this application discloses a video processing method, including:

[0035] S101. Obtain the initial video and editing information.

[0036] In this embodiment, the initial video can be a video recorded from the real world, or an animation or comic video, an AI - synthesized video, an edited video, etc. The editing information may include text descriptions of the corresponding modifications to be made to the video's tone, style, characters, animals, etc. For example, the editing information can be: control the video to increase the picture brightness to 20 within 3 minutes, add subtitles, the export format is MP4, and the export resolution is 1080P.

[0037] S102. Perform forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video.

[0038] S103. Perform backward optical flow processing on the initial video using the optical flow network to obtain a backward optical flow video.

[0039] In this embodiment, the optical flow network can be implemented based on structures such as convolutional neural networks. The optical flow network can calculate the optical flow value between each frame in the initial video and its next frame, and then convert the optical flow value into an optical flow map according to the amplitude. Based on all the optical flow maps, the corresponding optical flow video can be generated. Please refer to Figure 2 , the optical flow map converted from the optical flow value between the first frame image and the second frame image in the initial video is Figure 2 B in, where A is the first frame image. The Munsell hue system is adopted here. The forward optical flow video generates and converts optical flow maps in the forward order of each video frame in the initial video; the backward optical flow video generates and converts optical flow maps in the reverse order of each video frame in the initial video; that is to say, the generation ideas of the forward optical flow video and the backward optical flow video are similar, except that the backward optical flow video needs to arrange the video frames in reverse order before generating and converting the optical flow maps.

[0040] In one implementation, performing forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video includes: using the optical flow network to calculate the optical flow value between adjacent frames in the direction from the first frame to the last frame in the initial video, and converting the optical flow value into the corresponding optical flow map, and summarizing each optical flow map to obtain the forward optical flow video; correspondingly, performing backward optical flow processing on the initial video using the optical flow network to obtain a backward optical flow video includes: using the optical flow network to calculate the optical flow value between adjacent frames in the direction from the last frame to the first frame in the initial video, and converting the optical flow value into the corresponding optical flow map, and summarizing each optical flow map to obtain the backward optical flow video.

[0041] S104. Use a preset diffusion model to edit the initial video into a new video according to the editing information and the adjacent - frame optical flow motion information carried in the forward optical flow video and the backward optical flow video.

[0042] It should be noted that the diffusion model includes two main processes: the forward diffusion process (i.e., the noise-adding training process) and the reverse diffusion process (i.e., the denoising inference process). In the forward diffusion process, the model starts from the original data and gradually adds Gaussian noise to it until the data completely becomes pure Gaussian noise. The reverse diffusion process is the reverse operation of the forward process, that is, starting from pure Gaussian noise, gradually removing the noise, and finally restoring the original data. Specifically for model training, it can be achieved through a parameterized neural network (such as a noise predictor), which learns how to predict and remove the noise added at each step.

[0043] In this embodiment, the diffusion model perceives the adjacent-frame optical flow motion information in the real initial video, and the adjacent-frame optical flow motion information corresponds to the temporal motion-related information in the video. Then, the diffusion model edits the initial video into a new video according to the editing information and the adjacent-frame optical flow motion information, so that the diffusion model not only has the ability to edit real videos, but also improves the video editing accuracy by virtue of the temporal motion-related information in the video. This solution does not require adjusting model parameters, saves computing resources and model inference time, and thus also improves the editing efficiency of the diffusion model for videos.

[0044] Among them, the adjacent-frame optical flow motion information carried in the forward optical flow video and the reverse optical flow video can be determined with reference to the following process. In one implementation, before using the preset diffusion model to edit the initial video into a new video according to the editing information and the adjacent-frame optical flow motion information carried in the forward optical flow video and the reverse optical flow video, it further includes: generating forward optical flow features according to the first-frame edited frame and the forward optical flow video; the first-frame edited frame is obtained by editing the first frame of the initial video according to the editing information; generating reverse optical flow features according to the last-frame edited frame and the reverse optical flow video; the last-frame edited frame is obtained by editing the last frame of the initial video according to the editing information; splicing the forward optical flow features and the reverse optical flow features to obtain the adjacent-frame optical flow motion information. It can be seen that the adjacent-frame optical flow motion information includes optical flow motion information in both the forward and reverse directions; inputting the editing information, the encoded features of the initial video, the adjacent-frame optical flow motion information, and random noise into the preset diffusion model. The random noise is the pure Gaussian noise that needs to be added learned after the preset diffusion model is trained.

[0045] In order to make the forward optical flow features and the reverse optical flow features consistent in time sequence, in one implementation, splicing the forward optical flow features and the reverse optical flow features includes: arranging the reverse optical flow features in reverse order so that the forward optical flow features and the reverse optical flow features are consistent in time sequence, and then splicing the forward optical flow features and the reverse optical flow features arranged in reverse order.

[0046] It should be noted that, in order to ensure the effect of the newly edited video, the first frame and the last frame of the initial video can be edited separately first, and then the first-frame edited frame and the last-frame edited frame are shown to the user. If the user is satisfied with the effects of the first-frame edited frame and the last-frame edited frame, then the subsequent steps of this embodiment are executed to edit the entire initial video, thereby avoiding unnecessary resource waste. Any image editing model can be used here to edit the first frame and the last frame of the initial video separately. In one implementation, an editing model (which can be any image editing model) can be used to edit the first frame and the last frame of the initial video separately according to the editing information to obtain the first-frame edited frame and the last-frame edited frame; display the first-frame edited frame and the last-frame edited frame; after both the first-frame edited frame and the last-frame edited frame are confirmed, perform forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video; perform backward optical flow processing on the initial video using the optical flow network to obtain a backward optical flow video; use a preset diffusion model to edit the initial video into a new video according to the editing information and the adjacent-frame optical flow motion information carried in the forward optical flow video and the backward optical flow video. After receiving the denial instruction for the first-frame edited frame and / or the denial instruction for the last-frame edited frame, change the editing information and / or change the editing model, and re-edit the first frame and the last frame of the initial video, and display the re-edited first-frame edited frame and the last-frame edited frame, so that after the user is satisfied with the effects of the first-frame edited frame and the last-frame edited frame, the entire initial video is edited again.

[0047] The foregoing editing model can be an image editing model based on a generative adversarial network, or an image editing model based on a variational autoencoder, or an image editing model based on a diffusion model.

[0048] In one example, the training process of the preset diffusion model includes: constructing multiple triples including the original video, the editing instruction (i.e., the editing information), and the edited video; using the multiple triples to train the video editing ability of the initial diffusion model to obtain the preset diffusion model. Among them, one triple is a training data. In one example, one triple can include: O, P, Q, where O represents the original video, P represents the editing instruction for the original video, and Q represents the edited video obtained by modifying the original video O according to the editing instruction P.

[0049] In one implementation, multiple triples including the original video, editing instructions, and the edited video are constructed, which includes: using a video description model to respectively output corresponding text descriptions for multiple original videos to obtain multiple text descriptions; after respectively optimizing the multiple text descriptions, generating corresponding editing instructions according to each optimized text description to obtain multiple editing instructions; using a video reconstruction model to respectively determine the static features and temporal dynamic features in each original video; constructing multiple data groups including the original video, the corresponding editing instructions, the corresponding static features, and the corresponding temporal dynamic features; for each data group, editing the static features in the current data group according to the editing instructions in the current data group to obtain statically edited features (any image editing model can be used here to edit the static features in the current data group according to the editing instructions); mapping the temporal dynamic features in the current data group to the statically edited features to obtain the edited video corresponding to the original video in the current data group; constructing multiple triples including the original video, the corresponding editing instructions, and the corresponding edited video. If there are 10 original videos, then 10 text descriptions are correspondingly output using the video description model, and then these 10 text descriptions can be optimized using a natural language model. Using the natural language model, 10 corresponding editing instructions are respectively generated according to the optimized 10 text descriptions. Then, using the video reconstruction model for the 10 original videos, the static features and temporal dynamic features therein are determined; furthermore, the static features of the current original video are edited according to the editing instructions of each original video to obtain the statically edited features of the current original video, and the temporal dynamic features of the current original video are mapped to the statically edited features of the current original video, then the edited video corresponding to the current original video can be obtained; then 10 triples including the original video, the corresponding editing instructions, and the corresponding edited video can be constructed. In one implementation, respectively optimizing the multiple text descriptions includes: inputting the current text description and the corresponding combined description for generating the current text description into a natural language model so that the natural language model outputs the optimization result of the current text description.

[0050] It should be noted that the video description model may include: a video encoder and a natural language model. The video encoder can implement functions such as image encoding, vector alignment, temporal position encoding, video adaptation, and feature mapping to extract video features and align the dimensions of the video features with the input data of the natural language model. Here, the natural language model outputs, clusters, and merges segment descriptions for the input data. According to the merging of the segment descriptions, the corresponding video segments are further merged, and target frames are selected to generate text descriptions for the entire video frames. In one implementation, the process of using the video description model to output corresponding text descriptions for any original video includes: respectively outputting corresponding segment descriptions for k video segments in the current original video to obtain k segment descriptions, where k is a preset natural number; clustering the k segment descriptions to obtain a clustering result, for example, obtaining 8 categories; merging the k video segments and the k segment descriptions according to the clustering result to obtain S merged video segments and corresponding merged descriptions; generating corresponding text descriptions for the current original video based on the target frames in each merged video segment. If 8 categories are obtained, then S is equal to 8, and then target frames are respectively selected from the 8 merged video segments to generate text descriptions for the entire video frames. Among them, the target frame can be the video frame at the middle position (i.e., the median) of each merged video segment, or the video frame at the one-third position or other positions. In short, the target frame is one or more video frames determined from each merged video segment that can best represent the merged video segment.

[0051] In one implementation, generating corresponding text descriptions for the current original video based on the target frames in each merged video segment includes: respectively selecting at least one target frame from each merged video segment; encoding the image features and temporal order features of the selected target frames into a fused feature; obtaining corresponding text descriptions for the current original video based on the fused feature. For example, if one target frame is respectively selected from the 8 merged video segments, then 8 target frames are obtained, and based on the temporal order of these 8 target frames and the image features of these 8 target frames, corresponding text descriptions for the current original video can be obtained.

[0052] It should be noted that the video reconstruction model includes two parts: a deformation network and a static network. The static network is mainly used to extract static features, and the deformation network is mainly used to extract temporal dynamic features. In one implementation, the process of using the video reconstruction model to process each original video includes: extracting corresponding static features from each original video; splicing the static features of any original video with the time vector of the current original video, and performing deformation mapping on the splicing result according to the optical flow features of the current original video to obtain the temporal dynamic features of the current original video.

[0053] It can be seen that in this embodiment, the diffusion model is trained to perceive the temporal motion-related information in the real initial video, which not only endows the diffusion model with the ability to edit real videos, but also improves the video editing accuracy by leveraging the temporal motion-related information in the video. This solution does not require adjusting model parameters, saving computational resources and model inference time, thus also improving the editing efficiency of the diffusion model for videos.

[0054] Referring to the foregoing, the training dataset of the diffusion model is a triple {(x s , x0, c)}, where x s represents the original video, x0 represents the edited video obtained by editing the original video, and c represents the content of the editing information.

[0055] To construct this triple, real videos can be crawled from the network to form a real video data set S. Then, a video description model is used to predict the video description text for each video in the set, and the following output can be obtained . Among them, s i represents the i-th sample in the set S, represents the text description of the k-th video segment in the i-th sample predicted by the video description model, represents the starting position of the k-th video segment in s i . Please refer to Figure 3 . The video description model includes: a video encoder and a natural language model (i.e., a large language model). The video encoder can encode visual features of the input video and then input them into the large language model, enabling the large language model to output text descriptions corresponding to S (the number of preset clustering categories) segments respectively and a text description for the entire video. The S segments and their corresponding text descriptions can refer to Figure 3 the output result, in the format: time period + text.

[0056] Among them, the method for performing semantic similarity clustering on the text descriptions corresponding to each segment in is not limited. For example, a text encoder can be used to obtain the text features of the text descriptions corresponding to each segment, and then kmeans clustering can be performed on these text features to merge the repeated text contents, obtaining S classes where the text descriptions are located. The text descriptions in one class can be summarized into one, so it can be considered that S merged text descriptions are obtained . Then, according to the clustering result, the k video segments of each video are merged into S segments. With the help of the large language model, the middle frames in the S segments are input to the large language model to obtain the video text description of the entire video Figure 4 . The text description of the entire video can refer to

[0057] For example,Figure 4 As shown. After the image encoder in the video encoder encodes the image features, the initial parameter for query vector randomization is q, and q is used to determine the temporal dimension of S intermediate frames, obtaining the solidified vector h. For example, if the temporal dimension is 10, then the temporal dimension of each intermediate frame is 10. That is to say, the temporal length of the video frames taken from the S segments cannot exceed q. Of course, it is not necessarily only the intermediate frames that are taken, and any frames between the 1 / 3 position and the 1 / 2 position can also be taken. If the temporal dimension is 10, the frame vector B×T×D is mapped to B×10×D through the queried q. B represents the batch size, D represents the feature extraction vector of each frame, and T represents the temporal length of the frame vector. When querying the network, a transformer structure is used to perform cross-attention operations between the frame vector as the conditional vector and the initialized query vector to obtain the solidified vector h with a determined dimension. Among them, the video adapter, query network, and feature mapping network can adopt the training method of contrastive learning. After the entire video passes through the image encoder, query network, video adapter, and feature mapping, a unified dimension is obtained, such as B×10×d. B represents the batchsize, 10 represents the temporal dimension, and d represents the dimension of the input character encoding of the large language model, such as 4096. Then, contrastive loss learning is performed with the original text of the video to update the model parameters of the video adapter, query network, and feature mapping network.

[0058] Among them, the picture features of the intermediate frames in the video segment corresponding to the clustering sub-description and the query vector q pass through the query network to obtain a feature vector f with a fixed dimension. Temporal position encoding is added to f to obtain the fused feature fuse to strengthen the temporal order of the features; fuse passes through the video adapter and feature mapping to map the features to the dimension acceptable by the large language model, and finally the specific description of the entire video is obtained through the large language model.

[0059] The following is the optimization of the video text description: Combine and and input them into any natural language model to output the description t of the complete content of the entire video i . Among them, the prompt words input to the natural language model can be: "Now you are a text summarization expert. Given a series of sub-segment texts and a complete segment text, please summarize the sub-segment and complete segment texts and output the complete segment description. Sub-segments: {sub-segment 1}{sub-segment 2}{sub-segment 3}{sub-segment 4}{sub-segment 5}{sub-segment 6}{sub-segment 7}{sub-segment 8}. Complete segment: {complete segment text description}". In this step, the language logic of "a series of sub-segment texts and a complete segment text" is optimized using the natural language model to obtain a new text description t for the entire video i .

[0060] After that, input t i into any natural language model again to obtain an editing instruction . The prompt input to the natural language model for this step can be: "Now here is a piece of text, and you need to randomly replace the entities, styles, and behaviors in the text. The number of replacements should not exceed three times. Output the replaced text. Text description: {Given text description}". Based on the text description, use the natural language model to output the replaced text, which is an editing instruction for a segment of the entire video .

[0061] Next, for the given original video and the editing instruction , use the video reconstruction model and the editing model to determine the corresponding edited video . Among them, the video reconstruction model is used to determine the static features and temporal dynamic features in the original video; the editing model is used to edit the static features of the current original video according to the editing instruction of the original video to obtain the static editing features of the current original video, and map the temporal dynamic features of the current original video to the static editing features of the current original video, then the edited video corresponding to the current original video can be obtained

[0062] Please refer to Figure 5 , the training process of the video reconstruction model includes: inputting the original video . Here, the original video is regarded as an X, Y, T three-dimensional structure, where X represents the width of the video, Y represents the height of the video, and here the pixel is used as the basic unit; T represents the number of video frames in the video. There are a total of X×Y×T data sampling spatio-temporal points in such a video. The deformation network and the static network in the video reconstruction model are both simple multi-layer perceptron networks. The sampling spatio-temporal points are input into the deformation network to obtain the deformation vector ; if there are 3 scales of perceptrons in the deformation network, that is , then the feature encoding formula for each perceptron is: . If the output dimension of the perceptron is 256, then the obtained deformation vector is obtained by splicing the output features of the three scales. Finally, the deformation vector for each sampling spatio-temporal point has a dimension of 768. The static network outputs the spatial grid vector (i.e., the static feature) and the time vector for the sampling point . Splice these two, and then pass through the static network to obtain the combined dynamic and static feature of the entire video . This feature is used to predict the RGB three-channel color values of the corresponding sampling points. After that, according to the combined dynamic and static feature output by the static network and the color values of the original sampling spatio-temporal points Calculate the color reconstruction loss, and the formula for the color reconstruction loss is: , where L2 represents the L2 norm.

[0063] To enhance the accuracy of dynamic feature extraction, calculate the optical flow between two adjacent frames, and map the previously sampled spatio-temporal points according to the optical flow to , to determine the deformation of the points from the t-th frame to the (t + 1)-th frame. After passing the mapped spatio-temporal points through the deformation network, the mapped deformation vector (i.e., the temporal dynamic feature) is obtained. Then calculate the and similarity loss, and the formula for the similarity loss is , represents the L2 norm of the two deformation vectors, is the confidence probability of the displacement of each pixel point when calculating the optical flow between two adjacent frames, that is, the confidence of . Any optical flow algorithm can be used here. In this scheme, the video dynamic features and static features can be separated through the color reconstruction loss and the similarity loss of the deformation vectors, and the self-supervised reconstruction task of the video can be completed. The deformation vector represents the pixel deformation from the static content to each frame of the image.

[0064] According to the above, for multiple original videos, multiple triples {(x s , x0, c)} can be constructed, where x s represents the original video, x0 represents the edited video obtained by editing the original video, and c represents the content of the editing information. Next, use multiple triples {(x s , x0, c)} to train the diffusion model.

[0065] For the input triple {(x s , x0, c)}, the initial diffusion model first generates a series of Markov chain hidden vectors . This series of hidden vectors gradually add Gaussian noise to the video after editing. The formula for adding Gaussian noise is as follows: , where N represents the Gaussian distribution, represents the noise scheduling parameter, which is a value between 0 and 1 and represents the variance of the noise added at time step t. β t usually increases with the increase of t, indicating a gradual increase in the intensity of the noise; q represents the conditional probability distribution, representing the transition probability from to . The model learns the noise that needs to be added from this, and finally obtains the trained diffusion model.

[0066] Based on the hidden vector , the denoising inference process of the trained diffusion model is specifically formulated as follows: , is the mean of the model prediction, and is the variance of the model prediction (which can usually be simplified to a fixed value). The model gradually removes the noise by learning to predict the noise at each step and finally restores .

[0067] Specifically, during the video content editing process, the initial is randomly sampled from Gaussian noise , combined with the initial input of the original video , and the noise is gradually reduced to obtain the modified video . Usually, during the specific practice process, the deviation of the predicted noise is used as the specific optimization training function, and the specific function is as follows: , where E represents taking the expectation of the subsequent input function under the condition of satisfying the subscript; represents the original video sampled from the data distribution ; ; represents the noise ϵ sampled from the standard Gaussian distribution N(0, I); represents the time step t randomly sampled from the time step range [1, T]; where is the original input video, is the deviation (i.e., loss) obtained by neural network inference on the edited information c, the original input video , the input time step t, and the noisy image at time step t, are the parameters learned by the neural network.

[0068] Please refer to Figure 6 , the training process of the trained diffusion model includes:

[0069] 1. Parse the initial frame and the end frame from the video content.

[0070] 2. The editing instruction c and pass through the image editing model to obtain the edited picture . The editing instruction c is: A giraffe is walking on the moon wearing a spacesuit.

[0071] 3. The editing instruction c and pass through the image editing model to obtain the edited picture . Note that there is no restriction on which image editing model to use in steps 2 and 3, but in order to maintain the consistency of the generated data, generally the same image editing model is used.

[0072] 4. Calculate the forward optical flow for the original video to obtain the forward optical flow video .

[0073] 5. Reverse the video frames of the original video and calculate the backward optical flow to obtain the backward optical flow video .

[0074] 6. and respectively pass through the encoder to obtain the encoded vectors and , denotes the encoder network. and use the same encoder, which plays the role of feature compression. For example, the original resolution H×W is changed to H / 8×W / 8.

[0075] 7. Concatenate the encoded vectors and to obtain the forward branch vector .

[0076] 8. Perform steps 6 and 7 on and to obtain the backward branch vector .

[0077] 9. The original video passes through the encoder to obtain the video vector .

[0078] 10. Perform conditional vector concatenation according to the following formula to obtain the conditional concatenated vector : ; denotes Figure 6 the time-reversal operation in , that is, reverse the order of the backward branch vector

[0079] 11. Concatenate with random noise, and at the same time input the editing instruction c as the text condition, and input it into the diffusion model for denoising. Finally, obtain the denoised vector, and the denoised vector passes through the decoder to obtain the finally edited video.

[0080] It can be seen that the diffusion model obtained in this way is a general video content generation and editing model. The model can directly perform video editing operations on the video data and editing instructions input by the user without additional training on the data input by the user. In order to keep the model consistent in time series while maintaining spatial semantic changes, this solution designs a three-branch input. First, the initial frames of the original video are initially edited for spatial content. The results of this step can also be added to the user interaction experience during actual application. This step will generate an initial edit of the spatial content. If the user is not satisfied with the edit of the spatial content, they can modify the input conditions or replace the image editing model. This can greatly improve the user experience and avoid unnecessary resource consumption caused by excessive errors in the initial spatial content editing. After the user is satisfied with the spatial content edit of the first and last frames, the enhancement of time series content consistency continues. The forward optical flow video and the backward optical flow video of the original video are calculated. The forward optical flow video and the edited frame of the initial frame are compressed into latent vectors by the encoder, and then vector splicing is completed to obtain the forward branch vector. At the same time, the backward optical flow video and the edited frame of the end frame are compressed into latent vectors by the encoder, and then vector splicing is completed to obtain the backward branch vector. The backward branch vector (i.e., the backward optical flow feature) is added to the forward branch vector (i.e., the forward optical flow feature) after being reversed in time series to obtain the motion splicing vector (i.e., the optical flow motion information of adjacent frames). These two branches complete the enhancement of time series content consistency from the two dimensions of the forward and backward directions of the video.

[0081] To maintain the spatio-temporal consistency of the overall video, the third branch inputs the original video, which is also encoded into the latent vector space by the encoder. After vector splicing of the motion splicing vector, the original video vector, and the random noise vector, and with the editing instruction as the text condition, it is input into the diffusion model to complete the denoising of the corresponding noise, and finally, the final edited video is obtained through decoding by the decoder.

[0082] It can be seen that on the one hand, this solution introduces video construction training data based on real-world videos, and on the other hand, it designs a general video content generation and editing model with branches. Based on the high-quality construction of data samples in the previous step, the training of the general model is completed. After the training is completed, the user does not need to perform further model fine-tuning training, which greatly improves the model inference efficiency.

[0083] Next, a video processing device provided by an embodiment of the present application will be introduced. The video processing device described below can be mutually referred to with other embodiments described in this article.

[0084] An embodiment of the present application discloses a video processing device, including:

[0085] An acquisition module, configured to acquire an initial video and editing information;

[0086] A forward optical flow processing module, configured to perform forward optical flow processing on an initial video by using a preset optical flow network to obtain a forward optical flow video;

[0087] A backward optical flow processing module, configured to perform backward optical flow processing on the initial video by using the optical flow network to obtain a backward optical flow video;

[0088] An editing module, configured to edit the initial video into a new video by using a preset diffusion model according to the editing information and the adjacent frame optical flow motion information carried in the forward optical flow video and the backward optical flow video.

[0089] In one implementation, the forward optical flow processing module is specifically configured to:

[0090] Calculate the optical flow values between adjacent frames in the direction from the first frame to the last frame in the initial video by using the optical flow network, convert the optical flow values into corresponding optical flow maps, and summarize the optical flow maps to obtain the forward optical flow video;

[0091] Correspondingly, the backward optical flow processing module is specifically configured to:

[0092] Calculate the optical flow values between adjacent frames in the direction from the last frame to the first frame in the initial video by using the optical flow network, convert the optical flow values into corresponding optical flow maps, and summarize the optical flow maps to obtain the backward optical flow video.

[0093] In one implementation, it further includes:

[0094] A feature processing module, configured to generate forward optical flow features according to the first-frame edited frame and the forward optical flow video; the first-frame edited frame is obtained by editing the first frame of the initial video according to the editing information; generate backward optical flow features according to the last-frame edited frame and the backward optical flow video; the last-frame edited frame is obtained by editing the last frame of the initial video according to the editing information; splice the forward optical flow features and the backward optical flow features to obtain adjacent frame optical flow motion information; input the editing information, the encoded features of the initial video, the adjacent frame optical flow motion information, and random noise into the preset diffusion model.

[0095] In one implementation, the feature processing module is specifically configured to:

[0096] Reverse the order of the backward optical flow features, and splice the forward optical flow features and the backward optical flow features after reverse order arrangement.

[0097] In one implementation, it further includes:

[0098] A trial editing module is used to edit the first frame and the last frame of the initial video respectively according to the editing information by using an editing model, so as to obtain a first-frame edited frame and a last-frame edited frame; display the first-frame edited frame and the last-frame edited frame; after both the first-frame edited frame and the last-frame edited frame are confirmed, perform forward optical flow processing on the initial video by using a preset optical flow network to obtain a forward optical flow video; perform backward optical flow processing on the initial video by using the optical flow network to obtain a backward optical flow video; and edit the initial video into a new video by using a preset diffusion model according to the editing information and the adjacent-frame optical flow motion information carried in the forward optical flow video and the backward optical flow video.

[0099] In one implementation, the trial editing module is further used for:

[0100] After receiving a denial instruction for the first-frame edited frame and / or a denial instruction for the last-frame edited frame, change the editing information and / or change the editing model, and re-edit the first frame and the last frame of the initial video.

[0101] In one implementation, the training process of the preset diffusion model includes:

[0102] Construct a plurality of triples including the original video, the editing instruction, and the edited video;

[0103] Use the plurality of triples to train the video editing ability of the initial diffusion model to obtain the preset diffusion model.

[0104] In one implementation, constructing a plurality of triples including the original video, the editing instruction, and the edited video includes:

[0105] Use a video description model to respectively output corresponding text descriptions for a plurality of original videos to obtain a plurality of text descriptions;

[0106] After respectively optimizing the plurality of text descriptions, generate corresponding editing instructions according to the optimized text descriptions to obtain a plurality of editing instructions;

[0107] Use a video reconstruction model to respectively determine the static features and the time-domain dynamic features in each original video;

[0108] Construct a plurality of data groups including the original video, the corresponding editing instruction, the corresponding static feature, and the corresponding time-domain dynamic feature;

[0109] For each data group, edit the static feature in the current data group according to the editing instruction in the current data group to obtain a static edited feature; map the time-domain dynamic feature in the current data group to the static edited feature to obtain the edited video corresponding to the original video in the current data group;

[0110] Construct a plurality of triples including the original video, the corresponding editing instruction, and the corresponding edited video.

[0111] In one embodiment, the process of using a video description model to output a corresponding text description for any original video includes:

[0112] Outputting corresponding segment descriptions for k video segments in the current original video respectively to obtain k segment descriptions;

[0113] Clustering the k segment descriptions to obtain a clustering result;

[0114] Merging the k video segments and the k segment descriptions according to the clustering result to obtain S merged video segments and corresponding merged descriptions;

[0115] Generating a corresponding text description for the current original video according to the target frames in each merged video segment.

[0116] In one embodiment, generating a corresponding text description for the current original video according to the target frames in each merged video segment includes:

[0117] Selecting at least one target frame in each merged video segment respectively;

[0118] Encoding the image features and temporal order features of the selected target frames into a fused feature;

[0119] Obtaining a corresponding text description for the current original video based on the fused feature.

[0120] In one embodiment, optimizing multiple text descriptions respectively includes:

[0121] Inputting the current text description and the corresponding merged description for generating the current text description into a natural language model so that the natural language model outputs an optimized result of the current text description.

[0122] In one embodiment, the process of using a video reconstruction model to process each original video includes:

[0123] Extracting corresponding static features from each original video;

[0124] Concatenating the static features of any original video with the time vector of the current original video, and performing deformation mapping on the concatenated result according to the optical flow features of the current original video to obtain the temporal dynamic features of the current original video.

[0125] Wherein, for the more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0126] It can be seen that this embodiment provides a video processing device, which trains the ability of the diffusion model to perceive the temporal motion-related information in the real initial video. This not only enables the diffusion model to have the ability to edit real videos, but also improves the video editing accuracy by leveraging the temporal motion-related information in the video. This solution does not require adjusting model parameters, saving computing resources and model inference time, and thus also improves the editing efficiency of the diffusion model for videos.

[0127] Next, an electronic device provided by an embodiment of the present application will be introduced. The electronic device described below can be cross-referred to other embodiments described herein.

[0128] An embodiment of the present application discloses an electronic device, including:

[0129] A memory for storing a computer program;

[0130] A processor for executing the computer program to implement the method disclosed in any of the above embodiments.

[0131] Furthermore, an embodiment of the present application also provides an electronic device. Among them, the above electronic device can be either a Figure 7 server as shown in Figure 8 or a terminal as shown in Figure 7 and Figure 8 are both structural diagrams of electronic devices shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation on the scope of use of the present application.

[0132] Figure 7 This is a schematic structural diagram of a server provided by an embodiment of the present application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. Among them, the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the video processing disclosed in any of the foregoing embodiments.

[0133] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0134] In addition, as a carrier for resource storage, the memory can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon include an operating system, a computer program, and data, etc., and the storage method can be temporary storage or permanent storage.

[0135] Among them, the operating system is used to manage and control each hardware device and computer program on the server to enable the processor to perform operations and processing on the data in the memory. It can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the video processing method disclosed in any of the foregoing embodiments, the computer program can further include computer programs that can be used to complete other specific tasks. In addition to data such as update information of the application program, the data can also include data such as developer information of the application program.

[0136] Figure 8 It is a schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal may specifically include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.

[0137] Generally, the terminal in this embodiment includes: a processor and a memory.

[0138] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor can also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0139] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory is at least used to store the following computer program. After being loaded and executed by the processor, the computer program can implement the relevant steps in the video processing method executed by the terminal side disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of the application program.

[0140] In some embodiments, the terminal may further include a display screen, an input / output interface, a communication interface, sensors, a power supply, and a communication bus.

[0141] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the terminal, and it may include more or fewer components than shown in the figure.

[0142] Next, a computer-readable storage medium provided by an embodiment of the present application will be introduced. The computer-readable storage medium described below can be referred to each other with other embodiments described in this article.

[0143] A computer-readable storage medium is used to store a computer program. When the computer program is executed by a processor, it implements the video processing method disclosed in the foregoing embodiments.

[0144] Among them, the non-volatile storage medium is a computer-readable non-volatile storage medium. As a carrier for storing resources, it can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored on it include an operating system, a computer program, and data, etc., and the storage method can be transient storage or permanent storage. In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (abbreviated as ROM), random access memories (abbreviated as RAM), mobile hard disks, magnetic disks, or optical discs, etc., which are various media that can store computer programs.

[0145] Next, a computer program product provided by an embodiment of the present application will be introduced. The computer program product described below can be referred to each other with other embodiments described in this article.

[0146] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the aforementioned disclosed video processing method are implemented.

[0147] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned video processing methods is implemented.

[0148] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0149] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of non-volatile storage medium well-known in the technical field.

[0150] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A video processing method, characterized in that: include: Get initial video and editing information; Performing forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video; Using the optical flow network to perform reverse optical flow processing on the initial video to obtain a reverse optical flow video; Generate a forward optical flow feature according to the first edited frame and the forward optical flow video; The first frame editing frame is obtained by editing the first frame of the initial video according to the editing information; Generate a reverse optical flow feature according to the last frame editing frame and the reverse optical flow video; The tail frame editing frame is obtained by editing the tail frame of the initial video according to the editing information; The forward optical flow feature and the reverse optical flow feature are concatenated to obtain optical flow motion information of adjacent frames; The initial video is edited into a new video by using a preset diffusion model according to the editing information and the adjacent frame optical flow motion information.

2. The method according to claim 1, characterized in that Performing forward optical flow processing on the initial video using a preset optical flow network to obtain a forward optical flow video, including: Calculating the optical flow values ​​between adjacent frames in the direction from the first frame to the last frame in the initial video using the optical flow network, converting the optical flow values ​​into corresponding optical flow graphs, and summarizing the optical flow graphs to obtain the forward optical flow video; Accordingly, the optical flow network is used to perform reverse optical flow processing on the initial video to obtain a reverse optical flow video, including: The optical flow network is used to calculate the optical flow values ​​between adjacent frames in the direction from the last frame to the first frame in the initial video, and the optical flow values ​​are converted into corresponding optical flow maps, and the optical flow maps are summarized to obtain the reverse optical flow video.

3. The method according to claim 2, characterized in that The editing information, the encoding features of the initial video, the adjacent frame optical flow motion information and random noise are input into the preset diffusion model.

4. The method according to claim 3, characterized in that The forward optical flow feature and the reverse optical flow feature are concatenated, including: The reverse optical flow features are arranged in reverse order, and the forward optical flow features and the reverse optical flow features after the reverse order arrangement are spliced.

5. The method according to any one of claims 1 to 4, characterized in that: Also includes: Using the editing model, the first frame and the last frame of the initial video are edited according to the editing information to obtain a first frame editing frame and a last frame editing frame; Display the first frame editing frame and the last frame editing frame; After the first edited frame and the last edited frame are confirmed, forward optical flow processing is performed on the initial video using a preset optical flow network to obtain a forward optical flow video; Using the optical flow network to perform reverse optical flow processing on the initial video to obtain a reverse optical flow video; The step of editing the initial video into a new video by using a preset diffusion model according to the editing information and the adjacent frame optical flow motion information.

6. The method according to claim 5, characterized in that Also includes: After receiving the denial instruction of the first frame editing frame and / or the denial instruction of the last frame editing frame, the editing information and / or the editing model are changed, and the first frame and the last frame of the initial video are re-edited.

7. The method according to any one of claims 1 to 4, characterized in that: The training process of the preset diffusion model includes: Constructing multiple triplets including original videos, editing instructions, and edited videos; The video editing capability of the initial diffusion model is trained using the multiple triplets to obtain the preset diffusion model.

8. The method according to claim 7, characterized in that Construct multiple triplets including original video, editing instructions and edited video, including: Using the video description model, outputting corresponding text descriptions for the multiple original videos respectively to obtain multiple text descriptions; After respectively optimizing the multiple text descriptions, corresponding editing instructions are generated according to the optimized text descriptions to obtain multiple editing instructions; The video reconstruction model is used to respectively determine the static features and the temporal dynamic features in each original video; Constructing a plurality of data groups including original videos, corresponding editing instructions, corresponding static features, and corresponding temporal dynamic features; For each data group, the static features in the current data group are edited according to the editing instructions in the current data group to obtain the static editing features; the time domain dynamic features in the current data group are mapped to the static editing features to obtain the edited video corresponding to the original video in the current data group; Construct multiple triplets including original videos, corresponding editing instructions and corresponding edited videos.

9. The method according to claim 8, characterized in that The process of outputting a corresponding text description for any original video using the video description model includes: Output corresponding segment descriptions for k video segments in the current original video respectively, and obtain k segment descriptions; Clustering the k fragment descriptions to obtain a clustering result; Merging k video segments and k segment descriptions according to the clustering result to obtain S merged video segments and corresponding merged descriptions; Generate a text description corresponding to the current original video according to the target frames in each merged video segment.

10. The method according to claim 9, characterized in that Generate a text description of the current original video based on the target frames in each merged video segment, including: Selecting at least one target frame in each merged video segment; Encode the image features and time sequence features of the selected target frame into fusion features; A text description corresponding to the current original video is obtained based on the fusion features.

11. The method according to claim 9, characterized in that Optimizing the multiple text descriptions respectively includes: The current text description and the corresponding merged description generated by generating the current text description are input into a natural language model, so that the natural language model outputs an optimization result of the current text description.

12. The method according to claim 8, characterized in that The process of processing each original video using the video reconstruction model includes: Extract corresponding static features from each original video; The static features of any original video are spliced ​​with the time vector of the current original video, and the splicing result is deformed and mapped according to the optical flow features of the current original video to obtain the time domain dynamic features of the current original video.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Video data processing method and device, electronic equipment and readable storage medium

    CN118283297A