Video frame generation method and device and electronic equipment
By determining key feature points and their position change information in the video frame, and generating intermediate video frames using the target network model, the problems of poor quality of video frames and difficult to meet personalized needs in the existing technology generated in complex scenarios are solved, and high-quality, smooth and time-continuous intermediate video frame generation is achieved.
Patent Information
- Application Number
- CN202510241966.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video frame generation methods have shortcomings in processing complex scenarios and meeting user personalized needs, and it is difficult to generate high-quality, smooth and time-continuous intermediate video frames.
By determining the key feature points and their position change information in the video frame, the target network model is used to generate intermediate video frames, and the positions of key feature points are guided to conform to the predetermined change information, and the generation process of intermediate video frames is accurately controlled.
It realizes personalized creative intentions, generates high-quality, smooth and time-continuous intermediate video frames, improving the quality and coherence of video frame generation.
Smart Images

Figure CN120301987A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method, apparatus, and electronic device for generating video frames. Background Art
[0002] The generation of image sequences, especially intermediate frame interpolation technology, plays an important role in many fields. In related technologies, intermediate image sequences can be generated through interpolation algorithms. However, this method relies on the learning of motion information. If the image content changes drastically, the object motion trajectory is complex, or there is occlusion, etc., it will lead to poor effects of the finally obtained intermediate image sequences. In addition, the above method usually can only generate deterministic results and is difficult to meet the personalized needs of users for intermediate image sequences. In short, the existing interpolation methods still have deficiencies in dealing with complex scenarios and meeting the personalized needs of users. Summary of the Invention
[0003] In view of this, the purpose of the present disclosure is to provide a method, apparatus, and electronic device for generating video frames, which determine key feature points and the position change information of the key feature points in a video frame. Through the first video frame, the second video frame, and the position change information of the key feature points, it is possible to guide the position information of the key feature points in the intermediate video frames generated by the target network model to conform to the pre-determined position change information, accurately control the generation process of the intermediate video frames, realize personalized creative intentions, and generate high-quality, smooth, and temporally continuous intermediate video frames.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating video frames. The method includes: determining a first video frame and a second video frame from a video to be processed; determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point; wherein the position change information is used to indicate the motion trajectory of the key feature point between the first video frame and the second video frame; inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generating at least one intermediate video frame through the target network model; wherein the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame.
[0005] In a second aspect, embodiments of the present disclosure provide a video frame generation device, which includes: a video frame determination module for determining a first video frame and a second video frame from a video to be processed; a position change information determination module for determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point; wherein the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; an intermediate video frame generation module for inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generating at least one intermediate video frame through the target network model; wherein the position change information of the key feature point is the guiding information of the target network model, the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame.
[0006] In a third aspect, embodiments of the present disclosure provide an electronic device, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the video frame generation method according to any one of the first aspect.
[0007] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the video frame generation method according to any one of the first aspect.
[0008] Embodiments of the present disclosure bring the following beneficial effects:
[0009] The present disclosure provides a method, an apparatus, and an electronic device for generating video frames. The method includes: determining a first video frame and a second video frame from a video to be processed; determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point, where the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model to generate at least one intermediate video frame through the target network model, where the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame. By determining the key feature point and the position change information of the key feature point in the video frame, and guiding the target network model through the first video frame, the second video frame, and the position change information of the key feature point, the position information of the key feature point in the generated intermediate video frame conforms to the pre-determined position change information, which can accurately control the generation process of the intermediate video frame, realize the personalized creative intention, and generate high-quality, smooth, and temporally continuous intermediate video frames.
[0010] Other features and advantages of the present disclosure will be described in the following description, and some of them will be obvious from the description or learned by implementing the present disclosure. The objectives and other advantages of the present disclosure are achieved and obtained by the structures specifically pointed out in the description, the claims, and the drawings.
[0011] To make the above objectives, features, and advantages of the present disclosure more obvious and understandable, the following specific embodiments are given, and detailed descriptions are provided in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a flowchart of a method for generating a video frame provided by an embodiment of the present disclosure;
[0014] Figure 2 It is an example diagram of a guiding and regulating diagram in a method for generating a video frame provided by an embodiment of the present disclosure;
[0015] Figure 3 It is a schematic structural diagram of an apparatus for generating a video frame provided by an embodiment of the present disclosure;
[0016] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0018] Image sequence generation, especially intermediate frame interpolation technology, plays an important role in many fields. In some related technologies, an intermediate image sequence can be generated through an interpolation algorithm (such as an interpolation algorithm based on optical flow estimation or a difference algorithm based on motion compensation), usually relying on an accurate estimation of the pixel motion in the image sequence. However, when the image content changes drastically, the object motion trajectory is complex, or there is occlusion, etc., the accuracy of the image sequence will drop significantly, resulting in obvious visual artifacts or blurring in the interpolation result.
[0019] In some other related technologies, clearer intermediate frames can be generated through deep learning-based interpolation methods, but these methods still rely on the learning of motion information and usually can only generate deterministic results, which are difficult to meet the user's personalized control requirements for the transition process. Based on this, a video frame generation method, device, and electronic device provided by an embodiment of the present disclosure can be applied to devices such as mobile phones, computers, laptops, computers, servers, etc., and especially can be applied to devices with video generation functions.
[0020] To facilitate the understanding of this embodiment, a video frame generation method disclosed in an embodiment of the present disclosure will be introduced in detail first, as Figure 1 shown, the method includes the following steps:
[0021] Step S102, determine a first video frame and a second video frame from the video to be processed;
[0022] The aforementioned video to be processed generally refers to a video including multiple video frames, such as a 5-second video or a 10-second video. Two video frames, namely the aforementioned first video frame and second video frame, can be determined from the video to be processed as needed. Generally, the frame number of the first video frame is less than that of the second video frame, and the first video frame and the second video frame can be consecutive video frames or non-consecutive video frames. It is also possible to refer to the first video frame as the initial video frame and the second video frame as the target video frame or the end video frame.
[0023] Preferably, the first video frame is the starting frame of the video to be processed, and the second video frame is the ending frame of the video to be processed.
[0024] Step S104: Determine at least one key feature point in the first video frame and the second video frame, and determine the position change information of the key feature point; wherein, the position change information is used to indicate the motion trajectory of the key feature point between the first video frame and the second video frame.
[0025] Several representative feature points can be manually marked on the first video frame and the second video frame, such as key nodes of the human body, corner points of objects, etc. It is also possible to automatically detect the corresponding feature points between the first video frame and the second video frame through an image feature matching algorithm to obtain at least one key feature point.
[0026] The above-mentioned image feature matching algorithm can be SIFT (Scale-Invariant Feature Transform), ORB (Oriented FAST and Rotated BRIEF), etc.
[0027] Exemplarily, through the SIFT method, first convolve the first video frame and the second video frame with Gaussian kernels of different scales to generate multiple sets of images. Subtract adjacent-scale Gaussian images to obtain a DoG (Difference of Gaussians) image. Find local extreme points in the DoG image as candidate key points. Precise position the key points by fitting a three-dimensional quadratic function to remove low-contrast and edge response points. Use the Hessian matrix to eliminate edge response points. Calculate the gradient magnitude and direction within the area around the key point. Statistically determine the main direction of the gradient. Assign the main direction to the key point to ensure rotational invariance. Calculate the gradient direction histogram within the area around the key point to generate a 128-dimensional feature vector. Normalize the feature vector. Match the key points in different images by calculating the Euclidean distance between feature vectors. Use the nearest neighbor algorithm or ratio test to filter the matching points.
[0028] Exemplarily, through the ORB method, use FAST to detect the key points of the first video frame and the second video frame, calculate the direction of the key points by the centroid method, generate a rotation-invariant BRIEF descriptor, and use the Hamming distance for feature matching.
[0029] Finally, the first key feature point in the first video frame and the first key feature point in the second video frame can be obtained through the above methods. Finally, the same feature points between the first key feature point and the second key feature point are determined as the corresponding key feature points between the first video frame and the second video frame.
[0030] The position change information of key feature points can be obtained by manually specifying the positions of each key feature at different time points. For example, position editing can be performed by dragging operations on the key feature points.
[0031] The initial position information of key feature points can also be generated by interpolation methods.
[0032] Step S106: Input the first video frame, the second video frame, and the position change information of key feature points into a pre-trained target network model, and generate at least one intermediate video frame through the target network model; wherein, the position change information of key feature points is the guiding information of the target network model, the positions of key feature points in the intermediate video frame conform to the position change information of key feature points, and the intermediate video frame is a transitional video frame between the first video frame and the second video frame.
[0033] The above target network model can be a deep learning model or a diffusion model. The core principle of the diffusion model is to gradually add noise to the image through a forward diffusion process until the image completely becomes random noise. Then, through a reverse denoising process, a clear image is gradually restored from the noise. In this embodiment, the diffusion model can control the reverse denoising process according to the initial frame (the first video frame), the target frame (the second video frame), and the position change information of key feature points, so as to generate an intermediate video frame that meets the conditions. The conditions here refer to that the position information of key feature points in the intermediate video frame conforms to the position change information of key feature points.
[0034] Optionally, the first video frame, the second video frame, and the position change information of key feature points can be encoded to obtain an encoding result that can be recognized by the target network model. After the encoding results are fused, a fusion result is obtained, and the fusion result is input into the target network model. During the process of generating the intermediate video frame, the target network model will consider the first video frame, the second video frame, and the position change information of key feature points, so that the position information of key feature points in the finally generated intermediate video frame conforms to the position change information of key feature points.
[0035] An embodiment of the present disclosure provides a method for generating video frames. The method includes: determining a first video frame and a second video frame from a video to be processed; determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point, where the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generating at least one intermediate video frame through the target network model; where the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transitional video frame between the first video frame and the second video frame. This method determines the key feature point and the position change information of the key feature point in the video frame. Through the first video frame, the second video frame, and the position change information of the key feature point, it guides the position information of the key feature point in the intermediate video frame generated by the target network model to conform to the pre-determined position change information, can accurately control the generation process of the intermediate video frame, realizes the personalized creative intention, and generates high-quality, smooth, and temporally continuous intermediate video frames.
[0036] For the step of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point, a possible implementation manner is as follows: in response to a marking operation for a target feature point in the first video frame and the second video frame, determining the target feature point as the key feature point in the first video frame and the second video frame; and in response to a position adjustment operation for the key feature point, determining the position change information of the key feature point.
[0037] Optionally, in response to a first position adjustment operation for the key feature point, determining the position information of the key feature point at a first time point; in response to a second position adjustment operation for the key feature point, determining the position information of the key feature point at a second time point; and in response to a third position adjustment operation for the key feature point, determining the position information of the key feature point at a third time point, and determining the position information of the key feature point at the first time point, the second time point, and the third time point as the position change information of the key feature point.
[0038] Optionally, in response to a position adjustment operation for the key feature point, determining the position coordinate information of the adjusted key feature point; and according to the position coordinate information of the adjusted key feature point, determining the position change information of the key feature point; where the position change information of the key feature point includes the position information of the key feature point at different time points or different video frames.
[0039] The above position adjustment operations can be performed multiple times. For example, in response to at least one position adjustment operation for key feature points, determine the position coordinate information of the key feature points after each adjustment; according to the position coordinate information of the key feature points after each adjustment, determine the position change information of the key feature points.
[0040] In the above manner, the user can select key feature points according to their own needs and set the position change information of the key feature points, thereby guiding the model to focus on the key feature points during the generation process, making the position information of the key feature points conform to the above position change information, and then generating intermediate video frames that meet the needs of the player.
[0041] For the above steps of determining at least one key feature point in the first video frame and the second video frame and determining the position change information of the key feature point, another possible implementation manner: according to the image feature matching algorithm, determine the key feature points in the first video frame and the second video frame; determine the position information of the key feature points in the first video frame as the initial position information of the key feature points, and determine the position information of the key feature points in the second video frame as the target position information of the key feature points; according to the initial position information of the key feature points and the target position information of the key feature points, determine the position change information of the key feature points through the interpolation algorithm.
[0042] For the above step of determining the position change information of the key feature points through the interpolation algorithm, a possible implementation manner: for each key feature point, determine the position information of the key feature point on the intermediate video frame through the following method: P t = P start +(P end - P start )*t / N; where P t is the position information of the key feature point in the t-th frame of the intermediate video frame, P start is the initial position information of the key feature point, P end is the target position information of the key feature point, and N is the total number of frames of the intermediate video frame.
[0043] The above image feature matching algorithm can be SIFT (Scale-Invariant Feature Transform), ORB (Oriented FAST and Rotated BRIEF), etc.
[0044] The position change information of the above key feature points usually exists in the form of coordinate information. For each selected feature point, its coordinate position in each frame (or key frame) of the image sequence (intermediate sequence frames) will be recorded. Specific data structure: In a computer, a list or an array can be used to represent the trajectory (position change information) of a key feature point. For example, assume that 3 key feature points are selected and an image sequence consisting of 5 frames (including the starting frame and the target frame, with 3 intermediate video frames) is to be generated. Then, the position change information of the key feature points can be organized in the following form:
[0045] Key feature point 1: [(x1_frame1, y1_frame1), (x1_frame2, y1_frame2), (x1_frame3, y1_frame3), (x1_frame4, y1_frame4), (x1_frame5, y1_frame5)];
[0046] Key feature point 2: [(x2_frame1, y2_frame1), (x2_frame2, y2_frame2), (x2_frame3, y2_frame3), (x2_frame4, y2_frame4), (x2_frame5, y2_frame5)];
[0047] Key feature point 3: [(x3_frame1, y3_frame1), (x3_frame2, y3_frame2), (x3_frame3, y3_frame3), (x3_frame4, y3_frame4), (x3_frame5, y3_frame5)];
[0048] Among them, key feature points 1, 2, and 3 are determined key feature points, frame1, frame2, frame3, frame4, and frame5 represent the video frame numbers of the video, where frame1 is the first video frame and frame5 is the second video frame), and frame2, frame3, and frame4 are intermediate video frames; (x_ij, y_ij) represents the pixel coordinates of the i-th key feature point in the j-th video frame.
[0049] The most core information is the coordinate sequence of each key feature point in different intermediate video frames. These coordinate sequences describe the movement trajectory of the key feature point in the intermediate video frames.
[0050] For the step of inputting the first video frame, the second video frame, and the position change information of the key feature points into the pre-trained target network model to generate at least one intermediate video frame through the target network model, a possible implementation method:
[0051] (1) Determine the guiding and regulating map based on the position change information of key feature points; wherein, the guiding and regulating map serves as the guiding information for the target network model, and is used to prompt the areas that the target network model should focus on when generating intermediate video frames, so that the position information of the key feature points in the generated intermediate video frames conforms to the position change information of the key feature points;
[0052] Determine the position coordinate information of key feature points at different time points; for each key feature point, determine the heat map of the key feature point at the specified time point according to the position coordinate information of the key feature point at the specified time point; perform superposition processing on the heat maps of each key feature point at the specified time point to obtain the feature guiding map of the key feature point at the specified time point.
[0053] Optionally, generate a heat map for each key feature point at each target generation time point, with the position of the key feature point as the center and the pixel intensity conforming to a two-dimensional Gaussian distribution. The specific formula is:
[0054]
[0055] wherein, (x0, y0) is the position information (coordinate) of the key feature point, and σ is the standard deviation, which controls the influence range of the heat map. Obtain the heat maps of each key feature point at the same time point, and superimpose the heat maps of all key feature points at the same time point to generate a guiding and regulating map, which represents the areas that the model should focus on when generating a specific intermediate video frame.
[0056] Specifically, the guiding and regulating map itself is a grayscale image (such as an example diagram of a guiding and regulating map shown Figure 2 ), and its pixel intensity directly reflects the areas that the model should focus on. In the guiding and regulating map, the areas with higher (brighter) pixel intensity values represent the areas that the model should pay more attention to when generating images. These high-intensity areas are the above-mentioned key focus areas. The guiding and regulating map is generated by superimposing Gaussian heat maps. The center of each Gaussian heat map corresponds to the position of a feature point. Therefore, the bright peak areas in the guiding and regulating map usually correspond to the position change information of the key feature points.
[0057] The intensity of the Gaussian heat map is high at the center and gradually decays outward. This smooth change in intensity also means that the degree of attention of the model to the image area changes smoothly rather than abruptly. The closer to the center of the heat map, the higher the degree of attention; the farther away from the center, the lower the degree of attention.
[0058] (2) Input the first video frame, the second video frame, and the guidance control map into a pre-trained target network model, and generate at least one intermediate video frame through the target network model.
[0059] In order for the diffusion model to generate intermediate video frames that meet the conditions, it is necessary to incorporate the first video frame, the second video frame, and the guidance control map into the denoising process of the diffusion model.
[0060] Optionally, encode the first video frame and the second video frame through the first encoder in the target network model to obtain the first image feature and the second image feature; encode the guidance control map through the second encoder in the target network model to obtain the guidance feature; obtain the fused feature according to the first image feature, the second image feature, and the guidance feature, and input the fused feature into the denoising network of the target network model, and obtain at least one intermediate video frame through the denoising network.
[0061] The above first image feature and second image feature contain key information of the image, such as texture, shape, and semantic information. The above first encoder is an image encoder (such as a convolutional neural network), which encodes the first video frame and the second video frame into deep feature representations (i.e., the above first image feature and second image feature). The above second encoder is usually a control branch network similar to a part of the diffusion model structure (such as the encoder part of U-Net), and encodes the guidance control map through the control branch network to extract its deep features (i.e., the guidance feature). Then, fuse the first image feature, the second image feature, and the guidance feature, input them into the denoising network of the target network model, and obtain at least one intermediate video frame through the denoising network.
[0062] Among them, using different encoders to encode the video frame and the guidance control map can decouple the two functions of guidance information processing and image feature extraction, making the model design more modular, easier to understand and expand. Different encoders can be optimized specifically to adapt to different types of data. The encoder of the control branch network can be optimized specifically for the characteristics of the guidance control map. For example, a structure more suitable for processing sparse and highly local data such as heatmap can be designed. The first encoder can focus more on extracting the global visual features of the image. The independent control branch network makes the model more flexible and scalable, and can also easily replace or modify the structure of the control branch network without having a great impact on the backbone network of the diffusion model.
[0063] The above steps of fusing the first image feature, the second image feature, and the guidance feature to obtain the fused feature, inputting the fused feature into the denoising network of the target network model, and obtaining at least one intermediate video frame through the denoising network, a possible implementation method:
[0064] (1) Determine the noise input; perform fusion processing on the first image feature, the second image feature, and the noise input to obtain a first fusion feature;
[0065] (2) Input the first fusion feature into the denoising network of the target network model, and incorporate the guiding feature into the decoder part of the denoising network to obtain at least one intermediate video frame through the denoising network.
[0066] The above-mentioned noise input generally refers to the latent representation of the current frame being generated in each denoising step of the diffusion model. The noise input is a multi-dimensional array containing the noise information of the image.
[0067] Optionally, the first image feature, the second image feature, and the noise input can be concatenated through spatial concatenation to obtain a first fusion feature.
[0068] Optionally, the first image feature, the second image feature, and the noise input can be concatenated through channel concatenation to obtain a first fusion feature.
[0069] For image processing tasks, channel concatenation is a more commonly used method because it can increase the number of feature channels without changing the spatial dimensions of the feature map, thereby fusing feature information from different sources. In this embodiment, the implementation process of the channel concatenation method is mainly described:
[0070] Assume that the noise input, the first image feature, and the second image feature are all four-dimensional vectors with the shape of (B, C, H, W). In fact, B is the batch size, C is the number of channels, H is the height, and W is the width. Channel concatenation is to stack these three vectors along the channel dimension (C dimension) to form a new vector.
[0071] Shapes before concatenation: Noise input: (B, C1, H, W); First image feature: (B, C2, H, W); Second image feature: (B, C3, H, W); The first fusion feature obtained after channel concatenation: (B, C1 + C2 + C3, H, W). The number of channels of the concatenated vector (i.e., the first fusion feature) becomes (C1 + C2 + C3), while the height (H) and width (W) remain unchanged.
[0072] Through concatenation fusion, the noise state (i.e., the noise input) of the current denoising step, as well as the global information of the first video frame and the second video frame, are provided to the denoising network at the same time. This enables the denoising network to "perceive" the start and end states of generating the intermediate video frame when performing the denoising operation, so as to generate intermediate frames that better conform to the overall transition trend.
[0073] The first image feature and the second image feature are like providing "references" or "anchors" for the denoising network. The denoising network will "refer" to these features and combine the current noise input to "guide" the generation process, so that the generated intermediate video frames can not only maintain coherence with the first and last frames, but also gradually achieve the transition from the starting frame to the target frame.
[0074] The above ways of integrating the guiding features into the decoder part of the denoising network can be channel concatenation or adaptive affine transformation.
[0075] Optionally, the guiding features are integrated into the decoder part of the denoising network by means of channel concatenation.
[0076] Each layer of the decoder of the denoising network will receive the feature maps from the corresponding layer of the encoder. This connection is called a skip connection. The guiding features are concatenated with the feature maps of the skip connection and then input into the decoder layer.
[0077] Specifically, first, the guiding regulation map is encoded into a series of feature maps by the second encoder. Assuming that at the skip connection position corresponding to the i-th layer of the decoder of the denoising network, the feature map encoded by the second encoder is F(i), with a shape of {B, C, F(i), Hi, Wi}. At the same time, the feature map S(i) of the i-th layer of the U-Net encoder of the diffusion model is obtained, with a shape of {B, C, S(i), Hi, Wi}. These two feature maps are concatenated in channels to obtain {B, C, F(i)+S(i), Hi, Wi}, and the concatenated fused feature is obtained. The concatenated fused feature is used as the input of the i-th layer of the decoder of the denoising network for subsequent operations. This method is simple and direct, easy to implement. Channel concatenation does not introduce additional parameters, has low computational overhead, and can effectively integrate the guiding features into the features of the decoder.
[0078] Optionally, the guiding features are integrated into the decoder part of the denoising network by means of adaptive affine transformation.
[0079] Adaptive affine transformation is also commonly used at the "skip connection" of the U-Net decoder or inside the "ResBlock" (residual block) of the decoder layer. After the skip connection or ResBlock, an adaptive affine transformation module (Adaptive Affine Transformation Module, AdaAFF) is added. The role of the AdaAFF module is to dynamically adjust the "scale" and "bias" of the skip connection features (or ResBlock output features) according to the guiding features extracted by the regulation branch network. The feature map F(i) after adaptive affine transformation is used as the input of the i-th layer of the decoder of the denoising network.
[0080] The AdaAFF module can adaptively adjust the scale and offset of the skip connection features according to the guidance information, achieving more flexible and effective feature fusion. It can better control the influence degree of the guidance information on the decoder features.
[0081] The denoising process of the above target network model includes multiple denoising steps, and the multiple denoising steps are used to: gradually recover intermediate video frames from the noise through each denoising step; for the step of fusing the first image feature, the second image feature and the noise input to obtain the first fusion feature, a possible implementation manner: calculating the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature; determining the weights of the first image feature and the second image feature according to the correlation; fusing the first image feature, the second image feature and the noise input according to the weights to obtain the first fusion feature.
[0082] Specifically, the cross-attention mechanism is used to calculate the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature.
[0083] The cross-attention mechanism in this embodiment is like an "information retrieval module" or a "knowledge query module". Its role is to enable the denoising network to actively "query" and "refer to" the relevant information contained in the first image feature and the second image feature frames when generating the image features of the current step.
[0084] The core of the cross-attention mechanism is three key inputs (Query, Key, Value), and the current video frame being generated in the current denoising step acts as the Query. Query can be understood as "what do I want to know". Here, Query means "the current video frame being generated in the current denoising step of the denoising network, and what relevant information about the first video frame and the second video frame does it want to query and understand".
[0085] The first image feature and the second image feature together act as the Key and Value. Key and Value can be understood as a "knowledge base" or an "information source". Here, the first image feature and the second image feature as Key and Value mean "the feature information of the first video frame and the second video frame, which can be used as the knowledge source for the denoising network to query".
[0086] The core step of the cross-attention mechanism is to calculate attention scores, also known as relevance scores. The purpose of calculating attention scores is to measure the "degree of correlation" or "similarity" between the Query (the current video frame generated in the current denoising generation step) and the Key (the first image feature and the second image feature). The commonly used calculation method is Dot-Product Attention. Perform a dot product operation on the Query and the Key. If both the Query and the Key are feature vectors, the dot product is the vector inner product. If the Query and the Key are feature maps, the dot product operation can be performed at each position on the feature map.
[0087] To prevent the dot product result from being too large, the dot product result is usually scaled by dividing it by a scaling factor (e.g., the square root of the dimension size of the Key). Perform Softmax normalization on the scaled dot product result. The Softmax function can convert the dot product result into a probability distribution, making the sum of all attention scores equal to 1. The attention scores after Softmax represent the "degree of attention" or "correlation" of the Query (i.e., the current video frame) to different Keys (i.e., the first image feature and the second image feature). The higher the weight, the stronger the correlation and the higher the degree of attention.
[0088] After obtaining the attention scores, the cross-attention mechanism applies the attention scores to the Value (i.e., the first image feature and the second image feature). Specifically, it is to perform a weighted sum on the Value, and the weights are the attention scores. That is, perform a weighted sum on the first image feature and the second image feature according to the attention scores, and then fuse the weighted sum result with the noise input to obtain the first fused feature.
[0089] Since the position change information of the key feature points determined by the interpolation algorithm belongs to the initial trajectory, which is only a rough estimate and cannot perfectly reflect the motion trajectory of the feature points in the real scene. The position information of the key feature points can be adjusted through the following iterative optimization to make it fit the more real motion, thereby improving the quality of the generated video.
[0090] The denoising process of the above target network model includes multiple denoising steps, and the multiple denoising steps are used for: gradually recovering intermediate video frames from the noise through each denoising step; the above step of generating at least one intermediate video frame through the target network model, a possible implementation:
[0091] (1) For each denoising step of the target network model, in the current denoising step, obtain the current video frame being generated in the current denoising step;
[0092] The current video frame being generated in the current denoising step generally refers to the noise latent representation of the current video frame.
[0093] (2) Update the position information of the key feature points in the current video frame according to the position change information of the key feature points, and obtain the updated position information of the key feature points;
[0094] Obtain the feature map of the intermediate output of the denoising network decoder (i.e., the feature maps of the first video frame and the second video frame). Iteratively optimize the position information of the key feature points in the current video frame. Scale the feature map to the target image resolution using bilinear interpolation, where the target image resolution is the resolution of the video frame.
[0095] Optionally, for each key feature point in the current video frame, perform a nearest neighbor search on the feature map of the first video frame to calculate the position information of the first nearest neighbor point of the key feature point on the feature map of the first video frame; perform a nearest neighbor search on the feature map of the second video frame to calculate the position information of the second nearest neighbor point of the key feature point on the feature map of the second video frame; if the distance between the first nearest neighbor point and the second nearest neighbor point is less than a preset threshold, calculate the target position of the key feature point based on the position information of the first nearest neighbor point and the position information of the second nearest neighbor point; determine the target position as the updated position information of the key feature point.
[0096] The search formula is represents the neighborhood centered at p start with a radius of r1. is the position information of the first nearest neighbor point. Similarly, calculate the position information of the second nearest neighbor point through the above search formula
[0097] If and the distance between them is less than the preset threshold, it is considered that the match is reliable, and the target position of the key feature point is calculated through
[0098] (3) Denoise the current video frame according to the updated position information of the key feature points, and obtain the video frame being generated in the next denoising step of the current denoising step;
[0099] According to the updated position information of the key feature points, recalculate the trajectory of the key feature points in the current denoising step, and use it for the trajectory optimization of the next denoising step. That is, the updated position information p k of the key feature points will be immediately used to "guide" the subsequent denoising process of the current denoising step and the trajectory optimization of the next denoising step.
[0100] (4) Until all denoising steps are completed, the final denoising result is obtained, and the denoising result is decoded to obtain intermediate video frames.
[0101] All intermediate video frames are generated iteratively. The generation process of each intermediate video frame is accompanied by the step of "trajectory iterative optimization". Finally, all the generated intermediate video frames, the starting frame, and the target frame are concatenated in chronological order to obtain a complete image sequence.
[0102] Traditional automatic frame interpolation methods are prone to problems such as motion blur and object deformation when dealing with complex motions and large deformations, making it difficult to ensure the generation quality. In this embodiment, the feature of denoising frame by frame using the diffusion model is utilized to synchronously update the position information of key feature points in each denoising step. Specifically, the techniques of "nearest neighbor search of feature maps" and "bidirectional consistency verification" are used to "more accurately estimate the correspondence of feature points between the starting frame and the target frame", "reverse correct the deficiencies of the initial trajectory of linear interpolation", and "iteratively refine the trajectory to make it more conform to the real motion", significantly improving the frame interpolation quality and temporal coherence in the autonomous generation mode.
[0103] The above method further includes: determining a target video according to the first video frame, the second video frame, and at least one intermediate video frame; for each intermediate video frame in the target video, comparing the consistency of the intermediate video frame with the previous video frame to obtain a comparison result; and updating the model parameters of the target network model or performing image processing on the intermediate video frame according to the comparison result.
[0104] Optionally, the consistency of the intermediate video frame is compared with one or more previous video frames. Specifically, the sampling process of the target network model is updated, or the generation result is updated as a post-processing step.
[0105] Updating the sampling process of the target network model: For example, in each denoising step, a candidate frame is generated. Then, the candidate frame is compared with the previous frame for consistency. If the consistency does not meet the threshold requirement, the current candidate frame is rejected and a new candidate frame is resampled until a frame that meets the consistency requirement is found. This method can forcibly ensure temporal consistency but may reduce the generation efficiency.
[0106] Another example is that in the sampling process of the diffusion model, the result of the consistency comparison is used as additional conditional information to guide the sampling direction of the model. For example, the noise prediction network can be modified so that when predicting noise, it not only considers conditions such as the initial frame, the target frame, and the guidance control map, but also considers the consistency with the previous frame.
[0107] Update the generation result as a post - processing step: For example, after generating the target video, for each frame in the video, perform smoothing filtering on it with the previous frame (e.g., median filtering, Gaussian filtering, etc. in the time dimension). This can simply smooth the sudden changes in time and enhance temporal coherence. Another example is that for intermediate video frames that do not meet the consistency requirements, they can be weighted - fused with the previous frame to generate a new frame to replace the original frame.
[0108] The above - mentioned target network model is usually trained based on the following objective function: \(\mathcal{L}=\mathbb{E}_{z_t,z_0,z_n,t,\epsilon\sim\mathcal{N}(0,1)}\left[||\epsilon - \epsilon_{\theta}(z_t;t,z_0,z_n,C_{guide})||^2\right]\);
[0109] Among them, \((\mathcal{L})\) represents the value of the loss function, which is used to measure the gap between the model generation result and the true value. The goal is to minimize the value of the loss function by optimizing the model parameters. \((\mathbb{E})\) represents the expected value, which is the mean value obtained by integrating the random variable. \((z_t)\) represents the noisy image at the \(t\) - th step of the diffusion process, where \((t)\) is the time step of diffusion. \((z_t)\) can be obtained by gradually adding Gaussian noise to the original image. \((z_0)\) and \((z_n)\) represent the encoded features of the initial frame and the target frame, which can be used to guide the generation model. \((\epsilon)\) represents the noise input, which usually follows the standard normal distribution \(\mathcal{N}(0,1)\). \((\epsilon_{\theta})\) represents the noise predicted by the model, which is the output of the denoising network with parameters \((\theta)\). The training goal of the model is to make \((\epsilon_{\theta})\) as close as possible to the true noise \((\epsilon)\). \((C_{guide})\) represents the guiding control graph, which is used to guide the model to generate intermediate video frames that meet the user's intention. \((||\cdot||^2)\) represents the L2 - norm of the vector, which is used to calculate the difference size between two vectors.
[0110] In addition to the above - mentioned main loss term, additional loss terms or regularization terms for "director's guiding signal" or "trajectory correction" can be added to further optimize the generation effect. For example, the optical - flow consistency loss can be used to constrain the feature - point trajectory, or the contrast - learning loss can be used to enhance the visual quality of the generated image.
[0111] Corresponding to the above - mentioned method embodiment, the present disclosure embodiment provides a video - frame generation device, as Figure 3 shown, the device includes:
[0112] A video frame determination module 301 is configured to determine a first video frame and a second video frame from a video to be processed;
[0113] A position change information determination module 302 is configured to determine at least one key feature point in the first video frame and the second video frame, and determine the position change information of the key feature point; wherein, the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame;
[0114] An intermediate video frame generation module 303 is configured to input the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generate at least one intermediate video frame through the target network model; wherein, the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame.
[0115] The embodiment of the present disclosure provides a video frame generation device, which determines a first video frame and a second video frame from a video to be processed; determines at least one key feature point in the first video frame and the second video frame, and determines the position change information of the key feature point; wherein, the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; inputs the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generates at least one intermediate video frame through the target network model; wherein, the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame. This method determines the key feature point and the position change information of the key feature point in the video frame. Through the first video frame, the second video frame, and the position change information of the key feature point, the position information of the key feature point in the intermediate video frame generated by guiding the target network model conforms to the pre-determined position change information, which can accurately control the generation process of the intermediate video frame, realize the personalized creative intention, and generate high-quality, smooth and temporally continuous intermediate video frames.
[0116] The above-mentioned position change information determination module is further configured to: in response to a marking operation on a target feature point in the first video frame and the second video frame, determine the target feature point as the key feature point in the first video frame and the second video frame; in response to a position adjustment operation on the key feature point, determine the position change information of the key feature point.
[0117] The above-mentioned position change information determination module is further configured to: respond to a position adjustment operation for a key feature point, determine the position coordinate information of the adjusted key feature point; determine the position change information of the key feature point according to the position coordinate information of the adjusted key feature point; wherein, the position change information of the key feature point includes the position information of the key feature point at different time points or different video frames.
[0118] The above-mentioned position change information determination module is further configured to: determine the key feature points in the first video frame and the second video frame according to the image feature matching algorithm; determine the position information of the key feature points in the first video frame as the initial position information of the key feature points, and determine the position information of the key feature points in the second video frame as the target position information of the key feature points; determine the position change information of the key feature points through an interpolation algorithm according to the initial position information of the key feature points and the target position information of the key feature points.
[0119] The above-mentioned position change information determination module is further configured to: for each key feature point, determine the position information of the key feature point on the intermediate video frame in the following manner: P t = P start +(P end - P start )*t / N; where, P t is the position information of the key feature point in the t-th intermediate video frame, P start is the initial position information of the key feature point, P end is the target position information of the key feature point, and N is the total number of intermediate video frames.
[0120] The above-mentioned intermediate video frame generation module is further configured to: determine a guidance control map according to the position change information of the key feature points; wherein, the guidance control map is used as the guidance information of the target network model, and is used to prompt the area that the target network model should focus on when generating the intermediate video frame, so that the position information of the key feature points in the generated intermediate video frame conforms to the position change information of the key feature points; input the first video frame, the second video frame and the guidance control map into the pre-trained target network model, and generate at least one intermediate video frame through the target network model.
[0121] The above-mentioned intermediate video frame generation module is further configured to: determine the position coordinate information of the key feature points at different time points; for each key feature point, determine the heat map of the key feature point at the specified time point according to the position coordinate information of the key feature point at the specified time point; perform superposition processing on the heat maps of each key feature point at the specified time point to obtain the feature guidance map of the key feature point at the specified time point.
[0122] The above intermediate video frame generation module is further configured to: encode the first video frame and the second video frame through a first encoder in the target network model to obtain a first image feature and a second image feature; encode the guidance regulation map through a second encoder in the target network model to obtain a guidance feature; obtain a fusion feature according to the first image feature, the second image feature and the guidance feature, input the fusion feature into the denoising network of the target network model, and obtain at least one intermediate video frame through the denoising network.
[0123] The above intermediate video frame generation module is further configured to: determine a noise input; perform a fusion process on the first image feature, the second image feature and the noise input to obtain a first fusion feature; input the first fusion feature into the denoising network of the target network model, and incorporate the guidance feature into the decoder part of the denoising network, and obtain at least one intermediate video frame through the denoising network.
[0124] The above intermediate video frame generation module is further configured to: calculate the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature; determine the weights of the first image feature and the second image feature according to the correlation; perform a fusion process on the first image feature, the second image feature and the noise input according to the weights to obtain a first fusion feature.
[0125] The above intermediate video frame generation module is further configured to: for each denoising step of the target network model, in the current denoising step, obtain the current video frame being generated in the current denoising step; update the position information of the key feature points in the current video frame according to the position change information of the key feature points to obtain the updated position information of the key feature points; perform denoising processing on the current video frame according to the updated position information of the key feature points to obtain the video frame being generated in the next denoising step of the current denoising step; until all denoising steps are completed, obtain the final denoising result, and perform decoding processing on the denoising result to obtain the intermediate video frame.
[0126] The above intermediate video frame generation module is further configured to: for each key feature point in the current video frame, perform a nearest neighbor search on the feature map of the first video frame, and calculate the position information of the first nearest neighbor point of the key feature point on the feature map of the first video frame; perform a nearest neighbor search on the feature map of the second video frame, and calculate the position information of the second nearest neighbor point of the key feature point on the feature map of the second video frame; if the distance between the first nearest neighbor point and the second nearest neighbor point is less than a preset threshold, calculate the target position of the key feature point based on the position information of the first nearest neighbor point and the position information of the second nearest neighbor point; determine the target position as the updated position information of the key feature point.
[0127] The above device further includes a comparison model for: determining a target video according to a first video frame, a second video frame, and at least one intermediate video frame; for each intermediate video frame in the target video, comparing the intermediate video frame with the previous video frames for consistency to obtain a comparison result; and updating the model parameters of the target network model or performing image processing on the intermediate video frame according to the comparison result.
[0128] The video frame generation device provided by an embodiment of the present disclosure has the same technical features as the video frame generation method provided by the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.
[0129] This embodiment further provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above video frame generation method. The electronic device may be a server or a terminal device.
[0130] See Figure 4 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above video frame generation method.
[0131] Furthermore, Figure 4 the electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0132] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which may be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a bidirectional arrow is used in [description], but it does not mean that there is only one bus or one type of bus.
[0133] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step and logic block diagram disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0134] The processor in the above electronic device can implement the following operations in the above video frame generation method by executing machine-executable instructions:
[0135] Determine a first video frame and a second video frame from the video to be processed; determine at least one key feature point in the first video frame and the second video frame, and determine the position change information of the key feature point; wherein, the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; input the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generate at least one intermediate video frame through the target network model; wherein, the position change information of the key feature point is the guiding information of the target network model, the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transitional video frame between the first video frame and the second video frame. This method determines the key feature point and the position change information of the key feature point in the video frame. Through the first video frame, the second video frame, and the position change information of the key feature point, the position information of the key feature point in the intermediate video frame generated by guiding the target network model conforms to the pre-determined position change information, which can accurately control the generation process of the intermediate video frame, realize the personalized creative intention, and generate high-quality, smooth and temporally continuous intermediate video frames.
[0136] The steps of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point include: in response to the marking operation for the target feature point in the first video frame and the second video frame, determine the target feature point as the key feature point in the first video frame and the second video frame; in response to the position adjustment operation for the key feature point, determine the position change information of the key feature point.
[0137] The steps of determining the position change information of the key feature point in response to the position adjustment operation for the key feature point include: in response to the position adjustment operation for the key feature point, determine the position coordinate information of the key feature point after adjustment; according to the position coordinate information of the key feature point after adjustment, determine the position change information of the key feature point; wherein, the position change information of the key feature point includes the position information of the key feature point at different time points or different video frames.
[0138] The steps of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point include: determine the key feature points in the first video frame and the second video frame according to the image feature matching algorithm; determine the position information of the key feature point in the first video frame as the initial position information of the key feature point, and determine the position information of the key feature point in the second video frame as the target position information of the key feature point; according to the initial position information of the key feature point and the target position information of the key feature point, determine the position change information of the key feature point through the interpolation algorithm.
[0139] The steps of determining the position change information of key feature points through the interpolation algorithm include: for each key feature point, determining the position information of the key feature point on the intermediate video frame through the following method: P t = P start + (P end - P start ) * t / N; where P t is the position information of the key feature point in the t-th intermediate video frame, P start is the initial position information of the key feature point, P end is the target position information of the key feature point, and N is the total number of intermediate video frames.
[0140] The steps of inputting the first video frame, the second video frame, and the position change information of the key feature points into a pre-trained target network model to generate at least one intermediate video frame through the target network model include: determining a guidance and regulation map according to the position change information of the key feature points; where the guidance and regulation map is used as the guidance information of the target network model to prompt the areas that the target network model should focus on when generating the intermediate video frame, so that the position information of the key feature points in the generated intermediate video frame conforms to the position change information of the key feature points; inputting the first video frame, the second video frame, and the guidance and regulation map into the pre-trained target network model, and generating at least one intermediate video frame through the target network model.
[0141] The steps of determining the guidance and regulation map according to the position change information of the key feature points include: determining the position coordinate information of the key feature points at different time points; for each key feature point, determining the heat map of the key feature point at the specified time point according to the position coordinate information of the key feature point at the specified time point; performing superposition processing on the heat maps of each key feature point at the specified time point to obtain the feature guidance map of the key feature points at the specified time point.
[0142] The steps of inputting the first video frame, the second video frame, and the guidance and regulation map into the pre-trained target network model to generate at least one intermediate video frame through the target network model include: performing encoding processing on the first video frame and the second video frame through the first encoder in the target network model to obtain the first image feature and the second image feature; performing encoding processing on the guidance and regulation map through the second encoder in the target network model to obtain the guidance feature; obtaining the fusion feature according to the first image feature, the second image feature, and the guidance feature, and inputting the fusion feature into the denoising network of the target network model to obtain at least one intermediate video frame through the denoising network.
[0143] The step of fusing the first image feature, the second image feature, and the guiding feature to obtain a fused feature, and inputting the fused feature into the denoising network of the target network model to obtain at least one intermediate video frame through the denoising network includes: determining a noise input; fusing the first image feature, the second image feature, and the noise input to obtain a first fused feature; inputting the first fused feature into the denoising network of the target network model, and integrating the guiding feature into the decoder part of the denoising network to obtain at least one intermediate video frame through the denoising network.
[0144] The step of fusing the first image feature, the second image feature, and the noise input to obtain a first fused feature includes: calculating the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature; determining the weights of the first image feature and the second image feature according to the correlation; and fusing the first image feature, the second image feature, and the noise input according to the weights to obtain a first fused feature.
[0145] The step of generating at least one intermediate video frame through the target network model includes: for each denoising step of the target network model, in the current denoising step, obtaining the current video frame being generated in the current denoising step; updating the position information of the key feature points in the current video frame according to the position change information of the key feature points to obtain the updated position information of the key feature points; denoising the current video frame according to the updated position information of the key feature points to obtain the video frame being generated in the next denoising step of the current denoising step; until all denoising steps are completed, obtaining the final denoising result, and decoding the denoising result to obtain the intermediate video frame.
[0146] The step of updating the position information of the key feature points in the current video frame to obtain the updated position information of the key feature points includes: for each key feature point in the current video frame, performing a nearest neighbor search on the feature map of the first video frame and calculating the position information of the first nearest neighbor point of the key feature point on the feature map of the first video frame; performing a nearest neighbor search on the feature map of the second video frame and calculating the position information of the second nearest neighbor point of the key feature point on the feature map of the second video frame; if the distance between the first nearest neighbor point and the second nearest neighbor point is less than a preset threshold, calculating the target position of the key feature point based on the position information of the first nearest neighbor point and the position information of the second nearest neighbor point; and determining the target position as the updated position information of the key feature points.
[0147] The above method further includes: determining a target video based on a first video frame, a second video frame, and at least one intermediate video frame; for each intermediate video frame in the target video, comparing the intermediate video frame with the previous video frame for consistency to obtain a comparison result; and updating the model parameters of the target network model or performing image processing on the intermediate video frame according to the comparison result.
[0148] This embodiment further provides a machine-readable storage medium storing machine-executable instructions, which when called and executed by a processor, cause the processor to implement the above video frame generation method.
[0149] The machine-executable instructions stored in the above machine-readable storage medium can, by executing the machine-executable instructions, implement the following operations in the above video frame generation method:
[0150] Determining a first video frame and a second video frame from a video to be processed; determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point; wherein the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model to generate at least one intermediate video frame through the target network model; wherein the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transitional video frame between the first video frame and the second video frame. This method determines the key feature point and the position change information of the key feature point in the video frame, and through the first video frame, the second video frame, and the position change information of the key feature point, guides the position information of the key feature point in the intermediate video frame generated by the target network model to conform to the pre-determined position change information, can accurately control the generation process of the intermediate video frame, realizes the personalized creative intention, and generates high-quality, smooth and temporally continuous intermediate video frames.
[0151] The above step of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point includes: in response to a marking operation on the target feature point in the first video frame and the second video frame, determining the target feature point as the key feature point in the first video frame and the second video frame; and in response to a position adjustment operation on the key feature point, determining the position change information of the key feature point.
[0152] The steps for determining the position change information of the key feature points in response to the position adjustment operation for the key feature points include: in response to the position adjustment operation for the key feature points, determining the position coordinate information of the key feature points after adjustment; and determining the position change information of the key feature points according to the position coordinate information of the key feature points after adjustment; wherein, the position change information of the key feature points includes the position information of the key feature points at different time points or different video frames.
[0153] The steps for determining at least one key feature point in the first video frame and the second video frame and determining the position change information of the key feature points include: determining the key feature points in the first video frame and the second video frame according to the image feature matching algorithm; determining the initial position information of the key feature points as the position information of the key feature points in the first video frame, and determining the target position information of the key feature points as the position information of the key feature points in the second video frame; and determining the position change information of the key feature points through an interpolation algorithm according to the initial position information of the key feature points and the target position information of the key feature points.
[0154] The steps for determining the position change information of the key feature points through an interpolation algorithm include: for each key feature point, determining the position information of the key feature point on the intermediate video frame in the following manner: P t = P start + (P end - P start ) * t / N; where P t is the position information of the key feature point in the t-th intermediate video frame, P start is the initial position information of the key feature point, P end is the target position information of the key feature point, and N is the total number of intermediate video frames.
[0155] The steps for inputting the first video frame, the second video frame, and the position change information of the key feature points into a pre-trained target network model to generate at least one intermediate video frame through the target network model include: determining a guiding control graph according to the position change information of the key feature points; wherein, the guiding control graph serves as the guiding information of the target network model and is used to prompt the areas that the target network model should focus on when generating the intermediate video frame, so that the position information of the key feature points in the generated intermediate video frame conforms to the position change information of the key feature points; and inputting the first video frame, the second video frame, and the guiding control graph into the pre-trained target network model to generate at least one intermediate video frame through the target network model.
[0156] The steps of determining the guiding and regulating map according to the position change information of the key feature points are as follows: determining the position coordinate information of the key feature points at different time points; for each key feature point, determining the heat map of the key feature point at the specified time point according to the position coordinate information of the key feature point at the specified time point; superimposing the heat maps of each key feature point at the specified time point to obtain the feature guiding map of the key feature points at the specified time point.
[0157] The steps of inputting the first video frame, the second video frame and the guiding and regulating map into the pre-trained target network model to generate at least one intermediate video frame by the target network model are as follows: encoding the first video frame and the second video frame by the first encoder in the target network model to obtain the first image feature and the second image feature; encoding the guiding and regulating map by the second encoder in the target network model to obtain the guiding feature; obtaining the fusion feature according to the first image feature, the second image feature and the guiding feature, inputting the fusion feature into the denoising network of the target network model, and obtaining at least one intermediate video frame through the denoising network.
[0158] The steps of fusing the first image feature, the second image feature and the guiding feature to obtain the fusion feature, inputting the fusion feature into the denoising network of the target network model, and obtaining at least one intermediate video frame through the denoising network are as follows: determining the noise input; fusing the first image feature, the second image feature and the noise input to obtain the first fusion feature; inputting the first fusion feature into the denoising network of the target network model, and integrating the guiding feature into the decoder part of the denoising network, and obtaining at least one intermediate video frame through the denoising network.
[0159] The steps of fusing the first image feature, the second image feature and the noise input to obtain the first fusion feature are as follows: calculating the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature; determining the weights of the first image feature and the second image feature according to the correlation; fusing the first image feature, the second image feature and the noise input according to the weights to obtain the first fusion feature.
[0160] The steps of generating at least one intermediate video frame through the target network model described above include: for each denoising step of the target network model, in the current denoising step, obtaining the current video frame being generated in the current denoising step; updating the position information of the key feature points in the current video frame according to the position change information of the key feature points to obtain the updated position information of the key feature points; denoising the current video frame according to the updated position information of the key feature points to obtain the video frame being generated in the next denoising step of the current denoising step; until all denoising steps are completed to obtain the final denoising result, and decoding the denoising result to obtain the intermediate video frame.
[0161] The steps of updating the position information of the key feature points in the current video frame described above to obtain the updated position information of the key feature points include: for each key feature point in the current video frame, performing a nearest neighbor search on the feature map of the first video frame to calculate the position information of the first nearest neighbor point of the key feature point on the feature map of the first video frame; performing a nearest neighbor search on the feature map of the second video frame to calculate the position information of the second nearest neighbor point of the key feature point on the feature map of the second video frame; if the distance between the first nearest neighbor point and the second nearest neighbor point is less than a preset threshold, calculating the target position of the key feature point based on the position information of the first nearest neighbor point and the position information of the second nearest neighbor point; and determining the target position as the updated position information of the key feature point.
[0162] The method described above further includes: determining a target video according to the first video frame, the second video frame, and at least one intermediate video frame; for each intermediate video frame in the target video, comparing the intermediate video frame with the previous video frame for consistency to obtain a comparison result; and updating the model parameters of the target network model or performing image processing on the intermediate video frame according to the comparison result.
[0163] The computer program product of the video frame generation method, device, electronic device, and system provided by the embodiments of the present disclosure includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated here.
[0164] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0165] In addition, in the description of the embodiments of the present disclosure, unless otherwise clearly defined and limited, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.
[0166] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0167] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present disclosure. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0168] Finally, it should be noted that the above embodiments are only specific implementation manners of the present disclosure to illustrate the technical solutions of the present disclosure, rather than limitations thereto. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features within the technical scope disclosed by the present disclosure; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for generating video frames, characterized in that, The method includes: Determine a first video frame and a second video frame from the video to be processed; Determine at least one key feature point in the first video frame and the second video frame, and determine the position change information of the key feature point; wherein, the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; Input the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generate at least one intermediate video frame through the target network model; wherein, the position change information of the key feature point is the guiding information of the target network model, and the position of the key feature point in the intermediate video frame conforms to the position change information of the key feature point, and the intermediate video frame is a transition video frame between the first video frame and the second video frame.
2. The method according to claim 1, wherein The step of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point, includes: Respond to the marking operation for the target feature point in the first video frame and the second video frame, and determine the target feature point as the key feature point in the first video frame and the second video frame; Respond to the position adjustment operation for the key feature point, and determine the position change information of the key feature point.
3. The method according to claim 2, wherein The step of responding to the position adjustment operation for the key feature point and determining the position change information of the key feature point includes: Respond to the position adjustment operation for the key feature point, and determine the position coordinate information of the key feature point after adjustment; According to the position coordinate information of the key feature point after adjustment, determine the position change information of the key feature point; wherein, the position change information of the key feature point includes the position information of the key feature point at different time points or different video frames.
4. The method according to claim 1, wherein The step of determining at least one key feature point in the first video frame and the second video frame, and determining the position change information of the key feature point, includes: Determine the key feature points in the first video frame and the second video frame according to the image feature matching algorithm; Determine the position information of the key feature point in the first video frame as the initial position information of the key feature point, and determine the position information of the key feature point in the second video frame as the target position information of the key feature point; According to the initial position information of the key feature point and the target position information of the key feature point, determine the position change information of the key feature point through the interpolation algorithm.
5. The method according to claim 4, wherein The step of determining the position change information of the key feature point through the interpolation algorithm includes: For each key feature point, determine the position information of the key feature point on the intermediate video frame in the following manner: P t = P start + (P end - P start ) * t / N; Among them, P t is the position information of the key feature points in the middle video frame of the t-th frame, P start is the initial position information of the key feature points, P end is the target position information of the key feature points, and N is the total number of frames of the middle video frame.
6. The method according to claim 1, characterized in that, The step of inputting the first video frame, the second video frame, and the position change information of the key feature point into a pre-trained target network model, and generating at least one intermediate video frame through the target network model includes: Determine a guiding and regulating map according to the position change information of the key feature points; wherein, the guiding and regulating map is used as guiding information for the target network model, and is used to prompt the areas that the target network model should focus on when generating intermediate video frames, so that the position information of the key feature points in the generated intermediate video frames conforms to the position change information of the key feature points; Input the first video frame, the second video frame and the guiding and regulating map into a pre-trained target network model, and generate at least one intermediate video frame through the target network model.
7. The method according to claim 6, wherein The step of determining the guiding and regulating map according to the position change information of the key feature points includes: Determine the position coordinate information of the key feature points at different time points; For each key feature point, determine the heat map of the key feature point at the specified time point according to the position coordinate information of the key feature point at the specified time point; Perform superposition processing on the heat maps of each key feature point at the specified time point to obtain the feature guiding map of the key feature point at the specified time point.
8. The method according to claim 6, wherein The step of inputting the first video frame, the second video frame and the guiding and regulating map into a pre-trained target network model and generating at least one intermediate video frame through the target network model includes: Perform encoding processing on the first video frame and the second video frame through a first encoder in the target network model to obtain a first image feature and a second image feature; Perform encoding processing on the guiding and regulating map through a second encoder in the target network model to obtain a guiding feature; Obtain a fusion feature according to the first image feature, the second image feature and the guiding feature, input the fusion feature into the denoising network of the target network model, and obtain the at least one intermediate video frame through the denoising network.
9. The method according to claim 8, wherein The step of performing fusion processing on the first image feature, the second image feature and the guiding feature to obtain a fusion feature, inputting the fusion feature into the denoising network of the target network model, and obtaining the at least one intermediate video frame through the denoising network includes: Determine a noise input; Perform fusion processing on the first image feature, the second image feature and the noise input to obtain a first fusion feature; Input the first fusion feature into the denoising network of the target network model, and incorporate the guiding feature into the decoder part of the denoising network, and obtain the at least one intermediate video frame through the denoising network.
10. The method according to claim 9, characterized in that, The denoising process of the target network model includes a plurality of denoising steps, and the plurality of denoising steps are used for: gradually recovering the intermediate video frame from the noise through each denoising step; The step of performing fusion processing on the first image feature, the second image feature and the noise input to obtain a first fusion feature includes: Calculate the correlation between the current video frame being generated in the current denoising step and the first image feature and the second image feature; Determine the weights of the first image feature and the second image feature according to the correlation; Fuse the first image feature, the second image feature, and the noise input according to the weight to obtain a first fused feature.
11. The method according to claim 4, wherein The denoising process of the target network model includes multiple denoising steps, and the multiple denoising steps are used to: gradually recover intermediate video frames from the noise through each denoising step; The step of generating at least one intermediate video frame by the target network model includes: For each denoising step of the target network model, in the current denoising step, obtain the current video frame being generated in the current denoising step; Update the position information of the key feature points in the current video frame according to the position change information of the key feature points to obtain the updated position information of the key feature points; Perform denoising processing on the current video frame according to the updated position information of the key feature points to obtain the video frame being generated in the next denoising step of the current denoising step; Until all denoising steps are completed, obtain a final denoising result, and decode the denoising result to obtain an intermediate video frame.
12. The method according to claim 11, wherein The step of updating the position information of the key feature points in the current video frame to obtain the updated position information of the key feature points includes: For each key feature point in the current video frame, perform a nearest neighbor search on the feature map of the first video frame, and calculate the position information of the first nearest neighbor point of the key feature point on the feature map of the first video frame; Perform a nearest neighbor search on the feature map of the second video frame, and calculate the position information of the second nearest neighbor point of the key feature point on the feature map of the second video frame; If the distance between the first nearest neighbor point and the second nearest neighbor point is less than a preset threshold, calculate the target position of the key feature point based on the position information of the first nearest neighbor point and the position information of the second nearest neighbor point; Determine the target position as the updated position information of the key feature point.
13. The method according to claim 1, characterized in that The method further includes: Determine a target video according to the first video frame, the second video frame, and the at least one intermediate video frame; For each intermediate video frame in the target video, compare the intermediate video frame with the previous video frame for consistency to obtain a comparison result; Update the model parameters of the target network model or perform image processing on the intermediate video frame according to the comparison result.
14. A video frame generation device, characterized in that, The device includes: A video frame determination module, configured to determine a first video frame and a second video frame from a video to be processed; A position change information determination module, configured to determine at least one key feature point in the first video frame and the second video frame, and determine the position change information of the key feature point; wherein, the position change information is used to indicate the movement trajectory of the key feature point between the first video frame and the second video frame; An intermediate video frame generation module is configured to input the first video frame, the second video frame, and the position change information of the key feature points into a pre-trained target network model, and generate at least one intermediate video frame through the target network model; wherein, the position change information of the key feature points is the guiding information of the target network model, the positions of the key feature points in the intermediate video frame conform to the position change information of the key feature points, and the intermediate video frame is a transitional video frame between the first video frame and the second video frame.
15. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the video frame generation method according to any one of claims 1-13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the video frame generation method according to any one of claims 1-13.