Video stylization processing methods, devices, equipment and storage media
By identifying keyframes and non-keyframes in the video and utilizing motion vector mapping to style pixel blocks, the problem of high computational cost in deep learning models is solved, thus improving the efficiency and accuracy of video stylization processing.
Patent Information
- Application Number
- CN202610439674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, deep learning models require a large amount of computation to stylize video frames, resulting in low efficiency and long processing time for stylizing short dramas.
By identifying keyframes and non-keyframes in the video, and utilizing the feature similarity and motion vectors of adjacent frames, style pixel blocks of keyframes are mapped to non-keyframes to construct style video frames of non-keyframes, thereby reducing the number of frames processed and the computational load of the deep learning model.
It improves the efficiency of video stylization processing, reduces computational load and time consumption, and ensures the accuracy of stylization processing.
Smart Images

Figure CN122496674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video stylization processing method, apparatus, device, and storage medium. Background Technology
[0002] In recent years, the short drama industry has rapidly risen to prominence due to its characteristics of "short duration, fast pace, and tight plot," becoming an important content format to meet users' fragmented entertainment needs. In the promotion of short dramas, the industry generally adopts a model of "free chapters for user acquisition + paid chapters for monetization." If the same core story can be interpreted in different styles based on the free chapters of the same short drama, different types of user groups can be precisely attracted, improving commercial conversion efficiency. Therefore, secondary creation of short dramas has become a key link connecting users' personalized needs with the industry's commercial promotion, and its technological implementation and efficiency improvement are of great significance to the sustainable development of the short drama industry.
[0003] In existing technologies, stylizing short dramas to convert them into different styles is one of the important methods for secondary creation of short dramas. This mainly involves using pre-trained deep learning models to stylize the short drama frame by frame, obtaining multiple stylized video frames, which are then combined to form the stylized video of the short drama. However, the computational cost of stylizing video frames using deep learning models is high, resulting in a long time consumption for frame-by-frame style conversion of short dramas based on deep learning models, thus affecting the efficiency of stylization processing. Summary of the Invention
[0004] This application provides a video stylization processing method, apparatus, device, and storage medium to stylize keyframes in a video and map the stylized keyframe style pixel blocks to non-keyframes to achieve stylization processing of non-keyframes. This solves the problems of high computational load and time consumption caused by deep learning models converting video styles frame by frame in the prior art, and improves the efficiency of video stylization processing.
[0005] Firstly, this application provides a video stylization processing method, including: Based on the feature similarity of adjacent video frames in the original video, the key frames and non-key frames in the original video are determined; Each of the keyframes is stylized to obtain a first-style video frame; Based on the motion vector between the non-keyframe and the keyframe, the style pixel blocks in the first style video frame are mapped to the second style video frame of the non-keyframe to construct the second style video frame of the non-keyframe. The first style video frames and the second style video frames are sorted and summarized based on their corresponding timestamps to obtain the stylized video of the original video.
[0006] Secondly, this application provides a video stylization processing apparatus, comprising: The video frame classification module is configured to determine key frames and non-key frames in the original video based on the feature similarity of adjacent video frames in the original video. The first stylization processing module is configured to perform stylization processing on each of the keyframes to obtain a first style video frame. The second stylization processing module is configured to map the style pixel blocks in the first style video frame to the second style video frame of the non-key frame according to the motion vector between the non-key frame and the key frame, thereby constructing the second style video frame of the non-key frame. The stylized video generation module is configured to sort and summarize the first stylized video frame and the second stylized video frame based on their corresponding timestamps to obtain a stylized video of the original video.
[0007] Thirdly, this application provides a video stylization processing device, comprising: One or more processors; A memory that stores one or more programs that, when executed by one or more processors, cause the one or more processors to implement the video stylization processing method as described in the first aspect.
[0008] Fourthly, this application provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video stylization processing method as described in the first aspect.
[0009] In this application, keyframes and non-keyframes in the original video are determined by the feature similarity of adjacent video frames. Each keyframe is stylized to obtain a first-style video frame. Based on the motion vectors between the non-keyframes and keyframes, style pixel blocks in the first-style video frame are mapped to a second-style video frame, constructing a second-style video frame for the non-keyframes. The first-style and second-style video frames are sorted and summarized based on their corresponding timestamps to obtain the stylized video of the original video. Through these techniques, the style pixel blocks of the stylized keyframes can be mapped to the non-keyframes using the motion vectors between them, thus achieving style transfer from stylized keyframes to non-keyframes and ensuring the accuracy of video stylization. The computational cost of motion vector and pixel block mapping is far lower than that of deep learning models for converting video frame styles. Deep learning models effectively reduce the number of frames processed by using only keyframe stylization, thereby reducing the computational cost and time, and improving the efficiency of video stylization. Attached Figure Description
[0010] Figure 1 This is a flowchart of a video stylization processing method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the determination of keyframes and non-keyframes provided in an embodiment of this application; Figure 3 This is a schematic diagram showing the original video being expanded into multiple video frames according to an embodiment of this application; Figure 4 This is a flowchart of generating a first-style video frame provided in an embodiment of this application; Figure 5 This is a flowchart of the stylization process of non-key frames based on adjacent key frames provided in this application embodiment; Figure 6 This is a schematic diagram illustrating the construction process of the second-style video frame provided in an embodiment of this application; Figure 7 This is a flowchart of generating a second style video frame based on candidate style frame fusion provided in an embodiment of this application; Figure 8 This is a schematic diagram of the fusion process of the second style video frames provided in the embodiments of this application; Figure 9 A schematic diagram of the structure of a video stylization processing device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a video stylization processing device provided in an embodiment of this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. A process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0012] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0013] In a common existing approach, a pre-trained deep learning model is used to stylize the short drama frame by frame to obtain multiple stylized video frames, which are then combined to form the stylized video of the short drama. However, the computational cost of stylizing video frames using a deep learning model is high, resulting in a long processing time for converting the style of the short drama frame by frame based on the deep learning model, thus affecting the efficiency of the stylization process.
[0014] To address the aforementioned issues, this embodiment provides a video stylization processing method. This method stylizes keyframes in a video and maps the stylized keyframe style pixel blocks to non-keyframes, thereby achieving stylization processing of non-keyframes. By using stylization processing only on keyframes, the deep learning model effectively reduces the number of frames it processes, thus reducing the computational load and time of the deep learning model and improving the efficiency of video stylization processing.
[0015] The video stylization processing method provided in this embodiment can be executed by a video stylization processing device. This device can be implemented through software and / or hardware, and can consist of two or more physical entities, or a single physical entity. For example, the video stylization processing device can be a server or client providing video style conversion services, or a client connected to a server. For instance, some short video platforms provide video style conversion functions, in which case the stylization processing device can be a client of the short video platform or a backend server of the short video platform.
[0016] The video stylization processing device is equipped with at least one type of operating system, including but not limited to Android, Linux, and Windows. The device can install at least one application based on the operating system; this application can be a built-in application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the video stylization processing device has at least one application capable of executing video stylization processing methods. For example, a short video platform may have a video style conversion function; when implementing this function, the short video platform executes a video stylization processing method. The short video platform itself can also be an application executing the video stylization processing method.
[0017] For ease of understanding, this embodiment uses a server as the main entity executing the video stylization processing method as an example for description.
[0018] Figure 1 A flowchart of a video stylization processing method provided in an embodiment of this application is given. (Reference) Figure 1 The video stylization process specifically includes: S110. Based on the feature similarity of adjacent video frames in the original video, determine the key frames and non-key frames in the original video.
[0019] The original video is the video to be stylized. For example, users can upload the original video through the front-end interface of the video style conversion service provided by the server. The front-end transmits the uploaded original video to the back-end of the server, where the back-end performs stylization processing on the original video.
[0020] A keyframe can be understood as a scene transition frame in the original video, representing the start of a scene change in the original video from this frame. That is, when the content of a video frame differs significantly from the previous video frame, it can be identified as a keyframe. Therefore, the keyframe's status can be determined by the feature similarity between the current video frame and the previous video frame. Specifically, after the server obtains the original video, it extracts video frames sequentially according to the video frame time sequence. The first extracted video frame is designated as the keyframe. Then, for each extracted video frame, feature points are matched with the previously extracted video frame, and the feature similarity is determined based on the number of matched feature points. The feature similarity is compared with a preset similarity threshold. If the feature similarity is greater than or equal to the preset similarity threshold, it indicates that the content of the currently extracted video frame is highly similar to the previously extracted video frame, and the currently extracted video frame is determined to be a non-keyframe. If the feature similarity is less than the preset similarity threshold, it indicates that the content of the currently extracted video frame is dissimilar to the previously extracted video frame, and the currently extracted video frame is determined to be a keyframe.
[0021] It should be noted that this embodiment aims to prioritize the use of a deep learning model to stylize keyframes, and then transfer the visual style of the stylized keyframes to non-keyframes with similar content. In this case, the more similar the content of the keyframes and non-keyframes, the higher the integrity of the style transfer and the higher the stylization accuracy of the non-keyframes. Therefore, the more keyframes there are, the higher the accuracy of the video stylization processing, but the lower the processing efficiency. The number of keyframes can be set according to the video length, and the number of keyframes is determined by a similarity threshold. The higher the similarity threshold, the more keyframes there are. The similarity threshold can be set according to the range of keyframe numbers.
[0022] In addition, if the feature similarity between any two adjacent video frames in a video segment is greater than or equal to the similarity threshold, even if the adjacent video frames are highly similar, there will still be some differences between the first and last frames of the video segment. When the first or last frame of the video segment is used as a keyframe and its stylized video frame is used to stylize other video frames in the segment, video frames closer to the keyframe can still maintain a certain level of stylization accuracy. However, as the video frames continue, the number of frames between the video frame and the keyframe increases, leading to more obvious differences between the video frame and the keyframe, and subsequent stylized video frames cannot maintain high stylization accuracy. To further improve the stylization accuracy of video frames, when multiple highly similar video frames appear consecutively, the most recently extracted video frame can be designated as the keyframe.
[0023] Optional, Figure 2 This is a flowchart illustrating the determination of keyframes and non-keyframes provided in an embodiment of this application. For example... Figure 2 As shown, the steps for determining keyframes and non-keyframes specifically include S1101-S1104: S1101. Based on the chronological order of each video frame in the original video, traverse each video frame, determine the first traversed video frame as the key frame, and determine the feature similarity between each traversed video frame and the next video frame.
[0024] For example, video frames are extracted in chronological order from the original video, and the first and last video frames are identified as keyframes. Feature point matching is performed between the first and next video frames (i.e., the second frame), and the feature similarity between the first and second video frames is determined based on the number of matched feature points. The determination of whether the second video frame is a keyframe is based on steps S1102-S1104. Then, the feature similarity between the second and third video frames is calculated to determine whether the third video frame is a keyframe, and so on, until the last video frame is determined as a keyframe.
[0025] S1102. If the feature similarity between a video frame and the next video frame is less than a preset similarity threshold, then the video frame and the corresponding next video frame are determined as keyframes.
[0026] For example, Figure 3 This is a schematic diagram illustrating how the original video is expanded into multiple video frames according to an embodiment of this application. Figure 3 As shown, the first video frame is designated as keyframe A, and the last video frame is designated as keyframe E. Then, the feature similarity between each video frame and the next is calculated. For example, the feature similarity between the first and second video frames is 95%, between the second and third video frames is 92%, between the third and fourth video frames is 98%, and between the fifth video frame is 85%. Assuming a preset similarity threshold of 90%, the fourth and fifth video frames are designated as keyframe B and keyframe C, respectively.
[0027] S1103. If the feature similarity between a video frame and the next video frame is greater than or equal to a preset similarity threshold and the interval between the video frame and the previous key frame is a preset number of frames, then the video frame is determined as a key frame.
[0028] refer to Figure 3 Assuming that starting from the fifth video frame, the feature similarity between each subsequent video frame and the next video frame is greater than or equal to a preset similarity threshold, and the interval between the i-th video frame and the fifth video frame reaches a preset number of frames, then the i-th video frame is designated as a keyframe D. It's understandable that although the i-th video frame is highly similar to the (i-1)-th video frame, it differs from the fifth video frame. If the i-th video frame is designated as a non-keyframe, then when using the stylized video frame of the fifth video frame to stylize the i-th video frame, the differing parts cannot be successfully styled, affecting the stylization accuracy of the i-th video frame. Therefore, to ensure the accuracy of video frame stylization, once the currently traversed video frame is within a preset number of frames of the previous keyframe, it must be designated as a keyframe, even if it is highly similar to the previous video frame. The preset number of frames can be set according to the video stylization accuracy; for example, the higher the accuracy requirement, the smaller the preset number of frames.
[0029] S1104. If the feature similarity between a video frame and the next video frame is greater than or equal to a preset similarity threshold, but there is no preset number of frames between the video frame and the previous key frame, then the video frame is determined to be a non-key frame.
[0030] For example, if the feature similarity between a video frame and the next video frame is greater than or equal to a preset similarity threshold and the number of frames between the video frame and the previous keyframe is less than a preset number of frames, then the video frame is determined to be a non-keyframe. Figure 3 All video frames except keyframes are non-keyframes.
[0031] It needs to be explained that, Figure 3 The adjacent keyframes shown can be either consecutive or non-consecutive video frames. Keyframes characterize video transitions; therefore, if non-keyframes exist between two adjacent keyframes, it indicates that these non-keyframes belong to the same scene as the preceding and following keyframes. In other words, the non-keyframes are highly similar to their preceding and following keyframes, and these non-keyframes, along with their preceding and following keyframes, can form a video segment within the same scene. The preceding and following keyframes are precisely the first and last frames of this video segment. This embodiment aims to select the first and last frames of a video segment within the same scene as keyframes, so that subsequent stylized video frames based on the first and last frames can be used to stylize the non-keyframes in between. This bidirectional style visual transfer can effectively improve the accuracy of stylization processing for non-keyframes.
[0032] S120. Stylize each keyframe to obtain the first style video frame.
[0033] For example, a pre-trained deep learning model is used to stylize keyframes, and the first-style video frame output by the deep learning model is determined as the stylized video frame. For instance, the deep learning model can be a generative network, pre-trained with a large number of samples of the same style. The trained generative network can perform forward propagation calculations on the input keyframes to output the corresponding first-style video frame. In this case, one generative network is suitable for generating images of one style. If the server needs to implement multiple stylization processes, multiple generative networks can be set up to generate video frames of different styles.
[0034] Deploying multiple generative networks consumes excessive server resources and incurs high training costs. To address this, keyframes can be transformed into first-style video frames that satisfy the desired style conversion by leveraging content features in the keyframes and style features from the style image to be converted. Specifically, Figure 4 This is a flowchart illustrating the generation of a first-style video frame, provided in an embodiment of this application. For example... Figure 4 As shown, the steps for generating the first style video frame specifically include S1201-S1203: S1201. Extract the content features of keyframes and the style features of preset style images through a feature extraction network.
[0035] The preset style image is a style image with a pre-defined style to be converted. The feature extraction network can use the first 16 layers of the VGG19 network, with a convolution kernel size of 3×3 and padding of 1, ensuring that the size of the feature map output by the feature extraction network is consistent with the input image. Keyframes are input into the feature extraction network, which calculates the content features of the keyframes through forward propagation, using the feature maps of the conv4_2 and conv5_2 layers of the feature extraction network. The feature map of the conv4_2 layer retains the mid-to-high-level semantic information of the keyframes, such as object outlines and structural layout, while the feature map of the conv5_2 layer retains the high-level abstract information of the keyframes, such as scene category and subject position. The preset style image is input into the feature extraction network, which calculates the style features of the preset style image through forward propagation, using the feature maps of the conv1_1, conv2_1, conv3_1, and conv4_1 layers of the feature extraction network, so that the style features retain the low-frequency texture, mid-frequency structure, and high-frequency details of the style image.
[0036] S1202. The style features and content features are fused together using a style fusion network to obtain a fused feature map.
[0037] For example, the style fusion network is a residual network used to fuse style features and content features. Feature fusion can be achieved through the following steps: Input style features and content features into the residual network to obtain the fused feature map output by the intermediate layers of the residual network; calculate the content loss value of the content feature map and the fused feature map using mean squared error. Calculate the Gram matrix of the style features using the style feature map, which describes the correlation between different feature maps; calculate the Gram matrix of the fused features using the fused feature map; calculate the style loss value of the style feature Gram matrix and the fused feature Gram matrix using mean squared error; calculate the total loss value based on the style loss value and the content loss value; perform backpropagation based on the total loss value to iteratively optimize the network parameters of the residual network until the total loss value converges. After the total loss value converges, input the style features and content features into the residual network to obtain the fused feature map output by the residual network.
[0038] S1203. The first style video frame of the key frame is reconstructed based on the fused feature map by the image reconstruction network.
[0039] For example, the image reconstruction network is a deconvolutional network. The input fused feature map is upsampled step by step through the deconvolutional layers in the deconvolutional network, and the output feature map is the same size as the key frame. Then, the pixel values in the feature map are mapped to the [0,1] interval through the Sigmoid activation function, and then multiplied by 255 to convert it into an 8-bit RGB image format to obtain the first style video frame of the key frame.
[0040] This embodiment uses a feature extraction network, a style fusion network, and a network reconstruction network to transfer the visual style of a style image of any style to a keyframe, thereby enabling multiple stylization processes for the keyframe and avoiding the deployment of multiple generative networks on the server.
[0041] It should be noted that since all video frames are converted to the same style, the same style image is used for stylization of each keyframe. Therefore, after the feature extraction network extracts the style features of the style image for the first time, it can save the style features and the corresponding Gram matrix for use when processing other keyframes. This reduces the computational load on the server for stylizing keyframes and improves the efficiency of keyframe stylization. Furthermore, although the style fusion network requires iterative optimization, its network parameters are pre-trained after the first keyframe is stylized. Subsequent adjustments only require fine-tuning the network parameters to achieve convergence of the total loss value, further improving the efficiency of keyframe stylization. In summary, although the feature extraction network, style fusion network, and network reconstruction network involve the optimization process of the style fusion network during stylization, the server only consumes significant computational resources and time for processing the first keyframe; processing other keyframes does not excessively consume time and resources. However, processing every video frame in this way still places a huge computational burden on the server, potentially causing service congestion. Therefore, for non-key frames, stylization processing is still performed using step S130 to reduce the computational burden on the server and improve the server's operational stability.
[0042] S130. Based on the motion vector between non-keyframes and keyframes, map the style pixel blocks in the first style video frame to the second style video frame of the non-keyframes to construct the second style video frame of the non-keyframes.
[0043] In this application, the inventors discovered that keyframes and non-keyframes in the same scene are highly similar in content, meaning that most pixel blocks in the keyframe can be found in the non-keyframe in the same scene. When the keyframe has been converted into a first-style video frame, if the motion vector between the same pixel block in the keyframe and the non-keyframe is determined, the pixel block in the first-style video frame can be mapped to the correct position in the second-style video frame of the non-keyframe based on the motion vector. This achieves style transfer from the visual style of the keyframe to the style of the non-keyframe without the need for a deep learning model to stylize the non-keyframe, thus improving the efficiency of stylization processing of the non-keyframe.
[0044] For example, if the first video frame is determined to be a keyframe, and the next video frame among adjacent keyframes with a feature similarity greater than or equal to a similarity threshold is also a keyframe, then a non-keyframe is in the same scene as its preceding keyframe. In this case, motion vectors between the non-keyframe and its preceding keyframe can be calculated. Based on these motion vectors, style pixel blocks in the preceding keyframe are mapped to the corresponding positions in the second-style video frame of the non-keyframe, thus constructing the second-style video frame for the non-keyframe. Alternatively, if the last video frame is determined to be a keyframe, and the previous video frame among adjacent keyframes with a feature similarity greater than or equal to a similarity threshold is also a keyframe, then a non-keyframe is in the same scene as its following keyframe. In this case, motion vectors between the non-keyframe and its following keyframe can be calculated. Based on these motion vectors, style pixel blocks in the following keyframe are mapped to the corresponding positions in the second-style video frame of the non-keyframe, thus constructing the second-style video frame for the non-keyframe.
[0045] In addition, stylized non-keyframes can be used to stylize adjacent unstylized non-keyframes. This avoids significant pixel differences between non-keyframes and keyframes in the same scene due to large frame intervals, which could lead to numerous air gaps in the second-styled video frames of non-keyframes and affect the final processing effect. For example, if a non-keyframe and its preceding keyframe are in the same scene, the preceding keyframe is first used to stylize the subsequent non-keyframe, and then the stylized non-keyframe is used to stylize the subsequent non-keyframe, and so on, until all non-keyframes in the same scene as the preceding keyframe are stylized. Alternatively, if a non-keyframe and its following keyframe are in the same scene, the following keyframe is first used to stylize the preceding non-keyframe, and then the stylized non-keyframe is used to stylize the preceding non-keyframe, and so on, until all non-keyframes in the same scene as the following keyframe are stylized.
[0046] Depend on Figure 3 As shown in the illustration, this application can also simultaneously determine the first and last frames of a video segment in the same scene as keyframes. That is, the preceding and following keyframes of a non-keyframe are both in the same scene. In this case, the non-keyframe can be stylized using the first-style video frames of the preceding and following keyframes simultaneously to construct a second-style video frame for the non-keyframe. Specifically, Figure 5 This is a flowchart illustrating the stylization process of non-keyframes based on adjacent keyframes provided in this application. Figure 5 As shown, the step of stylizing non-keyframes based on adjacent keyframes specifically includes S1301-S1302: S1301. Obtain the non-key frame between two adjacent key frames, and determine the bidirectional motion vector between the non-key frame and the adjacent key frames before and after.
[0047] refer to Figure 3 Keyframe A and keyframe B are two adjacent keyframes. The non-keyframes between keyframe A and keyframe B can be obtained. Figure 3 The video frames 2, 3, and 4 are used to calculate the forward motion vector between the non-keyframe and keyframe A, and the reverse motion vector between the non-keyframe and keyframe B.
[0048] This embodiment describes the calculation of bidirectional motion vectors between video frame 2 and keyframes A and B as an example. When calculating the forward motion vector, video frame 2 is divided into multiple 16×16 pixel blocks. These blocks do not overlap. The coordinates of the top-left corner pixel of each block are determined as its coordinates. Since the range of motion of objects in keyframes A and B is limited, a larger rectangular region can be defined around each pixel block as the search area. For example, if the pixel block's coordinates are (x1, y1) and the rectangular region's size is [-p, p], then the pixel block in keyframe A is searched within the rectangular region of video frame 2 from (x1-p, y1-p) to (x1+p, y1+p). A 16×16 candidate block is extracted by sliding within the rectangular region of keyframe A with a step size of one pixel. The mean square error (MSE) between the pixel block in video frame 2 and the corresponding pixel grayscale value in each candidate block is calculated. The smaller the MSE, the higher the matching degree. The candidate block with the smallest MSE is determined as the first matching block. The coordinate difference between a pixel block and the first matching block is determined as the positive motion vector. For example, if the top-left corner coordinates of the first matching block are (x2, y2), then the positive motion vector is (x1-x2, y1-y2). This represents that when the pixel block moves from keyframe A to video frame 2, it moves x1-x2 pixels to the right and y1-y2 pixels down. Similarly, the first matching block and positive motion vector corresponding to each pixel block in video frame 2 can be determined. The coordinates of each first matching block and its corresponding positive motion vector are then combined to obtain the positive motion vector between keyframe A and video frame 2.
[0049] When calculating the reverse motion vector, a 16×16 candidate block is selected by sliding across the rectangular region of keyframe B with a step size of one pixel. The mean square error (MSE) of the grayscale values of the corresponding pixels in each candidate block is calculated between the pixel block in video frame 2 and the corresponding pixel in video frame 2. The smaller the MSE, the higher the matching degree. The candidate block with the smallest MSE is determined as the second matching block of the pixel block. The coordinate difference between the pixel block and the second matching block is determined as the reverse motion vector of the pixel block. For example, if the coordinates of the upper left corner of the second matching block are (x3, y3), then the reverse motion vector is (x1-x3, y1-y3). This represents that when the pixel block moves from keyframe B to video frame 2, it moves to the right by x1-x3 pixels and downwards by y1-y2 pixels. By analogy, the second matching block and the reverse motion vector corresponding to each pixel block in video frame 2 can be determined. The coordinates of each second matching block and the corresponding reverse motion vector are summed to obtain the reverse motion vector between keyframe B and video frame 2.
[0050] S1302. Based on the bidirectional motion vectors between the non-keyframe and the adjacent keyframes, map each style pixel block in the first style video frame of the adjacent keyframes to the second style video frame of the non-keyframe to construct the second style video frame of the non-keyframe.
[0051] For example, Figure 6 This is a schematic diagram illustrating the construction process of the second-style video frame provided in an embodiment of this application. For example... Figure 6As shown, the forward motion vector between pixel block 11 in video frame 2 and the first matching block 12 of the preceding keyframe A is (x1-x2, y1-y2), and the reverse motion vector between pixel block 11 in video frame 2 and the second matching block 13 of the following keyframe B is (x1-x3, y1-y3). Based on the coordinates (x2, y2) of the first matching block 12 in keyframe A, the corresponding first-style pixel block 15 can be determined in the corresponding first-style video frame A. The first-style pixel block 15 can be considered as a pixel block stylized from the first matching block 12. Similarly, based on the coordinates (x3, y3) of the second matching block 13 in keyframe B, the corresponding second-style pixel block 16 can be determined in the corresponding first-style video frame. The second-style pixel block 16 can be considered as a pixel block stylized from the second matching block 13. Based on the forward motion vector of the first matching block 12, the first style pixel block 15 can be mapped to the coordinates (x1, y1) of the second style video frame. Based on the reverse motion vector of the second matching block 13, the second style pixel block 16 can be mapped to the coordinates (x1, y1) of the second style video frame. At this time, the coordinates (x1, y1) contain both the first style pixel block 15 and the second style pixel block 16. The first style pixel block 15 and the second style pixel block 16 can be weighted and summed to obtain the target style pixel block of the second style video frame at coordinates (x1, y1), thus realizing the conversion of the pixel block at (x1, y1) in video frame 2 into the target style pixel block. By analogy, each pixel block in video frame 2 can be converted into the corresponding target style pixel block one by one, and each target style pixel block can construct the second style video frame of video frame 2.
[0052] This embodiment utilizes adjacent keyframes before and after non-keyframes to perform stylization processing on non-keyframes, thereby creating bidirectional motion and style constraints on non-keyframes, reducing local style deviations caused by a single keyframe, and improving the stylization processing accuracy of non-keyframes.
[0053] In another embodiment, candidate style frames corresponding to non-key frames can be constructed using the preceding and following keyframes, and then these candidate style frames can be fused. Specifically, Figure 7 This is a flowchart illustrating the generation of a second-style video frame based on candidate style frame fusion, provided in an embodiment of this application. For example... Figure 7 As shown, the step of generating a second style video frame based on candidate style frame fusion specifically includes S13021-S13024: S13021. Based on the positive motion vectors of the non-keyframe and the previous keyframe, map the style pixel block of the first style video frame of the previous keyframe to the first candidate style frame of the non-keyframe to construct the first candidate style frame of the non-keyframe.
[0054] For example, based on the coordinates of each first matching block of keyframe A and the corresponding positive motion vector, each first matching block is mapped to the first candidate style frame of video frame 2 one by one, and finally the first candidate style frame is constructed.
[0055] S13022. Based on the inverse motion vectors between the non-keyframe and the adjacent keyframe, map the style pixel blocks of the first style video frame of the adjacent keyframe to the second candidate style frame of the non-keyframe to construct the second candidate style frame of the non-keyframe.
[0056] For example, based on the coordinates of each second matching block in keyframe B and the corresponding reverse motion vector, each second matching block is mapped to the second candidate style frame in video frame 2 one by one, and finally the second candidate style frame is constructed.
[0057] S13023. Determine the fusion weight based on the number of frames between the non-key frame and the preceding and following key frames.
[0058] For example, the larger the interval frame number between a non-keyframe and its adjacent keyframe, the higher the correlation between the non-keyframe and its corresponding adjacent keyframe. This results in a higher visual consistency between the candidate style frames constructed from the adjacent keyframes and the non-keyframes. Therefore, the weight coefficients for fusing the first and second candidate style frames can be determined based on the interval frame number between the non-keyframe and its preceding and following keyframes. Specifically, the first interval frame number between the non-keyframe and its preceding keyframe, and the second interval frame number between the non-keyframe and its following keyframe can be calculated. The sum of the first and second interval frame numbers is then calculated. The ratio of the first interval frame number to the sum of the interval frame numbers is used as the fusion weight coefficient for the second candidate style frame, and the ratio of the second interval frame number to the sum of the interval frame numbers is used as the fusion weight coefficient for the first candidate style frame.
[0059] S13024. The first candidate style frame and the second candidate style frame are weighted and fused according to the fusion weight to obtain the second style video frame, which is a non-key frame.
[0060] For example, the first candidate style frame is multiplied by its corresponding fusion weight coefficient, and the second candidate style frame is multiplied by its corresponding fusion weight coefficient. Then, the pixel values at the same coordinates are summed to obtain the second style video frame. In this embodiment, the correlation between non-key frames and their preceding and following key frames is determined by the number of frames between the preceding and following key frames and non-key frames. Then, the fusion weight coefficient of the corresponding candidate style frame is determined based on the correlation, so as to increase the proportion of candidate style frames with higher correlation in the second style video frame and improve the visual consistency between the second style video frame and the generated first style video frame.
[0061] Optionally, if a pixel block in a non-keyframe has a mean square error greater than a preset error threshold when searching for a matching block within the search area of an adjacent keyframe, it indicates that the pixel block does not match any candidate block within the search area. In this case, it can be determined that the pixel block in the non-keyframe is a newly appearing pixel block relative to the adjacent keyframe, and therefore, it is determined that the pixel block does not have a matching block in the adjacent keyframe. Consequently, when constructing candidate style frames for non-keyframes based on the matching blocks of adjacent keyframes, there will be pixel gaps at the locations of pixel blocks that have not been matched. However, this embodiment aims to construct a second style video frame for non-keyframes using the preceding and following keyframes. When a newly appearing pixel block in a non-keyframe does not find a matching block in the preceding or following keyframe, it may find a matching block in the following or preceding keyframe. In this case, the style pixel block that exists alone as a matching block in the following or preceding keyframe can be directly mapped to the second style video frame to fill in the missing pixel blocks in the second style video frame.
[0062] Specifically, common pixel positions and independent pixel positions are determined in the first candidate style frame and the second candidate style frame; at the common pixel position, pixel values exist in both the first and second candidate style frames, while at the independent pixel position, pixel values exist only in the first or second candidate style frame; the pixel values at the common pixel position in the first and second candidate style frames are weighted and summed according to the fusion weight to obtain the fused pixel value, and the fused pixel value is mapped to the corresponding pixel position in the second style video frame; the pixel values at the independent pixel positions in the first or second candidate style frame are mapped to the corresponding pixel positions in the second style video frame. Figure 8 This is a schematic diagram illustrating the fusion process of second-style video frames provided in an embodiment of this application. For example... Figure 8As shown, the first pixel 21 in the first candidate style frame and the second pixel 20 in the second candidate style frame have the same pixel position and both have pixel values. Therefore, the pixel positions of the first pixel 21 and the second pixel 20 are determined as common pixel positions. Then, the pixel values of the first pixel 21 and the second pixel 20 are weighted and summed to obtain the pixel value at the corresponding common pixel position in the second style video frame. This pixel value is then mapped to the corresponding common pixel position in the second style video frame to obtain the third pixel 22. The fourth pixel 19 in the first candidate style frame and the fifth pixel 18 in the second candidate style frame have the same pixel position, but the fourth pixel 19 has no pixel value while the fifth pixel 18 has a pixel value. Therefore, the pixel positions of the fourth pixel 19 and the fifth pixel 18 are determined as independent pixel positions. Then, the pixel value of the fifth pixel 18 is mapped to the corresponding independent pixel position in the second style video frame to obtain the sixth pixel 17. That is, the pixel value of the sixth pixel 17 is equal to the pixel value of the fifth pixel 18. This embodiment directly maps the pixel values that exist independently in the candidate style frame to the second style video frame to fill in the missing pixel blocks in the second style video frame and improve the stylization processing accuracy of non-key frames.
[0063] It should be noted that if a pixel block in a non-keyframe has no matching block in either of the preceding or following keyframes, the candidate style frames constructed from the preceding and following keyframes will still have pixel blocks with missing pixel values after fusion. In this case, the pixel values around the missing pixel block can be used to fill in the internal pixel values of the missing pixel block, thereby constructing a complete second style video frame. For example, after the first and second candidate style frames are fused to construct the second style video frame, if the second style video frame has a hole region, the pixel gradient value of the hole region is determined based on the pixel region in the non-keyframe corresponding to the hole region; based on the pixel gradient value of the hole region, a diffusion algorithm is used to diffuse the pixel values of the edge pixels of the hole region in the second style video frame from the outer edge of the hole region to the interior. (Reference) Figure 8 In the first candidate style frame, pixel block 23 and pixel block 24 are at the same pixel location but neither has a pixel value. This results in pixel block 25 at the corresponding location in the second style video frame forming a hole region lacking pixel values. To fill the internal pixel values of the hole region, the corresponding pixel region can be located in the non-keyframe based on the location of the hole region. The pixel gradient value is calculated based on the pixel value of the pixel region. Then, the internal pixel values of the edge pixels of the hole region are determined layer by layer from the outside to the inside according to the pixel gradient value. Once the innermost pixel value is determined, the hole region is filled. This embodiment simulates filling the hole region with pixels of similar style to the surrounding pixels in the non-keyframe by simulating the texture in the non-keyframe and the style of the keyframe, so that the filled hole region integrates the texture of the non-keyframe and the style of the keyframe, improving the stylization processing accuracy of the non-keyframe.
[0064] S140. Sort and summarize the first style video frames and the second style video frames based on their corresponding timestamps to obtain the stylized video of the original video.
[0065] For example, the stylized video of the original video can be obtained by sorting the first style video frames and the second style video frames according to the order of their corresponding timestamps.
[0066] In summary, the video stylization processing method provided in this application determines keyframes and non-keyframes in the original video by using the feature similarity of adjacent video frames; stylizes each keyframe to obtain a first-style video frame; maps style pixel blocks in the first-style video frame to a second-style video frame of the non-keyframe based on the motion vectors between the non-keyframe and the keyframe, thus constructing a second-style video frame of the non-keyframe; and sorts and summarizes the first-style video frame and the second-style video frame based on their corresponding timestamps to obtain the stylized video of the original video. Through the above technical means, the style pixel blocks of the stylized keyframes can be mapped to the non-keyframes using the motion vectors between the keyframes and non-keyframes, thereby achieving style transfer from stylized keyframes to non-keyframes and ensuring the processing accuracy of video stylization. The computational cost of motion vector and pixel block mapping is far lower than the computational cost of deep learning models converting video frame styles. Deep learning models effectively reduce the number of frames processed by only stylizing keyframes, thereby reducing the computational cost and time of deep learning models and improving the efficiency of video stylization processing.
[0067] Based on the above embodiments, Figure 9 This is a schematic diagram of a video stylization processing apparatus provided in an embodiment of this application. (Reference) Figure 9 The video stylization processing device provided in this embodiment specifically includes: a video frame classification module 31, a first stylization processing module 32, a second stylization processing module 33, and a stylized video generation module 34.
[0068] Among them, the video frame classification module 31 is configured to determine the key frames and non-key frames in the original video based on the feature similarity of adjacent video frames in the original video. The first stylization processing module 32 is configured to stylize each keyframe to obtain a first style video frame. The second stylization processing module 33 is configured to map the style pixel blocks in the first style video frame to the second style video frame of the non-key frame according to the motion vector between the non-key frame and the key frame, so as to construct the second style video frame of the non-key frame. The stylized video generation module 34 is configured to sort and summarize the first style video frames and the second style video frames based on their corresponding timestamps to obtain a stylized video of the original video.
[0069] Based on the above embodiments, the video frame classification module 31 includes: a similarity determination unit, configured to traverse each video frame in the original video according to the chronological order of each video frame, determine the first traversed video frame and the last traversed video frame as key frames, and determine the feature similarity between each traversed video frame and the next video frame; a first key frame determination unit, configured to determine the video frame and the corresponding next video frame as key frames if the feature similarity between the video frame and the next video frame is less than a preset similarity threshold; a second key frame determination unit, configured to determine the video frame as a key frame if the feature similarity between the video frame and the next video frame is greater than or equal to a preset similarity threshold and there is a preset number of frames between the video frame and the previous key frame; and a non-key frame determination unit, configured to determine the video frame as a non-key frame if the feature similarity between the video frame and the next video frame is greater than or equal to a preset similarity threshold but there is no preset number of frames between the video frame and the previous key frame.
[0070] Based on the above embodiments, the first stylization processing module 32 includes: a feature extraction unit configured to extract the content features of the key frame and the style features of the preset style image through a feature extraction network; a feature fusion unit configured to fuse the style features and content features through a style fusion network to obtain a fused feature map; and a first stylization processing unit configured to reconstruct the first style video frame of the key frame based on the fused feature map through an image reconstruction network.
[0071] Based on the above embodiments, the second stylization processing module 33 includes: a motion vector determination unit, configured to acquire non-key frames between two adjacent key frames and determine bidirectional motion vectors between the non-key frames and adjacent key frames; and a style video frame construction unit, configured to map each style pixel block in the first style video frame of adjacent key frames to the second style video frame of the non-key frames according to the bidirectional motion vectors between the non-key frames and adjacent key frames, thereby constructing the second style video frame of the non-key frames.
[0072] Based on the above embodiments, the style video frame construction unit includes: a first candidate style frame construction subunit, configured to map the style pixel blocks of the first style video frame of the preceding key frame to the first candidate style frame of the non-key frame according to the forward motion vector of the non-key frame and the preceding key frame, thereby constructing the first candidate style frame of the non-key frame; a second candidate style frame construction subunit, configured to map the style pixel blocks of the first style video frame of the following key frame to the second candidate style frame of the non-key frame according to the reverse motion vector of the non-key frame and the following key frame, thereby constructing the second candidate style frame of the non-key frame; a fusion weight determination subunit, configured to determine the fusion weight according to the number of frames between the non-key frame and the preceding and following key frames; and a candidate style frame fusion subunit, configured to perform weighted fusion of the first candidate style frame and the second candidate style frame according to the fusion weight, thereby obtaining the second style video frame of the non-key frame.
[0073] Based on the above embodiments, the candidate style frame fusion subunit is specifically configured as follows: determining common pixel positions and independent pixel positions in the first candidate style frame and the second candidate style frame; at the common pixel position, both the first candidate style frame and the second candidate style frame have pixel values, and at the independent pixel position, only the first candidate style frame or the second candidate style frame has pixel values; according to the fusion weight, the pixel values of the first candidate style frame and the second candidate style frame at the common pixel position are weighted and summed to obtain the fused pixel value, and the fused pixel value is mapped to the corresponding pixel position of the second style video frame; the pixel values at the independent pixel positions in the first candidate style frame or the second candidate style frame are mapped to the corresponding pixel positions of the second style video frame.
[0074] Based on the above embodiments, the style video frame construction unit further includes: a hole region completion subunit, configured to, after mapping the pixel values at independent pixel positions in the first candidate style frame or the second candidate style frame to the corresponding pixel positions in the second style video frame, if the second style video frame has a hole region, determine the pixel gradient value of the hole region based on the pixel region in the non-key frame corresponding to the hole region; and based on the pixel gradient value of the hole region, diffuse the pixel values of the edge pixels of the hole region in the second style video frame from the outer edge of the hole region to the interior using a diffusion algorithm.
[0075] The video stylization processing apparatus provided in this application, as described above, determines keyframes and non-keyframes in the original video by using the feature similarity of adjacent video frames; it then performs stylization processing on each keyframe to obtain a first-style video frame; based on the motion vectors between non-keyframes and keyframes, it maps the style pixel blocks in the first-style video frame to a second-style video frame of the non-keyframes, constructing a second-style video frame of the non-keyframes; finally, it sorts and summarizes the first-style video frame and the second-style video frame based on their corresponding timestamps to obtain the stylized video of the original video. Through the above technical means, the style pixel blocks after stylization of keyframes can be mapped to non-keyframes using the motion vectors between keyframes and non-keyframes, thereby achieving style transfer from stylized keyframes to non-keyframes and ensuring the processing accuracy of video stylization. The computational cost of motion vector and pixel block mapping is far lower than the computational cost of deep learning models converting video frame styles. Deep learning models effectively reduce the number of frames processed by only stylizing keyframes, thereby reducing the computational cost and time of deep learning models and improving the efficiency of video stylization processing.
[0076] The video stylization processing apparatus provided in this application embodiment can be used to execute the video stylization processing method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0077] Figure 10 This is a schematic diagram of the structure of a video stylization processing device provided in an embodiment of this application, with reference to... Figure 10 The video stylization processing device includes a processor 41, a memory 42, a communication device 43, an input device 44, and an output device 45. The number of processors 41 and the number of memories 42 in the video stylization processing device can be one or more. The processor 41, memory 42, communication device 43, input device 44, and output device 45 of the video stylization processing device can be connected via a bus or other means.
[0078] The memory 42, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the video stylization processing methods of the video stylization code storage module 21, code acquisition module 22, and code sharing module 24 in any embodiment of this application (e.g., the video frame classification module 31, the first stylization processing module 32, the second stylization processing module 33, and the stylized video generation module 34 in the video stylization processing apparatus). The memory 42 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 42 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0079] The communication device 43 is used for data transmission.
[0080] The processor 41 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 42, thereby realizing the video stylization processing method described above.
[0081] Input device 44 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 45 may include display devices such as a display screen.
[0082] The video stylization processing device provided above can be used to execute the video stylization processing method provided in the above embodiments, and has corresponding functions and beneficial effects.
[0083] This application also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a video stylization processing method. The video stylization processing method includes: determining keyframes and non-keyframes in the original video based on the feature similarity of adjacent video frames in the original video; performing stylization processing on each keyframe to obtain a first style video frame; mapping style pixel blocks in the first style video frame to a second style video frame of the non-keyframe based on the motion vector between the non-keyframe and the keyframe to construct a second style video frame of the non-keyframe; and sorting and summarizing the first style video frame and the second style video frame based on their corresponding timestamps to obtain a stylized video of the original video.
[0084] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROM, floppy disk, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements, etc. Storage medium may also include other types of memory or combinations thereof. Furthermore, storage medium may reside in a first computer system in which the program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage medium may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.
[0085] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the video stylization processing method described above, but can also execute related operations in the video stylization processing method provided in any embodiment of this application.
[0086] The video stylization processing apparatus, storage medium, and video stylization processing device provided in the above embodiments can execute the video stylization processing method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the video stylization processing method provided in any embodiment of this application.
[0087] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application. The scope of this application is determined by the scope of the claims.
Claims
1. A video stylization processing method, characterized in that, include: Based on the feature similarity of adjacent video frames in the original video, the key frames and non-key frames in the original video are determined; Each of the keyframes is stylized to obtain a first-style video frame; Based on the motion vector between the non-keyframe and the keyframe, the style pixel blocks in the first style video frame are mapped to the second style video frame of the non-keyframe to construct the second style video frame of the non-keyframe. The first style video frames and the second style video frames are sorted and summarized based on their corresponding timestamps to obtain the stylized video of the original video.
2. The video stylization processing method according to claim 1, characterized in that, The step of determining keyframes and non-keyframes in the original video based on the feature similarity of adjacent video frames includes: Based on the chronological order of each video frame in the original video, each video frame is traversed, and the first and last video frames traversed are identified as keyframes. The feature similarity between each traversed video frame and the next video frame is also determined. If the feature similarity between the video frame and the next video frame is less than a preset similarity threshold, then the video frame and the corresponding next video frame are determined as keyframes. If the feature similarity between the video frame and the next video frame is greater than or equal to a preset similarity threshold and the video frame is separated from the previous key frame by a preset number of frames, then the video frame is determined as a key frame. If the feature similarity between the video frame and the next video frame is greater than or equal to a preset similarity threshold, but there is no preset number of frames between the video frame and the previous key frame, then the video frame is determined to be a non-key frame.
3. The video stylization processing method according to claim 1, characterized in that, The step of stylizing each keyframe to obtain a first-style video frame includes: The content features of the keyframes and the style features of the preset style image are extracted using a feature extraction network. The style features and content features are fused using a style fusion network to obtain a fused feature map; The first style video frame of the keyframe is reconstructed based on the fused feature map using an image reconstruction network.
4. The video stylization processing method according to claim 1, characterized in that, The step of mapping style pixel blocks in the first style video frame to the second style video frame of the non-key frame based on the motion vector between the non-key frame and the key frame, thereby constructing the second style video frame of the non-key frame, includes: Obtain the non-key frames between two adjacent key frames, and determine the bidirectional motion vectors between the non-key frames and the adjacent key frames before and after them. Based on the bidirectional motion vectors between the non-keyframe and the adjacent keyframes, each style pixel block in the first style video frame of the adjacent keyframes is mapped to the second style video frame of the non-keyframe, thus constructing the second style video frame of the non-keyframe.
5. The video stylization processing method according to claim 4, characterized in that, The step of mapping each style pixel block in the first style video frame of the adjacent key frames to the second style video frame of the non-key frame based on the bidirectional motion vectors of the non-key frame and the adjacent key frames to construct the second style video frame of the non-key frame includes: Based on the positive motion vectors of the non-key frame and the previous key frame, the style pixel block of the first style video frame of the previous key frame is mapped to the first candidate style frame of the non-key frame to construct the first candidate style frame of the non-key frame. Based on the inverse motion vectors of the non-keyframe and the adjacent keyframe, the style pixel blocks of the first style video frame of the adjacent keyframe are mapped to the second candidate style frame of the non-keyframe, thus constructing the second candidate style frame of the non-keyframe. The fusion weight is determined based on the number of frames between the non-key frame and the preceding and following key frames. The first candidate style frame and the second candidate style frame are weighted and fused according to the fusion weight to obtain the second style video frame of the non-key frame.
6. The video stylization processing method according to claim 5, characterized in that, The step of weighted fusing the first candidate style frame and the second candidate style frame according to the fusion weight to obtain the second style video frame of the non-key frame includes: In the first candidate style frame and the second candidate style frame, a common pixel position and an independent pixel position are determined; at the common pixel position, there is a pixel value in both the first candidate style frame and the second candidate style frame, and at the independent pixel position, there is a pixel value in only the first candidate style frame or the second candidate style frame. According to the fusion weight, the pixel values of the first candidate style frame and the second candidate style frame at the common pixel position are weighted and summed to obtain the fused pixel value, and the fused pixel value is mapped to the corresponding pixel position of the second style video frame; Map the pixel values at independent pixel positions in the first or second candidate style frame to the corresponding pixel positions in the second style video frame.
7. The video stylization processing method according to claim 6, characterized in that, After mapping the pixel values at independent pixel positions in the first candidate style frame or the second candidate style frame to the corresponding pixel positions in the second style video frame, the method further includes: If a hole region exists in the second style video frame, the pixel gradient value of the hole region is determined based on the pixel region in the non-key frame corresponding to the hole region. Based on the pixel gradient value of the hole region, a diffusion algorithm is used to diffuse the pixel values of the edge pixels of the hole region in the second style video frame from the outer edge of the hole region to the interior.
8. A video stylization processing device, characterized in that, include: The video frame classification module is configured to determine key frames and non-key frames in the original video based on the feature similarity of adjacent video frames in the original video. The first stylization processing module is configured to perform stylization processing on each of the keyframes to obtain a first style video frame. The second stylization processing module is configured to map the style pixel blocks in the first style video frame to the second style video frame of the non-key frame according to the motion vector between the non-key frame and the key frame, thereby constructing the second style video frame of the non-key frame. The stylized video generation module is configured to sort and summarize the first stylized video frame and the second stylized video frame based on their corresponding timestamps to obtain a stylized video of the original video.
9. A video stylization processing device, characterized in that, include: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the video stylization processing method as described in any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the video stylization processing method as described in any one of claims 1-7.