A Video Style Transfer Method Based on RAFT Optical Flow
By using the RAFT optical flow network to calculate the optical flow field and dynamic weight map in the video style migration technology, the video frames are decomposed into low-frequency and high-frequency layers for style migration, and motion alignment and timing smooth filtering are performed, which solves the problems of coherence and smooth transition between video frames, and achieves high-quality video style migration effect.
Patent Information
- Application Number
- CN202510480941.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing video style transfer technologies cannot effectively maintain the consistency and smooth transition between video frames, especially when dealing with complex backgrounds and fast motion scenes, local motion information cannot be accurately captured, resulting in poor style transfer results.
Using a video style migration method based on RAFT optical flow, the pixel displacement of adjacent video frames is calculated through the RAFT optical flow network, an optical flow field is generated, and a dynamic weight map is generated based on the optical flow field. Then, the video frame is decomposed into a low-frequency structure layer and a high-frequency detail layer, and global and local style transfers are performed, and motion alignment and timing smoothing filtering are performed through the optical flow field to generate a coherent stylized video.
Through accurate optical flow calculation and dynamic weight graph generation, it can effectively adapt to the motion of the video content, balance style presentation and detail retention, significantly improve the effect and quality of video style transfer, and avoid lag and incoherence of stylized videos.
Smart Images

Figure CN119991416B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly relates to a video style transfer method based on RAFT optical flow. Background Art
[0002] In the video style transfer task, the optical flow network plays a crucial role. By estimating the motion information between video frames, it ensures the consistency and smooth transition of style transfer in the time dimension. Traditional style transfer methods usually only process single-frame images, ignoring the dynamic changes between frames in the video, which may result in an incoherent or jumpy appearance of the style-transferred video. The optical flow network can accurately capture the motion patterns between each frame, maintain the relative motion of objects and backgrounds in the video, and thus avoid unnatural transitions or blurs during the style transfer process. Optical flow can not only handle the dynamic changes of fast motion and complex scenes, but also help solve problems caused by object occlusion or texture loss, ensuring that each frame after style transfer is consistent in motion and vision. In addition, the optical flow network can provide more accurate local motion estimation and global motion patterns, thus better adapting to the motion changes in multi-scale scenes and improving the effect and quality of style transfer. The optical flow network combined with deep learning automatically learns motion information through training, making the calculation efficiency higher and the estimation accuracy more accurate. This enables the application of optical flow in video style transfer to provide high-quality and stable transfer effects in long time series and complex scenes, and ultimately achieve a more natural and smooth video style conversion.
[0003] Traditional optical flow estimation methods, such as Lucas-Kanade and Horn-Schunck, rely on the assumption based on the relative motion of local pixels. These methods can provide basic optical flow estimation under some ideal conditions. However, in modern video style transfer tasks, these traditional methods face many challenges and perform poorly especially when dealing with complex and fast-changing scenes. First, traditional optical flow algorithms have low accuracy in high-dynamic-range videos or scenes with fast motion. This is because traditional methods usually cannot capture motion information at larger scales or finer granularities. Especially when there is a large amount of motion between frames, it is easy to cause optical flow estimation distortion. This will affect the quality of style transfer, resulting in unnatural motion blur or incorrect object positions in the generated video, thus destroying style consistency. Second, traditional optical flow methods are very sensitive to complex phenomena such as illumination changes, texture loss, occlusion, or reflection. In real-world scenes, video frames are often affected by environmental illumination, object reflection, and camera angle changes, which makes it difficult for traditional algorithms to accurately estimate optical flow. Especially in areas with less texture or low contrast, traditional optical flow methods often cannot provide effective estimates, resulting in large errors in optical flow in these areas, thus affecting the effect of video style transfer and making the transferred video style less natural or even showing obvious distortion. Summary of the Invention
[0004] The purpose of the present invention is to provide a video style transfer method based on RAFT optical flow to solve the following technical problems:
[0005] Existing video style transfer technologies cannot effectively maintain the coherence and smooth transition between video frames. When dealing with complex backgrounds and details, they cannot accurately capture local motion information and perform poorly when dealing with fast-moving scenes, affecting the overall transfer effect.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A video style transfer method based on RAFT optical flow includes the following steps:
[0008] Calculate the pixel displacement of adjacent frames in the video frame by frame through the RAFT optical flow network to generate an optical flow field containing the motion direction and magnitude of each pixel;
[0009] Generate a dynamic weight map according to the displacement amount of pixels in the optical flow field, and the weight value of each region is inversely proportional to the displacement amount;
[0010] Decompose the video frame into a low-frequency structure layer and a high-frequency detail layer through filtering, perform global style transfer on the low-frequency structure layer, and replace the color distribution and stroke features; perform local stylization on the high-frequency detail layer in combination with the dynamic weight map;
[0011] Fuse the processed low-frequency structure layer and high-frequency detail layer to generate a single-frame stylized image; align the stylized result of the current frame with adjacent frames according to the optical flow field, and adjust the pixel positions through back-projection and interpolation compensation;
[0012] Perform temporal smoothing filtering on the aligned multi-frame results to output a coherent stylized video.
[0013] As a further solution of the present invention: the generation process of the dynamic weight map includes:
[0014] Set a dynamic decision threshold based on the overall displacement distribution of the optical flow field, and mark the area exceeding the dynamic decision threshold as a high-dynamic area; perform non-linear attenuation on the weight values in the high-dynamic area, and the attenuation degree gradually increases with the increase of the displacement amount;
[0015] Perform morphological expansion operations on the boundary area of the weight map to smooth the transition boundary between the low-weight area and the high-weight area; the morphological expansion operation expands the coverage range of the low-weight area through the dilation algorithm to avoid weight jumps at the motion boundary;
[0016] Adopt an adaptive Gaussian kernel to smooth the weight map, and the radius of the Gaussian kernel is dynamically adjusted according to the severity of the local displacement change to ensure a natural transition of the weight distribution; the calculation method of the radius of the adaptive Gaussian kernel is: according to the displacement standard deviation in the local area of the optical flow field, the larger the standard deviation, the larger the radius of the Gaussian kernel, and the smaller the standard deviation, the smaller the radius.
[0017] As a further solution of the present invention: the decomposition process of the low-frequency structure layer and the high-frequency detail layer is specifically as follows:
[0018] Blur the original frame through a Gaussian filter to extract the low-frequency structure layer; subtract the low-frequency structure layer from the original frame to obtain the high-frequency detail layer, and truncate and compensate the negative value area of the high-frequency detail layer; the truncation and compensation operation sets the negative values of the high-frequency detail layer to zero and then superimposes a fixed proportion of the intensity of the original details;
[0019] Perform contrast stretching on the high-frequency detail layer. The contrast stretching expands the pixel value distribution of the high-frequency detail layer to a preset range through linear mapping to improve the editability of the texture, and the stretching range is adaptively adjusted according to the pixel distribution of the detail layer; retain the original geometric structure in the global style transfer of the low-frequency structure layer, and only replace the color and stroke features.
[0020] As a further solution of the present invention: the local stylization of the high-frequency detail layer includes:
[0021] Based on a preset weight threshold, divide the regions in the video frame into low-weight regions, medium-weight regions, and high-weight regions. For the high-frequency details in the low-weight regions, limit the stylization operation to the luminance channel and keep the original values of the chrominance channel; for the high-frequency details in the medium-weight regions, adopt a multi-scale fusion strategy and overlay coarse-grained strokes and fine-grained textures respectively; for the high-frequency details in the high-weight regions, fully apply the style texture and maintain the sharpness of the contour through an edge-preserving algorithm.
[0022] The multi-scale fusion strategy stacks the style textures layer by layer from coarse to fine by constructing style feature pyramids of different scales.
[0023] The edge-preserving algorithm detects the gradient information of the high-frequency details layer and retains the original contours of the regions with gradients higher than the set value during the stylization process.
[0024] As a further solution of the present invention: the fusion process of the low-frequency structure layer and the high-frequency details layer includes:
[0025] Dynamically adjust the feature fusion ratio between the low-frequency structure layer and the high-frequency details layer according to the motion information of the optical flow field. For the regions in the optical flow field where the displacement amount exceeds the dynamic determination threshold, increase the weight ratio of the high-frequency details layer in the fusion to the set ratio; for the regions where the displacement amount is lower than the dynamic determination threshold, increase the fusion weight ratio of the low-frequency structure layer to the set ratio.
[0026] Construct a cross-layer feature correlation graph, calculate the motion correlation between the low-frequency structure layer and the high-frequency details layer of adjacent frames through the optical flow field, and use the same fusion ratio for the regions with high correlation in the correlation graph.
[0027] Perform cross-frame consistency verification on the fused result, project the fusion ratio of the current frame along the optical flow field to adjacent frames, and ensure the transition of high and low frequency features between multiple frames.
[0028] As a further solution of the present invention: the motion alignment specifically includes:
[0029] Perform two-way consistency verification on the optical flow field. By calculating the cyclic error between the forward optical flow and the backward optical flow, mark the regions with errors exceeding the preset threshold as unreliable regions, and the rest as reliable regions; for the stylized features in the unreliable regions, fill them with weighted fusion of adjacent frames, and the weights are determined by the local smoothness of the optical flow field; perform sub-pixel alignment on the stylized features in the reliable regions, and adjust the interpolation weights through the fractional displacement values of the optical flow field.
[0030] The sub-pixel alignment adjusts the feature interpolation weights according to the fractional displacement values of the optical flow field through the bilinear interpolation algorithm to make the projection result match the pixel grid of the target frame.
[0031] As a further solution of the present invention: it also includes static region optimization:
[0032] Detect the areas in the video where the displacement amounts of consecutive multiple frames are lower than the dynamic determination threshold, and mark them as static areas; otherwise, mark them as dynamic areas. Perform high-precision style transfer on the first frame of the static area, reuse the result of the first frame for subsequent frames, and perform fine-tuning through the optical flow field; the fine-tuning operation adjusts the pixel positions of the reused area according to the displacement amount of the optical flow field to compensate for camera jitter or illumination changes.
[0033] Perform gradual blending at the junction of the static area and the dynamic area, and transition the stylization intensity through transparency superposition; the gradual blending forms a change in stylization intensity at the junction through the linear transition of the alpha channel.
[0034] As a further solution of the present invention: it further includes identifying key target areas in the video through a pre-trained model; setting a lower weight limit for the moving part of the key target area to limit the stylization intensity; setting an upper weight limit for the static part of the key target area to retain local details.
[0035] The pre-trained model adopts a semantic segmentation algorithm based on a convolutional neural network to output pixel-level target category labels; the setting of the lower weight limit and the upper weight limit is dynamically adjusted according to the importance of the target category, and the importance of the target category is obtained based on a preset target mapping table.
[0036] The beneficial effects of the present invention:
[0037] The present invention accurately calculates the pixel displacement between adjacent frames of the video through the RAFT optical flow network to generate an optical flow field. The dynamically generated weight map can flexibly adapt to the motion of the video content, overcoming the problem of lack of flexibility in weight setting in the prior art. In terms of video frame decomposition, the video frame is accurately decomposed into a low-frequency structure layer and a high-frequency detail layer by using Gaussian filtering and other methods, which can effectively balance the overall style presentation and detail retention, and solve the problem of inaccurate division in the prior art. In terms of stylization operations, global style transfer is performed on the low-frequency structure layer, and local style transfer is performed on the high-frequency detail layer in combination with the dynamic weight map, and different strategies are adopted for different weight areas. At the same time, motion alignment and temporal smoothing filtering are performed using the optical flow field, effectively avoiding stuttering and incoherence in the stylized video, significantly improving the viewing experience, and further improving the video style transfer effect through static area optimization, key target area processing, etc. Brief Description of the Drawings
[0038] The following further describes the present invention with reference to the drawings.
[0039] Figure 1 is a schematic flowchart of the present invention;
[0040] Figure 2 is a schematic structural diagram of the RAFT module of the present invention;
[0041] Figure 3 It is a schematic structural diagram of the style transfer network of the present invention;
[0042] Figure 4 It is the collaborative workflow of the style transfer network of the present invention;
[0043] Figure 5 It is the RGB histogram of the present invention;
[0044] Figure 6 It is the RGB histogram of the comparison method AdaIn;
[0045] Figure 7 It is the RGB histogram of the comparison method NNST;
[0046] Figure 8 It is the RGB histogram of the comparison method SANET;
[0047] Figure 9 It is a schematic structural diagram of the video style transfer based on RAFT optical flow of the present invention. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0049] Embodiment 1, please refer to Figure 1 As shown, a video style transfer method based on RAFT optical flow provided by the present invention includes the following steps:
[0050] 1. As an advanced and efficient optical flow calculation model, the RAFT optical flow network is based on a deep learning architecture and is trained using a large amount of video frame data, so as to accurately capture the pixel changes between adjacent frames. During the calculation process, the network carefully analyzes each frame of the video and outputs an optical flow field containing the movement direction and magnitude of each pixel through complex neural network operations. This optical flow field is like a detailed "movement map", providing a solid data basis for the subsequent dynamic analysis and processing of video content, enabling the present invention to clearly understand the movement trajectories of each pixel in the video between frames and providing a key basis for the dynamic adjustment of style transfer.
[0051] 2. Generate a dynamic weight map based on the displacement of pixels in the optical flow field. This process has rigorous and scientific steps. First, set a dynamic determination threshold based on the overall displacement distribution of the optical flow field. This determination threshold is not fixed, but is obtained through statistical analysis of the displacement of all pixels in the optical flow field. For example, calculate statistical quantities such as the mean and median of the displacement, and combine the characteristics of the video content (such as scene complexity, the proportion of moving objects, etc.) to determine a reasonable dynamic determination threshold. Areas exceeding this threshold are keenly marked as high-dynamic areas, which means that the pixels in these areas move more violently between frames. For the weight values of high-dynamic areas, adopt a non-linear attenuation strategy, and the attenuation degree gradually increases with the increase of the displacement. The logic behind this is that in areas with more violent motion, more features of the original video may need to be retained during the style transfer process to avoid content distortion caused by excessive style interference. Therefore, the weight values of areas with large displacements decrease, and this decrease is not linear, but decays at a faster rate as the displacement increases, so as to accurately control the intensity of stylization.
[0052] To make the weight map more reasonable and smooth, perform a morphological expansion operation on the boundary area of the weight map. Specifically, the dilation algorithm is used here to expand the coverage of the low-weight area. The dilation algorithm is a commonly used morphological operation in image processing. It expands the boundary pixels of the low-weight area outward, making the transition boundary between the low-weight area and the high-weight area smoother and more natural, and avoiding weight jumps at the motion boundary. Such jumps may cause obvious discontinuities in the stylization effect at the boundary, seriously affecting the visual effect. Through a carefully designed dilation algorithm, this potential flaw can be cleverly eliminated, making the transition of the weight map between different areas more smooth.
[0053] Finally, use an adaptive Gaussian kernel to smooth the weight map to further ensure the natural transition of the weight distribution. The radius of the adaptive Gaussian kernel is not a fixed value, but is dynamically adjusted according to the severity of the local displacement change. Its specific calculation method is: deeply analyze the displacement standard deviation of the local area in the optical flow field. The standard deviation is an important indicator to measure the degree of data dispersion. In the optical flow field, the larger the displacement standard deviation of the local area, the more violent the change in pixel displacement in this area. At this time, a larger Gaussian kernel radius is required for smoothing to better balance the weight distribution; on the contrary, the smaller the standard deviation, the relatively stable the pixel displacement change in this area, and the Gaussian kernel radius is correspondingly reduced. By dynamically adjusting the Gaussian kernel radius in this way, the weight map can be smoothly operated flexibly and accurately according to the actual motion situation of the video content, so that the finally generated dynamic weight map can highly fit the dynamic characteristics of the video.
[0054] 3. Reasonably decompose the video frames, and use a Gaussian filter to blur the original frames to extract the low-frequency structure layer. The Gaussian filter is essentially a linear smoothing filter, and its principle is based on the Gaussian function. It achieves the blurring effect by weighted averaging each pixel in the original frame and its neighboring pixels. Since low-frequency signals change slowly and appear as large similar areas in space, the Gaussian filter can effectively retain the information of these areas, while the fast-changing details corresponding to high-frequency signals (such as edges, textures, etc.) are smoothed out during the blurring process, thus successfully extracting the low-frequency structure layer that mainly contains information such as the outline of the scene main body, large-area color, and brightness change trends.
[0055] Subtract the original frame from the extracted low-frequency structure layer to obtain the high-frequency detail layer. However, negative value areas will appear in the high-frequency detail layer during this process. To solve this problem, truncation and compensation operations are required. Specifically, directly set the negative values in the high-frequency detail layer to zero to eliminate unreasonable negative pixel values. Subsequently, superimpose a fixed proportion of the intensity of the original details. This fixed proportion is not set arbitrarily but is determined based on a large number of experiments and the analysis of the video content, aiming to ensure that while removing negative values, as much original detail information as possible is retained, maintaining the integrity and accuracy of the high-frequency detail layer, and enabling subsequent processing to be based on more reliable detail data.
[0056] To further improve the editability of the texture of the high-frequency detail layer, contrast stretching needs to be performed on it. This operation uses the principle of linear mapping to expand the pixel value distribution of the high-frequency detail layer to a preset range. The preset range is not fixed but is adaptively adjusted according to the pixel distribution of the detail layer. For example, first count the minimum and maximum values of the pixel values in the high-frequency detail layer, and combine the characteristics of the video content (such as whether it is a high-contrast scene, etc.) to calculate reasonable mapping parameters, so that the pixel values can more clearly display texture details in the new range, providing a more distinct and easier-to-process texture basis for subsequent stylization operations. When performing global style transfer on the low-frequency structure layer, carefully retain the original geometric structure and only replace the color and brushstroke features. Through a specific style transfer algorithm (such as a neural network-based style transfer model), integrate the color distribution and brushstroke features of the target style into the low-frequency structure layer to ensure that while changing the style, the basic geometric shape and layout of the video scene are not damaged, maintaining the recognizability of the scene.
[0057] The local stylization operation of the high-frequency detail layer is meticulous and exquisite. Based on the preset weight threshold, the regions in the video frames are accurately divided into low-weight regions, medium-weight regions, and high-weight regions. For the high-frequency details in the low-weight regions, considering the preservation of the original color information, the stylization operation is strictly restricted to the luminance channel, maintaining the original values of the chrominance channel. In this way, while introducing stylistic changes to a certain extent, it ensures that the color information is as consistent as possible with the original video, avoiding the negative impact of color distortion on the visual effect. For the high-frequency details in the medium-weight regions, a multi-scale fusion strategy is adopted. Specifically, it is achieved by constructing style feature pyramids of different scales. Starting from the coarse-grained style features at the top layer of the pyramid, the fine-grained textures are stacked layer by layer downward. In this process, the coarse-grained strokes at the top layer can quickly outline the overall style contour, and the fine-grained textures at the bottom layer add rich details to the style. The organic combination of the two makes the stylization effect more natural and delicate. For the high-frequency details in the high-weight regions, the style texture is fully applied, and at the same time, an edge-preserving algorithm is used to maintain the sharpness of the contour. The edge-preserving algorithm works by detecting the gradient information of the high-frequency detail layer. A gradient threshold is preset. During the stylization process, for the regions where the gradient is higher than the set value, the original contour is retained. Because these high-gradient regions often correspond to important details such as the edges of objects, retaining the original contour can ensure that the stylized video presents the new style while not losing the detail clarity, improving the visual quality.
[0058] During the entire video processing process, static region optimization is also an indispensable part. By detecting the regions in the video where the displacement amount of consecutive multiple frames is lower than the dynamic determination threshold, they are keenly marked as static regions, and vice versa for dynamic regions. For static regions, high-precision style transfer is performed on their first frame, which means using a more complex and refined style transfer algorithm and investing more computing resources to ensure high-quality stylized results. The subsequent frames reuse the results of the first frame and are fine-tuned through the optical flow field. The fine-tuning operation adjusts the pixel positions of the reused regions according to the displacement amount of the optical flow field to compensate for camera jitter or light changes. Since static regions are not absolutely stationary, there may be pixel position changes caused by slight camera shaking or slow light changes. Through the accurate displacement information of the optical flow field, the pixel positions can be precisely adjusted to maintain the stability and coherence of the video. At the junction of static regions and dynamic regions, gradient blending is performed to achieve a natural transition of the stylization intensity. Specifically, it is achieved through the linear transition of the alpha channel, and the alpha channel is used to control the transparency of the image. At the junction, the alpha value is linearly adjusted according to the distance from the static region and the dynamic region, so that the stylization intensity smoothly transitions from the stylization degree of the static region to the stylization degree of the dynamic region, avoiding obvious style mutations and improving the overall visual fluency of the video.
[0059] In addition, key target regions in the video are identified through a pre-trained model. Here, a semantic segmentation algorithm based on a convolutional neural network is used to build the pre-trained model. This algorithm utilizes components such as convolutional layers, pooling layers, and fully connected layers to learn from a large amount of annotated video data, enabling it to output pixel-level target class labels and accurately identify various objects and regions in the video. For the identified key target regions, their moving parts and static parts are further distinguished. A lower weight limit is set for the moving parts to restrict the stylization intensity and prevent the loss of details or abnormal visual effects in the moving parts of the key targets due to excessive style changes. An upper weight limit is set for the static parts to retain local details and ensure that the static parts of the key targets can retain their original detailed features during the style transfer process, maintaining the recognizability and integrity of the targets. The setting of the lower and upper weight limits is not a fixed value but is dynamically adjusted according to the importance of the target category. The importance of the target category is obtained based on a preset target mapping table, which has pre-classified various targets (such as people, vehicles, buildings, etc.) and assigned weights according to their importance in the video content, enabling flexible and reasonable adjustment of the stylization weights of key target regions according to the characteristics of the video content in practical applications.
[0060] 4. Fusing the processed low-frequency structure layer and high-frequency detail layer to generate a single-frame stylized image is an important stage towards the final stylized video. To achieve a high-quality fusion effect, the feature fusion ratio between the low-frequency structure layer and the high-frequency detail layer needs to be dynamically regulated based on the motion information of the optical flow field. The dynamic decision threshold plays a crucial role here, which is determined based on the overall distribution of displacement amounts in the optical flow field. For regions in the optical flow field where the displacement amount exceeds the dynamic decision threshold, it means that the pixels in this region move relatively violently and the scene changes richly. At this time, to better present the details and style of the dynamic regions, the weight ratio of the high-frequency detail layer in the fusion is increased to a set ratio. This set ratio is not arbitrarily specified but is obtained through a large number of experiments and in-depth analysis of the video content, aiming to highlight the expressiveness of high-frequency details in rapidly changing regions, so that the stylized image can reflect the texture details of the new style in these regions and closely fit the dynamic characteristics of the original video. On the contrary, for regions where the displacement amount is lower than the dynamic decision threshold, the region is relatively stable and the scene changes less. Therefore, the fusion weight ratio of the low-frequency structure layer is increased to a set ratio to emphasize the overall structure of the scene and the consistency of the style, maintaining the visual coherence of the stable regions.
[0061] Constructing a cross-layer feature correlation graph is a key strategy to optimize the fusion effect. Through the powerful motion analysis ability of the optical flow field, the motion correlation between the low-frequency structure layer and the high-frequency detail layer of adjacent frames is calculated. The optical flow field can accurately capture the motion trajectories of pixels in each frame. Based on this, the similarity and correlation between different layers of adjacent frames can be quantified. In the correlation graph, for regions with high correlation, that is, those regions with relatively similar motion patterns between adjacent frames, the same fusion ratio is adopted. This ensures the coherence of the video in the time dimension and avoids visual jumps caused by large fluctuations in the fusion ratio between different frames, making the stylized video present a natural and smooth visual transition during motion.
[0062] To further ensure the quality of the stylized video, it is crucial to perform cross-frame consistency verification on the fused result. In this process, the fusion ratio of the current frame is projected onto adjacent frames along the optical flow field. The optical flow field acts as a link connecting the frames of the video. With its accurate motion information, the fusion ratio information of the current frame is transmitted to adjacent frames. In this way, the transition of high- and low-frequency features between multiple frames is smoothly achieved, avoiding the abruptness caused by inconsistent fusion ratios and greatly enhancing the visual smoothness and coherence of the stylized video, bringing a more comfortable viewing experience to the audience.
[0063] After the fusion operation, the stylized result of the current frame is motion-aligned with adjacent frames according to the optical flow field. This is a key step to ensure the natural dynamic effect of the stylized video. First, a two-way consistency check is performed on the optical flow field. The reliability of the optical flow field is evaluated by calculating the cyclic error between the forward optical flow and the backward optical flow. The forward optical flow describes the pixel motion from the current frame to the next frame, and the backward optical flow is vice versa. If the cyclic error between the two exceeds the preset threshold, it means that there may be a deviation in the optical flow calculation in this region, and this region is marked as an unreliable region; otherwise, it is a reliable region. For the stylized features in the unreliable region, weighted fusion filling of adjacent frames is used for optimization. The local smoothness of the optical flow field serves as a weight determinant here. The higher the local smoothness, the smoother the motion change in this region, and the lower the dependence on adjacent frames during weighted fusion; otherwise, it is higher. This weighted fusion method based on local smoothness can effectively compensate for the defects of optical flow calculation in the unreliable region and improve the stability of the stylized effect. For the stylized features in the reliable region, sub-pixel alignment is performed to achieve higher-precision motion matching. Sub-pixel alignment uses the bilinear interpolation algorithm to adjust the feature interpolation weights according to the fractional displacement value of the optical flow field. The bilinear interpolation algorithm determines the pixel value at the interpolation position by weighted averaging the four adjacent pixels in the target frame, and the fractional displacement value of the optical flow field precisely indicates the adjustment direction and degree of the interpolation weights, making the projection result accurately match the pixel grid of the target frame and ensuring that there are no pixel misalignments or blurs in the stylized video during motion, presenting a clear and smooth dynamic visual effect.
[0064] 5. The aligned multi-frame results are subjected to temporal smoothing filtering, which can eliminate the noise, fluctuations and discontinuities introduced by the previous steps, ensuring that the video presents a natural and smooth visual effect in the time dimension. Its principle is based on the time series analysis of multi-frame video data, and uses a specific filtering algorithm to adjust and optimize the pixel value or feature information of each frame. Common algorithms include mean filtering, median filtering and time-domain Gaussian filtering, each with its own characteristics and applicable scenarios.
[0065] Mean filtering calculates the average value of the same pixel position in several adjacent frames, replaces the original value of the corresponding pixel in the current frame, weakens the interference of single-frame data fluctuations, and makes the video picture transition smoother. Median filtering selects the median of the corresponding pixel set of adjacent frames as the pixel output value of the current frame, which can effectively remove impulse noise and avoid obvious defects in the video picture. Time-domain Gaussian filtering assigns different weights to the pixel values of adjacent frames based on the characteristics of the Gaussian function, and pays more attention to the reference of adjacent frame information, which can smooth noise fluctuations and retain dynamic details. In practical applications, it is necessary to select appropriate filtering algorithms and parameters based on the video frame rate, content complexity, degree of stylization and expected effect.
[0066] For example 2, please refer to Figure 2 The figure is a schematic diagram of the structure of a RAFT module provided by the present invention.
[0067] The RAFT (Recurrent All-Pairs Field Transforms) module is a deep learning model for optical flow estimation. Its core advantage is that it can efficiently and accurately estimate global optical flow. Compared with traditional optical flow methods, RAFT performs full-to-full pixel-level matching in the entire image and performs iterative optimization to capture more detailed and accurate optical flow information.
[0068] FeatureEncoder extracts feature maps of two adjacent frames, while ContextEncoder only extracts features from the first frame. Both are CNN-based networks and can be understood as shallow custom ResNets. The circle L represents the Look-up operation, and the series of boxes and arrows in the middle represent iterative optical flow estimation using GRU (a recurrent network).
[0069] The optical flow estimation process of the RAFT model starts with initializing the optical flow result to zero, that is, assuming that the initial optical flow field is 。Then, the ContextEncoder is used to extract global context information from the input first-frame image, and the FeatureEncoder is used to extract the features of the first and second frames. Then, through matrix multiplication operations, 4D correlation volumes (4DCorrelationVolumes) are calculated, which capture the matching information between each pair of pixels in the image. Based on these features and correlation information, the model uses the GRU (Gated Recurrent Unit) module to make a preliminary estimate of the optical flow, obtain the optical flow update amount, and update the optical flow field to , that is, by adding the initial zero optical flow, the optical flow estimate for the first step is obtained.
[0070] However, this preliminary estimate is usually not accurate enough. Therefore, RAFT further improves the accuracy of the optical flow estimate through iterative optimization. Specifically, the model looks up the correlation information in the 4D correlation volumes, takes the updated optical flow result as input and passes it to the GRU, and calculates the optical flow update amount again . With each iteration, the optical flow field is continuously updated to obtain a new optical flow estimate . This recursive process not only corrects the preliminary estimation result through repeated iterations, but also gradually optimizes the optical flow estimate, ultimately achieving an accurate optical flow field. Through multiple iterations, RAFT can gradually improve the optical flow estimate based on the more accurate motion information obtained in each iteration, greatly improving the accuracy of the final optical flow estimate in spatio-temporal motion.
[0071] 4DCorrelationVolumes are 4D voxels obtained by calculating the correlation of the feature maps of two adjacent frames pixel by pixel, with a size of , and the calculation method can be understood as transforming the feature map of the first frame into a matrix, transforming the feature map of the second frame into a matrix, and then performing matrix multiplication on the two to obtain , and adjusting the shape (reshape) to obtain the final result, which is expressed by the formula as follows:
[0072] ;
[0073] ;
[0074] where each element of 4DCorrelationVolumes can represent the correlation between the (i, j)th pixel of the first frame and the (k, l)th pixel of the second frame.
[0075] Therefore, RAFT can be divided into three stages:
[0076] Feature extraction: The network input consists of two consecutive frames. To extract features from these two images, the network uses two CNNs with shared weights. The architecture of the CNN consists of 6 residual layers, just like the layers of ResNet. Every other layer reduces the resolution by half while increasing the number of channels.
[0077] Visual Similarity: Visual similarity is calculated as the inner product of all pairs of feature maps. Thus, a four-dimensional tensor called the correlation volume is obtained, which provides key information about pixel displacements of different magnitudes. Then, the last two dimensions of this four-dimensional tensor are pooled with kernels of sizes 1, 2, 4, and 8 to construct a 4-layer correlation pyramid.
[0078] Iterative update: The iterative update is a sequence of gated recurrent units (GRUs) that combines all the data calculated previously in the present invention. The GRU units simulate an iterative optimization algorithm, but with an improvement - there are trainable convolutional layers with shared weights. Each update iteration produces a new optical flow update to make the prediction more accurate at each new step.
[0079] In the present invention, the RAFT module is applied to the optical flow estimation step of video frames. By using RAFT to estimate the optical flow of adjacent video frames, the motion information of each pixel is extracted as input data, accurately capturing the motion changes between video frames, searching for the best motion feature signals, thereby constraining the temporal consistency and spatial coherence of video frames. This precise motion estimation helps to optimize the frame-to-frame transitions during video style transfer, reducing unnatural effects caused by motion distortion or misalignment, and enhancing the coherence and stability of the video.
[0080] Then, the optical flow estimation result output by the RAFT module is optimized by the optical flow smoothing module to reduce noise and inconsistencies, ensuring the smoothness of the optical flow field. This module is implemented by minimizing a comprehensive energy function that combines spatial and temporal smoothing constraints.
[0081] In the video style transfer network of the present invention, the core role of the optical flow smoothing module is to optimize the optical flow estimation result and eliminate the problems caused by noise and inconsistencies in motion estimation. Specifically, the optical flow smoothing module not only focuses on the spatial smoothing of the optical flow field, but also combines temporal smoothing to ensure the coherence of the motion information between frames. The optical flow smoothing module mainly consists of spatial smoothing and temporal smoothing. Spatial smoothing mainly ensures that the optical flow field is smooth in the spatial dimension, that is, the motion information of adjacent pixels does not show drastic fluctuations. By constraining the optical flow vectors of adjacent pixels, discontinuities caused by feature noise or local estimation errors are avoided. Temporal smoothing ensures that the optical flow changes between adjacent video frames are continuous. That is to say, in the temporal dimension, the motion of an object does not exhibit jumps or unnatural changes. It makes the change of the optical flow field smoother by constraining the optical flow difference between adjacent frames.
[0082] In the network of the present invention, the optical flow smoothing module optimizes the optical flow estimation by minimizing a comprehensive energy function. The specific form of the objective function combines the constraints of spatial smoothing and temporal smoothing, and the general expression is:
[0083] ;
[0084] where is the optical flow vector of pixel point p in the optical flow field. is the spatial gradient of the optical flow field, representing the constraint of spatial smoothing. and are hyperparameters used to balance the weights of spatial smoothing and temporal smoothing. d is the time difference, representing the displacement between adjacent frames, which is used to constrain the optical flow smoothing in the temporal dimension. Ω is the set of all pixels in the image. The gradient of the optical flow field in the spatial dimension is calculated to ensure that the difference between the optical flow estimation values of adjacent pixels is as small as possible. This term makes the change of the optical flow smooth in space and avoids drastic changes caused by noise or estimation errors. Forces the optical flow changes between adjacent frames to be consistent. By constraining the optical flow difference between the current frame and adjacent frames, this term ensures that the optical flow field is smooth in the temporal dimension and avoids the incoherent motion of the object on the time axis.
[0085] In the video style transfer network of the present invention, the optical flow smoothing module optimizes the optical flow estimation result through spatial and temporal smoothing constraints, ensuring the stability and consistency of the optical flow field. This module can effectively eliminate the noise in the optical flow estimation, improve the coherence and smoothness of the video, so that the motion performance of the object in the style transfer process is more natural, avoiding inconsistencies and distortions.
[0086] Example three, please refer to Figures 3 - 9, which is a schematic structural diagram of a style transfer network provided by the present invention. The style transfer network uses an optimized method for style transfer, and this optimization process is achieved by calculating content loss and style loss. Specifically, the present invention uses the VGG-19 model to extract the features of the image and calculates the loss function based on these features to guide the optimization process.
[0087] The core idea of the style transfer module is to optimize an initial noise image to gradually approach the features of the target content image and the target style image. The present invention uses the VGG-19 model to extract the features of three key images (noise image, content image, style image), calculates the content loss and style loss based on these features, and then adjusts the noise image through an optimization algorithm until it minimizes both the content and style losses simultaneously.
[0088] For content, the present invention uses the features of the conv_4 layer, which is relatively deep in the model, enabling the present invention to capture more high-level features, that is, the general information of the scene. For style, the present invention retrieves features from the conv_1 to conv_5 layers, and these features contain general information as well as the texture details of the image. Then, the present invention calculates the content loss using the mean squared error (MSE) and calculates the style loss by calculating the Gram matrix and MSE.
[0089] In the video style transfer network of the present invention, the role of the style transfer module is to optimize the initial noise image so that it simultaneously inherits the structural information of the target content image and the texture features of the target style image. By using the VGG-19 model to extract content and style features and combining content loss and style loss for optimization, this module can accurately adjust the generated image to gradually approach the target style and content, thereby realizing video style transfer. Its significance lies in that by effectively fusing content and style features, the transferred image visually maintains the structure of the original scene and has the artistic features of the target style.
[0090] The specific implementation of the loss function calculation is as follows:
[0091] (1) Content loss and style loss
[0092] The content loss ensures that the content of the video frame remains consistent and avoids the destruction of the original image content due to style transfer. It is usually achieved by calculating the feature difference between the generated image and the reference image. The content loss extracts high-level features through a pre-trained VGG network to ensure that the semantic content of the generated image is consistent with the original video, where l represents the feature extraction layer in VGG-19. The content loss defined at layer l is the mean squared error between the feature map of the input frame xt and the feature map of its stylized output frame xt:
[0093] ;
[0094] Among them, represents the feature map of the l-th layer, is the dimension of the feature map of the l-th layer. The motivation for this content loss is that the high-level feature representations learned by the CNN abstract the content, and this content is what the present invention intends to preserve for the original input in the style transfer task. The content loss calculates the difference between the image features, encouraging the generated image to retain the high-level structure of the input image. In this way, it is ensured that the generated video does not lose the important content in the original image.
[0095] The style loss ensures that the generated video is consistent with the target style. This is achieved by measuring the difference between the generated image and the target image in the style space, usually by calculating the Gram matrix to measure the style similarity of the images:
[0096] ;
[0097] where is the feature map. The style loss is defined as the mean square error between the Gram matrices of the style image s and the stylized output frame :
[0098] ;
[0099] By comparing the style features of the generated image and the target image, the style loss can effectively keep the artistic style features of the image unchanged, so that the generated image or video not only meets the requirements in terms of content, but also can convey a specific artistic style or visual effect.
[0100] (2) Temporal consistency loss
[0101] The temporal consistency loss measures the consistency of the image content between video frames, especially the difference between adjacent frames. It calculates the pixel difference or feature difference between consecutive frames and maintains the coherence of the video content by minimizing these differences, preventing sudden changes or jumps between frames. Its core idea is that the frames in the video should transition smoothly visually, maintaining stable content and motion changes, rather than sudden drastic changes. During the style transfer process, unnatural flickers, misalignments or jumps may be introduced due to the stylization process, and this discontinuity will affect the viewing experience of the video. Therefore, through the temporal consistency loss, these unnatural changes can be reduced and the coherence of the video can be maintained. The present invention optimizes the style transfer network by calculating the MSE between the front and back frames and taking it as the temporal consistency loss:
[0102] ;
[0103] where and They are the i-th features extracted from the t-th and (t + 1)-th frames of the image in the network respectively. In the temporal consistency loss, MSE calculates the difference between the current frame and the previous frame, measuring their content consistency. A larger MSE value indicates a larger difference between the two frames, which may manifest as flickering, jumping, or other unnatural changes, while a smaller MSE value indicates that the content of the two frames is consistent and the visual transition is smoother.
[0104] (3) Optical flow smoothness loss
[0105] Although the temporal consistency loss helps to ensure the coherence of visual content between adjacent frames, it does not directly control the smoothness of motion between video frames. For example, if the optical flow in the video changes drastically, it may cause obvious jumps in the time series of the image, although the content itself remains consistent. The introduction of the optical flow smoothness loss is precisely to ensure that the motion between adjacent frames remains smooth and consistent.
[0106] Optical flow is a vector field that describes the change of each pixel in the image over time. It can represent the motion direction and speed of the image between different time points. In video style transfer, the goal of the optical flow smoothness loss is to make the pixel motion changes between consecutive frames smooth and consistent by constraining the change of the optical flow field, thus avoiding drastic motion changes or motion incoherence in video generation. The optical flow smoothness loss constrains the optical flow field between adjacent frames, making the motion vectors between adjacent frames consistent and not generating unnatural jumps or instantaneous changes. The formula is as follows:
[0107] ;
[0108] represents the gradient of the optical flow field u with respect to the spatial coordinates (x, y), measuring the change of the optical flow field in space. represents the difference between the optical flow fields at times t and t + 1, measuring the change of the optical flow field in time, is a hyperparameter used to balance the importance of the spatial and temporal smoothness terms.
[0109] The optical flow smoothness loss is jointly optimized with other losses to ensure that each generated frame of the image not only has the target style but also has no drastic motion changes in the time series of the video. By constraining the optical flow smoothness between video frames, the generated images will be more natural and consistent, avoiding discontinuous motion caused by style transfer.
[0110] In the present invention, the network structure of the present invention aims to achieve smooth transitions between video frames and ensure style consistency through the collaborative work of multiple modules. First, the network receives the current frame and the previous frame as inputs. The RAFT model is used to calculate the optical flow field between the two frames, capturing pixel-level motion information. Then, the calculated optical flow field is processed by the optical flow smoothing module to eliminate noise and inconsistencies, making the motion trajectory smoother and more natural.
[0111] Next, the network uses a feature extractor to extract the deep features of the current frame and the warped previous frame, identifying the subtle differences and commonalities between them. By calculating the mean squared error (MSE) between the features of the two frames, the network can quantify their differences and optimize the current frame through backpropagation to make its style and content more consistent with the previous frame. This optimization process ensures a more natural visual transition between video frames, avoiding drastic style changes, and finally generating an optimized output frame. The entire process not only improves the coherence of video content but also effectively reduces errors in optical flow estimation, ensuring smooth transitions between frames in tasks such as style transfer and video synthesis. The network structure is as Figure 4 shown.
[0112] The video style transfer algorithm of the present invention is based on the Pytorch deep learning framework, completed on an RTX4070 graphics card with Python version 3.11, and trained for 100 epochs. The transfer methods AdaIn, NeuralNeighborStyleTransfer, SANET are selected for comparison with the algorithm of this paper.
[0113] For objective metric analysis, RGB histograms of each method are plotted, and the RGB histogram superposition diagrams of each method are as Figures 5 - 8 shown. It can be seen from the histogram that the results of the present invention basically maintain the same RGB distribution between the front and back frames. At the same time, the method of the present invention also presents results with long-term style consistency on the visual comparison diagram.
[0114] The results show that the algorithm of the present invention has obtained the best scores in terms of deformation error, temporal error, and peak signal-to-noise ratio metrics. Especially in terms of temporal error and peak signal-to-noise ratio, it is better than other algorithms. The lower temporal error indicates that the algorithm of the present invention can better maintain continuity and dynamic consistency when processing image sequences, while the higher peak signal-to-noise ratio means that the quality difference between the reconstructed image and the original image is smaller, and the details and quality of the image are better preserved. This shows that the algorithm of the present invention not only outperforms other methods in terms of single-frame image quality but also can better handle changes and dynamics in time series, providing more accurate and clear results.
[0115] By analyzing the objective evaluation indexes and subjective visual feelings of multiple algorithms, the superiority of the video results migrated by this algorithm is verified. The method introduced in the present invention not only shows good effects in terms of subjective feelings, details, etc. in video style transfer, but also effectively reduces the occurrence of flicker phenomena.
[0116] The above has described an embodiment of the present invention in detail, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A video style transfer method based on RAFT optical flow, characterized in that: The following steps are involved: The RAFT optical flow network is used to calculate the pixel displacement of adjacent frames in the video frame by frame, generating an optical flow field containing the direction and size of each pixel's movement. Generate a dynamic weight map based on the displacement of pixels in the optical flow field, and the weight value of each area is inversely proportional to the displacement; Decompose the video frame into a low-frequency structure layer and a high-frequency detail layer through filtering, perform global style transfer on the low-frequency structure layer, and replace the color distribution and brushstroke features; Local stylization of high-frequency detail layers combined with dynamic weight maps; The processed low-frequency structure layer and high-frequency detail layer are fused to generate a single-frame stylized image; the stylized result of the current frame is motion-aligned with the adjacent frames according to the optical flow field, and the pixel position is adjusted through back-projection and interpolation compensation; Perform temporal smoothing filtering on the aligned multi-frame results to output a coherent stylized video; The generation process of the dynamic weight graph includes: The dynamic judgment threshold is set based on the overall displacement distribution of the optical flow field, and the area exceeding the dynamic judgment threshold is marked as a high dynamic area; the weight value of the high dynamic area is nonlinearly attenuated, and the attenuation degree gradually increases with the increase of displacement; A morphological expansion operation is performed on the boundary area of the weight map to smooth the transition boundary between the low-weight area and the high-weight area; the morphological expansion operation expands the coverage of the low-weight area through an expansion algorithm to avoid weight jumps at the motion boundary; An adaptive Gaussian kernel is used to smooth the weight map. The radius of the Gaussian kernel is dynamically adjusted according to the severity of the local displacement change to ensure a natural transition of the weight distribution. The radius of the adaptive Gaussian kernel is calculated as follows: according to the displacement standard deviation of the local area in the optical flow field, the larger the standard deviation, the larger the Gaussian kernel radius, and the smaller the standard deviation, the smaller the radius.
2. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The decomposition process of the low-frequency structure layer and the high-frequency detail layer is as follows: The original frame is blurred by a Gaussian filter to extract a low-frequency structure layer; the high-frequency detail layer is obtained by subtracting the original frame from the low-frequency structure layer, and the negative value area of the high-frequency detail layer is truncated and compensated; the truncation and compensation operation sets the negative value of the high-frequency detail layer to zero and then superimposes the fixed ratio intensity of the original detail; Performing contrast stretching on the high-frequency detail layer, wherein the contrast stretching expands the pixel value distribution of the high-frequency detail layer to a preset range through linear mapping, thereby improving the editability of the texture, and the stretching range is adaptively adjusted according to the pixel distribution of the detail layer; In the global style transfer of the low-frequency structure layer, the original geometric structure is retained and only the color and stroke features are replaced.
3. The video style transfer method based on RAFT optical flow according to claim 2, characterized in that: The local stylization of the high-frequency detail layer includes: Based on the preset weight threshold, the areas in the video frame are divided into low-weight areas, medium-weight areas and high-weight areas. For the high-frequency details in the low-weight areas, the stylization operation is limited to the luminance channel, and the original value of the chrominance channel is maintained; for the high-frequency details in the medium-weight areas, a multi-scale fusion strategy is adopted to superimpose coarse-grained brushstrokes and fine-grained textures respectively; for the high-frequency details in the high-weight areas, the style texture is fully applied, and the contour sharpness is maintained through the edge preservation algorithm; The multi-scale fusion strategy constructs style feature pyramids of different scales and superimposes style textures layer by layer from coarse to fine; The edge preservation algorithm detects the gradient information of the high-frequency detail layer and preserves the original contour of the area with a gradient higher than a set value during the stylization process.
4. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The fusion process of the low-frequency structure layer and the high-frequency detail layer includes: Dynamically adjust the feature fusion ratio between the low-frequency structure layer and the high-frequency detail layer according to the motion information of the optical flow field. For areas where the displacement in the optical flow field exceeds the dynamic judgment threshold, increase the weight ratio of the high-frequency detail layer in the fusion to the set ratio; for areas where the displacement is lower than the dynamic judgment threshold, increase the fusion weight ratio of the low-frequency structure layer to the set ratio. Construct a cross-layer feature correlation map, calculate the motion correlation between the low-frequency structure layer and the high-frequency detail layer of adjacent frames through the optical flow field, and use the same fusion ratio for high-correlation areas in the correlation map; The fused results are verified for consistency across frames, and the fusion ratio of the current frame is projected to the adjacent frames along the optical flow field to ensure the transition of high and low frequency features between multiple frames.
5. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The motion alignment specifically includes: The optical flow field is checked for bidirectional consistency. The cyclic error between the forward optical flow and the backward optical flow is calculated. The area where the error exceeds the preset threshold is marked as an unreliable area, and vice versa. The stylized features of the unreliable area are filled with weighted fusion of adjacent frames, and the weight is determined by the local smoothness of the optical flow field. The stylized features of the reliable area are aligned at the sub-pixel level, and the interpolation weight is adjusted by the fractional displacement value of the optical flow field. The sub-pixel alignment uses a bilinear interpolation algorithm to adjust the feature interpolation weights according to the fractional displacement value of the optical flow field so that the projection result matches the pixel grid of the target frame.
6. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: Also includes static area optimizations: Detect the area in the video where the displacement of multiple consecutive frames is lower than the dynamic judgment threshold and mark it as a static area; otherwise, it is a dynamic area, and high-precision style transfer is performed on the first frame of the static area. The subsequent frames reuse the first frame result and perform fine-tuning through the optical flow field; the fine-tuning operation adjusts the pixel position of the reused area according to the displacement of the optical flow field to compensate for camera jitter or lighting changes; Gradual blending is performed at the junction of the static area and the dynamic area, and the stylized intensity is transitioned by transparency superposition; the gradual blending forms a stylized intensity change at the junction through a linear transition of the alpha channel.
7. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: It also includes identifying key target areas in the video through pre-trained models; setting a lower weight limit for the moving parts of the key target areas to limit the stylization intensity; setting an upper weight limit for the static parts of the key target areas to retain local details; The pre-trained model adopts a semantic segmentation algorithm based on a convolutional neural network to output pixel-level target category labels; the setting of the lower and upper weight limits is dynamically adjusted according to the importance of the target category, and the importance of the target category is obtained based on a preset target mapping table.
Citation Information
Patent Citations
Motion guide mask method and pre-training method of visual transformer model
CN116168331A