Video style migration method based on RAFT optical flow

Through the video style transfer method based on RAFT optical flow, the problems of coherence and smooth transition between video frames are solved, efficient optical flow calculation and dynamic weight map generation are realized, and the effect and quality of video style transfer are significantly improved.

CN119991416AActive Publication Date: 2025-05-13NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510480941.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing video style transfer technology cannot effectively maintain the consistency and smooth transition between video frames, especially when dealing with complex backgrounds and fast motion scenes, it is impossible to accurately capture local motion information, resulting in poor style transfer results.

Method used

Using a video style migration method based on RAFT optical flow, the pixel displacement of adjacent video frames is calculated through the RAFT optical flow network, an optical flow field is generated, and a dynamic weight map is generated based on the optical flow field. Then, the video frame is decomposed into a low-frequency structure layer and a high-frequency detail layer, and global and local style transfers are performed, and motion alignment and timing smoothing filtering are performed through the optical flow field to generate a coherent stylized video.

Benefits of technology

Through accurate optical flow calculation and dynamic weight graph generation, it can effectively adapt to the motion of the video content, balance style presentation and detail retention, significantly improve the effect and quality of video style transfer, and avoid lag and incoherence in stylized videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991416A_ABST
    Figure CN119991416A_ABST
Patent Text Reader

Abstract

The invention discloses a video style migration method based on RAFT optical flow, which belongs to the technical field of video processing, and specifically comprises the following steps: calculating pixel displacement of adjacent frames in a video frame by frame through an RAFT optical flow network, and generating an optical flow field containing the motion direction and size of each pixel; generating a dynamic weight map according to the displacement of the pixels in the optical flow field; decomposing a video frame into a low-frequency structure layer and a high-frequency detail layer, and performing global style migration on the low-frequency structure layer; performing local stylization on the high-frequency detail layer in combination with the dynamic weight map; fusing the processed low-frequency structure layer and high-frequency detail layer to generate a single-frame stylized image; performing motion alignment on the stylization result of the current frame and the adjacent frame according to the optical flow field, and adjusting the pixel position through back projection and interpolation compensation; performing time sequence smooth filtering on the aligned multi-frame result, and outputting a coherent stylized video; according to the invention, the time consistency and motion fluency between video frames are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular to a video style migration method based on RAFT optical flow. Background Art

[0002] In the task of video style transfer, the optical flow network plays a vital role. It estimates the motion information between video frames to ensure that the style transfer is consistent and smoothly transitioned in the temporal dimension. Traditional style transfer methods usually only process single-frame images and ignore the dynamic changes between frames in the video, which may cause the style-transferred video to look incoherent or jumpy. The optical flow network can accurately capture the motion pattern between each frame and maintain the relative motion of objects and backgrounds in the video, thereby avoiding unnatural transitions or blurs during the style transfer process. Optical flow can not only handle fast motion and dynamic changes in complex scenes, but also help solve problems caused by object occlusion or texture loss, ensuring that each frame after style transfer is consistent in motion and vision. In addition, the optical flow network can also provide more accurate local motion estimation and global motion patterns, so as to better adapt to motion changes in multi-scale scenes and improve the effect and quality of style transfer. The optical flow network combined with deep learning can automatically learn motion information through training, making the calculation efficiency higher and the estimation accuracy more accurate. This enables the application of optical flow in video style transfer to provide high-quality and stable transfer effects in long time series and complex scenes, ultimately achieving a more natural and smooth video style transfer.

[0003] Traditional optical flow estimation methods, such as Lucas-Kanade and Horn-Schunck, rely on assumptions based on relative motion of local pixels. These methods can provide basic optical flow estimation under some ideal conditions. However, these traditional methods face many challenges in modern video style transfer tasks, especially when dealing with complex and rapidly changing scenes. First, traditional optical flow algorithms have low accuracy in videos with high dynamic range or scenes with fast motion. This is because traditional methods usually cannot capture large-scale or fine-grained motion information, especially when there is a large amount of motion between frames, which easily leads to distortion of optical flow estimation. This will affect the quality of style transfer, causing the generated video to have unnatural motion blur or object position errors, thereby destroying style consistency. Second, traditional optical flow methods are very sensitive to complex phenomena such as lighting changes, texture loss, occlusion or reflection. In real scenes, video frames are often affected by ambient lighting, object reflections, and changes in camera angles, which makes it difficult for traditional algorithms to accurately estimate optical flow. Especially in areas with less texture or lower contrast, traditional optical flow methods often cannot provide effective estimates, resulting in large errors in optical flow in these areas, which affects the effect of video style transfer, making the style of the transferred video unnatural or even obviously distorted. Summary of the invention

[0004] The purpose of the present invention is to provide a video style transfer method based on RAFT optical flow to solve the following technical problems: Existing video style transfer technology cannot effectively maintain the coherence and smooth transition between video frames. It cannot accurately capture local motion information when dealing with complex backgrounds and details, and it does not work well when dealing with fast-motion scenes, affecting the overall transfer effect.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A video style transfer method based on RAFT optical flow, comprising the following steps: The RAFT optical flow network is used to calculate the pixel displacement of adjacent frames in the video frame by frame, generating an optical flow field containing the direction and size of each pixel's movement. Generate a dynamic weight map based on the displacement of pixels in the optical flow field, and the weight value of each area is inversely proportional to the displacement; The video frames are decomposed into a low-frequency structure layer and a high-frequency detail layer through filtering. The low-frequency structure layer is globally transferred to replace the color distribution and stroke features. The high-frequency detail layer is locally stylized in combination with a dynamic weight map. The processed low-frequency structure layer and high-frequency detail layer are fused to generate a single-frame stylized image; the stylized result of the current frame is motion-aligned with the adjacent frames according to the optical flow field, and the pixel position is adjusted through back-projection and interpolation compensation; The aligned multi-frame results are temporally smoothed and filtered to output a coherent stylized video.

[0006] As a further solution of the present invention: the generation process of the dynamic weight map includes: The dynamic judgment threshold is set based on the overall displacement distribution of the optical flow field, and the area exceeding the dynamic judgment threshold is marked as a high dynamic area; the weight value of the high dynamic area is nonlinearly attenuated, and the attenuation degree gradually increases with the increase of displacement; A morphological expansion operation is performed on the boundary area of ​​the weight map to smooth the transition boundary between the low-weight area and the high-weight area; the morphological expansion operation expands the coverage of the low-weight area through an expansion algorithm to avoid weight jumps at the motion boundary; An adaptive Gaussian kernel is used to smooth the weight map. The radius of the Gaussian kernel is dynamically adjusted according to the severity of the local displacement change to ensure a natural transition of the weight distribution. The radius of the adaptive Gaussian kernel is calculated as follows: according to the displacement standard deviation of the local area in the optical flow field, the larger the standard deviation, the larger the Gaussian kernel radius, and the smaller the standard deviation, the smaller the radius.

[0007] As a further solution of the present invention: the decomposition process of the low-frequency structure layer and the high-frequency detail layer is specifically as follows: The original frame is blurred by a Gaussian filter to extract a low-frequency structure layer; the high-frequency detail layer is obtained by subtracting the original frame from the low-frequency structure layer, and the negative value area of ​​the high-frequency detail layer is truncated and compensated; the truncation and compensation operation sets the negative value of the high-frequency detail layer to zero and then superimposes the fixed ratio intensity of the original detail; The high-frequency detail layer is subjected to contrast stretching, which expands the pixel value distribution of the high-frequency detail layer to a preset range through linear mapping, thereby improving the editability of the texture. The stretching range is adaptively adjusted according to the pixel distribution of the detail layer. In the global style transfer of the low-frequency structure layer, the original geometric structure is retained, and only the color and stroke features are replaced.

[0008] As a further solution of the present invention: the local stylization of the high-frequency detail layer includes: Based on the preset weight threshold, the areas in the video frame are divided into low-weight areas, medium-weight areas and high-weight areas. For the high-frequency details in the low-weight areas, the stylization operation is limited to the luminance channel, and the original value of the chrominance channel is maintained; for the high-frequency details in the medium-weight areas, a multi-scale fusion strategy is adopted to superimpose coarse-grained brushstrokes and fine-grained textures respectively; for the high-frequency details in the high-weight areas, the style texture is fully applied, and the contour sharpness is maintained through the edge preservation algorithm; The multi-scale fusion strategy constructs style feature pyramids of different scales and superimposes style textures layer by layer from coarse to fine; The edge preservation algorithm detects the gradient information of the high-frequency detail layer and preserves the original contour of the area with a gradient higher than a set value during the stylization process.

[0009] As a further solution of the present invention: the fusion process of the low-frequency structure layer and the high-frequency detail layer includes: Dynamically adjust the feature fusion ratio between the low-frequency structure layer and the high-frequency detail layer according to the motion information of the optical flow field. For areas where the displacement in the optical flow field exceeds the dynamic judgment threshold, increase the weight ratio of the high-frequency detail layer in the fusion to the set ratio; for areas where the displacement is lower than the dynamic judgment threshold, increase the fusion weight ratio of the low-frequency structure layer to the set ratio. Construct a cross-layer feature correlation map, calculate the motion correlation between the low-frequency structure layer and the high-frequency detail layer of adjacent frames through the optical flow field, and use the same fusion ratio for high-correlation areas in the correlation map; The fused results are verified for consistency across frames, and the fusion ratio of the current frame is projected to the adjacent frames along the optical flow field to ensure the transition of high and low frequency features between multiple frames.

[0010] As a further solution of the present invention: the motion alignment specifically includes: The optical flow field is checked for bidirectional consistency. The cyclic error between the forward optical flow and the backward optical flow is calculated. The area where the error exceeds the preset threshold is marked as an unreliable area, and vice versa. The stylized features of the unreliable area are filled with weighted fusion of adjacent frames, and the weight is determined by the local smoothness of the optical flow field. The stylized features of the reliable area are aligned at the sub-pixel level, and the interpolation weight is adjusted by the fractional displacement value of the optical flow field. The sub-pixel alignment uses a bilinear interpolation algorithm to adjust the feature interpolation weights according to the fractional displacement value of the optical flow field so that the projection result matches the pixel grid of the target frame.

[0011] As a further solution of the present invention: static area optimization is also included: Detect the area in the video where the displacement of multiple consecutive frames is lower than the dynamic judgment threshold and mark it as a static area; otherwise, it is a dynamic area, and high-precision style transfer is performed on the first frame of the static area. The subsequent frames reuse the first frame result and perform fine-tuning through the optical flow field; the fine-tuning operation adjusts the pixel position of the reused area according to the displacement of the optical flow field to compensate for camera jitter or lighting changes; Gradual blending is performed at the junction of the static area and the dynamic area, and the stylized intensity is transitioned by transparency superposition; the gradual blending forms a stylized intensity change at the junction through a linear transition of the alpha channel.

[0012] As a further solution of the present invention: it also includes identifying key target areas in the video through a pre-trained model; setting a weight lower limit for the moving part of the key target area to limit the stylization strength; setting a weight upper limit for the static part of the key target area to retain local details; The pre-trained model adopts a semantic segmentation algorithm based on a convolutional neural network to output pixel-level target category labels; the setting of the lower and upper weight limits is dynamically adjusted according to the importance of the target category, and the importance of the target category is obtained based on a preset target mapping table.

[0013] Beneficial effects of the present invention: The present invention uses the RAFT optical flow network to accurately calculate the pixel displacement of adjacent frames of the video to generate an optical flow field. The dynamic weight map generated in this way can flexibly adapt to the motion of the video content, overcoming the problem of lack of flexibility in weight setting in the prior art. In terms of video frame decomposition, the video frame is accurately decomposed into a low-frequency structure layer and a high-frequency detail layer using methods such as Gaussian filtering, which can effectively balance the overall style presentation and detail retention, and solve the problem of inaccurate division in the prior art. In terms of stylization operations, the low-frequency structure layer is globally styled, and the high-frequency detail layer is locally stylized in combination with the dynamic weight map, and different strategies are adopted for different weight areas. At the same time, the optical flow field is used for motion alignment and temporal smoothing filtering, which effectively avoids stuttering and incoherence in the stylized video, significantly improves the viewing experience, and further improves the video style transfer effect through static area optimization and key target area processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present invention will be further described below in conjunction with the accompanying drawings.

[0015] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 is a schematic diagram of the structure of the RAFT module of the present invention; Figure 3 Schematic diagram of the structure of the style transfer network of the present invention; Figure 4 It is the collaborative workflow of the style transfer network of the present invention; Figure 5 is the RGB histogram of the present invention; Figure 6 It is the RGB histogram of the comparison method AdaIn; Figure 7 It is the RGB histogram of the comparison method NNST; Figure 8 It is the RGB histogram of the comparison method SANET; Fig. 9 It is a structural schematic diagram of the video style transfer based on RAFT optical flow of the present invention. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] For example, see Figure 1 As shown, a video style transfer method based on RAFT optical flow provided by the present invention includes the following steps: 1. As an advanced and efficient optical flow calculation model, the RAFT optical flow network works based on a deep learning architecture and uses a large amount of video frame data for training, so that it can accurately capture the changes in pixels between adjacent frames. During the calculation process, the network carefully analyzes each frame of the video, and outputs an optical flow field containing the movement direction and size of each pixel through complex neural network operations. This optical flow field is like a detailed "motion map", which provides a solid data foundation for the subsequent dynamic analysis and processing of video content, allowing the present invention to clearly understand the movement trajectory of each pixel in the video between frames, providing a key basis for the dynamic adjustment of style transfer.

[0018] 2. Generate a dynamic weight map based on the displacement of pixels in the optical flow field. This process has rigorous and scientific steps. First, a dynamic judgment threshold is set based on the overall displacement distribution of the optical flow field. This judgment threshold is not fixed, but is obtained through statistical analysis of the displacement of all pixels in the optical flow field. For example, the mean, median and other statistical quantities of the displacement are calculated, and a reasonable dynamic judgment threshold is determined in combination with the characteristics of the video content (such as scene complexity, the proportion of moving objects, etc.). Areas exceeding this threshold are keenly marked as high-dynamic areas, which means that the pixels in these areas move more violently between frames. For the weight values ​​of high-dynamic areas, a nonlinear attenuation strategy is adopted, and the degree of attenuation gradually increases with the increase of displacement. The logic behind this is that the more violent the movement of the area, the more features of the original video may need to be retained during the style transfer process to avoid excessive style interference and content distortion. Therefore, the weight value of the area with a large displacement is reduced, and this reduction is not linear, but decays at a faster rate as the displacement increases, so as to accurately control the intensity of stylization.

[0019] In order to make the weight map more reasonable and smooth, a morphological expansion operation is performed on the boundary area of ​​the weight map. Specifically, the dilation algorithm is used here to expand the coverage of the low-weight area. The dilation algorithm is a commonly used morphological operation in image processing. It expands the boundary pixels of the low-weight area outward to make the transition boundary between the low-weight area and the high-weight area smoother and more natural, avoiding weight jumps at the motion boundary. This jump may cause obvious discontinuity of the stylized effect at the boundary, seriously affecting the visual effect. Through a carefully designed dilation algorithm, this potential defect can be cleverly eliminated, making the transition between different areas of the weight map smoother.

[0020] Finally, the weight map is smoothed by an adaptive Gaussian kernel to further ensure the natural transition of the weight distribution. The radius of the adaptive Gaussian kernel is not a fixed value, but is dynamically adjusted according to the severity of the local displacement change. The specific calculation method is: in-depth analysis of the displacement standard deviation of the local area in the optical flow field. The standard deviation is an important indicator to measure the degree of data discreteness. In the optical flow field, the larger the standard deviation of the local area displacement, the more drastic the change of the pixel displacement in the area. At this time, a larger Gaussian kernel radius is required for smoothing to better balance the weight distribution; conversely, the smaller the standard deviation, the relatively stable pixel displacement change in the area, and the Gaussian kernel radius is correspondingly reduced. By dynamically adjusting the radius of the Gaussian kernel, the weight map can be flexibly and accurately smoothed according to the actual motion of the video content, so that the final generated dynamic weight map can be highly consistent with the dynamic characteristics of the video.

[0021] 3. Reasonably decompose the video frame and use the Gaussian filter to blur the original frame to extract the low-frequency structure layer. The Gaussian filter is essentially a linear smoothing filter. Its principle is based on the Gaussian function. The blurring effect is achieved by weighted averaging each pixel in the original frame and its neighboring pixels. Since the low-frequency signal changes slowly and appears as a large area of ​​similar areas in space, the Gaussian filter can effectively retain the information of these areas, while the fast-changing details corresponding to the high-frequency signal (such as edges, textures, etc.) are smoothed out during the blurring process, thereby successfully extracting the low-frequency structure layer that mainly contains the outline of the scene subject, large-area color and brightness change trends and other information.

[0022] Subtracting the original frame from the extracted low-frequency structure layer will yield the high-frequency detail layer. However, in this process, negative areas will appear in the high-frequency detail layer. To solve this problem, truncation and compensation operations are required. Specifically, the negative values ​​in the high-frequency detail layer are directly set to zero to eliminate unreasonable negative pixel values. Subsequently, the fixed ratio intensity of the original details is superimposed. This fixed ratio is not set arbitrarily, but is determined based on a large number of experiments and analysis of the video content. It aims to ensure that while removing negative values, the original detail information is retained as much as possible, maintaining the integrity and accuracy of the high-frequency detail layer, and making subsequent processing based on more reliable detail data.

[0023] In order to further improve the editability of the high-frequency detail layer texture, it is necessary to perform contrast stretching on it. This operation uses the principle of linear mapping to expand the pixel value distribution of the high-frequency detail layer to a preset range. The preset range is not fixed, but is adaptively adjusted according to the pixel distribution of the detail layer. For example, the minimum and maximum values ​​of the pixel values ​​of the high-frequency detail layer are first counted, and reasonable mapping parameters are calculated based on the characteristics of the video content (such as whether it is a high-contrast scene, etc.), so that the pixel values ​​can more clearly show the texture details within the new range, providing a more vivid and easier-to-process texture foundation for subsequent stylization operations. When performing global style transfer on the low-frequency structure layer, the original geometric structure is carefully preserved, and only the color and stroke features are replaced. Through a specific style transfer algorithm (such as a neural network-based style transfer model), the color distribution and stroke features of the target style are integrated into the low-frequency structure layer to ensure that the basic geometry and layout of the video scene are not destroyed while changing the style, and the recognizability of the scene is maintained.

[0024] The local stylization operation of the high-frequency detail layer is meticulous and exquisite. Based on the preset weight threshold, the regions in the video frame are accurately divided into low-weight regions, medium-weight regions, and high-weight regions. For the high-frequency details in the low-weight region, the stylization operation is strictly limited to the luminance channel to preserve the original color information, and the original value of the chrominance channel is maintained. In this way, while introducing style changes to a certain extent, the color information is ensured to be as consistent as possible with the original video, avoiding the negative impact of color distortion on the visual effect. For the high-frequency details in the medium-weight region, a multi-scale fusion strategy is adopted. Specifically, it is achieved by constructing a style feature pyramid of different scales, starting from the coarse-grained style features at the top layer of the pyramid, and superimposing fine-grained textures layer by layer. In this process, the coarse-grained brushstrokes at the top layer can quickly outline the overall style outline, and the fine-grained textures at the bottom layer add rich details to the style. The organic combination of the two makes the stylization effect more natural and delicate. For the high-frequency details in the high-weight region, the style texture is fully applied, and the sharpness of the outline is maintained with the help of the edge preservation algorithm. The edge preservation algorithm works by detecting the gradient information of the high-frequency detail layer, presetting a gradient threshold. During the stylization process, for areas with a gradient higher than the set value, the original contours are preserved. Because these high-gradient areas often correspond to important details such as the edges of objects, preserving their original contours can ensure that the stylized video presents a new style while maintaining detail clarity, thereby improving visual quality.

[0025] Static area optimization is also an indispensable part of the entire video processing process. By detecting the area in the video where the displacement of multiple consecutive frames is lower than the dynamic judgment threshold, it is keenly marked as a static area, and vice versa. For the static area, high-precision style transfer is performed on its first frame, which means using a more complex and sophisticated style transfer algorithm and investing more computing resources to ensure the generation of high-quality stylized results. Subsequent frames reuse the first frame results and are fine-tuned through the optical flow field. The fine-tuning operation adjusts the pixel position of the reused area according to the displacement of the optical flow field to compensate for camera shake or lighting changes. Since the static area is not absolutely still, there may be pixel position changes caused by slight camera shaking or slow changes in lighting. Through the precise displacement information of the optical flow field, the pixel position can be accurately adjusted to maintain the stability and coherence of the video. At the junction of the static area and the dynamic area, a gradient blending is performed to achieve a natural transition of the stylization intensity. This is achieved through the linear transition of the alpha channel, which is used to control the transparency of the image. At the junction, the alpha value is linearly adjusted according to the distance from the static area and the dynamic area, so that the stylization intensity smoothly transitions from the stylization degree of the static area to the stylization degree of the dynamic area, avoiding obvious style mutations and improving the overall visual smoothness of the video.

[0026] In addition, the key target areas in the video are identified through the pre-trained model. Here, a semantic segmentation algorithm based on a convolutional neural network is used to build a pre-trained model. The algorithm uses components such as convolutional layers, pooling layers, and fully connected layers to learn a large amount of annotated video data, so that it can output pixel-level target category labels and accurately identify various objects and areas in the video. For the identified key target areas, the moving part and the static part are further distinguished. A weight lower limit is set for the moving part to limit the stylization intensity to prevent excessive style changes from causing loss of details of the moving part of the key target or abnormal visual effects. An upper weight limit is set for the static part to retain local details, ensuring that the static part of the key target can retain its original detail features during the style transfer process and maintain the recognizability and integrity of the target. The setting of the lower and upper weight limits is not a fixed value, but is dynamically adjusted according to the importance of the target category. The importance of the target category is obtained based on a preset target mapping table, which pre-classifies and assigns weights to various targets (such as people, vehicles, buildings, etc.) according to their importance in the video content, so that in practical applications, the stylization weights of the key target areas can be flexibly and reasonably adjusted according to the characteristics of the video content.

[0027] 4. Fusion of low-frequency structure layer and high-frequency detail layer to generate a single-frame stylized image is an important stage towards the final stylized video. In order to achieve high-quality fusion effect, it is necessary to dynamically adjust the feature fusion ratio between low-frequency structure layer and high-frequency detail layer according to the motion information of optical flow field. The dynamic judgment threshold plays a key role here, which is determined based on the overall distribution of displacement in optical flow field. For the area where the displacement in the optical flow field exceeds the dynamic judgment threshold, it means that the pixel motion in this area is more intense and the scene changes are rich. At this time, in order to better present the details and style of the dynamic area, the weight ratio of the high-frequency detail layer in the fusion is increased to the set ratio. This set ratio is not arbitrarily specified, but is obtained after a large number of experiments and in-depth analysis of the video content. It aims to highlight the expressiveness of high-frequency details in rapidly changing areas, so that the stylized image in these areas can not only reflect the texture details of the new style, but also closely fit the dynamic characteristics of the original video. On the contrary, for the area where the displacement is lower than the dynamic judgment threshold, the area is relatively stable and the scene changes are small, so the fusion weight ratio of the low-frequency structure layer is increased to the set ratio to emphasize the consistency of the overall structure and style of the scene and maintain the visual coherence of the stable area.

[0028] Constructing a cross-layer feature correlation map is a key strategy to optimize the fusion effect. Through the powerful motion analysis capability of the optical flow field, the motion correlation between the low-frequency structure layer and the high-frequency detail layer of adjacent frames is calculated. The optical flow field can accurately capture the motion trajectory of pixels in each frame. Based on this, the similarity and correlation between different layers of adjacent frames can be quantified. In the correlation map, the same fusion ratio is used for high-correlation areas, that is, those areas with similar motion patterns between adjacent frames. This ensures the coherence of the video in the temporal dimension, avoids visual jumps caused by large fluctuations in the fusion ratio between different frames, and makes the stylized video present a natural and smooth visual transition during motion.

[0029] To further ensure the quality of stylized videos, it is crucial to verify the consistency of the fused results across frames. This process projects the fusion ratio of the current frame to the adjacent frames along the optical flow field. The optical flow field is like a link that connects the frames of the video. With its precise motion information, it transmits the fusion ratio information of the current frame to the adjacent frames. In this way, the transition of high- and low-frequency features between multiple frames can be smoothly achieved, avoiding the abruptness caused by inconsistent fusion ratios, greatly improving the visual smoothness and coherence of the stylized video, and bringing a more comfortable viewing experience to the audience.

[0030] After the fusion operation is completed, the stylized result of the current frame is motion aligned with the adjacent frames according to the optical flow field, which is a key link to ensure the natural dynamic effect of the stylized video. First, the optical flow field is bidirectionally checked for consistency. The reliability of the optical flow field is evaluated by calculating the cyclic error of the forward optical flow and the backward optical flow. The forward optical flow describes the pixel motion from the current frame to the next frame, while the backward optical flow is the opposite. If the cyclic error of the two exceeds the preset threshold, it means that there may be a deviation in the optical flow calculation of the area, and the area is marked as an unreliable area; otherwise, it is a reliable area. For the stylized features of the unreliable area, the weighted fusion filling of adjacent frames is used to optimize. The local smoothness of the optical flow field acts as a weight determining factor here. The higher the local smoothness, the smoother the motion change of the area, and the lower the dependence on adjacent frames during weighted fusion; otherwise, it is higher. This weighted fusion method based on local smoothness can effectively make up for the defects of optical flow calculation in unreliable areas and improve the stability of the stylized effect. For the stylized features of the reliable area, sub-pixel alignment is performed to achieve higher-precision motion matching. Sub-pixel alignment uses a bilinear interpolation algorithm to adjust the feature interpolation weights according to the fractional displacement value of the optical flow field. The bilinear interpolation algorithm determines the pixel value of the interpolation position by weighted average of the four adjacent pixels in the target frame, while the fractional displacement value of the optical flow field accurately indicates the direction and degree of adjustment of the interpolation weight, so that the projection result is accurately matched with the pixel grid of the target frame, ensuring that the stylized video will not have problems such as pixel dislocation or blur during motion, presenting a clear and smooth dynamic visual effect.

[0031] 5. The aligned multi-frame results are subjected to temporal smoothing filtering, which can eliminate the noise, fluctuations and discontinuities introduced by the previous steps, ensuring that the video presents a natural and smooth visual effect in the time dimension. Its principle is based on the time series analysis of multi-frame video data, and uses a specific filtering algorithm to adjust and optimize the pixel value or feature information of each frame. Common algorithms include mean filtering, median filtering and time-domain Gaussian filtering, each with its own characteristics and applicable scenarios.

[0032] Mean filtering calculates the average value of the same pixel position in several adjacent frames, replaces the original value of the corresponding pixel in the current frame, weakens the interference of single-frame data fluctuations, and makes the video picture transition smoother. Median filtering selects the median of the corresponding pixel set of adjacent frames as the pixel output value of the current frame, which can effectively remove impulse noise and avoid obvious defects in the video picture. Time-domain Gaussian filtering assigns different weights to the pixel values ​​of adjacent frames based on the characteristics of the Gaussian function, and pays more attention to the reference of adjacent frame information, which can smooth noise fluctuations and retain dynamic details. In practical applications, it is necessary to select appropriate filtering algorithms and parameters based on the video frame rate, content complexity, degree of stylization and expected effect.

[0033] For example 2, please refer to Figure 2 The figure is a schematic diagram of the structure of a RAFT module provided by the present invention.

[0034] The RAFT (Recurrent All-Pairs Field Transforms) module is a deep learning model for optical flow estimation. Its core advantage is that it can efficiently and accurately estimate global optical flow. Compared with traditional optical flow methods, RAFT performs full-to-full pixel-level matching in the entire image and performs iterative optimization to capture more detailed and accurate optical flow information.

[0035] FeatureEncoder extracts feature maps of two adjacent frames, while ContextEncoder only extracts features from the first frame. Both are CNN-based networks and can be understood as shallow custom ResNets. The circle L represents the Look-up operation, and the series of boxes and arrows in the middle represent iterative optical flow estimation using GRU (a recurrent network).

[0036] The optical flow estimation process of the RAFT model starts with initializing the optical flow result to zero, that is, assuming that the initial optical flow field is . Then, ContextEncoder is used to extract global context information from the first frame of the input image, and FeatureEncoder is used to extract features of the first and second frames. Then, through matrix multiplication operations, 4D correlation volumes (4DCorrelationVolumes) are calculated, which capture the matching information between each pair of pixels in the image. Based on these features and correlation information, the model uses the GRU (GatedRecurrentUnit) module to make a preliminary estimate of the optical flow, obtain the optical flow update, and update the optical flow field as , that is, by adding the initial zero optical flow, the optical flow estimation of the first step is obtained.

[0037] However, this initial estimate is usually not accurate enough, so RAFT further improves the accuracy of optical flow estimation through iterative optimization. Specifically, the model finds the correlation information in the 4D correlation volume, passes the updated optical flow result as input to the GRU, and calculates the optical flow update again. With each iteration, the optical flow field is continuously updated to obtain a new optical flow estimate This recursive process not only corrects the initial estimation results through repeated iterations, but also gradually optimizes the optical flow estimation and finally achieves an accurate optical flow field. Through multiple iterations, RAFT can gradually improve the optical flow estimation based on the more accurate motion information obtained in each iteration, so that the accuracy of the final optical flow estimation in spatiotemporal motion is greatly improved.

[0038] 4DCorrelationVolumes is a 4D volume pixel obtained by calculating the pixel-by-pixel correlation of the feature maps of two adjacent frames. The size is The calculation method can be understood as transforming the feature map of the first frame into The matrix of the second frame becomes The matrix of , and then the two are multiplied to get , adjust the shape (reshape) to get the final result, which can be expressed as follows: ; ; Each element of 4DCorrelationVolumes It can be expressed as the correlation between the (i, j)th pixel in the first frame and the (k, l)th pixel in the second frame.

[0039] Therefore, RAFT can be divided into three stages: Feature Extraction: The network input consists of two consecutive frames. To extract features from these two images, the network uses two CNNs with shared weights. The CNN architecture consists of 6 residual layers, just like the layers of ResNet, where the resolution is reduced by half with each subsequent layer, while the number of channels increases.

[0040] Visual Similarity: Visual similarity is calculated as the inner product of all feature map pairs. As a result, a 4D tensor called the correlation volume is obtained, which provides key information about the displacement of large and small pixels. Then, the last two dimensions of this 4D tensor are pooled with kernels of size 1, 2, 4, and 8 to construct a 4-layer correlation pyramid.

[0041] Iterative updates: An iterative update is a sequence of gated recurrent units (GRUs) that combines all the data previously calculated by the invention. The GRU unit simulates an iterative optimization algorithm, but with an improvement - trainable convolutional layers with shared weights. Each update iteration produces a new optical flow update to make the prediction more accurate with each new step. In the present invention, the RAFT module is applied to the optical flow estimation step of the video frame. By using RAFT to estimate the optical flow of adjacent video frames, the motion information of each pixel is extracted as input data, the motion changes between video frames are accurately captured, and the best motion feature signal is found, thereby constraining the temporal consistency and spatial coherence of the video frames. This precise motion estimation helps to optimize the inter-frame transition in the process of video style transfer, reduce the unnatural effects caused by motion distortion or dislocation, and improve the coherence and stability of the video.

[0042] The optical flow estimation result output by the RAFT module is then optimized by the optical flow smoothing module to reduce noise and inconsistency and ensure the smoothness of the optical flow field. This module is implemented by minimizing a comprehensive energy function that combines the constraints of spatial smoothness and temporal smoothness.

[0043] In the video style transfer network of the present invention, the core function of the optical flow smoothing module is to optimize the optical flow estimation results and eliminate the problems caused by noise and inconsistency in motion estimation. Specifically, the optical flow smoothing module not only focuses on the spatial smoothing of the optical flow field, but also combines temporal smoothing to ensure the consistency of motion information between frames. The optical flow smoothing module is mainly composed of spatial smoothing and temporal smoothing. Spatial smoothing mainly ensures that the optical flow field is smooth in the spatial dimension, that is, the motion information of adjacent pixels will not fluctuate violently. By constraining the optical flow vectors of adjacent pixels, discontinuities caused by feature noise or local estimation errors are avoided. Temporal smoothing ensures that the optical flow changes between adjacent video frames are continuous. That is to say, in the temporal dimension, the motion of the object will not jump or change unnaturally. It makes the changes in the optical flow field smoother by constraining the optical flow differences between adjacent frames.

[0044] In the network of the present invention, the optical flow smoothing module optimizes the optical flow estimation by minimizing a comprehensive energy function. The specific objective function form combines the constraints of spatial smoothing and temporal smoothing, and the general expression is: ; in is the optical flow vector of pixel p in the optical flow field. is the spatial gradient of the optical flow field, representing the constraint of spatial smoothness. and is a hyperparameter used to balance the weights of spatial smoothing and temporal smoothing. d is the time difference, which represents the displacement between adjacent frames and is used to constrain the smoothing of optical flow in the temporal dimension. Ω is the set of all pixels in the image. The gradient of the optical flow field in the spatial dimension is calculated to ensure that the difference between the optical flow estimates of adjacent pixels is as small as possible. This term makes the optical flow change smoothly in space and avoids drastic changes caused by noise or estimation errors. Enforces the optical flow changes between adjacent frames to be consistent. By constraining the difference in optical flow between the current frame and the adjacent frame, this term ensures that the optical flow field remains smooth in the time dimension and avoids incoherent motion of objects on the time axis.

[0045] The optical flow smoothing module in the video style transfer network of the present invention optimizes the optical flow estimation results through spatial and temporal smoothing constraints to ensure the stability and consistency of the optical flow field. This module can effectively eliminate noise in the optical flow estimation and improve the coherence and smoothness of the video, so that the object movement in the style transfer process is more natural and avoids inconsistency and distortion.

[0046] For example 3, please refer to Figure 3-Figure 9, which is a schematic diagram of the structure of a style transfer network provided by the present invention. The style transfer network uses an optimization method to perform style transfer, and this optimization process is achieved by calculating content loss and style loss. Specifically, the present invention uses the VGG-19 model to extract image features, and calculates the loss function based on these features to guide the optimization process.

[0047] The core idea of ​​the style transfer module is to optimize an initial noise image so that it gradually approaches the characteristics of the target content image and the target style image. The present invention uses the VGG-19 model to extract the features of three key images (noise image, content image, style image), and calculates the content loss and style loss based on these features, and then adjusts the noise image through the optimization algorithm until it minimizes the loss of content and style at the same time.

[0048] For content, we use the features of the conv_4 layer, which is relatively deep in the model, enabling us to capture higher-level features, i.e., general information of the scene. For style, we retrieve features from conv_1 to conv_5 layers, which contain general information as well as texture details of the image. We then calculate the content loss using mean squared error (MSE) and the style loss by calculating the Gram matrix and MSE.

[0049] In the video style transfer network of the present invention, the role of the style transfer module is to optimize the initial noise image so that it inherits the structural information of the target content image and the texture characteristics of the target style image at the same time. By using the VGG-19 model to extract content and style features, and combining content loss and style loss for optimization, the module can accurately adjust the generated image so that it gradually approaches the target style and content, thereby realizing video style transfer. Its significance lies in that by effectively integrating content and style features, the transferred image visually maintains the structure of the original scene and has the artistic characteristics of the target style.

[0050] The specific implementation of loss function calculation is as follows: (1) Content loss and style loss The content loss ensures that the content of the video frames remains consistent and avoids the destruction of the original image content due to style transfer. This is usually achieved by calculating the feature difference between the generated image and the reference image. The content loss extracts high-level features through a pre-trained VGG network to ensure that the semantic content of the generated image is consistent with the original video, where l represents the feature extraction layer in VGG-19. The content loss defined at layer l is the mean square error between the feature map of the input frame xt and its stylized output frame xt: ; in, represents the feature map of layer l, is the dimension of the feature map of layer l. This content loss is motivated by the observation that high-level features learned by CNNs represent abstract content, which is what we intend to preserve for the original input in the style transfer task. The content loss computes the difference between image features and encourages the generated image to preserve the high-level structure of the input image. In this way, it is ensured that the generated video does not lose important content in the original image.

[0051] The style loss ensures that the generated video is consistent with the target style. This is achieved by measuring the difference between the generated image and the target image in the style space, usually by calculating the Gram matrix to measure the style similarity of the image: ; in is the feature map. The style loss is defined as the style image s and the stylized output frame The mean squared error between the Gram matrices of : ; By comparing the style features of the generated image and the target image, the style loss can effectively keep the artistic style features of the image unchanged, so that the generated image or video not only meets the requirements in terms of content, but also conveys a specific artistic style or visual effect.

[0052] (2) Timing consistency loss Temporal consistency loss measures the consistency of image content between video frames, especially the differences between adjacent frames. It calculates the pixel differences or feature differences between consecutive frames, and maintains the coherence of video content by minimizing these differences, preventing mutations or jumps between frames. Its core idea is that there should be a visually smooth transition between frames in a video, maintaining stable content and motion changes, rather than sudden drastic changes. In the process of style transfer, unnatural flickers, dislocations or jumps may be introduced due to stylized processing, and this discontinuity will affect the viewing experience of the video. Therefore, through temporal consistency loss, these unnatural changes can be reduced and the coherence of the video can be maintained. The present invention optimizes the style transfer network by calculating the MSE between the previous and next frames and using it as the temporal consistency loss: ; in and are the i-th features extracted from the network for the t-th frame and the t+1-th frame respectively. In temporal consistency loss, MSE calculates the difference between the current frame and the previous frame, measuring their consistency in content. A larger MSE value indicates a larger difference between the two frames, which may appear as flickering, jumping, or other unnatural changes, while a smaller MSE value indicates that the content of the two frames is consistent and the visual transition is smoother.

[0053] (3) Optical flow smoothing loss Although temporal consistency loss helps ensure the coherence of visual content between adjacent frames, it does not directly control the smoothness of motion between video frames. For example, if the optical flow in a video changes dramatically, it may cause the image to jump noticeably in time, even though the content itself remains consistent. The introduction of optical flow smoothness loss is precisely to ensure that the motion between adjacent frames remains smooth and consistent.

[0054] Optical flow is a vector field that describes the change of each pixel in an image over time. It can represent the direction and speed of the image's movement between different time points. In video style transfer, the goal of optical flow smoothness loss is to constrain the change of the optical flow field so that the pixel motion changes between consecutive frames are smooth and consistent, thereby avoiding drastic motion changes or incoherent motion in video generation. Optical flow smoothness loss constrains the optical flow field between adjacent frames so that the motion vectors between adjacent frames remain consistent without unnatural jumps or instantaneous changes. The formula is as follows: ; It represents the gradient of the optical flow field u with respect to the spatial coordinate (x, y), and measures the change of the optical flow field in space. Represents the difference between the optical flow field at time t and t+1, and measures the change of the optical flow field over time. is a hyperparameter used to balance the importance of spatial and temporal smoothing terms.

[0055] The optical flow smoothness loss is optimized together with other losses to ensure that each generated frame not only has the target style, but also has no drastic motion changes in the timing of the video. By constraining the smoothness of the optical flow between video frames, the generated image will be more natural and consistent, avoiding discontinuous motion caused by style transfer.

[0056] In the present invention, the network structure of the present invention aims to achieve smooth transition between video frames and ensure the consistency of style through the collaborative work of multiple modules. First, the network receives the current frame and the previous frame as input. The RAFT model is used to calculate the optical flow field between the two frames to capture the pixel-level motion information. Then, the calculated optical flow field is processed by the optical flow smoothing module to eliminate noise and inconsistency, making the motion trajectory smoother and more natural.

[0057] Next, the network uses a feature extractor to extract deep features of the current frame and the distorted previous frame, identifying subtle differences and commonalities between them. By calculating the mean square error (MSE) between the features of the two frames, the network can quantify their differences and optimize the current frame through backpropagation to make its style and content more consistent with the previous frame. This optimization process ensures that the visual transition between video frames is more natural, avoids drastic changes in style, and ultimately generates an optimized output frame. The entire process not only improves the coherence of the video content, but also effectively reduces the error in optical flow estimation, ensuring a smooth transition between frames in tasks such as style transfer and video synthesis. The network structure is as follows Figure 4 shown.

[0058] The video style transfer algorithm of the present invention is based on the Pytorch deep learning framework, completed on the RTX4070 graphics card of python version 3.11, and trained for 100 cycles. The transfer methods AdaIn, NeuralNeighborStyleTransfer, and SANET are selected for comparison with the algorithm of this paper.

[0059] Objective indicator analysis, draw the RGB histogram of each method, and the RGB histogram overlay of each method is as follows Figure 5-Figure 8 As shown in the histogram, it can be seen that the RGB distribution of the results of the present invention is basically consistent between the previous and next frames. At the same time, the method of the present invention also presents a long-term style consistent result in the visual comparison diagram.

[0060] The results show that the algorithm of the present invention has obtained the best scores in deformation error, time error and peak signal-to-noise ratio indicators, especially better than other algorithms in time error and peak signal-to-noise ratio. The lower time error indicates that the algorithm of the present invention can better maintain continuity and dynamic consistency when processing image sequences, while the higher peak signal-to-noise ratio means that the quality difference between the reconstructed image and the original image is smaller, and the details and quality of the image are better preserved. This shows that the algorithm of the present invention is not only superior to other methods in terms of single-frame image quality, but also can better handle changes and dynamics in time series, providing more accurate and clear results.

[0061] By analyzing the objective evaluation indicators and subjective visual perception of multiple algorithms, the superiority of the algorithm in migrating video results is verified. The method introduced in the present invention not only shows good results in terms of subjective feeling and details in video style transfer, but also effectively reduces the occurrence of flickering.

[0062] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A video style transfer method based on RAFT optical flow, characterized in that: The following steps are involved: The RAFT optical flow network is used to calculate the pixel displacement of adjacent frames in the video frame by frame, generating an optical flow field containing the direction and size of each pixel's movement. Generate a dynamic weight map based on the displacement of pixels in the optical flow field, and the weight value of each area is inversely proportional to the displacement; Decompose the video frame into a low-frequency structure layer and a high-frequency detail layer through filtering, perform global style transfer on the low-frequency structure layer, and replace the color distribution and brushstroke features; Local stylization of high-frequency detail layers combined with dynamic weight maps; The processed low-frequency structure layer and high-frequency detail layer are fused to generate a single-frame stylized image; the stylized result of the current frame is motion-aligned with the adjacent frames according to the optical flow field, and the pixel position is adjusted through back-projection and interpolation compensation; The aligned multi-frame results are temporally smoothed and filtered to output a coherent stylized video.

2. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The generation process of the dynamic weight graph includes: The dynamic judgment threshold is set based on the overall displacement distribution of the optical flow field, and the area exceeding the dynamic judgment threshold is marked as a high dynamic area; the weight value of the high dynamic area is nonlinearly attenuated, and the attenuation degree gradually increases with the increase of displacement; A morphological expansion operation is performed on the boundary area of ​​the weight map to smooth the transition boundary between the low-weight area and the high-weight area; the morphological expansion operation expands the coverage of the low-weight area through an expansion algorithm to avoid weight jumps at the motion boundary; An adaptive Gaussian kernel is used to smooth the weight map. The radius of the Gaussian kernel is dynamically adjusted according to the severity of the local displacement change to ensure a natural transition of the weight distribution. The radius of the adaptive Gaussian kernel is calculated as follows: according to the displacement standard deviation of the local area in the optical flow field, the larger the standard deviation, the larger the Gaussian kernel radius, and the smaller the standard deviation, the smaller the radius.

3. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The decomposition process of the low-frequency structure layer and the high-frequency detail layer is as follows: The original frame is blurred by a Gaussian filter to extract a low-frequency structure layer; the high-frequency detail layer is obtained by subtracting the original frame from the low-frequency structure layer, and the negative value area of ​​the high-frequency detail layer is truncated and compensated; the truncation and compensation operation sets the negative value of the high-frequency detail layer to zero and then superimposes the fixed ratio intensity of the original detail; Performing contrast stretching on the high-frequency detail layer, wherein the contrast stretching expands the pixel value distribution of the high-frequency detail layer to a preset range through linear mapping, thereby improving the editability of the texture, and the stretching range is adaptively adjusted according to the pixel distribution of the detail layer; In the global style transfer of the low-frequency structure layer, the original geometric structure is retained and only the color and stroke features are replaced.

4. The video style transfer method based on RAFT optical flow according to claim 3, characterized in that: The local stylization of the high-frequency detail layer includes: Based on the preset weight threshold, the areas in the video frame are divided into low-weight areas, medium-weight areas and high-weight areas. For the high-frequency details in the low-weight areas, the stylization operation is limited to the luminance channel, and the original value of the chrominance channel is maintained; for the high-frequency details in the medium-weight areas, a multi-scale fusion strategy is adopted to superimpose coarse-grained brushstrokes and fine-grained textures respectively; for the high-frequency details in the high-weight areas, the style texture is fully applied, and the contour sharpness is maintained through the edge preservation algorithm; The multi-scale fusion strategy constructs style feature pyramids of different scales and superimposes style textures layer by layer from coarse to fine; The edge preservation algorithm detects the gradient information of the high-frequency detail layer and preserves the original contour of the area with a gradient higher than a set value during the stylization process.

5. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The fusion process of the low-frequency structure layer and the high-frequency detail layer includes: Dynamically adjust the feature fusion ratio between the low-frequency structure layer and the high-frequency detail layer according to the motion information of the optical flow field. For areas where the displacement in the optical flow field exceeds the dynamic judgment threshold, increase the weight ratio of the high-frequency detail layer in the fusion to the set ratio; for areas where the displacement is lower than the dynamic judgment threshold, increase the fusion weight ratio of the low-frequency structure layer to the set ratio. Construct a cross-layer feature correlation map, calculate the motion correlation between the low-frequency structure layer and the high-frequency detail layer of adjacent frames through the optical flow field, and use the same fusion ratio for high-correlation areas in the correlation map; The fused results are verified for consistency across frames, and the fusion ratio of the current frame is projected to the adjacent frames along the optical flow field to ensure the transition of high and low frequency features between multiple frames.

6. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: The motion alignment specifically includes: The optical flow field is checked for bidirectional consistency. The cyclic error between the forward optical flow and the backward optical flow is calculated. The area where the error exceeds the preset threshold is marked as an unreliable area, and vice versa. The stylized features of the unreliable area are filled with weighted fusion of adjacent frames, and the weight is determined by the local smoothness of the optical flow field. The stylized features of the reliable area are aligned at the sub-pixel level, and the interpolation weight is adjusted by the fractional displacement value of the optical flow field. The sub-pixel alignment uses a bilinear interpolation algorithm to adjust the feature interpolation weights according to the fractional displacement value of the optical flow field so that the projection result matches the pixel grid of the target frame.

7. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: Also includes static area optimizations: Detect the area in the video where the displacement of multiple consecutive frames is lower than the dynamic judgment threshold and mark it as a static area; otherwise, it is a dynamic area, and high-precision style transfer is performed on the first frame of the static area. The subsequent frames reuse the first frame result and perform fine-tuning through the optical flow field; the fine-tuning operation adjusts the pixel position of the reused area according to the displacement of the optical flow field to compensate for camera jitter or lighting changes; Gradual blending is performed at the junction of the static area and the dynamic area, and the stylized intensity is transitioned by transparency superposition; the gradual blending forms a stylized intensity change at the junction through a linear transition of the alpha channel.

8. The video style transfer method based on RAFT optical flow according to claim 1, characterized in that: It also includes identifying key target areas in the video through pre-trained models; setting a lower weight limit for the moving parts of the key target areas to limit the stylization intensity; setting an upper weight limit for the static parts of the key target areas to retain local details; The pre-trained model adopts a semantic segmentation algorithm based on a convolutional neural network to output pixel-level target category labels; the setting of the lower and upper weight limits is dynamically adjusted according to the importance of the target category, and the importance of the target category is obtained based on a preset target mapping table.

Citation Information

Patent Citations

  • Motion guide mask method and pre-training method of visual transformer model

    CN116168331A

  • Video denoising model processing method and apparatus, computer device, and storage medium

    WO2024217164A1

Cited By

  • Low-illumination image enhancement method and device

    CN121280302A