Video compression intelligent preprocessing method, system and device based on frequency domain perception optimization and medium

By introducing the frequency domain perception module and the optical flow collaboration module in the training stage of intelligent video compression, the problems of unbalanced video effects and damage to the frequency domain context in the prior art are solved, and efficient video preprocessing and compression effects are achieved.

CN120111245APending Publication Date: 2025-06-06XIDIAN UNIV

Patent Information

Application Number
CN202510262673.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video intelligent compression methods lead to unbalanced video effects when preprocessing and optimizing unimportant image frames, especially when complex or dense images are classified as unimportant frames, lack of optimization leads to degradation of picture quality. In addition, the perception-based video encoding preprocessing method ignores the modeling of the residual frame frequency domain information context, destroys the frequency domain context correlation across the transformation unit, and reduces the actual deployment effect.

Method used

By introducing a frequency domain perception module into the virtual encoder in the training stage, the context of the frequency domain information is modeled to eliminate statistical redundancy between transformation units; at the same time, an optical flow collaboration module is introduced to use high-level semantic information to repair the local distortion of the predicted frames in the feature domain, reducing the gap between training and actual deployment.

Benefits of technology

High-precision estimation of video motion information, predicted frame artifact repair and frequency domain statistical redundancy removal are realized, which significantly improves the performance and compression efficiency of video preprocessing, ensuring that video quality remains high in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111245A_ABST
    Figure CN120111245A_ABST
Patent Text Reader

Abstract

The invention discloses a video compression intelligent preprocessing method, system and device based on frequency domain perception optimization and a medium. The method comprises the following steps: constructing a preprocessing model; preprocessing the input coding frame and the reference frame; enabling the preprocessed coding frame and reference frame to pass through a virtual encoder to obtain a frequency domain coefficient of a residual frame, and calculating a code rate of the coding frame; enabling the frequency domain coefficient of the residual frame to pass through a virtual decoder to obtain a decoded frame; training a preprocessing model; testing the effect; the system, the equipment and the medium are used for implementing the method. According to the method, a frequency domain sensing module is introduced into a virtual encoder in a training stage, modeling is carried out on a frequency domain information context, and statistical redundancy between transformation units is eliminated; an optical flow collaboration module is introduced into a virtual encoder, local distortion of a prediction frame is repaired in a feature domain by using high-level semantic information, and the difference between training and actual deployment is effectively reduced; through the optimization, the preprocessing model can better learn the rate distortion characteristic of the standard encoder, and the quality of the preprocessed video is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video compression technology, and in particular to a video compression intelligent preprocessing method, system, device and medium based on frequency domain perception optimization. Background Art

[0002] As the demand for video content continues to grow, network bandwidth and storage space are facing increasing pressure. To meet these challenges, video compression technology is constantly improving, aiming to reduce bitrate by improving compression efficiency while maintaining or improving video quality. Video preprocessing technology refers to a series of processing operations performed on the video before encoding, such as denoising, adjusting brightness, etc., aiming to improve the efficiency of the encoding process and the final quality of the video.

[0003] The effect of using hand-designed filters to reduce video noise before encoding is poor and is no longer a mainstream method. Recently, preprocessing methods based on deep learning can optimize video content before encoding and improve the final visual quality of the encoded video. The video intelligent compression method based on image analysis uses neural networks to comprehensively consider multiple factors such as application scenarios, image features, and network status to determine the compression rate of the video. Further, based on the theme consistency, visual quality, and content prominence analysis of each image frame in the video, frames with low application value are screened out for contrast adjustment and their bit rate allocation is reduced, thereby retaining the clarity and details of key scenes in the video. The perception-based video coding preprocessing method optimizes the encoded bit rate while ensuring image quality by designing a loss function that includes fidelity and bit rate, ensuring that the compressed video effect is consistent with human visual perception. These intelligent preprocessing technologies can be closely integrated with the current popular video coding standards (such as H.264 / H.265, etc.) to significantly improve the performance and effect of video coding. However, these methods have some limitations: the video intelligent compression method based on image analysis only performs pre-processing optimization on unimportant image frames, which may lead to uneven final video effects, especially when some pictures with complex or dense details are classified as unimportant frames. The lack of optimization of these frames may lead to a decrease in picture quality; the perception-based video coding preprocessing method ignores the modeling of the frequency domain information context of the residual frame in the design of the training loop. This method adopts an independent coding strategy for the transform units in the frequency domain of the residual frame, which destroys the frequency domain context correlation across transform units, making it impossible to jointly model and compress redundant components, reducing the effect of using it in combination with the standard encoder in actual deployment.

[0004] The patent document with publication number CN118055235A discloses a video intelligent compression method based on image analysis. It determines the compression rate of the video by comprehensively considering multiple factors such as application scenarios, image features and network status, breaking through the limitations and one-sidedness of the prior art that only relies on application scenarios to determine the compression rate. However, since only unimportant image frames are preprocessed and optimized, the final video effect is uneven. For example, the optimization effect of unimportant frames in certain scenes may be better, but the visual experience of the entire video may be affected by incomplete optimization. In particular, when some pictures with more complex or dense details are classified as unimportant frames, the lack of detailed optimization of these frames may lead to a decrease in picture quality. Summary of the invention

[0005] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a video compression intelligent preprocessing method, system, device and medium based on frequency domain perception optimization, by introducing a frequency domain perception module in the virtual encoder in the training stage, the context of the frequency domain information is modeled to eliminate the statistical redundancy between the transformation units; at the same time, an optical flow collaboration module is introduced in the virtual encoder, and high-level semantic information is used to repair the local distortion of the predicted frame in the feature domain, thereby effectively reducing the gap between training and actual deployment. Through the above optimization, the preprocessing model can better learn the rate-distortion characteristics of the standard encoder, thereby achieving a significant improvement in the quality of the preprocessed video.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A video compression intelligent preprocessing method based on frequency domain perception optimization includes the following steps:

[0008] Step 1: Build a preprocessing model;

[0009] Step 2: preprocess the input coded frame and the reference frame using the preprocessing model constructed in step 1 to obtain the preprocessed coded frame and the reference frame;

[0010] Step 3: Pass the coded frame and the reference frame preprocessed in step 2 through a virtual encoder to obtain the frequency domain coefficients of the residual frame and calculate the coded frame bit rate;

[0011] Step 4: Pass the frequency domain coefficients of the residual frame obtained in step 3 through a virtual decoder to obtain a decoded frame;

[0012] Step 5: Use the coded frame rate calculated in step 3 and the decoded frame obtained in step 4 to train the preprocessing model constructed in step 1;

[0013] Step 6: Use the preprocessing model trained in step 5 to test the effect.

[0014] The specific method of step one includes:

[0015] The preprocessing model is built based on convolutional neural networks and spatial attention mechanisms, which is used to optimize the perception of input frames and improve the human eye perception quality of encoded videos; it includes: feature extraction convolutional layers, multiple stacked residual local feature blocks, 3×3 convolutional layers and reconstruction modules; among them, the feature extraction convolutional layers are composed of 3×3 convolutional layers; the residual local feature blocks include three layers of 3×3 convolutional layers, one 1×1 convolutional layer and an enhanced spatial attention mechanism; the three layers of 3×3 convolutional layers are composed of convolutional layers, activation functions and skip connections; the enhanced spatial attention mechanism includes the first 1×1 convolutional layer, 3×3 convolutional layers, maximum pooling, multiple 3×3 convolutions, upsampling, the second 1×1 convolutional layer, activation functions and skip connections; the reconstruction module is composed of 1×1 convolutional layers;

[0016] The input frame enters the preprocessing model, and first obtains coarse-grained features through feature extraction convolution. Then, multiple stacked residual local feature blocks are used to perform fine-grained learning optimization on the coarse-grained features to obtain fine-grained features. The fine-grained feature map is smoothed using a 3×3 convolutional layer; finally, the reconstruction module is used to map the fine-grained features to the pixel domain.

[0017] The specific method of step 2 includes:

[0018] Step 2.1, performing color space conversion on the input coded frame and the reference frame;

[0019] According to the conversion relationship between RGB and YUV color spaces, the color space conversion is first performed on the coded frame and the reference frame in the input RGB color space, and the RGB format is converted into the YUV format. Then, the converted YUV data is split into independent brightness component Y, chrominance component U and chrominance component V, where the brightness component Y represents the brightness information of the image, and the chrominance component U and chrominance component V represent the chrominance information of the image respectively;

[0020] Step 2.2, extracting the features of the brightness component Y of the coded frame and the reference frame;

[0021] In the preprocessing stage, the luminance component Y of the coded frame and the reference frame is first taken as input and passed to the preprocessing model. The preprocessing model performs preliminary feature extraction on the input coded frame and the reference frame respectively through the feature extraction convolution layer to capture the spatial structure and texture information in the frame; then, multiple stacked residual local feature blocks are used to refine the feature representation of the spatial structure and texture information, and a 3×3 convolution layer is used for smoothing operation to extract local detail information in the frame and retain high-dimensional features;

[0022] Step 2.3: Image reconstruction;

[0023] The high-dimensional features extracted and refined in step 2.1 are input into the reconstruction module, and are reconverted to the pixel domain using a 1×1 convolution operation mapping to perform image reconstruction to obtain the preprocessed coded frame and reference frame.

[0024] The specific method of step three includes:

[0025] Step 3.1, using the feature domain-based optical flow estimation on the coded frame and the reference frame preprocessed in step 2.3 to obtain an optical flow field;

[0026] Step 3.2, use the optical flow collaborative module to cooperate with the current frame to obtain the residual frame;

[0027] The optical flow collaborative module includes optical flow distortion and prediction frame post-processing. The optical flow distortion is used to map the pixel value in the reference frame pre-processed in step 2.3 to the corresponding position of the coded frame pre-processed in step 2.3 according to the motion vector in the optical flow field obtained in step 3.1 to generate a prediction frame; the prediction frame is input into the prediction frame post-processing for artifact repair to obtain a high-quality prediction frame; the coded frame pre-processed in step 2.3 is subtracted pixel by pixel from the high-quality prediction frame after artifact repair to obtain a residual frame;

[0028] Step 3.3, performing discrete cosine transform on the residual frame obtained in step 3.2 to obtain frequency domain coefficients of the residual frame;

[0029] Step 3.4, using a frequency domain perception module to remove the statistical correlation between the transform units from the frequency domain coefficients of the residual frame obtained in step 3.3;

[0030] The frequency domain perception module is composed of a feature extraction module, a linear mapping, a Transformer encoder, and a transformation unit reconstruction module. The feature extraction module is composed of a two-layer convolutional neural network. The Transformer encoder is composed of a multi-head self-attention mechanism, a first layer normalization layer, a first jump connection, a feedforward neural network layer, a second layer normalization layer and a second jump connection. The transformation unit reconstruction module is composed of a linear transformation and a deconvolution network.

[0031] Step 3.5, quantizing the frequency domain coefficients from which statistical correlation is removed in step 3.4;

[0032] Step 3.6: Input the frequency domain coefficients quantized in step 3.5 into the entropy model, estimate the probability distribution of the frequency domain coefficients, and calculate the coding frame rate.

[0033] The specific method of step 4 includes:

[0034] Step 4.1, dequantize and inverse discrete cosine transform the quantized frequency domain coefficients obtained in step 3.5 to obtain a decoded residual frame;

[0035] Step 4.2: Add the decoded residual frame obtained in step 4.1 to the input reference frame in step 2 to obtain a decoded frame.

[0036] The specific method of step five includes:

[0037] Design the loss function to train the preprocessing model constructed in step 1;

[0038] The loss function consists of bit rate loss and fidelity loss; the bit rate loss uses the coded frame bit rate obtained in step 3.6 as the loss, and the fidelity loss is calculated using the input coded frame and the decoded frame obtained in step 4; the formula is as follows:

[0039]

[0040] Among them, M represents The true distribution of represents the frequency domain coefficient; Indicates L 1 loss, Represents the multi-scale structural similarity loss; α and β are hyperparameters used to weigh L 1 and MS-SSIM loss, and α+β=1; in the fidelity loss, use L 1 The loss focuses on precise pixel-level alignment to ensure accurate restoration of the optimized video frame. MS-SSIM is used to ensure that the structural information of the optimized video frame remains unchanged from the original video frame. The final loss function is as follows:

[0041] L = l F +λl R

[0042] Among them, λ is a hyperparameter used to determine the proportion of bit rate loss.

[0043] The specific method of step six includes:

[0044] The preprocessed encoded video and the unpreprocessed encoded video are compared and analyzed using Video Multi-method Assessment Fusion (VMAF) and Structural Similarity (SSIM), and the VMAF and SSIM scores of the two are calculated and compared. The higher the score, the higher the video quality.

[0045] The present invention also provides a video compression intelligent preprocessing system based on frequency domain perception optimization, comprising:

[0046] A preprocessing model building module, used to build a preprocessing model;

[0047] A coding frame and reference frame preprocessing module, used for preprocessing the input coding frame and reference frame using a preprocessing model to obtain a preprocessed coding frame and reference frame;

[0048] The residual frame frequency domain coefficient and coding frame bit rate acquisition module is used to pass the pre-processed coding frame and the reference frame through the virtual encoder to obtain the residual frame frequency domain coefficient and calculate the coding frame bit rate;

[0049] A decoded frame acquisition module is used to obtain a decoded frame by passing the frequency domain coefficients of the residual frame through a virtual decoder;

[0050] The preprocessing model training module uses the encoding frame rate and the decoding frame to train the preprocessing model;

[0051] The preprocessing model effect testing module is used to perform effect testing using the trained preprocessing model.

[0052] The present invention also provides a video compression intelligent preprocessing device based on frequency domain perception optimization, comprising:

[0053] Memory: a computer program storing the above-mentioned video compression intelligent preprocessing method based on frequency domain perception optimization, which is a computer-readable device;

[0054] Processor: used to implement the video compression intelligent preprocessing method based on frequency domain perception optimization when executing the computer program.

[0055] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the video compression intelligent preprocessing method based on frequency domain perception optimization.

[0056] Compared with the existing pretreatment technology, the present invention has the following advantages:

[0057] 1. The present invention obtains the optical flow field by introducing the optical flow estimation based on the feature domain, which enhances the ability to resist illumination changes and provides higher optical flow field estimation accuracy and robustness. This innovation can effectively improve the effect of video preprocessing, so that the encoded video can still maintain high-quality visual performance in complex scenes.

[0058] 2. The present invention designs an optical flow collaborative module to repair artifacts in the predicted frames after optical flow distortion, thereby improving the quality and robustness of the predicted frames in complex scenes. Specifically, the optical flow collaborative module uses high-level semantic information to repair local distortions of the predicted frames in the feature domain, especially for artifacts caused by inaccurate optical flow estimation or motion boundary areas. The optical flow collaborative module can effectively identify and repair distorted areas in the predicted frames, such as motion blur, block effects, and texture loss, by combining the motion information of the optical flow field and the contextual features of the input frame. The optical flow collaborative module can significantly reduce the distortion of the predicted frames, so that the preprocessed video maintains a higher visual quality after encoding.

[0059] 3. The present invention designs a frequency domain perception module based on the Transformer encoder to perform context modeling on the frequency domain information of the residual frame and remove statistical correlation, thereby obtaining more realistic frequency domain information. This technical point makes the frame bit rate calculated by the entropy model closer to the result of the standard encoder, significantly reduces the gap between the virtual encoder and the standard encoder, and enhances the effect of the preprocessing model on the quality of the encoded video in actual deployment. Through this innovation, the present invention can more accurately simulate the rate-distortion characteristics of the standard encoder and further improve the efficiency and quality of video compression.

[0060] In summary, the present invention realizes high-precision estimation of video motion information, restoration of prediction frame artifacts, and removal of frequency domain statistical redundancy by introducing feature domain-based optical flow estimation, designing an optical flow collaboration module, and a frequency domain perception module based on a Transformer encoder. The feature domain optical flow estimation algorithm enhances the ability to resist illumination changes and provides more accurate motion information estimation; the optical flow collaboration module uses high-level semantic information to repair the predicted frames after optical flow distortion, significantly reducing distortion problems such as motion blur, block effects, and texture loss; the frequency domain perception module based on Transformer removes frequency domain statistical redundancy through context modeling, making the frame rate calculated by the entropy model closer to the result of the standard encoder. The present invention significantly improves the performance of video preprocessing, and can still maintain high-quality visual performance in complex scenes, while improving compression efficiency and quality, providing strong technical support for high-quality video compression. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic flow chart of the implementation method of the present invention.

[0062] Figure 2 is a network structure diagram of the preprocessing module in the present invention, wherein: Figure 2 (a) is the overall structure diagram of the preprocessing model. Figure 2 (b) is the residual local feature block network structure diagram, Figure 2 (c) is the network structure diagram of the enhanced spatial attention mechanism.

[0063] Figure 3 It is a structural diagram of the frequency domain perception module design in the present invention.

[0064] Figure 4 It is a graph of the experimental results of the present invention. DETAILED DESCRIPTION

[0065] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0066] The main technical idea of ​​the present invention is: use the preprocessing module to adaptively optimize the input frame, and then pass it through a virtual encoder constructed by motion estimation, motion compensation, transformation, quantization, entropy model and other technologies to obtain the input frame bit rate; after passing through the virtual encoder, use a virtual decoder constructed by inverse quantization, inverse transformation and other technologies to add it to the reference frame to obtain a reconstructed frame. The network parameters of the preprocessing module are optimized by the bit rate loss and distortion loss in the loss function. The evaluation indicators (Video Multi-Method Assessment Fusion (VMAF) and Structural Similarity (SSIM)) scores are used to evaluate the quality of the encoded video after preprocessing and compared with the encoded video without preprocessing.

[0067] like Figure 1 As shown, a video compression intelligent preprocessing method based on frequency domain perception optimization includes the following steps:

[0068] like Figure 2 As shown, step 1, building a preprocessing model;

[0069] The preprocessing model is built based on convolutional neural networks and spatial attention mechanisms to optimize the perception of input frames and improve the human eye perception quality of the encoded video; Figure 2 As shown in (a), it includes: a feature extraction convolution layer, multiple stacked residual local feature blocks, a 3×3 convolution layer and a reconstruction module; wherein the feature extraction convolution layer is composed of a 3×3 convolution layer; as shown in Figure 2 As shown in (b), the residual local feature block includes three 3×3 convolutional layers, one 1×1 convolutional layer and an enhanced spatial attention mechanism; the three 3×3 convolutional layers are composed of convolutional layers, activation functions and skip connections; Figure 2 As shown in (c), the enhanced spatial attention mechanism includes the first 1×1 convolution layer, 3×3 convolution layer, maximum pooling, multiple 3×3 convolutions, upsampling, the second 1×1 convolution layer, activation function and skip connection; the reconstruction module consists of a 1×1 convolution layer;

[0070] The input frame enters the preprocessing model, and first obtains coarse-grained features through the feature extraction convolution layer; then multiple stacked residual local feature blocks are used to perform fine-grained learning optimization on the coarse-grained features, that is, the three layers of 3×3 convolution layers in the residual local feature block are first used to progressively optimize the coarse-grained features and use jump connections to prevent the features from disappearing, and then the 1×1 convolution layer is used to increase the network's expressive power and nonlinear characteristics, and then the enhanced spatial attention mechanism is used to ensure that key feature information is paid attention to. Among them, the enhanced spatial attention mechanism uses the first 1×1 convolution layer to adjust the channel of the input feature map, and then uses a 3×3 convolution layer with a step size of 2 and a maximum pooling operation to extract global significant features, and then Then, local features are extracted through multiple 3×3 convolutions, and the feature map is restored to the original spatial size using an upsampling operation based on bilinear interpolation. The input feature map after channel adjustment is added using a skip connection. Then, the second 1×1 convolution layer is used to increase the network's expressiveness and nonlinear characteristics. The attention weight map is obtained after an activation function. The input feature map and the attention weight map are multiplied element by element to obtain a fine-grained feature map. The fine-grained feature map is smoothed using a 3×3 convolution layer, and then added to the coarse-grained features using a skip connection to prevent the features from disappearing. Finally, the fine-grained feature map is mapped to the pixel domain using the 1×1 convolution layer of the reconstruction module to obtain the preprocessed input frame.

[0071] Step 2: preprocess the input coded frame and the reference frame using the preprocessing model constructed in step 1 to obtain the preprocessed coded frame and the reference frame;

[0072] Step 2.1, performing color space conversion on the input coded frame and the reference frame;

[0073] According to the conversion relationship between RGB and YUV color space, the color space conversion is first performed on the coded frame and reference frame under the input RGB color space, and it is converted from RGB format to YUV format. Then, the converted YUV data is further split into independent brightness component Y, chrominance component U and chrominance component V, where the brightness component Y represents the brightness information of the image, and the chrominance component U and chrominance component V represent the chrominance information of the image respectively; through this conversion and splitting operation, the color information of the image can be separated into two parts, brightness and chrominance, providing a more suitable input data format for video compression processing for the subsequent preprocessing model, and facilitating more refined feature extraction and optimization processing for the brightness component Y. The brightness component Y required by the preprocessing model is obtained, and its formula is as follows:

[0074] Y=0.299R+0.587G+0.114B

[0075] U=-0.14713R-0.28886G+0.436B

[0076] V=0.615R-0.51499G-0.10001B

[0077] Step 2.2: Use the feature extraction convolution layer of the preprocessing model to extract features from the input coded frame and reference frame respectively. The specific network structure is as follows: Figure 2 As shown; use X t and P t As the input and output of the preprocessing model; first, a single 3×3 convolutional layer is used as a feature extraction convolution to extract the coarse-grained features of the image, and its formula is as follows:

[0078] F 0 =h ext (X t )

[0079] Among them, h ext (·) represents the convolution operation for feature extraction, F 0 is the extracted feature map; then, n residual local feature blocks are used in a cascaded manner to extract deep features. The process formula is:

[0080]

[0081] in, represents the nth residual local feature block, F n is the nth output feature map, and n=3 in the present invention. The residual local feature block is the core module of the preprocessing model, and the content of the coarse-grained features is deeply optimized by stacking multiple residual local feature blocks. Through this multi-level feature extraction mechanism, the preprocessing model can fully exploit the spatial correlation between the coded frame and the reference frame. Each residual local feature block is used to capture local features and generate dynamic attention weights. The image processing process can be expressed as the following formula:

[0082] F refined =ESA(Conv 1×1 (F conv +F input ))

[0083] Among them, F input represents the input features of the residual local feature block, F refined represents the output of the residual local feature block, F conv Represents the features extracted by convolution, Conv 1×1 Represents a 1×1 convolution operation. ESA represents the dynamic weighting of fused features to generate salient feature regions. This module optimizes input features by focusing on salient regions, effectively improving performance in complex dynamic scenes.

[0084] The process of enhancing the spatial attention mechanism to process features is as follows. First, the enhanced spatial attention mechanism module uses the first 1×1 convolution layer to adjust the channel of the input feature map, and then uses a 3×3 convolution layer with a stride of 2 and a maximum pooling operation to extract global salient features. After that, multiple 3×3 convolutions are used to extract local features, and the feature map is restored to the original spatial size using an upsampling operation based on bilinear interpolation. The formula is as follows:

[0085] F local =Conv(MaxPooling(Conv(Conv 1×1 (F in ))))

[0086] Among them, F in represents the input feature map, F local Represents local information, Conv represents a 3×3 convolution operation, MaxPooling represents a maximum pooling operation, Conv 1×1 Represents a 1×1 convolution operation. Then, the extracted local saliency features are restored to the original resolution through bilinear interpolation and fused with the input features after the first 1×1 convolution layer through the second 1×1 convolution layer to generate a dynamic attention weight map M. a , the formula is as follows:

[0087] M a =σ(Conv 1×1 (F up +Conv 1×1 (F in ))) (3-6)

[0088] Among them, σ represents the Sigmoid activation function, F up Represents the features after bilinear interpolation of local features, and the obtained attention weight map M a Used to dynamically weight the input feature map F in Finally, the attention weight map M a Multiply the input feature map pixel by pixel to enhance the features of the salient regions and suppress redundant information to obtain the output feature map F out , the formula is as follows:

[0089] F out =F in ⊙M a

[0090] A 3×3 convolutional layer is then used to smooth the gradually refined deep feature maps.

[0091] Step 2.3: Reconstruct the image based on the output feature map obtained in step 2.2. Image reconstruction in the feature space is conducive to preserving the key details of the image. The 1×1 convolutional layer of the reconstruction module is used to generate the final optimized image P. t , the formula is as follows:

[0092] P t =f rec ((f smooth (F n )+F 0 ))

[0093] Among them, f rec represents the reconstruction module, which consists of convolution operations, f smooth Represents a 3×3 convolution operation;

[0094] As an efficient linear transformation operation, 1×1 convolution can perform weighted combination of feature channels while maintaining spatial structure, thereby restoring abstract high-level semantic features to specific pixel values. This process can not only effectively integrate multi-channel feature information, but also retain image details and texture features during reconstruction, ensuring that the reconstructed image is as close as possible to the original input in visual quality.

[0095] Finally, the preprocessed coded frame and reference frame are obtained.

[0096] Step 3: Pass the coded frame and the reference frame preprocessed in step 2 through a virtual encoder to obtain the frequency domain coefficients of the residual frame and calculate the coded frame bit rate;

[0097] Step 3.1, use feature domain-based optical flow estimation on the coded frame and reference frame preprocessed in step 2.3 to obtain the optical flow field. Existing motion estimation methods use block matching algorithms or pixel domain optical flow estimation algorithms, which usually perform poorly when processing complex and fast motions, and have a large amount of calculations and limited accuracy. The present invention uses a Transfomer-based optical flow network to operate on the coded frame and the reference frame in the feature space to obtain motion information, and can more accurately capture complex motion patterns by extracting stable feature points for matching, and has strong resistance to noise and illumination changes, and is more adaptable to different image contents, providing higher estimation accuracy and robustness;

[0098] Step 3.2, use the optical flow collaborative module to cooperate with the current frame to obtain the residual frame;

[0099] The optical flow collaborative module includes optical flow distortion and prediction frame post-processing. The optical flow distortion is used to map the pixel values ​​in the reference frame pre-processed in step 2.3 to the corresponding positions of the coded frame pre-processed in step 2.3 according to the motion vector in the optical flow field obtained in step 3.1 to generate a prediction frame. In the optical flow distortion process, non-integer positions of pixels may be involved. In this case, an interpolation method (such as bilinear interpolation) is required to estimate the pixel values ​​of the corresponding positions to complete the smooth transition of the image. The prediction frame is input into the prediction frame post-processing for artifact repair to obtain a high-quality prediction frame. The artifacts and distortion generated during the optical flow distortion process are repaired to restore the image details and ensure that the prediction frame is more realistic and natural. The coded frame pre-processed in step 2.3 is subtracted pixel by pixel from the high-quality prediction frame after artifact repair to obtain a residual frame.

[0100] In this process, the optical flow field, as an estimation of spatiotemporal motion, describes the motion trajectory of each pixel in the image. By combining the reference frame preprocessed in step 2.3 with the optical flow field, the optical flow warp can perform spatiotemporal transformation on the reference frame preprocessed in step 2.3 to obtain a predicted frame; the predicted frame post-processing module focuses on repairing the artifact problem caused by the optical flow warp process. Based on the predicted frame generated by the traditional optical flow warp, the predicted frame post-processing module is no longer limited to post-processing in pixel space, but collaboratively optimizes multi-scale features through a convolutional neural network; the coded frame preprocessed in step 2.3 is subtracted pixel by pixel from the predicted frame after artifact repair to obtain a residual frame.

[0101] Step 3.3: Perform discrete cosine transform on the residual frame obtained in step 3.2 to obtain the frequency domain coefficients of the residual frame; its core function is to convert the pixel values ​​in the spatial domain to the frequency domain, and re-represent the information in the image or video in a more compact form. Through the transformation technology, the redundancy of the data can be significantly reduced, and most of the energy can be concentrated in a few transformation coefficients, thus creating conditions for subsequent quantization and entropy coding, and achieving a more efficient compression rate;

[0102] Step 3.4: Use the frequency domain perception module to model the context of the frequency domain coefficients of the residual frame obtained in step 3.3; Figure 3 As shown, the frequency domain perception module is mainly composed of four parts: feature extraction module, linear mapping, Transformer encoder, and transform unit reconstruction module. The feature extraction module is composed of two layers of convolutional neural network, the Transformer encoder is composed of a multi-head self-attention mechanism, the first layer normalization layer, the first jump connection, the feedforward neural network layer, the second layer normalization layer and the second jump connection, and the transform unit reconstruction module is composed of linear transformation and deconvolution network.

[0103] The frequency domain coefficients of the residual frame obtained in step 3.3 are input into the frequency domain perception module with a 4×4 transform unit. The features are first extracted through the two-layer convolutional neural network of the feature extraction module, and then linearly mapped to an embedding dimension suitable for the Transformer, and added to the position code and input into the Transformer encoder. The statistical correlation of the frequency domain coefficients is first removed through the multi-head self-attention layer, and then the input and output of the multi-head self-attention layer are input together to the first layer normalization layer to stabilize the gradient propagation and accelerate the model convergence. Then, the feedforward neural network layer is used to further enhance the nonlinear expression ability of the model, and then the input and output of the feedforward neural network layer are input together to the second layer normalization layer to stabilize the gradient propagation and accelerate the model convergence. Finally, the linear transformation of the transform unit reconstruction module is used to map the output size of the Transformer encoder to the feature map size before encoding, and then the deconvolution network corresponding to the feature extraction module is used to upsample the feature map to the original frequency domain coefficient transform unit size. The specific network structure is as follows Figure 3 As shown;

[0104] Step 3.5, quantizing the frequency domain coefficients from which statistical correlation is removed in step 3.4, so as to further compress the data volume and reduce the bit rate of storage and transmission; through quantization, the frequency domain coefficients are mapped to a lower numerical precision, most high-frequency components (such as image details and noise) are greatly compressed or even set to zero, and only low-frequency components (such as main structural information) are retained, thereby significantly reducing the data volume;

[0105] Step 3.6. The energy distribution of the quantized frequency domain coefficients obtained in step 3.5 has significant sparsity and non-uniformity. By using the entropy model to accurately model the probability distribution of the quantized frequency domain coefficients, the information entropy (i.e., the theoretical minimum bit rate) can be effectively estimated as the coding frame bit rate, which is used for the bit rate loss in the loss function to guide the training of the preprocessing model.

[0106] Step 4: Pass the frequency domain coefficients of the residual frame obtained in step 3 through a virtual decoder to obtain a decoded frame;

[0107] Step 4.1: Dequantize the quantized frequency domain coefficients obtained in step 3.5, and restore the approximate original value of the coefficient by multiplying the quantized frequency domain coefficients by the quantization step size to reconstruct the dynamic range of the frequency domain signal. Then perform an inverse discrete cosine transform operation to restore the frequency domain data to the pixel domain and generate a decoded residual frame;

[0108] Step 4.2: The decoded residual frame in step 4.1 is added to the input reference frame in step 2 to obtain a decoded frame. The decoded frame and the input coded frame are used to measure the distortion of the preprocessing module and can be used as the fidelity loss.

[0109] Step 5: Use the coded frame rate calculated in step 3 and the decoded frame obtained in step 4 to train the preprocessing model constructed in step 1;

[0110] Step 5.1, the loss function is divided into two parts: rate loss and fidelity loss, which are used to reduce the encoding rate under low distortion conditions. The rate loss uses the encoding frame rate obtained in step 3.6 as the loss, and the fidelity loss is calculated using the input encoding frame and the decoded frame obtained in step 4. The formula is as follows:

[0111]

[0112] Among them, M represents The true distribution of represents the frequency domain coefficient; Indicates L 1 loss, Represents the multi-scale structural similarity loss; α and β are hyperparameters used to weigh L 1 and MS-SSIM loss, and α+β=1; in the fidelity loss, use L 1 The loss focuses on precise pixel-level alignment to ensure accurate restoration of the optimized video frame. MS-SSIM is used to ensure that the structural information of the optimized video frame remains unchanged from the original video frame. The final loss function is as follows:

[0113] L = l F +λl R

[0114] Among them, λ is a hyperparameter used to determine the proportion of bit rate loss.

[0115] Step 6: Use the preprocessing model trained in step 5 to test the effect.

[0116] In order to comprehensively evaluate the quality of the pre-processed encoded video, two evaluation indicators, Video Multi-method Assessment Fusion (VMAF) and Structural Similarity (SSIM), are used for quantitative analysis. Video Multi-method Assessment Fusion (VMAF) is a comprehensive evaluation indicator based on machine learning. By integrating multiple visual quality features (such as detail loss, texture distortion, and motion blur, etc.), it can more accurately reflect the human eye's subjective perception of video quality; while Structural Similarity (SSIM) measures the similarity between the pre-processed video and the original video from three dimensions of brightness, contrast, and structure, and is particularly suitable for evaluating the structural fidelity of images.

[0117] In the experiment, the video multi-method assessment fusion (VMAF) and structural similarity (SSIM) were used to compare and analyze the preprocessed encoded video with the unpreprocessed encoded video. By calculating and comparing the video multi-method assessment fusion (VMAF) and structural similarity (SSIM) scores of the two, the effect of the preprocessing method in improving video quality can be quantified. The higher the score, the higher the video quality.

[0118] The present invention will be further described below in conjunction with experiments:

[0119] 1. Simulation experiment conditions:

[0120] The hardware environment of the simulation experiment of the present invention: Intel (R) Intel Core i9-12900KF CPU @ 3.20 GHz, 64 GB memory, RTX 3090 GPU; the software environment: Ubuntu 22.04 operating system, Python 3.8, Pytorch 1.12.0.

[0121] 2. Experimental content and results analysis:

[0122] The simulation experiment of the present invention performs a preprocessing effect test on the public video dataset UVG, and uses a rate-distortion curve based on the VMAF video evaluation index to evaluate the optimization effect of the present invention on the encoded video. At the same time, the experimental results in recent papers are cited and compared with the video preprocessing method proposed in the present invention to prove the advantages of the method of the present invention. The cited papers are shown in Table 1 below.

[0123] Table 1 is a list of papers citing the model

[0124] Model paper RPP Rate-PerceptionOptimizedPreprocessingforVideoCoding

[0125] The video compression intelligent preprocessing method proposed in this invention is used to conduct experiments on the public video dataset UVG. The results are as follows: Figure 4 As shown. In order to verify the performance advantages of the present invention, it is compared with the existing mainstream video preprocessing methods. The performance evaluation of video compression algorithms usually adopts rate-distortion performance analysis, and the rate-distortion curve is drawn to measure the algorithm's fidelity to video quality at different bit rates. The specific method is as follows: For the same video sequence, use multiple preset bit rates for compression, calculate the video quality index corresponding to each bit rate (such as video multi-method assessment fusion (VMAF)), and draw a mapping curve between bit rate and quality. If the rate-distortion curve of one algorithm is located in the upper left of another algorithm, it means that the algorithm can provide higher quality at the same bit rate, or can achieve reconstruction at a lower bit rate at the same quality. As Figure 4As shown in the figure, the rate-distortion curve of the present invention is located at the upper left of the comparison method, indicating that it significantly improves the video quality at the same bit rate, or saves the bit rate at the same quality. In addition, in the high bit rate area, the performance advantage of the present invention is more significant, which further verifies the robustness of the method of the present invention in complex scenes. Therefore, the preprocessing method proposed in the present invention is superior to the existing method in rate-distortion performance.

[0126] The present invention also provides a video compression intelligent preprocessing system based on frequency domain perception optimization, comprising:

[0127] A preprocessing model building module, used to implement the construction of the preprocessing model in step 1;

[0128] A coding frame and reference frame preprocessing module, used to implement step 2 to preprocess the input coding frame and reference frame using the preprocessing model constructed in step 1 to obtain preprocessed coding frame and reference frame;

[0129] The module for acquiring the frequency domain coefficient of the residual frame and the coding frame rate is used to implement the step 3 of passing the coding frame and the reference frame preprocessed in step 2 through a virtual encoder to obtain the frequency domain coefficient of the residual frame and calculate the coding frame rate;

[0130] A decoded frame acquisition module is used to implement step 4, passing the frequency domain coefficients of the residual frame obtained in step 3 through a virtual decoder to obtain a decoded frame;

[0131] A preprocessing model training module is used to implement the training of the preprocessing model constructed in step 1 in step 5 using the coded frame rate calculated in step 3 and the decoded frame obtained in step 4;

[0132] The preprocessing model effect testing module is used to implement the effect testing in step 6 using the preprocessing model trained in step 5.

[0133] The present invention also provides a video compression intelligent preprocessing device based on frequency domain perception optimization, comprising:

[0134] Memory: a computer program storing the above-mentioned video compression intelligent preprocessing method based on frequency domain perception optimization, which is a computer-readable device;

[0135] Processor: used to implement the video compression intelligent preprocessing method based on frequency domain perception optimization when executing the computer program.

[0136] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the video compression intelligent preprocessing method based on frequency domain perception optimization.

[0137] The above contents are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A video compression intelligent preprocessing method based on frequency domain perception optimization, characterized in that: The following steps are involved: Step 1: Build a preprocessing model; Step 2: preprocess the input coded frame and the reference frame using the preprocessing model constructed in step 1 to obtain the preprocessed coded frame and the reference frame; Step 3: Pass the coded frame and the reference frame preprocessed in step 2 through a virtual encoder to obtain the frequency domain coefficients of the residual frame and calculate the coded frame bit rate; Step 4: Pass the frequency domain coefficients of the residual frame obtained in step 3 through a virtual decoder to obtain a decoded frame; Step 5: Use the coded frame rate calculated in step 3 and the decoded frame obtained in step 4 to train the preprocessing model constructed in step 1; Step 6: Use the preprocessing model trained in step 5 to test the effect.

2. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step one includes: The preprocessing model is built based on convolutional neural networks and spatial attention mechanisms, which is used to optimize the perception of input frames and improve the human eye perception quality of encoded videos; it includes: feature extraction convolutional layers, multiple stacked residual local feature blocks, 3×3 convolutional layers and reconstruction modules; among them, the feature extraction convolutional layers are composed of 3×3 convolutional layers; the residual local feature blocks include three layers of 3×3 convolutional layers, one 1×1 convolutional layer and an enhanced spatial attention mechanism; the three layers of 3×3 convolutional layers are composed of convolutional layers, activation functions and skip connections; the enhanced spatial attention mechanism includes the first 1×1 convolutional layer, 3×3 convolutional layers, maximum pooling, multiple 3×3 convolutions, upsampling, the second 1×1 convolutional layer, activation functions and skip connections; the reconstruction module is composed of 1×1 convolutional layers; The input frame enters the preprocessing model, and first obtains coarse-grained features through feature extraction convolution. Then, multiple stacked residual local feature blocks are used to perform fine-grained learning optimization on the coarse-grained features to obtain fine-grained features. The fine-grained feature map is smoothed using a 3×3 convolutional layer; finally, the reconstruction module is used to map the fine-grained features to the pixel domain.

3. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step 2 includes: Step 2.1, performing color space conversion on the input coded frame and the reference frame; According to the conversion relationship between RGB and YUV color spaces, the color space conversion is first performed on the coded frame and the reference frame in the input RGB color space, and the RGB format is converted into the YUV format. Then, the converted YUV data is split into independent brightness component Y, chrominance component U and chrominance component V, where the brightness component Y represents the brightness information of the image, and the chrominance component U and chrominance component V represent the chrominance information of the image respectively; Step 2.2, extracting the features of the brightness component Y of the coded frame and the reference frame; In the preprocessing stage, the luminance component Y of the coded frame and the reference frame is first taken as input and passed to the preprocessing model. The preprocessing model performs preliminary feature extraction on the input coded frame and the reference frame respectively through the feature extraction convolution layer to capture the spatial structure and texture information in the frame; then, multiple stacked residual local feature blocks are used to refine the feature representation of the spatial structure and texture information, and a 3×3 convolution layer is used for smoothing operation to extract local detail information in the frame and retain high-dimensional features; Step 2.3: Image reconstruction; The high-dimensional features extracted and refined in step 2.1 are input into the reconstruction module, and are reconverted to the pixel domain using a 1×1 convolution operation mapping to perform image reconstruction to obtain the preprocessed coded frame and reference frame.

4. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step three includes: Step 3.1, using the feature domain-based optical flow estimation on the coded frame and the reference frame preprocessed in step 2.3 to obtain an optical flow field; Step 3.2, use the optical flow collaborative module to cooperate with the current frame to obtain the residual frame; The optical flow collaborative module includes optical flow distortion and prediction frame post-processing. The optical flow distortion is used to map the pixel value in the reference frame pre-processed in step 2.3 to the corresponding position of the coded frame pre-processed in step 2.3 according to the motion vector in the optical flow field obtained in step 3.1 to generate a prediction frame; the prediction frame is input into the prediction frame post-processing for artifact repair to obtain a high-quality prediction frame; the coded frame pre-processed in step 2.3 is subtracted pixel by pixel from the high-quality prediction frame after artifact repair to obtain a residual frame; Step 3.3, performing discrete cosine transform on the residual frame obtained in step 3.2 to obtain frequency domain coefficients of the residual frame; Step 3.4: using a frequency domain perception module to remove the statistical correlation between the transform units from the frequency domain coefficients of the residual frame obtained in step 3.3; The frequency domain perception module is composed of a feature extraction module, a linear mapping, a Transformer encoder, and a transformation unit reconstruction module. The feature extraction module is composed of a two-layer convolutional neural network. The Transformer encoder is composed of a multi-head self-attention mechanism, a first layer normalization layer, a first jump connection, a feedforward neural network layer, a second layer normalization layer and a second jump connection. The transformation unit reconstruction module is composed of a linear transformation and a deconvolution network. Step 3.5, quantizing the frequency domain coefficients from which statistical correlation has been removed in step 3.4; Step 3.6: Input the frequency domain coefficients quantized in step 3.5 into the entropy model, estimate the probability distribution of the frequency domain coefficients, and calculate the coding frame rate.

5. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step 4 includes: Step 4.1, dequantize and inverse discrete cosine transform the quantized frequency domain coefficients obtained in step 3.5 to obtain a decoded residual frame; Step 4.2: Add the decoded residual frame obtained in step 4.1 to the input reference frame in step 2 to obtain a decoded frame.

6. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step five includes: Design the loss function to train the preprocessing model constructed in step 1; The loss function consists of bit rate loss and fidelity loss; the bit rate loss uses the coded frame bit rate obtained in step 3.6 as the loss, and the fidelity loss is calculated using the input coded frame and the decoded frame obtained in step 4; the formula is as follows: Among them, M represents The true distribution of represents the frequency domain coefficient; represents L1 loss, Represents multi-scale structural similarity loss; α and β are hyperparameters used to weigh the weight of L1 and MS-SSIM losses, and α+β=1; in fidelity loss, L1 loss is used to focus on precise pixel-level alignment to ensure accurate restoration of optimized video frames, and MS-SSIM is used to ensure that the structural information of optimized video frames remains unchanged from that of the original video frames; the final loss function is as follows: L=l F +l R Among them, λ is a hyperparameter used to determine the proportion of bit rate loss.

7. The video compression intelligent preprocessing method based on frequency domain perception optimization according to claim 1 is characterized in that: The specific method of step six includes: The preprocessed encoded video and the unpreprocessed encoded video are compared and analyzed using Video Multi-method Assessment Fusion (VMAF) and Structural Similarity (SSIM), and the VMAF and SSIM scores of the two are calculated and compared. The higher the score, the higher the video quality.

8. A video compression intelligent preprocessing system based on frequency domain perception optimization based on the method according to any one of claims 1 to 7, characterized in that: include: A preprocessing model building module, used to build a preprocessing model; A coding frame and reference frame preprocessing module, used for preprocessing the input coding frame and reference frame using a preprocessing model to obtain a preprocessed coding frame and reference frame; The residual frame frequency domain coefficient and coding frame bit rate acquisition module is used to pass the pre-processed coding frame and the reference frame through the virtual encoder to obtain the residual frame frequency domain coefficient and calculate the coding frame bit rate; A decoded frame acquisition module is used to obtain a decoded frame by passing the frequency domain coefficients of the residual frame through a virtual decoder; The preprocessing model training module uses the encoding frame rate and the decoding frame to train the preprocessing model; The preprocessing model effect testing module is used to perform effect testing using the trained preprocessing model.

9. A video compression intelligent preprocessing device based on frequency domain perception optimization, characterized in that: include: Memory: a computer program storing a video compression intelligent preprocessing method based on frequency domain perception optimization as described in any one of claims 1 to 7, which is a computer-readable device; Processor: used to implement the video compression intelligent preprocessing method based on frequency domain perception optimization as described in any one of claims 1-7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the video compression intelligent preprocessing method based on frequency domain perception optimization described in any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent video compression method based on image analysis

    CN118055235A

Cited By

  • Perception optimization video coding quantization parameter prediction method based on attention mechanism

    CN120416478A

  • Remote sensing image compression method and device based on frequency domain enhancement and adaptive optimization

    CN120692401A

  • A Remote Sensing Image Compression Method and Apparatus Based on Frequency Domain Enhancement and Adaptive Optimization

    CN120692401B

  • Efficient compressed sensing target recognition system and method based on discrete cosine transform

    CN120997651A

  • Video data encoding and decoding method and device

    CN121262370A