End-to-End Deep Video Compression Method Based on Swin TransGAN
By introducing Swin Transformer and StyleSwin TransGAN networks into the DVC framework, combined with the wavelet condition discriminator, the problem that the classic DVC framework cannot effectively extract features and improve coding quality is solved, and higher quality video compression is achieved.
Patent Information
- Application Number
- CN202211406649.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-11-10
AI Technical Summary
The classic DVC framework is designed based on CNN and RNN, and cannot effectively extract features, resulting in the coding quality still needs to be improved.
The end-to-end deep video compression method based on Swin TransGAN is adopted, and features are extracted through Swin Transformer, and predicted frames are generated in combination with the StyleSwin TransGAN network, and finally the compression artifact is eliminated using a wavelet condition discriminator.
Improve the quality of video encoding, improve the encoding quality and reduce compression artifacts through effective feature extraction and high-quality image generation.
Smart Images

Figure CN115633180B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video coding, and more specifically, to an end-to-end deep video compression method based on Swin TransGAN. Background Art
[0002] Video coding refers to video compression technology, aiming to eliminate redundant information in video data. The effectiveness of a video coding scheme is generally evaluated from two aspects: compression efficiency and the quality of the reconstructed image. In recent years, the end-to-end deep video coding (DVC) framework has received increasing attention. The main directions for optimizing DVC are as follows: The first is to perform overall optimization by modifying the structure of DVC. For example, introducing a recurrent neural network or a GAN network enables the entire framework to use non-linear transformation to decorrelate and reduce the entropy of latent variables. The second is to optimize the overall end-to-end framework by adjusting the rate-distortion (R-D) model.
[0003] Although DVC has made great progress in the field of video coding, the classic DVC framework is still designed based on CNN and RNN. It mainly improves network performance through larger-scale and more complex network connections, but it cannot effectively extract features, and the coding quality still needs to be improved. Summary of the Invention
[0004] In view of this, the present invention provides an end-to-end deep video compression method based on Swin TransGAN, which can extract features through Swin Transformer and combine with the StyleSwin TransGAN network to achieve higher-quality image reconstruction.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An end-to-end deep video compression method based on Swin TransGAN, comprising:
[0007] S1. Combine the reconstruction value of the previous frame Perform motion estimation on the current frame X i To obtain the motion vector m of the current frame i ;
[0008] S2. Use Swin Transformer to perform motion and residual compression on the motion vector m i To obtain the compressed motion vector And the latent representation y i ;
[0009] S3. Combine the compressed motion vector And the reconstruction value of the previous frame Input to the generator of the pre - constructed StyleSwinTransGAN to generate a predicted frame Among them, the generator of StyleSwin TransGAN is a combination of StyleGAN and Swin Transformer, and a dual - attention mechanism is introduced;
[0010] S4. Add the predicted frame and the reconstructed value of the residual to obtain a reconstructed frame
[0011] S5. Use a wavelet conditional discriminator combined with the latent representation y i , and after performing hierarchical downsampling and discrete wavelet decomposition on the reconstructed frame , detect the frequency difference of the reconstructed frame at each scale through the Fourier spectrum to eliminate the compression artifacts of the reconstructed frame .
[0012] Furthermore, in S1, motion estimation is performed on the current frame X i through a pyramid optical flow network.
[0013] Furthermore, the pyramid optical flow network includes five - layer pyramids, each layer has five convolutional layers, the convolutional kernel size is 7x7, and the filter sizes are 32, 64, 65, 16, 2 respectively.
[0014] Furthermore, S2 includes:
[0015] S21. Divide the motion vector m i of the current frame into blocks of the same size and perform linear embedding;
[0016] S22. Input the linearly - embedded blocks into the first feature - extraction structure composed of Swin Transformer blocks and Patch Merging to extract features;
[0017] S23. Introduce temporal features through convolutional LSTM;
[0018] S24. Jointly input the temporal features and the features extracted by the first feature - extraction structure into the first feature - extraction structure composed of SwinTransformer blocks and Patch Merging to extract features, and input the quantized features to obtain the latent representation y i into the bitstream;
[0019] S25. Input the quantized latent representation y iAfter channel merging and the Swin Transformer block structure, the obtained features are input into a convolutional LSTM and then input into the channel merging and Swin Transformer block structure again to generate the reconstructed value of the motion vector for the current frame.
[0020] Further, S3 includes:
[0021] S31. The reconstructed motion vector is input into the generator of StyleSwin TransGAN after being normalized as a latent variable;
[0022] S32. The normalized motion vector is input into eight fully connected layers for decoupling; the decoupled motion vector is input into AdaIN through an affine transformation;
[0023] S33. The compressed motion vector and the reconstructed value of the previous frame are subjected to warping registration and downsampling to obtain an image
[0024] S34. The image is input into a feature extraction structure composed of AdaIN, dual attention, AdaIN, and a fully connected layer to extract features, obtaining an RGB image before upsampling;
[0025] S35. After the feature map extracted in S34 is upsampled and sinusoidal positional encoding is inserted, it is input into the next feature extraction structure composed of AdaIN, dual attention, AdaIN, and a fully connected layer to obtain an upsampled RGB image;
[0026] S36. The RGB image before upsampling is added to the upsampled RGB image to obtain the image for the current stage
[0027] S37. S34 - S36 are repeatedly executed a preset number of times to achieve multi-scale feature extraction and obtain the predicted frame
[0028] Further, in S32, the decoupled motion vector is input into AnaIN through an affine transformation by the following formula to obtain high-quality image generation:
[0029]
[0030]
[0031] where, represents the decoupled motion vector, Denotes the motion vector after affine transformation, A denotes the linear transformation, and b denotes the translation operation; Denotes the feature map obtained by Denotes the variance of; Denotes the mean value; Denotes the variance of; Denotes the mean value of.
[0032] Furthermore, S4 includes:
[0033] S41. Subtract the predicted frame from the original current frame X i to obtain the residual r i , and input the residual into the residual compression network to obtain the reconstructed value of the residual
[0034] S42. Add the reconstructed value of the residual to the predicted frame to obtain the reconstructed frame of the current frame
[0035] Furthermore, S5 includes:
[0036] S51. After convolving and upsampling the feature y i obtained in S24, combine it with the motion vector m i ;
[0037] S52. Combine the current frame X i , the reconstructed frame of the current frame , the previous frame X i-1 , the reconstructed frame of the previous frame and the combined feature vector in S51, and then input them into the linear projection;
[0038] S53. Input the linearly projected feature vector into the Fourier wavelet transform framework, input the output features into the convolutional LSTM, and at the same time downsample the combined feature vector in S52;
[0039] S54. Repeat S53 a preset number of times and then output the wavelet conditional discriminator result.
[0040] Furthermore, the loss function of the generator of StyleSwin TransGAN is:
[0041]
[0042] The loss function of the wavelet conditional discriminator is:
[0043]
[0044] Among them, L G and L D respectively represent the losses of the generator and the conditional discriminator; α, λ′, and β respectively represent hyperparameters for weighing the bit rate, distortion, and perceptual quality; X i , X i-1 , and respectively represent the current frame, the previous frame, the current reconstructed frame, and the previous reconstructed frame; the latent representation y i and the hidden state respectively represent the spatial feature and the temporal feature; m i represents the motion vector; D represents the conditional discriminator; i represents the current i-th frame; N represents the total number of frames in the current sequence.
[0045] Furthermore, the expression of the wavelet conditional discriminator is:
[0046]
[0047]
[0048] Among them, fWavelets(·) represents converting the image into a wavelet spectrum of multi-scale input to suppress compression artifacts, is the temporal feature.
[0049] Through the above technical solutions, it can be seen that compared with the prior art, the present invention discloses an end-to-end deep video compression method based on SwinTransGAN. First, in the optical flow compression network, based on the feature extraction ability of SwinTransformer, the present invention effectively compresses the optical flow feature map m i , and at the same time, uses the same network structure in the residual compression network to compress the residual feature map r i ; secondly, the present invention combines the high-quality image generation ability of StyleGAN and uses the generator of StyleSwin TransGAN to generate predicted frames; finally, the wavelet conditional discriminator uses the conditional discriminator and combines the Fourier spectrum and wavelet transform to eliminate the compression artifacts of high-resolution images, thereby obtaining high-quality reconstructed frames and achieving an improvement in coding quality. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0051] Figure 1 Flowchart of the end-to-end depth video compression method based on Swin TransGAN provided by the present invention;
[0052] Figure 2 Detailed architecture diagram of the end-to-end depth video compression method based on Swin TransGAN provided by the present invention. Specific embodiments
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] As Figure 1 shown, the embodiments of the present invention disclose an end-to-end depth video compression method based on Swin TransGAN, including:
[0055] S1. Combine the reconstruction value of the previous frame to perform motion estimation on the current frame X i to obtain the motion vector m of the current frame i ;
[0056] S2. Use Swin Transformer to perform motion and residual compression on the motion vector m i to obtain the compressed motion vector and the latent representation y i ;
[0057] S3. Input the compressed motion vector and the reconstruction value of the previous frame into the generator of the pre-constructed StyleSwinTransGAN to generate a predicted frame wherein, the generator of StyleSwin TransGAN is a combination of StyleGAN and Swin Transformer, and a dual attention mechanism is introduced;
[0058] S4. Input the predicted frame and the reconstructed value of the residual are added to obtain a reconstructed frame
[0059] S5. Use a wavelet conditional discriminator combined with the latent representation y i , and after performing hierarchical downsampling and discrete wavelet decomposition on the reconstructed frame , detect the frequency difference of the reconstructed frame through Fourier spectra at each scale to eliminate the compression artifacts of the reconstructed frame . The wavelet conditional discriminator uses the spatial features of the hidden state (i.e., temporal features, motion vectors, and quantized latent representation) as conditional inputs to ensure the temporality of the reconstructed image. Specifically, the loss function of the generator of StyleSwin TransGAN is as follows:
[0060]
[0061]
[0062] The loss function of the wavelet conditional discriminator is as follows:
[0063]
[0064] where L G and L D represent the losses of the generator and the wavelet conditional discriminator respectively; α, λ′, and β represent the hyperparameters for weighing the bit rate, distortion, and perceptual quality respectively; X i , X i-1 , and represent the current frame, the previous frame, the current reconstructed frame, and the previous reconstructed frame respectively; the latent representation y i and the hidden state represent the spatial feature and the temporal feature respectively; m i represents the motion vector; D represents the conditional discriminator; i represents the current frame i; N represents the total number of frames in the current sequence.
[0065] The above steps are further described below.
[0066] In one embodiment, as Figure 2 shown, in S1, the current frame X i and the reconstructed value of the previous frame are input into the pyramid optical flow network to extract the optical flow feature map (i.e., the motion vector m i ). The present invention uses a five-layer pyramid, with five convolutional layers in each layer, the convolutional kernel size is 7x7, and the filter sizes are 32, 64, 65, 16, 2 respectively, and finally outputs the motion vector m i . m iInput into the motion compression framework to achieve data volume compression
[0067] In one embodiment, S2 includes:
[0068] S21. Divide the motion vector m of the current frame i into blocks of the same size and perform linear embedding;
[0069] S22. Input the linearly embedded blocks into the first feature extraction structure composed of Swin Transformer blocks and Patch Merging to extract features;
[0070] S23. Introduce temporal features through convolutional LSTM;
[0071] S24. Jointly input the temporal features and the features extracted by the first feature extraction structure into the first feature extraction structure composed of Swin Transformer blocks and Patch Merging to extract features, and quantize the extracted features to obtain the feature y i Input into the bitstream;
[0072] S25. Quantize the feature y i Through channel merging and the Swin Transformer block structure, and input the obtained features into convolutional LSTM and then input into the channel merging and the Swin Transformer block structure again to generate the reconstructed value of the motion vector of the current frame Among them, channel merging is achieved by reducing the number of channels of the features and performing upsampling. The process of this channel merging can be regarded as the inverse process of block merging.
[0073] In a specific embodiment, S3 includes:
[0074] S31. Take the reconstructed motion vector as a latent variable, normalize it and input it into the generator of StyleSwinTransGAN;
[0075] S32. Input the normalized motion vector into eight fully connected layers for decoupling; Input the decoupled motion vector into AdaIN through affine transformation; Specifically, input the decoupled motion vector into AnaIN through affine transformation by the following formula to obtain high-quality image generation:
[0076]
[0077]
[0078] Among them, Represents the motion vector after decoupling. Represents the motion vector after affine transformation, where A represents the linear transformation and b represents the translation operation. Represents by The feature map obtained by normalizing. Represents The variance of. Represents The mean value. Represents The variance of. Represents The mean value of.
[0079] S33. The compressed motion vector and the reconstruction value of the previous frame are subjected to warping registration and downsampling to obtain the image Among them, the training loss function of the motion estimation network adopted in the warping registration process is as follows:
[0080]
[0081] Among them, W represents the warping registration operation, D represents the distortion calculation, and X i represents the current frame.
[0082] S34. The image is input into the feature extraction structure composed of AdaIN, dual attention, AdaIN and fully connected layers to extract features, and the RGB image before upsampling is obtained.
[0083] S35. After the feature map extracted in S34 is upsampled and sinusoidal positional encoding is inserted, it is input into the next feature extraction structure composed of AdaIN, dual attention, AdaIN and fully connected layers to obtain the upsampled RGB image.
[0084] S36. The RGB image before upsampling is added to the upsampled RGB image to obtain the image of the current stage This step can ensure that the image information of each stage is retained to ensure the stability of each stage of training.
[0085] S37. Repeat S34 - S36 for a preset number of times to achieve multi-scale feature extraction and obtain the predicted frame
[0086] Specifically, S3 implements the dual attention mechanism in the StyleSwin Transformer block through the following formula:
[0087] Double - Attention = Concat(head 1,...head h )W O
[0088] where W O ∈R C×C is a projection matrix used to mix the head and the output. head in the above formula i can be obtained by the following formula:
[0089]
[0090] where represents the projection matrices of query, key, and value in head i . x w and x sw represent non-overlapping blocks under regular window partitioning and sliding window partitioning, respectively.
[0091] In a specific embodiment, S4 includes:
[0092] S41. Subtract the predicted frame from the original current frame X i to obtain the residual r i , and input the residual into the residual compression network to obtain the reconstructed value of the residual
[0093] S42. Add the reconstructed value of the residual to the predicted frame to obtain the reconstructed frame of the current frame
[0094] In an embodiment, S5 includes:
[0095] S51. Combine the feature y i obtained in S24 after convolution and upsampling with the motion vector m i ;
[0096] S52. After combining the current frame X i , the reconstructed frame of the current frame , the previous frame X i-1 , the reconstructed frame of the previous frame with the combined feature vector in S51 again, input it into the linear projection to ensure the conditionality of the reconstructed frame generation (that is, the generated image has the same motion pattern as the original image. For example, if the apple rolls to the right in the original video, it is necessary to ensure that the apple in the reconstructed image also rolls to the right in the same way as the original image);
[0097] S53. Input the linearly projected feature vectors into the Fourier wavelet transform framework, and input the output features into the convolutional LSTM. At the same time, downsample the feature vectors recombined in S52.
[0098] S54. After repeating S53 for a preset number of times, output the results of the wavelet conditional discriminator.
[0099] Among them, the expression of the conditional discriminator is:
[0100]
[0101]
[0102] Among them, fWavelets(·) represents converting the image into a wavelet spectrum with multi-scale input to suppress compression artifacts, is the temporal feature.
[0103] The present invention not only solves the problem that the classical DVC framework only uses CNN or RNN for network design and cannot effectively extract features, but also designs a network framework combining Swin Transformer and GAN to improve the coding quality.
[0104] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0105] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An end-to-end deep video compression method based on Swin TransGAN, characterized in that, comprising: S1. Combine the reconstruction value of the previous frame Perform motion estimation on the current frame X i to obtain the motion vector m of the current frame i ; S2. Use Swin Transformer for motion vector m i to perform motion estimation and residual compression, obtaining the reconstructed motion vector and the latent representation y i ; S2 includes: S21. Divide the motion vector m of the current frame i into blocks of the same size and perform linear embedding; S22. Input the linearly embedded blocks into the first feature extraction structure composed of Swin Transformer blocks and Patch Merging to extract features; S23. Introduce temporal features through convolutional LSTM; S24. Co - input the temporal features and the features extracted by the first feature extraction structure into the second feature extraction structure composed of Swin Transformer blocks and Patch Merging to extract features, and obtain the latent representation y after quantifying the extracted features i into the input bitstream; S25. Quantize the latent representation y i Through channel merging and the Swin Transformer block structure, input the obtained features into the convolutional LSTM and then input them into the channel merging and the Swin Transformer block structure again to generate the reconstructed value of the motion vector of the current frame S3. Input the reconstructed motion vector and the reconstruction value of the previous frame into the generator of the pre - constructed StyleSwinTransGAN to generate a predicted frame wherein, the generator of StyleSwin TransGAN is a combination of StyleGAN and SwinTransformer, and a dual - attention mechanism is introduced; S3 includes: S31. Input the reconstructed motion vector into the generator of StyleSwin TransGAN after normalization as a latent variable; S32. Input the normalized motion vector into eight fully-connected layers for decoupling; input the decoupled motion vector into AdaIN through affine transformation; S33. Reconstructed motion vectors and the reconstructed values of the previous frame are subjected to warping registration and downsampling to obtain an image S34. Input the image into the feature extraction structure composed of AdaIN, dual attention, AdaIN, and fully connected layers to extract features, and obtain the RGB image before upsampling; S35. After upsampling the feature map extracted in S34 and inserting sinusoidal positional encoding, input it into the next feature extraction structure composed of AdaIN, dual attention, AdaIN and fully connected layers to obtain the upsampled RGB image; S36. Add the RGB image before upsampling and the RGB image after upsampling to obtain the image at the current stage S37. Repeat the execution of S34 - S36 for a preset number of times to achieve multi-scale feature extraction and obtain a predicted frame S4. Add the predicted frame and the reconstructed value of the residual to obtain the reconstructed frame S5. Use a wavelet conditional discriminator in combination with the latent representation y i , after performing hierarchical downsampling and discrete wavelet decomposition on the reconstructed frame , detect the frequency differences of the reconstructed frame at each scale through Fourier spectra to eliminate the compression artifacts of the reconstructed frame .
2. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that, In S1, the motion of the current frame X is estimated through a pyramid optical flow network i for motion estimation.
3. The end-to-end deep video compression method based on Swin TransGAN according to claim 2, characterized in that, The pyramid optical flow network includes five layers of pyramids, each layer has five convolutional layers, the convolutional kernel size is 7x7, and the filter sizes are 32, 64, 65, 16, 2 respectively.
4. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that, In S32, the decoupled motion vector is input into AnaIN through affine transformation by the following formula to obtain high-quality image generation: Among them, represents the motion vector after decoupling, represents the motion vector after affine transformation, A represents the linear transformation, and b represents the translation operation; represents the feature map obtained by normalizing ; represents variance; represents mean; represents variance; represents mean.
5. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that S4 comprising: S41. Subtract the predicted frame from the original current frame X i to obtain the residual r i , and input the residual into the residual compression network to obtain the reconstructed value of the residual S42. Reconstructed value of the residual is added to the predicted frame to obtain the reconstructed frame of the current frame 6. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that S5 comprising: S51. Combine the feature y obtained in S24 i with the motion vector m after convolution and upsampling i ; S52. Combine the current frame X i with the reconstructed frame of the current frame and the previous frame X i-1 with the reconstructed frame of the previous frame and then input the result after recombining with the combined feature vector in S51 into the linear projection; S53. Input the linearly projected feature vector into the Fourier wavelet transform framework, and input the output features into convolutional LSTM. At the same time, downsample the feature vector recombined in S52; After repeating S53 a preset number of times, eliminate the reconstructed frame compression artifacts.
7. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that, The loss function of the generator of StyleSwin TransGAN is: The loss function of the wavelet conditional discriminator is: Among them, and represent the losses of the generator and the wavelet conditional discriminator, respectively; α, λ′, and β represent hyperparameters for balancing the bit rate, distortion, and perceptual quality, respectively; X i , X i-1 , and represent the current frame, the previous frame, the current reconstructed frame, and the previous reconstructed frame, respectively; the latent representation y i and the hidden state represent the spatial feature and the temporal feature, respectively; m i represents the motion vector; D represents the conditional discriminator; i represents the current frame number i; N represents the total number of frames in the current sequence.
8. The end-to-end deep video compression method based on Swin TransGAN according to claim 1, characterized in that, The expression of the wavelet conditional discriminator is: Among them, fWavelets(·) represents converting an image into a wavelet spectrum of multi-scale input to suppress compression artifacts, which is the temporal feature.
Citation Information
Patent Citations
Video compression based on long range end-to-end deep learning
CN114450965A
Traditional image compression enhancement method based on reversible tone mapping network
CN114820354A