Swin-Transformer Architecture-Based Spatial-Frequency Domain Compressed Video Enhancement Method and Device

Through the empty frequency domain compressed video enhancement method based on the Swin-Transformer architecture, the frequency domain feature extraction and reconstruction is used to use video timing information to solve the problem of unremarkable texture performance of compressed video in the prior art, and a better video quality enhancement effect is achieved.

CN120147166BActive Publication Date: 2025-07-22HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510613168.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-22
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The compressed video quality enhancement method at this stage fails to effectively utilize video timing information, resulting in the unremarkable performance of the enhanced video frame details and textures.

Method used

The empty frequency domain compressed video enhancement method based on the Swin-Transformer architecture is adopted. By extracting and reconstructing the residual graphs of target frames and front and back frames, the Swin-Transformer block and Haar wavelet transformation are used, and the video quality is improved by combining multi-channel attention and global filtering module.

Benefits of technology

Effectively utilize video timing information, improve the detailed texture performance of compressed videos and achieve better enhancement effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147166B_ABST
    Figure CN120147166B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-frequency domain compressed video enhancement method and device based on the Swin-Transformer architecture, which relates to the field of image processing and includes: extracting the target frame and the information of the frames before and after the target frame, using a convolutional neural network to extract multi-frame feature information, aligning and fusing it into the residual map of the original video frame; using the Harr wavelet transform to extract the high-frequency component features and low-frequency component features of half the width and height of the residual map; using the Swin-Transformer module to model the texture and structural information for the extracted low-frequency component features, and using the U-net method to reconstruct the features for the extracted high-frequency component features; performing the inverse Haar wavelet transform on the reconstructed high-frequency and low-frequency component features and adding them to the original video frame to obtain the enhanced video frame. The present invention effectively solves the problem that in the enhancement of the quality of the defective compressed video at the present stage, less application of the frames before and after leads to the problem that the details and textures of the enhanced video frame are not prominent by extracting and reconstructing the frequency domain features of the residual map, and achieves a better enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and video enhancement, and particularly relates to a spatio-temporal frequency domain compressed video enhancement method and device based on the Swin-Transformer architecture. Background Art

[0002] Video quality enhancement refers to the process of removing artifacts generated in a video and improving the video picture quality, which involves a very wide range of fields, covering video super-resolution, video de-shaking, de-blurring, video de-raining and de-fogging, etc. With the development of digital technology, more and more high-definition and ultra-high-definition videos appear on the Internet. However, the current video coding technology can only be transmitted under limited bandwidth, so it is necessary to compress the video before sending it into the channel. Among them, compressed video quality enhancement is a specific research branch that focuses on the artifacts caused by the degradation process of video compression. And its application scenarios for subsequent visual tasks are also very wide, such as object recognition, target detection, target tracking, etc. The challenge of compressed video quality enhancement is to improve the visual quality of low-quality videos while ensuring the compression rate to ensure the user's experience of watching videos.

[0003] Traditional compressed video quality enhancement methods enhance the quality of a single-frame picture or optimize the transform coefficients in the compression standard to improve the quality of video frames. With the continuous progress of deep learning methods, more and more methods use convolutional neural networks to perform the task of improving image quality. Some methods use the method of residual non-local attention to remove image noise, and some methods use a dual-domain multi-path residual network to remove image artifacts. However, these methods all focus on a single video frame and cannot effectively utilize the temporal information in the video. Some methods use adjacent high-quality video frames to enhance the target frame, and some methods use deformable convolutions to extract the feature information of adjacent multiple video frames and align them to achieve the purpose of enhancing the target frame.

[0004] In addition, the Transformer architecture can achieve better feature extraction effects through its self-attention mechanism and parallel processing capabilities. The Transformer architecture makes up for the defect of the receptive field limitation of convolutional neural networks, can capture the feature information of distant frames, and makes the global perception ability of the neural network more effective. Therefore, some researchers also use methods based on Transformer for enhancing compressed videos. The window-based Transformer is used to aggregate the information in multiple channel feature maps and fuse the inter-frame information to improve the video quality. However, the compressed video quality enhancement model based on Transformer calculates the global similarity, and its computational complexity will grow exponentially with the expansion of the spatial resolution. When processing videos of larger sizes, it is often limited by hardware facilities. Summary of the Invention

[0005] The purpose of the present invention is to provide a spatio-frequency domain compressed video enhancement method and device based on the Swin-Transformer architecture. By extracting and reconstructing the frequency domain features of the residual map that fuses the target frame and several frames before and after the target frame, it effectively solves the problem in the enhancement of the defective compressed video quality at the present stage that the application of the frames before and after is less, resulting in the unremarkable performance of the detailed texture of the enhanced video frames, and achieves a better enhancement effect.

[0006] The present invention adopts the following technical solutions:

[0007] In the first aspect, a spatio-frequency domain compressed video enhancement method based on the Swin-Transformer architecture includes:

[0008] S101, obtaining a video sequence composed of the original target frame and several frames before and after the target frame, extracting multi-frame video features in the video sequence based on a convolutional neural network with a U-Net-like structure and converting them into multi-channel feature maps; using deformable convolution to perform feature alignment on the multi-channel feature maps and reconstructing them into residual maps;

[0009] S102, performing Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features;

[0010] S103, extracting and reconstructing the low-frequency component features by sequentially connecting multiple Swin-Transformer blocks, a multi-channel attention module, and a global filtering module to obtain the reconstructed low-frequency component features;

[0011] S104, respectively extracting high-frequency feature information of the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a U-Net-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and performing fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features;

[0012] S105, performing inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and adding the restored residual map to the original target frame to obtain the enhanced target frame.

[0013] Preferably, the S101 specifically includes:

[0014] Obtaining a video sequence composed of the original target frame and several frames before and after the target frame , the video sequence is transformed into a feature map with C channels through the Unet feature extraction module built based on the convolutional neural network; where, is the original target frame, and N is the number of front and back frames;

[0015] The feature map with C channels is aligned through deformable convolution, and multiple connected convolutional kernels with 2×C input and output channels, a convolutional kernel of 3, a padding of 1, and a stride of 1 are used for feature reconstruction. Finally, a convolutional kernel with 2×C input channels, 1 output channel, a convolutional kernel of 3, a padding of 1, and a stride of 1 is used to obtain a single residual map ; where, H and W respectively represent the height and width of the feature, represents the set of real numbers.

[0016] Preferably, in the S103, the processing process of multiple Swin-Transformer blocks is as follows:

[0017] The low-frequency component features are expanded to channels through a pointwise convolution, and then downsampled once through a convolution to obtain , and features are extracted through the first Swin-Transformer block, and then downsampled again in the same way to obtain ; for , features are extracted through the second Swin-Transformer block to obtain ; then, through the feature refinement module and the third Swin-Transformer block, the feature expression is further deepened, and the output is obtained; and are concatenated, and the feature expression ability of the intermediate layer is enhanced through residual connection. Through a convolution and passing through the fourth Swin-Transformer block, is obtained; is concatenated with , and through a convolution and the fifth Swin-Transformer block, an upsampling operation is performed using a transposed convolution to obtain ; and are concatenated, and again through a convolution, and through the sixth Swin-Transformer block and an upsampling operation, is obtained. is sent into the multi-channel attention module; where, H and W respectively represent the height and width of the feature, represents the set of real numbers.

[0018] ​Preferably, in S103, the processing process of the multi-channel attention module is as follows:

[0019] In the multi-channel attention module, first, layer normalization is used to ensure the stable distribution of data features in each channel. Then, pointwise convolution and depthwise separable convolution are used to fuse features and divide the result into Q, K, and V, that is , , , and the self-attention score is obtained by processing through a simplified formula for calculating the attention mechanism, as follows:

[0020] ;

[0021] Among them, represents the calculation method of the attention mechanism, which is used to measure the correlation between different regions; Q is Query, representing the feature representation of the current pixel region; K is Key, representing the features of all regions in the image; V is Value, which is used to provide the feature information corresponding to Key; represents the normalization process of the similarity result calculated by Q and K;

[0022] The self-attention score passes through a pointwise convolution and LeakyReLU, and the output of the feature is obtained through weighted connection to the global filtering module.

[0023] Preferably, in S103, the processing process of the global filtering module is as follows:

[0024] Receive the output of the multi-channel attention module, first perform a layer normalization process, then expand the result of the layer normalization into two branches for calculation through a grouped convolution, multiply the two branches element by element to obtain a residual result, and then pass through a grouped convolution and add the input to obtain the final low-frequency reconstruction result .

[0025] Preferably, S104 specifically includes:

[0026] For the vertical high-frequency component feature , the horizontal high-frequency component feature and the diagonal high-frequency component feature , respectively perform high-frequency feature information through the residual connection neural network of the Unet-like structure. First, for , and , perform a channel expansion once, expand the input one-dimensional feature map to layers, and obtain , and , and then perform a downsampling through a convolution respectively to obtain the downsampled high-frequency information , and , and perform downsampling operation again to obtain downsampled high-frequency information through a convolution , and , and obtain enhanced features through multiple convolutional layers , and , and respectively concatenate with , and , and obtain fused features through a convolution , and , then perform upsampling once through a transposed convolution to obtain , and , and concatenate with , and , and obtain , and through a convolution and a transposed convolution, and concatenate with , and , and obtain , and respectively through a convolution, and perform concatenation, and obtain three reconstructed high-frequency component features corresponding to the vertical high-frequency component , horizontal high-frequency component and diagonal high-frequency component .

[0027] In the second aspect, a spatio-temporal frequency domain compressed video enhancement device based on the Swin-Transformer architecture includes:

[0028] A multi-frame feature extraction and reconstruction module, configured to obtain a video sequence composed of an original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps; use deformable convolution to perform feature alignment on the multi-channel feature maps and reconstruct them into residual maps;

[0029] A high and low frequency component feature acquisition module, configured to perform Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features;

[0030] The low-frequency component feature reconstruction module is used to extract and reconstruct the low-frequency component features by sequentially connecting multiple Swin-Transformer blocks, a multi-channel attention module, and a global filtering module, and obtain the reconstructed low-frequency component features;

[0031] The high-frequency component feature reconstruction module is used to respectively extract high-frequency feature information of the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a Unet-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and fuse and process the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features;

[0032] The target frame enhancement module is used to perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.

[0033] In a third aspect, an electronic device includes:

[0034] One or more processors;

[0035] A storage device for storing one or more programs;

[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.

[0037] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements any of the spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.

[0038] In a fifth aspect, a computer program product includes a computer program, and when the computer program is executed by a processor, it implements any of the spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] (1) The present invention designs an image frequency domain processing architecture based on Swin-Transformer. Multiple stacked Swin-Transformer blocks can effectively extract low-frequency component feature information of the global structure; the high-frequency component features can retain and enhance the texture information in the image through information extraction by a residual connection neural network with a Unet-like structure;

[0041] (2) In the present invention, pointwise convolution is added between Swin-Transformer blocks, and residual fusion is performed from shallow to deep, and filtering and global attention mechanisms are added at the end, enabling effective utilization of low-frequency information; at the same time, in the Swin-Transformer block, a window attention mechanism is added, which can enhance the capture of local texture information. Attention calculation is performed in blocks through the window attention mechanism, reducing the computational complexity of the method using Transformer. Description of the Drawings

[0042] Figure 1 It is a flowchart of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture according to an embodiment of the present invention;

[0043] Figure 2 It is an overall flowchart block diagram of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture according to an embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;

[0045] Figure 4 It is a schematic diagram of the multi-channel attention module in the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the global filtering module in the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;

[0047] Figure 6 It is a schematic diagram of the high-frequency component feature enhancement process architecture according to an embodiment of the present invention;

[0048] Figure 7 It is a structural block diagram of the spatio-temporal frequency domain fusion compression video enhancement device based on the Swin-Transformer architecture according to an embodiment of the present invention;

[0049] Figure 8 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Embodiments

[0050] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0051] See Figure 1As shown in the figure, a spatio-temporal frequency domain compressed video enhancement method based on the Swin-Transformer architecture in this embodiment includes the following steps.

[0052] S101: Obtain a video sequence composed of the original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a U-Net-like structure, and convert them into multi-channel feature maps; use deformable convolution to perform feature alignment on the multi-channel feature maps and reconstruct them into residual maps.

[0053] S102: Perform Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features.

[0054] S103: Pass the low-frequency component features through a series of connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module for feature extraction and reconstruction to obtain the reconstructed low-frequency component features.

[0055] S104: Respectively pass the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a U-Net-like structure to extract high-frequency feature information and output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information. Perform fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features.

[0056] S105: Perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.

[0057] See Figure 2 As shown in the figure, the spatio-temporal frequency domain fusion compressed video enhancement method model based on the Swin-Transformer architecture in this embodiment includes three modules: a multi-frame feature extraction and reconstruction module, a Swin-Transformer low-frequency enhancement module, and a high-frequency component feature enhancement and fusion module. Figure 2First, obtain the original low-quality target frame and several frames before and after it, and send them together into the Unet feature extraction network to convert the video frames into feature maps of specific channels. Then, perform feature alignment on the feature maps through deformable convolution and reconstruct the residual map. Subsequently, enter the Swin-Transformer frequency domain enhancement module and the high-frequency component feature enhancement fusion module. Decompose the residual map through Haar wavelet transform to obtain high-frequency component features and low-frequency component features, and enhance them respectively through convolutional neural networks. Finally, obtain the enhanced high-quality target frame through inverse Haar wavelet transform of the reconstructed high- and low-frequency component features.

[0058] The following elaborates in detail on the process steps of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture.

[0059] The specific implementation of the multi-frame feature extraction and reconstruction module in S101 is as follows.

[0060] First, obtain a video sequence consisting of the original target frame and several frames before and after the target frame , and convert the video sequence into a feature map with C channels through the Unet feature extraction module built based on convolutional neural networks; where is the original target frame, and N is the number of frames before and after.

[0061] Then, perform feature alignment on the C-channel feature map through deformable convolution, and use a series of connected convolutional kernels with 2×C input and output channels, a convolutional kernel size of 3, a padding of 1, and a stride of 1 for feature reconstruction. Finally, use a convolutional kernel with 2×C input channels, 1 output channel, a convolutional kernel size of 3, a padding of 1, and a stride of 1 to obtain a single residual map , which serves as the input to the Haar wavelet transform; where H and W respectively represent the height and width of the feature, represents the set of real numbers.

[0062] See Figure 3 As shown, the Swin-Transformer low-frequency enhancement module includes multiple Swin-Transformer blocks, a multi-channel attention module, and a global filtering module. The input of the Swin-Transformer low-frequency enhancement module is the low-frequency component feature obtained through Haar wavelet decomposition. The window size of each Swin-Transformer block is 8, and it includes an embedding layer, a window attention layer, a multi-layer perceptron layer, a normalization layer, and a de-embedding layer. To obtain image information of different sizes and expand the receptive field, first let the low-frequency component pass through an output channel number of , pointwise convolution with a convolution kernel of 1 and a stride of 1 expands the original number of channels to , and then performs a downsampling operation through a convolution with a convolution kernel of 3, a stride of 2, and a padding of 1 to obtain , and extracts features through a Swin-Transformer block, and then performs another downsampling operation of the same kind to obtain . For , use the features extracted through a Swin-Transformer block to obtain , and then pass through a feature extraction module, which includes a depthwise convolution with an input and output number of channels of , a convolution kernel of 3, a stride of 1, and a number of groups of , a pointwise convolution with an input and output number of channels of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block to further deepen the feature expression and obtain the output , concatenate and , and enhance the feature expression ability of the intermediate layer through the way of residual connection. Through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block to obtain , concatenate with , through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block, and perform an upsampling operation through a transposed convolution with a convolution kernel of 4, a stride of 2, and a padding of 1 to obtain , and concatenate and , and once again through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and through the last Swin-Transformer block and an upsampling operation of the same kind to obtain , and send into the multi-channel attention module.

[0063] See Figure 4 As shown, in the multi-channel attention module, first use layer normalization to ensure the stable distribution of data features in each channel and prevent overfitting, and then use pointwise convolution and depthwise separable convolution with an output number of layers of to fuse features and separate the results into Q, K, and V, that is, , , , and by simplifying the formula of the computational attention mechanism, the model can effectively utilize global information while maintaining a low computational complexity, as follows:

[0064] ;

[0065] Among them, represents the computational method of the attention mechanism, which is used to measure the correlation between different regions; Q is the Query, representing the feature representation of the current pixel region; K is the Key, representing the features of all regions in the image; V is the Value, which is used to provide the feature information corresponding to the Key; represents the normalization processing of the similarity result calculated from Q and K, so as to generate attention weights and ensure that the sum of all weights is 1.

[0066] The finally calculated self-attention score is obtained. Finally, the score is passed through a pointwise convolution and LeakyReLU, and the feature is output to the global filtering module through weighted connection.

[0067] See Figure 5 As shown, in the global filtering module, first, a layer normalization process is performed, and then the result of the layer normalization is extended into two branches for calculation through a grouped convolution. Each branch contains a pointwise convolution and a depthwise separable convolution, and LeakyRelu is used as the activation layer in one of the layers. The residual result is obtained by multiplying the two branches element by element, and then a grouped convolution with a convolution kernel of 3 is performed, and the input is added to obtain the final low-frequency reconstruction result .

[0068] As Figure 6 shown, the input of the high-frequency component feature enhancement and fusion module is the three high-frequency components of the Haar wavelet, namely the vertical high-frequency component , the horizontal high-frequency component , and the diagonal high-frequency component . The feature maps of the three components are processed through a residual connection neural network similar to Unet to obtain an output feature map respectively. In the residual connection network similar to Unet, for the vertical high-frequency information , the horizontal high-frequency information , and the diagonal high-frequency information obtained by Haar wavelet decomposition, first, a channel expansion is performed on , , and to expand the input one-dimensional feature map to layers, obtaining , and and then perform downsampling once through a convolution with a convolution kernel of 3, padding of 1, stride of 2, and the number of input and output channels being to obtain downsampled high-frequency information 、 and and perform downsampling again. Obtain downsampled high-frequency information through a convolution with a convolution kernel of 3, padding of 1, stride of 2, and the number of input and output channels both being 、 and and obtain enhanced features through multiple convolutional layers with a convolution kernel of 3, padding of 1, stride of 1, and the number of input and output channels both being 、 and and concatenate them with 、 and respectively, and obtain fused features 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of 、 and and then perform upsampling once through a transposed convolution with a convolution kernel of 4, padding of 1, stride of 2, and the number of input and output channels both being to obtain 、 and and concatenate them with 、 and respectively, and then obtain 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of and a transposed convolution with a convolution kernel of 4, padding of 1, stride of 2, and the number of input and output channels both being 、 and and concatenate them with 、 and respectively, and obtain 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of 、 and respectively, and concatenate them. Pass through N convolutional layers with a convolution kernel of 3, padding of 1, stride of 1, and input channels of ​​​​​​​​​​​​​​​, the number of output channels is After convolution with , finally, through a convolution with a convolution kernel of 3, padding of 1, stride of 1, and the number of input channels being and the number of output channels being 3, three high-frequency feature reconstruction feature maps are finally obtained, corresponding to the vertical high-frequency component , the horizontal high-frequency component and the diagonal high-frequency component , and combined with the previous low-frequency reconstruction result

[0069] See Figure 7 As shown, the present invention also discloses a spatio-temporal frequency domain compressed video enhancement device based on the Swin-Transformer architecture, including:

[0070] The multi-frame feature extraction and reconstruction module 701 is used to obtain a video sequence composed of the original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps; use deformable convolution to perform feature alignment on the multi-channel feature maps and reconstruct them into residual maps;

[0071] The high and low frequency component feature acquisition module 702 is used to perform Haar wavelet transform on the residual map to obtain the high frequency component features and low frequency component features of the width and height dimensions of the residual map; the high frequency component features include vertical high frequency component features, horizontal high frequency component features, and diagonal high frequency component features;

[0072] The low frequency component feature reconstruction module 703 is used to perform feature extraction and reconstruction on the low frequency component features through a plurality of sequentially connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module to obtain the reconstructed low frequency component features;

[0073] The high frequency component feature reconstruction module 704 is used to respectively perform high frequency feature information extraction on the vertical high frequency component features, horizontal high frequency component features, and diagonal high frequency component features through a residual connection neural network with a Unet-like structure to output vertical high frequency feature information, horizontal high frequency feature information, and diagonal high frequency feature information, and perform fusion processing on the vertical high frequency feature information, horizontal high frequency feature information, and diagonal high frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high frequency component features;

[0074] The target frame enhancement module 705 is used to perform inverse Haar wavelet transform on the reconstructed high frequency component features and the reconstructed low frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.

[0075] The specific implementation of each module of a spatio-temporal frequency compressed video enhancement device based on the Swin-Transformer architecture is the same as that of a spatio-temporal frequency compressed video enhancement method based on the Swin-Transformer architecture, and will not be repeated in this embodiment.

[0076] See Figure 8 The following shows the schematic hardware structure of the electronic device provided by the embodiment of the present invention. Figure 8 In this embodiment, the electronic device includes: a processor 801 and a memory 802; wherein the memory 802 is used to store computer execution instructions; the processor 801 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0077] Optionally, the memory 802 can be either independent or integrated with the processor 801.

[0078] When the memory 802 is independently provided, the electronic device further includes a bus 803 for connecting the memory 802 and the processor 801.

[0079] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 801 executes the computer execution instructions, the above method is implemented.

[0080] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 801, the above method is implemented.

[0081] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0082] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0083] In addition, in each embodiment of the present invention, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.

[0084] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 801 to execute part of the steps of the methods in various embodiments of the present application.

[0085] It should be understood that the above processor 801 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 801 can also be any conventional processor 801, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 801, or implemented by the combination of the hardware and software modules in the processor 801.

[0086] The memory 802 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk, or an optical disc, etc.

[0087] The bus 803 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 803 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus 803 in the drawings of the present application is not limited to only one bus 803 or one type of bus 803.

[0088] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0089] An exemplary storage medium is coupled to the processor 801, enabling the processor 801 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor 801. The processor 801 and the storage medium can be located in an application specific integrated circuit (ASIC).

[0090] In this embodiment, tests are conducted on a hardware platform using Python 3.8 and PyTorch 1.13. The configuration of the hardware platform is such as Intel i9-13900K (32G memory), NVIDIA graphics card RTX4090, and Ubuntu operating system. The test dataset is the MFQEv2 dataset, and the test videos are 18 natural videos, divided into five categories: Category A (2560×1600), Category B (1920×1080), Category C (832×480), Category D (480×240), and Category E (1280×720). The number of frames of these videos ranges from 150 to 600. Compression is performed under different quantization parameters, and the quantization parameters are set to 22, 27, 32, 37, and 42. And a test video with a QP value of 37 is used for detailed analysis. The parameters for comparison are the peak signal-to-noise ratio ΔPSNR (dB) and the structural similarity ΔSSIM (10 -2 )), and the comparison target is the difference from the video compressed under the original platform (H.265 / HEVC).

[0091] Table 1 Comparison between the existing enhancement method and the method of the present invention;

[0092]

[0093] The test results are shown in Table 1. It can be seen from Table 1 that in terms of the objective indicators ΔPSNR and ΔSSIM, the method of the present invention is superior to other enhancement methods.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An air-frequency domain compressed video enhancement method based on the Swin-Transformer architecture, characterized in that Including: S101, obtaining a video sequence composed of an original target frame and several frames before and after the target frame, extracting multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and converting them into multi-channel feature maps; using deformable convolution to perform feature alignment on the multi-channel feature maps and reconstructing them into residual maps; S102, performing Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features; S103, extracting and reconstructing the low-frequency component features by passing the low-frequency component features through a plurality of sequentially connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module to obtain the reconstructed low-frequency component features; S104, respectively extracting high-frequency feature information of the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a Unet-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and performing fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features; S105, performing inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and adding the restored residual map to the original target frame to obtain the enhanced target frame; In S103, the processing process of the plurality of Swin-Transformer blocks is as follows: The low-frequency component features Pass through a pointwise convolution to expand the number of channels to C embed , and then perform a downsampling through a convolution to obtain And extract features through the first Swin-Transformer block, and then perform the same downsampling again to obtain For Use the features extracted through the second Swin-Transformer block to obtain Then pass through the feature refinement module and the third Swin-Transformer block to further deepen the feature expression and obtain the output Combine and through concatenation, improve the feature expression ability of the intermediate layer through residual connection, and obtain through a convolution and the fourth Swin-Transformer block Combine with through concatenation, through a convolution and the fifth Swin-Transformer block, and perform an upsampling operation using a transposed convolution to obtain Combine and through concatenation, again through a convolution, and obtain through the sixth Swin-Transformer block and an upsampling operation Send into the multi-channel attention module; where H and W represent the height and width of the feature respectively, and R represents the set of real numbers; In S103, the processing process of the multi-channel attention module is as follows: In the multi-channel attention module, layer normalization is first used to ensure the stable distribution of data features in each channel, and then pointwise convolution and depthwise separable convolution are used to fuse the features and separate the results into Q, K, and V, that is and the self-attention scores are obtained by processing through the simplified formula for calculating the attention mechanism; Passing the self-attention score through a pointwise convolution and LeakyReLU, and obtaining the output of the feature to the global filtering module through weighted connection; In S103, the processing process of the global filtering module is as follows: Receiving the output of the multi-channel attention module, first undergoes a layer normalization process, and then the result of the layer normalization is expanded into two branches for calculation through a grouped convolution. The two branches are multiplied element-wise to obtain a residual result, and then through a grouped convolution, and the input is added to obtain the final low-frequency reconstruction result 2. The spatio-temporal compression video enhancement method based on the Swin-Transformer architecture according to claim 1, wherein S101 specifically includes: Obtain a video sequence composed of the original target frame and a number of frames before and after the target frame Convert the video sequence through a Unet feature extraction module built based on a convolutional neural network into a feature map with C channels; where t I is the original target frame and N is the number of frames before and after Align the feature maps of the C channel through deformable convolution, and use multiple connected convolutional kernels with the number of input and output channels being 2×C, the convolutional kernel size being 3, the padding being 1, and the stride being 1 for feature reconstruction. Finally, use a convolutional kernel with the number of input channels being 2×C, the number of output channels being 1, the convolutional kernel size being 3, the padding being 1, and the stride being 1 to obtain a single residual map I t ′∈R (H ×W×1) ; where H and W respectively represent the height and width of the feature, and R represents the set of real numbers.

3. The spatio-temporal compression video enhancement method based on the Swin-Transformer architecture according to claim 1, wherein Processing through a simplified formula for calculating the attention mechanism to obtain the self-attention score, as follows: Attention(Q,K,V) = Softmax(QK T )V; Among them, Attention represents the calculation method of the attention mechanism, which is used to measure the correlation between different regions; Q is the Query, which represents the feature representation of the current pixel region; K is the Key, which represents the features of all regions in the image; V is the Value, which is used to provide the feature information corresponding to the Key; Softmax represents normalizing the similarity result calculated by Q and K.

4. The spatio-temporal compressive video enhancement method based on the Swin-Transformer architecture according to claim 1, wherein S104 specifically includes: Vertical high-frequency component features Horizontal high-frequency component features And diagonal high-frequency component features Respectively, the high-frequency feature information is processed through the residual connection neural network of the Unet-like structure. First, for I′ HL 、I′ LH and I′ HH Perform a channel expansion once to expand the input one-dimensional feature map to C hidden layers to obtain and Then, respectively, perform a downsampling through a convolution to obtain the downsampled high-frequency information and And perform a downsampling operation again to obtain the downsampled high-frequency information through a convolution and And through multiple convolutional layers, obtain the enhanced features and And respectively concatenate with and And obtain the fused features through a convolution and Then perform an upsampling through a transposed convolution once to obtain and And concatenate with and And perform a concatenation and obtain and And concatenate with and Perform a concatenation, and through a convolution, respectively obtain and Perform a concatenation, and through several convolutions, obtain three reconstructed high-frequency component features, corresponding to the vertical high-frequency component Horizontal high-frequency component And diagonal high-frequency component 5. An air-frequency domain compressed video enhancement device based on the Swin-Transformer architecture, characterized in that, Based on the spatio-temporal frequency domain compression video enhancement method based on the Swin-Transformer architecture as described in any one of claims 1 to 4, the device includes: A multi-frame feature extraction and reconstruction module, configured to obtain a video sequence composed of an original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps; using deformable convolution to perform feature alignment on the multi-channel feature maps and reconstructing them into residual maps; High and low frequency component feature acquisition module, configured to perform Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features; Low-frequency component feature reconstruction module, configured to perform feature extraction and reconstruction on the low-frequency component features through a plurality of sequentially connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module to obtain the reconstructed low-frequency component features; High-frequency component feature reconstruction module, configured to perform high-frequency feature information extraction on the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features respectively through a residual connection neural network with a Unet-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and perform fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features; Target frame enhancement module, configured to perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.

6. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-4.

7. A computer-readable storage medium, having a computer program stored thereon, and when the program is executed by a processor, the method according to any one of claims 1-4 is implemented.

8. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Image super-resolution method and device based on cross attention mechanism and Swin-Transform

    CN117237197A

  • Video super-resolution method and system based on deformable Transformer

    CN117593185A