Space-frequency domain compressed video enhancement method and device based on Swin-Transform architecture
By adopting the space frequency domain processing method based on the Swin-Transformer architecture in video quality enhancement, frequency domain feature extraction and reconstruction of the residual maps of video frames and front and back frames, the problem of insufficient utilization of front and back frame information in the prior art is solved, and a better video quality enhancement effect is achieved.
Patent Information
- Application Number
- CN202510613168.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
At this stage, the compressed video quality enhancement method is insufficient in utilizing the front and back frame information, resulting in the enhanced video frame details and texture performance not outstanding.
The empty frequency domain compressed video enhancement method based on the Swin-Transformer architecture is adopted. By extracting and reconstructing the residual maps of several frames before and after the fused target frame, the feature extraction and reconstruction is performed using multiple Swin-Transformer blocks, multi-channel attention modules and global filtering modules.
It effectively improves the performance of video frame detail texture, achieves better video quality enhancement effect, and reduces computing complexity.
Smart Images

Figure CN120147166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and video enhancement, and particularly relates to a spatio-frequency domain compressed video enhancement method and device based on the Swin-Transformer architecture. Background Art
[0002] Video quality enhancement refers to the process of removing artifacts generated in a video and improving the video picture quality. It involves a very wide range of fields, covering video super-resolution, video de-shaking, de-blurring, video de-raining and de-fogging, etc. With the development of digital technology, more and more high-definition and ultra-high-definition videos appear on the Internet. However, the current video coding technology can only transmit under limited bandwidth, so it is necessary to compress the video before sending it into the channel. Among them, compressed video quality enhancement is a specific research branch, focusing on the artifacts caused by the degradation process of video compression. And its application scenarios for subsequent visual tasks are also very extensive, such as object recognition, target detection, target tracking, etc. The challenge of compressed video quality enhancement is to improve the visual quality while ensuring the compression rate for low-quality videos, and to ensure the user experience of watching videos.
[0003] Traditional compressed video quality enhancement methods enhance the quality of a single frame of picture, or optimize the transform coefficients in the compression standard to improve the quality of video frames. With the continuous progress of deep learning methods, more and more methods use convolutional neural networks to perform the task of improving image quality. Some methods use the method of residual non-local attention to remove image noise, and some methods use a dual-domain multi-path residual network to remove image artifacts. However, these methods all focus on a single video frame and cannot effectively utilize the temporal information in the video. Some methods use adjacent high-quality video frames to enhance the target frame, and some methods use deformable convolutions to extract the feature information of adjacent multiple video frames and align them to achieve the purpose of enhancing the target frame.
[0004] In addition, the Transformer architecture can achieve better feature extraction effects through its self-attention mechanism and parallel processing capabilities. The Transformer architecture makes up for the defect of the receptive field limitation of convolutional neural networks and can capture the feature information of distant frames, making the global perception ability of the neural network more effective. Therefore, some researchers also use methods based on Transformer for enhancing compressed videos. The window-based Transformer is used to aggregate the information in multiple channel feature maps and fuse the inter-frame information to improve the video quality. However, the compressed video quality enhancement model based on Transformer calculates the global similarity, and its computational complexity will grow exponentially with the expansion of the spatial resolution. When processing videos of larger sizes, it is often limited by hardware facilities. Summary of the Invention
[0005] The purpose of the present invention is to provide a spatio-temporal frequency domain compressed video enhancement method and device based on the Swin-Transformer architecture. By extracting and reconstructing the frequency domain features of the residual map that fuses the target frame and several frames before and after the target frame, it effectively solves the problem in the enhancement of the defective compressed video quality at the present stage that the application of the frames before and after is less, resulting in the lack of prominent details and textures in the enhanced video frames, and achieves a better enhancement effect.
[0006] The present invention adopts the following technical solutions:
[0007] In the first aspect, a spatio-temporal frequency domain compressed video enhancement method based on the Swin-Transformer architecture includes:
[0008] S101. Obtain a video sequence composed of the original target frame and several frames before and after the target frame, extract the multi-frame video features in the video sequence based on a convolutional neural network with a U-Net-like structure and convert them into a multi-channel feature map; use deformable convolution to align the features of the multi-channel feature map and reconstruct them into a residual map;
[0009] S102. Perform Haar wavelet transform on the residual map to obtain the high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features;
[0010] S103. Pass the low-frequency component features through a plurality of Swin-Transformer blocks, a multi-channel attention module, and a global filtering module connected in sequence for feature extraction and reconstruction to obtain the reconstructed low-frequency component features;
[0011] S104. Respectively pass the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a U-Net-like structure to extract high-frequency feature information and output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and perform fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features;
[0012] S105. Perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.
[0013] Preferably, the S101 specifically includes:
[0014] Obtain a video sequence composed of the original target frame and several frames before and after the target frame , the video sequence is converted into a feature map with C channels through the Unet feature extraction module built based on the convolutional neural network; where, is the original target frame, and N is the number of front and back frames;
[0015] The feature map with C channels is subjected to feature alignment through deformable convolution, and multiple connected convolutional kernels with the number of input and output channels of 2×C, a convolutional kernel of 3, a padding of 1, and a stride of 1 are used for feature reconstruction. Finally, a convolutional kernel with an input channel number of 2×C, an output channel number of 1, a convolutional kernel of 3, a padding of 1, and a stride of 1 is used to obtain a single residual map ; where, H and W respectively represent the height and width of the feature, represents the set of real numbers.
[0016] Preferably, in the S103, the processing process of multiple Swin-Transformer blocks is as follows:
[0017] The low-frequency component features are expanded to channels through a pointwise convolution, and then downsampled once through a convolution to obtain , and features are extracted through the first Swin-Transformer block, and then downsampled once again to obtain ; for , features are extracted through the second Swin-Transformer block to obtain ; then, through the feature refinement module and the third Swin-Transformer block, the feature expression is further deepened, and the output is obtained; and are concatenated, and the intermediate layer feature expression ability is improved through the residual connection method, and is obtained through a convolution and passing through the fourth Swin-Transformer block; is concatenated with , and through a convolution and the fifth Swin-Transformer block, an upsampling operation is performed using a transposed convolution to obtain ; and are concatenated, and again through a convolution, and through the sixth Swin-Transformer block and an upsampling operation to obtain , and is sent into the multi-channel attention module; where, H and W respectively represent the height and width of the feature, represents the set of real numbers.
[0018] Preferably, in S103, the processing process of the multi-channel attention module is as follows:
[0019] In the multi-channel attention module, first, layer normalization is used to ensure the stable distribution of data features in each channel. Then, pointwise convolution and depthwise separable convolution are used to fuse features and the result is separated into Q, K, and V, that is , , , and the self-attention score is obtained through processing by simplifying the formula for calculating the attention mechanism, as follows:
[0020] ;
[0021] Among them, represents the calculation method of the attention mechanism, which is used to measure the correlation between different regions; Q is Query, representing the feature representation of the current pixel region; K is Key, representing the features of all regions in the image; V is Value, which is used to provide the feature information corresponding to Key; represents the normalization processing of the similarity result calculated from Q and K;
[0022] The self-attention score passes through a pointwise convolution and LeakyReLU, and the output of the feature is obtained through weighted connection to the global filtering module.
[0023] Preferably, in S103, the processing process of the global filtering module is as follows:
[0024] Receive the output of the multi-channel attention module, first perform a layer normalization process, then expand the result of the layer normalization into two branches for calculation through a grouped convolution. The two branches are multiplied element by element to obtain a residual result, and then through a grouped convolution, and the input is added to obtain the final low-frequency reconstruction result .
[0025] Preferably, S104 specifically includes:
[0026] For the vertical high-frequency component feature , the horizontal high-frequency component feature and the diagonal high-frequency component feature , respectively, perform high-frequency feature information through the residual connection neural network of the Unet-like structure. First, for , and , perform a channel expansion once, expand the input one-dimensional feature map to layers to obtain , and , and then perform a downsampling through a convolution respectively to obtain the downsampled high-frequency information , and , and perform downsampling operation again to obtain downsampled high-frequency information through a convolution , and , and obtain enhanced features through multiple convolutional layers , and , and respectively concatenate with , and , and obtain fused features through a convolution , and , and then perform upsampling once through a transposed convolution to obtain , and , and concatenate with , and , and concatenate and obtain through a convolution and a transposed convolution , and , and concatenate with , and , and concatenate, and respectively obtain through a convolution , and , concatenate, and obtain three reconstructed high-frequency component features through several convolutions, corresponding to the vertical high-frequency component , the horizontal high-frequency component and the diagonal high-frequency component .
[0027] In the second aspect, a spatio-temporal frequency domain compressed video enhancement device based on the Swin-Transformer architecture includes:
[0028] A multi-frame feature extraction and reconstruction module, configured to obtain a video sequence composed of an original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps; perform feature alignment on the multi-channel feature maps using deformable convolutions and reconstruct them into residual maps;
[0029] A high and low frequency component feature acquisition module, configured to perform Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features;
[0030] The low-frequency component feature reconstruction module is used to extract and reconstruct the low-frequency component features by sequentially connecting multiple Swin-Transformer blocks, a multi-channel attention module, and a global filtering module, and obtain the reconstructed low-frequency component features;
[0031] The high-frequency component feature reconstruction module is used to extract high-frequency feature information of the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features respectively through a residual connection neural network with a Unet-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and fuse and process the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features;
[0032] The target frame enhancement module is used to perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.
[0033] In a third aspect, an electronic device includes:
[0034] One or more processors;
[0035] A storage device for storing one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.
[0037] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements any of the above-mentioned spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.
[0038] In a fifth aspect, a computer program product includes a computer program, and when the computer program is executed by a processor, it implements any of the above-mentioned spatio-temporal frequency domain compressed video enhancement methods based on the Swin-Transformer architecture.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] (1) The present invention designs an image frequency domain processing architecture based on Swin-Transformer. The stacked multiple Swin-Transformer blocks can effectively extract the low-frequency component feature information of the global structure; the high-frequency component features are extracted through a residual connection neural network with a Unet-like structure, which can retain and enhance the texture information in the image;
[0041] (2) In the present invention, pointwise convolution is added between Swin-Transformer blocks and residual fusion is performed from shallow to deep, and filtering and global attention mechanisms are added at the end, enabling effective utilization of low-frequency information; meanwhile, in the Swin-Transformer block, a window attention mechanism is added, which can enhance the capture of local texture information. Attention calculation is performed in blocks through the window attention mechanism, reducing the computational complexity of the method using Transformer. Description of the Drawings
[0042] Figure 1 It is a flowchart of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture according to an embodiment of the present invention;
[0043] Figure 2 It is an overall flowchart block diagram of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture according to an embodiment of the present invention;
[0044] Figure 3 It is a schematic diagram of the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of the multi-channel attention module in the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;
[0046] Figure 5 It is a schematic diagram of the global filtering module in the low-frequency component feature enhancement process architecture according to an embodiment of the present invention;
[0047] Figure 6 It is a schematic diagram of the high-frequency component feature enhancement process architecture according to an embodiment of the present invention;
[0048] Figure 7 It is a structural block diagram of the spatio-temporal frequency domain fusion compression video enhancement device based on the Swin-Transformer architecture according to an embodiment of the present invention;
[0049] Figure 8 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Embodiments
[0050] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0051] See Figure 1As shown in the figure, an air-frequency domain compressed video enhancement method based on the Swin-Transformer architecture in this embodiment includes the following steps.
[0052] S101: Obtain a video sequence composed of the original target frame and several frames before and after the target frame. Extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps. Use deformable convolution to perform feature alignment on the multi-channel feature maps and reconstruct them into residual maps.
[0053] S102: Perform Haar wavelet transform on the residual map to obtain high-frequency component features and low-frequency component features with half the width and height of the residual map. The high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features.
[0054] S103: Pass the low-frequency component features through a series of connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module for feature extraction and reconstruction to obtain the reconstructed low-frequency component features.
[0055] S104: Respectively pass the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a Unet-like structure to extract high-frequency feature information and output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information. Perform fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features.
[0056] S105: Perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map. Add the restored residual map to the original target frame to obtain the enhanced target frame.
[0057] See Figure 2 As shown in the figure, the air-frequency domain fusion compressed video enhancement method model based on the Swin-Transformer architecture in this embodiment includes three modules: a multi-frame feature extraction and reconstruction module, a Swin-Transformer low-frequency enhancement module, and a high-frequency component feature enhancement and fusion module. Figure 2In this method, first, the original low-quality target frame and several frames before and after it are obtained, and they are sent together into the Unet feature extraction network to convert the video frames into feature maps of specific channels. Then, the feature maps are aligned through deformable convolution and the residual map is reconstructed. Subsequently, it enters the Swin-Transformer frequency domain enhancement module and the high-frequency component feature enhancement fusion module. The residual map is decomposed by Haar wavelet transform to obtain high-frequency component features and low-frequency component features, which are enhanced through convolutional neural networks respectively. Finally, the reconstructed high- and low-frequency component features are obtained through the inverse Haar wavelet transform to get the enhanced high-quality target frame.
[0058] The process steps of the spatio-temporal frequency domain fusion compression video enhancement method based on the Swin-Transformer architecture are elaborated in detail as follows.
[0059] The specific implementation of the multi-frame feature extraction and reconstruction module in S101 is as follows.
[0060] First, a video sequence composed of the original target frame and several frames before and after the target frame is obtained , and the video sequence is transformed into a feature map with channel C through the Unet feature extraction module built based on the convolutional neural network; where is the original target frame, and N is the number of frames before and after.
[0061] Then, the feature map of C channels is aligned through deformable convolution, and feature reconstruction is performed using a series of connected convolutional kernels with the number of input and output channels being 2×C, the convolutional kernel size being 3, the padding being 1, and the stride being 1. Finally, a single residual map is obtained using a convolutional kernel with the number of input channels being 2×C, the number of output channels being 1, the convolutional kernel size being 3, the padding being 1, and the stride being 1 , which is used as the input of the Haar wavelet transform; where H and W respectively represent the height and width of the feature, represents the set of real numbers.
[0062] See Figure 3 As shown, the Swin-Transformer low-frequency enhancement module includes multiple Swin-Transformer blocks, a multi-channel attention module, and a global filtering module. The input of the Swin-Transformer low-frequency enhancement module is the low-frequency component features obtained after Haar wavelet decomposition . The window size of each Swin-Transformer block is 8, and it includes an embedding layer, a window attention layer, a multi-layer perceptron layer, a normalization layer, and a de-embedding layer. In order to obtain image information of different sizes and expand the receptive field, first, the low-frequency component passes through an output channel number of , pointwise convolution with a convolution kernel of 1 and a stride of 1 expands the original number of channels to , and then performs a downsampling through a convolution with a convolution kernel of 3, a stride of 2, and a padding of 1 to obtain , and extracts features through a Swin-Transformer block, and then performs another downsampling of the same kind to obtain . For , use the features extracted through a Swin-Transformer block to obtain , and then pass through a feature extraction module, which includes a depthwise convolution with an input and output number of channels of , a convolution kernel of 3, a stride of 1, and a number of groups of , a pointwise convolution with an input and output number of channels of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block to further deepen the feature expression and obtain the output . Concatenate and , and improve the feature expression ability of the intermediate layer through residual connection. Through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block to obtain . Concatenate with , and through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and a Swin-Transformer block, and perform an upsampling operation using a transposed convolution with a convolution kernel of 4, a stride of 2, and a padding of 1 to obtain , and concatenate and , and again through a convolution with an input channel number of , an output channel number of , a convolution kernel of 3, a stride of 1, and through the last Swin-Transformer block and an upsampling operation of the same kind to obtain , and send into the multi-channel attention module.
[0063] See Figure 4 As shown, in the multi-channel attention module, first use layer normalization to ensure the stable distribution of data features in each channel and prevent overfitting, and then use pointwise convolution and depthwise separable convolution with an output layer number of to fuse features and separate the results into Q, K, and V, that is, , , , and by simplifying the formula of the computational attention mechanism, the model can effectively utilize global information while maintaining a low computational complexity, as follows:
[0064] ;
[0065] Among them, represents the computational method of the attention mechanism, which is used to measure the correlation between different regions; Q is the Query, representing the feature representation of the current pixel region; K is the Key, representing the features of all regions in the image; V is the Value, which is used to provide the feature information corresponding to the Key; represents the normalization process of the similarity result calculated from Q and K, so as to generate attention weights and ensure that the sum of all weights is 1.
[0066] The finally calculated self-attention score is obtained. Finally, the score is passed through a pointwise convolution and LeakyReLU, and the feature is output to the global filtering module through a weighted connection method.
[0067] See Figure 5 As shown, in the global filtering module, first, a layer normalization process is performed, and then the result of the layer normalization is expanded into two branches for calculation through a grouped convolution. Each branch contains a pointwise convolution and a depthwise separable convolution, and LeakyRelu is used as the activation layer in one of the layers. The residual result is obtained by multiplying the two branches element by element, and then a grouped convolution with a convolution kernel of 3 is performed, and the input is added to obtain the final low-frequency reconstruction result .
[0068] As Figure 6 shown, the input of the high-frequency component feature enhancement and fusion module is the three high-frequency components of the Haar wavelet, namely the vertical high-frequency component , the horizontal high-frequency component , and the diagonal high-frequency component . The feature of each of the three components is processed through a residual connection neural network similar to Unet, and an output feature map is obtained respectively. In the residual connection network similar to Unet, for the vertical high-frequency information , the horizontal high-frequency information , and the diagonal high-frequency information obtained by the Haar wavelet decomposition, first, a channel expansion is performed on , , and . The input one-dimensional feature map is expanded to layers, and is obtained, and and then perform downsampling once through a convolution with a convolution kernel of 3, padding of 1, stride of 2, and the number of input and output channels being to obtain downsampled high-frequency information 、 and and perform downsampling again. Obtain downsampled high-frequency information through a convolution with a convolution kernel of 3, padding of 1, stride of 2, and the number of input and output channels both being 、 and and obtain enhanced features through multiple convolutional layers with a convolution kernel of 3, padding of 1, stride of 1, and the number of input and output channels both being 、 and and concatenate them with 、 and respectively, and obtain fused features 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of 、 and and then perform upsampling once through a transposed convolution with a convolution kernel of 4, padding of 1, stride of 2, and the number of input and output channels both being to obtain 、 and and concatenate them with 、 and respectively, and then obtain 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of and a transposed convolution with a convolution kernel of 4, padding of 1, stride of 2, and the number of input and output channels both being 、 and and concatenate them with 、 and respectively, and obtain 、 through a convolution with a convolution kernel of 3, padding of 1, stride of 1, input channels of 、 and respectively, and concatenate them, and pass through N convolutional layers with a convolution kernel of 3, padding of 1, stride of 1, and input channels of , the number of output channels is After convolution, finally, through a convolution with a kernel size of 3, padding of 1, stride of 1, and the number of input channels being and the number of output channels being 3, three high-frequency feature reconstruction feature maps are finally obtained, corresponding to the vertical high-frequency component , the horizontal high-frequency component and the diagonal high-frequency component , and combined with the previous low-frequency reconstruction result to obtain the finally enhanced high-quality video frame through the inverse Haar wavelet transform.
[0069] See Figure 7 As shown, the present invention also discloses a spatio-temporal frequency domain compressed video enhancement device based on the Swin-Transformer architecture, including:
[0070] The multi-frame feature extraction and reconstruction module 701 is used to obtain a video sequence composed of the original target frame and several frames before and after the target frame, extract multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and convert them into multi-channel feature maps; use deformable convolution to perform feature alignment on the multi-channel feature maps and reconstruct them into residual maps;
[0071] The high and low frequency component feature acquisition module 702 is used to perform Haar wavelet transform on the residual map to obtain the high-frequency component features and low-frequency component features of the width and height dimensions of the residual map; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features;
[0072] The low-frequency component feature reconstruction module 703 is used to perform feature extraction and reconstruction on the low-frequency component features through a plurality of sequentially connected Swin-Transformer blocks, a multi-channel attention module, and a global filtering module to obtain the reconstructed low-frequency component features;
[0073] The high-frequency component feature reconstruction module 704 is used to respectively perform high-frequency feature information extraction on the vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features through a residual connection neural network with a Unet-like structure to output vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and perform fusion processing on the vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features;
[0074] The target frame enhancement module 705 is used to perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain the restored residual map, and add the restored residual map to the original target frame to obtain the enhanced target frame.
[0075] The specific implementation of each module of a spatio-temporal frequency domain compressed video enhancement device based on the Swin-Transformer architecture is the same as that of a spatio-temporal frequency domain compressed video enhancement method based on the Swin-Transformer architecture, and will not be repeated in this embodiment.
[0076] See Figure 8 The following shows a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 8 In this embodiment, the electronic device includes: a processor 801 and a memory 802; wherein the memory 802 is used to store computer execution instructions; the processor 801 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0077] Optionally, the memory 802 can be either independent or integrated with the processor 801.
[0078] When the memory 802 is independently provided, the electronic device further includes a bus 803 for connecting the memory 802 and the processor 801.
[0079] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 801 executes the computer execution instructions, the above method is implemented.
[0080] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 801, the above method is implemented.
[0081] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0082] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0083] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing unit, or each module can exist physically alone, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0084] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 801 to execute some steps of the methods in various embodiments of the present application.
[0085] It should be understood that the above processor 801 can be a central processing unit (Central Processing Unit, abbreviated as CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 801 can also be any conventional processor 801, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 801, or implemented by the combination of hardware and software modules in the processor 801.
[0086] The memory 802 may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0087] The bus 803 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 803 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus 803 in the drawings of the present application is not limited to only one bus 803 or one type of bus 803.
[0088] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0089] An exemplary storage medium is coupled to the processor 801, enabling the processor 801 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor 801. The processor 801 and the storage medium can be located in an application specific integrated circuit (ASIC).
[0090] In this embodiment, Python 3.8 and PyTorch 1.13 are used to conduct tests on the hardware platform. The configuration of the hardware platform is such as Intel i9-13900K (32G memory), NVIDIA RTX4090 graphics card, and Ubuntu operating system. The test data set is the MFQEv2 data set, and the test videos are 18 natural videos, divided into five categories: Category A (2560×1600), Category B (1920×1080), Category C (832×480), Category D (480×240), and Category E (1280×720). The number of frames of these videos ranges from 150 to 600. And compression is performed under different quantization parameters. The quantization parameters are set to 22, 27, 32, 37, and 42, and a test video with a QP value of 37 is used for detailed analysis. The parameters for comparison are the peak signal-to-noise ratio ΔPSNR (dB) and the structural similarity ΔSSIM (10 -2 )), and the comparison target is the difference from the video compressed under the original platform (H.265 / HEVC).
[0091] Table 1 Comparison between the existing enhancement method and the method of the present invention;
[0092]
[0093] The test results are shown in Table 1. It can be seen from Table 1 that in terms of the objective indicators ΔPSNR and ΔSSIM, the method of the present invention is superior to other enhancement methods.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A space-frequency domain compressed video enhancement method based on Swin-Transformer architecture, characterized in that: include: S101, obtaining a video sequence consisting of an original target frame and several frames before and after the target frame, extracting multi-frame video features in the video sequence based on a convolutional neural network with a Unet-like structure and converting them into a multi-channel feature map; aligning the multi-channel feature map using deformable convolution and reconstructing it into a residual map; S102, performing Haar wavelet transform on the residual image to obtain high-frequency component features and low-frequency component features of half the width and height of the residual image; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features, and diagonal high-frequency component features; S103, extracting and reconstructing the low-frequency component features through a plurality of Swin-Transformer blocks, a multi-channel attention module and a global filtering module connected in sequence to obtain reconstructed low-frequency component features; S104, extracting high-frequency feature information of the vertical high-frequency component features, the horizontal high-frequency component features, and the diagonal high-frequency component features through a residual connection neural network with a Unet-like structure, outputting vertical high-frequency feature information, horizontal high-frequency feature information, and diagonal high-frequency feature information, and fusing the vertical high-frequency feature information, the horizontal high-frequency feature information, and the diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain reconstructed high-frequency component features; S105, performing inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain a restored residual image, and adding the restored residual image to the original target frame to obtain an enhanced target frame.
2. The method for enhancing video in space-frequency domain compression based on Swin-Transformer architecture according to claim 1, characterized in that: The S101 specifically includes: Get a video sequence consisting of the original target frame and several frames before and after the target frame , the video sequence is extracted by the Unet feature extraction module based on the convolutional neural network Transformed into a feature map with channel C; where is the original target frame, N is the number of previous and next frames; The feature map of the C channel is aligned through deformable convolution, and multiple connected convolution kernels with input and output channels of 2×C, convolution kernel of 3, padding of 1 and stride of 1 are used for feature reconstruction. Finally, a single residual map is obtained using a convolution kernel with input channels of 2×C, output channels of 1, convolution kernel of 3, padding of 1 and stride of 1. ; where H and W represent the height and width of the feature, respectively. Represents the set of real numbers.
3. The method for enhancing video in space-frequency domain compression based on Swin-Transformer architecture according to claim 1, characterized in that: In S103, the processing process of multiple Swin-Transformer blocks is as follows: The low frequency component features After a point-by-point convolution, the number of channels is expanded to , and then downsampled once through a convolution to get , and extract features through the first Swin-Transformer block, and then go through the same downsampling to get ;right Using the features extracted by the second Swin-Transformer block ; Then the feature extraction module and the third Swin-Transformer block further deepen the feature expression and get the output ;Will and The feature expression ability of the intermediate layer is improved by concatenating them, and the fourth Swin-Transformer block is used to obtain ;Will and Concatenate, pass through a convolution and the fifth Swin-Transformer block, and use a transposed convolution to perform an upsampling operation to obtain ;Will and Concatenate, pass through a convolution again, and pass through the sixth Swin-Transformer block and an upsampling operation to obtain ,Will Sent to the multi-channel attention module; where H and W represent the height and width of the feature respectively, Represents the set of real numbers.
4. The method for enhancing video in space-frequency domain compression based on Swin-Transformer architecture according to claim 3, characterized in that: In S103, the processing process of the multi-channel attention module is as follows: In the multi-channel attention module, layer normalization is first used to ensure the stable distribution of data features of each channel, and then point-by-point convolution and depth-separable convolution are used to fuse features and separate the results into Q, K, and V, that is, , , , and the self-attention score is obtained by simplifying the formula for calculating the attention mechanism, as follows: ; in, Indicates the calculation method of the attention mechanism, which is used to measure the correlation between different regions; Q is Query, which represents the feature representation of the current pixel area; K is Key, which represents the features of all regions in the image; V is Value, which is used to provide feature information corresponding to the Key; Indicates that the similarity results calculated by Q and K are normalized; The self-attention score is passed through a point-by-point convolution and LeakyReLU, and the feature output is obtained by weighted connection to the global filtering module.
5. The method for enhancing space-frequency domain compressed video based on Swin-Transformer architecture according to claim 4, characterized in that: In S103, the processing process of the global filtering module is as follows: The output of the multi-channel attention module is received and first normalized by a layer. Then, a group convolution is performed to expand the result of the layer normalization into two branches for calculation. The two branches are element-wise multiplied to obtain the residual result, which is then convolved by a group convolution and added to the input to obtain the final low-frequency reconstruction result. .
6. The method for enhancing video in space-frequency domain compression based on Swin-Transformer architecture according to claim 1, characterized in that: The S104 specifically includes: Vertical high frequency component characteristics , horizontal high frequency component characteristics And the diagonal high frequency component characteristics The high-frequency feature information is obtained by using a residual connection neural network with a Unet-like structure. , and Perform a channel expansion to expand the input one-dimensional feature map to layer, get , and , and then downsample once through a convolution to obtain the downsampled high-frequency information , and , and downsample again, and get the downsampled high frequency information through a convolution , and , and through multiple convolutional layers, the enhanced features are obtained , and , and respectively with , and Concatenate and obtain the fused features through a convolution , and , and then upsampled once through a transposed convolution to obtain , and , and with , and Concatenate and pass through a convolution and a transposed convolution to get , and , and with , and Splice and pass a convolution to get , and , and then concatenate them. Through several convolutions, we get three reconstructed high-frequency component features, corresponding to the vertical high-frequency components. , horizontal high frequency component And the diagonal high frequency components .
7. A space-frequency domain compressed video enhancement device based on Swin-Transformer architecture, characterized in that: include: The multi-frame feature extraction and reconstruction module is used to obtain a video sequence consisting of the original target frame and several frames before and after the target frame. The convolutional neural network based on the Unet structure extracts the multi-frame video features in the video sequence and converts them into multi-channel feature maps; the deformable convolution is used to align the multi-channel feature maps and reconstruct them into residual maps; A high- and low-frequency component feature acquisition module is used to perform Haar wavelet transform on the residual image to obtain high-frequency component features and low-frequency component features of half the width and height of the residual image; the high-frequency component features include vertical high-frequency component features, horizontal high-frequency component features and diagonal high-frequency component features; A low-frequency component feature reconstruction module is used to extract and reconstruct the low-frequency component features through a plurality of Swin-Transformer blocks, a multi-channel attention module and a global filtering module connected in sequence to obtain the reconstructed low-frequency component features; A high-frequency component feature reconstruction module is used to extract high-frequency feature information of vertical high-frequency component features, horizontal high-frequency component features and diagonal high-frequency component features through a residual connection neural network with a Unet-like structure, output vertical high-frequency feature information, horizontal high-frequency feature information and diagonal high-frequency feature information, and fuse the vertical high-frequency feature information, horizontal high-frequency feature information and diagonal high-frequency feature information through a multi-layer convolutional neural network to obtain the reconstructed high-frequency component features; The target frame enhancement module is used to perform inverse Haar wavelet transform on the reconstructed high-frequency component features and the reconstructed low-frequency component features to obtain a restored residual image, and add the restored residual image to the original target frame to obtain an enhanced target frame.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image super-resolution method and device based on cross attention mechanism and Swin-Transform
CN117237197A
Video super-resolution method and system based on deformable Transformer
CN117593185A
Low-light image enhancement method and device based on wavelet transform and retinex-net
US20250078230A1