A lightweight social platform image restoration method, system and readable storage medium based on dual-branch frequency and space fusion
By employing a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, and utilizing generative adversarial networks trained with generators and discriminators, the high computational complexity and limited reconstruction performance of image restoration on social media platforms are addressed, achieving efficient and accurate image restoration results.
Patent Information
- Application Number
- CN202411664926.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-20
AI Technical Summary
Existing image restoration methods on social media platforms suffer from high computational complexity, long inference time, and limited performance improvement, especially in blind super-resolution techniques where accuracy and model training efficiency need to be improved.
A lightweight social platform image restoration method based on dual-branch frequency and spatial fusion is adopted. Low-resolution image patches are generated using a pre-set second-order degradation model. The image restoration process is optimized by combining a feature fusion network in the frequency and spatial domains and training a generative adversarial network with a generator and a discriminator.
It achieves a smoother and more efficient image restoration experience on mobile devices, improves image clarity and detail recovery capabilities, stabilizes the training process, and enhances the robustness and accuracy of the model.
Smart Images

Figure CN119599915B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image restoration technology, and more specifically, relates to a lightweight social platform image restoration method, system, and readable storage medium based on dual-branch frequency and spatial fusion. Background Technology
[0002] With the widespread adoption of social media platforms, people increasingly use handheld devices to communicate with friends, share updates, and engage in interactive activities. However, when these images are uploaded to social media platforms, they are often affected by various compression artifacts and unpredictable noise degradation, leading to a significant decline in image quality. Therefore, how to achieve excellent visual effects while effectively reducing computational complexity and accelerating inference time, providing users with a smoother and more efficient user experience on mobile devices, has become an important and meaningful task. In recent years, numerous super-resolution (SR) models based on convolutional neural networks (CNNs) have explored various methods to reduce model complexity. For example, FSRCNN and ESPCN employ post-upsampling strategies to predefine inputs, thereby reducing computational burden. ShuffleMixer reduces model complexity and latency by utilizing large kernel convolutions and channel shuffling techniques. BSRN enhances model capabilities by replacing standard convolutions with carefully designed depthwise separable convolutions and integrating two effective attention mechanisms. Spatial Adaptive Feature Hybrid Networks develop an efficient SR model by integrating the principles of convolution and self-attention. While these methods have made substantial progress in model efficiency, there is still potential for further improvement in reconstruction performance. Balancing computational efficiency with performance improvement is an important and promising research area.
[0003] Blind super-resolution (BSR) is also a key method for solving this type of problem. Blind super-resolution generally aims to recover high-resolution images from low-resolution images with unknown degradation. It does not require prior knowledge of the specific type or extent of image degradation, can handle complex and diverse degradation scenarios, and plays an important role in image restoration and visual quality improvement, especially suitable for scenarios where the degradation process cannot be accurately modeled. Most methods address this problem either explicitly or implicitly: implicit methods aim to learn degradation networks directly from real-world low-resolution images, while explicit methods attempt to approximate real low-resolution images by designing synthetic degradation processes. BSRGAN and Real-ESRGAN design complex degradation pipelines to synthesize degraded images, thereby enhancing the generalization ability of existing SR models, demonstrating significant progress in the field of blind super-resolution.
[0004] The existing invention patent with publication number CN118822846A proposes a method for improving lightweight social platform image restoration based on dual-branch frequency and spatial fusion, belonging to the field of lightweight social platform image restoration based on dual-branch frequency and spatial fusion. The method includes: selecting a suitable image dataset; dividing the dataset according to an appropriate ratio; improving the shallow feature extraction module; improving the deep feature extraction module accordingly, employing a multi-scale feature extraction module to extract features from the feature map, enhancing the network's feature extraction capability; using an improved efficient channel attention network to calculate the importance of each channel in the feature map, improving the algorithm's feature representation capability; employing a multi-level feature fusion mechanism to strengthen the network's utilization of intermediate layer feature information, enhancing the algorithm's reconstruction capability; setting relevant training parameters, training the model, and saving the weights; loading the training weights with the minimum loss function on the validation set, and testing the model's reconstruction performance. This scheme can improve the quality of reconstructed images to a certain extent, obtaining more realistic super-resolution images. However, its accuracy and the efficiency of model training still need improvement. Summary of the Invention
[0005] To overcome the problem of how to improve the accuracy and efficiency of image restoration networks in the prior art, this invention provides a lightweight social platform image restoration method, system, and readable storage medium based on dual-branch frequency and spatial fusion.
[0006] The first aspect of this invention provides a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, comprising the following steps:
[0007] The training data is cropped into small image patches, and the small image patches are processed using a pre-defined second-order degradation model to obtain low-resolution image patches, which are used as the training dataset.
[0008] A lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is constructed by using a pre-defined lightweight restoration network as the generator and a frequency-domain-based discriminative gating network as the discriminator.
[0009] A lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is trained using a training dataset. The generator is used to convert the image to be processed into a high-quality image. The discriminator is used to discriminate the high-quality images generated by the generator during model training, and the discrimination result is fed back and used to adjust the generator.
[0010] Low-quality images are input into a lightweight social platform image restoration network model trained on dual-branch frequency and spatial fusion to restore high-quality images.
[0011] Furthermore, the preset second-order degradation model includes: two blurring modules, two size adjustment modules, two noise-adding modules, and two hybrid compression modules. The hybrid compression module is constructed by combining the JPEG compression method and a prediction-based algorithm with the Real-ESRGAN synthetic degradation algorithm to build the hybrid compression module; and reducing the strength of the default parameters of the Real-ESRGAN algorithm in the hybrid compression module to obtain the adjusted hybrid compression module.
[0012] Processing small image patches using a pre-defined second-order degradation model includes the following steps:
[0013] Small image blocks are sequentially input into the first blur module, the first size adjustment module, and the first noise module, and the noise-added image blocks are output. The noise-added image blocks are then input into the first hybrid compression module, and the processed small image blocks are output.
[0014] The processed small image blocks are sequentially input into the second blur module, the second size adjustment module, and the second noise-adding module to obtain a noisy image block. The noisy image block is then input into the second hybrid compression module to output a low-resolution image block.
[0015] Furthermore, the preset lightweight restoration network includes: a convolutional kernel, a number of parallel dual-branch frequency and spatial domain basis feature fusion modules, and an upsampling module. The image to be processed is sequentially passed through the convolutional kernel, the number of parallel dual-branch frequency and spatial domain basis feature fusion modules, and the upsampling module to output a high-quality image.
[0016] Furthermore, the dual-branch frequency and spatial basis feature fusion module includes a parallel-running spatial adaptive feature hybrid network and a frequency basis discriminative gated network. The spatial adaptive feature hybrid network is used to extract spatial features of image features, and the frequency basis discriminative gated network is used to extract frequency features of image features.
[0017] Furthermore, the method for converting the image to be processed into a high-quality image includes the following steps:
[0018] The image to be processed is input into the convolution kernel for convolution, and the primary features of the image are output.
[0019] The primary image features are input into a spatial adaptive feature mixing network, which outputs spatial features in the spatial domain; the primary image features are input into a frequency domain basis discriminant gated network, which outputs frequency features in the frequency domain.
[0020] By fusing the spatial features in the spatial domain and the frequency features in the frequency domain, a fused feature is obtained.
[0021] The fused features are input into the upsampling module, and residual connections are used to capture high-frequency details of the image to obtain a high-quality image.
[0022] Furthermore, the extraction of frequency domain features from primary image features using a frequency domain basis discriminant gating network includes the following steps:
[0023] The primary features X of the input image are subjected to a 1×1 pointwise convolution to double the number of channels, and the spatial local context is encoded using a 3×3 depthwise convolution. The expression is as follows:
[0024] X1 = DW_Conv 3x3 (Conv 1x1 (X))
[0025] Where X1 represents the features after 1x1 convolution and 3x3 depthwise convolution, and Conv 1x1 (·) represents a 1×1 pointwise convolution, DW_Conv 3x3 (·) represents a 3×3 depthwise convolution;
[0026] The features are transformed to the frequency domain, and then the Fast Fourier Transform is used to segment the features into two independent components along the channel dimension, as expressed by:
[0027] X F =F(X1)
[0028]
[0029] Among them, X F This represents the features transformed into the frequency domain using a Fast Fourier Transform. F(·) represents the features after segmentation along the channel dimension, F(·) represents the Fast Fourier Transform, and Split(·) represents the channel segmentation operation.
[0030] For a component after splitting A 1×1 pointwise convolution is applied, and the convolution result is used as a gate for the second component to selectively preserve relevant frequency information. The expression is as follows:
[0031]
[0032] Where ⊙ represents pixel-wise dot product. This represents the characteristics output after the above steps;
[0033] The GELU activation function is applied for nonlinear mapping, and the expression is:
[0034]
[0035] Among them, X out F represents the frequency characteristics of the final output. -1 (·) represents the inverse fast Fourier transform, and G(·) represents the GELU function.
[0036] Furthermore, the method for constructing the frequency domain-based discriminative gating network is as follows: a frequency domain-based discriminative gating module is added between the convolution module and the layer normalization module in the discriminator based on the U-net structure, and the layer normalization module is replaced by the spectral normalization module.
[0037] Furthermore, the training strategy of the lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is based on the training strategy of generative adversarial networks.
[0038] A second aspect of the present invention provides a lightweight social platform image restoration system based on dual-branch frequency and spatial fusion, comprising a memory and a processor. The memory includes a program for a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. When the processor executes the program for the lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion.
[0039] A third aspect of the present invention provides a computer-readable storage medium comprising a program for a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. When the program is executed by a processor, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion.
[0040] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0041] This invention designs a feature fusion network with dual-branch frequency and spatial domain bases as the generator for a lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion. It integrates a frequency-domain gated dynamic network and spatial domain feature selection techniques, effectively exploring both frequency and spatial features simultaneously to achieve superior image restoration results and significantly restore image clarity. Furthermore, this invention incorporates a discriminator into the lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion. The discriminator evaluates the accuracy and efficiency of the generator's image restoration, providing feedback on the results. Based on this feedback, adjustments are made to both the generator and discriminator, effectively stabilizing the training dynamics and improving the model's robustness and ability to restore detailed textures. Attached Figure Description
[0042] To make the objectives and technical solutions of this invention clearer, the following drawings are provided and described:
[0043] Figure 1A flowchart of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion provided for embodiments of the present invention;
[0044] Figure 2 A flowchart of synthetic degradation data based on a second-order degradation model provided in an embodiment of the present invention;
[0045] Figure 3 This is a schematic diagram of a feature fusion module for dual-branch frequencies and spatial domain bases provided in an embodiment of the present invention;
[0046] Figure 4 This is a schematic diagram of a frequency domain discriminative gating network provided in an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of a frequency domain-based discriminator provided in an embodiment of the present invention;
[0048] Figure 6 The diagram shows the effect of a qualitative comparison between the dual-branch frequency and spatial fusion network and the blind super-resolution model provided in the embodiment of the present invention. Detailed Implementation
[0049] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0050] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0051] Example 1:
[0052] This invention provides a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, such as... Figure 1 The diagram shows a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. The specific steps are as follows:
[0053] S1: The training data is cropped into small image patches, and then the corresponding low-resolution (LR) image patches are generated using the second-order degradation model as the training dataset.
[0054] In one specific embodiment, the training sets of DIV2K, Flickr2K, and 600 collected images are used as training data. HR images are cropped into 256×256 pixel image patches, and the corresponding low-resolution image patches are generated using the second-order degradation model proposed in this invention. Through this step, synthetic low-quality images with a realistic appearance are generated for training and optimizing the image restoration model.
[0055] More specifically, the preset second-order degradation model includes: two blurring modules, two size adjustment modules, two noise-adding modules, and two hybrid compression modules. The hybrid compression module is constructed by combining the JPEG compression method and a prediction-based algorithm with the Real-ESRGAN synthetic degradation algorithm to build the hybrid compression module; reducing the strength of the default parameters of the Real-ESRGAN algorithm in the hybrid compression module to avoid excessive blurring or noise in the synthesized degradation image, thus obtaining the adjusted hybrid compression module.
[0056] Processing small image patches using a pre-defined second-order degradation model includes the following steps:
[0057] Small image blocks are sequentially input into the first blur module, the first size adjustment module, and the first noise module, and the noise-added image blocks are output. The noise-added image blocks are then input into the first hybrid compression module, and the processed small image blocks are output.
[0058] The processed small image blocks are sequentially input into the second blur module, the second size adjustment module, and the second noise-adding module to obtain a noisy image block. The noisy image block is then input into the second hybrid compression module to output a low-resolution image block.
[0059] It should be noted that the DIV2K dataset is a high-quality dataset specifically designed for image super-resolution tasks and is widely used in research and development in the field of computer vision. This dataset contains 800 high-resolution (HR) training images as the training set and 100 high-resolution validation images as the validation set. Each image has extremely high sharpness, making it ideal for training and evaluating super-resolution algorithms.
[0060] Flickr2K is a high-quality dataset for image restoration tasks, containing 2650 2K resolution images. These images are primarily used for training and evaluating super-resolution algorithms, particularly for research on converting low-resolution images to high-resolution images.2 The Flickr2K dataset, provided by Flickr, is a large, expanded dataset containing 800 training images, 100 validation images, and 100 test images.
[0061] Real-ESRGAN stands for Enhanced Super-Resolution Generative Adversarial Network.
[0062] It should be noted that this invention provides a second-order degradation model, and the process for synthesizing degradation data based on the second-order degradation model is as follows: Figure 2 As shown, this invention first combines JPEG compression methods and prediction-based algorithms to construct a broad compression degradation space, termed "hybrid compression." Furthermore, the strength of default parameters in Real-ESRGAN is reduced to avoid excessive blurring or noise in synthesized degraded images. Additionally, a dataset of 650 high-quality selfie images is collected, enhancing the network's robustness when processing selfie images from social media platforms.
[0063] S2: Using a pre-defined lightweight restoration network as the generator and a frequency-domain-based discriminative gating network as the discriminator, a lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is constructed.
[0064] More specifically, since not all frequency-domain features contribute to image restoration, exploring effective frequency information is crucial for improving image restoration quality. To effectively distinguish and utilize these key frequency features, this invention provides a frequency-domain-based discriminative gating network, i.e., a frequency-domain-based discriminator, which can effectively determine the low-frequency and high-frequency information that should be retained when restoring potentially sharp images. Specifically, the frequency-domain-based discriminator involves adding a frequency-domain-based discriminative gating module between the convolutional module and the layer normalization module in a U-net-based discriminator, and replacing the layer normalization module with a spectral normalization module to form the frequency-domain-based discriminator. Figure 5 As shown.
[0065] It should be noted that in the field of existing generative adversarial network super-resolution technology, although most methods can generate realistic images that conform to human perception, generating high-quality images requires a discriminator that can provide accurate gradient feedback to ensure the restoration of details in local image textures. Traditional methods typically use U-Net-based discriminators to evaluate realism at the pixel level and provide detailed pixel-level feedback to the generator. However, directly applying such discriminators may lead to instability during training, mainly due to the difference in model capacity between the generator and the discriminator, causing the generator to be unable to effectively keep up with the discriminator's ability to distinguish realism. To alleviate training instability, this invention attempts to reduce the number of channels in each convolutional layer of the discriminator, but this weakens the discriminator's ability to accurately identify real details. Therefore, this invention proposes a frequency domain-based discriminator network, such as... Figure 4 As shown, this network is designed to capture subtle differences in the frequency domain during image restoration, thereby more accurately evaluating the quality of image detail restoration. Through fine-tuning of frequency features, the frequency-domain-based discriminator significantly enhances the model's performance in image detail discrimination, thus improving the stability and effectiveness of the entire GAN system in generating high-quality super-resolution images.
[0066] The preset lightweight restoration network includes: a convolutional kernel, several parallel dual-branch frequency and spatial feature fusion modules, and an upsampling module. The dual-branch frequency and spatial feature fusion module includes a parallel-running spatial adaptive feature hybrid network and a frequency domain basis discriminative gating network.
[0067] More specifically, the method for generating a high-quality image from an image to be processed includes the following steps:
[0068] The image to be processed is input into the convolution kernel for convolution, and the primary features of the image are output.
[0069] The primary image features are input into a spatial adaptive feature mixing network, which outputs spatial features in the spatial domain; the primary image features are input into a frequency domain basis discriminant gated network, which outputs frequency features in the frequency domain.
[0070] By fusing the spatial features in the spatial domain and the frequency features in the frequency domain, a fused feature is obtained.
[0071] The fused features are input into the upsampling module, and residual connections are used to capture high-frequency details of the image to obtain a high-quality image.
[0072] It should be noted that, in order to fully utilize the advantages of spatial feature interaction, this invention proposes a feature fusion module with dual-branch frequencies and spatial domain basis, such as... Figure 3As shown, this module integrates a spatial adaptive feature fusion network with a frequency domain basis discriminative gated network. The spatial adaptive feature fusion network is a lightweight network module that learns long-range dependencies from multi-scale feature representations to better explore features more useful for high-quality image reconstruction. The dual-branch frequency and spatial basis feature fusion module of this invention enables the simultaneous extraction of spatial and frequency domain features to achieve a more comprehensive image restoration effect.
[0073] More specifically, the method of extracting frequency domain features from primary image features using a frequency domain basis discriminant gating network, such as... Figure 4 As shown, it includes the following steps:
[0074] The primary features X of the input image are subjected to a 1×1 pointwise convolution to double the number of channels, and the spatial local context is encoded using a 3×3 depthwise convolution. The expression is as follows:
[0075] X1 = DW_Conv 3x3 (Conv 1x1 (X))
[0076] Where X1 represents the features after 1x1 convolution and 3x3 depthwise convolution, and Conv 1x1 (·) represents a 1×1 pointwise convolution, DW_Conv 3x3 (·) represents a 3×3 depthwise convolution;
[0077] The features are transformed to the frequency domain, and then the Fast Fourier Transform is used to segment the features into two independent components along the channel dimension, as expressed by:
[0078] X F =F(X1)
[0079]
[0080] Among them, X F This represents the features transformed into the frequency domain using a Fast Fourier Transform. F(·) represents the features after segmentation along the channel dimension, F(·) represents the Fast Fourier Transform, and Split(·) represents the channel segmentation operation.
[0081] A 1×1 pointwise convolution is applied to one of the segmented components, and the convolution result is used as a gate for the second component to selectively retain relevant frequency information. The expression is as follows:
[0082]
[0083] Where ⊙ represents pixel-wise dot product. This represents the characteristics output after the above steps;
[0084] The GELU activation function is applied for nonlinear mapping, and the expression is:
[0085]
[0086] Among them, X out F represents the frequency characteristics of the final output. -1 (·) represents the inverse fast Fourier transform, and G(·) represents the GELU function.
[0087] S3: The lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is trained using the training dataset. The generator is used to convert the image to be processed into a high-quality image. The discriminator is used to discriminate the high-quality images generated by the generator during model training, and the discrimination results are fed back and used to adjust the generator.
[0088] More specifically, the training strategy of the lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is based on the training strategy of generative adversarial networks.
[0089] The training is based on a lightweight social platform image restoration network model that fuses frequency and spatial data in a dual-branch manner, and includes the following steps:
[0090] S3.1: Train the generator and discriminator using the training dataset.
[0091] A generator with a dual-branch frequency and spatial basis feature fusion module is trained using a training dataset. First, features are extracted through 3x3 convolution. The spatial adaptive feature fusion network dynamically selects representative image features in the spatial domain. The frequency basis discriminative gating network selects frequency information that is helpful for image restoration. Finally, the two types of information are fused to better assist in subsequent image restoration work.
[0092] Then, in the second stage, a frequency-domain-based discriminative gating network is used to train the generator, improving the quality and detail recovery capabilities of the generated images and stabilizing the training process.
[0093] S3.2: Collect data for model testing and build a test set.
[0094] S3.3: Input the test set into the lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion to test the generator's restoration effect.
[0095] In one specific embodiment, 100 DIV2K validation images and 60 additional high-quality selfie images collected by us were used as the test set. These were transmitted via WeChat to obtain real LR images. These were then used to test and evaluate the effectiveness of our method against other methods, providing both qualitative and quantitative comparisons.
[0096] In one specific embodiment, PSNR (Peak signal-to-noise ratio) and SSIM (structural similarity) are used to measure the fidelity of image reconstruction, while LPIPS (Learned Perceptual Image Patch Similarity), NIQE (Natural Image Quality Evaluator), and MUSIQ (Multi-scale Image Quality Transformer) are used to evaluate perceptual quality.
[0097] In addition, this embodiment uses the number of model parameters (#Params), the number of floating-point operations (#Flops), and the average running time (#Avg.Time) to verify the model efficiency.
[0098] S4: Input low-quality images into the trained lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion to restore high-quality images.
[0099] In a specific embodiment, such as Figure 6 As shown in the figure, on the DIV2K dataset, after the synthesis degradation process of this invention, the dual-branch frequency and spatial fusion network (DBFSNet) and the existing blind super-resolution model were qualitatively compared. From the images, it can be seen that the image restoration effect of the method described in this invention is more accurate.
[0100] This invention first develops a dual-branch frequency and spatial fusion network (DBFSNet). This network integrates a frequency-domain gated dynamic network and a spatial-domain feature selection technique, effectively exploring both frequency and spatial features simultaneously to achieve superior image restoration and significantly restore image clarity. Secondly, addressing the training instability issue caused by traditional GAN discriminators in lightweight networks, this invention provides a lightweight frequency-domain discriminator, improving the model's training stability and detail recovery capabilities. Furthermore, this invention adjusts the second-order degradation model for synthetic degradation data to better simulate real-world social media image degradation scenarios. A small dataset suitable for this invention was collected, and experiments verified the efficiency and accuracy of the proposed method.
[0101] The lightweight social platform image restoration model based on dual-branch frequency and spatial fusion provided by this invention achieves restoration results that are competitive with advanced large-scale model solutions in terms of both quality and fidelity. This invention achieves the best results at the lowest cost, striking a good balance between model efficiency and restoration effectiveness.
[0102] Example 2:
[0103] This embodiment provides a lightweight social platform image restoration system based on dual-branch frequency and spatial fusion, including a memory and a processor. The memory includes a program for a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. When the processor executes the program for the lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion as described in Embodiment 1.
[0104] Example 3:
[0105] This embodiment provides a computer-readable storage medium, which includes a program for a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. When the program is executed by a processor, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion as described in Embodiment 1.
[0106] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, characterized in that, Includes the following steps: The training data is cropped into small image patches, and the small image patches are processed using a pre-defined second-order degradation model to obtain low-resolution image patches, which are used as the training dataset. A lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is constructed by using a pre-defined lightweight restoration network as the generator and a frequency-domain-based discriminative gating network as the discriminator. A lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is trained using a training dataset. The generator is used to convert the image to be processed into a high-quality image. The discriminator is used to discriminate the high-quality images generated by the generator during model training, and the discrimination result is fed back and used to adjust the generator. Low-quality images are input into a lightweight social platform image restoration network model trained on dual-branch frequency and spatial fusion to restore high-quality images. The preset second-order degradation model includes: two blurring modules, two size adjustment modules, two noise-adding modules, and two hybrid compression modules. The hybrid compression module is constructed by combining the JPEG compression method and a prediction-based algorithm with the Real-ESRGAN synthetic degradation algorithm to build the hybrid compression module; and reducing the strength of the default parameters of the Real-ESRGAN algorithm in the hybrid compression module to obtain the adjusted hybrid compression module. The preset lightweight restoration network includes: a convolutional kernel, several parallel dual-branch frequency and spatial domain basis feature fusion modules, and an upsampling module; The dual-branch frequency and spatial basis feature fusion module includes a parallel-running spatial adaptive feature mixing network and a frequency basis discriminative gated network. The spatial adaptive feature mixing network is used to extract the spatial features of the primary features of the image, and the frequency basis discriminative gated network is used to extract the frequency features of the primary features of the image. The method for constructing the frequency domain-based discriminative gating network is as follows: a frequency domain-based discriminative gating module is added between the convolution module and the layer normalization module in the discriminator based on the U-net structure, and the layer normalization module is replaced by the spectral normalization module.
2. The lightweight social platform image restoration method based on dual-branch frequency and spatial fusion according to claim 1, characterized in that, Processing small image patches using a pre-defined second-order degradation model includes the following steps: Small image blocks are sequentially input into the first blur module, the first size adjustment module, and the first noise module, and the noise-added image blocks are output. The noise-added image blocks are then input into the first hybrid compression module, and the processed small image blocks are output. The processed small image blocks are sequentially input into the second blur module, the second size adjustment module, and the second noise-adding module to obtain a noisy image block. The noisy image block is then input into the second hybrid compression module to output a low-resolution image block.
3. The lightweight social platform image restoration method based on dual-branch frequency and spatial fusion according to claim 1, characterized in that, A method for processing images using a pre-defined lightweight restoration network includes the following steps: The image to be processed is input into the convolution kernel, and the primary features of the image are output. The primary features of the image are input into several parallel feature fusion modules with dual-branch frequency and spatial domain basis, and the fused features are output. The fused features are input into the upsampling module to output a high-quality image.
4. The lightweight social platform image restoration method based on dual-branch frequency and spatial fusion according to claim 1, characterized in that, A method for converting an image to be processed into a high-quality image includes the following steps: The image to be processed is input into the convolution kernel for convolution, and the primary features of the image are output. The primary image features are input into a spatial adaptive feature mixing network, which outputs spatial features in the spatial domain; the primary image features are input into a frequency domain basis discriminant gated network, which outputs frequency features in the frequency domain. By fusing the spatial features in the spatial domain and the frequency features in the frequency domain, a fused feature is obtained. The fused features are input into the upsampling module, and residual connections are used to capture high-frequency details of the image to obtain a high-quality image.
5. The lightweight social platform image restoration method based on dual-branch frequency and spatial fusion according to claim 4, characterized in that, Extracting frequency features in the frequency domain from primary image features using a frequency domain basis discriminant gated network includes the following steps: The primary features X of the input image are subjected to a 1×1 pointwise convolution to double the number of channels, and the spatial local context is encoded using a 3×3 depthwise convolution. The expression is as follows: X1=DW_Conv 3x3 (Conv. 1x1 (X)) Where X1 represents the features after 1x1 convolution and 3x3 depthwise convolution, and Conv 1x1 (·) represents a 1×1 pointwise convolution, DW_Conv 3x3 (·) represents a 3×3 depthwise convolution; The features are transformed to the frequency domain, and then the Fast Fourier Transform is used to segment the features into two independent components along the channel dimension, as expressed by: X F =F(X1) Among them, X F This represents the features transformed into the frequency domain using a Fast Fourier Transform. F(·) represents the features after segmentation along the channel dimension, F(·) represents the Fast Fourier Transform, and Split(·) represents the channel segmentation operation. For a component after splitting A 1×1 pointwise convolution is applied to obtain the convolution result, which is then used as a gate for the second component to selectively retain relevant frequency information. The expression is as follows: Where ⊙ represents pixel-wise dot product. This represents the characteristics output after the above steps; The GELU activation function is applied for nonlinear mapping, and the expression is: Among them, X out F represents the frequency characteristics of the final output. -1 (·) represents the inverse fast Fourier transform, and G(·) represents the GELU function.
6. The lightweight social platform image restoration method based on dual-branch frequency and spatial fusion according to claim 1, characterized in that, The training strategy of the lightweight social platform image restoration network model based on dual-branch frequency and spatial fusion is based on the training strategy of generative adversarial networks.
7. A lightweight social platform image restoration system based on dual-branch frequency and spatial fusion, characterized in that, The device includes a memory and a processor. The memory includes a program for a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion. When the processor executes the program for the lightweight social platform image restoration method based on dual-branch frequency and spatial fusion, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a lightweight social platform image restoration method program based on dual-branch frequency and spatial fusion. When the lightweight social platform image restoration method program based on dual-branch frequency and spatial fusion is executed by a processor, it implements the steps of a lightweight social platform image restoration method based on dual-branch frequency and spatial fusion as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method for improving image super-resolution
CN118822846A
Image processing method and device, electronic equipment and readable storage medium
CN111899177A
Underwater structure disease identification method and system, and storage medium
CN117011688A