A method for realizing face super-resolution based on a practical degradation model

By adopting a practical degradation model and a dual-branch attention mechanism in the super-resolution of face images, we simulate multiple image degradation effects and extract image features, solving the problem of insufficient generalization ability and robustness in the prior art, achieving higher quality image reconstruction and better adaptability.

CN118799189BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411287556.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-06-13
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

When processing face images in real scenes, it is difficult for the prior art to fully capture the complexity of image degradation, resulting in insufficient generalization and robustness of super-resolution algorithms.

Method used

The method based on practical degradation model is used to simulate the image degradation effect in actual scenes, including Gaussian noise, Rayleigh noise, motion blur, salt and pepper noise and mean blur, and the degradation space is expanded through random strategies. At the same time, a dual-branch attention mechanism is designed to extract and fuse the global and local features of the image through the multi-head self-attention mechanism and the double-layer routing attention mechanism.

Benefits of technology

It significantly improves the super-resolution reconstruction effect of face images, enhances the robustness and adaptability of the model, and can better handle image degradation in real and complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799189B_ABST
    Figure CN118799189B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for realizing face super-resolution based on a practical degradation model, which relates to the technical fields of digital image processing and pattern recognition. The method uses a practical degradation model to simulate the image degradation effect in the actual scene to obtain a simulated image; uses convolutional blocks with different convolutional kernels to extract the feature information of different scales of the simulated image to obtain hierarchical features at different scales; uses a dual-branch network structure to perform deep feature extraction on the hierarchical features at different scales, fuses the extracted feature information, and performs a convolutional operation through an SCConv convolutional block to obtain a fused feature map, thereby completing the reconstruction of the face image. This method not only improves the reconstruction quality and visual effect of the image, but also has good robustness and adaptability, and can be widely applied to face image processing and enhancement tasks in actual scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing and pattern recognition, and more specifically, to a method for realizing face super-resolution based on a practical degradation model. Background Art

[0002] Face super-resolution (FSR) technology aims to reconstruct a low-resolution (LR) face image into a high-resolution (HR) image through algorithms, which is crucial for applications such as face recognition and face segmentation. Due to shooting conditions, device limitations, and compression during transmission, the acquired face images often suffer from insufficient resolution, leading to an increase in the complexity of subsequent processing tasks.

[0003] Early FSR methods mainly relied on traditional image processing techniques, such as upsampling techniques like bicubic interpolation and bilinear interpolation. Although these methods improved the image resolution to a certain extent, they often failed to effectively restore the detailed information of the image, especially in terms of image blurring and noise problems.

[0004] With the development of statistical modeling and learning inference techniques, methods such as Bayesian methods, convex optimization, and principal component analysis (PCA) have been introduced into FSR to improve the effect of image reconstruction. However, these traditional methods still have the problem of limited representation ability when dealing with face images in complex scenarios.

[0005] In recent years, the rise of deep learning techniques, especially the development of convolutional neural networks (CNNs), has brought new breakthroughs to face super-resolution. Researchers have proposed a variety of CNN-based network models to achieve super-resolution reconstruction of face images. These methods significantly improve the quality of the reconstructed image by learning the mapping relationship between low-resolution and high-resolution images. At the same time, face prior information, such as facial landmarks and heatmaps, is also integrated into the model to enhance the reconstruction effect of the face contour.

[0006] Nevertheless, the existing technologies still face challenges when dealing with face images in real-world scenarios. The degradation of face images in the real world is usually the result of the combined action of multiple factors, including noise, blurring, compression artifacts, etc. Existing degradation models often use fixed degradation kernels or simple degradation processes, and these models cannot comprehensively capture the complexity of image degradation in the real world. In addition, a single degradation method or a model with fixed parameters is difficult to adapt to the changing actual scenarios, restricting the generalization ability and robustness of super-resolution algorithms. Summary of the Invention

[0007] In view of this, the present invention provides a method for face super-resolution based on a practical degradation model. By simulating the image degradation effects in the actual scenario, including Gaussian noise, Rayleigh noise, motion blur, salt-and-pepper noise, and mean blur, etc., different degradation processes are selected by a random strategy. In addition, random parameters are adopted during the degradation process to adjust the degree of degradation, expanding the degradation space and improving the model upper limit. The present invention also includes a dual-branch attention mechanism, where the upper branch uses a multi-head self-attention mechanism to capture global and local correlation information, and the lower branch uses a two-layer routing attention mechanism to refine key features. The combination of the two improves the robustness of super-resolution reconstruction.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A method for face super-resolution based on a practical degradation model, comprising the following steps:

[0010] Utilize the practical degradation model to simulate the image degradation effect in the actual scenario to obtain a simulated image;

[0011] Use convolutional blocks with different convolutional kernels to extract feature information of different scales of the simulated image to obtain hierarchical features at different scales;

[0012] Use a dual-branch network structure to perform deep feature extraction on the hierarchical features at different scales, fuse the extracted feature information, and perform a convolutional operation through an SCConv convolutional block to obtain a fused feature map, completing the reconstruction of the face image.

[0013] Optionally, the practical degradation model selects different degradation processes by a random strategy, including Gaussian noise, Rayleigh noise, motion blur, salt-and-pepper noise, and mean blur.

[0014] Optionally, in the upper branch of the dual-branch structure, a multi-head attention model is adopted.

[0015] Optionally, a two-layer routing attention mechanism is adopted in the lower branch.

[0016] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides a method for face super-resolution based on a practical degradation model, having the following beneficial effects:

[0017] 1. The method of the present invention is reasonably designed, convenient to implement, and has good migration.

[0018] 2. The practical degradation model proposed by the present invention is more in line with the degradation of face images in the real scenario and is more suitable for the super-resolution network of face images.

[0019] 3. The face super-resolution of the present invention has achieved good quantitative and qualitative results on widely used Helen, SCUT_FBP, and PFHQ datasets, and has broad practical application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0021] Figure 1 It is the network structure diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0023] The embodiment of the present invention discloses a method for realizing face super-resolution based on a practical degradation model, including the following steps:

[0024] Use the practical degradation model to simulate the image degradation effect in the actual scene to obtain a simulated image;

[0025] Use convolutional blocks with different convolutional kernels to extract the feature information of different scales of the simulated image to obtain hierarchical features at different scales;

[0026] Use a double-branch network structure to perform deep feature extraction on the hierarchical features at different scales, fuse the extracted feature information, and perform a convolutional operation through spatial channel reconstruction convolution (SCConv) to obtain a fused feature map, thereby completing the reconstruction of the face image.

[0027] Use the practical degradation model to degrade the high-resolution face image to simulate the image distortion effect in the actual scene. The degradation process includes: using a random combination of various noises and degradation methods to more accurately simulate the face image degradation effect in the real world. Specifically, the present invention uses 5 representative degradation methods, namely Gaussian noise, Rayleigh noise, motion blur, salt and pepper noise, and mean blur.

[0028] Due to the influence of factors such as electronic components and transmission media, Gaussian noise is a common interference noise in images. In the method of the present invention, three types of Gaussian noise are selected, namely, color Gaussian noise, grayscale Gaussian noise, and multi-dimensional Gaussian noise. The present invention determines which type of Gaussian noise to add according to the generated random number N. For color Gaussian noise and grayscale Gaussian noise, the present invention uses the same mean and standard deviation:

[0029] (1);

[0030] where x represents the noise pixel value, represents the mean, represents the standard deviation. For multi-dimensional Gaussian noise, the present invention introduces more variable factors. First, a 3×3 random orthogonal matrix and a 3×3 matrix with random numbers on the diagonal are randomly generated. Then, a covariance matrix is calculated through the combination of these matrices for generating multi-dimensional Gaussian noise.

[0031] (2);

[0032] where X represents the noise vector, represents the mean vector, represents the covariance matrix, k represents the dimension of the noise vector, represents the determinant of the covariance matrix.

[0033] Rayleigh noise occupies an important position in the field of image processing, especially in images acquired in special environments. The presence of this noise is usually closely related to situations such as long-distance shooting, low-light conditions, or wireless sensor networks. The characteristics of Rayleigh noise are that it has a non-negative amplitude, and when the amplitude is 0, its probability density is also 0. Its probability density function shows an exponentially decaying trend, where the probability of larger amplitude values is lower. The present invention uses the built-in functions of PyTorch to generate noise following the Rayleigh distribution and adds it to the image according to the specified intensity parameter S. This method provides a good degradation effect when simulating the noise impact in special image processing scenarios. This process can be represented by formula (3):

[0034] (3);

[0035] where, represents the input image tensor, S represents the specified Rayleigh noise intensity parameter, and R represents the random tensor generated following the standard normal distribution.

[0036] Motion blur is a blurring effect caused by relative motion between an object or device and the scene being photographed during image capture. Motion blur is mainly divided into two categories: camera motion blur and object motion blur, both of which have a negative impact on image quality. First, camera motion blur is caused by the movement or jitter of the camera during exposure, and this blurring effect makes the entire image appear blurred or stretched of the object or scene. Second, object motion blur is caused by the movement of the object being photographed during exposure. This causes the details in the image to become blurred. The present invention comprehensively considers these two types of motion blur to achieve a more accurate motion blur effect. The present invention generates a motion blur kernel matrix and uses a convolution operation to achieve the effect of motion blur. In addition, the processed image is also normalized. Among them, the generation of the motion blur kernel can be represented by formula (4):

[0037] (4);

[0038] where D represents the size of the blur kernel and M represents the affine transformation matrix.

[0039] In digital image processing, sudden bright or dark spots often appear on the image. This is a common and significant type of noise, namely salt-and-pepper noise, also known as impulse noise. This kind of noise is usually caused by sensor failures, data transmission errors or other problems in the image acquisition and processing links. Compared with Gaussian noise, salt-and-pepper noise is more extreme, causing some pixel values to become the brightest or darkest, and thus forming obvious black and white spots in the image. In the task of face image super-resolution, the preservation and enhancement of color information are crucial. Therefore, by introducing salt-and-pepper noise with colored spots, the present invention simulates the complex noise situations that face images may encounter in real scenarios, thereby improving the robustness of the super-resolution model of the present invention. Specifically, the present invention independently adds salt-and-pepper noise to each channel of the face image, controls the introduction probabilities of pepper noise and salt noise to ensure that the generated noise shows a moderate random distribution in the image. The density of the introduced salt-and-pepper noise is the key factor affecting the noise intensity of the generated image, and a higher density will make the noise more obvious in the image.

[0040] Mean blur is a classic technique in the field of image processing, which mainly achieves image smoothing by applying a convolution kernel with uniform weights to the image. Specifically, the present invention first creates a convolution kernel with uniform weights, and the elements of its weight matrix are all 1 to achieve the purpose of uniform distribution. The size of the convolution kernel is specified manually, which in turn determines the degree of blur. After verification, the present invention selects a 5×5 convolution kernel. Then, the present invention performs a convolution operation on the input image while adjusting the padding to keep the spatial dimensions unchanged. As shown in formula (5):

[0041] (5);

[0042] Where I is the input image, O is the output image, and W is the weight matrix of the convolutional kernel, with a size of (m, n). is the pixel value at position (i, j) in the output image.

[0043] Due to the complexity and diversity of degradation situations in the real environment, the single degradation methods adopted in previous works are insufficient in simulating real-world degradation. Therefore, the present invention introduces a method of random degradation in the face image super-resolution task based on a random shuffling strategy. This strategy aims to better simulate the image quality fluctuations in the real environment, including noise, blur, and other possible interference factors. Specifically, the degradation process of each LR image is random. To improve the generalization ability of the degradation effect, multiple degradation strategies are used for each LR image. At the same time, the hyperparameters of each degradation process are not fixed. For example, in Gaussian noise, the present invention sets a random value N to determine the addition of different Gaussian noises to the image. When this random value N (ranging from 0 to 1) is greater than 0.6, colored Gaussian noise is selected to be added. When it is less than 0.4, gray Gaussian noise is added. In other cases, multi-dimensional Gaussian noise is added. Through this random degradation strategy, the generalization ability of the model can be greatly enhanced, and the robustness of the image super-resolution algorithm in actual scenarios can be improved.

[0044] For HR images, the present invention uses a random degradation strategy to generate corresponding LR images. These LR images are all generated under the guidance of variable hyperparameters and are all derived from multiple degradation methods.

[0045] To achieve a balance between network operation efficiency and performance, the present invention selects three ordinary convolutional blocks as the shallow feature extraction module. Specifically, the input original LR image passes through three convolutional blocks with different convolutional kernel sizes respectively, and feature information of different scales is extracted. The three convolutional blocks are K5N64S2P2, K3N32S2P1, and K3N16S2P1, where K represents the convolutional kernel size, N represents the output channels, S represents the stride, and P represents the padding. The whole process is shown in the following formula:

[0046] (6);

[0047] Where, represents the input original LR image, and H, W, and C represent the height, width, and channels of the image respectively. F h 、F m 、F l represent features of three different scales respectively. Then, each shallow feature map passes through a GELU activation function. The process can be represented by the following formula (7):

[0048] (7);

[0049] Through the shallow feature extraction stage, the present invention obtains hierarchical features at three different scales {F h , F m , F l}. After that, they are sent to the next stage, namely the deep feature extraction stage with a dual-branch structure.

[0050] Different from the networks used for image restoration and super-resolution in the past, the present invention uses a unique dual-branch network structure for deep feature extraction. Specifically, the feature information at three levels extracted by the shallow feature extraction stage is fused twice respectively to be delivered to the deep feature extraction stage. The present invention selects to downsample F h so as to perform pointwise addition fusion with F m ; downsample F m so as to perform pointwise addition fusion with F l . After that, the fused feature maps are respectively subjected to convolution operations through SCConv convolution blocks. Finally, the fused feature maps and are obtained. This process can be expressed by formula (8):

[0051] (8);

[0052] In the upper branch, the present invention adopts a multi-head self-attention module (Multi-Head Self-Attention, MHSA) to enhance the model's ability to understand the correlation relationship between channels. At the same time, it captures the long-distance dependence relationship between different positions, enabling the model to better understand the global structure. This process can be expressed by the following formula (9):

[0053] (9);

[0054] In the lower branch, a bi-level routing attention mechanism (Bi-Level Routing Attention, BRA) is adopted.

[0055] F low passes through the BRA module, and the model extracts valuable information. This helps to improve the global attention efficiency and enables the model to focus more on the features of important regions. This process can be expressed by the following formula (10):

[0056] (10);

[0057] When the deep feature information is respectively extracted from the upper and lower branches, the extracted feature information and Perform the same point - by - point addition fusion operation as before. In the present invention, the feature information of the two branches is selected to be added element - by - element, and then fused through a 3×3 ordinary convolution. At the same time, the residual connection method is adopted to improve the feature representation ability and reusability, enabling the network to better utilize low - level information, thereby improving the model's perception ability of multi - scale and multi - level features.

[0058] The whole process can be represented by the following formula (11):

[0059] (11);

[0060] In addition, as Figure 1 shown, the reconstruction module is alternately composed of three convolutional layers and two sub - pixel convolutional layers for reconstructing the HR face image.

[0061] Experiments are carried out on the Helen, SCUT_FBP and PFHQ datasets for verification.

[0062] The peak signal - to - noise ratio (PSNR), structural similarity index (SSIM) and learnable perceptual image patch similarity (LPIPS) are used as evaluation metrics to evaluate the reconstruction effect of the model.

[0063] Through the above - mentioned specific implementation manners, the technical solution of the present invention can significantly improve the super - resolution reconstruction effect of face images and has the robustness to process real and complex scenes. Specifically, the present invention has achieved significant technological progress in the following aspects:

[0064] First, a practical degradation model is adopted to simulate the common image distortion effects in the actual scene, including Gaussian noise, Rayleigh noise, motion blur, salt - and - pepper noise and mean blur. This makes the training data more diverse and representative, ensuring that the model can show good adaptability and robustness when processing low - quality images in the real world.

[0065] Second, the dual - branch attention super - resolution network structure designed in the present invention fully exploits the global and local features in the image through the multi - head self - attention mechanism and the two - layer routing attention mechanism. The multi - head self - attention mechanism can capture the long - range dependencies of the image, while the two - layer routing attention mechanism focuses on the extraction and enhancement of key features. The collaborative work of this dual - branch structure has significantly improved the details and overall perception of the reconstructed image.

[0066] Furthermore, the experimental results show that the present invention performs excellently on multiple publicly available datasets such as Helen, SCUT_FBP, and PFHQ. Under different degradation methods, the model of the present invention can effectively cope with them, showing relatively high peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learnable perceptual image patch similarity (LPIPS). This indicates that the technical solution of the present invention can not only handle a single type of image degradation but also has the ability to cope with various complex degradation situations.

[0067] In summary, by innovatively combining a practical degradation model and a dual-branch attention super-resolution network, the present invention proposes an effective method for face image super-resolution reconstruction. This method not only improves the reconstruction quality and visual effect of the image but also has good robustness and adaptability, and can be widely applied to face image processing and enhancement tasks in practical scenarios.

[0068] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0069] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for realizing face super-resolution based on a practical degradation model, characterized in that: The following steps are involved: Using a practical degradation model to simulate image degradation effects in actual scenes, a simulated image is obtained; Convolution blocks with different convolution kernels are used to extract feature information of different scales of the simulated image and obtain hierarchical features at different scales; The dual-branch network structure is used to extract deep features from hierarchical features at different scales, and the extracted feature information is fused and convolved through the SCConv convolution block to obtain the fused feature map and complete the face image reconstruction. Add salt and pepper noise independently to each channel of the face image, and control the probability of introducing pepper noise and salt noise to ensure that the generated noise presents a moderate random distribution in the image; In the upstream branch of the dual-branch structure, a multi-head attention model is used; A two-layer routing attention mechanism is adopted in the downstream branch.

2. The method for realizing face super-resolution based on a practical degradation model according to claim 1, characterized in that: The practical degradation model selects different degradation processes with a random strategy, including Gaussian noise, Rayleigh noise, motion blur, salt and pepper noise, and mean blur.

Citation Information

Patent Citations

  • High-resolution reconstruction method for remote sensing image under complex degradation model

    CN117474764A