Face image enhancement method for compensating Internet transmission degradation
By designing a face image enhancement model that includes shallow feature extraction, deep feature extraction and subpixel convolution reconstruction, the problem of degradation of face image quality in Internet transmission is solved, and high-quality image reconstruction and low computational complexity are achieved.
Patent Information
- Application Number
- CN202411542869.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-31
AI Technical Summary
During the Internet transmission process, face images are susceptible to degradation processing such as low resolution, blur and compression artifacts, resulting in a decline in image quality. It is difficult for the prior art to achieve a good balance between computing efficiency and image quality.
A face image enhancement model consisting of shallow feature extraction, deep feature extraction and subpixel convolution reconstruction is designed. Convolutional neural network and residual convolutional neural network are used for feature extraction and reconstruction, combining window multi-head self-attention mechanism and sliding window multi-head self-attention mechanism to achieve efficient feature conversion and high-quality image reconstruction.
High-quality facial image reconstruction is achieved, with high overall quality, real effects, fewer artifacts, relatively small calculations, small parameters, and low time complexity, suitable for deployment on mobile devices.
Smart Images

Figure CN119991473A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to a method for enhancing facial images for Internet transmission degradation. Background Art
[0002] During Internet transmission, facial images will be degraded by low resolution, blur and compression artifacts. The purpose of face image enhancement technology for Internet transmission degradation is to reconstruct low-quality face images caused by factors such as network bandwidth limitations, transmission delays, and device performance differences during Internet application transmission into high-quality ones. The possible degradation factors are mainly low resolution, noise, blur and compression artifacts. This technology is widely used in various Internet software transmission low-quality face image enhancement tasks.
[0003] For the task of face image enhancement, past work can be divided into two categories. The first category mainly uses deep learning networks, and uses skip connections in convolutional neural networks to introduce residuals into the model to speed up the convergence of deeper network training, or uses generative adversarial networks to obtain better texture features, or improves the global feature learning ability of the model based on the Transformer structure. However, as the image reconstruction quality improves, the network depth and model complexity also increase significantly, making computational efficiency gradually become a serious problem and challenge. The huge number of parameters and computational complexity are not conducive to deployment on mobile platforms. The other category combines the characteristics of face images to perform special processing on them, such as adding facial prior information as branch constraints to the network, iterative collaboration between face key point calibration and image reconstruction, and establishing a deep face dictionary network. However, in order to further improve the subjective visual effect of reconstructed face images, most of the current face image enhancement methods have integrated face fantasy technology. Some face images have obvious differences in facial features compared to the original face images, and even seem to be a fictional face that does not exist. Summary of the invention
[0004] The purpose of the present invention is to propose a facial image enhancement method that compensates for the degradation of Internet transmission, and to design a facial image enhancement model composed of shallow feature extraction, deep feature extraction and sub-pixel convolution reconstruction to reconstruct low-quality images compressed during network transmission.
[0005] In one aspect, the present invention provides a method for enhancing a facial image for compensating for Internet transmission degradation, comprising:
[0006] S1: Input low-quality face images degraded via the Internet;
[0007] S2: Construct a convolutional neural network and use the constructed convolutional neural network to perform shallow feature extraction on low-quality images;
[0008] S3: construct a residual convolutional neural network, and use the constructed residual convolutional neural network to extract local and global deep features. The residual convolutional neural network includes two residual jump connection modules and a convolutional layer. The shallow feature extraction output features are used as input, and the deep feature extraction output features are obtained through the residual convolutional neural network. The two residual jump connection modules included in the residual convolutional neural network are a jump connection module with a window multi-head self-attention mechanism and a jump connection module with a sliding window multi-head self-attention mechanism. The calculation results of the input features obtained by the two residual jump connection modules are convolved through the convolutional layer to obtain the output features of the deep feature extraction part.
[0009] S4: Construct a sub-pixel convolutional network, using the fusion of the output of the shallow feature extraction obtained by S2 and the output of the deep feature extraction obtained by S3 as input to achieve the conversion of features from low resolution to high resolution, and obtain the sub-pixel convolution reconstructed output image.
[0010] In some embodiments, the jump connection module with the window multi-head self-attention mechanism further includes a window multi-head self-attention block, three multi-layer perceptron blocks connected in series, and the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected. The input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block. The output of the window multi-head self-attention block, the output of the first multi-layer perceptron block, and the input / output addition result of the window multi-head self-attention block are added as window result one, and result one is obtained by the first The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, and the output of the second multi-layer perceptron block are added together as the window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, the window result two, and the output of the third multi-layer perceptron block are added together as the window result three;
[0011] Among them, the jump connection module with the sliding window multi-head self-attention mechanism further includes a sliding window multi-head self-attention block, three multi-layer perceptron blocks connected in series, the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected, the input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block, the output of the window multi-head self-attention block, the output of the first multi-layer perceptron block and the input / output addition result of the window multi-head self-attention block are added and processed as sliding window result one, and the sliding window result one is composed of the first multi-layer perceptron block. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, result one, and the output of the second multi-layer perceptron block are added together as sliding window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, sliding window result one, sliding window result two, and the output of the third multi-layer perceptron block are added together as sliding window result three.
[0012] In some embodiments, the sub-pixel convolutional network includes a first convolutional layer, a LeakyReLU activation function, two sub-pixel upsampling modules and a second convolutional layer connected in series, and the sub-pixel upsampling module consists of a convolutional layer and a pixel reorganization block.
[0013] In some embodiments, the residual convolutional neural network further includes a loss function The formula is as follows:
[0014]
[0015] Among them, I RE is a high-resolution image reconstructed through network training, I HR The original high-resolution image.
[0016] In some embodiments, according to the method for enhancing facial images to compensate for Internet transmission degradation according to claim 1, it is characterized in that the deep feature extraction output feature is shown in the following formula:
[0017]
[0018] F D =D XONV (F2),
[0019] Among them, F irepresents the image features after the i-th RSSTB module, Indicates the operation of the ith RSSTB module, D CONv represents the convolutional layer operation in deep feature extraction D, F D It represents the output after deep feature extraction D.
[0020] In another aspect, the present invention provides a facial image enhancement system for compensating for Internet transmission degradation, comprising:
[0021] An input module, used for inputting low-quality face images degraded by transmission via the Internet;
[0022] A shallow feature extraction module is used to construct a convolutional neural network and use the constructed convolutional neural network to perform shallow feature extraction on low-quality images;
[0023] A deep feature extraction module constructs a residual convolutional neural network, and uses the constructed residual convolutional neural network to extract local and global deep features. The residual convolutional neural network includes two residual jump connection modules and a convolutional layer, and uses the shallow feature extraction output features as input, and obtains the deep feature extraction output features through the residual convolutional neural network; the two residual jump connection modules included in the residual convolutional neural network are respectively a jump connection module with a window multi-head self-attention mechanism and a jump connection module with a sliding window multi-head self-attention mechanism. After the calculation results of the input features obtained by the two residual jump connection modules are convolved through the convolutional layer, the output features of the deep feature extraction part are obtained;
[0024] The sub-pixel convolution reconstruction module is used to construct a sub-pixel convolution network, and uses the fusion of the output of the shallow feature extraction obtained by the shallow feature extraction module and the output of the deep feature extraction obtained by the deep feature extraction module as input to achieve the conversion of features from low resolution to high resolution, and obtain a sub-pixel convolution reconstructed output image.
[0025] In some embodiments, the jump connection module with the window multi-head self-attention mechanism further includes a window multi-head self-attention block, three multi-layer perceptron blocks connected in series, and the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected. The input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block. The output of the window multi-head self-attention block, the output of the first multi-layer perceptron block, and the input / output addition result of the window multi-head self-attention block are added as window result one, and result one is obtained by the first The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, and the output of the second multi-layer perceptron block are added together as the window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, the window result two, and the output of the third multi-layer perceptron block are added together as the window result three;
[0026] Among them, the jump connection module with the sliding window multi-head self-attention mechanism further includes a sliding window multi-head self-attention block, three multi-layer perceptron blocks connected in series, the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected, the input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block, the output of the window multi-head self-attention block, the output of the first multi-layer perceptron block and the input / output addition result of the window multi-head self-attention block are added and processed as sliding window result one, and the sliding window result one is composed of the first multi-layer perceptron block. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, result one, and the output of the second multi-layer perceptron block are added together as sliding window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, sliding window result one, sliding window result two, and the output of the third multi-layer perceptron block are added together as sliding window result three.
[0027] In some embodiments, the sub-pixel convolutional network includes a first convolutional layer, a LeakyReLU activation function, two sub-pixel upsampling modules and a second convolutional layer connected in series, and the sub-pixel upsampling module consists of a convolutional layer and a pixel reorganization block.
[0028] In some embodiments, using The loss function optimizes the network parameters, and the formula is as follows:
[0029]
[0030] Among them, I RE is a high-resolution image reconstructed through network training, I HR The original high-resolution image.
[0031] In some embodiments, the deep feature extraction output feature is shown as follows:
[0032]
[0033] F D =D XONV (F2),
[0034] Among them, F i represents the image features after the i-th RSSTB module, Indicates the operation of the ith RSSTB module, D CONv represents the convolutional layer operation in deep feature extraction D, F D It represents the output after deep feature extraction D.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1) The designed shallow feature extraction maps the input image to a higher-dimensional feature space. The deep feature extraction part captures the multi-dimensional local and global information of the image, providing better detail features for image reconstruction. The sub-pixel convolution reconstruction part realizes the conversion of features from low resolution to high resolution, and maps it from the high-dimensional feature space back to the three-color channel for observation.
[0037] 2) The overall quality of the image reconstructed by the model is high, the effect is more realistic, and there are fewer artifacts, which makes the human eye feel comfortable; the model has relatively less calculation and parameters, and has lower time complexity;
[0038] 3) We independently developed and designed a residual convolutional neural network, which includes two residual skip connection modules and a convolutional layer. This overcomes the complex computational complexity caused by the high similarity of the continuous intermediate layers MSA. The calculation results of the multi-head self-attention mechanism of the continuous intermediate layers have a high degree of similarity, achieving a good balance between excellent face image reconstruction effect and low parameter amount, computational complexity, and time complexity. It can be used to enhance face images degraded by Internet transmission, which is conducive to subsequent deployment in mobile devices such as mobile phones and tablets;
[0039] 4) It can be used to solve the degradation problems of facial images such as reduced resolution and increased noise during transmission in Internet applications due to factors such as network bandwidth limitations, transmission delays, and device performance differences;
[0040] 5) Bring more stable optimization results to the final image reconstruction through dimensionality increase operation;
[0041] 6) The sub-pixel convolutional network realizes the conversion of features from low resolution to high resolution through specific convolution operations and multi-channel reorganization of features, which can better preserve the detail information of the image; the final convolution layer maps the image from the high-dimensional feature space back to the three-color channels for observation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is an overall flow chart of a facial image enhancement method for compensating for Internet transmission degradation of the present invention;
[0043] Figure 2 is a block diagram of a face image enhancement model in a specific implementation manner of the present invention;
[0044] Figure 3 It is a comparative schematic diagram of the RSSTB module jump connection in a specific implementation manner of the present invention;
[0045] Figure 4 Schematic diagram of a residual convolutional neural network in a specific embodiment of the present invention;
[0046] Figure 5 Schematic diagram of the combined structure of ordinary convolution and sub-pixel convolution in a specific implementation manner of the present invention;
[0047] Figure 6 It is a schematic diagram of a sub-pixel convolution 2x upsampling module in a specific implementation manner of the present invention;
[0048] Reference numerals:
[0049] 1. Shallow feature extraction part, 2. Deep feature extraction part, 3. Sub-pixel convolution reconstruction part. DETAILED DESCRIPTION
[0050] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and embodiments.
[0051] like Figure 1 As shown, the implementation process of a face image enhancement method for compensating for Internet transmission degradation of the present invention specifically includes the following steps:
[0052] S1. Input low-quality face images that have been degraded through Internet transmission. For example, high-quality face images are degraded through WeChat software transmission to obtain a low-quality face image dataset and use it as input for the network model. The low-quality face dataset is composed of low-quality face images compressed through WeChat software transmission of the CelebAMask high-definition face image dataset, and is used for face image enhancement model training.
[0053] S2. Construct a convolutional neural network and use the constructed convolutional neural network to extract shallow features of the low-quality image. Specifically, the convolutional neural network in this step is constructed by a 3×3 convolutional layer, wherein the shallow feature extraction process includes taking a low-resolution face image I degraded by Internet applications LR ∈R H×W×C As the input feature of the 3×3 convolutional layer, H, W, and C are the height, width, and number of channels of the input image respectively. LR Mapping from a smaller dimension space to a higher dimensional feature space, the shallow feature extraction output feature F0 is obtained through the convolutional neural network, and the expression is as follows:
[0054] F0=S(I LR )
[0055] Among them, S() represents the shallow feature extraction operation function;
[0056] S3, construct a residual convolutional neural network, and use the constructed residual convolutional neural network to extract local and global deep features. Specifically, the residual convolutional neural network in this step includes two residual skip connection (SwinTransformer) modules (Residual Skip Swin Transformer Block, RSSTB) and a 3×3 convolutional layer, wherein the deep feature extraction process includes taking the shallow feature extraction output feature F0 as input, and obtaining the deep feature extraction output feature F through the residual convolutional neural network. i , the expression is as follows:
[0057]
[0058] F D =D CONV (F2),
[0059] in, Indicates the operation of the ith RSSTB module. The ith RSSTB module is used to operate the output features of the previous RSSTB module. CONV represents the convolutional layer operation in the deep feature extraction D, FD represents the output feature after the deep feature extraction D, and F i-1 represents the output feature after the i-1th residual module (RSSTB), and F2 represents the output feature of the second deep feature extraction;
[0060] like Figure 3 As shown in the figure, the Transformer module consists of a multi-layer perceptron block MLP and a multi-head self-attention block MSA, each with a normalization layer. The conventional connection method is that the input and output of the MSA are added together as the input of the MLP connected to it, and the output of the MLP is added together as the input of the next level MSA, and so on. However, research results show that the calculation results of the multi-head self-attention mechanism of consecutive intermediate layers are highly similar. In order to use this feature to reduce the amount of model calculation, the RSSTB module is redesigned as a skip connection method in this step; as shown in the figure. Figure 4 As shown, the two RSSTB modules in this step are jump connections with window multi-head self-attention mechanism ( sw inTransformer) modules and skip connections with sliding window multi-head self-attention mechanisms ( sw in Transformer) module, with a skip connection with a windowed multi-head self-attention mechanism ( swThe Transformer module further includes a window multi-head self-attention block (W-MSA), three serially connected multi-layer perceptron blocks, and the window multi-head self-attention block and the three serially connected multi-layer perceptron blocks are jump-connected. ① The input and output of the window multi-head self-attention block (W-MSA) are added and processed as the input of the first multi-layer perceptron block (MLP), and the output of the window multi-head self-attention block (W-MSA) is jump-connected to the output of the first multi-layer perceptron block (MLP). The output of the window multi-head self-attention block (W-MSA), the output of the first multi-layer perceptron block (MLP), and the input / output addition result of the window multi-head self-attention block (W-MSA) are added as window result one, and result one is output by the first multi-layer perceptron block (MLP); ② The first multi- ③ The output of the second multi-layer perceptron block (MLP) is used as the input of the second multi-layer perceptron block (MLP), the output of the window multi-head self-attention block (W-MSA) is jump-connected to the output of the second multi-layer perceptron block (MLP), and the output of the window multi-head self-attention block (W-MSA), the window result one and the output of the second multi-layer perceptron block (MLP) are added as the window result two; ③ The output of the second multi-layer perceptron block (MLP) is used as the input of the third multi-layer perceptron block (MLP), the output of the window multi-head self-attention block (W-MSA) is jump-connected to the output of the third multi-layer perceptron block (MLP), and the output of the window multi-head self-attention block (W-MSA), the window result one, the window result two and the output of the third multi-layer perceptron block (MLP) are added as the window result three. Similarly, the jump connection with the sliding window multi-head self-attention mechanism ( swThe Transformer module further includes a sliding window multi-head self-attention block (SW-MSA), three serially connected multi-layer perceptron blocks, and the window multi-head self-attention block and the three serially connected multi-layer perceptron blocks are jump-connected. ① The input and output of the window multi-head self-attention block (SW-MSA) are added and processed as the input of the first multi-layer perceptron block (MLP). The output of the window multi-head self-attention block (SW-MSA) is jump-connected to the output of the first multi-layer perceptron block (MLP). The output of the window multi-head self-attention block (SW-MSA), the output of the first multi-layer perceptron block (MLP), and the input / output addition result of the window multi-head self-attention block (SW-MSA) are added as the sliding window result one, and the sliding window result one is output by the first multi-layer perceptron block (MLP); ② The first The output of the multilayer perceptron block (MLP) is used as the input of the second multilayer perceptron block (MLP), the output of the window multi-head self-attention block (SW-MSA) is jump-connected to the output of the second multilayer perceptron block (MLP), and the output of the window multi-head self-attention block (SW-MSA), result one and the output of the second multilayer perceptron block (MLP) are added as sliding window result two; ③ The output of the second multilayer perceptron block (MLP) is used as the input of the third multilayer perceptron block (MLP), the output of the window multi-head self-attention block (SW-MSA) is jump-connected to the output of the third multilayer perceptron block (MLP), and the output of the window multi-head self-attention block (SW-MSA), sliding window result one, sliding window result two and the output of the third multilayer perceptron block (MLP) are added as sliding window result three. The above content can be summarized as follows: in three basic Transformer modules with window multi-head self-attention mechanisms or sliding window multi-head self-attention mechanisms, the attention of the first window multi-head self-attention mechanism or sliding window multi-head self-attention mechanism is calculated, and its calculation results are directly inserted into the outputs of the multi-layer perceptrons in the first and second layers by using a jump connection method, and the inputs of the multi-layer perceptrons in the second and third layers are added to obtain the calculation results of the multi-head attention mechanism. After the calculation results of the multi-head attention mechanism pass through the final 3×3 convolutional layer, the final deep feature extraction output feature F is obtained. i ;
[0061] Formal definition: Let the height, width, and number of channels of the given input feature be H, W, and C respectively. First, the single-channel feature is divided into There are M×M non-overlapping local windows, that is, the input features of size H×W×C are divided into Among them, W-MSA (Window-based Multi-head Self-Attention) is a common window multi-head self-attention mechanism, SW-MSA (Shifted Windo w -based Multi-headSelf-Attention) for window sliding The multi-head self-attention mechanism for each local window feature in the corresponding module The standard self-attention mechanism is calculated separately, and the formula is as follows:
[0062] Q = XP Q , K = XP K , V = XP V
[0063]
[0064] Among them, P Q , P K , P V is the projection matrix shared across different windows, B is a learnable relative position code. After multiple parallel calculations, the results of the multi-head self-attention modules are connected as the final result;
[0065] For an image of size H×W×C, after being divided into blocks of each window size of M×M, the corresponding self-attention mechanism MSA and window multi-head self-attention mechanism W-MSA are calculated as shown in the following formula:
[0066] Ω(MSA)=4HWC 2 +2(HW) 2 C
[0067] Ω(W-MSA)=4HWC 2 +2M 2 HWC
[0068] The method of dividing the window to calculate the local self-attention mechanism greatly reduces the amount of model calculation. In addition, the method of alternating the connection between the window multi-head self-attention mechanism and the sliding window multi-head self-attention mechanism can ensure that the information at the window connection is not missed, and better focus on the global information while capturing the local information in the window.
[0069] After each (sliding) window multi-head self-attention mechanism, three multi-layer perceptrons (MLP) with normalization layers are used for further feature transformation. MLP calculates the gradient through the back-propagation algorithm, and propagates the error from the output layer back to the input layer to update the network parameters. A normalization (LayerNorm, LN) layer is added before each MSA and MLP module, and residual connections are used between modules. Assuming the input is X, the module can be expressed by the following formula:
[0070] X MSA =MSA(LN(X))
[0071]
[0072]
[0073] Among them, MSA() represents the (sliding) window multi-head self-attention mechanism operation, MLP() is a two-layer perceptron operation including GELU nonlinear transformation, and LN() represents the normalization layer operation. is the output after the i-th MLP operation, is the final output of the jumping window multi-head self-attention mechanism module;
[0074] S4, construct a sub-pixel convolutional network, and use the sub-pixel convolutional network to reconstruct a high-quality image. The input of the sub-pixel convolutional network R is the output F0 of the shallow feature extraction part and the output F D The fusion of the two images is the final reconstructed image I. RE , the formula is:
[0075] I RE =R(F0+F D )
[0076] Among them, the shallow feature extraction results mainly contain low-frequency information in the image, while the deep feature extraction focuses more on recovering the high-frequency information lost in low-quality images. The model SSFSR directly transmits the low-frequency information output by the shallow feature extraction part into the sub-pixel convolution reconstruction part through jump connections, prompting the deep feature extraction part to focus on the recovery of high-frequency information and improve the stability of training;
[0077] like Figure 5 and Figure 6As shown in the figure, the sub-pixel convolution network of this step consists of a 3×3 convolution layer and a LeakyReLU activation function, two sub-pixel convolution 2x upsampling modules and a 3×3 convolution layer; the output feature maps of the shallow feature extraction part and the deep feature extraction part are fused and input to the first convolution layer in the sub-pixel reconstruction part. The channel dimension of the feature map can be scaled for the first time here, which is determined by the number of input channels C. in Scaling to the reconstruction dimension C set in the sub-pixel convolution reconstruction part RE , a LeakyReLU activation function is connected after the first convolutional layer; then, the convolutional layer of the first sub-pixel convolution 2x upsampling module changes the channel dimension of the feature map again, increasing it by 4 times, so that after the subsequent pixel reorganization layer increases the length and width of the feature map from H×W to 2H×2W, the feature map can still maintain the same channel dimension as before the sub-pixel convolution; after this, the convolutional layer of the second sub-pixel convolution 2x upsampling module also increases the feature dimension of the image from C RE Increase to 4×C RE , the pixel reorganization layer that follows doubles the length and width of the feature map again, from 2H×2W to 4H×4W, thus achieving a 4-fold increase in the resolution of the feature map; finally, the channel dimension is C RE The feature map of is passed through the fourth convolution layer. The convolution operation transforms the number of channels of the feature map into the number of channels C required for the output image (the number of channels here is 3 when the output is an RGB color image), thereby obtaining the final sub-pixel convolution reconstructed output image I with a spatial dimension of 4H×4W×C. RE .
[0078] The sub-pixel convolution module realizes the conversion of features from low resolution to high resolution through specific convolution operations and multi-channel reorganization of features, which can better preserve the detail information of the image; the final convolution layer maps the image from the high-dimensional feature space back to the three-color channels for observation.
[0079] like Figure 2 As shown in the model block diagram of a face image enhancement method for compensating for Internet transmission degradation of the present invention:
[0080] S1, shallow feature extraction part, is used to map the input image to a higher-dimensional feature space for subsequent feature extraction and image reconstruction.
[0081] S2, deep feature extraction part, is used to capture the multi-dimensional local and global information of the image and provide better detail features for image reconstruction.
[0082] S3, the sub-pixel convolution reconstruction part, is used to realize the conversion of features from low resolution to high resolution, and map them from the high-dimensional feature space back to the three-color channels for observation.
[0083] In summary, the present invention provides a facial image enhancement method for compensating for Internet transmission degradation, comprising a low-quality facial data set compressed by WeChat software transmission and a facial image enhancement model composed of a shallow feature extraction part, a deep feature extraction part, and a sub-pixel convolution reconstruction part. In view of the problem of complex calculation of the self-attention mechanism, the present invention has a high degree of similarity in the calculation results of the multi-head self-attention mechanism based on the continuous intermediate layer, which saves some complex self-attention mechanism calculations; in view of the problem that the window multi-head self-attention mechanism has defects in extracting information between windows, the present invention adopts the method of alternating connection of the window multi-head self-attention mechanism and the sliding window multi-head self-attention mechanism, so that the information at the window connection is not missed, and the global information is better focused while capturing the local information in the window.
[0084] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications made without departing from the principle of the present invention should also be regarded as falling within the scope of protection of the present invention.
Claims
1. A facial image enhancement method for compensating for Internet transmission degradation, characterized in that: include: S1: Input low-quality face images degraded via the Internet; S2: Construct a convolutional neural network and use the constructed convolutional neural network to perform shallow feature extraction on low-quality images; S3: construct a residual convolutional neural network, and use the constructed residual convolutional neural network to extract local and global deep features. The residual convolutional neural network includes two residual jump connection modules and a convolutional layer. The shallow feature extraction output features are used as input, and the deep feature extraction output features are obtained through the residual convolutional neural network. The two residual jump connection modules included in the residual convolutional neural network are a jump connection module with a window multi-head self-attention mechanism and a jump connection module with a sliding window multi-head self-attention mechanism. The calculation results of the input features obtained by the two residual jump connection modules are convolved through the convolutional layer to obtain the output features of the deep feature extraction part. S4: Construct a sub-pixel convolutional network, using the fusion of the output of the shallow feature extraction obtained by S2 and the output of the deep feature extraction obtained by S3 as input to achieve the conversion of features from low resolution to high resolution, and obtain the sub-pixel convolution reconstructed output image.
2. A facial image enhancement method for compensating for Internet transmission degradation according to claim 1, characterized in that: The jump connection module with the window multi-head self-attention mechanism further includes a window multi-head self-attention block, three series-connected multi-layer perceptron blocks, and the window multi-head self-attention block and the three series-connected multi-layer perceptron blocks are jump-connected. The input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block. The output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block. The output of the window multi-head self-attention block, the output of the first multi-layer perceptron block, and the input / output addition result of the window multi-head self-attention block are added as window result one, and result one is obtained by the first multi-layer perceptron. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one and the output of the second multi-layer perceptron block are added together as the window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, the window result two and the output of the third multi-layer perceptron block are added together as the window result three; Among them, the jump connection module with the sliding window multi-head self-attention mechanism further includes a sliding window multi-head self-attention block, three multi-layer perceptron blocks connected in series, the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected, the input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block, the output of the window multi-head self-attention block, the output of the first multi-layer perceptron block and the input / output addition result of the window multi-head self-attention block are added and processed as sliding window result one, and the sliding window result one is composed of the first multi-layer perceptron block. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, result one, and the output of the second multi-layer perceptron block are added together as sliding window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, sliding window result one, sliding window result two, and the output of the third multi-layer perceptron block are added together as sliding window result three.
3. The method for enhancing facial images for compensating for Internet transmission degradation according to claim 1, characterized in that: The sub-pixel convolutional network includes a first convolutional layer, a LeakyReLU activation function, two sub-pixel upsampling modules and a second convolutional layer connected in series, and the sub-pixel upsampling module is composed of a convolutional layer and a pixel reorganization block.
4. The method for enhancing facial images for compensating for Internet transmission degradation according to claim 1, characterized in that: The residual convolutional neural network further includes a loss function The formula is as follows: Among them, I RE is a high-resolution image reconstructed through network training, I HR Original high-resolution image.
5. The method for enhancing facial images to compensate for Internet transmission degradation according to claim 1, characterized in that: The output features of deep feature extraction are shown as follows: F D =D CONV (F2), Among them, F i represents the image features after the i-th RSSTB module, Indicates the operation of the ith RSSTB module, D CONV represents the convolutional layer operation in deep feature extraction D, F D It represents the output after deep feature extraction D.
6. A facial image enhancement system for compensating for Internet transmission degradation, characterized in that: include: An input module, used for inputting low-quality face images degraded by transmission via the Internet; A shallow feature extraction module is used to construct a convolutional neural network and use the constructed convolutional neural network to perform shallow feature extraction on low-quality images; A deep feature extraction module is used to construct a residual convolutional neural network, and the constructed residual convolutional neural network is used to extract local and global deep features. The residual convolutional neural network includes two residual jump connection modules and a convolutional layer. The deep feature extraction process includes taking the output features of the shallow feature extraction as input, and obtaining the output features of the deep feature extraction through the residual convolutional neural network. The two residual jump connection modules are respectively a jump connection module with a window multi-head self-attention mechanism and a jump connection module with a sliding window multi-head self-attention mechanism. The calculation result of the multi-head self-attention mechanism is obtained. After the calculation result of the multi-head self-attention mechanism passes through the convolutional layer, the output of the convolutional layer of the shallow feature extraction part is added to the output of the deep feature extraction part to perform deep feature extraction; The sub-pixel convolution reconstruction module is used to construct a sub-pixel convolution network, and uses the fusion of the output of the shallow feature extraction obtained by the shallow feature extraction module and the output of the deep feature extraction obtained by the deep feature extraction module as input to achieve the conversion of features from low resolution to high resolution, and obtain a sub-pixel convolution reconstructed output image.
7. A facial image enhancement method for compensating for Internet transmission degradation according to claim 6, characterized in that: The jump connection module with the window multi-head self-attention mechanism further includes a window multi-head self-attention block, three series-connected multi-layer perceptron blocks, and the window multi-head self-attention block and the three series-connected multi-layer perceptron blocks are jump-connected. The input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block. The output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block. The output of the window multi-head self-attention block, the output of the first multi-layer perceptron block, and the input / output addition result of the window multi-head self-attention block are added as window result one, and result one is obtained by the first multi-layer perceptron. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one and the output of the second multi-layer perceptron block are added together as the window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, the window result one, the window result two and the output of the third multi-layer perceptron block are added together as the window result three; Among them, the jump connection module with the sliding window multi-head self-attention mechanism further includes a sliding window multi-head self-attention block, three multi-layer perceptron blocks connected in series, the window multi-head self-attention block and the three multi-layer perceptron blocks connected in series are jump-connected, the input and output of the window multi-head self-attention block are added and processed as the input of the first multi-layer perceptron block, the output of the window multi-head self-attention block is jump-connected to the output of the first multi-layer perceptron block, the output of the window multi-head self-attention block, the output of the first multi-layer perceptron block and the input / output addition result of the window multi-head self-attention block are added and processed as sliding window result one, and the sliding window result one is composed of the first multi-layer perceptron block. The output of the first multi-layer perceptron block is used as the input of the second multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the second multi-layer perceptron block, and the output of the window multi-head self-attention block, result one, and the output of the second multi-layer perceptron block are added together as sliding window result two; the output of the second multi-layer perceptron block is used as the input of the third multi-layer perceptron block, and the output of the window multi-head self-attention block is jump-connected to the output of the third multi-layer perceptron block, and the output of the window multi-head self-attention block, sliding window result one, sliding window result two, and the output of the third multi-layer perceptron block are added together as sliding window result three.
8. A facial image enhancement system for compensating for Internet transmission degradation according to claim 6, characterized in that: The window size M of the window multi-head self-attention block or the sliding window multi-head self-attention block is 8, the number of channels C of the feature size in the feature extraction process is 180, and the number of heads of the multi-head self-attention mechanism is 6.
9. A facial image enhancement system for compensating for Internet transmission degradation according to claim 6, characterized in that: use The loss function optimizes the network parameters, and the formula is as follows: Among them, I RE is a high-resolution image reconstructed through network training, I HR Original high-resolution image.
10. The facial image enhancement system for compensating for Internet transmission degradation according to claim 6, characterized in that: The output features of deep feature extraction are shown as follows: F D =D CONV (F2), Among them, F i represents the image features after the i-th RSSTB module, Indicates the operation of the ith RSSTB module, D CONV represents the convolutional layer operation in deep feature extraction D, F D It represents the output after deep feature extraction D.
Citation Information
Patent Citations
Light-weight multi-scale infrared image super-resolution reconstruction method
CN114092330A
Image super-resolution reconstruction model and method based on residual mixed attention network
CN115222601A
Image super-resolution reconstruction method and device based on self-attention mechanism and medium
CN115496654A
Method for removing Gibbs artifacts of magnetic resonance image
CN116309910A
Face super-resolution reconstruction method based on prior information and attention mechanism
CN117315735A