Facial image super-resolution method and device based on diversity style transfer
By constructing a diverse style transfer network model to generate high- and low-quality face sample pairs, the problem of poor super-resolution of low-quality face images in existing technologies is solved, and more efficient super-resolution effects and recognition accuracy are achieved.
Patent Information
- Application Number
- CN202410820136.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-06-24
AI Technical Summary
Existing low-quality face image super-resolution methods perform poorly in real low-quality face image super-resolution scenarios, with low reconstruction accuracy and poor results from traditional degradation simulation methods.
A diversity style transfer network model is constructed. By obtaining a dataset of high-definition facial content maps and low-quality facial style maps, the diversity style transfer network model is iteratively trained to generate high- and low-quality facial sample pairs, and the super-resolution effect is improved through a low-quality facial image super-resolution model.
The super-resolution effect of the low-quality face image super-resolution model on real low-quality face images is improved, and high-low quality face sample pairs that are closer to the degradation effect of real low-quality faces are generated, thereby improving recognition accuracy.
Smart Images

Figure CN118761906B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital image technology, and in particular to a facial image super-resolution method, device, storage medium, and electronic device. Background Art
[0002] In the security surveillance field, surveillance cameras often capture low-quality facial images that are difficult to identify. Image super-resolution methods can transform these low-quality facial images into high-definition images, enabling better facial identification. However, the effectiveness of super-resolution of low-quality facial images depends on the accuracy of degradation simulation.
[0003] In recent years, deep learning technology has been increasingly used in the super-resolution of low-quality facial images. Most of the existing low-quality facial super-resolution methods are based on various priors and different network architectures. For the degradation simulation methods of low-quality facial images, traditional degradation simulation methods are still used, such as downsampling, adding Gaussian noise, adding Gaussian blur and adding JPEG compression. These methods perform poorly in the super-resolution scenarios of real low-quality facial images and have low reconstruction accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a facial image super-resolution method, apparatus, storage medium, and electronic device, which can improve the recognition accuracy of super-resolution facial images.
[0005] The present invention provides a facial image super-resolution method, comprising:
[0006] Obtain a dataset of high-definition facial content images and a dataset of low-quality facial style images;
[0007] Constructing a diversity style transfer network model, inputting the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model to obtain a style transfer result;
[0008] Calculating an overall loss function, and iteratively training the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model;
[0009] Inputting the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determining high-low quality face sample pairs based on the style transfer results;
[0010] A low-quality face image super-resolution model is provided, and the high-low quality face sample pairs are input into the trained low-quality face image super-resolution model to obtain a high-definition face image.
[0011] Furthermore, in the above-mentioned facial image super-resolution method, the diversity style transfer network model includes two encoders and one decoder;
[0012] The step of inputting the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model to obtain a style transfer result includes:
[0013] Inputting the high-definition face content image in the high-definition face content image dataset into a first encoder to obtain content features of different sizes and content features of different combinations;
[0014] Inputting the low-quality facial style map in the low-quality facial style map dataset into a second encoder to obtain style features of different sizes and style features of different combinations;
[0015] The content features of different sizes, the content features of different combinations, the style features of different sizes, the style features of different combinations, and the style random perturbation amount are input into a decoder to obtain a style transfer result.
[0016] Furthermore, in the above-mentioned facial image super-resolution method, the decoder includes a facial style perturbation module, a multi-layer perceptron, and a convolutional layer;
[0017] The step of inputting the content features of different sizes, the content features of different combinations, the style features of different sizes, the style features of different combinations, and the style random perturbation amount into a decoder to obtain a style transfer result includes:
[0018] Inputting the content features of different combinations, the style features of different sizes, and the style features of different combinations into a convolution layer respectively to obtain a first parameter, a second parameter, and a third parameter;
[0019] Inputting the style random perturbation amount into the multi-layer perceptron to obtain a style perturbation factor;
[0020] Calculating the attention mean and attention variance after disturbance based on the first parameter, the second parameter, the third parameter and the style disturbance factor input into the face style disturbance module;
[0021] Stylized degradation is performed on the content features of different sizes based on the attention mean, the attention variance, and the style perturbation factor to obtain a style transfer result.
[0022] Furthermore, in the above-mentioned face image super-resolution method, the overall loss function is:
[0023]
[0024] in, represents the super-resolution feedback loss, represents the style loss, Indicates content loss, represents the super-resolution feedback loss weight, Represents the weight of style loss.
[0025] Furthermore, the above-mentioned facial image super-resolution method further comprises:
[0026] Based on the style transfer results of different styles and the same content, the style difference is performed to obtain the style difference results;
[0027] Normalizing the high-definition face image and the high-definition face content map and then calculating the super-resolution feedback loss;
[0028] The super-resolution feedback loss weight is calculated based on the attribute information of the high-definition face content map, the style transfer result and the high-definition face content map.
[0029] Furthermore, the above-mentioned facial image super-resolution method further comprises:
[0030] Calculating the mean and variance of the style transfer result and the low-quality facial style map respectively through a convolutional neural network, and calculating the style loss based on the mean and variance of the style transfer result and the mean and variance of the low-quality facial style map;
[0031] Calculating the style loss weight based on the style random perturbation amount;
[0032] The content loss is calculated based on the style migration result and the high-definition face content map.
[0033] The present application also provides a facial image super-resolution device, comprising:
[0034] Acquisition module, used to obtain high-definition face content map dataset and low-quality face style map dataset;
[0035] A first execution module is configured to construct a diverse style transfer network model, input the high-definition face content map dataset and the low-quality face style map dataset into the diverse style transfer network model, and obtain a style transfer result;
[0036] A training module, configured to calculate an overall loss function and iteratively train the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model;
[0037] A second execution module is configured to input the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determine a high-quality-low-quality face sample pair based on the style transfer results;
[0038] The third execution module is used to provide a low-quality face image super-resolution model, input the high-low quality face sample pairs into the trained low-quality face image super-resolution model, and obtain a high-definition face image.
[0039] Furthermore, in the above-mentioned facial image super-resolution device, the first execution module includes a first encoder, a second encoder and a decoder;
[0040] The first encoder is configured to obtain content features of different sizes and content features of different combinations based on the high-definition face content graph in the high-definition face content graph dataset;
[0041] The second encoder is used to obtain style features of different sizes and style features of different combinations based on the low-quality face style maps in the low-quality face style map dataset;
[0042] The decoder is used to obtain a style transfer result based on the content features of different sizes, the content features of different combinations, the style features of different sizes, the style features of different combinations, and the style random perturbation amount.
[0043] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor to execute any of the above-mentioned facial image super-resolution methods.
[0044] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to perform the steps in any of the above-mentioned facial image super-resolution methods.
[0045] The present application provides a facial image super-resolution method, device, storage medium, and electronic device. The present application obtains style transfer results by constructing a diverse style transfer network, and obtains high-low quality facial sample pairs based on the style transfer results, that is, high-low quality facial degradation sample pairs that are closer to the degradation effect of real low-quality faces. A low-quality facial image super-resolution model is trained based on the high-low quality facial sample pairs, which can improve the super-resolution effect of the low-quality facial image super-resolution model on real low-quality facial images. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.
[0047] Figure 1 A flowchart of the facial image super-resolution method provided in an embodiment of the present application.
[0048] Figure 2 Schematic diagram of the diverse style transfer network model structure provided in the embodiments of this application.
[0049] Figure 3 A schematic diagram of the structure of the facial feature extraction module provided in an embodiment of the present application.
[0050] Figure 4 This is a schematic diagram of the structure of the IRSE module provided in an embodiment of the present application.
[0051] Figure 5 A flowchart of obtaining style transfer results through a decoder provided in an embodiment of the present application.
[0052] Figure 6 Flowchart of the training of a diverse style transfer network model provided in an embodiment of the present application.
[0053] Figure 7 A schematic diagram of the structure of a facial image super-resolution device provided in an embodiment of the present application.
[0054] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0055] Figure 9 Another structural schematic diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0057] Most of the existing low-quality face super-resolution methods are based on various priors and different network architectures, and for the degradation simulation methods of low-quality face images, traditional degradation simulation methods are still used, which makes them perform poorly in the super-resolution scenarios of real low-quality face images.
[0058] To address the above-mentioned issues, embodiments of the present application provide a facial image super-resolution method, apparatus, storage medium, and electronic device. The facial image super-resolution apparatus provided in embodiments of the present application can be integrated into an electronic device, such as a terminal or server. The terminal may include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other device.
[0059] See also Figure 1 , Figure 1 This is a flow chart of a facial image super-resolution method provided in an embodiment of the present application, which is applied to an electronic device. The facial image super-resolution method includes the following steps:
[0060] S1, obtain a high-definition face content map dataset and a low-quality face style map dataset.
[0061] Specifically, we took a number of real low-quality face images captured by several surveillance cameras. We took one real low-quality face image from each camera as the real low-quality face style map dataset, and the remaining low-quality face images from all cameras were used as the low-quality face style image test dataset for testing. We also took a large number of high-definition faces as the high-definition face content map dataset.
[0062] After the dataset is prepared, a face detection tool is used to detect the position of the face and its facial features in each image and crop it, retaining only the face part and removing irrelevant background. Then, according to the agreed position of the key points of the facial features, the cropped faces are radially transformed so that all facial features are aligned to the agreed key points of the facial features.
[0063] S2, build a diverse style transfer network model, input the high-definition face content map dataset and the low-quality face style map dataset into the diverse style transfer network model to obtain the style transfer results.
[0064] In one embodiment, if Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of a diverse style transfer network model provided in an embodiment of the present application. The diverse style transfer network model includes two encoders and one decoder.
[0065] Step S2 can be specifically expressed by the following formula:
[0066]
[0067] in, For high-definition face content images, is a low-quality face style map, For the encoder, For content features of different sizes and combinations, Stylistic features in different sizes and combinations.
[0068] In one embodiment, step S2 may include the following steps:
[0069] S21, inputting the high-definition face content image in the high-definition face content image dataset into a first encoder to obtain content features of different sizes and content features of different combinations.
[0070] S22: Inputting the low-quality facial style map in the low-quality facial style map dataset into a second encoder to obtain style features of different sizes and style features of different combinations.
[0071] Specifically, the encoder includes a face feature extraction module, Figure 3 This is a structural diagram of the face feature extraction module provided in the embodiment of the present application, as shown in FIG. Figure 3 As shown in Figure 1, the facial feature extraction module consists of three convolutional blocks and 24 IRSE (Improved Residual Squeeze-and-Excitation) modules. The convolutional block consists of a 3×3 convolutional layer, a BatchNorm layer, and a PReLU activation layer. Figure 4 The structural diagram of the IRSE module provided in the embodiment of the present application is as follows: Figure 4 As shown in Figure 1, the IRSE module consists of two 3×3 convolutional layers, two BatchNorm layers, a PReLU activation layer, and an SE attention module. The SE attention module includes two fully connected layers, a mean pooling layer, a Sigmoid activation layer, and a ReLU activation layer. The facial feature extraction module also includes four downsampling operations and precedes the IRSE modules at four different levels. There are two facial feature extraction modules in the network encoder, one for extracting facial features from the high-definition facial content map and the other for the low-quality facial style map.
[0072] In one embodiment, the high-definition face content image and real low-quality face style map Input two face feature extraction modules respectively 、 , get content features of different sizes and style characteristics , which can be specifically expressed by the following formula:
[0073]
[0074] Downsample and concatenate style features and content features of different sizes to obtain content features of different combinations and style characteristics , which can be specifically expressed by the following formula:
[0075]
[0076] in, represents downsampling, Indicates splicing.
[0077] S23 , inputting content features of different sizes, content features of different combinations, style features of different sizes, style features of different combinations, and style random perturbations into a decoder to obtain a style transfer result.
[0078] Specifically, the decoder includes a facial style perturbation module, a multi-layer perceptron, and convolutional layers. The facial style perturbation module includes an attention weighting module and an InstanceNorm (instance normalization) layer. The attention weighting module consists of three 1×1 convolutional layers and two InstanceNorm layers. The attention weighting module calculates the attention matrix based on the input high-definition facial content map features and low-quality facial style features. This matrix is used to calculate the attention mean and attention variance of the true low-quality facial style features. The attention mean and attention variance of the true low-quality facial style features are used to stylize the high-definition facial content map features after InstanceNorm. The facial style perturbation module receives the attention matrix perturbation, mean perturbation, and variance perturbation output by the multi-layer perceptron and performs perturbations to prevent intra-style degradation. The multi-layer perceptron consists of four fully connected layers. The input of the multi-layer perceptron is the style random perturbation, and the output is the attention matrix perturbation, mean perturbation and variance perturbation, which are used to weight the attention matrix, attention mean and attention variance in each face style perturbation module.
[0079] In one embodiment, Figure 5 For a flow chart of obtaining the style transfer result through the decoder provided in the embodiment of this application, please refer to Figure 5 , step S23 may include the following steps:
[0080] S231 , inputting different combinations of content features, different sizes of style features, and different combinations of style features into the convolution layer respectively to obtain a first parameter, a second parameter, and a third parameter.
[0081] Specifically, the content features of different combinations, style features of different sizes, and style feature instances of different combinations are normalized and then passed through a 1×1 convolution layer to obtain the first parameter , the second parameter With the third parameter , which can be specifically expressed by the following formula:
[0082]
[0083] in, represents instance normalization, Represents a 1×1 convolutional layer.
[0084] S232: Input the style random perturbation amount into the multi-layer perceptron to obtain the style perturbation factor.
[0085] Specifically, the style random perturbation Input Multilayer Perceptron In the figure, after being mapped by four fully connected layers, The partition function is divided into several forms such as The collection of is input as the style perturbation factor into the face style perturbation module of each layer of the decoder. The four elements of the style perturbation factor have different shapes, sizes and uses, among which 、 The multiplication results in a shape of The matrix is used to perturb the mean weight matrix, 、 , and after the broadcast operation, the stylized mean and stylized variance are weighted respectively.
[0086] S233, based on the first parameter, the second parameter, the third parameter and the style perturbation factor, input into the face style perturbation module to calculate the perturbed attention mean and attention variance.
[0087] Specifically, use 、 、 and style disturbance factors Calculate the mean attention after perturbation and attention variance .
[0088]
[0089] S=
[0090] =
[0091] in, represents matrix multiplication, is the transpose of the matrix, is the normalization function.
[0092] S234: Based on the attention mean, attention variance, and style perturbation factor, content features of different sizes are stylized and degraded to obtain style transfer results.
[0093] Specifically, the style perturbation factor attention mean is used , attention variance and style disturbance factors Perform stylized degradation on the content features to obtain the stylized degraded facial features (The style-degraded facial features are the style transfer results), which can be specifically expressed by the following formula:
[0094]
[0095] S3 calculates the overall loss function and iteratively trains the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model.
[0096] In one embodiment, after calculating the overall loss function, back propagation is performed to update the parameters of the diversity style transfer network model, where the overall loss function is:
[0097]
[0098] in, represents the super-resolution feedback loss, represents the style loss, Indicates content loss, represents the super-resolution feedback loss weight, Represents the weight of style loss.
[0099] In one embodiment, Figure 6 The flowchart of the diversity style transfer network model training provided in the embodiment of this application is as follows: Figure 6 As shown, step S3 includes the following steps:
[0100] S31, performing style difference based on style migration results of different styles and the same content to obtain style difference results.
[0101] Specifically, the inter-style difference method is used to calculate the inter-style difference result, which can be calculated by the following formula:
[0102]
[0103]
[0104] in, express Style transfer results under style, express Style transfer results under style, express Style transfer results under style and The interpolation result of the style transfer result under the style is performed linearly pixel by pixel. Style transfer results under style , the weight of its pixel during interpolation is , relatively, Style transfer results under style , the weight of its pixel during interpolation is In the process of training the diversity style transfer network, the present invention randomly selects pictures with the same content and different styles in a batch according to the weight set shown. Perform inter-style interpolation to generate multiple inter-style interpolation results. Generally, .
[0105] S32, normalize the high-definition face image and the high-definition face content map and then calculate the super-resolution feedback loss.
[0106] Specifically, the super-resolution feedback loss can be calculated by the following formula:
[0107]
[0108] in, is the style transfer result, Represents a high-definition face content map, represents any low-quality face image super-resolution model trained on paired high-low quality face sample pairs, It is a pixel normalization operation, which is used to normalize the pixel values of the image to between.
[0109] S33, calculating the super-resolution feedback loss weight based on the attribute information of the high-definition face content map, the style transfer result, and the high-definition face content map.
[0110] Specifically, the super-resolution feedback loss weight is an adaptive weight that can adjust the relationship between stylization and super-resolution feedback as the current stylization progresses. It can be calculated using the following formula:
[0111]
[0112] in, 、 、 Represents the number of channels, height, and width of the high-definition face content image respectively, and divides the numerator and denominator of the fraction by It can be found that the denominator represents the average difference between the pixels of the restored image and the high-definition image.
[0113] S34, respectively calculates the mean and variance of the style transfer result and the low-quality face style map through a convolutional neural network, and calculates the style loss based on the mean and variance of the style transfer result and the mean and variance of the low-quality face style map.
[0114] Specifically, it can be calculated by the following formula:
[0115]
[0116] in, Indicates the calculation of the mean, Indicates the calculation of variance, ReLU that extracts the input image using the pre-trained VGG16 convolutional neural network In the embodiment of the present application, the style of the image is represented by the mean and variance of the features, and the style loss is calculated by the style transfer result. and style map The mean and variance measures of the respective VGG features.
[0117] S35, calculating the style loss weight based on the style random perturbation amount.
[0118] Specifically, the style loss weight and style random perturbation amount The modulus length hook can be used to control the degree of stylization degradation, which can be calculated by the following formula:
[0119]
[0120] in, Indicates the number of dimensions for obtaining random disturbances of wind turbines, express The mold length.
[0121] S36, calculates content loss based on the style transfer results and the high-definition face content map.
[0122] Specifically, it can be calculated by the following formula:
[0123]
[0124] in, ReLU representing the input image extracted with pre-trained VGG16 _1 layer features.
[0125] S4: Input the high-definition face content map and the low-quality face style map into the trained diversity style transfer network model to obtain the corresponding style transfer results, and determine the high-low quality face sample pairs based on the style transfer results.
[0126] Among them, the style transfer result is a stylized low-quality face image, and the high-low quality Renliang sample pair is a high-quality face content image-stylized low-quality face image sample pair.
[0127] When generating high-quality and low-quality face sample pairs, the style random perturbation is used to generate face sample pairs with intra-style degradation diversity, and the inter-style interpolation is used to generate face sample pairs with inter-style degradation diversity. The style random perturbation is a random vector of fixed dimension, and its elements are (i∈Z,1≤i≤d(Z)) obeys the standard normal distribution ∼N(0,1), during the training of the diversity style transfer network model, the style random perturbation amount is randomly generated. When using the trained diversity style transfer network model to generate sample pairs, the style random perturbation amount can be randomly generated or the modulus length can be manually controlled to control the degree of stylization degradation.
[0128] S5 provides a low-quality face image super-resolution model, inputs the high-low quality face sample pairs into the trained low-quality face image super-resolution model, and obtains a high-definition face image.
[0129] The low-quality face image super-resolution model may be any low-quality face image super-resolution model trained based on paired high- and low-quality face sample pairs, such as a GFPGAN model.
[0130] In one embodiment, before step S5, a low-quality face image super-resolution model needs to be trained using high-quality and low-quality face sample pairs to obtain a trained low-quality face image super-resolution model.
[0131] The embodiments of this application construct a diverse style transfer network model that utilizes style transfer degradation to generate high- and low-quality face sample pairs with both intra- and inter-style degradation diversity. Compared to traditional degradation methods, this method achieves degradation results that are closer to those of real low-quality face images. The embodiments of this application utilize the high- and low-quality face sample pairs generated by the diverse style transfer network model to train a low-quality face image super-resolution model, successfully improving the super-resolution effect of the low-quality face image super-resolution model on real low-quality face images.
[0132] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a facial image super-resolution device. The facial image super-resolution device can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, a server, or other devices. The terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro processing box, or other devices.
[0133] See also Figure 7 , Figure 7 The present invention specifically describes a facial image super-resolution device provided by an embodiment of the present invention, which is applied to an electronic device. The facial image super-resolution device may include:
[0134] Acquisition module, used to obtain high-definition face content map dataset and low-quality face style map dataset;
[0135] The first execution module is used to build a diversity style transfer network model, input the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model, and obtain the style transfer result;
[0136] The training module is used to calculate the overall loss function and iteratively train the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model;
[0137] The second execution module is used to input the high-definition face content map and the low-quality face style map into the trained diversity style transfer network model to obtain corresponding style transfer results, and determine high-quality and low-quality face sample pairs based on the style transfer results;
[0138] The third execution module is used to provide a low-quality face image super-resolution model, input the high-low quality face sample pairs into the trained low-quality face image super-resolution model, and obtain a high-definition face image.
[0139] During specific implementation, the above modules and / or units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above modules and / or units can refer to the previous method embodiments. The specific beneficial effects that can be achieved can also be found in the beneficial effects in the previous method embodiments, which will not be repeated here.
[0140] In addition, an embodiment of the present application also provides an electronic device, which may be a computer, a tablet computer, or other device. Figure 8 As shown, the electronic device 400 includes a processor 401 and a memory 402. The processor 401 is electrically connected to the memory 402.
[0141] The processor 401 is the control center of the electronic device 400. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or loading applications stored in the memory 402 and calling data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.
[0142] In this embodiment, the processor 401 in the electronic device 400 loads instructions corresponding to one or more application processes into the memory 402 according to the following steps, and the processor 401 runs the application stored in the memory 402 to implement various functions:
[0143] Obtain a dataset of high-definition facial content images and a dataset of low-quality facial style images;
[0144] Constructing a diversity style transfer network model, inputting the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model to obtain a style transfer result;
[0145] Calculating an overall loss function, and iteratively training the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model;
[0146] Inputting the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determining high-low quality face sample pairs based on the style transfer results;
[0147] A low-quality face image super-resolution model is provided, and the high-low quality face sample pairs are input into the trained low-quality face image super-resolution model to obtain a high-definition face image.
[0148] This electronic device can implement the steps in any embodiment of the facial image super-resolution method provided in the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any facial image super-resolution method provided in the embodiments of the present invention. Please refer to the previous embodiments for details and will not be repeated here.
[0149] Figure 9 The following is a block diagram of the structure of an electronic device provided in an embodiment of the present invention, which can be used to implement the facial image super-resolution method provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor, or other device.
[0150] RF circuit 510 is used to receive and transmit electromagnetic waves, converting them into electrical signals, thereby enabling communication with a communications network or other devices. RF circuit 510 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, memory, and the like. RF circuit 510 can communicate with various networks, such as the Internet, an intranet, or a wireless network, or with other devices via a wireless network. These wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The wireless networks may utilize various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE802.11g, and / or IEEE802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messaging, and any other suitable communication protocols, including those currently undeveloped.
[0151] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above-mentioned embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizing functions such as taking pictures with the front camera, processing the captured images, and switching the display color of the displayed content on the display screen. The memory 520 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 520 may further include a memory remotely located relative to the processor 580, and these remote memories may be connected to the electronic device 500 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0152] The input unit 530 may be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function control.
[0153] The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces. These graphical user interfaces can be composed of graphics, text, icons, videos, or any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like.
[0154] Audio circuit 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuit 560 converts received audio data into electrical signals and transmits them to speaker 561, which then converts them into sound signals for output. Microphone 562, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 560 and converted into audio data. The audio data is then processed by output processor 580 and transmitted via RF circuit 510 to, for example, another terminal. Alternatively, the audio data may be output to memory 520 for further processing. Audio circuit 560 may also include an earphone jack to allow communication between external headphones and electronic device 500.
[0155] Electronic device 500, through a transmission module 570 (e.g., a Wi-Fi module), can help users receive requests, send information, and so on, providing users with wireless broadband Internet access. Although the figure shows transmission module 570, it is understood that it is not a required component of electronic device 500 and can be omitted as needed without changing the essence of the invention.
[0156] Processor 580 is the control center of electronic device 500. It connects all components of the phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 520 and accessing data stored in memory 520, it executes various functions of electronic device 500 and processes data, thereby providing overall monitoring of the electronic device. Optionally, processor 580 may include one or more processing cores. In some embodiments, processor 580 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 580.
[0157] Electronic device 500 also includes a power supply 590 (e.g., a battery) for powering various components. In some embodiments, the power supply can be logically connected to processor 580 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 590 can also include any components, such as one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0158] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:
[0159] Obtain a dataset of high-definition facial content images and a dataset of low-quality facial style images;
[0160] Constructing a diversity style transfer network model, inputting the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model to obtain a style transfer result;
[0161] Calculating an overall loss function, and iteratively training the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model;
[0162] Inputting the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determining high-low quality face sample pairs based on the style transfer results;
[0163] A low-quality face image super-resolution model is provided, and the high-low quality face sample pairs are input into the trained low-quality face image super-resolution model to obtain a high-definition face image.
[0164] During specific implementation, the above modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. The specific implementation of the above modules can be found in the previous method embodiments and will not be repeated here.
[0165] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be accomplished through instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the embodiments of the facial image super-resolution method provided in the embodiments of the present invention.
[0166] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0167] Since the instructions stored in the storage medium can execute the steps in any embodiment of the facial image super-resolution method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any facial image super-resolution method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0168] The above is a detailed introduction to a facial image super-resolution method, device, storage medium and electronic device provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A facial image super-resolution method, characterized in that: The method comprises: Obtain a dataset of high-definition facial content images and a dataset of low-quality facial style images; Construct a diversity style transfer network model, input the high-definition face content map dataset and the low-quality face style map dataset into the diversity style transfer network model, and obtain a style transfer result; wherein the diversity style transfer network model includes two encoders and one decoder; the processing process of the diversity style transfer network model includes: Inputting the high-definition face content map in the high-definition face content map dataset into a first encoder to obtain content features of different sizes and content features of different combinations; inputting the low-quality face style map in the low-quality face style map dataset into a second encoder to obtain style features of different sizes and style features of different combinations; inputting the content features of different sizes, the content features of different combinations, the style features of different sizes, the style features of different combinations, and the style random perturbation amount into a decoder to obtain a style transfer result; The decoder includes a face style perturbation module, a multi-layer perceptron, and a convolutional layer; the processing process of the decoder includes: Inputting the content features of different combinations, the style features of different sizes, and the style features of different combinations into the convolution layer respectively to obtain a first parameter, a second parameter, and a third parameter; inputting the style random perturbation amount into the multi-layer perceptron to obtain a style perturbation factor; inputting the first parameter, the second parameter, the third parameter, and the style perturbation factor into the face style perturbation module to calculate the perturbed attention mean and attention variance; performing stylized degradation on the content features of different sizes based on the attention mean, the attention variance, and the style perturbation factor to obtain a style transfer result; Calculating an overall loss function, and iteratively training the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model; Inputting the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determining high-low quality face sample pairs based on the style transfer results; A low-quality face image super-resolution model is provided, and the high-low quality face sample pairs are input into the trained low-quality face image super-resolution model to obtain a high-definition face image.
2. The facial image super-resolution method according to claim 1, wherein: The overall loss function is: in, represents the super-resolution feedback loss, represents the style loss, Indicates content loss, represents the super-resolution feedback loss weight, Represents the weight of style loss.
3. The facial image super-resolution method according to claim 2, characterized in that: The method further comprises: Based on the style transfer results of different styles and the same content, the style difference is performed to obtain the style difference results; Normalizing the high-definition face image and the high-definition face content map and then calculating the super-resolution feedback loss; The super-resolution feedback loss weight is calculated based on the attribute information of the high-definition face content map, the style transfer result and the high-definition face content map.
4. The facial image super-resolution method according to claim 2, wherein: The method further comprises: Calculating the mean and variance of the style transfer result and the low-quality facial style map respectively through a convolutional neural network, and calculating the style loss based on the mean and variance of the style transfer result and the mean and variance of the low-quality facial style map; Calculating the style loss weight based on the style random perturbation amount; The content loss is calculated based on the style migration result and the high-definition face content map.
5. A facial image super-resolution device, the facial image super-resolution device being used to implement the facial image super-resolution method according to claim 1, characterized in that: include: Acquisition module, used to obtain high-definition face content map dataset and low-quality face style map dataset; A first execution module is configured to construct a diverse style transfer network model, input the high-definition face content map dataset and the low-quality face style map dataset into the diverse style transfer network model, and obtain a style transfer result; A training module, configured to calculate an overall loss function and iteratively train the diversity style transfer network model based on the overall loss function to obtain a trained diversity style transfer network model; A second execution module is configured to input the high-definition face content map and the low-quality face style map into a trained diversity style transfer network model to obtain corresponding style transfer results, and determine a high-quality-low-quality face sample pair based on the style transfer results; The third execution module is used to provide a low-quality face image super-resolution model, input the high-low quality face sample pairs into the trained low-quality face image super-resolution model, and obtain a high-definition face image.
6. The facial image super-resolution device according to claim 5, characterized in that: The first execution module includes a first encoder, a second encoder and a decoder; The first encoder is configured to obtain content features of different sizes and content features of different combinations based on the high-definition face content graph in the high-definition face content graph dataset; The second encoder is used to obtain style features of different sizes and style features of different combinations based on the low-quality face style maps in the low-quality face style map dataset; The decoder is used to obtain a style transfer result based on the content features of different sizes, the content features of different combinations, the style features of different sizes, the style features of different combinations, and the style random perturbation amount.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor to execute the facial image super-resolution method according to any one of claims 1 to 4.
8. An electronic device, characterized in that: It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the face image super-resolution method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Low-resolution face image super-resolution method for recognition.
CN112288627A