Method and apparatus for repairing target face image
By using an encoder in image repair technology that is inversely related to the sampling process of the StyleGan generator, a high-definition version of the face image is generated, which solves the problem of poor face image repair effect in traditional methods, and achieves high-definition and smooth face image repair.
Patent Information
- Application Number
- CN202010369826.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-04-30
AI Technical Summary
Traditional image repair methods are difficult to generate high-definition images of human faces, and there are problems of fixing blurred details and unsmooth images.
Through an encoder that is inversely related to the sampling process of the StyleGan generator, image features are extracted from the target face image, a one-dimensional feature vector corresponding to the specified target face image is generated, and the feature vector is input to the StyleGan generator to generate a high-definition face image.
It realizes efficient repair of face images, and the generated high-definition face images are clear in details and have high smoothness, solving the blur and non-smooth problems existing in traditional methods.
Smart Images

Figure CN113592724B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image restoration, and in particular, to a method and device for restoring a target face image. Background Art
[0002] Image restoration technology is an important branch problem in the field of image processing. Image restoration refers to the restoration and reconstruction of missing image information caused during image retention or the restoration after removing redundant objects in the image. Nowadays, image restoration has been widely applied to specific scenarios such as old photo restoration, cultural relic protection, and removing redundant objects.
[0003] Due to the inherent blur and complexity of natural images, especially for face images, there may be problems such as blurred restoration details and non-smooth restored images during restoration. The restoration effect of traditional methods is not good, and it is impossible to generate high-definition images of faces. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for restoring a target face image. An encoder that is inverse to the sampling process of the StyleGan generator is used to generate a one-dimensional feature vector corresponding to the specified target face image, and then the one-dimensional feature vector is input into the generator to generate a high-definition image of the specified target face, thereby realizing the restoration of the face image.
[0005] In a first aspect, an embodiment provides a method for restoring a target face image, including: extracting image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, where the trained encoder is inverse to the sampling process of the trained StyleGan generator; inputting the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-definition face image after restoring the target face image.
[0006] In an optional implementation manner, the StyleGan generator first passes through a fully connected layer structure and then through a preset number of convolutional and upsampling structures, and the encoder first passes through a preset number of convolutional and downsampling structures and then through a fully connected layer structure, where the convolutional kernel size, the number of fully connected layer structures, and the preset number of the encoder are the same as those of the StyleGan generator.
[0007] In an optional implementation manner, the step of extracting image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image includes:
[0008] Extract the first image feature from the target face image through the first convolution and downsampling structure in the trained encoder to generate the first feature map; extract the i-th image feature from the (i-1)-th feature map through the i-th convolution and downsampling structure in the trained encoder to generate the i-th feature map, and repeat the generation step of the i-th feature map until all the convolution and downsampling structures in the encoder are traversed to obtain the target feature map, where i is an integer greater than 1 and not exceeding a preset number, and i takes values from small to large; map the target feature map through the trained encoder to obtain the first one-dimensional feature vector corresponding to the target face image.
[0009] In an alternative embodiment, the step of inputting the first one-dimensional feature vector into the trained StyleGan generator to obtain the first high-definition face image after repairing the target face image includes: fusing the one-dimensional feature vector and the k-th feature map through the first convolution and upsampling structure in the trained StyleGan generator to generate the first repaired feature map; fusing the (j-1)-th repaired feature map and the (k-1)-th feature map through the j-th convolution and upsampling structure in the trained StyleGan generator to generate the j-th repaired feature map, and repeating the generation step of the j-th repaired feature map until all the convolution and upsampling structures in the generator are traversed to obtain the target repaired feature map, where j is an integer greater than 1 and not exceeding a preset number, k is an integer greater than 1 and not exceeding a preset number, j takes values from small to large, and k takes values from large to small; obtain the first high-definition face image after repairing the target face image according to the target repaired feature map.
[0010] In an alternative embodiment, the method further includes: determining a training sample set, where each group of training samples in the training sample set includes a low-definition face image sample, a one-dimensional feature vector sample, and a high-definition face image sample; training the initial encoder based on the low-definition face image samples and one-dimensional feature vector samples in the training sample set to obtain an intermediate encoder; training the intermediate encoder and the pre-trained intermediate StyleGan generator based on the low-definition face image samples and high-definition face image samples in the training sample set to obtain the trained encoder and the trained StyleGan generator.
[0011] In an alternative embodiment, the step of training the initial encoder based on the low-definition face image samples and one-dimensional feature vector samples in the training sample set to obtain an intermediate encoder includes: extracting the second one-dimensional feature vector from the low-definition face image samples through the initial encoder, determining the first loss function according to the difference between the second one-dimensional feature vector and the one-dimensional feature vector sample, if the first loss function does not meet the first preset condition, then optimize the initial encoder and continue training until the first loss function meets the first preset condition, and determine the encoder corresponding to the first loss function that meets the first preset condition as the intermediate encoder.
[0012] In an alternative embodiment, the step of obtaining the trained encoder and the trained StyleGan generator based on the low-resolution face image samples and the high-resolution face image samples in the training sample set includes:
[0013] Extracting a third one-dimensional feature vector from the low-resolution face image samples through the intermediate encoder; repairing the third one-dimensional feature vector through the pre-trained intermediate StyleGan generator to obtain a second high-resolution face image, determining a second loss function based on the difference between the second high-resolution face image and the high-resolution face image samples, and if the second loss function does not meet the second preset condition, optimizing the intermediate encoder and the intermediate StyleGan generator and continuing the training until the second loss function meets the second preset condition, and determining the encoder and the StyleGan generator corresponding to the second loss function that meets the second preset condition as the trained encoder and the trained StyleGan generator.
[0014] In an alternative embodiment, for each group of training samples in the training sample set, it is determined through the following steps: inputting the pre-determined one-dimensional feature vector samples into the intermediate StyleGan generator to generate high-resolution face image samples; adjusting the size of the high-resolution face image samples by shrinking and enlarging them, and adding Gaussian noise of random size to obtain low-resolution face image samples.
[0015] In a second aspect, an embodiment provides a target face image repair device, including: a generation module, configured to extract image features from a target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, where the trained encoder is inverse to the sampling process of the trained StyleGan generator; a repair module, configured to input the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-resolution face image after repairing the target face image.
[0016] In a third aspect, an embodiment provides an electronic device, including a memory, a processor, and a program stored on the memory and capable of running on the processor, where the processor implements the target face image repair method according to any one of the foregoing embodiments when executing the program.
[0017] In a fourth aspect, an embodiment provides a computer-readable storage medium, where a computer program is stored, and when the computer program is executed, the target face image repair method according to any one of the foregoing embodiments is implemented.
[0018] An embodiment of the present invention provides a method and apparatus for repairing a target face image. An encoder trained to be inverse to the sampling process of a trained StyleGan generator extracts image features from the target face image, generates a one-dimensional feature vector corresponding to the target face image, and uses this one-dimensional feature vector as the input of the trained StyleGan generator to generate a first high-definition image corresponding to the target face. In this way, the StyleGan generator can generate a high-definition face image for a specified low-definition face image, thereby realizing the repair of the face image.
[0019] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0020] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, details are described as follows. Description of the Drawings
[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a flowchart of a method for repairing a target face image provided by an embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of the functional modules of a device for repairing a target face image provided by an embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of the hardware architecture of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0026] The current face generation model StyleGan can generate a high-definition face image through a one-dimensional low-precision vector of random input. For example, by inputting a one-dimensional feature vector with a precision of 1*512, a high-definition face image with a precision of 1024*1024 pixels can be obtained. At this time, the high-definition face image corresponds to the input random vector and belongs to a random face. It can be understood that since the input feature vector is random, the output face is also random, and the face generation model StyleGan does not have the function of repairing a specific face image.
[0027] In the field of face image repair, generally, it is necessary to improve the precision of a specified face image, so as to achieve the purpose of repairing the specified face image. However, the current face generation model cannot obtain the one-dimensional feature vector corresponding to the specified target face, and thus cannot generate a high-definition image of the specified target face, which affects the repair application of face images.
[0028] Based on this, a method and device for repairing a target face image provided by an embodiment of the present invention generate a one-dimensional feature vector corresponding to a specified target face image through an encoder that is inverse to the StyleGan generator architecture, and then input the one-dimensional feature vector into the generator to generate a high-definition image of the specified target face, thereby realizing the repair of the face image.
[0029] To facilitate the understanding of this embodiment, a method for repairing a target face image disclosed by an embodiment of the present invention applies a StyleGan generator with the ability to generate high-definition faces to the field of face repair. Among them, it can be applied to a generative adversarial network, which generally includes a generative model and a discriminative model. The method provided by an embodiment of the present invention can be specifically applied to the generative model in the generative adversarial network. Here, a method for repairing a target face image disclosed by an embodiment of the present invention will be introduced in detail first.
[0030] Figure 1 It is a flowchart of a method for repairing a target face image provided by an embodiment of the present invention.
[0031] Referring to Figure 1 , a method for repairing a target face image provided by the embodiment includes the following steps:
[0032] Step S102, extract image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, where the trained encoder is inverse to the sampling process of the trained StyleGan generator;
[0033] Step S104, input the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-definition face image after repairing the target face image.
[0034] In a preferred embodiment of the actual application, an encoder that is inverse to the sampling process of the StyleGan generator is used to extract image features from the target face image, generate a one-dimensional feature vector corresponding to the target face image, and use this one-dimensional feature vector as the input of the StyleGan generator. The StyleGan generator generates a corresponding high-definition face image according to this one-dimensional feature vector, thereby realizing customized high-definition face generation, which can be applied to the restoration of face images. In this way, the StyleGan generator can generate a high-definition face image for a specified low-definition face image, thereby realizing the restoration of the face image.
[0035] It can be understood that the target face image used to generate the one-dimensional feature vector belongs to a low-definition face image, and this low-definition face image is restored to a high-definition face image by the restoration method provided by the embodiments of the present invention.
[0036] For the StyleGan generator in the above steps, the StyleGan generator includes a fully connected layer structure, a preset number of convolutions, and an upsampling structure. The running order of the StyleGan generator is to first pass through the fully connected layer structure and then through the preset number of convolutions and the upsampling structure.
[0037] For the encoder in the above steps, the encoder may include a fully connected layer structure, a preset number of convolutions, and an upsampling structure. The running order of the encoder is to first pass through the preset number of convolutions and the downsampling structure and then through the fully connected layer structure.
[0038] Among them, the convolutional kernel size, the number of fully connected layer structures, and the preset number of the encoder are the same as those of the StyleGan generator. For example, if the StyleGan generator first passes through the fully connected layer structure and then through the convolution + upsampling structure + convolution + upsampling structure, the encoder first passes through the convolution + downsampling structure + convolution + downsampling structure, and then through the fully connected layer structure. At this time, the preset number is two.
[0039] It should be noted that the structure of the StyleGan generator can be changed accordingly according to the specific application scenario. Therefore, there may be other forms of the design of the encoder applied in the embodiments of the present invention. The structure of the encoder is inverse to the structure of the StyleGan generator to achieve the purpose of providing a one-dimensional feature vector corresponding to the target face image for the StyleGan generator.
[0040] For example, the StyleGan generator structure can be determined according to the size of the generated face image. If it is necessary to enlarge the face image from a precision size of 4*4 to a precision size of 512*512, for the StyleGan generator, the image with a precision size of 4*4 can be enlarged through 7 upsampling processes (determined by the size of the convolutional kernel used) and magnified 7 times; among them, each time it is doubled. Due to the structural characteristics of the StyleGan generator that magnifies step by step, the generated face image has higher precision.
[0041] In order to be able to more clearly understand the process of the encoder generating the feature map, the following will be further introduced in combination with specific examples. As an example, the above step S102 may include the following steps:
[0042] Step 1.1), extract the first image feature from the target face image through the first convolution and downsampling structure in the trained encoder to generate the first feature map;
[0043] Step 1.2), extract the i-th image feature from the (i - 1)-th feature map through the i-th convolution and downsampling structure in the trained encoder to generate the i-th feature map, and repeat the generation step of the i-th feature map until all the convolution and downsampling structures in the encoder are traversed to obtain the target feature map, where i is an integer greater than 1 and not exceeding the preset number, and i takes values from small to large;
[0044] Step 1.3), map the target feature map through the trained encoder to obtain the first one-dimensional feature vector corresponding to the target face image, where this mapping is implemented through the fully connected layer structure in the encoder.
[0045] Here, if the encoder includes 5 "convolution + downsampling structures", because the encoder structure is the inverse process of the StyleGan generator, the hyperparameters such as the corresponding convolutional kernel size of the encoder are also the same as those of the StyleGan generator, that is, the output of the encoder convolutional layer is a feature map of the same size as the input layer of the StyleGan generator, so that the StyleGan generator can be used for subsequent fusion with the feature map output by the encoder. Among them, downsampling reduces the size of the face image and is used to obtain the face feature map at a small scale.
[0046] Specifically, taking the above embodiment as an example, the preset number is 5. The first image feature is extracted from the target face image through the first convolution and downsampling structure in the encoder to generate the first feature map. The second image feature is extracted from the first feature map through the second convolution and downsampling structure in the encoder to generate the second feature map. The generation step of the second feature map is repeatedly executed, that is, the third image feature is extracted from the second feature map through the third convolution and downsampling structure in the encoder to generate the third feature map, until all the convolution and downsampling structures in the encoder are traversed, and the fifth feature map is generated.
[0047] It can be understood that the StyleGan generator in the embodiment of the present invention can be understood as a decoder. After passing through 5 "convolution + downsampling structures" in the encoder, the feature map is mapped to a 1*512 feature vector through 3 fully connected layers, which is used as the input of the StyleGan generator, thereby realizing the generation process of high-definition face images.
[0048] In order to more clearly illustrate the process of the generator generating high-definition face images, the following will be further introduced in combination with specific examples. As an example, step S104 above may specifically include the following steps:
[0049] Step 2.1), the one-dimensional feature vector and the k-th feature map are fused through the first convolution and upsampling structure in the trained StyleGan generator to generate the first repaired feature map;
[0050] Step 2.2), the (j-1)-th repaired feature map and the (k-1)-th feature map are fused through the j-th convolution and upsampling structure in the trained StyleGan generator to generate the j-th repaired feature map. The generation step of the j-th repaired feature map is repeatedly executed until all the convolution and upsampling structures in the generator are traversed to obtain the target repaired feature map. j is an integer greater than 1 and not exceeding the preset number, and k is an integer greater than 1 and not exceeding the preset number, where j takes values from small to large and k takes values from large to small;
[0051] Step 2.3), the first high-definition face image after the target face image is repaired is obtained according to the target repaired feature map.
[0052] Here, each feature map generated during the encoder convolution downsampling process is added and fused with the feature map of the corresponding size in the StyleGan generator, and the encoder and the StyleGan generator are jointly trained, so as to improve the similarity between the high-definition face generated by the generator and the high-definition face features corresponding to the original low-definition face, so as to realize the repair of low-definition face images. At this time, the input of the generation model is a pre-made low-definition face image, and the output after training the generation model is a high-definition face image. Among them, the generation model includes an encoder and a StyleGan generator.
[0053] Specifically, taking the above embodiment as an example, the preset number is 5. The one-dimensional feature vector and the fifth feature map are fused by the first convolution and upsampling structure in the generator to generate the first repaired feature map. The first repaired feature map and the fourth feature map are fused by the second convolution and upsampling structure in the generator to generate the second repaired feature map. The generation step of the second repaired feature map is repeatedly executed, that is, the second repaired feature map and the third feature map are fused by the third convolution and upsampling structure in the generator to generate the third repaired feature map until all the convolution and upsampling structures in the generator are traversed to generate the fifth repaired feature map.
[0054] In an alternative embodiment, to obtain a more accurate high-definition face repaired image, the repair network needs to be trained. The repair network includes an encoder and a StyleGan generator. As an example, the method may further include the following steps:
[0055] Step 3.1), determining a training sample set. Wherein, the training sample set may include multiple groups of training samples, and each group of training samples in the training sample set includes a low-definition face image sample, a one-dimensional feature vector sample, and a high-definition face image sample;
[0056] Step 3.2), training the initial encoder based on the low-definition face image samples and the one-dimensional feature vector samples in the training sample set to obtain an intermediate encoder.
[0057] Step 3.3), training the intermediate encoder and the pre-trained intermediate StyleGan generator based on the low-definition face image samples and the high-definition face image samples in the training sample set to obtain a trained encoder and a trained StyleGan generator.
[0058] For the above step 3.2), the second one-dimensional feature vector can be obtained by feature extraction of the low-definition face image sample through the initial encoder, and the first loss function is determined according to the difference between the second one-dimensional feature vector and the one-dimensional feature vector sample. The first loss function is judged. If the first loss function does not meet the first preset condition, the initial encoder is optimized and then training continues; the above steps are iteratively executed until the first loss function meets the first preset condition, and the encoder corresponding to the first loss function that meets the first preset condition is determined as the intermediate encoder.
[0059] For step 3.3) above, the intermediate encoder can be used to extract features from the low-resolution face image samples to obtain the third one-dimensional feature vector; the pre-trained intermediate StyleGan generator can be used to repair the third one-dimensional feature vector to obtain the second high-resolution face image, and the second loss function can be determined based on the difference between the second high-resolution face image and the high-resolution face image samples; the second loss function is judged. If the second loss function does not meet the second preset condition, the intermediate encoder and the intermediate StyleGan generator are optimized and then continue to be trained until the second loss function meets the second preset condition. The encoder and the StyleGan generator corresponding to the second loss function that meets the second preset condition are determined as the trained encoder and the trained StyleGan generator.
[0060] Among them, the pre-trained intermediate StyleGan generator can be a pre-trained StyleGan generator that can generate high-resolution face images based on one-dimensional feature vectors.
[0061] Here, the second loss function can also be determined by introducing the difference in the intermediate quantities of the repair network. Specifically, the second loss function can be determined by aggregating the difference between the one-dimensional feature vector generated by the encoder (the repaired one-dimensional feature vector corresponding to the repaired high-resolution face image obtained through the generator combined with the encoder) and the one-dimensional feature vector corresponding to the high-resolution face image generated by the pre-trained intermediate StyleGan generator (the high-resolution face image corresponding to the training sample).
[0062] For example, the second loss function calculates the difference loss by taking the difference or the squared difference between the above two variables respectively according to the difference between the generated high-resolution face image and the sample and the difference between the generated one-dimensional feature vector and the sample.
[0063] For step 3.1) above, for each group of training samples in the training sample set, it can be determined through the following steps:
[0064] Step 4.1), input the pre-determined one-dimensional feature vector samples into the intermediate StyleGan generator to generate high-resolution face image samples;
[0065] Step 4.2), adjust the size of the high-resolution face image samples by shrinking and enlarging, and add Gaussian noise with random size to obtain low-resolution face image samples.
[0066] As another alternative implementation, the training sample set can be pre-set high-resolution face images and the corresponding one-dimensional feature vectors and low-resolution face images.
[0067] In the embodiments of the present invention, by introducing an encoder structure into the generative model and combining the encoder with the StyleGan generator, a customized face restoration function is achieved.
[0068] As Figure 2 shown, the embodiments of the present invention provide a target face image restoration device, including:
[0069] A generation module 201, configured to extract image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, where the trained encoder is inverse to the sampling process of the trained StyleGan generator;
[0070] A restoration module 202, configured to input the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-definition face image after restoration of the target face image.
[0071] In some embodiments, the StyleGan generator first passes through a fully connected layer structure and then through a preset number of convolutional and upsampling structures, and the encoder first passes through a preset number of convolutional and downsampling structures and then through a fully connected layer structure, where the convolutional kernel size, the number of fully connected layer structures, and the preset number of the encoder are consistent with those of the StyleGan generator.
[0072] In some embodiments, the generation module is specifically further configured to extract first image features from the target face image through the first convolutional and downsampling structures in the trained encoder to generate a first feature map; extract the i-th image features from the (i - 1)-th feature map through the i-th convolutional and downsampling structures in the trained encoder to generate the i-th feature map, and repeat the generation step of the i-th feature map until all the convolutional and downsampling structures in the encoder are traversed to obtain a target feature map, where i is an integer greater than 1 and not exceeding the preset number, and i takes values from small to large; map the target feature map through the trained encoder to obtain a first one-dimensional feature vector corresponding to the target face image.
[0073] In some embodiments, the restoration module is specifically further configured to fuse the first one-dimensional feature vector and the k-th feature map through the first convolutional and upsampling structures in the trained StyleGan generator to generate a first restored feature map; fuse the (j - 1)-th restored feature map and the (k - 1)-th feature map through the j-th convolutional and upsampling structures in the trained StyleGan generator to generate the j-th restored feature map, and repeat the generation step of the j-th restored feature map until all the convolutional and upsampling structures in the generator are traversed to obtain a target restored feature map, where j is an integer greater than 1 and not exceeding the preset number, k is an integer greater than 1 and not exceeding the preset number, j takes values from small to large, and k takes values from large to small; obtain a first high-definition face image after restoration of the target face image according to the target restored feature map.
[0074] In some embodiments, the device further includes a training module, specifically for: determining a training sample set, where each group of training samples in the training sample set includes a low-resolution face image sample, a one-dimensional feature vector sample, and a high-resolution face image sample; training an initial encoder based on the low-resolution face image samples and the one-dimensional feature vector samples in the training sample set to obtain an intermediate encoder; and training the intermediate encoder and a pre-trained intermediate StyleGan generator based on the low-resolution face image samples and the high-resolution face image samples in the training sample set to obtain a trained encoder and a trained StyleGan generator.
[0075] In some embodiments, the training module is specifically further for:
[0076] Extracting features from the low-resolution face image samples through the initial encoder to obtain a second one-dimensional feature vector, determining a first loss function according to the difference between the second one-dimensional feature vector and the one-dimensional feature vector samples. If the first loss function does not meet the first preset condition, optimize the initial encoder and continue training until the first loss function meets the first preset condition, and determine the encoder corresponding to the first loss function that meets the first preset condition as the intermediate encoder;
[0077] Extracting features from the low-resolution face image samples through the intermediate encoder to obtain a third one-dimensional feature vector; repairing the third one-dimensional feature vector through a pre-trained intermediate StyleGan generator to obtain a second high-resolution face image, determining a second loss function based on the difference between the second high-resolution face image and the high-resolution face image samples. If the second loss function does not meet the second preset condition, optimize the intermediate encoder and the intermediate StyleGan generator and continue training until the second loss function meets the second preset condition, and determine the encoder and the StyleGan generator corresponding to the second loss function that meets the second preset condition as the trained encoder and the trained StyleGan generator.
[0078] In some embodiments, the training module is further for:
[0079] Inputting the pre-determined one-dimensional feature vector samples into the intermediate StyleGan generator to generate high-resolution face image samples;
[0080] Adjusting the size of the high-resolution face image samples by reducing and then enlarging them, and adding Gaussian noise of random size to obtain low-resolution face image samples.
[0081] Further, as Figure 3As shown in the figure, it is a schematic diagram of an electronic device 300 for implementing the target face image restoration method provided by an embodiment of the present invention. In this embodiment, the electronic device 300 may be, but is not limited to, a computer device with analysis and processing capabilities such as a personal computer (PC), a laptop computer, a monitoring device, a server, etc. As an alternative embodiment, the electronic device 300 may be the target face image restoration method.
[0082] Figure 3 It is a schematic diagram of the hardware architecture of the electronic device 300 provided by an embodiment of the present invention. As Figure 3 shown, the electronic device 300 includes a memory 301 and a processor 302. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method provided in the above embodiment.
[0083] Referring to Figure 3 , the electronic device further includes: a bus 303 and a communication interface 304. The processor 302, the communication interface 304, and the memory 301 are connected through the bus 303. The processor 302 is used to execute an executable module stored in the memory 301, such as a computer program.
[0084] Among them, the memory 301 may include a high-speed random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 304 (which may be wired or wireless), a communication connection between the system network element and at least one other network element can be realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0085] The bus 303 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0086] Among them, the memory 301 is used to store programs. After receiving an execution instruction, the processor 302 executes the program. The method executed by the device defined by the process disclosed in any embodiment of the present application can be applied to the processor 302 or implemented by the processor 302.
[0087] The processor 302 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 302 or instructions in the form of software. The above-mentioned processor 302 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 301, and the processor 302 reads the information in the memory 301 and combines its hardware to complete the steps of the above method.
[0088] Corresponding to the above cross-blockchain communication method, an embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and run by a processor, the machine-executable instructions cause the processor to run the steps of the above face image restoration method.
[0089] The face image restoration device provided by the embodiments of the present application may be specific hardware on a device or software or firmware installed on the device, etc. For the device provided by the embodiments of the present application, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing-described systems, devices, and units can all refer to the corresponding processes in the above method embodiments, and will not be repeated here.
[0090] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0091] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, each functional unit in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0093] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code. A module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the cross-blockchain communication method in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0095] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0096] Finally, it should be noted that: the above embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application.
Claims
1. A method for repairing a target face image, characterized in that, it includes: extracting image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, wherein the trained encoder is inverse to the sampling process of the trained StyleGan generator, and the target face image belongs to a low-resolution face image; inputting the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-resolution face image after repairing the target face image; the StyleGan generator first passes through a fully connected layer structure and then through a preset number of convolutional and upsampling structures, and the encoder first passes through a preset number of convolutional and downsampling structures and then through a fully connected layer structure, wherein the convolutional kernel size of the encoder and the StyleGan generator, the number of fully connected layer structures, and the preset number are the same; the step of extracting image features from the target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image includes: extracting first image features from the target face image through the first convolutional and downsampling structure in the trained encoder to generate a first feature map; extracting the i-th image feature from the (i-1)-th feature map through the i-th convolutional and downsampling structure in the trained encoder to generate the i-th feature map, and repeating the generation step of the i-th feature map until all the convolutional and downsampling structures in the encoder are traversed to obtain a target feature map, where i is an integer greater than 1 and not exceeding the preset number, and i takes values from small to large; mapping the target feature map through the trained encoder to obtain a first one-dimensional feature vector corresponding to the target face image.
2. The method according to claim 1, characterized in that, the step of inputting the first one-dimensional feature vector into the trained StyleGan generator to obtain a first high-resolution face image after repairing the target face image includes: fusing the first one-dimensional feature vector and the k-th feature map through the first convolutional and upsampling structure in the trained StyleGan generator to generate a first repaired feature map; fusing the (j-1)-th repaired feature map and the (k-1)-th feature map through the j-th convolutional and upsampling structure in the trained StyleGan generator to generate the j-th repaired feature map, and repeating the generation step of the j-th repaired feature map until all the convolutional and upsampling structures in the trained StyleGan generator are traversed to obtain a target repaired feature map, where j is an integer greater than 1 and not exceeding the preset number, k is an integer greater than 1 and not exceeding the preset number, j takes values from small to large, and k takes values from large to small; obtaining a first high-resolution face image after repairing the target face image according to the target repaired feature map.
3. The method according to claim 1, characterized in that, the method further includes: determining a training sample set, and each group of training samples in the training sample set includes a low-resolution face image sample, a one-dimensional feature vector sample, and a high-resolution face image sample; Based on the low-resolution face image samples and one-dimensional feature vector samples in the training sample set, train the initial encoder to obtain an intermediate encoder; Based on the low-resolution face image samples and high-resolution face image samples in the training sample set, train the intermediate encoder and the pre-trained intermediate StyleGan generator to obtain a trained encoder and a trained StyleGan generator.
4. The method according to claim 3, wherein, The step of training the initial encoder based on the low-resolution face image samples and one-dimensional feature vector samples in the training sample set to obtain an intermediate encoder includes: extracting features from the low-resolution face image samples through the initial encoder to obtain a second one-dimensional feature vector, determining a first loss function according to the difference between the second one-dimensional feature vector and the one-dimensional feature vector samples, if the first loss function does not meet the first preset condition, then optimize the initial encoder and continue training until the first loss function meets the first preset condition, and determine the encoder corresponding to the first loss function that meets the first preset condition as the intermediate encoder; The step of training the intermediate encoder and the pre-trained intermediate StyleGan generator based on the low-resolution face image samples and high-resolution face image samples in the training sample set to obtain a trained encoder and a trained StyleGan generator includes: extracting features from the low-resolution face image samples through the intermediate encoder to obtain a third one-dimensional feature vector; repairing the third one-dimensional feature vector through the pre-trained intermediate StyleGan generator to obtain a second high-resolution face image, determining a second loss function based on the difference between the second high-resolution face image and the high-resolution face image samples, if the second loss function does not meet the second preset condition, then optimize the intermediate encoder and the intermediate StyleGan generator and continue training until the second loss function meets the second preset condition, and determine the encoder and StyleGan generator corresponding to the second loss function that meets the second preset condition as the trained encoder and the trained StyleGan generator.
5. The method according to claim 3, wherein, For each group of training samples in the training sample set, it is determined through the following steps: Input the pre-determined one-dimensional feature vector samples into the intermediate StyleGan generator to generate the high-resolution face image samples; Adjust the size of the high-resolution face image samples by shrinking and enlarging, and add Gaussian noise with a random size to obtain the low-resolution face image samples.
6. A target face image repair device, wherein, It includes: A generation module, configured to extract image features from a target face image through a trained encoder to generate a first one-dimensional feature vector corresponding to the target face image, wherein the trained encoder is inverse to the sampling process of the trained StyleGan generator, and the target face image belongs to a low-resolution face image; A repair module, configured to input the first one-dimensional feature vector into the trained StyleGan generator to obtain the first high-definition face image after repairing the target face image; The StyleGan generator first passes through a fully-connected layer structure and then through a preset number of convolutional and upsampling structures. The encoder first passes through a preset number of convolutional and downsampling structures and then through a fully-connected layer structure, where the convolutional kernel size of the encoder and the StyleGan generator, the number of fully-connected layer structures, and the preset number are the same; Specifically, the generation module is further configured to extract first image features from the target face image through the first convolutional and downsampling structures in the trained encoder to generate a first feature map; extract the i-th image features from the (i-1)-th feature map through the i-th convolutional and downsampling structures in the trained encoder to generate the i-th feature map, and repeat the generation step of the i-th feature map until all the convolutional and downsampling structures in the encoder are traversed to obtain the target feature map, where i is an integer greater than 1 and not exceeding the preset number, and i takes values from small to large; map the target feature map through the trained encoder to obtain the first one-dimensional feature vector corresponding to the target face image.
7. An electronic device, Characterized in that, It includes a memory, a processor, and a program stored on the memory and capable of running on the processor. When the processor executes the program, it implements the target face image repair method according to any one of claims 1 to 5.
8. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the target face image repair method according to any one of claims 1-5.
Citation Information
Patent Citations
A semantic image restoration method based on a DenseNet generative adversarial network
CN109559287A