A face image restoration method and system based on the balance of generation and discrimination confrontation

Through the LesT-GAN network combined with the locally enhanced sliding window Transformer module and the mask-guided patch-level discriminator, the problem of structural distortion and texture blur in image repair is solved, and high-quality image repair and strong generalization capabilities are achieved.

CN116468638BActive Publication Date: 2025-07-11SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310463824.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-07-11
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

Existing image repair methods have challenges in generating reasonable structural and fine-grained textures, resulting in distorted structural or blurred textures from the generated repair results.

Method used

Using a LesT-GAN network based on generation and identification balanced adversarial, combined with a locally enhanced sliding window Transformer module and a mask-guided patch-level discriminator, the generator network is trained by reconstructing loss, perceived loss and style loss to generate high-quality image repair results.

Benefits of technology

显著提高了图像修复的质量和泛化能力,生成器网络训练更稳定,生成更高质量的局部细粒度纹理,减少计算需求并保持较大的感受野。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468638B_ABST
    Figure CN116468638B_ABST
Patent Text Reader

Abstract

The present invention relates to a face image restoration method and system based on the balance of generation and discrimination confrontation. The method includes: S1: Processing the original real image x r and the mask m to obtain the masked image x m ; S2: Inputting x m into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. x m passes through the generator network to generate the restored image x f ; Step S3: Inputting x f , x r and m into the discriminator network to output a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in x f or x r is real or fake; S4: Constructing a total loss function through the reconstruction loss, perceptual loss, style loss, and adversarial loss of the generator network for training the LesT-GAN, and synthesizing reasonable content and clear texture for the missing area of the to-be-restored x m . The method proposed by the present invention makes the quality of image restoration higher and has strong generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a face image restoration method and system based on balanced generation and discrimination confrontation. Background Art

[0002] Image restoration is an important task of synthesizing and replacing visually realistic and semantically consistent content in large missing areas, which requires understanding the semantic structure of the image and performing image generation. It has a wide range of applications in photo and video editing, removal, and restoration. The purpose of edge detection is to capture areas where the brightness changes sharply, and these areas are usually what we focus on. In an image, the areas of two-degree discontinuity are usually one of the following: the image depth discontinuity, the image (gradient) orientation discontinuity, the image illumination (intensity) discontinuity, and the texture change.

[0003] Currently, there are two image restoration methods in computer vision. One is the traditional technology based on diffusion or patches, and the other is the deep learning technology based on deep convolutional networks (Convolutional Neural Network, CNN) and generative adversarial networks (Generative Adversarial Networks, GAN). The former method can synthesize reasonable textures but cannot understand high-level semantics, making it challenging to reconstruct local complex details. The latter method uses the way of jointly training a convolutional-based encoder-decoder network and an adversarial network to generate perceptually and semantically reasonable image content and has achieved a series of remarkable results. However, there are still some problems with the model of this method: the generated restoration results either cannot reconstruct a reasonable structure or cannot reconstruct fine-grained textures, often resulting in structural distortion or texture blur. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a face image restoration method and system based on balanced generation and discrimination confrontation.

[0005] The technical solution of the present invention is as follows: A face image restoration method based on balanced generation and discrimination confrontation, including:

[0006] Step S1: Process the original real image x r and the mask m to obtain the masked image x m ;

[0007] Step S2: Input the masked image x m into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder, and input the masked image x mThrough the generator network, a repaired image x is generated f ;

[0008] Step S3: Input the repaired image x f , the original real image x r and the mask m into the discriminator network, and output a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in x f or x r is real or fake;

[0009] Step S4: Construct a total loss function through reconstruction loss, perceptual loss, style loss, and adversarial loss of the generator network for training the LesT-GAN network, and synthesize reasonable content and clear texture for the missing area of the masked image x m to be repaired.

[0010] Compared with the prior art, the present invention has the following advantages:

[0011] 1. The present invention discloses a face image repair method based on balanced adversarial generation and discrimination, and proposes a locally enhanced sliding window Transformer, namely the LesT module, which restricts the calculation of self-attention within each window, and at the same time better interacts with other windows through sliding window operations, greatly reducing the computational requirements while still maintaining a large receptive field; on the other hand, due to the limitations of existing standard Transformers in extracting low-level features to obtain local dependencies, the present invention adds a depth convolution module between the two fully connected layers of the LeFF module in the LesT module to better capture local context and thus achieve the effect of local enhancement.

[0012] 2. Traditional image repair methods, such as image repair methods based on deep convolutional networks and generative adversarial networks, often result in structural distortion or texture blur due to the problem that the generated repair results either cannot reconstruct a reasonable structure or cannot reconstruct fine-grained textures. The discriminator network proposed by the present invention can be guided by a patch-level mask, enabling the generator network training to directly rely on real pictures, promoting more stable training of the generator and generating higher-quality local fine-grained textures.

[0013] 3. Through ablation studies, it is found that the generator network and discriminator network of the LesT-GAN proposed by the present invention are significantly superior to existing methods, and the proposed LesT-GAN is further evaluated in practical applications. The results show that the model can also achieve promising results in the real world, and at the same time generalizes well to a resolution higher than that of the images during training, with strong generalization ability. Description of the Drawings

[0014] Figure 1 This is a flowchart of a face image restoration method based on balanced generation and discrimination confrontation in an embodiment of the present invention;

[0015] Figure 2 This shows the mask image x in an embodiment of the present invention m Schematic diagram of the generation process;

[0016] Figure 3 This is a schematic diagram of the structure of the LesT-GAN network in an embodiment of the present invention;

[0017] Figure 4 This is a schematic diagram of the structure of a group of LesT modules in an embodiment of the present invention;

[0018] Figure 5 This is a schematic diagram of the structure of the LeFF module in an embodiment of the present invention;

[0019] Figure 6 This is a schematic diagram of the discrimination result of the discriminator network in an embodiment of the present invention;

[0020] Figure 7 This is a structural block diagram of a face image restoration system based on balanced generation and discrimination confrontation in an embodiment of the present invention. Detailed implementation manners

[0021] The present invention provides a face image restoration method based on balanced generation and discrimination confrontation, with higher image restoration quality and stronger generalization ability.

[0022] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention through specific embodiments and in conjunction with the accompanying drawings.

[0023] To accurately describe the content of the present invention, the following terms and their meanings are explained in the present invention:

[0024] Image restoration: Utilize the edges of the damaged areas, that is, the colors and structures of the edges, and infer the information content of the damaged information areas based on the information left in these images, and then fill the damaged areas to achieve image repair.

[0025] Gaussian filtering: Gaussian filtering is a linear smoothing filter suitable for eliminating Gaussian noise. In fact, it is a process of weighted averaging the entire image, and the value of each pixel point is obtained by weighted averaging its own value and the values of other pixels in the neighborhood.

[0026] Mask: A mask is a string of binary codes that performs a bitwise AND operation on the target field to mask the current input bits.

[0027] Skip connection: It skips some layers in the neural network and uses the output of one layer as the input of the next layer. In relatively deep networks, it solves the problems of gradient explosion and gradient disappearance during training.

[0028] Embodiment 1

[0029] As Figure 1 shown, a face image restoration method based on the balance of generation and discrimination confrontation provided by an embodiment of the present invention includes the following steps:

[0030] Step S1: Process the original real image x r and the mask m to obtain the masked image x m ;

[0031] Step S2: Input the masked image x m into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder. Pass the masked image x m through the generator network to generate the restored image x f ;

[0032] Step S3: Input the restored image x f , the original real image x r and the mask m into the discriminator network, and output a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in x f or x r is real or fake;

[0033] Step S4: Construct a total loss function through reconstruction loss, perceptual loss, style loss, and generator network adversarial loss to train the LesT-GAN network, and synthesize reasonable content and clear texture for the missing area of the masked image x m to be restored.

[0034] In one embodiment, the above step S1: Process the original real image x r and the mask m to obtain the masked image x m , specifically including:

[0035] Step S11: Acquire the original real image x r , and at the same time randomly generate a binary mask m, where m = 1 represents the missing area and m = 0 represents the known area;

[0036] The real face images on the CelebA-HQ dataset are selected as the images to be repaired in the embodiments of the present invention, and irregular binary masks m (m = 1 represents the missing area, m = 0 represents the known area) are randomly generated from the standard convolutional layer and the sigmoid function in PConv, and are used for training and testing according to common settings. These masks are divided into different categories according to the ratio of the masked area to the size of the whole image. In the embodiments of the present invention, the sizes of all masks and masked images used for training and testing are set to 256×256;

[0037] Step S12: Use x m = x r ⊙(1 - m) to represent the masked image x m ; where ⊙ represents pixel-wise multiplication.

[0038] As Figure 2 shown, the generation process of the masked image x m is demonstrated.

[0039] As Figure 3 shown, in one embodiment, in the above step S2: input the masked image x m into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder. After passing the masked image x m through the generator network, a repaired image x f is generated, which specifically includes:

[0040] Step S21: Construct an encoder-decoder based on Transformer as the generator network:

[0041] The encoder of the generator network is composed of 3 downsampling modules. Each downsampling module contains a group of LesT modules and a downsampling layer. Each group of LesT modules is composed of two consecutive LesT blocks. One LesT block contains a multi-head self-attention W-MSA module and a LeFF module, and the other LesT block contains a sliding window-based multi-head self-attention SW-MSA module and a LeFF module;

[0042] The calculation process of the LesT module is as follows:

[0043]

[0044]

[0045]

[0046]

[0047] Among them, and respectively represent the output features of the W-MSA module and the SW-MSA; represents the output feature of the LeFF module; LN() represents layer normalization, and l represents the layer number;

[0048] The window-based multi-head self-attention (W-MSA) does not use global self-attention like ordinary Transformers, but calculates self-attention within a local window. The image is evenly divided into non-overlapping windows, and calculating self-attention for each window separately significantly reduces the computational cost. The window-based self-attention module lacks cross-window connections, which limits its overall modeling ability. To achieve cross-window information interaction while maintaining the efficient calculation of non-overlapping windows, the sliding window-based multi-head self-attention (SW-MSA) is used to implement the sliding window operation. Therefore, as Figure 4 shown, a group of LesT modules includes 2 consecutive LesT blocks, where the W-MSA and SW-MSA modules are alternately used to calculate self-attention, and relative position encoding is applied to the attention module; the Norm module in the LesT module represents the layer normalization module.

[0049] The LeFF module includes: the Tokens2Img module is used to convert the output of the W-MSA module or the SW-MSA into a feature map, the 3×3 depth convolution module is used to perform feature fusion on the feature map, and the Img2Tokens module is used to restore the initial dimension of the input;

[0050] The present invention designs a locally enhanced feed-forward network (LeFF) module to replace the feed-forward neural network (FFN) module of the traditional Transformer. The LeFF module can combine the advantage of the CNN in extracting local information with the ability of the Transformer to establish long-term dependencies. As Figure 5 shown, a 3×3 depth convolution layer is added between the two fully connected layers FC of the LeFF module to better capture local context and thus achieve the effect of local enhancement. To correspond to the convolution operation, first, the Tokens2Img module applies linear projection and spatial restoration to the sequence generated by the W-MSA module or the SW-MSA to convert it into a feature map. Then, feature fusion is performed using a 3×3 depth convolution to capture local information. Finally, the Img2Tokens module flattens the feature map into a sequence and performs linear projection to match the input initial dimension;

[0051] The decoder of the generator network consists of 3 upsampling modules, and each upsampling module contains an upsampling layer and a group of LesT modules; the LesT modules of the decoder have the same structure as those of the encoder, which will not be elaborated here.

[0052] Each upsampling module and its corresponding downsampling module are skip-connected to solve the problems of gradient explosion and gradient disappearance during training;

[0053] Six groups of LesT modules are stacked between the encoder and the decoder to better capture long-range context information;

[0054] Step S22: Input the masked image x m into the generator network, encode it through the encoder to obtain the encoded image, and then pass the encoded image through the decoder to obtain the restored image x f .

[0055] The LesT module proposed in the present invention benefits from the rapid growth of the receptive field, thereby obtaining higher quality and parameter efficiency. It restricts the calculation of self-attention within each window, and at the same time better interacts with other windows through the sliding window operation, greatly reducing the computational requirements while still maintaining a large receptive field. On the other hand, due to the limitations of existing standard Transformers in extracting low-level features to obtain local dependencies, the present invention adds a depth convolution layer between two fully connected layers in the LeFF module of the LesT module to better capture local context and thus achieve the effect of local enhancement.

[0056] In one embodiment, the above step S3: Input the restored image x f , the original real image x r and the mask m into the discriminator network, and output a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in x f or x r is real or fake, specifically including:

[0057] Step S31: Construct a mask-guided and patch-based relative average discriminator as the discriminator network. Among them, the discriminator network consists of several downsampling convolutional layers, and each downsampling convolution simulates pixel diffusion around the boundary of the missing area through a patch mask obtained by Gaussian filtering;

[0058] Step S32: Input x f , x r and m into the discriminator network, and output a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in the restored image or the original real image is real or fake;

[0059] The present invention proposes to use a mask-guided patch-level relative average discriminator (MRA-PatchGAN) to replace the traditional standard discriminator in the image inpainting task. To simulate pixel propagation around the boundary of the missing region, the present invention uses a soft patch-level mask for training, where the soft mask is obtained by Gaussian filtering. The MRA-PatchGAN discriminator utilizes both real data and fake data when predicting the authenticity of samples, and predicts the probability of being relatively true instead of absolutely true or false, which can lead to a stronger discriminator. As Figure 6 shown, each pixel in the prediction map of the real image represents the relative probability value that the prediction of the N×N patch in the input image is more real than the inpainted image; each pixel in the prediction map of the inpainted image represents the relative probability value that the prediction of the N×N patch in the input image is more real than the real image. Comparing the MRA-PatchGAN discriminator of the present invention with the existing PatchGAN, HM-PatchGAN, and SM-PatchGAN discriminators, the MRA-PatchGAN discriminator can reduce the evaluation of noise and unimportant regions in the image by the discriminator, and at the same time uses the gradient penalty technique to solve the mode collapse and gradient vanishing problems, which can make the training smoother, reduce unstable gradient updates, help reduce the mode collapse problem during the training process, and improve the stability of the adversarial generative network training. By using a mask to restrict the discriminator to only evaluate the missing region of the image, the discriminator can be more focused on evaluating the details and structures generated by the generator, thereby enabling the generator to generate higher-quality images, promoting the generator to generate more real and detail-rich images, and improving the quality of the generator. MRA-PatchGAN divides the image into multiple patches and evaluates each patch in the discriminator to identify aspects that need improvement, providing a more fine-grained evaluation result. At the same time, since MRA-PatchGAN considers the relative difference between the generated samples and the real samples, rather than just the absolute difference, this makes the evaluation result more objective and accurate.

[0060] The MRA-PatchGAN discriminator proposed by the present invention enables the training of the generator network to directly rely on real pictures, promoting more stable training of the generator and generating higher-quality local fine-grained textures;

[0061] Step S33: Define the adversarial loss of the generator network as:

[0062]

[0063]

[0064] where C(.) represents the output of the discriminator without the sigmoid function at the end;

[0065] where represents xr Extract the expected value of the feature distribution function Denote x f Extract the expected value of the feature distribution function Denote x f And the expected value of the feature distribution function of m;

[0066] Define the adversarial loss of the discriminator network as:

[0067]

[0068] In one embodiment, the above step S4: Construct a total loss function through the reconstruction loss, perceptual loss, style loss, and adversarial loss of the generator network for training the LesT-GAN network, for the masked image x to be repaired m Synthesize reasonable content and clear texture for the missing region, specifically including:

[0069] Step S41: Calculate the reconstruction loss L rec , by calculating the repaired image x f And the original real image x r Measure the reconstruction accuracy of the pixel-level difference by the similarity between them:

[0070] L rec =||x r -x f ||1

[0071] Where, ||||1 represents the first norm;

[0072] Step S42: Calculate the perceptual loss L per , which uses the perceptual loss based on the pre-trained network to make x f Semantically closer to x r :

[0073]

[0074] Where, Is the activation map of the i-th layer from the pre-trained network, N I Is The number of elements in;

[0075] In the embodiment of the present invention, the VGG-19 network is pre-trained using the ImageNet dataset, and the layers pool1, pool2, and pool3 of VGG-19 are used for loss calculation;

[0076] Step S43: Calculate the style loss L sty , use the style loss based on the pre-trained network to make the texture of x f Similar to that of x rSimilar, that is, to make the L1 distance between the Gram matrix of the deep features of the restored image and the original real image closer:

[0077]

[0078] Step S44: Calculate the total loss function:

[0079] L total = λ adv L G + λ rec L rec + λ per L per + λ sty L sty

[0080] Wherein, λ adv 、λ rec 、λ per and λ sty are trade-off parameters; in the embodiments of the present invention, λ adv = 0.01, λ rec = 1, λ per = 0.1, and λ sty = 250.

[0081] To verify the effectiveness of the LesT block proposed in the present invention, the dataset CelcbA-HQ is used, and based on the metrics: L1 loss, PSNR, SSIM, and FID, where L1 is the mean square error, and the real image and the restored image are compared pixel by pixel. The smaller the value, the closer the restored image is to the original image, and vice versa. The peak signal-to-noise ratio (PSNR) is usually solved based on MSE, so it is still a pixel-by-pixel image comparison. The closer two images are, the smaller the PSNR value, and vice versa. The structural similarity (Structural Similarity Index, SSIM) is mainly used to compare the brightness difference, contrast difference, and structural difference between the real image and the restored image. FID (Fréchet Inception Distance) is a metric used to evaluate the image quality of a generative model. It can be used to compare the difference between the image generated by the generative model and the real image. The smaller the value of FID, the smaller the distance between the generated image and the real image in the feature space, and the better the performance of the generative model. An ablation experiment was conducted on the LesT block of the present invention with the GatedConv block, AOT block, and Swin Transformer block, and the comparison results are shown in Table 1:

[0082] Table 1 Ablation experiment results of the LesT block on the CelebA-HQ dataset

[0083]

[0084] Note: “↑” indicates that the higher the value, the better, “↓” indicates that the lower the value, the better, and the bold font is the optimal value for each line.

[0085] It can be seen from Table 1 that the LesT block proposed by the present invention has obtained significant improvements in terms of L1 loss, PSNR, SSIM, and FID.

[0086] Similarly, in order to verify the effectiveness of the discriminator MRA-PatchGAN proposed by the present invention, the dataset CelebA-HQ was used, and an ablation experiment was conducted on the MRA-PatchGAN of the present invention with PatchGAN and SM-PatchGAN blocks based on the metrics: L1 loss, PSNR, SSIM, and FID. The comparison results are shown in Table 2:

[0087] Table 2 Ablation experiment results of MAR-PatchGAN on the CelebA-HQ dataset

[0088]

[0089] Note: “↑” indicates that the higher the value, the better, “↓” indicates that the lower the value, the better, and the bold font is the optimal value for each line.

[0090] According to the comparison results in Table 2, the MRA-PatchGAN of the present invention is superior to PathGAN and SM-PatchGAN in all evaluation metrics. The MRA-PatchGAN of the present invention can distinguish between missing region patches and real region patches, enabling the generator to directly utilize real images during training. This optimization forces the discriminator to capture realistic textures, thereby promoting the generator to synthesize clearer textures for high-resolution image inpainting.

[0091] In addition, the face image inpainting method based on the balance of generation and discrimination provided by the present invention has better generalization ability. The qualitative and quantitative comparison results of LesT-GAN in three actual application datasets show that LesT-GAN is significantly superior to the state-of-the-art methods. In addition, the role of each component of LesT-GAN was analyzed through ablation studies to verify the effectiveness of each component. The proposed LesT-GAN was further evaluated in actual applications, and the results show that the model can also achieve promising results in the real world and can be well generalized to resolutions higher than the resolution of the images during training, with strong generalization ability.

[0092] Example 2

[0093] As Figure 7As shown in the figure, an embodiment of the present invention provides a face image restoration system based on balanced generation and discrimination confrontation, including the following modules:

[0094] A mask image generation module 81 for processing the original real image x r and the mask m to obtain the mask image x m ;

[0095] A restored image generation module 82 for inputting the mask image x m into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder. The mask image x m passes through the generator network to generate the restored image x f ;

[0096] A prediction map generation module 83 for inputting the restored image x f , the original real image x r and the mask m into the discriminator network and outputting a prediction map. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in x f or x r is real or fake;

[0097] A loss function construction module 84 for constructing a total loss function through reconstruction loss, perceptual loss, style loss, and generator network adversarial loss, for training the LesT-GAN network to synthesize reasonable content and clear texture for the missing area of the mask image x m to be restored.

[0098] The above embodiments are provided only for the purpose of describing the present invention and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. All equivalent substitutions and modifications made without departing from the spirit and principle of the present invention shall be covered within the scope of the present invention.

Claims

1. A face image restoration method based on balanced adversarial generation and discrimination, characterized in that Including: Step S1: Process the original real image and the mask to obtain a masked image ; Step S2: Take the masked image as the input of the LesT-GAN network, which includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder. Take the masked image as the input of the generator network, and generate a restored image . Specifically, it includes: Step S21: Construct an encoder-decoder based on Transformer as the generator network: The encoder of the generator network consists of 3 downsampling modules. Each downsampling module contains a group of LesT modules and a downsampling layer. Each group of LesT modules consists of two consecutive LesT blocks. One LesT block contains a multi-head self-attention W-MSA module and a LeFF module, and the other LesT block contains a sliding window-based multi-head self-attention SW-MSA module and a LeFF module; The calculation process of the LesT module is as follows: Among them, and represent the output features of the W-MSA module and the SW-MSA respectively; represents the output feature of the LeFF module; () represents layer normalization, represents the number of layers; The LeFF module includes: a Tokens2Img module for converting the output of the W-MSA module or SW-MSA into a feature map, a 3×3 depth convolution module for performing feature fusion on the feature map, and an Img2Tokens module for restoring the initial dimension of the input; The decoder of the generator network consists of 3 upsampling modules. Each upsampling module contains an upsampling layer and a group of the LesT modules; Each group of the upsampling modules and the corresponding group of the downsampling modules are respectively subjected to skip connections; 6 groups of the LesT modules are stacked between the encoder and the decoder; Step S3: Input the repaired image , the original real image , and the mask into the discriminator network to output a prediction map, where each pixel of the prediction map represents whether the prediction of the N×N pixel block in or is real or fake; Step S4: Construct a total loss function through reconstruction loss, perceptual loss, style loss, and generator network adversarial loss for training the LesT-GAN network to synthesize reasonable content and clear texture for the missing regions of the masked image to be repaired. ​ 2. The face image restoration method based on the balance confrontation of generation and discrimination according to claim 1, wherein The said step S1: Process the original real image and the mask to obtain a masked image , specifically including: Step S11: Collect the original real image , and simultaneously randomly generate a binary mask , where = 1 indicates the missing area, = 0 indicates the known area; Step S12: Use = (1 - m) to represent the masked image ; where represents pixel-level multiplication.

3. The face image restoration method based on the generative and discriminative balance confrontation according to claim 2, wherein The said step S2: the masked image is input into the LesT-GAN network, which includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder. The masked image passes through the generator network to generate a repaired image , and further includes the following steps: Step S22: Input the masked image into the generator network, encode it through the encoder to obtain an encoded image, and then pass the encoded image through the decoder to obtain a restored image .

4. The face image restoration method based on the balance of generation and discrimination confrontation according to claim 3, characterized in that, The said step S3: the repaired image , the original real image , and the mask are input into the discriminator network to output a prediction map, and each pixel of the prediction map represents whether the prediction of the N×N pixel block in or is real or false, specifically including: Step S31: Construct a mask-guided and patch-based relative average discriminator as the discriminator network. Among them, the discriminator network consists of several downsampling convolutional layers. Each downsampling convolution simulates pixel diffusion around the boundary of the missing area through a patch mask obtained by Gaussian filtering; Step S32: Input , and into the discriminator network, and output a prediction map, where each pixel of the prediction map represents whether the prediction of the N×N pixel block in the restored image or the original real image is real or fake; Step S33: Define the adversarial loss of the generator network as: Among them, C(.) represents the output of the discriminator without the sigmoid function at the end; Among them, represents the expected value of the extracted feature distribution function, represents the expected value of the extracted feature distribution function, represents and the expected value of the extracted feature distribution function of; Define the adversarial loss of the discriminator network as: 。 5. The face image restoration method based on the balance of generation and discrimination confrontation according to claim 4, characterized in that, Step S4: Construct a total loss function through reconstruction loss, perceptual loss, style loss, and generator network adversarial loss for training the LesT-GAN network to synthesize reasonable content and clear texture for the missing region of the masked image to be repaired. Specifically, it includes: Step S41: Calculate the reconstruction loss , by calculating the similarity between the repaired image and the original ground truth image to measure the reconstruction accuracy of pixel-level differences: Among them, represents the first normal form; Step S42: Calculate the perceptual loss , which uses the perceptual loss based on the pre-trained network to make semantically closer to : where is the activation map from the i-th layer of the pre-trained network, is the number of elements in; Step S43: Calculate the style loss , and use the style loss based on the pre-trained network to make 's texture similar to : Step S44: Calculate the total loss function: wherein, is a trade-off parameter.

6. A face image restoration system based on the balance of generation and discrimination adversarial, characterized in that, Including the following modules: A mask image generation module for processing an original real image and a mask to obtain a mask image ; A module for generating a repaired image, which is used to process the masked image and input it into the LesT-GAN network. The LesT-GAN network includes a generator network and a discriminator network. Among them, the generator network is composed of an encoder and a decoder, and processes the masked image through the generator network to generate a repaired image , specifically including: Step S21: Construct an encoder-decoder based on Transformer as the generator network: The encoder of the generator network consists of 3 downsampling modules. Each downsampling module contains a group of LesT modules and a downsampling layer. Each group of LesT modules consists of two consecutive LesT blocks. One LesT block contains a multi-head self-attention W-MSA module and a LeFF module, and the other LesT block contains a sliding window-based multi-head self-attention SW-MSA module and a LeFF module; The calculation process of the LesT module is as follows: Among them, and represent the output features of the W-MSA module and the SW-MSA respectively; represents the output feature of the LeFF module; ( ) represents layer normalization, represents the number of layers; The LeFF module includes: a Tokens2Img module for converting the output of the W-MSA module or SW-MSA into a feature map, a 3×3 depth convolution module for performing feature fusion on the feature map, and an Img2Tokens module for restoring the initial dimension of the input; The decoder of the generator network consists of 3 upsampling modules. Each upsampling module contains an upsampling layer and a group of the LesT modules; Each group of the upsampling modules and the corresponding group of the downsampling modules are respectively subjected to skip connections; Six groups of the LesT modules are stacked between the encoder and the decoder; A prediction map generation module, for the repaired image , the original real image and the mask are input into the discriminator network, and a prediction map is output. Each pixel of the prediction map represents whether the prediction of the N×N pixel block in or is true or false; Construct a loss function module for constructing a total loss function through reconstruction loss, perceptual loss, style loss, and generator network adversarial loss for training the LesT-GAN network for the masked image to be repaired to synthesize reasonable content and clear texture for the missing area.