Face image restoration method and system based on mask auto-encoder

Through the face image recovery method based on mask autoencoder, the problem that deep forgery detection algorithm in the prior art is difficult to provide evidence links, and highly interpretable face recovery and detection are achieved, which improves identity similarity.

CN119941895AActive Publication Date: 2025-05-06JINAN UNIVERSITY

Patent Information

Application Number
CN202510007605.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing deep falsification detection algorithms are difficult to provide a link of evidence that matches the detection results, and lack highly interpretable detection methods.

Method used

The face image recovery method based on the mask autoencoder is adopted, and the shading features and high-level semantic features of the source face image are extracted, and the feature stitching and segmentation probability map generation is performed. The fake face image is restored with the mask autoencoder to reconstruct the source face and the target face.

Benefits of technology

The effective recovery of deep-fake face images is achieved, providing a link of evidence matching the detection results, improving the interpretability of the detection algorithm, and the identity similarity on multiple deep-fake data sets reaches 75.82%, surpassing the existing algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941895A_ABST
    Figure CN119941895A_ABST
Patent Text Reader

Abstract

The invention discloses a face image restoration method and system based on a mask auto-encoder, and belongs to the technical field of deep forgery detection, and the method comprises the steps: employing a RetinaFace model to cut a forgery face image, and obtaining a source face image and a target face image corresponding to the forgery face image; extracting features of the source face image, performing up-sampling on high-level semantic features, then splicing the high-level semantic features with shading features, and obtaining a segmentation probability graph based on a spliced feature graph; multiplying the segmentation probability graph by the source face image to obtain source face information, and restoring and reconstructing the source face based on the source face information to obtain a reconstructed source face; partitioning the segmentation probability graph, calculating the mean value of each segmentation probability graph, calculating the weight based on the mean value, and performing masking based on the weight to obtain a covered image; and inputting the potential target face features in the reconstructed source face and the covering image into a mask auto-encoder to obtain a face image recovery image of the forged face image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep fake detection, and in particular to a face image restoration method and system based on a mask autoencoder. Background Art

[0002] The development of artificial intelligence technology has brought convenience to people, but it has also brought potential threats. Deep fake technology can achieve realistic face replacement, and the popularity of short video platforms has exacerbated the impact of deep fake content on society, allowing false information to spread rapidly, posing a huge threat to individuals and society. At present, there is an urgent need to build an effective deep fake defense and detection algorithm to control the current deep fake data from a technical level.

[0003] Current research mainly regards deepfake detection as a binary classification task, aiming to improve the generalization detection performance of various types of forged data. Existing detection algorithms can be roughly divided into data-driven detection methods and model optimization-based detection methods. Data-driven detection methods simulate the face-changing process to generate diverse forged data to improve generalization performance. Model optimization-based methods detect deepfakes by optimizing model structure, designing feature consistency analysis, and time continuity analysis. Although these algorithms have high detection accuracy, they can only provide classification or positioning results, and cannot provide a chain of evidence that matches the detection results. Therefore, how to improve the interpretability of detection algorithms is an urgent problem that needs to be solved. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a face image restoration method based on a mask autoencoder, the method comprising:

[0005] Step S1, collecting a forged face image, cropping the forged face image using a RetinaFace model, and obtaining a source face image and a target face image corresponding to the forged face image;

[0006] Step S2, extracting the background texture features and high-level semantic features of the source face image, upsampling the high-level semantic features and splicing them with the background texture features, and obtaining a segmentation probability map based on the spliced ​​feature map;

[0007] Step S3, multiplying the segmentation probability map by the source face image to obtain source face information, and restoring and reconstructing the source face based on the source face information to obtain a reconstructed source face;

[0008] Step S4, dividing the segmentation probability map into blocks and calculating the mean of each segmentation probability map, assigning weights based on the mean results and calculating the weights, performing masking based on the weights, and obtaining a masked image;

[0009] Step S5: input the potential target facial features in the reconstructed source face and the mask image into the mask autoencoder to obtain a facial image restoration image of the forged facial image.

[0010] Optionally, in step S2, extracting the background texture features and high-level semantic features of the source face image, upsampling the high-level semantic features and splicing them with the background texture features, and obtaining the segmentation probability map based on the spliced ​​feature map specifically includes:

[0011] Use EfficientNet-B0 to extract low-level texture features and high-level semantic features;

[0012] The high-level semantic features are upsampled using transposed convolution, the upsampled high-level semantic features are spliced ​​with the underlying texture features, and the spliced ​​features are input into a transposed convolution layer to generate a segmentation probability map.

[0013] Optionally, in step S3, the source face information is obtained by multiplying the segmentation probability map with the source face image, and the source face is restored and reconstructed based on the source face information to obtain the reconstructed source face. Specifically, the process includes:

[0014] S31, based on the source face information obtained by multiplying the segmentation probability map with the source face image, down-sampling the source face information through 4 convolution-pooling layers to extract a feature map of high-level semantic features;

[0015] S32, dividing the feature map of the high-level semantic features into source face features and potential target face features, upsampling the source face features using transposed convolution, and convolving the potential target face features using a convolution layer to obtain a feature map with a preset number of channels;

[0016] S33, using residual connection to splice the feature map of the high-level semantic feature with the feature map of the preset number of channels, and using convolution to restore the number of channels of the spliced ​​feature map;

[0017] S34. Repeat S31-S33 4 times to restore the feature map of the number of channels, restore the source face image, calculate the identity loss and the perception loss based on the source face image and the source face image in step S1, and obtain the reconstructed source face.

[0018] Optionally, in step S4, the segmentation probability map is divided into blocks and the mean of each segmentation probability map is calculated, weights are assigned and calculated based on the mean results, and masking is performed based on the weights to obtain the masked image, which specifically includes:

[0019] The mask block number is calculated based on the segmentation probability map, the forged face is divided into blocks based on the mask block number, and the forged face after the division is covered by the mask block number to obtain a covered image.

[0020] Optionally, in step S5, inputting the potential target face features in the reconstructed source face and the mask image into the mask autoencoder to obtain the target face image restoration image of the forged face image specifically comprises:

[0021] Inputting the potential target face features in the reconstructed source face and the uncovered image blocks in the masked image into a masked autoencoder for flattening and merging into a tensor, inputting the tensor into a fully connected layer, and obtaining a tensor image feature vector;

[0022] Input the tensor image feature vector into the Transformer module, and calculate the global features of the image through a multi-head attention mechanism;

[0023] The mask block filling is performed on the global features of the image, and the filled global features of the image are input into a decoder to perform an operation of restoring the image block content, thereby obtaining a facial image restoration image of the forged facial image.

[0024] The present invention also discloses a face image restoration system based on a mask autoencoder, the system comprising: an image cropping module, a feature map splicing module, a source face restoration module, an image mask module and an autoencoder restoration module;

[0025] The image cropping module is used to collect a forged face image, crop the forged face image using a RetinaFace model, and obtain a source face image and a target face image corresponding to the forged face image;

[0026] The feature map splicing module is used to extract the background features and high-level semantic features of the source face image, upsample the high-level semantic features and splice them with the background features, and obtain a segmentation probability map based on the spliced ​​feature map;

[0027] The source face restoration module is used to obtain source face information by multiplying the segmentation probability map with the source face image, and to restore and reconstruct the source face based on the source face information to obtain a reconstructed source face;

[0028] The image mask module is used to divide the segmentation probability map into blocks and calculate the mean of each segmentation probability map, assign weights based on the mean results and calculate the weights, and perform masking based on the weights to obtain a masked image;

[0029] The autoencoder restoration module is used to input the target facial features in the reconstructed source face and the mask image into the mask autoencoder to obtain a facial image restoration image of the forged facial image.

[0030] Optionally, the workflow of the feature map splicing module specifically includes:

[0031] Use EfficientNet-B0 to extract low-level texture features and high-level semantic features;

[0032] The high-level semantic features are upsampled using transposed convolution, the upsampled high-level semantic features are spliced ​​with the underlying texture features, and the spliced ​​features are input into a transposed convolution layer to generate a segmentation probability map.

[0033] Optionally, the source face restoration module includes: a semantic feature extraction submodule, a channel convolution submodule, a channel restoration submodule and an image reconstruction submodule;

[0034] The semantic feature extraction submodule is used to obtain source face information based on the multiplication of the segmentation probability map and the source face image, downsample the source face information through four convolution-pooling layers, and extract a feature map of high-level semantic features;

[0035] The channel convolution submodule is used to divide the feature map of the semantic feature into source face features and potential target face features, upsample the source face features using transposed convolution, and convolve the potential target face features using a convolution layer to obtain a feature map with a preset number of channels;

[0036] The channel restoration submodule is used to splice the feature map of high-level semantic features with the feature map of the preset number of channels using residual connection, and restore the number of channels of the spliced ​​feature map using convolution;

[0037] The image reconstruction submodule is used to repeat the feature map of the restored channel number 4 times S31-S33, restore the source face image, calculate the identity loss and the perceptual loss based on the source face image and the source face image in step S1, and obtain the reconstructed source face.

[0038] Optionally, the workflow of the image mask module specifically includes: calculating the mask block number based on the segmentation probability map, dividing the forged face into blocks based on the mask block number, and using the mask block number to mask the forged face after the block division to obtain a masked image.

[0039] Optionally, the workflow of the autoencoder recovery module specifically includes:

[0040] Inputting the potential target face features in the reconstructed source face and the uncovered image blocks in the masked image into a masked autoencoder for flattening and merging into a tensor, inputting the tensor into a fully connected layer, and obtaining a tensor image feature vector;

[0041] Input the tensor image feature vector into the Transformer module, and calculate the global features of the image through a multi-head attention mechanism;

[0042] The mask block filling is performed on the global features of the image, and the filled global features of the image are input into a decoder to perform an operation of restoring the image block content, thereby obtaining a facial image restoration image of the forged facial image.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention proposes a deep fake face restoration algorithm based on a masked autoencoder, and uses the restored source face and target face as evidence to support the detection results. First, an identity information segmentation module is constructed to segment the source / target face information to obtain an identity segmentation map; the source face information is further segmented and input into the source face restoration module to reconstruct the source face and extract potential target identity features; finally, the segmentation map is used to generate a mask to cover the image to be detected, and a target face restoration module based on a masked autoencoder is constructed to restore the masked area and reconstruct the target face. After testing on 3 deep fake data sets and 6 types of deep fake algorithms, the target face restored by the technical solution of the present invention has an identity similarity of up to 75.82% with the original target face, surpassing the existing algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0046] Figure 1 This is a framework diagram of the deep fake face recovery algorithm based on masked autoencoder;

[0047] Figure 2 This is the Transformer module architecture diagram used in the masked self-encoder. Specific implementation methods

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Embodiment 1

[0052] A face image restoration method based on mask autoencoder, such as Figure 1 As shown, specifically including:

[0053] Step S1: collect a forged face image, and use a RetinaFace model to crop the forged face image to obtain a source face image and a target face image corresponding to the forged face image.

[0054] RetinaFace is used to locate and align the forged face images in the dataset, and then a 224×224 face image is cropped based on the coordinates of the facial feature points. Similarly, the corresponding source and target face images are obtained.

[0055] Step S2, extracting the texture features and high-level semantic features of the source face image, upsampling the high-level semantic features and splicing them with the texture features, and obtaining a segmentation probability map based on the spliced ​​feature map, specifically comprising: using the pre-trained EfficientNet-B0 to extract features, taking the first layer features as the underlying texture features, taking the second-to-last layer features as the high-level semantic features, using 4 transposed convolutions to upsample the high-level semantic features, splicing them with the underlying texture features, generating a segmentation map through the transposed convolution layer again, and constraining it to the range of 0-1 using the Sigmoid function.

[0056] Step S3, multiplying the segmentation probability map by the source face image to obtain source face information, and restoring and reconstructing the source face based on the source face information to obtain a reconstructed source face, including: S31, based on the source face information obtained by multiplying the segmentation probability map by the source face image, down-sampling the source face information through 4 convolution-pooling layers to extract a feature map of high-level semantic features; S32, segmenting the high-level semantic features into source face features and potential target face features, using transposed convolution to downsample the source face features; The feature is upsampled, and the potential target face feature is convolved by a convolution layer to obtain a feature map with a preset number of channels; S33, the feature map of the high-level semantic feature is spliced ​​with the feature map with the preset number of channels by residual connection, and the number of channels is restored by convolution of the spliced ​​feature map; S34, S31-S33 are repeated 4 times for the feature map with the restored number of channels, the source face image is restored, and the identity loss and the perceptual loss are calculated based on the source face image and the source face image in step S1 to obtain a reconstructed source face. In this embodiment, we set the number of channels to 98.

[0057] Specifically, S31, inputting the source face information into the source face recovery module, performing downsampling operation through 4 convolution-pooling layers, and extracting high-level semantic features of the image;

[0058] S32, split the feature map of the high-level semantic features obtained in S31 into two parts: source face features and potential target face features, use a transposed convolution layer to upsample the source face features; use a convolution layer to process the potential target face features so that the number of channels is consistent with the number of channels in the target face image. In this embodiment, we set the number of channels to 98.

[0059] S33, concatenate the feature map calculated by the downsampling layer operation with the corresponding feature in the upsampling stage through residual connection, and restore the number of channels by convolution operation; in this embodiment, the number of restored channels is the same as the number of channels of the feature map calculated by the downsampling operation. For example: the feature obtained by downsampling is 512 channels, and the corresponding upsampling stage obtains 512 channels. After concatenation, the two are 1024 channels, and then the convolution operation is used to restore them to 512 channels.

[0060] S34. Restore the source face image through four transposed convolution-residual connection-convolution operations, and then calculate the identity loss and perceptual loss with the original source face image to optimize the source face recovery module.

[0061] Step S4, divide the segmentation probability map into blocks and calculate the mean of each segmentation probability map, assign weights based on the mean result and calculate the weights, perform masking based on the weights, and obtain a masked image, specifically including: calculating the mask block number based on the segmentation probability map, dividing the forged face into blocks based on the mask block number, and using the mask block number to mask the forged face after the block division to obtain a masked image. S41, divide the segmentation probability map into blocks, calculate the mean of the segmentation probability map, and the mean calculation process of the probability map is: because the values ​​of the probability map are all probabilities of 0-1, add the values ​​of the probability maps of all segmentation maps and calculate the mean. If the mean is less than 0.5, set the weight of each block to a random value between 0-1; if the mean is greater than or equal to 0.5, calculate the mean of each block as the weight; S42, sort the weights calculated in S41 from large to small, define the mask ratio as β, and mask the image blocks with larger weights. In this embodiment, we set β to 0.5.

[0062] Step S5, inputting the potential target facial features in the reconstructed source face and the masked image into the masked autoencoder to obtain a facial image restoration image of the forged facial image, specifically comprising: inputting the potential target facial features in the reconstructed source face and the uncovered image blocks in the masked image into the masked autoencoder for flattening, merging into a tensor, inputting the tensor into a fully connected layer to obtain a tensor image feature vector; inputting the tensor image feature vector into a Transformer module, calculating the image global features through a multi-head attention mechanism; performing mask block filling on the image global features, inputting the filled image global features into a decoder to perform an operation of restoring the image block content, and obtaining a facial image restoration image of the forged facial image.

[0063] S51, dividing the image blocks into blocks, and masking the image blocks according to the mask numbers calculated in S42 (pixels are set to 0); in this embodiment, we set the block size to 16, each image is divided into 196 image blocks, and 98 image blocks are retained after masking;

[0064] S52: Input the uncovered image block into the masked autoencoder to restore the image block content corresponding to the target face image.

[0065] S521, flatten the uncovered image blocks and merge them into a tensor; input the tensor into a fully connected layer to obtain the corresponding image feature vector;

[0066] S522, input the image feature vector into the Transformer module (such as Figure 2As shown in Figure 2, the global features of the image are calculated through a multi-head self-attention mechanism; the final global features are obtained by repeating the Transformer module calculation L times. The global feature calculation process is as follows:

[0067] f m =F(W·x sv )+b)+p sv );

[0068] Among them, F represents the encoding network composed of L Transformer modules, x sv represents the unmasked image tensor, W and b represent the weight and bias of the fully connected layer respectively, and p sv Represents x sv The corresponding position code. In this embodiment, L is set to 7.

[0069] S523, using the mask block sequence number, fill the global features with mask blocks to make them correspond to the complete image block sequence; fill the unmasked image block features to make them consistent with the number of image blocks. Assuming that the number of unmasked image blocks is 52, the number of complete image block features is 64, and the empty features with a value of 0 are used to fill 52 to 64.

[0070] S524, input the filled global features into the decoder (composed of n Transformer modules), restore the image block content through multi-head attention mechanism, normalization layer and other operations, and calculate the identity loss, perception loss, attribute loss and pixel recovery loss with the original target face image. The process of restoring the image content is as follows:

[0071] x rec =Sigmoid(D(f m ||f t ));

[0072] Where D represents the decoder consisting of n Transformer modules, f m represents the features obtained by encoding the uncovered image, f t represents the potential target face features extracted in the source face restoration module, and the Sigmoid function is used to constrain the output to the range of 0 to 1. In this embodiment, n is set to 3.

[0073] Embodiment 2

[0074] The scheme was experimented on three datasets: FaceForensics++, CelebaMegaFS, and FFHQ-E4S. Among them, Faceforensics++ contains two existing deep fake face-changing methods, DeepFake (DF) and FaceShifter (FShi). In addition, another set of fake data was generated using the LCR algorithm. These three types of fake images are split into corresponding training sets, validation sets, and test sets according to the division strategy specified by FaceForensics++. CelebaMegaFS contains three deep fake methods: IDInjection (IDI), FTM, and LCR, generating 30,038 FTM images, 30,441 IDInjection images, and 30,010 LCR images. FFHQ-E4S is built on the FFHQ dataset. We randomly selected 10,000 images from FFHQ and used the E4S algorithm to generate 10 swapped images for each original image. For CelebaMegaFS and FFHQ-E4S, each dataset is divided into training set, validation set and test set in the ratio of 8:1:1.

[0075] In the experiment, FID is used to evaluate the quality of image restoration, and IDSim is used to evaluate the identity similarity between the restored image and the original target face image. Table 1 shows the restoration performance on the three datasets. The "None" method in the first row shows the FID score between the forged face and the original target face, and the next six rows show the values ​​of the corresponding indicators between the restored face and the original target face. The image restoration results on FaceForensics++, CelebaMegaFS, and FFHQ-E4S datasets are analyzed. Compared with existing face restoration methods (MAT, Repaint) and deep fake restoration methods (RECCE, Delocate, DFI), this scheme achieves the best FID score and identity similarity in different types of deep fake face restoration, and can restore more realistic target face images. It can be seen that this scheme can restore the source face while restoring the target face, and the effect is better than the existing algorithm. Table 1 shows the target face restoration effect on various types of deep fake data.

[0076] Table 1

[0077]

[0078] Embodiment 3

[0079] A face image restoration system based on a mask autoencoder, the system comprising: an image cropping module, a feature map splicing module, a source face restoration module, an image mask module and an autoencoder restoration module;

[0080] The image cropping module is used to collect fake face images, crop the fake face images using the RetinaFace model, and obtain the source face image and target face image corresponding to the fake face image. RetinaFace is used to locate and align the fake face images in the dataset, and further crop a 224×224 face image based on the coordinates of the facial feature points. Similarly, the corresponding source face and target face images are obtained.

[0081] The feature map splicing module is used to extract the texture features and high-level semantic features of the source face image, upsample the high-level semantic features and splice them with the texture features, and obtain the segmentation probability map based on the spliced ​​feature map. For the obtained cropped forged face image, it is input into the identity segmentation module to calculate the segmentation probability map; the pre-trained EfficientNet-B0 is used to extract features, and the first layer features are used as the underlying texture features, and the second to last layer features are used as high-level semantic features. The high-level semantic features are upsampled by 4 transposed convolutions, spliced ​​with the underlying texture features, and the segmentation map is generated again by the transposed convolution layer, and the Sigmoid function is used to constrain it to the range of 0-1.

[0082] The source face restoration module is used to obtain source face information by multiplying the segmentation probability map with the source face image, and restore and reconstruct the source face based on the source face information to obtain a reconstructed source face. The source face restoration module includes: a semantic feature extraction submodule, a channel convolution submodule, a channel restoration submodule and an image reconstruction submodule;

[0083] The semantic feature extraction submodule is used to obtain source face information based on the multiplication of the segmentation probability map and the source face image, downsample the source face information through 4 convolution-pooling layers, and extract the feature map of high-level semantic features; the channel convolution submodule is used to divide the feature map of high-level semantic features into source face features and potential target face features, upsample the source face features using transposed convolution, and convolve the potential target face features using a convolution layer to obtain a feature map with a preset number of channels; the channel restoration submodule is used to splice the feature map of high-level semantic features with the feature map with the preset number of channels using residual connection, and restore the number of channels of the spliced ​​feature map using convolution; the image reconstruction submodule is used to repeat the feature map with the restored number of channels 4 times S31-S33, restore the source face image, calculate the identity loss and perceptual loss based on the source face image and the source face image, and obtain the reconstructed source face.

[0084] The obtained segmentation probability map is multiplied by the forged face image to obtain the source face information, and the source face information is input into the source face recovery module to restore and reconstruct the source face. The source face information is input into the source face recovery module, and down-sampling operations are performed through 4 convolution-pooling layers to extract high-level semantic features of the image; the obtained high-level semantic features are divided into two parts: source face features and potential target face features, and the source face features are up-sampled using the transposed convolution layer; the potential target face features are processed using the convolution layer so that the number of channels is consistent with the number of channels in the target face recovery module. In this embodiment, we set the number of channels to 98. The feature map calculated by the downsampling layer operation is spliced ​​with the corresponding features of the upsampling stage through residual connection, and the number of channels is restored by convolution operation; the source face image is restored through 4 transposed convolution-residual connection-convolution operations, and then the identity loss and perceptual loss are calculated with the original source face image to optimize the source face recovery module.

[0085] The image mask module is used to divide the segmentation probability map into blocks and calculate the mean of each segmentation probability map, assign weights based on the mean results and calculate the weights, perform masking based on the weights, and obtain a masked image. The obtained segmentation probability map is input into the mask generation module, the mask block number is calculated, and the forged face is divided into blocks, and the corresponding image blocks are masked according to the mask block number; the segmentation probability map is divided into blocks, and the mean of the segmentation probability map is calculated. If the mean is less than 0.5, the weight of each block is set to a random value between 0 and 1; if the mean is greater than or equal to 0.5, the mean of each block is calculated as the weight; the calculated weights are sorted from large to small, and the mask ratio is defined as β, and the image blocks with larger weights are masked. In this embodiment, β is set to 0.5.

[0086] The autoencoder restoration module is used to input the potential target facial features in the reconstructed source face and the mask image into the mask autoencoder to obtain a facial image restoration image of the forged facial image.

[0087] The image blocks are divided into blocks, and the image blocks are masked (pixels are set to 0) according to the mask sequence number; in this embodiment, we set the block size to 16, each image is divided into 196 image blocks, and 98 image blocks are retained after masking;

[0088] The unmasked image block is input into the masked autoencoder to restore the content of the image block corresponding to the target face image.

[0089] Flatten the uncovered image blocks and merge them into a tensor; input it into the fully connected layer to obtain the corresponding image feature vector;

[0090] The image feature vector is input into the Transformer module (such as Figure 2As shown in Figure 2, the global features of the image are calculated through a multi-head self-attention mechanism; the final global features are obtained by repeating the Transformer module calculation L times. The calculation process is as follows:

[0091] f m =F(W·x sv +b)+p sv );

[0092] Among them, F represents the encoding network composed of L Transformer modules, x sv represents the unmasked image tensor, W and b represent the weight and bias of the fully connected layer respectively, and p sv Represents x sv The corresponding position code. In this embodiment, L is set to 7.

[0093] Using the mask block sequence number, the mask block of the global feature is filled to make it correspond to the complete image block sequence;

[0094] The padded global features are input into the decoder (composed of n Transformer modules), and the image block content is restored through multi-head attention mechanism, normalization layer and other operations, and the identity loss, perception loss, attribute loss and pixel restoration loss are calculated with the original target face image. The process of restoring the image content is as follows:

[0095] x rec =Sigmoid(D(f m ||f t ));

[0096] Where D represents the decoder consisting of n Transformer modules, f m represents the features obtained by encoding the uncovered image, f t represents the potential target face features extracted in the source face restoration module, and the Sigmoid function is used to constrain the output to the range of 0 to 1. In this embodiment, n is set to 3.

[0097] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A face image restoration method based on mask autoencoder, characterized in that: The method comprises: Step S1, collecting a forged face image, cropping the forged face image using a RetinaFace model, and obtaining a source face image and a target face image corresponding to the forged face image; Step S2, extracting the background texture features and high-level semantic features of the source face image, upsampling the high-level semantic features and splicing them with the background texture features, and obtaining a segmentation probability map based on the spliced ​​feature map; Step S3, multiplying the segmentation probability map by the source face image to obtain source face information, and restoring and reconstructing the source face based on the source face information to obtain a reconstructed source face; Step S4, dividing the segmentation probability map into blocks and calculating the mean of each segmentation probability map, assigning weights based on the mean results and calculating the weights, performing masking based on the weights, and obtaining a masked image; Step S5: input the potential target facial features in the reconstructed source face and the mask image into the mask autoencoder to obtain a facial image restoration image of the forged facial image.

2. The face image restoration method based on masked autoencoder according to claim 1, characterized in that: In step S2, the process of extracting the shading features and high-level semantic features of the source face image, upsampling the high-level semantic features and splicing them with the shading features, and obtaining the segmentation probability map based on the spliced ​​feature map specifically includes: Use EfficientNet-B0 to extract low-level texture features and high-level semantic features; The high-level semantic features are upsampled using transposed convolution, the upsampled high-level semantic features are spliced ​​with the underlying texture features, and the spliced ​​features are input into a transposed convolution layer to generate a segmentation probability map.

3. The face image restoration method based on masked autoencoder according to claim 1, characterized in that: In step S3, the source face information is obtained by multiplying the segmentation probability map with the source face image, and the source face is restored and reconstructed based on the source face information to obtain the reconstructed source face. Specifically, the process includes: S31, based on the source face information obtained by multiplying the segmentation probability map with the source face image, down-sampling the source face information through 4 convolution-pooling layers to extract a feature map of high-level semantic features; S32, dividing the feature map of the high-level semantic features into source face features and potential target face features, upsampling the source face features using transposed convolution, and convolving the potential target face features using a convolution layer to obtain a feature map with a preset number of channels; S33, using residual connection to splice the feature map of the high-level semantic feature with the feature map of the preset number of channels, and using convolution to restore the number of channels of the spliced ​​feature map; S34. Repeat S31-S33 4 times to restore the feature map of the number of channels, restore the source face image, calculate the identity loss and the perception loss based on the source face image and the source face image in step S1, and obtain the reconstructed source face.

4. The face image restoration method based on masked autoencoder according to claim 3, characterized in that: In step S4, the segmentation probability map is divided into blocks and the mean of each segmentation probability map is calculated, weights are assigned and calculated based on the mean results, and masking is performed based on the weights to obtain a covered image, which specifically includes: The mask block number is calculated based on the segmentation probability map, the forged face is divided into blocks based on the mask block number, and the forged face after the division is covered by the mask block number to obtain a covered image.

5. The face image restoration method based on masked autoencoder according to claim 4, characterized in that: In step S5, inputting the potential target face features in the reconstructed source face and the mask image into the mask autoencoder to obtain the target face image restoration image of the forged face image specifically includes: Inputting the potential target face features in the reconstructed source face and the uncovered image blocks in the masked image into a masked autoencoder for flattening and merging into a tensor, inputting the tensor into a fully connected layer, and obtaining a tensor image feature vector; Input the tensor image feature vector into the Transformer module, and calculate the global features of the image through a multi-head attention mechanism; The mask block filling is performed on the global features of the image, and the filled global features of the image are input into a decoder to perform an operation of restoring the image block content, thereby obtaining a facial image restoration image of the forged facial image.

6. A face image restoration system based on a masked autoencoder, the system being used to implement the image restoration method according to any one of claims 1 to 4, characterized in that: The system includes: Image cropping module, feature map splicing module, source face restoration module, image mask module and autoencoder restoration module; The image cropping module is used to collect a forged face image, crop the forged face image using a RetinaFace model, and obtain a source face image and a target face image corresponding to the forged face image; The feature map splicing module is used to extract the background features and high-level semantic features of the source face image, upsample the high-level semantic features and splice them with the background features, and obtain a segmentation probability map based on the spliced ​​feature map; The source face restoration module is used to obtain source face information by multiplying the segmentation probability map with the source face image, and to restore and reconstruct the source face based on the source face information to obtain a reconstructed source face; The image mask module is used to divide the segmentation probability map into blocks and calculate the mean of each segmentation probability map, assign weights based on the mean results and calculate the weights, and perform masking based on the weights to obtain a masked image; The autoencoder restoration module is used to input the target facial features in the reconstructed source face and the mask image into the mask autoencoder to obtain a facial image restoration image of the forged facial image.

7. The face image restoration system based on masked autoencoder according to claim 6, characterized in that: The workflow of the feature map splicing module specifically includes: Use EfficientNet-B0 to extract low-level texture features and high-level semantic features; The high-level semantic features are upsampled using transposed convolution, the upsampled high-level semantic features are spliced ​​with the underlying texture features, and the spliced ​​features are input into a transposed convolution layer to generate a segmentation probability map.

8. The face image restoration system based on masked autoencoder according to claim 6, characterized in that: The source face restoration module includes: a semantic feature extraction submodule, a channel convolution submodule, a channel restoration submodule and an image reconstruction submodule; The semantic feature extraction submodule is used to obtain source face information based on the multiplication of the segmentation probability map and the source face image, downsample the source face information through four convolution-pooling layers, and extract a feature map of high-level semantic features; The channel convolution submodule is used to divide the feature map of the semantic feature into source face features and potential target face features, upsample the source face features using transposed convolution, and convolve the potential target face features using a convolution layer to obtain a feature map with a preset number of channels; The channel restoration submodule is used to splice the feature map of high-level semantic features with the feature map of the preset number of channels using residual connection, and restore the number of channels of the spliced ​​feature map using convolution; The image reconstruction submodule is used to repeat the feature map of the restored channel number 4 times S31-S33, restore the source face image, calculate the identity loss and the perceptual loss based on the source face image and the source face image in step S1, and obtain the reconstructed source face.

9. The face image restoration system based on mask autoencoder according to claim 6, characterized in that: The workflow of the image mask module specifically includes: calculating the mask block number based on the segmentation probability map, dividing the forged face into blocks based on the mask block number, and using the mask block number to cover the forged face after the block division to obtain a covered image.

10. The face image restoration system based on mask autoencoder according to claim 6, characterized in that: The workflow of the autoencoder recovery module specifically includes: Inputting the potential target face features in the reconstructed source face and the uncovered image blocks in the masked image into a masked autoencoder for flattening and merging into a tensor, inputting the tensor into a fully connected layer, and obtaining a tensor image feature vector; Input the tensor image feature vector into the Transformer module, and calculate the global features of the image through a multi-head attention mechanism; The mask block filling is performed on the global features of the image, and the filled global features of the image are input into a decoder to perform an operation of restoring the image block content, thereby obtaining a facial image restoration image of the forged facial image.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and readable storage medium

    CN111340188A

  • Real-time small face detection method based on improved YOLOv5

    CN116092154A

  • Image segmentation model training method and device, electronic equipment and storage medium

    CN117744733A

  • Self-supervision pre-training method in distribution network line self-adaptive inspection based on unmanned aerial vehicle

    CN118587563A

Cited By

  • Pipeline deformation visual measurement method and system based on detection robot

    CN121304938A