Living body face attribute editing method and system based on mask supervision target feature decoupling
Through the target feature decoupling method of mask supervision, combined with attention module and logistic regression classifier, the precise control problem of face attribute editing in the existing technology is solved, and high-precision target attribute editing and image quality improvement are achieved.
Patent Information
- Application Number
- CN202510337063.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
The existing face attribute editing technology is difficult to accurately control the editing of a single attribute, and there is entanglement between different face attributes, resulting in insufficient detail accuracy and realism of the generated image.
The target feature decoupling method based on mask supervision is adopted, and the latent vectors are decoupled through the target attribute and non-target attribute attention module, combined with the latent vector self-fusion and reconstruction of the input face training module, and the logistic regression classifier and focus loss optimization model are used to achieve high-precision editing of the target attributes.
High-precision editing of the target attribute area is achieved, the quality of generated images and the ability to control attribute changes is improved, interference between attributes is avoided, and the accuracy and controllability of facial attribute editing is improved.
Smart Images

Figure CN120375438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face attribute generation and confrontation, and particularly relates to a live face attribute editing method and system based on mask supervision and target feature decoupling. Background Art
[0002] The face attribute editing technology constructed by using the generative adversarial network can generate diverse feature images based on real face data and change the face attributes according to actual needs, such as simulating face aging, expressions in different scenarios, hair style changes, etc. In the training of the face liveness detection model, on the one hand, this technology can provide diverse facial data, enhance the robustness and generalization ability of the model in complex scenarios, and avoid identity misjudgment; on the other hand, it can greatly reduce the workload of manually collecting data, protect user privacy, and avoid data leakage. By exploring the matching relationship between latent vectors of different resolutions and different face attributes, researchers can simulate the situations of faces in different states, which helps to improve the generalization of the model, understand its working mechanism, and further analyze the impact of facial editing on the recognition result.
[0003] The existing face editing technology is represented by StyleGAN. Thanks to the rich semantic information contained in its latent space, high-quality semantic editing can be achieved. Traditional methods usually assume that the latent space has decoupling properties and achieve attribute editing by adjusting specific directions, but parameters need to be manually adjusted. Currently, most research focuses on the exploration and editing of the latent space, such as using direction vectors to represent specific attribute changes (such as smiling, hair color); in addition, aiming at the mutual interference between attributes, research shows that by optimizing the latent space structure (such as Style Space) or predicting the offset of the latent space, more accurate and independent attribute editing can be achieved.
[0004] However, there are still some deficiencies in current face attribute editing. Since the dimension of the latent vector and its control over different resolutions are not completely independent and clearly divided, it is difficult to precisely control the attributes of each dimension; at the same time, there are entanglement restrictions between different face attributes, which limits the precise editing of a single attribute. In addition, when the original StyleGAN controls image generation, it often operates on the entire image, and the definition of the local area of the target attribute is blurred, making it difficult to precisely control the local detail changes of the RGB image. Therefore, there is an urgent need for a face attribute editing technology that can focus on efficiently decoupling the latent space, precisely controlling the editing of a single attribute, avoiding interference between attributes, improving the detail accuracy and realism of the generated image, especially the balance between global and local area editing. Summary of the Invention
[0005] To overcome the defects and deficiencies of the existing technologies, the present invention provides a live face attribute editing method and system based on mask-supervised target feature decoupling. The present invention, according to the publicly available face dataset and the corresponding face attribute mask dataset, divides the training set and the test set according to a ratio; uses the target attribute attention module and the non-target attribute attention module to decouple the target attribute and the irrelevant attributes in the latent space based on the attention mechanism, and optimizes the latent vector embedding; uses the target attribute region mask supervision module to use the local mask supervision of the target attribute to enhance the ability of the attention network to optimize the target attribute region intensively and obtain a more compact feature representation ability; uses the latent vector self-fusion module to obtain the latent vector after feature decoupling optimization; uses the training module based on reconstructing the input face to minimize the overall loss function to obtain the best model network and parameter weights, and trains a logistic regression binary classifier with focal loss supervision to obtain the direction vectors of common attributes; uses the face attribute editing and evaluation module to interpolate the latent vector and the attribute direction vectors to edit different attributes of the face, and makes a comparison with the existing networks through qualitative and quantitative results.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a live face attribute editing method based on mask-supervised target feature decoupling, including the following steps:
[0008] Obtain a face dataset and a mask dataset corresponding to the face attributes;
[0009] Combine the attribute segmentation masks provided by the mask dataset into a complete mask image;
[0010] The face data is encoded by the StyleGAN style encoder to obtain a latent vector, and the latent vector generates an attention-enhanced latent vector w based on the multi-channel attention mechanism t and the latent vector w b ;
[0011] The latent vector w t and the latent vector w b are passed through the local generator to generate the corresponding mask depth prediction map, and the attention network is optimized based on the attribute mask labels of the mask dataset;
[0012] The latent vector w t and the latent vector w b are spliced and fused into a latent vector w * ;
[0013] Train the overall network based on the reconstructed input face, minimize the total loss function of the model, obtain the direction vectors of different attributes according to different attribute category labels and the logistic regression classification model, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model;
[0014] Based on the latent vector w * Interpolate with the direction vectors of different attributes to obtain the final latent vector w target , and use it as the input of the pre-trained StyleGAN generator. After the step-by-step resolution output of the StyleGAN generator, obtain the generated face after editing the target attribute.
[0015] As a preferred technical solution, the latent vector generates the attention-enhanced latent vector w t and the latent vector w b based on the multi-channel attention mechanism, specifically including:
[0016] Construct a target attribute attention module and a non-target attribute attention module to perform self-attention reinforcement on the latent vectors of different resolutions. The target attribute attention module uses a 1×1 convolutional layer, and the non-target attribute attention module uses a 3×3 convolutional layer;
[0017] Generate a query matrix, a key matrix, and a value matrix. After the steps of feature concatenation of the query matrix and the key matrix, vector size adjustment, and Softmax normalization, obtain the weight matrix, and multiply the weight matrix by the value matrix to generate the attention-enhanced latent vector w t and the latent vector w b , specifically expressed as:
[0018]
[0019] w i =Softmax(Reshape([Q i ,K i W conv )]V i
[0020] where i has two values. When i takes the value of t, it represents the features related to the target attribute attention module. When i takes the value of b, it represents the features related to the non-target attribute attention module. V i ,K i ,Q i respectively represent the value matrix, key matrix, and query matrix corresponding to the above two modules, are the weight coefficients corresponding to the value matrix, key matrix, and query matrix respectively, are the feature mapping networks corresponding to the value matrix, key matrix, and query matrix respectively, w L , w M, w H respectively correspond to the features of the input latent vector w at low, medium, and high resolutions. [·] represents the operation of feature concatenation, and W conv represents the parameters of the convolutional layer, and Reshape represents adjusting the feature size to be consistent with the latent vector w.
[0021] As a preferred technical solution, the backbone network of the local generator is a ResNet model, using IdentityBlock and Conv Block as residual connections. The convolutional layer has two types of convolutional kernels, 1×1 and 3×3. The 1×1 convolutional kernel changes the number of output channels, and the 3×3 convolution changes the size of the output feature map. BatchNorm is used as the batch normalization operation and ReLU is used as the activation function. The input dimension and output dimension of the Identity Block are the same, which is used to deepen the network. The input dimension and output dimension of the Conv Block are different, which is used to change the feature dimension of the network.
[0022] As a preferred technical solution, the latent vector w t and the latent vector w b are concatenated and fused into the latent vector w * , specifically including:
[0023] Based on the vector fusion method of mean and variance statistical features, the latent vector w t and the latent vector w b are concatenated and fused into the latent vector w * , expressed as:
[0024] μ(w mix ) = αμ(w t ) + (1 - α)μ(w b ) + n μ
[0025] σ 2 (w mix ) = βσ 2 (w t ) + (1 - β)σ 2 (w b ) + n σ
[0026] w * = MLP[μ(w mix ), σ 2 (w mix )]
[0027] where, n μ and n σ represent random noise, sampled from the difference space between the latent vector w t and the latent vector w b nμ obeys a Gaussian distribution with a mean of |μ(w t ) - μ(w b )| and a variance of 1, and n σ obeys a Gaussian distribution with a mean of |σ 2 (w t ) - σ 2 (w b )| and a variance of 1. [·] represents the operation of feature concatenation, MLP represents a multi - layer perceptron network, α is the weight coefficient of the mean feature, β is the weight coefficient of the variance feature, w mix is the latent vector after feature fusion, and w * is the finally output latent vector.
[0028] As a preferred technical solution, the total loss function is expressed as:
[0029] L = λ1L latent + λ2L mask + λ3L att ++ λ4L id
[0030] L latent = 1 - CosSim(G(w * ), G(w))
[0031]
[0032] where L latent represents the latent vector consistency loss, L mask represents the attribute mask loss, L att represents the attention regularization loss, L id represents the identity perception loss, λ1, λ2, λ3, and λ4 represent the corresponding weight coefficients, G represents the StyleGAN generator, N represents the number of images in a batch, g(·) represents the mask depth prediction map, M t (·) represents the true attribute mask label, M(x i ) represents the global mask label of the face, α and β represent the weight magnitudes of the target and non - target attributes, H and W are the height and width of the image respectively, m ij represents the pixel value size of the attention map at coordinates (i, j), I input represents the input face, I rec represents the reconstructed face, and respectively represent the output features of I input and I rec after the l - th layer of ArcFace, ijk represents the result of the output activation of the i - th convolutional kernel at the (j, k) position after passing through the intermediate layer, C l 、Hl 、W l is the number of channels, height, and width of the feature map of the l-th layer.
[0033] As a preferred technical solution, direction vectors of different attributes are obtained according to different attribute category labels and a logistic regression classification model, expressed as:
[0034] e i = argmax w P(y = w·x + b), y ∈ {0, 1}
[0035] where e i represents the direction vector of different attributes i;
[0036] The binary classifier model is a logistic regression model, represented by the following conditional probability distribution:
[0037]
[0038] where x represents the input, each sample contains n features, Y ∈ {0, 1} represents the output; w represents the weight vector, and b represents the bias.
[0039] As a preferred technical solution, the focal loss is expressed as:
[0040]
[0041] where α is the balance factor, (1 - p) γ is the modulation factor, and γ is the focusing parameter.
[0042] As a preferred technical solution, based on the interpolation method between the latent vector w * and the direction vectors of different attributes, the final latent vector w target is obtained, specifically expressed as:
[0043] w target = w * + λe i
[0044] Through the hierarchical output of the generator, the generated face after editing the target attribute is obtained:
[0045] I edit = G style (w target )
[0046] where e i represents the direction vector of different attributes i, λ is the interpolation coefficient, G style is the generator of StyleGAN, and I edit is the image with the target attribute edited.
[0047] As a preferred technical solution, the pre-trained StyleGAN generator adopts a multi-resolution step-by-step generation process in multiple stages, which is jointly controlled by the final latent vector w target and the adaptive instance normalization AdaIN layer. Each stage includes an upsampling step and a convolutional step based on a 3×3 convolutional kernel. The AdaIN layer is expressed as:
[0048]
[0049] where x represents the content feature input, which is the feature map from the previous layer in the generation network, y represents the style feature input, and μ(y) and σ(y) are the scaling coefficient and bias coefficient obtained by w target through the radiometric transformation;
[0050] Random noise obeying the Gaussian distribution is added to each channel before the AdaIN layer.
[0051] The present invention provides a live face attribute editing system based on mask supervised target feature decoupling, including: a dataset acquisition module, a dataset preprocessing module, a StyleGAN style encoder, a target attribute attention module, a non-target attribute attention module, a target attribute region mask supervision module, a latent vector self-fusion module, a reconstructed input face training module, a StyleGAN generator, and a face attribute editing picture output module;
[0052] The dataset acquisition module is used to acquire a face dataset and a mask dataset corresponding to the face attributes;
[0053] The dataset preprocessing module is used to combine the attribute segmentation masks provided by the mask dataset into a complete mask image;
[0054] The StyleGAN style encoder is used to convert face data into a latent vector;
[0055] The target attribute attention module is used to output an attention-enhanced latent vector w t ;
[0056] The non-target attribute attention module is used to output an attention-enhanced latent vector w b ;
[0057] The target attribute region mask supervision module is used to generate corresponding mask depth prediction maps for the latent vectors w t and w b through a local generator, and optimize the attention network based on the attribute mask labels of the mask dataset;
[0058] The latent vector self-fusion module is used to fuse the latent vector wt and the latent vector w b are concatenated and fused into the latent vector w * ;
[0059] The reconstructed input face training module is used to train the overall network based on the reconstructed input face, minimize the total model loss function, obtain the direction vectors of different attributes according to different attribute category labels and the logistic regression classification model, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model;
[0060] The StyleGAN generator is used to reconstruct the input face;
[0061] The face attribute editing picture output module is used to obtain the final latent vector w * by interpolating the latent vector w target with the direction vectors of different attributes, and use it as the input of the pre-trained StyleGAN generator. Through the hierarchical output of the StyleGAN generator, the generated face after editing the target attribute is obtained.
[0062] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0063] (1) The present invention combines the StyleGAN backbone network and the attention mechanism to construct the target attribute attention module and the non-target attribute attention module, and trains each attribute independently, which can achieve the precise decoupling of the feature of the target attribute area, avoid the unexpected changes to other attributes when modifying a certain attribute in the latent space, and thus improve the accuracy and controllability of attribute editing.
[0064] (2) The present invention makes full use of the face dataset and the face mask dataset. By introducing the target attribute mask label as the supervision signal of the attention network, the attention network focuses on optimizing the specified target attribute and focuses on the compact area related to the target attribute rather than the whole image, significantly improving the optimization effect for the target attribute area, realizing the high-precision editing of the target attribute area, and avoiding the interference problem caused by global optimization.
[0065] (3) The present invention combines the logistic regression classifier and the focal loss to train the corresponding attribute direction vectors for common face attributes (such as hair color, eye size, lip change, bang style, etc.). On this basis, through the interpolation operation of the latent vector and the attribute direction vector, the fine-grained editing of the target attribute can be realized. The present invention not only improves the quality of the generated image, but also has the precise control ability for attribute changes, providing an efficient solution for complex face attribute editing tasks. Description of the Drawings
[0066] Figure 1 Schematic flowchart of the in - vivo face attribute editing method based on mask - supervised target feature decoupling of the present invention;
[0067] Figure 2 Schematic diagram of the overall architecture of the in - vivo face attribute editing system based on mask - supervised target feature decoupling of the present invention;
[0068] Figure 3(a) is a schematic diagram of the network architecture of the target attribute attention module of the present invention;
[0069] Figure 3(b) is a schematic diagram of the network architecture of the non - target attribute attention module of the present invention;
[0070] Figure 4 Schematic diagram of the network architecture of the latent vector self - fusion module of the present invention;
[0071] Figure 5 Schematic diagram of the face attribute editing process in the training and testing phases of the present invention;
[0072] Figure 6 Schematic diagram of the face attribute editing result in the testing phase of the present invention. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0074] Embodiment 1
[0075] In this embodiment, the celebrity face image dataset CelebA - HQ and its corresponding facial mask dataset CelebAMask - HQ are used for training and testing, and the specific implementation process of the present invention is introduced in detail. The CelebA - HQ dataset contains 30,000 high - resolution celebrity face images. These images have high resolution, rich texture details, and detailed annotations of face key points and attribute features (such as facial features, skin color, hair color, pose, etc.), and are suitable for a variety of face processing tasks. In this embodiment, an image resolution of 512×512 pixels is used; at the same time, the CelebAMask - HQ database provides corresponding 512×512 - pixel manual segmentation masks for each image in CelebA - HQ. These masks accurately annotate facial components and accessories including skin, nose, eyes, etc. The provided mask data reaches the pixel level, presenting the boundaries and details of each part of the face components exquisitely. In this embodiment, for the four common attributes: bangs, eyes, lips, and hair color are used as the regions for attribute editing.
[0076] This embodiment is deployed on the Linux system, using Pytorch 1.9.0 as the deep learning framework, with the CUDA version being 12.0 and the graphics card driver version being 525.89.02. In terms of hardware configuration, two Tesla V100 graphics cards are used, and the container environment is Conda. The main dependent libraries include scikit-image, tensorboard, matplotlib, pillow, tqdm, etc.;
[0077] As Figure 1 shown, this embodiment provides a method for editing living face attributes based on mask-supervised target feature decoupling, including the following steps:
[0078] S1: Data preprocessing. According to the publicly available face dataset and the corresponding face attribute mask dataset, the attribute segmentation masks provided by the mask dataset are combined into a complete mask image, which is used for the target attribute region mask supervision module and the non-target attribute mask supervision module to calculate the mask loss function and supervise the generation of the attention map;
[0079] And the training set and the test set are divided according to a certain proportion. In this embodiment, the pictures numbered 0 - 27999 are used as the training set, while the pictures numbered 28000 - 29999 are used as the validation set. Since the 19 attribute label images contained in the mask dataset are single-channel, in this embodiment, the pixel values of these single-channel images are mapped to a set of predefined RGB color lists to be converted into color images. Subsequently, all attribute images are stacked in order onto one image to generate a complete face mask segmentation image, which is saved to the specified mask folder. To more efficiently enable the model to load image data and reduce a large amount of memory occupancy, in this embodiment, the image dataset and the mask dataset are adjusted and cropped to a size of 512×512 pixels according to the principle of one-to-one matching, and multi-process acceleration is used. Finally, the RGB image data is converted into LMDB format data, and each image is stored in the form of key-value pairs, where the key is the label of the image and the value is the byte data of the image;
[0080] S2: Construct the target attribute attention and non-target attribute attention modules, and decouple the latent vectors of different resolutions into target attribute embedding vectors and non-target attribute embedding vectors based on the multi-channel attention mechanism;
[0081] In this embodiment, the target attribute attention and non-target attribute attention modules are used to perform self-attention enhancement on the low, medium, and high resolutions of the latent vector, realizing the decoupling of the latent vector features. For the input latent vector w, the target attribute attention module uses a 1×1 convolutional layer with a smaller receptive field to strengthen the model's attention to local features, and the non-target attribute attention module uses a 3×3 convolutional layer with a larger receptive field to perceive global features, generating query matrices, key matrices, and value matrices. After that, through steps such as feature concatenation of the query matrix and the key matrix, vector size adjustment, and Softmax normalization, a weight matrix is obtained. Finally, the weight matrix is multiplied by the value matrix to generate the attention-enhanced latent vector w t and w b ;
[0082] The specific process is as follows: The model inputs face data with a batch size of N. First, it passes through the StyleGAN style encoder. Through the non-linear mapping network, the high-dimensional face data is mapped into a low-dimensional latent vector w in the latent space. w contains high-dimensional semantic features of the face (such as face shape, pose, skin, facial features, etc.), and the dimension is 18×512. This style encoder is constructed based on a convolutional neural network (CNN), and the activation function is the ReLU function. After that, w is respectively input into the target attribute attention and non-target attribute attention modules, which are used to perform self-attention enhancement on the three resolutions of the latent vector (4×4 - 16×16 resolution, 16×16 - 64×64 resolution, 64×64 - 1024×1024 resolution), realizing the decoupling of the latent vector features. Among them, the first dimension to the fourth dimension of the w vector correspond to the features w of the 4×4 - 16×16 resolution L , with a size of 4×512, the fifth dimension to the eighth dimension correspond to the features w of the 16×16 - 64×64 resolution M , with a size of 4×512, and the ninth dimension to the eighteenth dimension correspond to the features w of the 64×64 - 1024×1024 resolution H , with a size of 10×512.
[0083] The target attribute attention module uses a 1×1 convolutional layer with a smaller receptive field to strengthen the model's attention to local features and a 3×3 convolutional layer with a larger receptive field to perceive global features for non-target attributes. For the target attribute attention module, w L 、w M 、w H simultaneously pass through three feature mapping networks. Among them the network is a convolutional network with a convolutional kernel size of 1×1 and 64 convolutional kernels consists of two consecutive convolutional networks and activation layers, with a convolutional kernel size of 1×1, 64 convolutional kernels, a stride of 1, and the activation function is ReLU. These three networks respectively generate the query matrix Q corresponding to the latent vector wt , key matrix K i , value matrix V i . Then, the query matrix Q t and the key matrix V t are concatenated by features, the size of the feature vector is adjusted to 18×512 through a 1×1 convolutional layer, and the Softmax function is used for normalization to enhance the focusing ability on important information, obtaining the weight matrix. Finally, the weight matrix and the value matrix V t are multiplied, and through the feature extraction of the fully connected layer FC, the potential vector w with enhanced attention is output t .
[0084] Similarly, for the non-target attribute attention module, w L , w M , w H simultaneously pass through three feature mapping networks, where the network and have the same network structure , consisting of a convolutional network and an activation layer, where the convolutional kernel size is 3×3, the number of convolutional kernels is 64, the stride is 1, and the rest of the structure is the same as that of the target attribute attention module. Finally, the potential vector w with enhanced attention is generated b . The calculation formula for this process is as follows:
[0085]
[0086] w i = Softmax(Reshape([Q i , K i W conv )]V i
[0087] Among them, i has two values. When i takes the value of t, it represents the features related to the target attribute attention module. When i takes the value of b, it represents the features related to the non-target attribute attention module. V i , K i , Q i respectively represent the value matrix, key matrix, and query matrix corresponding to the above two modules are the weight coefficients corresponding to the value matrix, key matrix, and query matrix respectively, and they are learnable parameters are the feature mapping networks corresponding to the three respectively, w L , w M , w HCorresponding to the features of the input latent vector w at three resolutions (including 4×4 - 16×16 resolution, 16×16 - 64×64 resolution, 64×64 - 1024×1024 resolution), [·] represents the operation of feature concatenation, W conv represents the parameters of the 1×1 convolutional layer, Reshape represents adjusting the feature size to be consistent with w, w i represents the final latent feature output by the model after attention enhancement;
[0088] S3: Construct a target attribute region mask supervision module, use the local mask label of the target attribute in the mask dataset as supervision, combine the attribute mask loss and the attention regularization loss to enhance the ability of the attention network to optimize the target attribute region intensively, and obtain a more compact target attribute feature representation ability;
[0089] In this embodiment, the target attribute region mask supervision module optimizes the attention module based on the attribute mask label of the mask dataset, so that the attention network can keep the other irrelevant regions unchanged while focusing on the region to be edited. The latent vectors w t and w b are respectively input into the local generator to generate the corresponding mask depth prediction maps, and the model parameters are optimized by minimizing the attribute mask loss and the attention regularization loss.
[0090] The specific process is as follows. As shown in Figures 3(a) and 3(b), the attention-enhanced latent vectors w t and w b are respectively input into the local generator to obtain the predicted attention depth maps of the target attribute region and the non-target attribute region and The backbone network of this local generator is a ResNet model, which mainly uses Identity Block and Conv Block as residual connections. The convolutional layer has two types of convolutional kernels, 1×1 and 3×3. Among them, the 1×1 convolutional kernel changes the number of output channels, and the 3×3 convolutional kernel changes the size of the output feature map. BatchNorm is used as the batch normalization operation and ReLU is used as the activation function. The input dimension and output dimension (size, number of channels) of the Identity Block are the same, and multiple Identity Blocks are connected in series, mainly used to deepen the network, while the input dimension and output dimension of the Conv Block are different, mainly used to change the feature dimension of the network.
[0091] To enhance the ability of the attention network to optimize the target attribute region intensively, minimize the attribute mask loss L mask and the attention regularization loss L att to optimize the model parameters. The attribute mask loss L maskBy means of pixel-by-pixel comparison, the predicted depth map of the target attribute region is made as similar as possible to the actual target mask label map. The predicted depth map of the non-target attribute region and the actual target mask are spliced to restore the predicted mask map of the global face, making the global predicted mask map as similar as possible to the actual face mask map, so as to achieve the attention decoupling of the target region and the non-target region. The specific calculation formula of this process is as follows:
[0092]
[0093] Among them, N represents the number of images in a batch, g(·) represents the mask depth prediction map, and M t (·) represents the true attribute mask label, and M(x i ) represents the global mask label of the face. α and β represent the weight sizes of the target and non-target attributes;
[0094] To further constrain the predicted depth map of the target attribute region to be concentrated in a smaller range, the mask attention regularization loss L att is adopted. This loss is constructed based on the L2 loss, minimizing the global pixel value size of the predicted depth map, encouraging the attention network to focus on more compact regions related to the target attribute rather than the entire region of the image. The specific calculation formula of this loss function is as follows:
[0095]
[0096] Among them, H and W are the height and width of the image respectively, and m ij represents the pixel value size of the attention map at coordinates (i, j), and the value range is [0, 1];
[0097] S4. Construct a latent vector self-fusion module to fuse the target attribute embedding vector and the non-target attribute embedding vector after feature enhancement by the target attribute and non-target attribute modules, and optimize in combination with the latent vector consistency loss to obtain the latent vector after feature decoupling optimization;
[0098] In this embodiment, the latent vector self-fusion module is used to re-splice and fuse the latent vectors w t and w b into w * . The encoding of the target attribute in w * is more explicit, with better feature dimension decoupling effect, and can better capture the mutual relationship between features, providing a generation basis for subsequent face attribute editing.
[0099] The specific process is as follows. As Figure 4 shown, the latent vector w t output by the target attribute attention module and the latent vector wb , as the two inputs of the potential vector self-fusion module, the method of vector fusion is based on the method of mean and variance statistical features. First, the mean and variance are calculated, and then the mean and variance vectors of all channels are concatenated. Assuming to calculate the mean and variance of the feature vector x, the specific calculation formulas are as follows:
[0100]
[0101] μ(x) = [μ1, μ2, …, μ C
[0102]
[0103]
[0104] where x is the feature vector input to the model, N is the batch size of the input images of the model, H and W are the height and width of the feature vector x respectively, both with a value of 64, C represents the number of channels of the feature vector x, with a value of 128, μ c represents the mean of the c-th channel, and σ c represents the standard deviation of the c-th channel, c ∈ {1, 2, 3, …, C}, and [·] represents feature concatenation along the channel dimension.
[0105] Next, calculate the means and variances of the potential vectors w t and w b respectively according to the above formulas. μ(·) and σ 2 (·) respectively represent the operations of calculating the mean and variance of the feature vector along the channel dimension. After that, the means and variances of w t and w b are fused by weighted summation respectively. α is the weight coefficient of the mean feature, and β is the weight coefficient of the variance feature. They are both hyperparameters. In this embodiment, the value of α is 0.8, and the value of β is 0.8. After that, the fused potential vector is passed through a multi-layer perceptron network, so that the feature fusion vector w * has the same dimension as the original potential vector w. Finally, the potential vector self-fusion module outputs w * , with a size of 18 × 512. The specific calculation formula for this process is as follows:
[0106] μ(w mix ) = αμ(w t ) + (1 - α)μ(w b ) + n μ
[0107] σ 2 (w mix ) = βσ 2 (w t )+(1-β)σ 2 (w b )+n σ
[0108] w * =MLP[μ(w mix ),σ 2 (w mix )]
[0109] where n μ and n σ are random noises sampled from the difference space between w t and w b . n μ follows a Gaussian distribution with mean |μ(w t ) - μ(w b )| and variance 1, and n σ follows a Gaussian distribution with mean |σ 2 (w t ) - σ 2 (w b )| and variance 1. [·] represents the operation of feature concatenation, and MLP represents a multi-layer perceptron network, which is constructed by a 3-layer fully connected neural network with the ReLU function as the activation function. w mix is the latent vector after feature fusion, and w * is the latent vector finally output by the self-fusion module. Further, in order to ensure that there is no significant facial distortion when reconstructing the input face from w * and the initial w through the pre-trained StyleGAN generator, the latent vector consistency loss is minimized, the visual attributes of the input face image are explicitly retained, and the consistency of the face identity is maintained.
[0110] S5. Construct a training module based on the reconstructed input face, train the overall network based on the reconstructed input face, minimize the total loss function of the model, obtain the direction vectors of different attributes according to the different attribute category labels and the logistic regression classification model in the CelebA-HQ dataset, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model;
[0111] In this embodiment, the training module based on the reconstructed input face mainly includes two steps: reconstructing the input and obtaining the direction vectors of attribute features. Reconstructing the input is used to obtain the optimal network structure and parameter weights, and obtaining the direction vectors of attribute features is used for subsequent face attribute editing. The attribute features are four common face attribute features: hair, eyes, lips, and bangs.
[0112] The specific training process is as Figure 5As shown in ①, the model parameters are initialized using the Xavier parameter initialization method. Based on the optimization algorithm of mini-batch gradient descent, the batch size is set to 8, the training epoch is 200000, and the model parameters are saved every 10000 epochs. num_workers is set to 8, that is, 8 independent subprocesses are enabled to load data in parallel. Adam is used as the model optimizer, the initial learning rate lr is set to 0.02, the weight decay parameter is 0.002, and the value of the momentum parameter momentum is set to 0.8.
[0113] Step 1 is based on the input face image I input Reconstruct the input image I rec , focusing on optimizing the feature extraction of the face local area by the target attribute attention module, non-target attribute module, region mask supervision module, etc., to obtain the latent vector w after feature decoupling * , the main loss functions optimized in this step are the latent vector consistency loss L latent , the attribute mask loss L mask , the attention regularization loss L att , the identity perception loss L id , the overall loss function of Step 1 is expressed as the following equation:
[0114] L = λ1L latent + λ2L mask + λ3L att ++ λ4L id
[0115] L latent = 1 - CosSim(G(w * ), G(w))
[0116]
[0117]
[0118] where L latent is used to explicitly retain the visual characteristics of the input image, and to ensure that the face generated based on w * will not produce large facial distortions. λ1 is its corresponding weight coefficient, G represents the StyleGAN generator, L mask is used for the supervision of the target attribute area, and λ2 is its corresponding weight coefficient. L att is used to keep the attention area within a compact range, and λ3 is its corresponding weight coefficient. In L id , I input represents the input face, and I rec is the reconstructed face. It is the pre-trained face recognition model ArcFace. l represents the l-th layer of the ArcFace network. and respectively represent I input and I rec the output features after passing through the l-th layer of ArcFace. ijk represents the result of the output activation of the i-th convolutional kernel at the (j,k) position after passing through the intermediate layer. C l 、H l 、W l are the number of channels, height, and width of the feature map of the l-th layer. By reducing the cosine similarity between the original image and the edited image in the feature space of the pre-trained ArcFace face recognition network, and by optimizing the angular distribution of the feature vectors, the distances between faces belonging to the same identity are minimized as much as possible within the feature space, while the distances between faces belonging to different identities are maximized as much as possible, maintaining the consistency of the input face identity information. L id pays more attention to the face identity features and is closer to the visual perception of humans themselves. λ4 is its corresponding weight coefficient. In this embodiment, the values of λ1, λ2, λ3, and λ4 are 1, 0.8, 0.8, and 1 respectively;
[0119] Step 2: According to different face attribute category labels and the logistic regression binary classifier model, train to obtain an interpretable attribute feature space E, which contains the direction vectors e i of different attributes i, representing the change direction of the attribute in the latent space. Based on this change direction, attribute editing is achieved, which can be expressed by the following equation:
[0120] e i =argmax w P(y=w·x+b),y∈{0,1}
[0121] The binary classifier model is a logistic regression model, which can be represented by the following conditional probability distribution:
[0122]
[0123] where, x∈R n represents the input, and each sample contains n features; Y∈{0, 1} represents the output; w∈R n represents the weight vector, b∈R represents the bias, and they are the parameters that the classifier needs to learn. For each input x, by comparing the magnitudes of P(Y=0∣x) and P(Y=1∣x), x is classified into the category with the larger conditional probability. The maximum likelihood estimation method is used to estimate the model parameters w and b, and these parameters can accurately predict and fit the conditional probability. When the convergence condition of the algorithm is satisfied or the set maximum number of iterations is reached, the network training stops, and the final network and parameter weights are saved.
[0124] The loss function in Step 2 is constructed based on the binary cross-entropy loss function. Its goal is to find a high-dimensional hyperplane to separate positive and negative samples of different classes with the same attribute. To further address the problem of imbalance between positive and negative samples in the dataset and reduce the impact of some difficult-to-classify samples on the performance of the classifier, in this embodiment, the training of the model is achieved by combining focal loss. The loss function optimized in this step is expressed as the following equation:
[0125]
[0126] where α is a balancing factor, and its value range is [0, 1]. It is mainly used to balance the proportion of positive and negative samples. In this embodiment, the value for samples with a larger number is 0.3, and (1 - p) γ is a modulation factor, where γ is the focusing parameter, which is an adjustable hyperparameter and takes a non-negative value. In this embodiment, its value is 2. When γ = 0, the loss function degenerates into a weighted cross-entropy loss function. As γ increases, the model will pay more attention to difficult-to-classify samples.
[0127] S6. Construct a face attribute editing and evaluation module to implement face attribute editing and evaluation. Interpolate the direction vectors of different attributes to achieve face attribute editing, and evaluate the performance of the model by comparing the qualitative and quantitative results with existing networks;
[0128] In this embodiment, the face attribute editing and evaluation module is used to generate a face after target attribute editing and make a visual qualitative and quantitative comparison of the generated image.
[0129] The specific process is as follows. As shown in process ②, to maintain the consistency of the non-target attribute area, fix the parameters of the non-target attribute attention module, and based on the optimized latent vector w Figure 5 and the interpolation method with the direction vector e * of different attributes i i to obtain the final latent vector w target . Finally, use w target as the input of the pre-trained StyleGAN generator. Through the hierarchical output of the generator, the generated face after target attribute editing can be obtained. The specific formula for this process is as follows:
[0130] w target = w * + λe i
[0131] I edit = G style (w target )
[0132] where λ is the interpolation coefficient, with a value range of [0, 1]. When λ = 0, the original input image can be output. When λ = 1, the target attribute edited image can be output. G style is the generator of StyleGAN, and I edit is the target attribute edited image, which is obtained through the step-by-step refinement and step-by-step stylization of the generator network. It is synthesized by the interaction between the feature maps and the latent vector w target output at each layer;
[0133] The pre-trained StyleGAN generator adopts a multi-resolution step-by-step generation process with 8 stages. The network starts from a resolution of 4×4 and gradually increases to 8×8, 16×16, 32×132 resolutions, and finally to an image with a resolution of 512×512. This process is jointly controlled by w target and the adaptive instance normalization AdaIN layer. The initial input of the StyleGAN generator is a constant of 4×512×512. Each stage consists of an upsampling UpSample step and a convolution step based on a 3×3 convolution kernel. The expression of the AdaIN layer is as follows:
[0134]
[0135] where x represents the content feature input, which is the feature map from the previous layer in the generation network and contains the basic content information of the generated image. y represents the style feature input, and μ(y) and σ(y) are the scaling coefficient and bias coefficient obtained by w target through the affine transformation. At the same time, before the AdaIN module, random noise following a Gaussian distribution is added to each channel to enrich the details of the image, and finally a high-quality and delicate face edited image is generated.
[0136] This embodiment evaluates the designed network from two dimensions: qualitative and quantitative:
[0137] (1) Qualitative evaluation: The experimental results are as Figure 6 shown. From left to right are the input face, the reconstructed face, the edited face with the "hair" attribute, the edited face with the "lips" attribute, the edited face with the "eyes" attribute, and the edited face with the "bangs" attribute. The value of λ is 1. It can be seen that based on the 4 kinds of face attribute editing examples given, the present invention can realize the editing of specific face attributes. The edited face only changes in the target attribute area, and there are no obvious artifacts and distortions in the irrelevant features, identity, background, pose, etc. of the face, thus proving the effectiveness and feasibility of the present invention.
[0138] (2) Quantitative evaluation: To further quantitatively evaluate the quality of the generated images, in this embodiment, two commonly used metrics, FID and LPIPS, are used as standards, and compared with existing generation algorithms to prove the superiority of the method in this embodiment. FID calculates the distribution difference in the feature space between the generated images and the real images to describe the similarity between images. The calculation formula is as shown in the following equation:
[0139]
[0140] where μ r , ∑ r are the mean matrix and covariance matrix of the original images respectively, and μ g , ∑ g are the mean matrix and covariance matrix of the generated images respectively. Tr represents the trace of the matrix. The lower the FID, the better the quality of the synthesized images.
[0141] LPIPS measures the similarity between images from the perspective of human perception. The images are decomposed into blocks of different resolutions, the distances between features at each resolution are calculated, and the similarity between images is obtained by weighted summation. The calculation formula is as shown in the following equation:
[0142]
[0143] where I1 and I2 represent the original image and the generated image respectively, i represents the i-th layer of the pre-trained VGGFace network, is the normalized representation of the feature vector of the input image I at the i-th layer of the VGGFace network, and w i represents the weight coefficient, and ‖·‖2 represents the L2 norm. The lower the LPIPS, the better the perceptual similarity of the synthesized images.
[0144] In this embodiment, the above two commonly used metrics, FID and LPIPS, are used to evaluate the generated images. The generation models based on the StyleGAN backbone network, which are currently widely used, are selected for comparison, including the HiSD model (2021), the VecGAN model (2023), the FEAT model (2023), and the StyleMapGAN model (2023). The experimental results are shown in Tables 1 and 2 below. The bolded data are the optimal results;
[0145] Table 1 Comparison of FID evaluation index results
[0146]
[0147] Table 2 Comparison of LPIPS evaluation index results
[0148]
[0149] In this embodiment, quantitative evaluations are respectively carried out for four face attributes, and the average value of the four attributes is obtained to get the evaluation value of the overall image. As can be seen from Table 1 and Table 2, the values of FID and LPIPS of the present invention are 130.23 and 21.54 respectively. The overall performance of the model is better than that of the existing generation models, and excellent performance has been achieved in both the quality of the generated images and the consistency of the face identities in the generated images, thus verifying the effectiveness of the present invention.
[0150] The present invention overcomes the entanglement limitation of different face attribute feature vectors in the latent space, introduces local attribute mask supervision, and solves the problem of poor precise control ability for different attribute regions of face images. It provides face samples with diverse attributes for the face liveness detection model, can effectively reduce the misjudgment cases caused by facial attribute differences, and helps to enhance the ability of the face liveness detection model in facial attribute recognition.
[0151] Embodiment 2
[0152] As Figure 2 shown, this embodiment provides a live face attribute editing system based on mask supervision target feature decoupling, including: a dataset acquisition module, a dataset preprocessing module, a StyleGAN style encoder, a target attribute attention module, a non-target attribute attention module, a target attribute region mask supervision module, a latent vector self-fusion module, a reconstructed input face training module, a StyleGAN generator, and a face attribute editing picture output module;
[0153] In this embodiment, the dataset acquisition module is used to acquire a face dataset and a mask dataset corresponding to the face attributes;
[0154] In this embodiment, the dataset preprocessing module is used to combine the attribute segmentation masks provided by the mask dataset into a complete mask image;
[0155] In this embodiment, the StyleGAN style encoder is used to convert the face data into a latent vector;
[0156] In this embodiment, the target attribute attention module is used to output the attention-enhanced latent vector w based on the multi-channel attention mechanism t ;
[0157] In this embodiment, the non-target attribute attention module is used to output the attention-enhanced latent vector w based on the multi-channel attention mechanism b ;
[0158] In this embodiment, the target attribute region mask supervision module is used to combine the latent vector w t and w bGenerate the corresponding mask depth prediction map through the local generator, and optimize the attention network based on the attribute mask labels of the mask dataset;
[0159] In this embodiment, the latent vector self-fusion module is used to fuse the latent vector w t and the latent vector w b by splicing to form the latent vector w * ;
[0160] In this embodiment, the reconstructed input face training module is used to train the overall network based on the reconstructed input face, minimize the total model loss function, obtain the direction vectors of different attributes according to different attribute category labels and the logistic regression classification model, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model;
[0161] In this embodiment, the StyleGAN generator is used to reconstruct the input face;
[0162] In this embodiment, the face attribute editing picture output module is used to obtain the final latent vector w * by interpolating the latent vector w target with the direction vectors of different attributes, and use it as the input of the pre-trained StyleGAN generator. Through the hierarchical output of the StyleGAN generator, the generated face after editing the target attribute is obtained.
[0163] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent substitution methods and are all included in the protection scope of the present invention.
Claims
1. A live face attribute editing method based on mask supervision for target feature decoupling, characterized in that It includes the following steps: Obtain a face dataset and a mask dataset corresponding to face attributes; Combine the attribute segmentation masks provided by the mask dataset into a complete mask image; The face data is processed by the StyleGAN style encoder to obtain a latent vector, and the latent vector generates an attention-enhanced latent vector w based on the multi-channel attention mechanism. t and the latent vector w b ; The latent vector w t and the latent vector w b are passed through a local generator to generate corresponding masked depth prediction maps, and the attention network is optimized based on the attribute mask labels of the masked dataset; Concatenate the latent vector w t and the latent vector w b to fuse them into the latent vector w * ; Based on the reconstructed input face, train the overall network, minimize the total loss function of the model, obtain the direction vectors of different attributes according to different attribute category labels and the logistic regression classification model, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model; Based on the latent vector w * Interpolate with the direction vectors of different attributes to obtain the final latent vector w target , and use it as the input of the pre-trained StyleGAN generator. After the StyleGAN generator outputs layer by layer, a generated face with edited target attributes is obtained.
2. The method for editing living face attributes based on mask supervision and target feature decoupling according to claim 1, characterized in that The latent vector generates an attention-enhanced latent vector w based on a multi-channel attention mechanism t and the latent vector w b , specifically including: Construct a target attribute attention module and a non-target attribute attention module to perform self-attention reinforcement on latent vectors of different resolutions. The target attribute attention module uses a 1×1 convolutional layer, and the non-target attribute attention module uses a 3×3 convolutional layer; Generate a query matrix, a key matrix, and a value matrix. After the steps of feature concatenation of the query matrix and the key matrix, vector size adjustment, and Softmax normalization, a weight matrix is obtained. Multiply the weight matrix with the value matrix to generate the attention-enhanced latent vector w t and the latent vector w b , which is specifically expressed as: w i = Softmax(Reshape([Q i , K i W conv )]V i Among them, i has two values. When i takes the value of t, it represents the features related to the target attribute attention module. When i takes the value of b, it represents the features related to the non-target attribute attention module, V i , K i , Q i respectively represent the value matrix, key matrix, and query matrix corresponding to the above two modules. are the weight coefficients corresponding to the value matrix, key matrix, and query matrix respectively. are the feature mapping networks corresponding to the value matrix, key matrix, and query matrix respectively, w L , w M , w H correspond to the features of the low, medium, and high resolutions of the input latent vector w respectively. [·] represents the operation of feature concatenation, W conv represents the parameters of the convolutional layer. Reshape represents adjusting the feature size to be consistent with the latent vector w.
3. The method for editing living face attributes based on mask-supervised target feature decoupling according to claim 1, wherein The backbone network of the local generator is a ResNet model. Identity Block and Conv Block are used as residual connections. The convolutional layer has two types of convolutional kernels: 1×1 and 3×3. The 1×1 convolutional kernel changes the number of output channels, and the 3×3 convolution changes the size of the output feature map. BatchNorm is used as the batch normalization operation, and ReLU is used as the activation function. The input dimension and output dimension of IdentityBlock are the same, which is used to deepen the network. The input dimension and output dimension of Conv Block are different, which is used to change the feature dimension of the network.
4. The method for live face attribute editing based on mask supervision and target feature decoupling according to claim 1, wherein Concatenate the latent vector w t and the latent vector w b to fuse them into the latent vector w * , specifically including: Vector fusion method based on mean and variance statistical features, fusing the latent vector w t and the latent vector w b by concatenation into the latent vector w * , expressed as: μ(w mix ) = αμ(w t )+(1 - α)μ(w b )+n μ σ 2 (w mix ) = βσ 2 (w t )+(1 - β)σ 2 (w b )+n σ w * = MLP[μ(w mix ), σ 2 (w mix )] where n μ and n σ represent random noise sampled from the difference space between the latent vector w t and the latent vector w b . n μ follows a Gaussian distribution with mean |μ(w t ) - μ(w b )| and variance 1, and n σ follows a Gaussian distribution with mean |σ 2 (w t ) - σ 2 (w b )| and variance 1. [·] represents the operation of feature concatenation, MLP represents a multi-layer perceptron network, α is the weight coefficient of the mean feature, β is the weight coefficient of the variance feature, w mix is the latent vector after feature fusion, and w * is the finally output latent vector.
5. The method for live face attribute editing based on mask-supervised target feature decoupling according to claim 1, wherein The total loss function is expressed as: L = λ1L latent + λ2L mask + λ3L att ++ λ4L id L latent = 1 - CosSim(G(w * ), G(w)) Among them, L latent represents the latent vector consistency loss, L mask represents the attribute mask loss, L att represents the attention regularization loss, L id represents the identity perception loss, λ1, λ2, λ3, and λ4 represent the corresponding weight coefficients, G represents the StyleGAN generator, N represents the number of images in a batch, g(·) represents the mask depth prediction map, M t (·) represents the true attribute mask label, M(x i ) represents the global mask label of the face, α and β represent the weight magnitudes of the target and non-target attributes, H and W are the height and width of the image respectively, m ij represents the pixel value size of the attention map at the coordinates (i, j), I input represents the input face, I rec represents the reconstructed face, and respectively represent the output features of I input and I rec after passing through the l-th layer of ArcFace, ijk represents the result of the output activation of the i-th convolutional kernel at the (j, k) position after passing through the intermediate layer, C l 、H l 、W l are the number of channels, height, and width of the l-th layer feature map.
6. The method for editing living face attributes based on mask-supervised target feature decoupling according to claim 1, wherein The direction vectors of different attributes are obtained according to different attribute category labels and the logistic regression classification model, which is expressed as: e i = argmax w P(y = w·x + b), y ∈ {0, 1} Among them, e i represents the direction vector of different attributes i; The binary classifier model is a logistic regression model, which is represented by the following conditional probability distribution: Where x represents the input, each sample contains n features, Y∈{0, 1} represents the output; w represents the weight vector, and b represents the bias.
7. The method for editing living face attributes based on mask-supervised target feature decoupling according to claim 1, wherein The focal loss is expressed as: where α is the balance factor, (1 - p) γ is the modulation factor, and γ is the focusing parameter.
8. The method for editing living face attributes based on mask-supervised target feature decoupling according to claim 1, wherein Based on the latent vector w * Interpolate with the direction vectors of different attributes to obtain the final latent vector w target , which is specifically expressed as: w target = w * + λe i Through the output of the generator layer by layer, obtain the generated face after editing the target attribute: I edit = G style (w target ) Among them, e i represents the direction vector of different attributes i, λ is the interpolation coefficient, and G style is the generator of StyleGAN, and I edit is the target attribute edited image.
9. The live face attribute editing method based on mask-supervised target feature decoupling according to claim 1, wherein The pre-trained StyleGAN generator adopts a multi-resolution step-by-step generation process in multiple stages, which is jointly controlled by the final latent vector w target and the adaptive instance normalization (AdaIN) layer. Each stage includes an upsampling step and a convolution step based on a 3×3 convolution kernel. The AdaIN layer is expressed as: Among them, x represents the content feature input, which is the feature map from the previous layer in the generation network, y represents the style feature input, and μ(y) and σ(y) are the scaling coefficient and bias coefficient obtained through affine transformation; target The scaling coefficient and bias coefficient obtained through affine transformation; Add random noise that follows a Gaussian distribution to each channel before the AdaIN layer.
10. A live face attribute editing system based on mask supervision and target feature decoupling, characterized in that, It includes: A dataset acquisition module, a dataset preprocessing module, a StyleGAN style encoder, a target attribute attention module, a non-target attribute attention module, a target attribute region mask supervision module, a latent vector self-fusion module, a reconstructed input face training module, a StyleGAN generator, and a face attribute editing picture output module; The dataset acquisition module is used to obtain a face dataset and a mask dataset corresponding to face attributes; The dataset preprocessing module is used to combine the attribute segmentation masks provided by the mask dataset into a complete mask image; The StyleGAN style encoder is used to convert face data into latent vectors; The target attribute attention module is used to output the latent vector w with enhanced attention based on the multi-channel attention mechanism t ; The non-target attribute attention module is used to output a potentially enhanced vector w based on the multi-channel attention mechanism b ; The target attribute region mask supervision module is used to generate corresponding mask depth prediction maps for the latent vectors w t and w b through the local generator, and optimize the attention network based on the attribute mask labels of the mask dataset; The potential vector self-fusion module is used to splice and fuse the potential vector w t and the potential vector w b into the potential vector w * ; The reconstructed input face training module is used to train the overall network based on the reconstructed input face, minimize the total loss function of the model, obtain the direction vectors of different attributes according to different attribute category labels and the logistic regression classification model, optimize the performance of the classifier based on the focal loss, and save the best network and weights of the model; The StyleGAN generator is used to reconstruct the input face; The face attribute editing picture output module is used to obtain the final latent vector w * by interpolating the direction vectors of different attributes with the latent vector w, target and use it as the input of the pre-trained StyleGAN generator. After the StyleGAN generator outputs layer by layer, a generated face with edited target attributes is obtained.
Citation Information
Cited By
Vehicle tank body surface texture analysis method and system based on deep learning
CN121724987A
A vehicle tank surface texture analysis method and system based on deep learning
CN121724987B