A method for extracting feature latent codes, a computer device, and a storage medium
By extracting the three-dimensional reconstruction images of the face image and inputting them into the feature latent code extraction network and the style confrontation generation network, the problem of difficulty in extracting the feature latent code of the face image in the prior art is solved, and accurate restoration of the face image and controllable attribute changes are achieved.
Patent Information
- Application Number
- CN202111027269.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-02
AI Technical Summary
The prior art is difficult to extract feature codes in real scenes from face images, which limits the application of style confrontation generation networks, such as the inability to generate images of changes in facial expressions or aging changes in actual users.
By obtaining the first face image of the feature latent code to be extracted and its corresponding three-dimensional reconstruction image, input it into the feature latent code extraction network and the style confrontation generation network, a second face image with three-dimensional spatial features is generated, and according to the difference between the first face image and the second face image, the feature latent code extraction network is iteratively adjusted until it converges.
It realizes the accurate extraction of feature latent codes from face images, so that the style confrontation generation network can effectively restore face images, and expands the application scope of the network, such as generating images of user's facial expression changes and aging changes.
Smart Images

Figure CN113763535B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of image processing, and in particular, to a method for extracting feature latent codes, a computer device, and a storage medium. Background Art
[0002] The style generative adversarial network (StyleGAN) is a network structure used for face image generation in the field of artificial intelligence recently. By adding a mapping network to the generator network, feature latent codes related to face attributes and noise are introduced into the feature map, and then face images can be generated.
[0003] The inventors of the present application found in the process of implementing the embodiments of the present application that currently, randomly generated feature latent codes are input into the style generative adversarial network to generate images. Since the feature latent codes of images in the real scene cannot be obtained, the application of the adversarial generative network is limited. For example, it is impossible to use the adversarial generative network to generate images of the facial expression changes and aging changes of actual users. Therefore, how to extract feature latent codes from face images is an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by the embodiments of the present application is to provide a method for extracting feature latent codes, a computer device, and a storage medium. The method can effectively extract the feature latent codes that can reflect the face images from the face images, that is, make the feature latent codes accurate and be able to effectively restore the face images, which is beneficial to expanding the application of the style generative adversarial network.
[0005] To solve the above technical problem, in a first aspect, the embodiments of the present application provide a method for extracting feature latent codes, including:
[0006] Obtain a first face image for which feature latent codes are to be extracted and a three-dimensional reconstruction image corresponding to the first face image, where the three-dimensional reconstruction image includes the three-dimensional spatial features of the face in the first face image;
[0007] Input the first face image into a preset feature latent code extraction network to obtain the feature latent codes output by the feature latent code extraction network;
[0008] Input the feature latent codes and the three-dimensional reconstruction image into the style generative adversarial network to generate a second face image fused with the three-dimensional spatial features;
[0009] Iteratively adjust the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image until the feature latent code extraction network converges;
[0010] Use the feature latent code output by the converged feature latent code extraction network as the feature latent code of the first face image.
[0011] In some embodiments, obtaining the three-dimensional reconstruction image corresponding to the first face image includes:
[0012] Perform three-dimensional face reconstruction on the face in the first face image to obtain first three-dimensional reconstruction parameters;
[0013] Perform two-dimensional rendering on the first three-dimensional reconstruction parameters to obtain the three-dimensional reconstruction image corresponding to the first face image.
[0014] In some embodiments, the style adversarial generation network includes a mapping network and a plurality of sequentially arranged generation networks, and the plurality of generation networks are respectively used to output feature maps of different sizes. Among them, the target generation network is used to generate and output a second feature map according to the input first feature map, and the size of the second feature map is larger than that of the first feature map. The target generation network is any one of the plurality of sequentially arranged generation networks;
[0015] Input the feature latent code and the three-dimensional reconstruction image into the style adversarial generation network to generate a second face image fused with the three-dimensional space feature, including:
[0016] Input the feature latent code into the mapping network to decouple the features of the feature latent code and generate an intermediate vector;
[0017] Extract features from the three-dimensional reconstruction image according to the target size to obtain a three-dimensional reconstruction feature map. The target size is the size of the second feature map, and the size of the three-dimensional reconstruction feature map is the target size;
[0018] Input the intermediate vector, the first feature map, the three-dimensional reconstruction feature map, and random noise into the target generation network for fusion to output the second feature map;
[0019] Determine the second feature map output by the last generation network in the plurality of sequentially arranged generation networks as the second face image.
[0020] In some embodiments, the target generation network includes at least one convolutional layer and a first fusion layer, and a second fusion layer is configured after each convolutional layer.
[0021] Input the intermediate vector, the first feature map, the three-dimensional reconstruction feature map, and random noise into the target generation network for fusion to output the second feature map, including:
[0022] Input the intermediate vector into each second fusion layer. Each second fusion layer performs an affine transformation on the intermediate vector according to the target size to obtain a feature factor, where the feature factor is adapted to the target size;
[0023] The target second fusion layer uses the feature factor to fuse the first intermediate feature map input to the target second fusion layer with the random noise to obtain the second intermediate feature map output by the target second fusion layer, where the target second fusion layer is any one of the second fusion layers, and the first intermediate feature map is the feature map output by the convolutional layer in the previous layer of the target second fusion layer;
[0024] Input the second intermediate feature map output by the last second fusion layer and the three-dimensional reconstruction feature map into the first fusion layer for fusion to obtain the second feature map.
[0025] In some embodiments, the feature factor includes a scaling factor and a bias factor;
[0026] The target second fusion layer uses the feature factor to fuse the first intermediate feature map input to the target second fusion layer with the random noise to obtain the second intermediate feature map output by the target second fusion layer, including:
[0027] The second intermediate feature map output by the target second fusion layer is calculated using the following formula,
[0028] y ij = y (s,i) *(T ij + B i ) + y (b,i) ;
[0029] where i is the label of the target generation network, 1 ≤ i ≤ N, N is the number of generation networks, j is the label of the target second fusion layer, 1 ≤ j ≤ M, M is the number of second fusion layers in the target generation network, Tij is the first intermediate feature map input to the target second fusion layer, Bi is the random noise, y(s,i) is the scaling factor, and y(b,i) is the bias factor;
[0030] The input of the second intermediate feature map output by the last second fusion layer and the three-dimensional reconstruction feature map into the first fusion layer for fusion to obtain the second feature map includes:
[0031] The second feature map is calculated using the following formula;
[0032] X i+1 = y ij * h i , j = M
[0033] Among them, y ij is the second intermediate feature map output by the last second fusion layer, h i is the three-dimensional reconstruction feature map, and X i+1 is the second feature map.
[0034] In some embodiments, the difference between the first face image and the second face image includes the structural difference between the first face image and the second face image;
[0035] Before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, it further includes:
[0036] Calculating the brightness similarity value, contrast similarity value, and structural similarity value between the first face image and the second face image;
[0037] Taking the product of the brightness similarity value, the contrast similarity value, and the structural similarity value as the structural difference.
[0038] In some embodiments, the difference between the first face image and the second face image includes the pixel difference between the first face image and the second face image;
[0039] Before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, it further includes:
[0040] Calculating the pixel difference between the pixel points at the target position in the first face image and the pixel points at the target position in the second face image to obtain the pixel difference corresponding to the target position, where the target position is any position in the first face image or the second face image;
[0041] Taking the sum of the pixel differences corresponding to each position in the first face image and the second face image as the pixel difference.
[0042] In some embodiments, the difference between the first face image and the second face image includes the three-dimensional reconstruction parameter difference between the first face image and the second face image;
[0043] Before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, it further includes:
[0044] Calculate the parameter difference between the first 3D reconstruction parameter and the second 3D reconstruction parameter, and use the parameter difference as the 3D reconstruction parameter difference. The first 3D reconstruction parameter is the 3D reconstruction parameter obtained by performing 3D face reconstruction on the face in the first face image, and the second 3D reconstruction parameter is the 3D reconstruction parameter obtained by performing 3D face reconstruction on the face in the second face image.
[0045] To solve the above technical problem, in a second aspect, an embodiment of the present application provides a computer device, including a memory and one or more processors. The one or more processors are configured to execute one or more computer programs stored in the memory. When the one or more processors execute the one or more computer programs, the computer device implements the method described in the first aspect above.
[0046] To solve the above technical problem, in a third aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the method described in the first aspect above.
[0047] The beneficial effects of the embodiments of the present application: Different from the prior art, the feature latent code extraction method provided by the embodiments of the present application obtains the first face image for which the feature latent code is to be extracted and the corresponding 3D reconstruction image of the first face image, inputs the feature latent code of the first face image and the 3D reconstruction image extracted by the feature latent code extraction network into the style generation adversarial network to generate a second face image fused with the 3D spatial features of the first face image. Then, according to the difference between the first face image and the second face image, the feature latent code extraction network is iteratively tuned until the feature latent code extraction network converges. Finally, the feature latent code output by the converged feature latent code extraction network is used as the feature latent code of the first face image.
[0048] Since the feature latent code output by the feature latent code extraction network is used as the feature latent code of the first face image after the feature latent code extraction network converges, it is ensured that the second face image is sufficiently similar to the first face image, that is, it is ensured that the feature latent code finally output by the feature latent code extraction network can accurately reflect the feature attributes of the first face image, so that the finally output feature latent code can restore the first face image. Therefore, changing the feature latent code in the style adversarial generation network can realize controllable changes to the attributes of the first face image.
[0049] In addition, in the process of generating the second face image using the style adversarial generation network, the three-dimensional reconstruction image reflecting the three-dimensional spatial features of the face in the first face image is input into the style generation adversarial network. Thus, in the process of the style adversarial generation network generating the second face image, the three-dimensional reconstruction image can play a role in supervising the position of the entire face and the distribution of facial features, enabling the feature maps output by the style adversarial generation network to have targeted responses at positions with different geometric information, so that the second face image incorporates three-dimensional spatial features and is closer to the original first face image. This can effectively reduce the error introduced by the style adversarial network to the second face image, that is, the second face image can accurately express the feature latent code, which helps the feature latent code extraction network output more accurate feature latent codes after convergence. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the figures in the drawings do not constitute a scale limitation.
[0051] Figure 1 Partial structural schematic diagram of the style adversarial generation network provided by an embodiment of the present application;
[0052] Figure 2 Flow schematic diagram of the feature latent code extraction method provided by an embodiment of the present application;
[0053] Figure 3 Structural schematic diagram of the style adversarial generation network provided by another embodiment of the present application;
[0054] Figure 4 For Figure 2 Sub-flow schematic diagram of step S23 in the method shown;
[0055] Figure 5 For Figure 4 Sub-flow schematic diagram of step S233 in the method shown;
[0056] Figure 6 Flow schematic diagram of the feature latent code extraction method provided by another embodiment of the present application;
[0057] Figure 7 Flow schematic diagram of the feature latent code extraction method provided by another embodiment of the present application;
[0058] Figure 8 Structural block diagram of the computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.
[0060] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. In addition, the terms "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.
[0062] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in this specification in the description of the present application are only for the purpose of describing specific embodiments and are not used to limit the present application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0063] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0064] To facilitate the understanding of the technical solution of the present application, the relevant principles of the style adversarial generation network and the feature latent code involved in the present application will be introduced first.
[0065] See Figure 1 , Figure 1 which is a partial structural schematic diagram of the style adversarial generation network provided by the embodiment of the present application. As Figure 1As shown, the style adversarial generative network includes a mapping network S1 and an image generator S2. Among them, the mapping network S1 is used to perform feature decoupling on the composite features contained in the feature latent code, so as to map the feature latent code into multiple sets of feature control vectors input to the image generator, and the multiple sets of feature control vectors obtained by mapping are input into the image generator S2 to perform style control on the image generator (i.e., facial attribute control). The image generator S2 is used to perform style control and processing on the constant tensor based on the control vector input by the mapping network S1, so as to generate an image. When the feature latent code reflects facial features, the image generated by the image generator S2 is a facial image.
[0066] Among them, Figure 1 As shown, the mapping network S1 includes 8 fully connected layers, which are connected in sequence to perform nonlinear mapping on the feature latent code to obtain an intermediate vector w. The intermediate vector w can reflect various facial features, such as eye features, mouth features, or nose features.
[0067] The image generator S2 includes N generative networks arranged in sequence. The first generative network includes a constant tensor const and a convolution layer. The constant tensor and the convolution layer are both configured with an adaptive instance normalization layer. It can be understood that the constant tensor const is equivalent to an initial data for generating an image. Any of the remaining generative networks except the first generative network includes two convolution layers, and each convolution layer is configured with an adaptive instance normalization layer. In the image generator S2, each generative network outputs a feature map, which serves as the input of the next generative network. As the generative network is recursive, the size of the output feature map becomes larger and larger. The feature map output by the last generative network is the generated face image. It can be understood that the target size of the feature map output by each generative network is set, for example Figure 1 The size of the feature map output by the first generation network is 4*4, and the size of the feature map output by the second generation network is 8*8. If the size of the final generated face image is 1024*1024, the size of the feature map output by the last generation network is 1024*1024.
[0068] In a generative network, convolutional layers and adaptive instance normalization layers are interleaved, and the output of the previous layer is the input of the next layer. Figure 1 As shown in the figure, the convolution layer includes a 3*3 convolution kernel, which is used to perform deconvolution operation on the input image, and the output size is increased to obtain a feature map. The feature map output by the convolution layer is then input to the instance normalization layer, and the intermediate vector w and random noise B are also input to the instance normalization layer. The processing process of the instance normalization layer (i.e., AdaIN) is as follows: Figure 1Shown as follows: The intermediate vector w is expanded into a scaling factor y(s,i) and a bias factor y(b,i) through a learnable affine transformation (i.e., a fully connected layer). These two factors are weighted and summed with the feature map output by the previous convolutional layer after normalization and the random noise B. Thus, the intermediate vector w affects the style of the feature map output by the convolutional layer, that is, style control can be achieved. Secondly, the random noise B is used to enrich the details of the feature map.
[0069] Above is the process of generating an image by the style adversarial generation network. There are various application scenarios for the style adversarial generation network, for example, it can be used in the image generation scenario.
[0070] Among them, the feature latent code is a kind of feature vector, which can also be called a feature map. Specifically, it can be a multi-dimensional vector, and each value in the vector is within the range of [-1,1]. For example, it can be a 18*512 vector, and each value in this vector is within the range of [-1,1]. It can be understood that by inputting the feature latent code into the style adversarial generation network, the image corresponding to the feature latent code can be generated. It can be understood that the feature latent code can also be understood as the feature of the image extracted from the image based on the neural network. The feature latent code can represent the image. When the feature latent code is determined, the image generated based on this feature latent code is also determined. And from another perspective, the feature latent code can also be understood as the vector output after the image passes through the convolutional layer in the neural network.
[0071] However, currently, randomly generated feature latent codes are input into the style adversarial generation network to generate images. Since the feature latent codes of images in the real scenario cannot be obtained, the application of the adversarial generation network is restricted. For example, it is impossible to use the adversarial generation network to generate images of the facial expression changes and aging changes of actual users.
[0072] In view of this, an embodiment of the present application proposes a method for extracting feature latent codes. Please refer to Figure 2 , and this method includes but is not limited to the following steps:
[0073] S21: Obtain a first face image for which the feature latent code is to be extracted and a three-dimensional reconstruction image corresponding to the first face image. The three-dimensional reconstruction image includes the three-dimensional spatial features of the face in the first face image.
[0074] S22: Input the first face image into a preset feature latent code extraction network to obtain the feature latent code output by the feature latent code extraction network.
[0075] S23: Input the feature latent code and the three-dimensional reconstruction image into the style adversarial generation network to generate a second face image fused with three-dimensional spatial features.
[0076] S24: Iteratively adjust the parameters of the feature latent code extraction network according to the differences between the first face image and the second face image until the feature latent code extraction network converges.
[0077] S25: Use the feature latent code output by the converged feature latent code extraction network as the feature latent code of the first face image.
[0078] In step S21, the first face image is an image including a face. For example, the first face image can be an ID photo, etc. In this embodiment, it is necessary to extract the feature latent code from the first face image, that is, obtain a one-dimensional vector that can reflect the features of the first face image from the first face image. Equivalently, the feature latent code is the vector representation form of the first face image.
[0079] The three-dimensional reconstruction image corresponding to the first face image is another representation form of the first face image. This three-dimensional reconstruction image includes the three-dimensional spatial features of the face in the first face image. The three-dimensional spatial features can be understood as solid geometry features. In some embodiments, the first face image is converted into a three-dimensional face (constituted by three-dimensional point cloud data), and then after two-dimensional rendering of the three-dimensional face, a two-dimensional rendered image is obtained, that is, the three-dimensional reconstruction image is obtained.
[0080] In some embodiments, the steps of obtaining the three-dimensional reconstruction image corresponding to the first face image specifically include:
[0081] A) Perform three-dimensional face reconstruction on the face in the first face image to obtain the first three-dimensional reconstruction parameters.
[0082] B) Perform two-dimensional rendering on the first three-dimensional reconstruction parameters to obtain the three-dimensional reconstruction image corresponding to the first face image.
[0083] In this embodiment, a 3DMM model can be used to calculate the first three-dimensional reconstruction parameters of the first face image, that is, to perform three-dimensional face reconstruction on the face in the first face image to obtain the first three-dimensional reconstruction parameters.
[0084] In the 3DMM model, all three-dimensional faces can be represented by the same number of point clouds (spatial coordinate positions), and the points with the same serial number represent the same semantics. For example, for each face, the k-th point cloud is the left outer canthus point. The 3DMM model statistically analyzed the facial laser scanning data of 200 face samples (100 young men and 100 young women) to obtain a face statistical model. The three-dimensional reconstruction parameters of the face can be calculated according to the following formula:
[0085]
[0086]
[0087] Among them, S 2Represents the average shape of 200 face samples, that is, the average of the spatial coordinate positions of the point clouds included in the 200 face samples, T 2 Represents the average texture of 200 face samples, that is, the average texture of the point clouds included in the 200 face samples; S i Is the orthogonal position feature basis vector obtained after performing principal component analysis (PCA) on these 200 face samples, T i Is the color feature basis vector obtained after performing principal component analysis (PCA) on these 200 face samples; α i Is S i 's coefficient, β i Is T i 's coefficient, where, α i =(α 1 ,α 2 ,......,α 199 ), β i =(β 1 ,β 2 ,......,β 199 ). It should be noted that here i is the label of the vector dimension, representing the position in the dimension. Thus, any face can obtain its 3D reconstruction parameters by adjusting the coefficients α and β. For example, for the first face image, the first 3D reconstruction parameters corresponding to the first face image can also be obtained by adjusting the coefficients α and β.
[0088] Then, input the first 3D reconstruction parameters into the differentiable renderer to generate a 2D rendered image, which is the 3D reconstruction image corresponding to the first face image. Among them, the differentiable renderer can be mesh_renderer in the TensorFlow framework (i.e., differentiable 3D network renderer), which belongs to the prior art and will not be described in detail here. It can be understood that when rendering, the size of the generated 2D rendered image can be set to 256*256 for convenient subsequent processing.
[0089] In this embodiment, through the above method, the 3D reconstruction image corresponding to the first face image is obtained, so that the 3D reconstruction image includes the 3D spatial features of the face in the first face image.
[0090] In step S22, the feature latent code extraction network is used to extract the features of the first face image. The feature latent code extraction network includes a convolutional layer, an activation function layer, and a normalization layer to reduce the dimension of the input first face image and output a feature latent code (one-dimensional vector). It can be understood that the mathematical expression of the feature latent code extraction network is as follows:
[0091]
[0092] Where, fm l represents the m-th feature map of the l-th layer, f m l+1 represents the n-th feature map of the (l + 1)-th layer, W represents the convolutional kernel, B represents the bias term, σ(·) represents the ReLU activation function, and IN represents normalization. By setting the parameters of each convolutional layer (such as the convolutional kernel size, number, and corresponding stride), the first face image is dimensionally reduced to generate a feature map. The feature map e output by the last convolutional layer is input into the fully connected layer, and the fully connected layer transforms the feature map e into a one-dimensional feature vector, such as a 9216 * 1 feature vector. This one-dimensional feature vector is the feature latent code. In some embodiments, the one-dimensional feature vector is also transformed into a multi-dimensional vector, and this multi-dimensional vector is the feature latent code.
[0093] It can be understood that the feature latent code extraction network can be an existing Mobilenet algorithm, VGG algorithm, etc. The specific structure of the feature latent code extraction network is not limited herein, as long as feature extraction and dimensional reduction are achieved.
[0094] In step S23, the feature latent code and the three-dimensional reconstruction image are input into the style adversarial generation network to generate a second face image fused with three-dimensional spatial features. From the principle of the above-mentioned style adversarial generation network, it can be known that the feature latent code is deconvolved to generate a series of feature maps with gradually increasing sizes. The feature map output by the last layer of the network has the largest size and is the second face image. During the process of generating a series of feature maps, the three-dimensional reconstruction image can be fused with at least one feature map, such as linear fusion or non-linear fusion, etc., so that the second face image is fused with the three-dimensional spatial features included in the three-dimensional reconstruction image.
[0095] That is, during the process of generating the second face image using the style adversarial generation network, the three-dimensional reconstruction image reflecting the three-dimensional spatial features of the face in the first face image is input into the style generation adversarial network together. Thus, during the process of the style adversarial generation network generating the second face image, the three-dimensional reconstruction image can play a role in supervising the position of the entire face and the distribution of facial features, enabling the feature maps output by the style adversarial generation network to have targeted responses at positions with different geometric information, making the second face image fused with three-dimensional spatial features and closer to the original first face image. It can effectively reduce the error brought to the second face image by the style adversarial network, that is, the second face image can accurately express the feature latent code, which helps the feature latent code extraction network output a more accurate feature latent code after subsequent convergence.
[0096] It can be understood that if the feature latent code is accurate, that is, the feature latent code can truly reflect the features of the first face image, the more similar the second face image generated from the feature latent code is to the first face image. Thus, the feature latent code extraction network can be iteratively tuned according to the difference between the first face image and the second face image until the feature latent code extraction network converges. It can be understood that the convergence of the feature latent code extraction network here can mean that at a certain parameter, the difference between the first face image and the second face image output by the network is less than a preset threshold or fluctuates within a certain range. When the feature latent code extraction network converges, it indicates that the first face image and the second face image are highly similar, that is, the feature latent code output by the converged feature latent code extraction network can well restore the first face image. Therefore, using the feature latent code output by the converged feature latent code extraction network as the feature latent code of the first face image makes the feature latent code of the first face image more accurate and reasonable.
[0097] In some embodiments, the adam algorithm is used to optimize the model parameters. For example, the number of iterations is set to 1000 times, the initial learning rate is set to 0.001, the weight decay of the learning rate is set to 0.0005, and every 10 iterations, the learning rate decays to 1 / 10 of the original. Among them, the learning rate, and the difference between the first face image and the second face image can be input into the adam algorithm to obtain the adjustment parameters output by the adam algorithm, and the next training is performed using these adjustment parameters until after the training is completed, the model parameters of the converged feature latent code extraction network are output.
[0098] It should be noted that in the embodiments of the present application, the first face image is an image. Based on this image, the feature latent code extraction network is trained to obtain a feature latent code extraction model (the converged feature latent code extraction network) corresponding to the first face image. This feature latent code extraction model is more suitable for obtaining the feature latent code of the first face image, that is, this feature latent code extraction model only corresponds to the first face image and is not suitable for other face images. Moreover, further, the training of the feature latent code extraction network is not to obtain a general feature latent code extraction model, but to extract the feature latent code from the first face image. During the training process, as the feature latent code extraction network gradually converges, the difference between the second face image restored by the feature latent code and the first face image becomes smaller and smaller, and finally, a feature latent code with high accuracy is obtained.
[0099] It can be understood that through the above method, it is possible to effectively extract the feature latent code from the face image, making it possible to extract the feature latent code. Further, since the feature latent code can be extracted from the face image, it is possible to change the features of the face in the face image or realize the age change of the face by modifying the feature latent code, providing users with richer technical solutions. For example, using the above method to extract the feature latent code a2 of the face image a1 of user A, where the expression of user A in the face image a1 is serious. Modify the feature latent code a2 to obtain the modified feature latent code a3. Input the feature latent code a3 into the style generative adversarial network, and a new face image a4 can be obtained. The expression of user A in the face image a4 is a smile, so that by modifying the feature latent code, more and richer images belonging to the same user can be obtained.
[0100] In the embodiment of the present application, by using the style generative adversarial network, the first face image, and the three-dimensional reconstruction image including the three-dimensional spatial features of the face in the first face image to train the feature latent code extraction network, during the training process, the feature latent code extracted from the first face image is continuously optimized until the feature latent code extraction network converges. The feature latent code output by the converged feature latent code extraction network is used as the feature latent code of the first face image. Since the feature latent code output by the feature latent code extraction network is used as the feature latent code of the first face image after the feature latent code extraction network converges, it is ensured that the second face image is similar enough to the first face image, that is, it is ensured that the feature latent code finally output by the feature latent code extraction network can accurately reflect the feature attributes of the first face image, so that the finally output feature latent code can restore the first face image. Therefore, changing the feature latent code in the style adversarial generation network can realize the controllable change of the attributes of the first face image.
[0101] In addition, during the process of generating the second face image using the style adversarial generation network, the three-dimensional reconstruction image reflecting the three-dimensional spatial features of the face in the first face image is input into the style generative adversarial network together. Thus, during the process of the style adversarial generation network generating the second face image, the three-dimensional reconstruction image can play a role in supervising the position of the entire face and the distribution of facial features, making the feature map output by the style adversarial generation network have targeted responses at different geometric information positions, so that the second face image incorporates three-dimensional spatial features and is closer to the original first face image, effectively reducing the error brought by the style adversarial generation network to the second face image, that is, the second face image can accurately express the feature latent code, which helps the converged feature latent code extraction network output more accurate feature latent codes.
[0102] In some embodiments, please refer to Figure 3, the style adversarial generation network includes a mapping network and multiple sequentially arranged generation networks. The multiple generation networks are respectively used to output feature maps of different sizes. The feature map output by the previous generation network is used as the input to the subsequent generation network. As the generation network increases, the size of the output feature map continuously increases, and the finer the features reflected by the feature map with a larger size.
[0103] The following takes the target generation network as an example to exemplarily illustrate the structure and processing process of the generation network. It can be understood that the structures and processing processes of each generation network are the same. The target generation network is any one of the multiple sequentially arranged generation networks. For the convenience of description, the label of the target generation network is denoted as i, that is, the i-th generation network, and the total number of generation networks is N. The target generation network is used to generate and output a second feature map according to the input first feature map, where the first feature map is the feature map output by the previous generation network (i - 1) of the target generation network, and the size of the second feature map is larger than that of the first feature map.
[0104] That is, as Figure 3 shown, each generation network outputs a feature map, and for two adjacent generation networks, the output of the previous generation network is the input of the subsequent generation network. And, as i increases, the size of the output second feature map becomes larger and larger. For example, if N = 9, it includes 9 generation networks. If the size of the second feature map output by the first generation network is 4 * 4, then the size of the first feature map input to the second generation network is 4 * 4, and the size of the second feature map output by the second generation network is 8 * 8, and so on. The size of the second feature map output by the 9th generation network is 1024 * 1024. Each generation network performs a transposed convolution operation on the input first feature map to increase the dimension of the first feature map.
[0105] It can be understood that the earlier the target generation network i (i.e., the smaller i is), the smaller the size of the output second feature map, and the coarser the features that can be affected. When the size of the output second feature map does not exceed 8^2, it mainly affects the pose, hairstyle, or facial shape, etc. When the size of the output second feature map is greater than 16^2 and less than 32^2, it affects finer facial features, hairstyle, opening or closing of eyes, etc. When the size of the output second feature map is greater than 64^2 and less than or equal to 1024, it affects the color of eyes, hair, or skin, as well as microscopic features.
[0106] It should be noted that the multiple sequentially arranged generation networks are the Figure 1 image generator S2 in the shown embodiment, and the mapping network is the Figure 1 mapping network in the shown embodiment. The relevant principles of the style adversarial generation network will not be elaborated one by one here.
[0107] In this embodiment, please refer toFigure 4 , step S23 specifically includes:
[0108] S231: Input the feature latent code into the mapping network to decouple the features of the feature latent code and generate an intermediate vector.
[0109] S232: According to the target size, extract features from the three-dimensional reconstruction image to obtain a three-dimensional reconstruction feature map. The target size is the size of the second feature map, and the size of the three-dimensional reconstruction feature map is the target size.
[0110] S233: Input the intermediate vector, the first feature map, the three-dimensional reconstruction feature map, and random noise into the target generation network for fusion to output the second feature map.
[0111] S234: Determine the second feature map output by the last generation network in the multiple sequentially arranged generation networks as the second face image.
[0112] First, input the feature latent code into the mapping network to decouple the features of the feature latent code and generate an intermediate vector w. Since the mapping network includes multiple fully connected layers, the feature decoupling process is the operation process of convolutional operation downsampling, that is, reducing the dimension of the feature latent code into a one-dimensional intermediate vector w. This intermediate vector w will be input into each generation network respectively, and each generation network obtains 2 control vectors according to this intermediate vector W, so that different elements of the control vector can control different visual features, so that when adjusting a certain control vector, other control vectors will not change, that is, it will not affect other features and there will be no feature entanglement.
[0113] In order to fuse the features in the three-dimensional reconstruction image into the second feature map and make the fused features adapt to the second feature map, according to the target size (i.e., the size of the second feature map), extract features from the three-dimensional reconstruction image to obtain a three-dimensional reconstruction feature map. Among them, the specific way of feature extraction can be convolutional operation downsampling. Since the size of the three-dimensional reconstruction feature map is the target size, that is, the size of the three-dimensional reconstruction feature map is the same as the size of the second feature map, the roughness of the features reflected by the three-dimensional reconstruction feature map is the same as the roughness of the features reflected by the second feature map.
[0114] As Figure 3 shown, the features of the first three-dimensional reconstruction feature map (4*4*512) are rougher than those of the second three-dimensional reconstruction feature map (8*8*512), and the features of the seventh three-dimensional reconstruction feature map (256*256*512) are the finest.
[0115] It can be understood that the number of three-dimensional reconstruction feature maps is the same as the number of generation networks, and the three-dimensional reconstruction feature maps correspond one-to-one with the generation networks, that is, one generation network inputs a three-dimensional reconstruction feature map with a matching size.
[0116] Then, the intermediate vector, the first feature map, the three-dimensional reconstruction feature map, and the random noise are input into the target generation network for fusion to output the second feature map.
[0117] Among them, the fusion calculation process of the intermediate vector, the first feature map, and the random noise is the same as that of the fusion calculation process of a generation network in the Figure 1 illustrated embodiment. That is, in this embodiment, the module in the target generation network for fusing the intermediate vector, the first feature map, and the random noise also has the same layer structure as the corresponding module in the Figure 1 illustrated embodiment.
[0118] In this embodiment, the size of the three-dimensional reconstruction feature map is the same as the size of the second feature map output by the target generation network. Thus, the three-dimensional reconstruction feature map can be linearly or non-linearly fused with the second feature map before output. Since the size of the three-dimensional reconstruction feature map is adapted to the target generation network, the features of the three-dimensional reconstruction image can be incorporated into each second feature map output by each generation network in batches according to the coarseness and fineness of the features. At each level of generating the second feature map, it can play a role in supervising the position of the entire face and the distribution of facial features, enabling the second feature maps output by each generation network to have targeted responses at positions with different geometric information.
[0119] It can be understood that linear fusion is to perform a linear operation on two feature maps through a linear function, such as addition or subtraction, etc., and non-linear fusion is to perform a non-linear operation on two feature maps through a non-linear function, such as multiplication, division, or taking the logarithm, etc.
[0120] Based on the principle of the style adversarial generation network, the second feature map output by the last generation network among multiple sequentially arranged generation networks is determined as the second face image. This second face image incorporates three-dimensional spatial features and is closer to the original first face image, which can effectively reduce the error brought to the second face image by the style adversarial network. That is, the second face image can accurately express the feature latent code, which helps the feature latent code extraction network after convergence to output a more accurate feature latent code.
[0121] In this embodiment, through the above method, the features of the three-dimensional reconstruction image can be incorporated into each second feature map output by each generation network in batches according to the coarseness and fineness of the features. At each level of generating the second feature map, it can play a role in supervising the position of the entire face and the distribution of facial features, enabling the second feature maps output by each generation network to have targeted responses at positions with different geometric information. That is, by fusing according to the coarseness and fineness of the features, the second face image can more accurately express the feature latent code, which helps the feature latent code extraction network after convergence to output a more accurate feature latent code.
[0122] In some embodiments, the target generation network includes at least one convolutional layer and a first fusion layer, and a second fusion layer is configured after each convolutional layer. The second fusion layer is the Figure 1 adaptive instance normalization layer in the illustrated embodiment. It can be understood that the first fusion layer is used to fuse the three-dimensional reconstruction feature maps.
[0123] Based on the above network structure, in this embodiment, please refer to Figure 5 , step S233 specifically includes:
[0124] S2331: Input the intermediate vector into each second fusion layer. Each second fusion layer performs an affine transformation on the intermediate vector according to the target size to obtain a feature factor, where the feature factor is adapted to the target size.
[0125] S2332: The target second fusion layer uses the feature factor to fuse the first intermediate feature map input to the target second fusion layer with random noise to obtain the second intermediate feature map output by the target second fusion layer, where the target second fusion layer is any one of the second fusion layers, and the first intermediate feature map is the feature map output by the convolutional layer in the previous layer of the target second fusion layer.
[0126] S2333: Input the second intermediate feature map output by the last second fusion layer and the three-dimensional reconstruction feature map into the first fusion layer for fusion to obtain a second feature map.
[0127] It can be understood that the coarseness of the features reflected by the second feature maps output by different generation networks is different. Thus, each second fusion layer in the target generation network performs an affine transformation on the intermediate vector according to the corresponding coarseness of the features corresponding to the target generation network (i.e., the size of the second feature map, the target size) to obtain a feature factor. It can be understood that each generation network corresponds to a feature factor. Feature factor i corresponds to target generation network i, and the feature factor is adapted to the size of the second feature map. That is, the coarseness of the features that the feature factor can control is roughly the same as the coarseness of the features reflected by the second feature map.
[0128] Since there is at least one second fusion layer in the target generation network, and the processing process of each second fusion layer is the same, any one of the second fusion layers (i.e., the target second fusion layer) is taken as an example for exemplary illustration here. The target second fusion layer can be an instance normalization layer (AdaIN), which uses a feature factor to fuse the input first intermediate feature map with random noise to obtain the output second intermediate feature map, so that the second intermediate feature map is fused with the features reflected by the random noise. It can be understood that the first intermediate feature map is the feature map output by the convolutional layer in the previous layer of the target second fusion layer. The fusion implemented by this second fusion layer is a linear calculation or a non - linear calculation between the feature factor, the first intermediate feature map, and the random noise. It can be understood that the features reflected by the random noise can be features that do not affect the human face identity, such as hair, freckles, or beards.
[0129] Then, the second intermediate feature map output by the last second fusion layer and the 3D reconstruction feature map are input into the first fusion layer for fusion to obtain the second feature map. The fusion implemented by this first fusion layer is a linear calculation or a non - linear calculation between the second intermediate feature map output by the last second fusion layer and the 3D reconstruction feature map.
[0130] In this embodiment, by fusing the 3D reconstruction feature map including 3D spatial features of the corresponding scale with the last second intermediate feature map in the target generation network, that is, the 3D reconstruction feature map does not go through the convolutional operation in the target generation network and is directly fused, so that the 3D reconstruction feature map can better play the role of supervising the position of the entire human face and the distribution of facial features, and enables the second feature map output by the generation network to have targeted responses at positions with different geometric information and be fused with the corresponding 3D spatial features.
[0131] In some embodiments, the feature factor includes a scaling factor and a bias factor.
[0132] In this embodiment, step S2332 specifically includes:
[0133] The second intermediate feature map output by the target second fusion layer is calculated using the following formula,
[0134] y ij =y (s,i) *(T ij +B i )+y (b,i) ;
[0135] where i is the label of the target generation network, 1 ≤ i ≤ N, N is the number of generation networks, j is the label of the target second fusion layer, 1 ≤ j ≤ M, M is the number of second fusion layers in the target generation network, T ijis the first intermediate feature map input to the second fusion layer of the input target, Bi is random noise, y(s,i) is the scaling factor i, and y(b,i) is the bias factor i.
[0136] That is, for each first intermediate feature map Tij output by the convolutional layer, the above formula is used to fuse it with random noise to optimize the details of the first intermediate feature map.
[0137] In this embodiment, step S2333 specifically includes:
[0138] The second feature map is calculated using the following formula;
[0139] X i+1 = y ij * h i , j = M
[0140] where y ij is the second intermediate feature map output by the last second fusion layer, h i is the three-dimensional reconstruction feature map, and X i+1 is the second feature map.
[0141] In this embodiment, in the form of a product, the three-dimensional reconstruction feature map reflecting different feature scales is added to the second feature map, so that the added second feature map has a targeted response at positions with different geometric information.
[0142] In some embodiments, the difference between the first face image and the second face image includes the structural difference between the first face image and the second face image.
[0143] Please refer to Figure 6 , before step S24, it further includes:
[0144] S31: Calculate the brightness similarity value, contrast similarity value, and structural similarity value between the first face image and the second face image.
[0145] S32: Use the product of the brightness similarity value, contrast similarity value, and structural similarity value as the structural difference.
[0146] Among them, the brightness similarity value is an index used to reflect the similarity degree between the brightness of two images. The contrast similarity value is an index used to reflect the similarity degree between the contrasts of two images. The structural similarity value can be the Structural SIMilarity (SSIM) index, and the SSIM index is an index used to measure the similarity degree of two images. The brightness similarity value r(Y,Y’), contrast similarity value c(Y,Y’), and structural similarity value s(Y,Y’) of the first face image and the second face image can be calculated according to the following formula.
[0147]
[0148]
[0149]
[0150] Among them, Y represents the first face image, Y' represents the second face image, and μ Y represents the mean pixel value of the first face image, and μ Y' represents the mean pixel value of the second face image, and σ Y represents the variance value of the pixel values of the first face image, and σ Y' represents the variance value of the pixel values of the second face image;
[0151] Among them, σ YY' represents the covariance value between the first face image and the second face image, and c1 = (k1, *LO) 2 , c2 = (k2, *LO) 2 , where k1 and k2 represent preset constants. For example, k1 and k2 can be 0.01 and 0.03 respectively, and LO is the range of pixel values, which can usually take the value of 255.
[0152] Furthermore, the above-mentioned structural difference Ls(Y, Y') is obtained according to the following formula:
[0153] Ls(Y, Y') = r(Y, Y') * c(Y, Y') * s(Y, Y')
[0154] In this embodiment, through the above method, the difference between the first face image and the second face image includes the structural difference between the first face image and the second face image. Thus, the second face image is closer to the first face image in terms of structure (brightness, contrast, and SSIM), that is, the feature latent code can better restore the first face image in terms of structure (brightness, contrast, and SSIM), that is, the feature latent code output by the converged feature latent code extraction network can also accurately reflect the structural features (brightness, contrast, and SSIM) of the first face image. Furthermore, in terms of global features, the face image can be better restored, improving the restoration degree of the face image.
[0155] In some embodiments, the difference between the first face image and the second face image includes the pixel difference between the first face image and the second face image.
[0156] Please refer to Figure 7 , and before step S24, it further includes:
[0157] S33: Calculate the pixel difference between the pixel points at the target position in the first face image and the pixel points at the target position in the second face image to obtain the pixel difference corresponding to the target position, where the target position is any position in the first face image or the second face image.
[0158] S34: Take the sum of the pixel differences corresponding to each position in the first face image and the second face image as the pixel difference.
[0159] Specifically, the calculation formula for the pixel difference Lp(Y, Y’) between the first face image and the second face image is as follows:
[0160]
[0161] where Y i,j represents the pixel value of the pixel point with coordinates (i, j) in the first face image Y, and Y i,j ’ represents the pixel value of the pixel point with coordinates (i, j) in the second face image Y’, and P*Q represents the resolution size of the first face image and the second face image. For example, P can be 1024 and Q can also be 1024, so the resolution of the first face image and the second face image is 1024*1024.
[0162] In this embodiment, through the above method, the difference between the first face image and the second face image includes the pixel difference between the first face image and the second face image, making the second face image closer to the first face image in terms of pixel values, that is, enabling the feature latent code to better restore the first face image in terms of pixel values, that is, enabling the feature latent code output by the converged feature latent code extraction network to also reflect the pixel features of the first face image, and thus being able to better restore the face image in terms of local features and improving the restoration degree of the face image.
[0163] In some embodiments, the difference between the first face image and the second face image includes the 3D reconstruction parameter difference between the first face image and the second face image. That is, the difference between the first 3D reconstruction parameter of the first face image and the second 3D reconstruction parameter of the second face image.
[0164] Before step S24, it further includes:
[0165] S35: Calculate the parameter difference between the first 3D reconstruction parameter and the second 3D reconstruction parameter, and take the parameter difference as the 3D reconstruction parameter difference.
[0166] where the first 3D reconstruction parameter is the 3D reconstruction parameter obtained by performing 3D face reconstruction on the face in the first face image, and the second 3D reconstruction parameter is the 3D reconstruction parameter obtained by performing 3D face reconstruction on the face in the second face image.
[0167] Specifically, the calculation formula for the 3D reconstruction parameter difference Lv(Y, Y') between the first face image and the second face image is as follows:
[0168]
[0169] Wherein, Vi represents the i-th parameter in the 3D reconstruction parameters corresponding to the first face image, Vi' represents the i-th parameter in the 3D reconstruction parameters corresponding to the second face image, and G represents the structural size of the 3D reconstruction parameters of the first face image and the second face image. For example, G can be 398.
[0170] In this embodiment, through the above method, the difference between the first face image and the second face image includes the 3D reconstruction parameter difference between the first face image and the second face image, making the second face image closer to the first face image in terms of 3D geometric features. That is, the feature latent code can better restore the first face image in terms of 3D geometric features, and the feature latent code output by the converged feature latent code extraction network can also reflect the 3D geometric features of the first face image. Furthermore, the face image can be better restored in terms of depth features, improving the restoration degree of the face image.
[0171] In some embodiments, the total difference between the first face image and the second face image can be calculated according to the following formula:
[0172] L = αLs(Y, Y') + βLp(Y, Y') + γLv(Y, Y')
[0173] Wherein, α, β, and γ are the weights of the structural difference Ls, the pixel difference Lp, and the 3D reconstruction parameter difference, respectively.
[0174] It can be understood that it is equivalent to constructing a multi-dimensional similarity loss function. This multi-dimensional similarity loss function includes the structural difference reflecting the global difference, the pixel difference reflecting the local difference, and the 3D reconstruction parameter difference reflecting the depth. Thus, the texture features (texture depth and texture color) of the second face image are closer to those of the first face image, and the feature latent code output by the converged feature latent code extraction network can reflect the texture features of the first face image, enabling better restoration of the first face image in terms of global features, local features, and depth features, and improving the restoration degree.
[0175] An embodiment of the present application provides a computer device. Please refer to Figure 8 , which is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. Specifically, as Figure 8 shown, the computer device 10 includes at least one processor 11 and a memory 12 connected by communication ( Figure 8connected by a bus, taking one processor as an example).
[0176] Among them, the processor 11 is used to provide computing and control capabilities to control the computer device 10 to execute corresponding tasks. For example, it controls the computer device 10 to execute the feature latent code extraction method of any feature in the above method embodiments.
[0177] It can be understood that the processor 11 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0178] The memory 12, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the feature latent code extraction method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory 12, the processor 11 can implement the feature latent code extraction method in any of the following method embodiments. Specifically, the memory 12 can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 12 can also include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and their combinations.
[0179] An embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program. The computer program includes program instructions, and when the program instructions are executed by the processor, the processor is caused to execute each step in the above method embodiments.
[0180] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for extracting characteristic latent codes, characterized in that, it includes: Obtain a first face image for which characteristic latent codes are to be extracted and a three-dimensional reconstruction image corresponding to the first face image, where the three-dimensional reconstruction image includes three-dimensional spatial features of the face in the first face image; Input the first face image into a preset characteristic latent code extraction network to obtain the characteristic latent codes output by the characteristic latent code extraction network; Input the characteristic latent codes and the three-dimensional reconstruction image into a style adversarial generation network, and the style adversarial generation network performs deconvolution on the characteristic latent codes to obtain characteristic maps of different sizes, and fuses the three-dimensional reconstruction image with at least one of the characteristic maps to generate a second face image fused with the three-dimensional spatial features; According to the difference between the first face image and the second face image, iteratively adjust the parameters of the characteristic latent code extraction network until the characteristic latent code extraction network converges; Use the characteristic latent codes output by the converged characteristic latent code extraction network as the characteristic latent codes of the first face image.
2. The method according to claim 1, characterized in that, the obtaining the three-dimensional reconstruction image corresponding to the first face image includes: Perform three-dimensional face reconstruction on the face in the first face image to obtain first three-dimensional reconstruction parameters; Perform two-dimensional rendering on the first three-dimensional reconstruction parameters to obtain the three-dimensional reconstruction image corresponding to the first face image.
3. The method according to claim 1, characterized in that, the style adversarial generation network includes a mapping network and a plurality of sequentially arranged generation networks, and the plurality of generation networks are respectively used to output characteristic maps of different sizes. Among them, the target generation network is used to generate and output a second characteristic map according to the input first characteristic map, and the size of the second characteristic map is larger than the size of the first characteristic map. The target generation network is any one of the plurality of sequentially arranged generation networks; the inputting the characteristic latent codes and the three-dimensional reconstruction image into the style adversarial generation network, and the style adversarial generation network performing deconvolution on the characteristic latent codes to obtain characteristic maps of different sizes, and fusing the three-dimensional reconstruction image with at least one of the characteristic maps to generate a second face image fused with the three-dimensional spatial features includes: Input the characteristic latent codes into the mapping network to perform feature decoupling on the characteristic latent codes to generate an intermediate vector; According to the target size, perform feature extraction on the three-dimensional reconstruction image to obtain a three-dimensional reconstruction feature map, where the target size is the size of the second characteristic map, and the size of the three-dimensional reconstruction feature map is the target size; Input the intermediate vector, the first characteristic map, the three-dimensional reconstruction feature map, and random noise into the target generation network for fusion to output the second characteristic map; Determine the second characteristic map output by the last generation network among the plurality of sequentially arranged generation networks as the second face image.
4. The method according to claim 3, characterized in that, the target generation network includes at least one convolutional layer and a first fusion layer, and a second fusion layer is configured after each convolutional layer, Inputting the intermediate vector, the first feature map, the three-dimensional reconstruction feature map, and the random noise into the target generation network for fusion to output the second feature map includes: Inputting the intermediate vector into each second fusion layer, and each second fusion layer performs an affine transformation on the intermediate vector according to the target size to obtain a feature factor, where the feature factor is adapted to the target size; The target second fusion layer uses the feature factor to fuse the first intermediate feature map input to the target second fusion layer with the random noise to obtain the second intermediate feature map output by the target second fusion layer, where the target second fusion layer is any one of the second fusion layers, and the first intermediate feature map is the feature map output by the convolutional layer in the previous layer of the target second fusion layer; Inputting the second intermediate feature map output by the last second fusion layer and the three-dimensional reconstruction feature map into the first fusion layer for fusion to obtain the second feature map.
5. The method according to claim 4, wherein, the feature factor includes a scaling factor and a bias factor; The target second fusion layer uses the feature factor to fuse the first intermediate feature map input to the target second fusion layer with the random noise to obtain the second intermediate feature map output by the target second fusion layer, including: Calculating the second intermediate feature map output by the target second fusion layer using the following formula, y ij = y (s,i) *(T ij + B i ) + y (b,i) ; where i is the label of the target generation network, 1 ≤ i ≤ N, N is the number of generation networks, j is the label of the target second fusion layer, 1 ≤ j ≤ M, M is the number of second fusion layers in the target generation network, and T ij is the first intermediate feature map input to the target second fusion layer, Bi is the random noise, y(s,i) is the scaling factor, and y(b,i) is the bias factor; Inputting the second intermediate feature map output by the last second fusion layer and the three-dimensional reconstruction feature map into the first fusion layer for fusion to obtain the second feature map, including: Calculating the second feature map using the following formula; X i+1 = y ij * h i , j = M Among them, y ij is the second intermediate feature map output by the last second fusion layer, h i is the three-dimensional reconstruction feature map, X i+1 is the second feature map.
6. The method according to any one of claims 1-5, wherein, the difference between the first face image and the second face image includes the structural difference between the first face image and the second face image; Before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, further including: Calculating the brightness similarity value, the contrast similarity value, and the structural similarity value between the first face image and the second face image; Taking the product of the brightness similarity value, the contrast similarity value, and the structural similarity value as the structural difference.
7. The method according to any one of claims 1-5, wherein, the difference between the first face image and the second face image includes the pixel difference between the first face image and the second face image; Before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, further including: Calculating the pixel difference between the pixel points at the target position in the first face image and the pixel points at the target position in the second face image to obtain the pixel difference corresponding to the target position, where the target position is any position in the first face image or the second face image; The sum of the pixel differences corresponding to each position in the first face image and the second face image is used as the pixel difference.
8. The method according to any one of claims 1-5, wherein, the difference between the first face image and the second face image includes the difference in three-dimensional reconstruction parameters between the first face image and the second face image; before iteratively adjusting the parameters of the feature latent code extraction network according to the difference between the first face image and the second face image, further comprising: calculating the parameter difference between the first three-dimensional reconstruction parameter and the second three-dimensional reconstruction parameter, and using the parameter difference as the difference in three-dimensional reconstruction parameters, where the first three-dimensional reconstruction parameter is the three-dimensional reconstruction parameter obtained by performing three-dimensional face reconstruction on the face in the first face image, and the second three-dimensional reconstruction parameter is the three-dimensional reconstruction parameter obtained by performing three-dimensional face reconstruction on the face in the second face image.
9. A computer device, wherein, comprising a memory and one or more processors, the one or more processors are configured to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, the computer device implements the method according to any one of claims 1-8.
10. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 1-8.
Citation Information
Patent Citations
Feature latent code extraction method and device, equipment and storage medium
CN113077379A
All-head texture network structure based on single face image and generation method
CN113095149A