Face posture editing method in double-submerged space
Through the dual latent space method, combining the feature pyramid and full connection layer to extract the latent vector, the problems of face distortion and attribute loss in face pose editing are solved, high-quality face pose editing is achieved, and the flexibility and accuracy of editing are enhanced.
Patent Information
- Application Number
- CN202510588528.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-19
AI Technical Summary
The existing facial posture editing methods are prone to face distortion and loss of attributes during larger posture conversion, and are insufficient in terms of diversity and flexibility.
The dual latent space method is used to extract face expressions, postures and contour parameters through ResNet-50 and fully connected layer, and additional latent vectors are extracted in combination with feature pyramids and fully connected layer. The edited face image is generated using the inversion model, and identity loss, face loss and regularization loss are trained to ensure the consistency and quality of face features.
It effectively reduces the loss of facial attributes and improves the accuracy and flexibility of facial posture editing. The generated images naturally present the required posture while maintaining the original identity characteristics, improving the diversity and reality of the editing.
Smart Images

Figure CN120510341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of digital image processing and relates to a facial posture editing method in a dual-latent space. Background Art
[0002] Face pose editing manipulates the pose of a face image, generating a new face with the desired pose while preserving other details. As a current research hotspot in artificial intelligence, face pose editing is widely used in fields such as face recognition and video editing. This has led to a growing interest in face pose editing among researchers. For example, face pose editing can not only help straighten faces, improving face recognition efficiency, but also assist in editing face videos, enhancing the fidelity of video editing.
[0003] Facial pose editing, as a key method for generating and controlling virtual avatars, plays an important role in building "real-time, anthropomorphic" virtual interactive experiences. Current facial pose editing can be divided into two stages, primarily based on StyleGAN. Prior to StyleGAN, facial pose editing models were typically based on generative adversarial networks (GANs) and typically required training using multiple datasets to address inherent limitations. For example, in FNM, to improve face editing results under various conditions, the authors used the Multi-PIE dataset, a large-scale, controlled dataset commonly used to evaluate face recognition in constrained scenes with varying poses, lighting, and expressions. Because the facial poses in this dataset are derived from 15 fixed camera angles, further validation of its performance under various facial poses using other unconstrained datasets, such as CASIA-WebFace, MS-Celeb1M, and LFW, is often necessary. In addition, before the emergence of StyleGAN, traditional facial pose editing methods usually directly synthesized the edited face through a given editing purpose (such as face frontalization) during actual editing, which greatly limited the diversity of facial pose editing.
[0004] The emergence of StyleGAN has advanced solutions to these problems and, to a certain extent, enabled flexible control of facial attributes, including pose. Subsequently, a generative network based on a dual latent space was proposed, which achieved a better separation between content and style, enabling the specific modification of single attributes, including pose, by manipulating the value of the latent vector. However, when editing faces with large poses, problems such as facial distortion and loss of facial attributes often occur, affecting the editing results.
[0005] In summary, face pose editing still faces several pressing challenges. For example, accurately modeling complex transformations on input facial images is extremely challenging, as it requires maintaining the input identity while making large, yet accurate, changes to facial features and head shape. However, existing methods for pose conversion have fallen short of expectations. This is partly because existing methods suffer from face distortion when facing large pose transformations, and partly because existing methods inevitably alter head shape during the conversion process. Summary of the Invention
[0006] In view of this, an object of the present invention is to provide a facial posture editing method in a dual latent space.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] A facial posture editing method in a dual-latent space, the method comprising:
[0009] S1. Obtain a facial image dataset, divide the dataset and perform preprocessing operations;
[0010] S2, extracting the expression, posture and contour parameters of the original face and the target face, replacing the posture parameters of the original face with the posture parameters of the target face, and performing rough face reconstruction;
[0011] S3, extracting additional latent vectors from the concatenated features of the original face and the coarsely reconstructed face through feature pyramid and fully connected layers;
[0012] S4, inverting the original face into a latent vector through the inversion model, combining the extracted additional latent vector to obtain the actual edited latent vector, and generating the edited face image according to the latent vector through the generator;
[0013] S5. Use the sum of identity loss, face loss and regularization loss as the loss function to train the above process until the model converges.
[0014] Furthermore, in step S1, data set partitioning refers to dividing the data set into a training set and a test set according to a preset ratio;
[0015] The preprocessing operation includes: rotating and mirror symmetric processing of the face image, and normalizing the processed face image.
[0016] Further, in step S2, it includes the following steps:
[0017] S21. Use ResNet-50 network to extract the original face x and the target face x t The expression, posture and contour parameters are obtained by vector v and vt ; Use two fully connected layers and an activation function to obtain the actual face parameters:
[0018] V1=Linear2(ReLU(Linear1(v)))
[0019] V2=Linear2(ReLU(Linear1(v t )))
[0020] Among them, Linear1(·) and Linear2(·) represent two fully connected layers. Linear1(·) is used to scale the vector size, and Linear2(·) is used to restore the vector size. ReLU(·) represents the activation function.
[0021] S22. Extract the expression θ, posture β, and contour parameter ψ of the original face from vector V1, and extract the expression θ′, posture β′, and contour parameter ψ′ of the target face from vector V2, and form vectors R1 and R2 for rough reconstruction of the face respectively:
[0022] R1=(θ,β,ψ)
[0023] R2=(θ′,β′,ψ′)
[0024] Substitute the posture parameters in R2 into the new face parameter vector R3 obtained in R1:
[0025] R3=(θ,β′,ψ)
[0026] S23, the facial parameter vector R3 is passed through a 3D face reconstruction model to obtain a rough reconstructed face x r .
[0027] Further, in step S3, it includes the following steps:
[0028] S31, the original face x and the coarse reconstructed face x r To perform the fusion:
[0029] x c =Concat(x,x r )
[0030] Among them, Concat(·) represents the fusion operation, x C represents the fused image;
[0031] S32, extracting a multi-layer feature map of the fused image through multiple bottleneck layers;
[0032] S33, select the feature maps of the 6th, 20th and 23rd layers of the multi-layer feature map, respectively construct two feature pyramids with the same structure but independent parameters, and extract two sets of potential vectors V through the fully connected layeres and V cs .
[0033] Furthermore, in step S32, the extraction process of the multi-layer feature map is expressed as:
[0034] T1 i =BN(Conv1(T i ))
[0035] T2 i =BN(Conv3(ReLU(Conv3(BN(T i )))))
[0036] N=Sigmoid(Fc2(ReLU(Fc1(Avg(T2 i )))))
[0037] T3 i =N·T2 i
[0038] T i+1 =T1 i +T3 i
[0039] Among them, T i Represents the feature map of the i-th layer; if i is 1, then T i Represented as x c ;T1 i Indicates T i The first intermediate result after feature extraction, T2 i Indicates T i The second intermediate result after feature extraction, T3 i Indicates T2 i The result obtained after the attention mechanism, N represents the attention weight matrix; T i+1 Represents the feature map of the i+1th layer; Conv3(·) represents convolution and convolution operations of size 3; Avg(·) represents the average pooling operation on the channel; Fc1(·) represents the fully connected layer that compresses the number of channels; Fc2(·) represents the fully connected layer that restores the number of channels to the pre-compression state; Sigmoid(·) represents the activation function;
[0040] Furthermore, in step S33, the steps of extracting the latent vector from the multi-layer feature map are as follows:
[0041] Top-down process of feature pyramid: the new 23rd layer feature map is consistent with the original 23rd layer feature map
[0042] T2′3=T 23
[0043] Perform residual connection operation:
[0044] T' 20 =T 20 +Conv1(T' 23 )
[0045] T'6 = T6 + Conv1 (T' 20 )
[0046] When the feature map after convolution, that is, Temp i When the size is not 1, do the following:
[0047] Extract features from the feature maps of layers 6, 20, and 23 and reduce their size:
[0048] Temp=LeakyReLU(Conv3(T') i ))),i=6,20,23
[0049] Continue to extract features and reduce the feature map size through convolution operations with a convolution kernel size of 3×3 until Temp i Size 1:
[0050] Temp=LeakyReLU(Conv3(Temp)))
[0051] When Temp i When the size is 1, perform the following operations to get the eigenvector:
[0052] v i =LeakyReLU(Linear(Temp))
[0053]
[0054]
[0055] In the above process, T′ i represents the i-th layer feature map used for feature extraction, and i is 6, 20, 23; Linear(·) represents the full convolution layer; LeakyReLU(·) represents the activation function; V es represents the extracted additional style latent vector; V cs Represents the extracted additional content latent vector.
[0056] Furthermore, in step S4, the original face image is inverted into the latent space using the inversion network to obtain its style latent vector V S and content latent vector V C , and are respectively combined with the corresponding additional latent vector V es and V cs Add together to get the edited latent vector:
[0057] V edi_c =V c +V ex
[0058] V edi_s =V s +V es
[0059] Then, the edited face image y is obtained through the generative model:
[0060]
[0061] Among them, V edi_c Represents the edited content latent vector, V edi_s represents the edited style latent vector; G(·) represents the generative model.
[0062] Furthermore, in step S5, a loss function including identity loss, face loss and regularization loss is used for training. The loss function is specifically expressed as:
[0063] L total =L face +L ID +L reg
[0064] Extract expression θ from the final generated face image y y , posture β y and the contour ψ y Parameters, get the vector group representation R4:
[0065] R4=(θ y ,β y ,ψ y )
[0066] Then the face loss L face Expressed as:
[0067]
[0068] Where MSE(·) represents the mean square error, represents the i-th value of vector R3, represents the i-th value of vector R3, and n represents the length of the vector;
[0069] Identity loss L ID It is used to measure the identity change of the face before and after editing, which is expressed as:
[0070] L ID =w(Δ pose )·(1-Similarity(R(x),R(y)))
[0071] Among them, Δ pose represents the pose difference between the two images, w(·) represents the pose weight, R(·) is the pre-trained ArcFace network, and Similarity(·) represents the similarity function; there are:
[0072] w(Δ pose )=1-Δ pose
[0073]
[0074]
[0075] Among them, α s ,α t Denote the pitch, β and s ,β t Denote the heading in the deflection angle, γ s ,γ t Represents the roll in the deflection angle; A, B are two different face images, A i represents the i-th pixel value of A, B i represents the i-th pixel value of B;
[0076] Regularization loss is used to improve the quality of generated images and control the generated faces not to deviate too much from the real faces. It is expressed as:
[0077]
[0078] Among them, lantent i represents the potential vector obtained after the i-th image passes through the inversion network, lantent avg Represents the calculated average latent vector of the faces in the dataset.
[0079] The beneficial effects of the present invention are:
[0080] First, the present invention demonstrates superior performance in addressing the issues of facial distortion and attribute loss in larger facial poses. Traditional methods often suffer from severe facial distortion and attribute loss when faced with large pose transitions. The present invention successfully overcomes this problem by designing a dual-latent space facial pose editing network and a method for guiding facial pose editing through coarse facial reconstruction. During the coarse facial reconstruction phase, ResNet-50 and fully connected layers are used to extract key parameters, and coarse reconstruction is performed in conjunction with the target facial pose parameters. This provides accurate guidance for features such as contours for subsequent editing, effectively maintaining the consistency of facial contours and other features before and after editing. This reduces the loss of facial attributes during the editing process, avoids facial distortion caused by large pose changes, and ensures that the edited facial image retains its original identity while naturally and realistically presenting the desired pose.
[0081] Secondly, the present invention further enriches the information required for face pose editing through the additional latent vector extraction stage. After splicing the original face image and the coarsely reconstructed face, a module combining a feature pyramid and a fully connected layer is used to extract additional latent vectors. These latent vectors contain rich information for guiding face pose editing, providing strong support for generating high-quality edited face images. In the face pose editing stage, the original face is inverted into a latent vector and combined with the additional latent vector. The edited face image is generated through a generator, so that the generated face image can better integrate the features of the original face and the requirements of the target pose, enhancing the flexibility and accuracy of editing, and can generate a face pose image that better meets the expected requirements.
[0082] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0084] Figure 1 Schematic diagram of the overall network structure of a facial posture editing method in a dual-latent space according to an embodiment of the present invention;
[0085] Figure 2 Schematic diagram of the network structure for coarse face reconstruction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0086] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0087] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0088] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0089] See also Figures 1 and 2 , which is a face pose editing method in dual latent space.
[0090] Specifically, this embodiment first provides a detailed process of a face pose editing method in a dual latent space, such as Figure 1 As shown, it specifically includes the following steps:
[0091] S1: Data preparation stage: obtain a high-quality face image dataset, divide the dataset into training set and test set, and perform corresponding preprocessing;
[0092] S2: Coarse face reconstruction stage: ResNet-50 and fully connected layers are combined to extract expression, posture, and contour parameters from the face pose. The posture parameters are replaced with the posture parameters of the target face image. These sets of parameters are then used to coarsely reconstruct a face image after the pose change through the 3D face reconstruction model. This image can provide guidance for subsequent face pose editing and maintain the consistency of facial contours and other features before and after editing.
[0093] S3: Additional latent vector extraction stage: The original face image and the coarsely reconstructed face are spliced together, and a module combining a feature pyramid and a fully connected layer is designed to extract the additional latent vector required for face editing. This latent vector contains rich information to guide face pose editing.
[0094] S4: Face pose editing stage: The original face is inverted into a latent vector through the model, and combined with the additional latent vector obtained in S3 to obtain the actual edited latent vector. The latent vector is then generated by the generator to obtain the edited face image;
[0095] S5: Model training phase: The sum of identity loss, face loss, and regularization loss is used as the loss function. The contour and expression parameters of the original face image and the posture parameters of the target face are extracted and compared with the generated result. The loss is calculated and the model parameters are trained through backpropagation until the model parameters converge.
[0096] In step S1 of this embodiment, pre-processing the face image specifically includes the following steps:
[0097] S11: Performing rotation, mirror symmetry, and other processing on the face image.
[0098] S12: Performing a normalization operation on the processed face image.
[0099] In step S2 of this embodiment, as Figure 2 As shown, the specific steps include:
[0100] S21: ResNet-50 is used to extract various information of the face image, and the result is a vector v of size 1×2048; here, the original face x and the target face x are used. t Various information, get vectors v and v t ; Use two fully connected layers and an activation function to obtain the actual face parameters:
[0101] V1=Linear2(ReLU(Linear1(v))) (1)
[0102] V2=Linear2(ReLU(Linear1(v t ))) (2)
[0103] Linear1(·) and Linear2(·) represent two fully connected layers, and ReLU(·) represents the activation function. Linear1(·) converts a 1×2048 vector into a 1×1024 vector. Linear2(·) converts it back to 1×2048.
[0104] S22: Extract the required expression, posture, and contour parameters from the V1 vector and V2 vector obtained in step S21 as the subsequent coarse reconstruction data of the face, and name them as θ, β, ψ respectively; the vector used for coarse reconstruction of the face is represented by R:
[0105] R1=(θ,β,ψ) (3)
[0106] R2=(θ′,β′,ψ′) (4)
[0107] R3 = (θ, β′, ψ) (5) where R1 represents the expression, posture, and contour parameters extracted from the original face image. R2 represents the parameters extracted from the target face. R3 represents the new vector group obtained by replacing the posture parameters in R2 with those in R1, which contains the expression and contour information of the original face and the posture information of the target face.
[0108] S23: The face parameter vector R3 obtained in step S22 is passed through a 3D face reconstruction model to obtain the required rough reconstructed face x r .
[0109] In step S3 of this embodiment, it specifically includes the following steps:
[0110] S31: Fusing the original face image with the roughly reconstructed face image obtained in step S23 to achieve information exchange. The specific steps are as follows:
[0111] x c =Concat(x,x r ) (6)
[0112] Among them, x represents the original face image, x r represents the rough reconstruction of the face, Concat(·) represents the fusion operation, x C Represents the fused image.
[0113] S32: Use multiple bottleneck layers to extract multi-layer feature maps of the fused image. The steps of extracting feature maps by the bottleneck layer are as follows:
[0114] T1 i =BN(Conv1(T i )) (7)
[0115] T2 i =BN(Conv3(ReLU(Conv3(BN(T i ))))) (8)
[0116] N=Sigmoid(Fc2(ReLU(Fc1(Avg(T2 i ))))) (9)
[0117] T3 i =N·T2 i (10)
[0118] T i+1 =T1 i +T3 i (11) Where, T i Represents the feature map of the i-th layer; if i is 1, then T i Represented as x c ;T1 i Indicates T i The first intermediate result after feature extraction, T2 i Indicates T i The second intermediate result after feature extraction, T3 i Indicates T2 i The result obtained after the attention mechanism, N represents the attention weight matrix; T i+1 Represents the feature map of the i+1th layer; Conv3(·) represents convolution and convolution operations with size 3, stride 1, and padding 1; Avg(·) represents the average pooling operation on the channel; Fc1(·) represents a fully connected layer, and the number of channels will be compressed; Fc2(·) represents a fully connected layer, and the number of channels will be restored to the original state; Sigmoid(·) represents the activation function.
[0119] S33: Select the feature maps of the 6th, 20th and 23rd layers in step S32, respectively, and construct two feature pyramids with the same structure but independent parameters, and extract two sets of latent vectors through the fully connected layer. The steps for extracting latent vectors from feature maps are as follows:
[0120] Top-down process of feature pyramid: the new 23rd layer feature map is consistent with the original 23rd layer feature map
[0121] T' 23 =T 23
[0122] Perform residual connection operation:
[0123] T' 20 =T 20 +Conv1(T' 23 )
[0124] T'6 = T6 + Conv1 (T' 20 )
[0125] When the feature map after convolution, that is, Temp i When the size is not 1, do the following:
[0126] Extract features from the feature maps of layers 6, 20, and 23 and reduce their size:
[0127] Temp=LeakyReLU(Conv3(T i ′))),i=6,20,23
[0128] Continue to extract features and reduce the feature map size through convolution operations with a convolution kernel size of 3×3 until Temp i Size 1:
[0129] Temp=LeakyReLU(Conv3(Temp)))
[0130] When Temp i When the size is 1, perform the following operations to get the eigenvector:
[0131] v i =LeakyReLU(Linear(Temp)) (17)
[0132]
[0133]
[0134] In the above process, T′ i represents the i-th layer feature map used for feature extraction, and i is 6, 20, 23; Linear(·) represents the full convolution layer; LeakyReLU represents the activation function; V es represents the extracted additional style latent vector; V ec Represents the extracted additional content latent vector.
[0135] In step S4 of this embodiment, it specifically includes the following steps:
[0136] S41: First, use the inversion network to invert the original face image into the latent space to obtain its style latent vector V S and content latent vector V C , add it to the additional latent vector obtained in step S33 to obtain the edited latent vector, and finally obtain the edited face image through the generative model. The specific steps are as follows:
[0137] V edi_c =V c +V ec (20)
[0138] V edi_s =V s +V es (twenty one)
[0139]
[0140] Among them, V edi_c Represents the edited content latent vector, V edi_s represents the edited style latent vector; G(·) represents the generative model, and y represents the edited face image.
[0141] In step S5 of this embodiment, it specifically includes the following steps:
[0142] The loss function contains the identity loss L ID , face loss L face and regularization loss L reg The formula is as follows:
[0143] L total =L face +L ID +L reg (twenty three)
[0144] The modules in step S21 are used to extract the parameters of the generated result graph and intercept the expression, posture and contour parameters. The formula is as follows:
[0145] R3=(θ,β′,ψ) (24)
[0146] R4=(θ y ,β y ,ψ y ) (25)
[0147] L face =MSE(R3,R4) (26)
[0148]
[0149] Among them, R4 represents the facial expression, posture and contour parameters extracted from the generated image; MSE represents the mean square error, represents the i-th value of vector R3, Represents the i-th value of vector R3, and n represents the length of the vector.
[0150] Identity loss is used to measure the identity change of the face before and after editing. The formula is as follows:
[0151] L ID =w(Δ pose )·(1-Similarity(R(x),R(y))) (28)
[0152]
[0153]
[0154] w(Δ pose )=1-Δ pose(31)
[0155] Among them, α s ,α t Denote the pitch, β and s ,β t Denote the heading in the deflection angle, γ s ,γ t Represents the roll in the deflection angle; A, B are two different face images, A i represents the i-th pixel value of A, B i represents the i-th pixel value of B. In this embodiment, Δ pose The maximum value is 0.5.
[0156] Regularization loss is used to improve the quality of generated images and to control the generated faces from deviating too much from the real faces. Its formula is as follows:
[0157]
[0158] Among them, latent i Represents the potential vector obtained after the i-th image passes through the inversion network, latent avg Represents the calculated average latent vector of the faces in the dataset.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A facial pose editing method in a dual-latent space, characterized by: The method comprises: S1. Obtain a facial image dataset, divide the dataset and perform preprocessing operations; S2, extracting the expression, posture and contour parameters of the original face and the target face, replacing the posture parameters of the original face with the posture parameters of the target face, and performing rough face reconstruction; S3, extracting additional latent vectors from the concatenated features of the original face and the coarsely reconstructed face through feature pyramid and fully connected layers; S4, inverting the original face into a latent vector through the inversion model, combining the extracted additional latent vector to obtain the actual edited latent vector, and generating the edited face image according to the latent vector through the generator; S5. Use the sum of identity loss, face loss and regularization loss as the loss function to train the above process until the model converges.
2. The facial pose editing method in a dual-latent space according to claim 1, characterized in that: In step S1, data set division refers to dividing the data set into a training set and a test set according to a preset ratio; The preprocessing operation includes: rotating and mirror symmetric processing of the face image, and normalizing the processed face image.
3. The facial pose editing method in a dual-latent space according to claim 1, characterized in that: In step S2, it includes the following steps: S21. Use ResNet-50 network to extract the original face x and the target face x t The expression, posture and contour parameters are obtained by vector v and v t ; Use two fully connected layers and an activation function to obtain the actual face parameters: V1=Linear2(ReLU(Linear1(v))) V2=Linear2(ReLU(Linear1(v t ))) Among them, Linear1(·) and Linear2(·) represent two fully connected layers. Linear1(·) is used to scale the vector size, and Linear2(·) is used to restore the vector size. ReLU(·) represents the activation function. S22. Extract the expression θ, posture β, and contour parameter ψ of the original face from vector V1, and extract the expression θ′, posture β′, and contour parameter ψ′ of the target face from vector V2, and form vectors R1 and R2 for rough reconstruction of the face respectively: R1=(θ,β,ψ) R2=(θ′,β′,ψ′) Substitute the posture parameters in R2 into the new face parameter vector R3 obtained in R1: R3=(θ,β′,ψ) S23, the facial parameter vector R3 is passed through a 3D face reconstruction model to obtain a rough reconstructed face x r .
4. The method for editing facial pose in a dual-latent space according to claim 3, wherein: In step S3, it includes the following steps: S31, the original face x and the coarse reconstructed face x r To perform the fusion: x c =Concat(x,x t ) Among them, Concat(·) represents the fusion operation, x C represents the fused image; S32, extracting a multi-layer feature map of the fused image through multiple bottleneck layers; S33, select the feature maps of the 6th, 20th and 23rd layers of the multi-layer feature map, respectively construct two feature pyramids with the same structure but independent parameters, and extract two sets of potential vectors V through the fully connected layer ed and V vd .
5. The method for editing facial pose in a dual-latent space according to claim 4, characterized in that: In step S32, the extraction process of the multi-layer feature map is expressed as: T1 i =BN(Conv1(T i )) T2 i =BN(Conv3(ReLU(Conv3(BN(T i ))))) N=Sigmoid(Fc2(ReLU(Fc1(Avg(T2 i ))))) <h2 style=";text-align:left;direction:ltr">T3<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (N·T2)<h2 style=";text-align:left;direction:ltr"> i T i+1 =T1 i +T3 i Among them, T i Represents the feature map of the i-th layer; if i is 1, then T i Represented as x c ;T1 i Indicates T i The first intermediate result after feature extraction, T2 i Indicates T i The second intermediate result after feature extraction, T3 i Indicates T2 i The result obtained after the attention mechanism, N represents the attention weight matrix; T i+1 Represents the feature map of the i+1th layer; Conv3(·) represents the convolution operation with a convolution kernel size of 3×3; BN(·) represents the batch normalization operation; Conv3(·) represents the convolution and convolution operations of size 3; Avg(·) represents the average pooling operation on the channel; Fc1(·) represents the fully connected layer that compresses the number of channels; Fc2(·) represents the fully connected layer that restores the number of channels to the pre-compression state; Sigmoid(·) represents the activation function.
6. The method for editing facial pose in a dual-latent space according to claim 5, characterized in that: In step S33, the steps of extracting the latent vector from the multi-layer feature map are as follows: The top-down process of the feature pyramid: the new 23rd layer feature map is consistent with the original 23rd layer feature map: T′ 23 =T 23 Perform residual connection operation: T′ 20 =T 20 +Conv1(T′ 23 ) T′6=T6+Conv1(T′ 20 ) When the feature map after convolution, that is, Temp i When the size is not 1, do the following: Extract features from the feature maps of layers 6, 20, and 23 and reduce their size: Temp=LeakyReLU(Conv3(T′ i ))),i=6,20,23 Continue to extract features and reduce the feature map size through convolution operations with a convolution kernel size of 3×3 until Temp i Size 1: Temp=LeakyReLU(Conv3(Temp))) When Temp i When the size is 1, perform the following operations to get the eigenvector: v i =LeakyReLU(Linear(Temp)) In the above process, T′ i represents the i-th layer feature map used for feature extraction, and i is 6, 20, 23; Linear(·) represents the full convolution layer; LeakyReLU(·) represents the activation function; V es represents the extracted additional style latent vector; V cs Represents the extracted additional content latent vector.
7. The method for editing facial pose in a dual-latent space according to claim 4, characterized in that: In step S4, the original face image is inverted into the latent space using the inversion network to obtain its style latent vector V S and content latent vector V C , and are respectively combined with the corresponding additional latent vector V es and V cs Add together to get the edited latent vector: V edi_c =V c +V ec V edi_s =V s +V es Then, the edited face image y is obtained through the generative model: Among them, V edi_c Represents the edited content latent vector, V edi_s represents the edited style latent vector; G(·) represents the generative model.
8. The method for editing facial pose in a dual-latent space according to claim 7, wherein: In step S5, a loss function including identity loss, face loss and regularization loss is used for training. The loss function is specifically expressed as: L total =L face +L ID +L reg Extract expression θ from the final generated face image y y , posture β y and the contour ψ y Parameters, get the vector group representation R4: R4=(θ y ,b y ,ψ y ) Then the face loss L face Expressed as: Where MSE(·) represents the mean square error, represents the i-th value of vector R3, represents the i-th value of vector R3, and n represents the length of the vector; Identity loss L ID It is used to measure the identity change of the face before and after editing, which is expressed as: L ID =w(Δ pose )·(1-Similarity(R(x),R(y))) Among them, Δ pose represents the pose difference between the two images, w(·) represents the pose weight, R(·) is the pre-trained ArcFace network, and Similarity(·) represents the similarity function; there are: w(D pose )=1-D pose Among them, α s ,α t Denote the pitch, β and s ,β t Denote the heading in the deflection angle, γ s ,γ t Represents the roll in the deflection angle; A, B are two different face images, A i represents the i-th pixel value of A, B i represents the i-th pixel value of B; Regularization loss is used to improve the quality of generated images and control the generated faces not to deviate too much from the real faces. It is expressed as: Among them, lantent i represents the potential vector obtained after the i-th image passes through the inversion network, lantent avg Represents the calculated average latent vector of the faces in the dataset.