Method and system for hairstyle attribute transfer based on deep adaptation

By combining monocular facial ranging and key point detection with the StyleGANv2 generator, the problem of hairstyle and facial distortion caused by facial depth misalignment was solved, enabling precise editing of hairstyle attributes and improving the compositing effect.

CN115660946BActive Publication Date: 2025-11-25FUZHOU UNIV ZHICHENG COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211343425.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-11-25
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing hairstyle attribute editing algorithms based on generative adversarial networks are prone to causing distortion of the hairstyle and face when the facial depth is misaligned, which affects the synthesis effect.

Method used

The depth difference of the input image is estimated by a monocular facial ranging algorithm, facial key points are detected, facial reference points and offsets are calculated, facial depth is aligned, and hairstyle attributes are edited in the latent feature space through a fast hairstyle attribute editing module and a StyleGANv2 generator to output the target image.

Benefits of technology

It enables precise editing of hairstyle attributes without retraining, avoiding distortion caused by misalignment of facial depth and improving the effect of hairstyle synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660946B_ABST
    Figure CN115660946B_ABST
Patent Text Reader

Abstract

The application provides a hairstyle attribute migration method and system based on deep adaptation, comprising the following steps: step S1: estimating the estimated depth difference of the input image through a face monocular distance measurement algorithm; step S2: detecting the face key points of the image to be adjusted through a face key point detection model RCPR; step S3: calculating the facial reference point and the offset through the depth difference and the face key points; step S4: aligning the face depth through the facial reference point and the offset; step S5: editing the hairstyle of the input image through a fast hairstyle attribute editing module, and outputting the latent code of the target image; and step S6: mapping the latent code of the target image to the image domain through a StyleGANv2 generator to obtain the target image. The application optimizes the feature latent code of the input image, and through the method of reconstructing the image through the pre-trained generation network, the precise hairstyle attribute editing effect can be achieved without retraining, thereby meeting the basic needs of users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and particularly relates to a hairstyle attribute migration method and system based on deep adaptation. BACKGROUND

[0002] As one of the important factors of facial attributes, hairstyle influences people's overall temperament to some extent. Different hairstyles can easily represent a person's age, gender, social class, cultural level, fashion preferences, etc., and are an important part of personalized modeling. Different hairstyles on the same person will bring different visual experiences to the onlookers, affecting people's temperament. With the advent of the artificial intelligence era, deep learning, machine vision and other frontier technologies have developed and matured, marking the arrival of the intelligent era. Deep learning technology is used in various fields to solve practical problems. In particular, the advent of generative adversarial networks has made breakthroughs in face attribute synthesis technology. However, there are still many problems with hairstyle attribute editing algorithms based on generative adversarial networks, such as the problem of hairstyle and face distortion caused by hairstyle editing when the face depth is not aligned. SUMMARY

[0003] In view of the defects and deficiencies of the prior art, the purpose of the present application is to provide a hairstyle attribute migration method and system based on deep adaptation, which takes into account the relationship between input images and avoids the mismatch problem of hairstyle and face features after fusion under the condition of face depth misalignment, and can further improve the hairstyle synthesis effect.

[0004] It mainly includes the following steps: step S1: estimating the estimated depth difference of the input image by a face monocular distance measurement algorithm; step S2: detecting the face key points of the image to be adjusted by a face key point detection model RCPR; step S3: calculating the face reference point and the offset by the depth difference and the face key points; step S4: aligning the face depth by the face reference point and the offset; step S5: editing the hairstyle of the input image by a fast hairstyle attribute editing module, and outputting the latent code of the target image; step S6: mapping the latent code of the target image to the image domain by a StyleGANv2 generator to obtain the target image. The present application optimizes the feature latent code of the input image, and can achieve accurate hairstyle attribute editing effect without retraining by using the method of reconstructing the image by the pre-trained generation network, thereby meeting the basic needs of users.

[0005] And the system obtained based on the above method: a user inputs a face image and a hairstyle reference image, the system calls a deep adaptive alignment module to align the depth of the input image, edits attribute features in the latent feature space of the image, and finally outputs a target image through a generation network, so as to obtain a target image containing the identity features of the face image and the hairstyle features of the hairstyle reference image.

[0006] The technical scheme adopted by the present application to solve its technical problems is:

[0007] A hairstyle attribute editing method based on deep adaptive alignment, characterized in that it comprises the following steps:

[0008] Step S1: estimating the estimated depth difference of input images {I1, I2} through a face monocular distance measurement algorithm

[0009] Step S2: detecting the face key points p of the image to be adjusted through a face key point detection model RCPR n ;

[0010] Step S3: calculating the face reference point c and the offset L=(x, y) through the depth difference and the face key points p n ;

[0011] Step S4: aligning the face depth through the face reference point c and the offset L=(x, y)

[0012] Step S5: editing the hairstyle of the input image through a fast hairstyle attribute editing module, and outputting the latent encoding C blend of the target image

[0013] Step S6: mapping the latent encoding C blend of the target image to the image domain through a StyleGANv2 generator to obtain the target image I.

[0014] Further, in step S1, the estimated depth difference of input images {I1, I2} is estimated through a face monocular distance measurement algorithm A pre-trained face detection model RFBnet is used to locate the face positions in I1 and I2, and the face frame size [W face ,H face ] is recorded

[0015] The estimated depth difference of input images {I1, I2} is estimated through a face monocular distance measurement algorithm Specifically, the depth value is estimated through the principle of triangulation, and according to the principle of similar triangles, In the formula, x is the face imaging width, w is the actual face width, and d is the camera focal length. The face depth estimation value is obtained by solving the equation:

[0016]

[0017] After the difference, the estimated depth difference of the input image is obtained as

[0018] Further, in step S2, 68 facial key points p of the image to be adjusted are detected by a facial key point detection model RCPR n At least including: 24 key points related to the eyes and mouth of the face.

[0019] Further, the specific calculation steps of step S3 are as follows:

[0020]

[0021] Wherein And The average value of the 6 key points of the left and right eyes, that is, the key point of the eye with the average value, The average value of the 12 mouth key points, Indicates the center point position of the two eyes;

[0022] Then, the vector Indicates the line from the right eye to the left eye, Indicates the line from the mouth to the eye center point, and the estimated position of the reference point c is obtained according to the center point position of the two eyes and the eye-mouth vector:

[0023]

[0024] Then, the position of the best face alignment frame is calculated according to the reference point c and the hairstyle height H, wherein the hairstyle height H is obtained by a pre-trained semantic segmentation network, and the Point hair = max (Segment (Z)), wherein Segment is a trained face semantic segmentation model Faceprasing, and Z is an input image;

[0025] Finally, the normalized scale is defined as:

[0026]

[0027] At this time, the offset Since the input is square in the application scenario of, it is stipulated that The reference point offset.

[0028] Further, in step S4, the face depth is aligned by the face reference point c and the offset L=(x, y), and the face alignment coordinates are defined as

[0029] The face region is extracted from the source image by coordinates, that is, image information within the range is extracted; when the image has edge holes, the holes are filled by linear interpolation.

[0030] Further, in step S5, the hairstyle of the input image is edited by a fast hairstyle attribute editing module, and the input image group {I1, I2} is projected into the FS latent space of StyleGANv2 by a fast image embedding algorithm;

[0031] The fast image embedding algorithm first encodes the image into the W+ latent space of StyleGANv2 through a pre-trained encoder, and then maps the image to the FS latent space through an iterative optimization method, for subsequent latent code editing, that is:

[0032]

[0033] where L F is the 2-norm of the optimized structure tensor F and the initialized The 2-norm is used to ensure that the final FS latent code is close to the effective area of the StyleGAN latent space;

[0034] The hairstyle of the input image is edited by the fast hairstyle attribute editing module, and the latent code of each input image group is edited as follows: The editing of the hairstyle attribute feature is constrained by the semantic segmentation map of the input image:

[0035] The different semantic regions of the input image are identified by a pre-trained face semantic segmentation model BiseNet, and the target image semantic map M is obtained by recombining the semantic map after portrait semantic segmentation. The face and background of M correspond to the semantic map M1 of the input image I1, and the hair shape corresponds to the semantic map M2 of the input image I2.

[0036] Further, for the target image semantic map M obtained by recombining the semantic map after portrait semantic segmentation, for each face image, the orientation and head shape are different, and many misalignments and holes may occur. The hairstyle or face mask is inflated by an inflation operation until the holes or misalignments are filled.

[0037] Further, the hairstyle of the input image is edited by the fast hairstyle attribute editing module, the input image latent code output by the image embedding module is fine-tuned by the LatentEdit module, and the target image semantic segmentation map M is used as a constraint condition. The fine-tuned latent code is guaranteed to be similar to the image latent code by using a style loss function Style loss.

[0038] The style loss function Style loss is specifically defined as:​ wherein is the activation map of the l-th layer of the VGG network, K is the Jacobi matrix, and I is defined as k (Z) = Segment(Z) | k is the k-th semantic region map of the target image Z, then the style loss L style is calculated as follows:

[0039]

[0040] wherein I k (Z k ) is the k-th semantic region map of the target image Z, and Z k sets all parts other than the k-th region to zero;

[0041] The Latent Edit module extracts the corresponding latent code according to the semantic segmentation α k , mixes the latent codes to generate the target latent code C blend ; for C blend , a set of weight matrices μ k is expected to be found so that The latent code fusion is realized by iteratively optimizing LPIPS under the conditions that ∑ k μ k = 1 and μ k > 0, that is:

[0042]

[0043] wherein, is the activation map output by the l-th layer of the VGG model, and is normalized.

[0044] Further, in step S6, the latent code C blend of the target image is mapped to the image domain by the StyleGANv2 generator to obtain the target image I, and the latent code is reconstructed into an image by the pre-trained StyleGANv2 generator to obtain the target image after hairstyle editing.

[0045] In addition, a hairstyle attribute migration system based on deep adaptation is provided according to the hairstyle attribute migration method based on deep adaptation described above; a user inputs a face image and a hairstyle reference image, the system calls the deep adaptive alignment module to align the depth of the input image, edits the attribute features in the latent feature space of the image, and finally outputs a target image through a generation network to obtain a target image containing the identity features of the face image and the hairstyle features of the hairstyle reference image.

[0046] Compared with the prior art, the main design points and advantages of the present application and the preferred schemes thereof include:

[0047] 1. Under the current premise of hairstyle attribute editing based on latent code, a deep adaptive algorithm is proposed, which is beneficial to the synthesis effect of hairstyle attribute editing task through the depth difference between the input face images constrained by the face depth of monocular ranging;

[0048] 2. The hybrid method of encoder and reverse iteration is adopted in the hidden code acquisition method. The hidden code obtained by the encoder method is often higher than the LPIPS of the reverse iteration, and a large amount of calculation is required by the iterative method, which consumes time. Therefore, the hybrid method is proposed to save time and efficiency while obtaining the optimal hidden code. BRIEF DESCRIPTION OF DRAWINGS

[0049] The present application will be further described in detail below in combination with the accompanying drawings and specific embodiments:

[0050] Figure 1 The deep adaptive method flowchart of the embodiment of the present application.

[0051] Figure 2 The flowchart of the fast hairstyle editing module of the embodiment of the present application.

[0052] Figure 3 The hairstyle attribute migration diagram of the embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the features and advantages of the patent more obvious and easy to understand, the following examples are specifically described as follows:

[0054] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the specification have the same meaning as understood by those skilled in the art to which the present application belongs.

[0055] It should be noted that the terms used herein are only for the purpose of describing the specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form, and in addition, it should be understood that when the terms "comprise" and / or "include" are used in the specification, they indicate the presence of a feature, step, operation, device, component and / or their combination.

[0056] As shown in Figures 1-3 The embodiment provides a hairstyle attribute editing method based on deep adaptive alignment, which specifically comprises the following steps:

[0057] S1, estimating the estimated depth difference of input images {I1, I2} by face monocular ranging algorithm

[0058] S2, detecting 68 facial key points p of the image to be adjusted by a facial key point detection model RCPR n ;

[0059] S3, calculating the facial reference point c and the offset L=(x, y) by the depth difference between the facial key points p n ;

[0060] S4, aligning the facial depth by the facial reference point c and the offset L=(x, y) ;

[0061] S5, editing the hairstyle of the input image by a fast hairstyle attribute editing module to output the latent code C of the target image blend ;

[0062] S6, mapping the latent code C of the target image to the image domain by a StyleGANv2 generator to obtain the target image I blend .

[0063] As a preferred, in the embodiment, the step of estimating the estimated depth difference of the input image {I1, I2} by a facial monocular ranging algorithm is specifically: using a pre-trained face detection model RFBnet to locate the face position in I1 and I2, and recording the face frame size [W face , H face ];

[0064] The estimated depth difference of the input image {I1, I2} is estimated by a facial monocular ranging algorithm According to the principle of similar triangles,

[0065]

[0066] Therefore, the estimated depth difference of the input image can be directly obtained as For the depth alignment direction, the image with deeper depth is adjusted.

[0067] As a preferred, in the embodiment, the step of detecting 68 facial key points p of the image to be adjusted by a facial key point detection model RCPR is specifically: inputting the image to be processed into the pre-trained RCPR model to output the 68 key point position matrix of the face n ;

[0068] For the 68 facial key points p n , more attention is paid to the 24 key points of the eyes and the mouth of the face

[0069] Further, the depth difference is calculated by with the face key points p n The step of calculating the face reference point c and the offset L=(x, y) is specifically: first, the average position of the key points of the two eyes and the mouth is calculated according to the following formula, i.e.

[0070]

[0071] wherein and are the average values of the 6 key points of the left and right eyes, i.e. the key points of the eyes with the average values, is the average value of the 12 key points of the mouth, represents the center point position of the two eyes;

[0072] The vector represents the line connecting the right eye to the left eye, represents the line connecting the mouth to the eye center point, and the estimated position of the reference point c is obtained according to the center point position of the two eyes and the eye-mouth vector:

[0073]

[0074] The position of the best face alignment frame is calculated according to the reference point c and the hairstyle height H, wherein the hairstyle height H is obtained by a pre-trained semantic segmentation network, and the Point hair = max(Segment(Z)), wherein Segment is a trained face semantic segmentation model BiSeNet, and Z is an input image;

[0075] The normalized scale is defined as:

[0076]

[0077] At this time, the offset L=(x, y) can be calculated

[0078] As a preferred, in the embodiment, since the input is square in the main application scenario, it is stipulated that is the reference point offset.

[0079] The specific steps of aligning the face depth by the face reference point c and the offset L=(x, y) are: the face region is extracted from the source image by coordinates, i.e. the image information within the range is extracted;

[0080] Since the calculated reference offset may be greater than the image coordinates, when the image has edge holes, the holes can be filled by linear interpolation.

[0081] As preferred, in the present embodiment, the hairstyle of the input image is edited by the fast hairstyle attribute editing module, and the latent code C of the output target image is output blend The specific steps are: projecting the input image group {I1, I2} into the FS latent space of StyleGANv2 through the fast image embedding algorithm;

[0082] Wherein, the image is first encoded into the W+ latent space of StyleGANv2 through the pre-trained encoder, and then mapped to the FS latent space through the inverse calculation iterative optimization method, which is used for subsequent latent code editing, that is:

[0083]

[0084] Wherein, L F The 2-norm of the optimized structure tensor F and the initialized 2-norm to ensure that the final FS latent code is close to the effective area of the StyleGAN latent space;

[0085] For each group of input image latent codes The editing of the hairstyle attribute feature is constrained by the semantic segmentation map of the input image;

[0086] For the input image, the pre-trained BiSeNet is used to obtain the semantic segmentation map of the image, and for the input two images, the semantic segmentation maps M1 and M2 are obtained respectively.

[0087] The target image semantic map M is obtained by recombining the semantic map after portrait semantic segmentation, and the face and background of M correspond to the semantic map M1 of the input image I1, and the hair shape corresponds to the semantic map M2 of the input image I2.

[0088] Further, due to the different orientations and head shapes of each face image, many misalignments and holes will occur, in order to solve this problem, the hairstyle or face mask is inflated through the inflation operation until the holes or misalignments are filled;

[0089] In the present embodiment, the input image latent code output by the image embedding module is fine-tuned by the Latent Edit module, and the target image semantic segmentation map M is used as a constraint condition, and the style loss function Style loss is used to ensure that the fine-tuned latent code is similar to the image latent code ;

[0090] Further, define Wherein is the activation map of the lth layer of the VGG network, K is the Jacobi matrix, and I k (Z) = Segment (Z) | k is the k semantic region map of the target image Z, and the style loss Lstyle The calculation formula is as follows:

[0091]

[0092] In the formula, I k (Z k )⊙Z k Set all parts except the k region to zero.

[0093] Further, by semantic segmentation α k Extract the corresponding latent code, and mix the latent code to generate the target latent code C blend For C blend , a set of weight matrices μ k is expected to be found so that By iteratively optimizing LPIPS under the condition that ∑ k μ k = 1 and μ k > 0, the latent code fusion is realized, that is:

[0094]

[0095] Wherein, is the activation map output by the VGG model l layer, and is normalized.

[0096] In this embodiment, the specific steps of mapping the latent code C blend of the target image to the image domain to obtain the target image I by the StyleGANv2 generator are as follows: the latent code is reconstructed into an image by the pre-trained StyleGANv2 generator to obtain the target image after hairstyle editing.

[0097] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0098] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in one or more flow or blocks.

[0099] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in one or more flow or blocks.

[0100] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in one or more flow or blocks.

[0101] The above descriptions are only preferred embodiments of the present application, and are not intended to limit the present application to other forms. Any person skilled in the art can make modifications or improvements on the basis of the above disclosed technical content without departing from the technical scope of the present application. Any simple modification, equivalent change and improvement made on the basis of the above embodiments according to the technical essence of the present application shall still fall within the protection scope of the present application.

[0102] The present application is not limited to the above best mode, and anyone can derive other various forms of hairstyle attribute transfer methods and systems based on deep adaptation based on the inspiration of the present application. Any equivalent change and modification made within the scope of the present application shall fall within the scope of the present application.

Claims

1. A method for editing a hairstyle attribute based on deep adaptive alignment, characterized in that, The method comprises the following steps: Step S1 : estimating the estimated depth difference of the input images {I1, I2} by a face monocular range finding algorithm Step S2: detecting the face key points p of the image to be adjusted by a face key point detection model RCPR n ; Step S3: calculating the depth difference between the face key point p and the face reference point c with the face key point p n calculating the face reference point c and the offset L = (x, y) The specific calculation steps of step S3 are as follows: wherein with are the average values of the 6 key points of the left and right eyes, i.e. the key points of the eyes are taken as average values, is the average value of the 12 mouth key points, denotes the center point position of both eyes; After that, the vector represents the line from the right eye to the left eye, represents the line from the mouth to the center of the eye, the estimated position of the reference point c is obtained according to the center point positions of the two eyes and the eye-mouth vector: After that, the position of the best face alignment frame is calculated according to the reference point c and the hairstyle height H, where the hairstyle height H is obtained by a pre-trained semantic segmentation network, and the definition of Point hair = max(Segment(Z)), where Segment is a trained face semantic segmentation model Faceprasing, and Z is an input image. Finally, the normalized scale is defined: At this time, the offset amount is calculated Since the input is a square in the application scenario of, it is stipulated that is the reference point offset amount; Step S4: Align the face depth by the facial reference point c and the offset L=(x, y); In step S4, the face is aligned by aligning the face landmark c with the offset L = (x, y) to the face depth, and the face alignment coordinates are defined as The face region is extracted from the source image through coordinates, that is, the image information within the range is extracted; when the image has edge holes, the holes are filled by a linear interpolation method; Step S5: editing the hairstyle of the input image through the quick hair style attribute editing module, outputting the latent code C of the target image blend ; Step S6: generating a target image I by the StyleGANv2 generator from the latent encoding C of the target image blend mapping to the image domain.

2. The deep adaptive alignment based hairstyle attribute editing method of claim 1, wherein: In step S1, the estimated depth difference of the input image {I1, I2} is estimated by a face monocular ranging algorithm The face position in I1, I2 is located using a pre-trained face detection model RFBnet, and the face frame size [W face ,H face ] is recorded; Estimating the estimated depth difference of the input image {I1, I2} by a human face monocular distance measuring algorithm Specifically: estimating the depth value by the principle of triangulation, according to the principle of similar triangles, In the formula, x is the imaging width of the human face, w is the actual width of the human face, and d is the focal length of the camera. The estimated depth value of the human face is obtained by solving the equation: After differencing, the estimated depth difference of the input image is obtained as 3.The deep-adaptive-alignment-based hairstyle attribute editing method of claim 1, wherein, In step S2, 68 facial key points p of the image to be adjusted are detected by a facial key point detection model RCPR n At least including: 24 key points related to the eyes, mouth of the face.

4. The deep adaptive alignment based hairstyle attribute editing method according to claim 1, wherein, In step S5, the hairstyle of the input image is edited through a fast hairstyle attribute editing module, and the input image set {I1, I2} is projected into the FS latent space of StyleGANv2 through a fast image embedding algorithm; The fast image embedding algorithm first encodes the image into the W+ latent space of StyleGANv2 through a pre-trained encoder, and then maps the image into the FS latent space through an iterative optimization method, for subsequent latent code editing, that is: where L F is the optimized structure tensor F and the initialized 2-norm of to ensure that the final FS latent code representation is close to the effective region of the StyleGAN latent space. editing the hairstyle of the input image by the fast generation type attribute editing module, for each group of potential encoding of the input image editing of the hairstyle attribute features is constrained by the semantic segmentation map of the input image: The different semantic regions of the input image are identified through a pre-trained face semantic segmentation model BiseNet, and the target image semantic graph M is obtained by recombining the semantic graph after portrait semantic segmentation, wherein the face and background of M correspond to the semantic graph M1 of the input image I1, and the hair shape corresponds to the semantic graph M2 of the input image I2.

5. The deep adaptive alignment based hairstyle property editing method of claim 4, wherein: For the target image semantic graph M obtained by recombining the semantic graph after portrait semantic segmentation, for different orientations and head shapes of each face image, many misalignments and holes may occur, and the hairstyle or face mask is expanded through an expansion operation until the holes or misalignments are filled.

6. The deep adaptive alignment based hairstyle property editing method of claim 4, wherein, The hair style of the input image is edited through the fast hair style attribute editing module, the input image latent code output by the image embedding module is fine-tuned through the Latent Edit module, and the fine-tuned latent code is constrained by the target image semantic segmentation map M, and a style loss function Styleloss is used to ensure that the fine-tuned latent code is similar to the image latent code ; The style loss function Style loss is specifically defined as follows: Definition wherein is the activation map of the lth layer of the VGG network, K is the Jacobi matrix, and definition I k (Z) = Segment(Z) | k is the k semantic region map of the target image Z, and the calculation formula of the style loss L style is as follows: wherein I k (Z k )⊙Z k set all but the k region to zero; The Latent Edit module is segmented according to semantics alpha k Extract the corresponding latent code, and generate the target latent code C by mixing the latent code blend ; For C blend , a set of weight matrices mu k is expected to be found By iterating and optimizing LPIPS under the condition that mu k mu k = 1 and mu k > 0, the latent code fusion is realized, that is: wherein, is the activation map of the output of the VGG model at the l-th layer, normalized.

7. The deep adaptive alignment based hairstyle attribute editing method according to claim 1, wherein, In step S6, the latent encoding C of the target image is generated by the StyleGANv2 generator. blend The target image I is obtained by mapping to the image domain. The latent code is reconstructed into an image by the generator of the pre-trained StyleGANv2, resulting in the target image after hairstyle editing.

8. A deep adaptation based hairstyle attribute transfer system, comprising: The hairstyle attribute transfer method based on deep adaptation according to any one of claims 1-7; a user inputs a face image and a hairstyle reference image, a deep adaptation alignment module is called by the system to align the depth of the input image, attribute features are edited in the latent feature space of the image, and finally a target image containing the identity features of the face image and the hairstyle features of the hairstyle reference image is output by a generation network.