Method and device for processing hair region in image, electronic equipment and storage medium

By adjusting the image angle and using mask stitching technology, the problem of mismatched hairstyle details in hairstyle migration was solved, achieving a more natural hair migration effect.

CN116128902BActive Publication Date: 2025-12-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310119312.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-07
Publication Date
2025-12-09
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

Existing technologies often result in significant differences between the details of the original hairstyle and the actual hairstyle when transferring other hairstyles into photos, leading to unnatural effects.

Method used

By determining the first image and the second image, the object in the first image is adjusted based on the facial angle information of the second object, and mask segmentation and stitching are performed to generate the target mask. The image is then reconstructed based on the spatial information of hair details to ensure that the hair shape and color match.

Benefits of technology

This achieves a more natural match between the hairstyle and details in the transplanted image and the original image, improving the naturalness of the hair transplant effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116128902B_ABST
    Figure CN116128902B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for processing a hair region in an image, an electronic device and a storage medium. The method comprises: performing angle adjustment on a first object in a first image based on angle information of a face of a second object in a second image to obtain a third image; performing mask segmentation on the second image and the third image respectively to obtain a first hair mask corresponding to a hair region of the first object and a body mask corresponding to a body region; splicing the first hair mask and the body mask to obtain a target mask; determining a fourth image based on the second image, the third image and the target mask; and reconstructing the fourth image based on hair detail spatial information of the third image to obtain a fifth image, wherein the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image. The present application can make the hairstyle and details of the hair in the fifth image after hair transfer more match the hairstyle and details of the hair in the first image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of Internet, and particularly relates to a method and device for processing hair region in an image, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of digital technology in the past two or three decades, it has become a fashion for people to record the details of life by taking photos. Nowadays, electronic devices that can be used for taking photos are more and more, such as mobile phones, digital cameras, tablet computers, etc. In order to get satisfactory and interesting photos, users usually make some adjustments to some elements in their photos when using these photo-taking devices to take photos, such as adjusting their hairstyle, background, etc.

[0003] In some examples of adjusting the hairstyle, other hairstyles are usually used to replace the user's hairstyle by migrating to the user's photo. However, after migrating other hairstyles to the user's photo, the new photo may present a case that the hairstyle details are quite different from the original hairstyle details. SUMMARY

[0004] The present disclosure provides a method and device for processing hair region in an image, electronic equipment and storage medium, and the technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a method for processing hair region in an image is provided, comprising:

[0006] determining a first image and a second image; the first image comprising a hair region of a first object; the second image comprising a body region of a second object; the body region being a region left after removing the hair region of the second object from a region of the second object;

[0007] adjusting the angle of the first object in the first image based on angle information of the face of the second object in the second image to obtain a third image; the angle information of the first object in the third image matches the angle information of the second object in the second image;

[0008] performing mask segmentation on the second image to obtain a body mask corresponding to the body region, and performing mask segmentation on the third image to obtain a first hair mask corresponding to the hair region of the first object;

[0009] splicing the first hair mask and the body mask to obtain a target mask;

[0010] determining a fourth image based on the second image, the third image and the target mask; the body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape of the target mask, and the hair color in the fourth image matches the hair color in the third image.

[0011] reconstructing the fourth image based on the hair detail spatial information of the third image to obtain a fifth image; the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image.

[0012] In some possible embodiments, the method further includes:

[0013] masking segmenting the second image to obtain a second hair mask corresponding to the hair region of the second object.

[0014] In some possible embodiments, splicing the first hair mask and the body mask to obtain the target mask includes:

[0015] If the hair of the second object is thicker than the hair of the first object, determining a third hair mask from the second hair mask based on the first hair mask; the third hair mask is a mask corresponding to the hair region of the second object that is different in thickness on the outside; in the case that the hair region of the first object and the hair region of the second object are aligned at the hair bottom edge, the thickness different on the outside is the thickness of the hair region of the second object that exceeds the hair region of the first object.

[0016] Splicing the first hair mask, the third hair mask, and the body mask to obtain the target mask.

[0017] In some possible embodiments, before determining the fourth image based on the second image, the third image, and the target mask, the method further includes:

[0018] If there is a to-be-eliminated region in the hair region of the second object compared with the hair region of the target mask, eliminating the to-be-eliminated region in the hair region of the second object in the second image to obtain a first mixed latent variable.

[0019] The to-be-eliminated region is a region that does not exist in the hair region of the target mask.

[0020] In some possible embodiments, eliminating the to-be-eliminated region in the hair region of the second object in the second image to obtain the first mixed latent variable includes:

[0021] respectively determining a first latent variable of the second image in a feature editing space and a second latent variable of the second image in a feature retention space;

[0022] determining a first fitted latent variable based on the first latent variable and the target mask; the first fitted latent variable is a variable in the feature editing space;

[0023] mapping the first fitted latent variable into a second fitted latent variable; the second fitted latent variable is a variable in the feature retention space;

[0024] determine a first mixed latent variable based on the second fitted latent variable and the second latent variable; the first mixed latent variable is a variable in the feature retention space.

[0025] In some possible embodiments, determining the fourth image based on the second image, the third image and the target mask comprises:

[0026] respectively determining a third latent variable of the third image in the feature editing space and a fourth latent variable of the third image in the feature retention space;

[0027] determine a third fitted latent variable based on the third latent variable and the target mask; the third fitted latent variable is a variable in the feature editing space;

[0028] mapping the third fitted latent variable into a fourth fitted latent variable; the fourth fitted latent variable is a variable in the feature retention space;

[0029] determine a second mixed latent variable based on the first mixed latent variable and the fourth fitted latent variable; the second mixed latent variable is a variable in the feature retention space;

[0030] generating an image according to the second mixed latent variable to obtain the fourth image.

[0031] In some possible embodiments, the fourth latent variable indicates hair detail space information of the third image;

[0032] reconstructing the fourth image based on the hair detail space information of the third image to obtain a fifth image, comprising:

[0033] determine a third mixed latent variable based on the fourth latent variable and the second mixed latent variable; the third mixed latent variable is a variable in the feature retention space;

[0034] generating an image according to the third mixed latent variable to obtain the fifth image.

[0035] According to a second aspect of the embodiments of the present disclosure, a device for processing a hair region in an image is provided, comprising:

[0036] an image obtaining module configured to determine a first image and a second image; the first image comprises a hair region of a first object; the second image comprises a body region of a second object; the body region is a region of the second object after removing the hair region of the second object;

[0037] an angle adjusting module configured to perform angle adjustment on the first object in the first image based on angle information of a face of the second object in the second image to obtain a third image; the angle information of the first object in the third image matches the angle information of the second object in the second image;

[0038] The mask segmentation module is configured to perform mask segmentation on the second image to obtain a body mask corresponding to a body region, and perform mask segmentation on the third image to obtain a first hair mask corresponding to a hair region of the first object;

[0039] The mask determination module is configured to perform splicing of the first hair mask and the body mask to obtain a target mask;

[0040] The first fusion module is configured to determine a fourth image based on the second image, the third image and the target mask; the body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape of the target mask, and the hair color in the fourth image matches the hair color in the third image;

[0041] The second fusion module is configured to reconstruct the fourth image based on the hair detail spatial information of the third image to obtain a fifth image; the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image.

[0042] In some possible embodiments, the mask segmentation module is configured to perform:

[0043] mask segmentation on the second image to obtain a second hair mask corresponding to a hair region of the second object.

[0044] In some possible embodiments, the mask determination module is configured to perform:

[0045] if the hair of the second object is thicker than the hair of the first object, determining a third hair mask from the second hair mask based on the first hair mask; the third hair mask is a mask corresponding to a hair region with a difference in thickness on the outside in the hair region of the second object; in the case that the hair base edges of the hair region of the first object and the hair region of the second object are aligned, the difference in thickness on the outside is the thickness of the hair region of the second object beyond the hair region of the first object;

[0046] splicing the first hair mask, the third hair mask and the body mask to obtain the target mask.

[0047] In some possible embodiments, the device further comprises a hair elimination module configured to perform:

[0048] if there is an elimination region in the hair region of the second object compared with the hair region of the target mask, eliminating the elimination region in the hair region of the second object in the second image to obtain a first mixed latent variable;

[0049] The elimination region is a region that does not exist in the hair region of the target mask.

[0050] In some possible embodiments, the hair elimination module is configured to perform:

[0051] respectively determining a first latent variable of the second image in the feature editing space and a second latent variable of the second image in the feature preserving space;

[0052] determining a first fitted latent variable based on the first latent variable and the target mask; the first fitted latent variable is a variable in the feature editing space;

[0053] mapping the first fitted latent variable into a second fitted latent variable; the second fitted latent variable is a variable in the feature preserving space;

[0054] determining a first mixed latent variable based on the second fitted latent variable and the second latent variable; the first mixed latent variable is a variable in the feature preserving space.

[0055] In some possible embodiments, the first fusion module is configured to perform:

[0056] respectively determining a third latent variable of the third image in the feature editing space and a fourth latent variable of the third image in the feature preserving space;

[0057] determining a third fitted latent variable based on the third latent variable and the target mask; the third fitted latent variable is a variable in the feature editing space;

[0058] mapping the third fitted latent variable into a fourth fitted latent variable; the fourth fitted latent variable is a variable in the feature preserving space;

[0059] determining a second mixed latent variable based on the first mixed latent variable and the fourth fitted latent variable; the second mixed latent variable is a variable in the feature preserving space.

[0060] generating the image according to the second mixed latent variable to obtain a fourth image.

[0061] In some possible embodiments, the fourth latent variable indicates the hair detail space information of the third image.

[0062] The second fusion module is configured to perform:

[0063] determining a third mixed latent variable based on the fourth latent variable and the second mixed latent variable; the third mixed latent variable is a variable in the feature preserving space.

[0064] generating the image according to the third mixed latent variable to obtain a fifth image.

[0065] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method of any one of the above first aspect.

[0066] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which enables an electronic device to perform the method of any one of the first aspect of the embodiments of the present disclosure when instructions in the computer-readable storage medium are executed by a processor of the electronic device.

[0067] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, which includes a computer program stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs the method of any one of the first aspect of the embodiments of the present disclosure.

[0068] The technical solutions provided by the embodiments of the present disclosure at least have the following beneficial effects:

[0069] The first image and the second image are determined, the first image includes a hair region of a first object, and the second image includes a body region of a second object, the body region being a region of the second object after removing a hair region of the second object, the first object in the first image is angle-adjusted based on angle information of a face of the second object in the second image, to obtain a third image, angle information of the first object in the third image matches the angle information of the second object in the second image, the second image is mask segmented to obtain a body mask corresponding to the body region, and the third image is mask segmented to obtain a first hair mask corresponding to the hair region of the first object, the first hair mask and the body mask are spliced to obtain a target mask, a fourth image is determined based on the second image, the third image and the target mask, the body region in the fourth image matches the body region in the second image, a hair form in the fourth image matches the hair form of the target mask, and a hair color in the fourth image matches a hair color in the third image, the fourth image is reconstructed based on hair detail spatial information of the third image to obtain a fifth image, and the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image. The hair of the first image after angle adjustment is used to transfer the second object, so that the target mask determined based on the first hair mask and the body mask can be more natural, and thus the hair style and details of the hair in the fifth image after hair transfer are more matched with the hair style and details in the first image.

[0070] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.

[0072] Figure 1 is a schematic diagram of an application environment according to an exemplary embodiment;

[0073] Figure 2 is a flowchart of a method for processing a hair region in an image according to an exemplary embodiment;

[0074] Figure 3 is a schematic diagram of a first image according to an exemplary embodiment;

[0075] Figure 4 is a schematic diagram of a second image according to an exemplary embodiment;

[0076] Figure 5 is a schematic diagram of a target mask according to an exemplary embodiment;

[0077] Figure 6 is a flowchart of determining a target mask according to an exemplary embodiment;

[0078] Figure 7 is a schematic diagram of a target mask according to an exemplary embodiment;

[0079] Figure 8 is a flowchart of determining a first mixed latent variable according to an exemplary embodiment;

[0080] Figure 9 is a step diagram of performing region elimination according to an exemplary embodiment;

[0081] Figure 10 is a flowchart of determining a fourth image according to an exemplary embodiment;

[0082] Figure 11 is a step diagram of determining a fourth image according to an exemplary embodiment;

[0083] Figure 12 is a flowchart of determining a fifth image according to an exemplary embodiment;

[0084] Figure 13 is a block diagram of a processing device for a hair region in an image according to an exemplary embodiment;

[0085] Figure 14 is a block diagram of a method for processing a hair region in an image according to an example embodiment. DETAILED DESCRIPTION

[0086] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0087] It should be noted that the terms "first", "second" and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in other than the order illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0089] Please refer to Figure 1 , Figure 1 is a schematic diagram of an application environment of a method for processing a hair region in an image according to an example embodiment, as Figure 1 shown, the application environment can include a server 01 and a client 02.

[0090] In some possible embodiments, the server 01 can include a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud image hair region processing, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. Basic cloud computing services. The operating system running on the server can include but is not limited to Android system, IOS system, linux, windows, Unix, etc.

[0091] In some possible embodiments, the client 02 described above can include, but is not limited to, a smart phone, a desktop computer, a tablet computer, a notebook computer, a smart speaker, a digital assistant, an augmented reality (AR) / virtual reality (VR) device, a smart wearable device, and the like. It can also be a software running on the client, such as an application, a mini program, and the like. Optionally, the operating system running on the client can include, but is not limited to, an Android system, an IOS system, Linux, Windows, Unix, and the like.

[0092] In some possible embodiments, the server 01 or the client 02 can determine a first image and a second image, the first image including a hair region of a first object, the second image including a body region of a second object, the body region being a region of the second object excluding the hair region of the second object, perform angle adjustment on the first object in the first image based on angle information of a face of the second object in the second image to obtain a third image, the angle information of the first object in the third image matching the angle information of the second object in the second image, perform mask segmentation on the second image to obtain a body mask corresponding to the body region, and perform mask segmentation on the third image to obtain a first hair mask corresponding to the hair region of the first object, splice the first hair mask and the body mask to obtain a target mask, determine a fourth image based on the second image, the third image, and the target mask, the body region in the fourth image matching the body region in the second image, a hair form in the fourth image matching the hair form of the target mask, and a hair color in the fourth image matching the hair color in the third image, and reconstruct the fourth image based on hair detail spatial information of the third image to obtain a fifth image, the hair detail spatial information of the fifth image matching the hair detail spatial information of the third image.

[0093] In some possible embodiments, the client 02 and the server 01 can be connected through a wired link or a wireless link.

[0094] In an exemplary implementation, the client, the server, and the database corresponding to the server can all be node devices in a blockchain system, capable of sharing the information obtained and generated to other node devices in the blockchain system, realizing information sharing among multiple node devices. The multiple node devices in the blockchain system can be configured with the same blockchain, which is composed of multiple blocks, and the adjacent blocks have an association relationship, so that when the data in any block is tampered with, it can be detected through the next block, thereby avoiding the data in the blockchain being tampered with, and ensuring the security and reliability of the data in the blockchain.

[0095] Figure 2is a flow chart of a method for processing a hair region in an image according to an example embodiment, as shown in Figure 2 The method for processing a hair region in an image can be applied to a client, such as a mobile client, or to other node devices, such as a server. The following is described with the server as an example, and the method includes the following steps:

[0096] In step S201, a first image and a second image are determined; the first image includes a hair region of a first object; and the second image includes a body region of a second object. The body region is a region of the second object after removing the hair region of the second object.

[0097] In the embodiments of the present application, the device end can determine the first image and the second image. In an alternative embodiment, when the device end is a server, the server can obtain the first image and the second image through the client, or the server can obtain the first image and the second image from the online. In another alternative embodiment, when the device end is a client (such as a mobile phone), the client can download the first image and the second image from the online, or can obtain the first image and / or the second image through the image acquisition module (such as a camera) provided on the client.

[0098] In the embodiments of the present application, the first image includes the hair region and other regions of the first object. Figure 3 is a schematic diagram of a first image according to an example embodiment, as shown in Figure 3 The first image includes a first object, which can include a hair region 301 and other regions 302. Alternatively, the other regions 302 can refer to the face region (including eyes, nose, eyebrows, lips, etc.) of the first object. Alternatively, as shown in Figure 3 The other regions 302 can include not only the face region of the first object, but also the neck region of the first object. Alternatively, the regions included in the other regions can be determined according to the actual first image, which is not limited here.

[0099] In the embodiments of the present application, the second image includes the hair region and the body region of the second object. Figure 4 is a schematic diagram of a second image according to an example embodiment, as shown in Figure 4 The second image includes a second object, which can include a hair region 401 and a body region 402, and can optionally include a background region 403. The body region 402 is a region of the second object after removing the hair region of the second object, such as the face region, neck region and shoulder region shown in Figure 4 The face region can include eyes, nose, eyebrows, lips, etc.

[0100] In the embodiments of the present application, the hair region in the first image is used to replace the hair region in the second image, so that the hair of the first object can be presented on the second object in the second image.

[0101] In step S203, the first object in the first image is angle-adjusted based on the angle information of the face of the second object in the second image, to obtain a third image; the angle information of the first object in the third image matches the angle information of the second object in the second image.

[0102] In the embodiments of the present application, the first image and the second image have an angle inconsistency, for example, the face of the first object in the first image is turned 40 degrees to the right, and the face of the second object in the second image is straight, that is, the face of the second object has no angle offset. If the device directly processes the first image and the second image, and replaces the hair region of the first image with the hair region of the second image, an unnatural hairstyle caused by the angle problem will occur. Based on the above reasons, the device can angle-adjust the first object in the first image based on the angle information of the second object in the second image to obtain a third image, wherein the third image includes the first object. Specifically, the device can angle-adjust the first object in the first image based on the angle information of the head or face of the second object in the second image, so that the angle information of the first object in the third image matches the angle information of the second object in the second image. For example, the head of the first object or the entire first object is angle-adjusted so that the head of the first object or the entire first object is straight in the first image, and the adjusted first image is referred to as the third image.

[0103] In an optional embodiment, the device can angle-adjust the first object in the first image based on the angle information of the face of the second object in the second image by using an image correction component. Optionally, the image correction component is established on the basis of Pivotal Tuning for Latent-based Editing of Real Images (PTI) and InterFaceGAN algorithms, and PTI and InterFaceGAN are used to align the first object in the first image and the second object in the second image to solve the problem of mismatch between the generated hair of the second object and the second object. Optionally, InterFaceGAN is used to perform angle adjustment of the first object in the first image in the alignment process of the first object and the second object, and PTI is used to ensure the similarity between the third image obtained by adjustment and the first image.

[0104] In an optional embodiment, after the device side performs angle adjustment on the first object in the first image based on angle information of the face of the second object in the second image to obtain an adjusted first image, the face in the first object in the adjusted first image can be deformed. To solve the problem of face deformation in the first object in the adjusted first image, the device side can align the key points of the face in the first object in the adjusted first image with the key points of the face in the second object in the second image through image affine transformation.

[0105] Specifically, the device side can perform key point recognition on the face in the first object in the adjusted first image to determine the preset number of key points contained in the face in the first object in the adjusted first image, and perform key point recognition on the face of the second object in the second image to determine the preset number of key points contained in the face of the second object in the second image. Then, the device side can adjust the relationship between the preset number of key points contained in the face in the first object in the adjusted first image based on the relationship between the preset number of key points contained in the face of the second object in the second image, using image affine transformation, to alleviate or solve the problem of face deformation in the first object in the adjusted first image, to obtain a third image.

[0106] In the embodiments of the present application, image affine transformation can realize a plurality of operations such as translation and rotation through a series of geometric transformations. The transformation can maintain the flatness and parallelism of the image. Flatness means that a straight line remains a straight line after the image is subjected to affine transformation, and parallelism means that parallel lines remain parallel after the image is subjected to affine transformation.

[0107] In step S205, the second image is subjected to mask segmentation to obtain a body mask corresponding to a body region, and the third image is subjected to mask segmentation to obtain a first hair mask corresponding to a hair region of the first object.

[0108] In the embodiments of the present application, the device side can perform mask segmentation on the second image to obtain a body mask corresponding to a body region in the second image. The device side can perform mask segmentation on the third image to obtain a first hair mask corresponding to a hair region of the first object in the third image.

[0109] Specifically, the device side can perform target recognition on the third image to determine the pixels of each target in the third image. In combination with the first hair mask corresponding to the hair region of the first object in the third image, the device side can determine the pixels of the hair region of the first object in the third image. Figure 3The first object shown, the target can refer to the hair, face and neck of the first object in the third image, and optionally, the target can also refer to the eyes, nose, eyebrows, lips, etc. in the face of the first object. Subsequently, the device end can perform binaryzation processing on the third image based on the pixels of each target in the third image, to obtain the mask corresponding to each target in the third image, that is, the first hair mask corresponding to the hair region of the first object, the face mask corresponding to the face region of the first object, the neck mask corresponding to the neck region of the first object, and even the mask corresponding to the eyes in the face of the first object, the mask corresponding to the eyebrows in the face of the first object and the mask corresponding to the lips in the face of the first object, etc.

[0110] Specifically, the device end can perform target recognition on the second image to determine the pixels of each target in the second image. Figure 4 The second object shown, the target can refer to the hair, face, neck and shoulder of the second object in the second image, and optionally, the target can also refer to the eyes, nose, eyebrows, lips, etc. in the face of the second object. The device end can perform binaryzation processing on the second image based on the pixels of each target in the second image, to obtain the mask corresponding to each target in the second image, that is, the second hair mask corresponding to the hair region of the second object, the face mask corresponding to the face region of the second object, the neck mask corresponding to the neck region of the second object and the shoulder mask corresponding to the shoulder region of the second object, and even the mask corresponding to the eyes in the face of the second object, the mask corresponding to the eyebrows in the face of the second object and the mask corresponding to the lips in the face of the second object, etc.

[0111] Optionally, the target in the second image can also include the body of the second object (i.e. the part of the second object in the second image after removing the hair of the second object), and the device end can perform binaryzation processing based on the pixels of the body in the second image to obtain the body mask corresponding to the body region in the second image.

[0112] As can be seen from the foregoing, the device end can perform mask segmentation on the second image to obtain the second hair mask corresponding to the hair region of the second object and the body mask corresponding to the body region of the second object. At the same time, the device end can perform mask segmentation on the third image to obtain the first hair mask corresponding to the hair region of the first object.

[0113] In step S207, the first hair mask and the body mask are spliced to obtain a target mask.

[0114] In an optional embodiment, the device end can splice the first hair mask corresponding to the hair region of the first object and the body mask corresponding to the body region in the second image to obtain a target mask. Figure 5is a schematic diagram of a target mask according to an example embodiment, as shown in Figure 5 including the target mask obtained by splicing the first hair mask 501 and the body mask 502.

[0115] Since the target mask is to replace the hair region of the first object in the third image to the second object, the hair region in the target mask needs to consider some features in the hair region of the second object, such as hair thickness, which can be manifested as the thickness of the top of the hair. Based on this, the embodiment of the present application proposes a way to obtain a target mask. Figure 6 is a flowchart of determining a target mask according to an example embodiment, as shown in Figure 6 including:

[0116] In step S601, if the hair of the second object is thicker than the hair of the first object, a third hair mask is determined from the second hair mask based on the first hair mask; wherein the third hair mask is the mask corresponding to the hair region with the outer side difference thickness in the hair region of the second object; in the case of aligning the hair bottom edges of the hair region of the first object and the hair region of the second object, the outer side difference thickness is the thickness of the hair region of the second object exceeding the hair region of the first object.

[0117] If the hair of the second object is thicker than the hair of the first object, that is, the hair of the second object is relatively thick and more, but the hair of the first object is relatively thin, if the first hair mask and the body mask are directly spliced, the target mask obtained may not be natural (the first hair mask in the target mask may not be natural relative to the body mask), resulting in that the final image obtained subsequently does not meet the demand. Therefore, in the case that the hair of the second object is thicker than the hair of the first object, the device can determine a third hair mask from the second hair mask based on the first hair mask, wherein the third hair mask is the mask corresponding to the hair region with the outer side difference thickness in the hair region of the second object.

[0118] Specifically, the device can align the hair bottom edges of the first hair mask and the second hair mask, at this time, the part of the second hair mask exceeding the first hair mask is the third hair mask. Since the device aligns the hair bottom edges of the first hair mask and the second hair mask, the thickness corresponding to the part of the second hair mask exceeding the first hair mask can be defined as the outer side difference thickness, at this time, the mask corresponding to the hair region with the outer side difference thickness is the third hair mask.

[0119] In step S3603, the first hair mask, the third hair mask and the body mask are spliced to obtain a target mask.

[0120] In the embodiments of the present application, the device side can splice the first hair mask, the third hair mask and the body mask to obtain the target mask. Figure 7 is a schematic diagram of a target mask according to an exemplary embodiment, as shown in Figure 7 includes a first hair mask 701, a third hair mask 702 and a body mask 703. Among them, Figure 7 The body mask shown in the above formula (1) refers to the blank area surrounded by the first hair mask 701, and does not include the blank area outside the first hair mask.

[0121] In the embodiments of the present application, the hair mask part of the target mask obtained by splicing the first hair mask and the third hair mask can be more in line with the head shape of the second object (the difference in hair thickness caused by the different head shapes of the first object and the second object), which can alleviate or even solve the problem that the head shape of the second object generated subsequently is greatly different from the actual head shape of the second object.

[0122] In step S209, a fourth image is determined based on the second image, the third image and the target mask. The body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape in the target mask, and the hair color in the fourth image matches the hair color in the third image.

[0123] In the embodiments of the present application, the device side can determine a fourth image based on the second image, the third image and the target mask, wherein the body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape in the target mask, and the hair color in the fourth image matches the hair color in the third image. Optionally, the background region in the fourth image matches the background region in the second image.

[0124] In an optional embodiment, the matching in the above paragraph can mean that the similarity is greater than or equal to a preset value, for example, the similarity of the body region in the fourth image and the body region in the second image is greater than or equal to 95%. On this basis, when the device side determines the fourth image based on the second image, the third image and the target mask, the body region in the fourth image is the same as the body region in the second image, the background region in the fourth image is the same as the background region in the second image, the hair shape in the fourth image is the same as the hair shape in the target mask, and the hair color in the fourth image is the same as the hair color in the third image.

[0125] Optionally, the hair shape mentioned above can refer to the hairstyle and the hair amplitude, wherein the hair amplitude can refer to the thickness of the hair.

[0126] From the above Figure 3 and Figure 4It can be seen from the corresponding first image and second image respectively that the hair of the second object in the second image has a bang, but the hair of the first object in the first image does not have a bang, that is, the hair of the second object has a part that does not exist in the hair of the first object (different from the roof thickness affecting the head shape in the above), and therefore, the part of the hair of the second object that does not exist in the hair of the first object needs to be eliminated.

[0127] In the embodiment of the application, before determining the fourth image, if there is a to-be-eliminated region in the hair region of the second object compared with the hair region of the target mask, the device end can eliminate the to-be-eliminated region in the hair region of the second object in the second image to obtain a first mixed latent variable, and the first mixed latent variable is used to generate a sixth image. Wherein, the to-be-eliminated region is a region that does not exist in the hair region of the target mask. Figure 8 is a flowchart of determining a first mixed latent variable according to an example embodiment, as shown in Figure 8 as shown, comprising:

[0128] In step S801, a first latent variable of the second image in a feature editing space and a second latent variable in a feature retention space are determined respectively.

[0129] Figure 9 is a step diagram of performing region elimination according to an example embodiment, as shown in Figure 9 as shown, in the embodiment of the application, the device end can determine the first latent variable of the second image in the feature editing space and the second latent variable in the feature retention space.

[0130] In an optional embodiment, the feature editing space is a W+ space, and the first latent variable is an 18*512 vector. In the embodiment of the application, the StyleGan network includes a mapping network (G_mapping), which mainly generates style parameters, that is, samples data Z space from a normal distribution, generates 512-dimensional data W space through multiple fully connected layers, and W space is copied 18 times to be called W+ space. W+ space is a high-editable latent space, and the coupling degree of each feature in W+ space is low, so that when one of the feature data is edited and changed, it will not affect other feature data. For example, if you want to make the person in the input image fatter, you can determine m values in the 18*512 vector that represent fat and thin, and adjust the m values to get a fatter person.

[0131] In an optional embodiment, the feature retention space is an FS latent space, which is better at retaining details and can better encode spatial information. Wherein, the structure tensor F in the FS latent space can roughly control the spatial position of the feature, and the appearance code S can accurately control the global style attribute.

[0132] In step S803, a first fitted latent variable is determined based on the first latent variable and the target mask; the first fitted latent variable is a variable in a feature editing space.

[0133] Optionally, as shown in Figure 9 , the device end can use the first latent variable with strong editability to fit the target mask to obtain the first fitted latent variable in the feature editing space. In the process of fitting the target mask with the first latent variable with strong editability, the feature data in the first fitted latent variable indicating the length of the hair can be edited and changed (from short to long), and the feature data indicating whether the hair has bangs can be edited and changed (from having to not having), so that the first fitted latent variable in the feature editing space is obtained.

[0134] Optionally, as shown in Figure 9 , the device end can input the first fitted latent variable in the feature editing space into the StyleGan generator to obtain a visualized image. Compared with the second image, the visualized image has longer hair and the bangs are removed.

[0135] In step S805, the first fitted latent variable is mapped into a second fitted latent variable; the second fitted latent variable is a variable in a feature retention space.

[0136] According to the above description, since the feature editing space has strong editability but poor ability to retain spatial detail information of the image, as shown in Figure 9 , the device end can map the first fitted latent variable into the second fitted latent variable in the feature retention space.

[0137] In step S807, a first mixed latent variable is determined based on the second fitted latent variable and a second latent variable; the first mixed latent variable is a variable in the feature retention space.

[0138] As shown in Figure 9 , the device end can mix the second fitted latent variable and the second latent variable that retains more spatial detail information of the second image to obtain the first mixed latent variable in the feature retention space.

[0139] Optionally, the device end generates a sixth image based on the first mixed latent variable, so that the sixth image can be visually displayed. Optionally, the device end can input the first mixed latent variable in the feature retention space into the StyleGan generator to obtain a visualized sixth image. Compared with the second image, the sixth image has the bangs removed.

[0140] Figure 10 is a flowchart for determining a fourth image according to an exemplary embodiment, as shown in Figure 10 , comprising:

[0141] In step S1001, the third image is determined in the third latent variable of the feature editing space and the fourth latent variable of the feature retention space.

[0142] Figure 11 is a step of determining a fourth image according to an exemplary embodiment, as Figure 11 As shown, in the embodiment of the application, the device end can determine the third image in the third latent variable of the feature editing space and the fourth latent variable of the feature retention space.

[0143] In an optional embodiment, the feature editing space is W+ space, and the first latent variable is an 18*512 vector. In the embodiment of the application, the StyleGan network includes a mapping network (G_mapping), which mainly generates style parameters, that is, samples data Z space from a normal distribution, generates 512-dimensional data W space through multiple fully connected layers, and W space is copied 18 times to be called W+ space. W+ space is a high-editable latent space, and the coupling degree of each feature in W+ space is low, so that when one of the feature data is edited and changed, it will not affect other feature data. For example, if you want to make the person in the input image fat, you can determine m values in the 18*512 vector that represent fat and thin, and adjust the m values to get a fat person.

[0144] In an optional embodiment, the feature retention space is FS latent space, which is better at retaining details and can better encode spatial information. Among them, the structure tensor F in the FS latent space can roughly control the spatial position of the feature, and the appearance code S can accurately control the global style attribute.

[0145] In step S1003, the third fitting latent variable is determined based on the third latent variable and the target mask; and the third fitting latent variable is a variable in the feature editing space.

[0146] Optionally, as shown in Figure 11 The device end can use the third latent variable with strong editability to fit the target mask to obtain the third fitting latent variable in the feature editing space. Among them, in the process of using the third latent variable with strong editability to fit the target mask, the feature data in the third latent variable that identifies the edge of the hair (such as the thickness of the roof) can be edited and changed to obtain the third fitting latent variable in the feature editing space.

[0147] Optionally, as shown in Figure 11 The device end can input the third fitting latent variable in the feature editing space into the StyleGan generator to obtain a visual new image, and compared with the third image, the edge of the hair and the thickness of the roof in the visual new image are changed.

[0148] In step S1005, the third fitting latent variable is mapped into a fourth fitting latent variable in the feature retention space.

[0149] According to the above, since the feature editing space has strong editability but poor ability to retain spatial details of the image. As shown in Figure 11 The device end can map the third fitting latent variable into the fourth fitting latent variable in the feature retention space.

[0150] In step S1007, a second mixed latent variable is determined based on the first mixed latent variable and the fourth fitting latent variable; the second mixed latent variable is a variable in the feature retention space.

[0151] As shown in Figure 11 In the embodiment of the present application, the device end can determine the second mixed latent variable based on the first mixed latent variable and the fourth fitting latent variable, wherein the second mixed latent variable retains the body region in the second image, the background region in the second image, the hair shape of the target mask, and the hair color in the third image.

[0152] In step S1009, an image is generated according to the second mixed latent variable to obtain a fourth image.

[0153] Optionally, as shown in Figure 11 The device end can input the second mixed latent variable into the StyleGan generator to obtain the fourth image.

[0154] In step S211, the fourth image is reconstructed based on the hair detail spatial information of the third image to obtain a fifth image; the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image.

[0155] In the embodiment of the present application, the device end can reconstruct the fourth image based on the hair detail spatial information of the third image to obtain the fifth image, wherein the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image.

[0156] In an optional embodiment, the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image means that the similarity between the hair detail spatial information of the fifth image and the hair detail spatial information of the third image is greater than or equal to a preset value, such as 95%. Optionally, the hair detail spatial information of the fifth image and the hair detail spatial information of the third image can be the same.

[0157] Figure 12 is a flowchart for determining the fifth image according to an exemplary embodiment, as shown in Figure 12 includes:

[0158] In step S1201, a third mixed latent variable is determined based on the fourth latent variable and the second mixed latent variable; the third mixed latent variable is a variable in the feature retention space.

[0159] In the embodiment of the application, the fourth latent variable indicates the hair detail space information of the third image. As shown in Figure 11 The device end can add the hair detail space information of the third image to the third mixed latent variable to obtain the third mixed latent variable in the feature retention space.

[0160] In step S1203, an image is generated according to the third mixed latent variable to obtain a fifth image.

[0161] Optionally, as shown in Figure 11 The device end can input the third mixed latent variable into the StyleGan generator to obtain the fifth image. The fifth image retains the body region in the second image, the background region in the second image, the hair shape of the target mask, the hair color in the third image, and the space detail information in the third image. The space detail information represents the hair curl degree, the hair thickness, the hair level, and the hair flow sense of each sub-region in the hair region.

[0162] In summary, through the above-described embodiments, the hair region of the first object in the first image can be migrated to the second object of the second image, and the hair migration can retain more detailed features of the hair of the first object, and on the basis of ensuring the realism of the hair region, the hair region of the first object and the face region of the second user can also be better fused.

[0163] Figure 13 is a block diagram of a processing device for a hair region in an image according to an example embodiment. The device has the function of implementing the data processing method in the above-mentioned method embodiments, which can be implemented by hardware or corresponding software executed by hardware. Referring to Figure 13 The device includes an image acquisition module 1301, an angle adjustment module 1302, a mask segmentation module 1303, a mask determination module 1304, a first fusion module 1305, and a second fusion module 1306:

[0164] The image acquisition module 1301 is configured to determine a first image and a second image; the first image includes a hair region of a first object; the second image includes a body region of a second object; the body region is a region of the second object after removing the hair region of the second object from the second object;

[0165] The angle adjustment module 1302 is configured to perform angle adjustment on the first object in the first image based on angle information of the second object in the second image, to obtain a third image; angle information of the first object in the third image matches angle information of the second object in the second image.

[0166] The mask segmentation module 1303 is configured to perform mask segmentation on the second image to obtain a body mask corresponding to a body region, and perform mask segmentation on the third image to obtain a first hair mask corresponding to a hair region of the first object.

[0167] The mask determination module 1304 is configured to perform splicing of the first hair mask and the body mask to obtain a target mask.

[0168] The first fusion module 1305 is configured to perform determination of a fourth image based on the second image, the third image and the target mask; a body region in the fourth image matches a body region in the second image, a hair shape in the fourth image matches a hair shape of the target mask, and a hair color in the fourth image matches a hair color in the third image.

[0169] The second fusion module 1306 is configured to perform reconstruction of the fourth image based on hair detail spatial information of the third image to obtain a fifth image; hair detail spatial information of the fifth image matches hair detail spatial information of the third image.

[0170] In some possible embodiments, the mask segmentation module is configured to perform:

[0171] Mask segmentation on the second image to obtain a second hair mask corresponding to a hair region of the second object.

[0172] In some possible embodiments, the mask determination module is configured to perform:

[0173] If the hair of the second object is thicker than the hair of the first object, determining a third hair mask from the second hair mask based on the first hair mask; the third hair mask is a mask corresponding to a hair region with a difference in thickness on the outside in the hair region of the second object; in the case that the hair region of the first object and the hair region of the second object are aligned at the hair bottom edge, the difference in thickness on the outside is the thickness of the hair region of the second object beyond the hair region of the first object.

[0174] Splicing the first hair mask, the third hair mask and the body mask to obtain the target mask.

[0175] In some possible embodiments, the device further includes a hair elimination module configured to perform:

[0176] If there is a to-be-eliminated region in the hair region of the second object compared with the hair region of the target mask, the to-be-eliminated region in the hair region of the second object in the second image is eliminated to obtain a first mixed latent variable;

[0177] The to-be-eliminated region is a region that does not exist in the hair region of the target mask.

[0178] In some possible embodiments, the hair elimination module is configured to perform:

[0179] respectively determine a first latent variable of the second image in the feature editing space and a second latent variable of the second image in the feature retention space;

[0180] determine a first fitted latent variable based on the first latent variable and the target mask; the first fitted latent variable is a variable in the feature editing space;

[0181] map the first fitted latent variable into a second fitted latent variable; the second fitted latent variable is a variable in the feature retention space;

[0182] determine a first mixed latent variable based on the second fitted latent variable and the second latent variable; the first mixed latent variable is a variable in the feature retention space.

[0183] In some possible embodiments, the first fusion module is configured to perform:

[0184] respectively determine a third latent variable of the third image in the feature editing space and a fourth latent variable of the third image in the feature retention space;

[0185] determine a third fitted latent variable based on the third latent variable and the target mask; the third fitted latent variable is a variable in the feature editing space;

[0186] map the third fitted latent variable into a fourth fitted latent variable; the fourth fitted latent variable is a variable in the feature retention space;

[0187] determine a second mixed latent variable based on the first mixed latent variable and the fourth fitted latent variable; the second mixed latent variable is a variable in the feature retention space.

[0188] perform image generation according to the second mixed latent variable to obtain a fourth image.

[0189] In some possible embodiments, the fourth latent variable indicates hair detail space information of the third image.

[0190] The second fusion module is configured to perform:

[0191] determine a third mixed latent variable based on the fourth latent variable and the second mixed latent variable; the third mixed latent variable is a variable in the feature retention space.

[0192] According to the third mixed latent variable, the image generation is performed to obtain a fifth image.

[0193] It should be noted that the apparatus provided in the above embodiments is only used as an example for the division of the above functional modules in realizing the functions thereof, and in actual application, the above functions can be completed by different functional modules according to the needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be described here.

[0194] Figure 14 is a block diagram of a processing apparatus 3000 for a hair region in an image according to an example embodiment. For example, the apparatus 3000 can be a mobile phone, computer, digital broadcast terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, and the like.

[0195] Referring to Figure 14 , the apparatus 3000 can include one or more of the following components: a processing component 3002, a memory 3004, a power supply component 3006, a multimedia component 3008, an audio component 3010, an input / output (I / O) interface 3012, a sensor component 3014, and a processing component for a hair region in an image 3016.

[0196] The processing component 3002 generally controls the overall operation of the apparatus 3000, such as operations associated with display, phone calls, data communication, camera operations, and recording operations. The processing component 3002 can include one or more processors 3020 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 3002 can include one or more modules to facilitate interaction between the processing component 3002 and other components. For example, the processing component 3002 can include a multimedia module to facilitate the interaction between the multimedia component 3008 and the processing component 3002.

[0197] The memory 3004 is configured to store various types of data to support the operation of the device 3000. Examples of these data include instructions for any application or method operating on the apparatus 3000, contact data, phonebook data, messages, pictures, videos, and the like. The memory 3004 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0198] Power component 3006 provides power to various components of device 3000. Power component 3006 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for device 3000.

[0199] Multimedia component 3008 includes a screen providing an output interface between device 3000 and a user. In some embodiments, the screen includes a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, multimedia component 3008 includes a front camera and / or a rear camera. When device 3000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0200] Audio component 3010 is configured to output and / or input audio signals. For example, audio component 3010 includes a microphone (MIC) configured to receive external audio signals when device 3000 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode. The received audio signals can be further stored in memory 3004 or transmitted via image processing component 3016. In some embodiments, audio component 3010 also includes a speaker for outputting audio signals.

[0201] I / O interface 3012 provides an interface between processing component 3002 and peripheral interface modules, which can be a keyboard, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0202] The sensor component 3014 includes one or more sensors for providing status assessments of various aspects of the device 3000. For example, the sensor component 3014 can detect an open / closed state of the device 3000, relative positioning of components of the device 3000, such as a display and keypad of the device 3000, a change in position of the device 3000 or a component of the device 3000, presence or absence of user contact with the device 3000, orientation or acceleration / deceleration of the device 3000, and temperature changes of the device 3000. The sensor component 3014 can include a proximity sensor configured to detect presence of a nearby object without any physical contact. The sensor component 3014 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 3014 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0203] The image hair region processing component 3016 is configured to facilitate wired or wireless image hair region processing between the device 3000 and other devices. The device 3000 can access a wireless network based on image hair region processing criteria, such as WiFi, a carrier network (e.g., 2G, 3G, 4G, or 5G), or a combination thereof. In an example embodiment, the image hair region processing component 3016 receives broadcast signals or broadcast-related information from an external broadcast managing system via a broadcast channel. In an example embodiment, the image hair region processing component 3016 also includes a near field image hair region processing (NFC) module to facilitate short-range image hair region processing. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0204] In example embodiments, the device 3000 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, for performing the methods described above.

[0205] Embodiments of the present application also provide a computer-readable storage medium, which can be arranged in an electronic device to store at least one instruction or at least one program for implementing an image hair region processing method. The at least one instruction or the at least one program is loaded and executed by the processor to implement the image hair region processing method provided by the method embodiments described above.

[0206] Embodiments of the present application also provide a computer program product, the computer program product comprising a computer program stored in a readable storage medium, at least one processor of a computer device reading and executing the computer program from the readable storage medium, so that the computer device executes the method in any one of the first aspect of the embodiments of the present application.

[0207] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments. And the above-mentioned description is for specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0208] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0209] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.

[0210] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of processing a region of hair in an image, characterized in that, The method comprises the following steps: determining a first image and a second image; the first image comprises a hair region of a first object; the second image comprises a body region of a second object; the body region is a region of the second object after removing the hair region of the second object; adjusting the angle of the first object in the first image based on the angle information of the face of the second object in the second image to obtain a third image; the angle information of the first object in the third image matches the angle information of the second object in the second image; masking and segmenting the second image to obtain a body mask corresponding to the body region, and masking and segmenting the third image to obtain a first hair mask corresponding to the hair region of the first object; splicing the first hair mask and the body mask to obtain a target mask; determining a fourth image based on the second image, the third image and the target mask; the body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape of the target mask, and the hair color in the fourth image matches the hair color in the third image; reconstructing the fourth image based on the hair detail spatial information of the third image to obtain a fifth image; the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image; Before the step of determining the fourth image based on the second image, the third image and the target mask, the method further comprises the following steps: if there is a to-be-eliminated region in the hair region of the second object compared with the hair region of the target mask, eliminating the to-be-eliminated region in the hair region of the second object in the second image to obtain a first mixed latent variable; the to-be-eliminated region is a region that does not exist in the hair region of the target mask.

2. The method of processing a region of hair in an image according to claim 1, characterized in that, The method further comprises the following steps: masking and segmenting the second image to obtain a second hair mask corresponding to the hair region of the second object.

3. The method of processing a region of hair in an image according to claim 2, characterized in that, The step of splicing the first hair mask and the body mask to obtain a target mask comprises the following steps: if the hair of the second object is thicker than the hair of the first object, determining a third hair mask from the second hair mask based on the first hair mask; wherein the third hair mask is a mask corresponding to the hair region on the outside with a thickness difference in the hair region of the second object; in the case that the hair region of the first object and the hair region of the second object are aligned at the bottom edge, the thickness difference on the outside is the thickness of the hair region of the second object beyond the hair region of the first object; splicing the first hair mask, the third hair mask and the body mask to obtain the target mask.

4. The method of processing a hair region in an image according to claim 1, characterized by, The step of eliminating the to-be-eliminated region in the hair region of the second object in the second image to obtain a first mixed latent variable comprises the following steps: respectively determining a first latent variable of the second image in a feature editing space and a second latent variable of the second image in a feature retention space; determine a first fitted latent variable based on the first latent variable and the target mask; the first fitted latent variable is a variable in the feature editing space; map the first fitted latent variable into a second fitted latent variable; the second fitted latent variable is a variable in the feature preserving space; determine a first mixed latent variable based on the second fitted latent variable and the second latent variable; the first mixed latent variable is a variable in the feature preserving space.

5. The method of processing a region of hair in an image according to claim 4, characterized in that, the fourth image is determined based on the second image, the third image and the target mask, comprising: respectively determine a third latent variable of the third image in the feature editing space and a fourth latent variable of the third image in the feature preserving space; determine a third fitted latent variable based on the third latent variable and the target mask; the third fitted latent variable is a variable in the feature editing space; map the third fitted latent variable into a fourth fitted latent variable; the fourth fitted latent variable is a variable in the feature preserving space; determine a second mixed latent variable based on the first mixed latent variable and the fourth fitted latent variable; the second mixed latent variable is a variable in the feature preserving space; generate an image according to the second mixed latent variable to obtain the fourth image.

6. The method of processing a region of hair in an image according to claim 5, characterized in that, the fourth latent variable indicates the hair detail space information of the third image; the fourth image is reconstructed based on the hair detail space information of the third image to obtain a fifth image, comprising: determine a third mixed latent variable based on the fourth latent variable and the second mixed latent variable; the third mixed latent variable is a variable in the feature preserving space; generate an image according to the third mixed latent variable to obtain the fifth image.

7. An apparatus for processing a region of hair in an image, characterized in that, comprising: an image acquisition module configured to determine a first image and a second image; the first image includes a hair region of a first object; the second image includes a body region of a second object; the body region is the region of the second object after removing the hair region of the second object; an angle adjustment module configured to perform angle adjustment on the first object in the first image based on angle information of the face of the second object in the second image to obtain a third image; the angle information of the first object in the third image matches the angle information of the second object in the second image; a mask segmentation module configured to perform mask segmentation on the second image to obtain a body mask corresponding to the body region, and perform mask segmentation on the third image to obtain a first hair mask corresponding to the hair region of the first object; a mask determination module configured to splice the first hair mask and the body mask to obtain a target mask; a first fusion module configured to determine a fourth image based on the second image, the third image and the target mask; the body region in the fourth image matches the body region in the second image, the hair shape in the fourth image matches the hair shape of the target mask, and the hair color in the fourth image matches the hair color in the third image; a second fusion module configured to perform reconstruction of the fourth image based on hair detail spatial information of the third image, to obtain a fifth image; the hair detail spatial information of the fifth image matches the hair detail spatial information of the third image; Before the determining the fourth image based on the second image, the third image and the target mask, the device further comprises a hair removal module configured to perform: if there is a region to be removed in the hair region of the second object compared with the hair region of the target mask, removing the region to be removed in the hair region of the second object in the second image to obtain a first mixed latent variable; the region to be removed is a region not existing in the hair region of the target mask.

8. An electronic device, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method for processing a hair region in an image according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, When the instructions in the computer readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method for processing a hair region in an image according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Hair style conversion method and device and storage medium

    CN115660948A