Method for synthesizing image by using artificial intelligence and method for matching hair style designer based on synthesized image

By learning the hair model of GAN structure and combining the learning of encoder and discriminator, a synthetic image that retains the user's appearance characteristics is generated, which solves the problems of type limitation and feature retention in the prior art, and achieves efficient matching with the hair stylist.

CN120077404AActive Publication Date: 2025-05-30QUANTUMLEAP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380073371.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-12
Publication Date
2025-05-30
Estimated Expiration
2043-10-12

AI Technical Summary

Technical Problem

Existing image synthesis technologies have problems such as type limitations and difficulty in retaining original photo features, and when matching hairstyle designers, they mainly rely on conditions such as schedule or fees.

Method used

By learning the hair model of the GAN structure, the user's images are processed using the mask of the hair region, and combined with the learning of the encoder and discriminator, a synthetic image that retains the user's appearance characteristics is generated while matching the appropriate hairstyle designer.

Benefits of technology

It realizes the generation of synthetic images of the desired hairstyle from the actual photos of the user, while only changing the hairstyle while retaining personal characteristics, improving the accuracy and user experience of the hairstyle designer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077404A_ABST
    Figure CN120077404A_ABST
Patent Text Reader

Abstract

The present specification relates to an image conversion technology, in a method for synthesizing an image by an image synthesis device, a hair model of a Generative Advanced Network (GAN) structure is learned using a plurality of pieces of learning data relating to a hair style, a hair image including an image of a user and a new hair style is received, the image of the user is masked using a mask for a hair region, and the image is synthesized by using an image conversion device. The learned hair model is used to generate a composite image synthesized on the basis of the masked image of the user and the hair image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present specification relates to image transformation technology, and more specifically, to an image synthesis method for obtaining a new photo based on a user's photo using machine learning, and a method for generating an image of changing the user's hairstyle based on the image synthesis method and matching a suitable hairstylist to the user. Background Art

[0002] Currently, there are various image transformation and synthesis technologies for generating new images using original images. Depending on what kind of data the original and new images are, the purpose of the transformation, and even the degree of transformation, the type of technology selected will also be different. In recent years, with the development of artificial intelligence technology, such artificial intelligence technology has also been applied to image transformation and synthesis, and GAN (Generative Adversarial Ne tworks) or autoencoders disclosed in the following prior art documents have become representative means.

[0003] "Generative Adversarial Networks", Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, 2014.

[0004] In the image generation and transformation technology using GAN, artificial neural networks receive various noise inputs to generate new images that did not exist before or convert them into images or videos with other forms or information. Previous deep learning technology usually learns a multi-layer artificial neural network for learning data, but GAN uses a generative neural network that generates virtual images that are difficult to distinguish from real images through the interaction of two artificial neural networks.

[0005] On the other hand, most image synthesis services currently sold on the market can only synthesize pre-selected fixed hairstyle types, so the export of images with styles similar to real people is limited. Although there are research examples that can derive relatively more natural and excellent results when synthesizing images using GAN models, they have not reached the level of practical services that use real people as objects. Summary of the invention

[0006] Technical issues

[0007] The technical problem to be solved by the embodiments of this specification is to address the weakness that the conventional fixed image synthesis technology has limitations on the synthesis type, and to solve the problem that when applying deep learning technologies such as GAN, it is difficult to obtain an image of the desired form from the original photos of actual users, or the main features or information of the original photos are lost. Furthermore, it overcomes the limitation that most of the conventional technologies for matching users and hairstylists only focus on conditions such as schedules or costs.

[0008] Means for Solving the Technology

[0009] To solve the above technical problem, a method for an image synthesis device including at least one processor in an embodiment of this specification to synthesize an image includes the following steps: The image synthesis device learns a hair model of the GAN (Generative Adversarial Networks) structure using a plurality of learning data about hairstyles; the image synthesis device receives an image including a user and a hair image of a new hairstyle; the image synthesis device masks the image of the user using a mask for the hair region; and the image synthesis device generates a synthetic image synthesized based on the masked image of the user and the hair image using the learned hair model.

[0010] In a method for image synthesis in an embodiment, the step of learning the hair model includes the following steps: A generator receives a latent vector in a latent space to generate a fake image; and a discriminator receives the fake image and a real image to calculate a loss regarding their difference. The generator is learned to generate a fake image similar to the real image based on the loss, and the discriminator is learned to determine whether the loss is within a critical value based on the loss.

[0011] In a method for image synthesis in an embodiment, the step of learning the hair model further includes the following steps: Using an encoder to invert the semantic features of hairstyles from actual images including multiple hairstyles to generate a latent space in which similar hairstyle distributions are in adjacent spaces.

[0012] In an image synthesis method of an embodiment, the discriminator includes: a first discriminator that determines whether the virtual image and the real image are of the same face; and a second discriminator that determines whether the virtual image and the real image are of the same hairstyle, and provides the losses calculated by the first discriminator and the second discriminator to the generator to simultaneously guide the learning of the face and the hairstyle. In addition, the first discriminator is learned based on multiple face photos of the same person, and the second discriminator is learned based on multiple hairstyle photos of the same hairstyle.

[0013] To solve the above technical problem, a method for matching a hairstylist based on image synthesis in another embodiment of this specification, which includes at least one processor, includes the following steps: the matching system receives an image of a user; the matching system sets a desired hairstyle input by the user, and uses an image synthesis algorithm to generate a synthetic image synthesized according to the hairstyle from the image of the user; and recommends a hairstylist corresponding to the hairstyle of the synthetic image generated by the matching system. The image synthesis algorithm uses multiple learning data about hairstyles to learn a hair model of the GAN (Generative Adversarial Networks) structure, receives an image of the user and a hair image of a new hairstyle, masks the image of the user using a mask for the hair region, and uses the learned hair model to generate a synthetic image synthesized based on the masked image of the user and the hair image.

[0014] In a method for matching a hairstylist in another embodiment, the step of recommending the hairstylist includes the following steps: considering at least one of the practice areas and resumes of multiple hairstylists and displaying at least one or more hairstylist candidates.

[0015] In a method for matching a hairstylist in another embodiment, the step of recommending the hairstylist further includes the following steps: displaying at least one of the service fees, service areas, and available service dates of the displayed hairstylist candidates together, so as to guide a service appointment between the user and the hairstylist candidates.

[0016] In the hairstylist matching method of another embodiment, in the above image synthesis algorithm, according to the encoder, the semantic features of the hairstyles are inverted from the actual images including multiple hairstyles to generate a latent space where similar hairstyles are distributed in adjacent spaces. The generator receives the latent vectors in the latent space to generate fake images, and the discriminator receives the above fake images and real images to calculate the loss regarding their differences. The above generator is learned in such a way as to generate fake images similar to the real images based on the above loss, and the discriminator is learned in such a way as to determine whether the above loss is within a critical value based on the above loss, thereby learning the above hair model.

[0017] In the hairstylist matching method of another embodiment, the above discriminator includes: a first discriminator that determines whether the above fake image and the above real image are of the same face; and a second discriminator that determines whether the above fake image and the above real image are of the same hairstyle. The losses calculated through the above first discriminator and the above second discriminator are provided to the above generator to simultaneously guide the learning of the face and the hairstyle. The above first discriminator is learned based on multiple face photos of the same person, and the above second discriminator is learned based on multiple hairstyle photos of the same hairstyle.

[0018] On the other hand, a computer-readable recording medium is provided below, which records a program for causing a computer to execute the above image synthesis method and hairstylist matching method.

[0019] Advantages of the Invention

[0020] The embodiments of this specification can apply deep learning technology to generate a synthetic image with a desired hairstyle from the actual photos of the user. In particular, a hair model that provides masking of the hair area and separately learns the face and the hairstyle is provided, which can change only the hairstyle while retaining the user's own inherent appearance features. By introducing image synthesis technology on a platform that connects users and hairstylists, it is possible to guide the matching of hairstylists based on the changed hairstyles of the users. Brief Description of the Drawings

[0021] To help understand this specification, the accompanying drawings, which are included as part of the detailed description, provide the embodiments of this specification and explain the technical features of this specification together with the detailed description.

[0022] Figure 1This is a diagram showing the basic idea of the image synthesis method disclosed in the embodiments of this specification.

[0023] Figure 2 This is a diagram showing the basic structure of GAN (Generative Adversarial Networks).

[0024] Figure 3 This is a diagram schematically showing the processing procedure of image synthesis disclosed in the embodiments of this specification.

[0025] Figure 4 This is a flowchart showing the method for synthesizing an image according to an embodiment of this specification.

[0026] Figure 5 This is a diagram showing the structures of a generator and a discriminator for image synthesis according to an embodiment of this specification.

[0027] Figure 6 This is a diagram for explaining the hair model learning process according to an embodiment of this specification.

[0028] Figure 7 This is a flowchart showing the method for matching a hairstylist based on image synthesis according to another embodiment of this specification.

[0029] Figures 8a to 12 This is a diagram illustrating the processing flow of an application embodying the hairstylist matching method according to another embodiment of this specification.

[0030] Figure 13 This is a block diagram showing the hairstylist matching system according to another embodiment of this specification.

[0031] <Symbol Explanation>

[0032] 10: Hairstylist (hairstylist terminal)

[0033] 20: User (user terminal)

[0034] 30: Matching system

[0035] 31: Communication unit

[0036] 32: Processor

[0037] 33: Memory Detailed Implementation Manner

[0038] Next, embodiments of the present specification will be specifically described with reference to the accompanying drawings. However, in the following description and the accompanying drawings, detailed descriptions of well-known functions or structures that would obscure the gist of the embodiments are omitted. Also, throughout the specification, unless otherwise specifically stated to the contrary, "including" a certain component does not mean excluding other components, but rather means that other components may also be included.

[0039] In addition, terms such as first, second, etc. are used to describe various components, but the above terms do not limit the above components. The above terms are used to distinguish one component from other components. For example, without departing from the scope of the claims of the present invention, the first component may be named the second component, and similarly, the second component may also be named the first component.

[0040] The terms used in this specification are only used to describe specific embodiments and do not limit this specification. In the case where no other definition is explicitly provided, a singular expression includes a plural meaning. In this application, terms such as "including" or "comprising" are used to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and do not preclude the existence or additional possibility of one or more other features, numbers, steps, actions, components, parts, or combinations thereof in advance.

[0041] Unless otherwise specifically defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those skilled in the art. The meaning of terms defined in advance and used in general has the same meaning as that in the context of the relevant technology. In the case where no clear definition is provided in this specification, it cannot be interpreted as an abnormal or overly formal meaning.

[0042] Figure 1 As a diagram showing the basic idea of the image synthesis method disclosed in the embodiments of this specification, it is used to generate a composite image (C) by referring to a target image (B) from an original image (A). At this time, the original image (A) may be an actual photo of the user, and the target image (B) may be a photo with a different hairstyle from the original image (A). The finally generated composite image (C) may be a photo that only reflects the hairstyle features included in the target image (B) in the person (user) of the original image (A). For this purpose, an operation is required to replace the information about the hairstyle in the original image (A) with the information about the hairstyle in the target image (B). Although such a hairstyle change is made, in order to achieve natural image synthesis, it is necessary to maintain the personal characteristics of the user in the original image (A).

[0043] Figure 2FIG. 0 is a diagram showing the basic structure of a GAN (Generative Adversarial Networks). A GAN is a generative model in which two neural networks, a generator 210 that learns a probability distribution and a discriminator 230 that discriminates between different sets, compete with each other to learn. The generator 210 performs training to create virtual examples that deceive the discriminator to the maximum extent, and the discriminator 230 performs training to accurately discriminate between the virtual examples and actual examples presented by the generator 210 to the maximum extent. In this way, the generator 210 is trained adversarially to deceive the discriminator 230, so that the GAN generates replicas (virtual products similar to actual products) that are very similar to actual examples through a process in which the two neural networks develop adversarially. Due to such characteristics, the GAN has received evaluations suitable for image generation and even synthesis.

[0044] However, several problems have been found in the image synthesis technology using GANs.

[0045] First, as Figure 2 illustrated, when random noise is given as the input value to the generator 210 and an image is to be synthesized using the user's actual photo, it is difficult to generate an image of the desired form from the original photo. Therefore, it is necessary to construct the input value before the generator 210 so that the random noise is not simply randomly generated image information, but extracts a certain feature from the original photo and projects that feature.

[0046] Second, the drawback of taking over all the overall characteristics included in the image has been found. For example, when considering the Figure 1 object shown, if only the hairstyle is to be deformed, but the skin color, makeup, and other features of the face are also reflected in the synthesized image, a synthesized image that is somewhat different from the original face is generated. As a result, the user feels that unnecessary features that are not their own face are also mixed in, and it is necessary to further supplement the weaknesses of the deformed image.

[0047] Third, the problem that when synthesizing an image, the characteristics (e.g., hairstyle) of the context image cannot be maintained according to the scale of the target image appears. That is, it is necessary to appropriately control the feature parts involved in the synthesized image according to the purpose.

[0048] Embodiments of the present specification developed in recognition of the above problems provide a technical means as follows: using the user's actual photo as the input value, while preserving the features of the target image, only concentrating on synthesizing the features to be changed among the features of each region included in the image.

[0049] Figure 3 It is a diagram schematically showing the processing procedure of image synthesis disclosed in the embodiments of this specification.

[0050] When the original image 310 of an actual photo of the user being photographed is input, a preprocessing procedure (from cropping 320 to aligning process) is performed in consideration of synthesis performance. For example, when the original image 310 is a full-body photo or in a situation where a part of the area is biased within the photo area or mixed with multiple other objects, from the perspective of hairstyle synthesis targeted in this embodiment, it is preferably cropped and aligned so that only the face and hair area are placed in the center.

[0051] Then, a pre-designed mask 330 is received in such a way that only a specific hair area within the image is used to perform a masking process on the image. At the same time, the hair image 340 to be transformed is received. At this time, the hairstyle of the hair image 340 is different from that of the original image 310, and the target hairstyle that the user wants to change can be input.

[0052] In this way, a synthetic image 360 is generated based on the pre-learned hair model 350 from the masked image and the hair image 340 to be changed and output. Here, regarding how to learn a deep learning model and apply it to image synthesis, it will be specifically described later through Figures 5 to 6 and will be specifically described.

[0053] Figure 4 It is a flowchart showing the method of synthesizing an image according to an embodiment of this specification. From the perspective of implementation, an image synthesis device including at least one processor can execute Figure 4 the processing procedure defined by each step, and drive software including commands for each step through the above-mentioned processor.

[0054] In step S410, the image synthesis device learns a hair model of the GAN (Generative Adversarial Networks) structure using a plurality of learning data regarding hairstyles.

[0055] First, as learning data for hair model learning, hairstyle images of Koreans, i.e., existing data, were collected. During the implementation process of this embodiment, a total of 500,000 photos consisting of 440,000 photos of just having beauty treatments, 10,000 photos of hairstyles with hair tied up or coiled, and 50,000 photos of normal styles were used. These photos were labeled and distinguished according to the same criterion and used as a single data-set or used according to the purpose of the service. All data-sets were augmented by 100%. The importance of a data-set of data collected in a relevant manner for performing a specific task is self-evident. In particular, the individual data types constituting the data-set, the data form of that type, and the quality of the data have a great impact on artificial intelligence learning and prediction performance. Therefore, the data-set disclosed in this embodiment will be specifically disclosed below.

[0056] [Table 1]

[0057]

[0058]

[0059] The data types in Table 1 will be described separately in sequence as follows.

[0060] 1) The beauty salon uniform data-set can provide neat hairstyles just after the treatment in the form of images (extension png), excel files (extension csv), and JSON (JavaScript Object Notation; java script object notation). JSON, as a text-based standard format for representing data structured in the syntax of JavaScript objects, is used when transmitting data in web applications. Since the exif data (exchangeable image file format data) may have different patterns due to differences in exif tag versions, it is preferably provided in json rather than csv.

[0061] 2) The beauty salon longtail data-set can provide neat hairstyles with low likelihood of treatment just after the treatment in the form of images (extension png), excel files (extension csv), and JSON. Longtail data is required when training an AI model, but it may also be inconvenient to use this data. The word 'longtail' from statistics refers to the phenomenon that multiple events with low occurrence probabilities are distributed longer on one side of the statistical distribution. The longtail has a great impact on the design and operation of AI systems. Existing AI systems are particularly weak in longtail data because it is not included in AI learning data that requires a large amount of data due to its low occurrence probability.

[0062] 3) The daily hairstyle dataset can be provided in the form of images (extension: png), excel files (extension: csv), or JSON. It contains data where, more than two weeks after coming back from the beauty salon, the style cannot be immediately distinguished at a glance, and the backgrounds and lighting of the photos are diverse with a lot of noise.

[0063] 4) The special hairstyle dataset can be provided in the form of images (extension: png), excel files (extension: csv), or JSON. It contains hairstyles (such as ponytails, hair loss, very long hair, etc.) that are not done at the beauty salon but are maintained by most people.

[0064] When designing the dataset, the most important thing to consider is data balance. It needs to be designed in such a way that the data is evenly distributed according to appropriate classification criteria, minimizing the data bias that can be expected during learning. In this embodiment, the dataset is constructed in such a way that it includes data with a longtail effect for hairstyles with a relatively large number of actual orders, achieving both trends and an even distribution.

[0065] In addition, the hairstyle images collected in this embodiment, namely the new collected data, are collected by barbershops and hairstylists, which are the applicable areas of the technology, by taking pre- and post-treatment photos of customers, maintaining the same pattern information as the existing data (Korean hairstyle images). An example of the file structure for the new collected data is shown below.

[0066] The "Annotation.csv" file can have a structure like Table 2 below.

[0067] [Table 2]

[0068]

[0069]

[0070] Annotation means the operation of appending each metadata such as the object or image category used to explain the original data to the dataset in the form of 'labels'. That is, it is equivalent to the operation of displaying annotations on the original data so that artificial intelligence can understand the content of the data. The explanatory information data can display various forms and explanatory information according to functional purposes. Here, it provides the hairstyle name, hairstyle type, hair length, hair color, bangs, hair loss degree, side style, age, front representative 2D shot, left and right angles, up and down angles, color, hair part type, gender, special hairstyle distinction, Segment rgb average, etc. in csv form.

[0071] The "Meta - Annotation.csv" file can have a structure like Table 3 below.

[0072] [Table 3]

[0073] Name Type Description id categorical Temporary ID path (Path) string Image file path source (Source) categorical Shooting set id collect-type (Collection type) categorical Procurement category author (Author) string Photographer collect-date (Collection date) timestamp Shooting date polygon1 (Polygon 1) json list Hair segment coordinates polygon2 (Polygon - 2) json list Face segment coordinates before-after (Before - After) categorical Before / after treatment or not height (Height) integer Vertical resolution width (Width) integer Horizontal resolution device (Device) string Shooting equipment

[0074] Metadata, as structured data about data, that is, data used to describe other data, is data assigned to content according to specified rules in order to effectively find and utilize the information sought in a large amount of information. Metadata refers to the information appended to a certain piece of data, that is, structured information, for the purpose of analyzing, classifying, and adding additional information to it. From the perspective of data, it is equivalent to labeling for the purpose of describing data. Labeling means adding information about an object, that is, metadata, when identifying an object in an image, and can provide the path to the picture file of the hairstyle, the shooting set, the photographer, the shooting date, the hair-face segment coordinates, the resolution, the shooting device, etc. in csv format.

[0075] The "optional-Annotation.csv" file can have a structure as shown in Table 4 below.

[0076] [Table 4]

[0077] Name Type Description source (Source) categorical Shooting set id hair-width (Hair width) categorical Hair thickness water-repellency (Water repellency) categorical Hydrophobic hair natural-curl (Natural curl) categorical Natural curl or not damage (Damage) categorical Degree of damage

[0078] Optional annotations, as data providing additional descriptions about hair, can provide the shooting set, hair thickness, hydrophobic hair, whether it is naturally curly, the degree of damage, etc. in csv format.

[0079] The "exifData.csv" file can have a structure as shown in Table 5 below.

[0080] [Table 5]

[0081] Name Type Description path (Path) string File path

[0082] Table 5 can provide the path for storing data in csv format.

[0083] As described above, to solve the problem of being unable to adjust detailed features during image synthesis, the embodiments of this specification introduce an inversion process of generating noise while preserving the features of the target image, i.e., semantic knowledge. This noise reflects the features input into the image to be generated from an actual photo. That is, an encoder that converts an image (actual photo) into noise for image-to-image transformation is disclosed. The encoder can perform various functions such as generating a latent vector that reflects the features of the image, transforming the pose, expression, etc. of the image, or interpolating two images to generate an averaged image. This embodiment aims to derive a latent vector focused on the features regarding the hairstyle.

[0084] Figure 5 FIG. is a diagram showing the structures of a generator and a discriminator for image synthesis according to an embodiment of this specification, and more specifically shows the process (S410) of learning Figure 4 the hair model.

[0085] The generator 510 receives a latent vector in the latent space to generate a fake image. In addition, the discriminators 531, 533 receive the above-mentioned fake image and real image and calculate the loss regarding their differences. The generator 510 learns to generate a fake image similar to the real image based on the above loss, and the discriminators 531, 533 learn to determine whether the above loss is within a critical value based on the above loss.

[0086] However, different from the GAN technology that usually has one discriminator, the embodiments of this specification include at least two discriminators 531, 533. When changing the hairstyle using the conventional GAN, there is a problem that the appearance of the user's face is also changed. The discriminators of this embodiment not only simply judge how similar the synthesized fake photo is to the real photo, but also distinguish the face and the hairstyle in the photo separately for judgment. For this purpose, the discriminator is divided into two types. One discriminator is used to determine whether the face (person) in the predicted photo is the same as the current user's face, and the other discriminator is used to determine whether the hairstyle in the predicted photo is the same as the target hairstyle.

[0087] Refer to Figure 5, two discriminators are illustrated. The first discriminator 531 functions to determine whether a virtual image and a real image are of the same face, and the second discriminator 533 functions to determine whether a virtual image and a real image are of the same hairstyle. Then, the losses calculated by the first discriminator 531 and the second discriminator 533 respectively are provided to the generator 510 to simultaneously guide the learning of the face and the hairstyle.

[0088] On the other hand, since the first discriminator 531 needs to learn based on multiple face photos of the same person, it can receive multiple photos of the same person (for example, person1_photo1, person1_photo2, person2_photo1, person2_photo2,...) as the learning dataset. In addition, since the second discriminator 533 needs to learn based on multiple hairstyle photos of the same hairstyle, it can receive multiple photos of the same hairstyle (for example, target_hairstyle1_photo1, target_hairstyle1_photo2, same_hairstyle2_photo1, same_hairstyle2_photo2,...) as the learning dataset.

[0089] Observing the two types of discriminators 531 and 533 described above, it can be seen that learning for separately judging the identity of the face and the identity of the hairstyle is required. Therefore, in the above Figure 4 In the S410 step, in the learning data required by the image synthesis device to learn the hair model, not only images of the hairstyle but also images of the face are needed. For this reason, the learning data may include: image data including a face area for face learning, image data including a hair area for hairstyle learning, and data for masking the hair area.

[0090] Figure 6 As a diagram for explaining the hair model learning process of an embodiment of this specification, the learning using an encoder 610 and a decoder 630 is illustrated.

[0091] First, in order to extract features from an actual photo, when a photo is input to the encoder 610, an encoded feature is output. Of course, at this time, the input photo is a photo of a hairstyle, and only after preprocessing and input can the hair model of the target hairstyle be learned. After that, the decoder 630 receives this feature and then analogizes the original photo. When such a series of processes are performed multiple times for each photo, the encoded features of photos with similar hairstyles are shown as adjacent points in the latent space and are learned. In the case of photos with different hairstyles, the result will be that they appear far apart in the latent space.

[0092] In summary, during the process of learning the hair model, the semantic features of the hairstyles are inverted from the actual images including multiple hairstyles by using the encoder 610, so that a latent space where similar hairstyles are distributed in adjacent spaces can be generated. Through such a process, the problem that in the conventional GAN technology, random noise cannot project the features of actual photos can be solved.

[0093] In this way, after the learning of the encoder 610 is completed, no matter which photo is input, the encoded feature includes information about the hairstyle. Therefore, according to the convenience of implementation, only this feature can be provided to the Generator of the GAN.

[0094] When the learning of the hair model is completed according to the above process, return to Figure 4 and describe the remaining structure of this embodiment.

[0095] In step S430, the image synthesis device receives an image including the user and a hair image of a new hairstyle. At this time, the image of the user can be an actual photo that hopes to preserve various features of the user's appearance intact.

[0096] In step S450, the above image synthesis device masks the above user image by using a mask for the hair region. In this process, it is controlled that various features of the appearance are preserved intact in the user's actual photo, and only the hairstyle is changed in the way of controlling the deformation domain.

[0097] In step S470, the above image synthesis device generates a synthesized image based on the masked above user image and the above hair image by using the learned above hair model. The previously learned hair model includes 1 generator and 2 discriminators. In particular, the learning of the generator is guided simultaneously by the first discriminator that discriminates the identity of the face and the second discriminator that discriminates the identity of the hairstyle. Therefore, the synthesized image generated by the hair model disclosed in this embodiment can obtain the result that the features in the original image (the actual photo of the user) are preserved except for the features of the target hairstyle (for example, skin color or makeup), and only the hairstyle is naturally reflected.

[0098] Next, the platform application technology applying the above image synthesis method for hairstyles will be introduced.

[0099] Figure 7 It is a flowchart showing a method for matching a hairstylist based on image synthesis according to another embodiment of this specification. From the perspective of implementation, a matching system including at least one processor can executeFigure 7 The processing procedure defined by each step, and drives software including commands for each step through the above-mentioned processor. Regarding the processing procedure of image synthesis, it has been described in detail above through Figures 4 to 6 Therefore, in order to omit the repeated description, only a brief overview will be given here.

[0100] In step S710, the matching system receives the user's image. For example, a user who hopes to change their hairstyle takes a real photo of themselves and provides the user's image to the matching system.

[0101] In step S730, the above-mentioned matching system pre-learns the hair model applied in the image synthesis algorithm. Or the above-mentioned matching system can only receive the result (hair model) learned by other devices physically separated individually.

[0102] In step S750, the above-mentioned matching system sets the desired hairstyle input by the user, and uses the image synthesis algorithm to generate a synthetic image according to the above-mentioned hairstyle from the above-mentioned user's image. Here, in the above-mentioned image synthesis algorithm, a hair model of the GAN (Generative Adversarial Networks) structure is learned using a plurality of learning data regarding the hairstyle, and receives the user's image and the hair image of the new hairstyle, masks the above-mentioned user's image using a mask for the hair area, and generates a synthetic image synthesized based on the masked above-mentioned user's image and the above-mentioned hair image using the learned above-mentioned hair model.

[0103] In addition, in the image synthesis algorithm, an encoder is used to invert the semantic features of the hairstyle from the actual images including a plurality of hairstyles to generate a latent space where similar hairstyles are distributed in adjacent spaces. The generator receives a latent vector in the latent space and generates a fake image. The discriminator receives the above-mentioned fake image and the real image and calculates the loss regarding the difference therebetween. The above-mentioned generator is learned in such a way that based on the above-mentioned loss, a fake image similar to the real image is generated. The discriminator is learned in such a way that based on the above-mentioned loss, it determines whether the above-mentioned loss is within a critical value, so that the above-mentioned hair model can be learned.

[0104] Furthermore, preferably, the discriminator includes a first discriminator that determines whether the virtual image and the real image are of the same face, and a second discriminator that determines whether the virtual image and the real image are of the same hairstyle. The losses calculated by the first discriminator and the second discriminator are respectively provided to the generator to simultaneously guide the learning of the face and the hairstyle. The first discriminator is learned based on multiple face photos of the same person, and the second discriminator is learned based on multiple hairstyle photos of the same hairstyle.

[0105] In step S770, the matching system recommends a hairstylist corresponding to the hairstyle of the composite image generated in step S750. For this purpose, the matching system can be a cooperation platform that connects barbershops and hairstylists working in barbershops to the user. That is, multiple hairstylists can be registered in the matching system, and hairstylists meeting the user's requirements can be recommended based on the items that the hairstylist can perform and various treatment conditions. When the user selects a recommended hairstylist, a convenient function for handling treatment reservations and settlements within one platform can be provided.

[0106] In summary, in step S770 of recommending a hairstylist, at least one hairstylist candidate is displayed considering at least one of the treatment fields and resumes of multiple hairstylists. Furthermore, at least one of the treatment fees, treatment areas, and available treatment dates of the displayed hairstylist candidates is also displayed, thereby guiding a treatment reservation between the user and the hairstylist candidates.

[0107] Figures 8a to 12 It is a diagram illustrating a processing flow of an application of a hairstylist matching method embodying another embodiment of the present specification.

[0108] Figure 8a and Figure 8b Illustrates a user interface of the matching application. First, in Figure 8a , the user takes a photo of himself / herself and displays it on the screen to select various items to be changed in the current hairstyle. For example, hair length, waves, hair shape, and even hair color can be presented as selectable items. Second, in Figure 8b , the user finds a photo of himself / herself from the storage device of the terminal (e.g., a smartphone) and finds a photo of another person (e.g., an artist) as the target photo and presents it. Then, the user selects the result prediction to confirm the composite image generated from the user's own photo and the target photo. This can be an image predicting the result when a hairstyle treatment is performed.

[0109] Figure 9 Shows according to the above Figure 8a orFigure 8b The case where an item selected by the user interface is used to display a composite image on the screen. Compared with Figure 8a or Figure 8b the original photo (user image), it can be confirmed that Figure 9 the composite image of Figure 9 saves all the characteristics of the same person while very naturally showing an image in which the hair length and waves change. In this way, the user can select the designer search button in the

[0110] Figure 10 screen to perform the matching service of the matching system (platform).

[0111] Figure 11 Shows hairstylists who can be matched according to the user's conditions. These hairstylists are hairstylists who can perform procedures on the previously generated composite image. If necessary, the results with additional search conditions can be displayed. For example, only the filtered results are displayed with the resume range or popularity desired by the user as additional conditions. At this time, the user can select a hairstylist and enter the detailed screen.

[0111] Figure 11 Shows the services that the selected hairstylist can perform. The user selects at least one service provided by this hairstylist and enters the Figure 12 reservation screen.

[0112] Figure 12 Shows the available treatment times of the selected hairstylist. If necessary, as an example, it is displayed on one screen including each of the previously retrieved hairstylists, so as to guide other selections by the user. In this way, the user can complete the reservation at a specific available treatment time. If necessary, enter the settlement screen and provide various options for pre-settlement.

[0113] Figure 13 As a block diagram of a hairstylist matching system showing another embodiment of this specification, the matching method of Figure 7 is reconstructed from the perspective of the hardware structure.

[0114] The hairstylist 10 can be the terminal held by the hairstylist or the reservation terminal of the barbershop, and can be connected to the matching system 30 through the network.

[0115] The user 20 uses the terminal or PC held by the user and connects to the matching system 30 through the network.

[0116] The matching system 30 includes a communication unit 31 for connecting to the hairstylist 10 and the user 20 via a network. The communication unit 31 mediates the matching and reservation of the barbershop for the user. The matching system 30 downloads or stores in the memory 33 a matching software including commands defining a series of processes for receiving a matching request from the user 20 and processing the same, and includes a processor 32 for executing the matching software downloaded or stored in the memory 33.

[0117] The matching software includes commands for receiving an image of the user, setting a desired hairstyle received from the user 20, generating a composite image of the hairstyle from the image of the user using an image synthesis algorithm, and recommending the hairstylist 10 corresponding to the hairstyle of the generated composite image. Here, the image synthesis algorithm learns a hair model of the GAN (Generative Adversarial Networks) structure using a plurality of learning data on hairstyles, receives an image of the user and a hair image of a new hairstyle, masks the image of the user using a mask for the hair region, and generates a composite image synthesized from the masked image of the user and the hair image using the learned hair model.

[0118] By Figure 13 The matching system prompted stores personalized data using photo data obtained from the customer, and also stores a large amount of learning data on other hairstyles using the treatment result photos input by the hairstylist. In this case, the designer actively provides the results of his or her treatment to the matching system for marketing purposes, so as to achieve the goal of showing to the customer, and from the perspective of the matching system, it can be an opportunity to obtain high-quality learning data.

[0119] On the other hand, the embodiments of the present specification can be embodied as computer-readable code in a computer-readable recording medium. The computer-readable recording medium includes all kinds of recording devices storing data readable by a computer system.

[0120] Examples of the computer-readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. In addition, the computer-readable recording medium is dispersed in computer systems connected by a network, stores the computer-readable code in a distributed manner, and executes the same. And functional programs, codes, and code segments for embodying the embodiments can be easily derived by programmers in the technical field to which the present specification pertains.

[0121] As described above, this specification has been explained centering on its various embodiments. Those skilled in the art should understand that each embodiment can be implemented in a form that can be deformed without departing from the essential characteristics of this specification. Therefore, the disclosed embodiments should be considered from an illustrative point of view rather than a limiting point of view. The scope of this specification should not be defined according to the above description, but should be defined according to the claims, and all differences within the equivalent scope are included in this specification.

[0122] Industrial Applicability

[0123] According to the embodiments of the present specification described above, deep learning technology can be applied to generate a synthetic image in which the hairstyle is changed to a desired hairstyle from the actual photo of the user. In particular, a hair model that masks the hair area and separately learns the face and hairstyle is provided. While maintaining one's own inherent appearance characteristics, only the hairstyle is changed, and an image synthesis technology is introduced into the platform that connects the user and the hairstylist, thereby guiding the matching of the hairstylist based on the changed hairstyle of the user.

Claims

1. An image synthesis method, which is a method for an image synthesis device including at least one processor to synthesize an image, and includes the following steps: The image synthesis device learns a hair model of a GAN, i.e., a generative adversarial network structure, using a plurality of learning data, and the plurality of learning data includes: image data of a face area for face learning, image data of a hair area for hairstyle learning, and data obtained by masking the hair area; The above image synthesis device receives an image of a user and a hair image of a new hairstyle; The above image synthesis device masks the above user's image using a mask for the hair area; and The above image synthesis device uses the learned above hair model to generate a synthetic image synthesized based on the masked above user's image and the above hair image, In the step of learning the above hair model, The generator generates a virtual image, and the discriminator receives the above virtual image and the real image to distinguish the difference therebetween, and the two compete with each other to learn, The above discriminator includes a first discriminator and a second discriminator, so as to separately distinguish the face and the hairstyle in the photo for judgment, wherein, The first discriminator receives multiple face photos of the same person as a learning data set for learning, so as to judge whether the above virtual image and the above real image are of the same face; and The second discriminator receives multiple hairstyle photos of the same hairstyle as a learning data set for learning, so as to judge whether the above virtual image and the above real image are of the same hairstyle.

2. The image synthesis method according to claim 1, wherein, The step of learning the above hair model includes the following steps: The generator receives a latent vector in the latent space to generate a virtual image; and The discriminator receives the above virtual image and the real image to calculate a loss regarding the difference therebetween, The above generator learns in such a way as to generate a virtual image similar to the real image based on the above loss, The above discriminator learns in such a way as to determine whether the above loss is within a critical value based on the above loss.

3. The image synthesis method according to claim 2, wherein, The step of learning the above hair model further includes the following steps: Using an encoder, reverse-transform the semantic features of the hairstyle from the actual images including multiple hairstyles to generate a latent space in which similar hairstyle distributions are in adjacent spaces.

4. The image synthesis method according to claim 2, wherein, The above discriminator provides the losses calculated respectively through the above first discriminator and the above second discriminator to the above generator to simultaneously guide the learning of the face and the hairstyle.

5. The image synthesis method according to claim 1, wherein, The above learning data includes: A first data set, which provides a hairstyle image just after the treatment; A second data set, which is longtail data, provides a hairstyle image with a low possibility of treatment just after the treatment; A third data set, which provides a daily hairstyle image where the hairstyle cannot be immediately distinguished; and A fourth data set, which provides a hairstyle image maintained by most people although no treatment has been performed, Construct a data set in such a way that it includes data on parts of hairstyles with a long-tail effect that have a large number of actual orders.

6. A hairstylist matching method, which is a method for a matching system including at least one processor to match hairstylists based on image synthesis, and includes the following steps: The matching system receives an image of the user. The matching system sets a desired hairstyle input by the user, and generates a synthetic image based on the hairstyle from the image of the user using an image synthesis algorithm; and Recommends a hairstylist corresponding to the hairstyle of the synthetic image generated by the matching system, The image synthesis algorithm uses masking of multiple learning data to learn a hair model of a GAN, i.e., a generative adversarial network structure, Wherein, The multiple learning data includes: image data including a face area for face learning, image data including a hair area for hairstyle learning, and data with masking processing on the hair area, and receives an image of the user and a hair image of a new hairstyle, masks the image of the user using a mask for the hair area, and generates a synthetic image synthesized based on the masked image of the user and the hair image using the learned hair model, The generator generates a virtual image, and the discriminator receives the virtual image and the real image to distinguish the difference between them, and the two compete with each other to learn the hair model, The discriminator includes a first discriminator and a second discriminator, so as to separately distinguish the face and the hairstyle in the photo for judgment, wherein, The first discriminator receives multiple face photos of the same person as a learning data set for learning, so as to judge whether the virtual image and the real image are the same face; and The second discriminator receives multiple hairstyle photos of the same hairstyle as a learning data set for learning, so as to judge whether the virtual image and the real image are the same hairstyle.

7. The hairstylist matching method according to claim 6, Wherein, The step of recommending the hairstylist includes the following steps: Display at least one or more hairstylist candidates by considering at least one of the practice fields and resumes of multiple hairstylists.

8. The hairstylist matching method according to claim 7, Wherein, The step of recommending the hairstylist further includes the following steps: Display at least one of the service fees, service areas, and available service dates of the displayed hairstylist candidates together to guide a service appointment between the user and the hairstylist candidates.

9. The hairstylist matching method according to claim 6, Wherein, In the above image synthesis algorithm, an encoder is used to reverse-transmit the semantic features of the hairstyle from the actual image including multiple hairstyles, thereby generating a latent space in which similar hairstyles are distributed in adjacent spaces. The generator receives the latent vector in the latent space and generates a virtual image. The discriminator receives the virtual image and the real image and calculates the loss regarding the difference therebetween. The generator is learned to generate a virtual image similar to the real image based on the loss. The discriminator is learned to determine whether the loss is within a critical value based on the loss, and the losses calculated by the first discriminator and the second discriminator are respectively provided to the generator to simultaneously guide the learning of the face and the hairstyle, thereby learning the hair model.

10. The hairstylist matching method according to claim 6, wherein, the learning data includes: a first data set that provides hairstyle images immediately after the treatment; a second data set that is long-tail data and provides hairstyle images with a low likelihood of treatment immediately after the treatment; a third data set that provides daily hairstyle images where the hairstyle cannot be immediately distinguished; and a fourth data set that provides hairstyle images maintained by most people although no treatment has been performed, constituting a data set in such a manner as to collectively include data of a part having a long-tail effect of hairstyles with a large number of actual orders.

Citation Information

Patent Citations

  • Generative adversarial network model-based hairstyle changing method

    CN107527318A

  • Image conversion method and device, computer equipment and storage medium

    CN111489287A

  • Modifying an appearance of hair

    US20220207807A1

  • KR20200115706A