A method for synthesizing images using artificial intelligence and a method for matching hair designers based on image synthesis
The method uses GANs to synthesize images with changed hairstyles while preserving user features, addressing limitations in conventional technologies and enabling effective hair designer matching.
Patent Information
- Application Number
- JP2025522206
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-12
- Publication Date
- 2025-11-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional image synthesis technologies are limited in their ability to synthesize realistic hairstyles from user photos and often lose the main features of the original image, while existing user-hair designer matching systems focus only on non-image-based criteria like schedule and cost.
A method using Generative Adversarial Networks (GAN) to learn a hair model from multiple learning data, masking user images to change hairstyles while preserving facial features, and a matching system to connect users with suitable hair designers based on synthesized images.
Generates natural images with changed hairstyles while maintaining user features, and facilitates effective hair designer matching based on synthesized images.
Smart Images

Figure 2025537087000001_ABST
Abstract
Description
[Technical Field]
[0001] This specification relates to an image conversion technology, and more particularly to an image synthesis method that uses machine learning to acquire a new photo from a user's photo, and a method that generates an image in which the user's hairstyle is changed based on the image synthesis method and matches the user with a hair designer that suits the user. [Background technology]
[0002] There are various image conversion and synthesis techniques for generating a new image using an original image. The type of technique selected may vary depending on the type of data in the original image and the new image, the purpose of the conversion, or the degree of conversion. With the recent development of artificial intelligence technology, such artificial intelligence technology has also been used in image conversion and synthesis, and representative techniques include generative adversarial networks (GANs) and autoencoders, as presented in the following prior art documents:
[0003] "Generative Adversarial Networks", Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, 2014.
[0004] Image generation and conversion technology using GAN allows an artificial neural network to generate new images that do not exist before, or convert them into images or videos with different forms or information, based on various noise inputs. While conventional deep learning technology typically trains a single multi-layered artificial neural network on training data, GAN utilizes a single generative neural network that ultimately generates fake images that are difficult to distinguish from the real thing through the interaction of two artificial neural networks.
[0005] Meanwhile, most currently available image synthesis services are limited in their ability to synthesize images based on pre-selected fixed hairstyles, limiting their ability to derive realistic style images. Although there are research examples showing that synthesizing images using GAN models can produce relatively more natural and superior results, they have not yet reached the level of practical services targeting actual people. Summary of the Invention [Problem to be solved by the invention]
[0006] The technical problem that the embodiments of the present specification aim to solve is to overcome the weakness of conventional fixed-method image synthesis technologies, which have limitations on synthesis types, and to solve the problem that even when applying deep learning technologies such as GAN, it is difficult to obtain a morphological image from an actual user's original photo, or the main features and information of the original photo are lost.Furthermore, it aims to overcome the limitation that conventional technologies for matching users and hair designers are mostly focused only on conditions such as schedule and cost. [Means for solving the problem]
[0007] In order to solve the technical problem, according to one embodiment of the present specification, a method for synthesizing an image using an image synthesis device including at least one processor includes the steps of: the image synthesis device learning a hair model with a Generative Adversarial Networks (GAN) structure using multiple learning data related to hairstyles; the image synthesis device receiving an input of a hair image including a user's image and a new hairstyle; the image synthesis device masking the user's image using a hair region mask; and the image synthesis device generating a synthetic image based on the masked user's image and the hair image using the learned hair model.
[0008] In an image synthesis method according to an embodiment, the step of training the hair model includes a step of a generator receiving a latent vector in a latent space as an input and generating a fake image, and a step of a discriminator receiving the fake image and a real image as input and calculating a loss related to the difference between them, wherein the generator trains to generate a fake image similar to the real image based on the loss, and the discriminator trains to determine whether the loss is within a threshold based on the loss.
[0009] In an image synthesis method according to an embodiment, the step of learning the hair model may further include the step of generating a latent space in which similar hairstyles are distributed in a neighboring space by inverting semantic features of hairstyles from an actual image including a plurality of hairstyles using an encoder.
[0010] In one embodiment of the image synthesis method, the classifier includes a first classifier that determines whether the fake image and the real image have the same face, and a second classifier that determines whether the fake image and the real image have the same hairstyle, and the losses calculated by the first classifier and the second classifier can be provided to the generator to simultaneously learn faces and hairstyles. Furthermore, the first classifier can be trained based on multiple face photos of the same person, and the second classifier can be trained based on multiple hairstyle photos of the same hairstyle.
[0011] To solve the above technical problem, according to another embodiment of the present specification, a method for matching a hair designer based on image synthesis by a matching system including at least one processor includes the steps of receiving an input of a user's image by the matching system, setting a desired hairstyle input by the user, and generating a synthetic image from the user's image according to the hairstyle using an image synthesis algorithm, and the matching system recommending a hair designer corresponding to the hairstyle of the generated synthetic image, wherein the image synthesis algorithm learns a hair model with a Generative Adversarial Networks (GAN) structure using a large amount of learning data related to hairstyles, receives input of a user's image and a hair image including a new hairstyle, masks the user's image using a mask for a hair region, and generates a synthetic image based on the masked user's image and the hair image using the learned hair model.
[0012] In another embodiment of the hair designer matching method, the step of recommending a hair designer may include a step of displaying at least one or more hair designer candidates, taking into consideration at least one of the fields of practice and careers of the multiple hair designers.
[0013] In another embodiment of the hair designer matching method, the step of recommending a hair designer may further include a step of encouraging the user to make a treatment appointment between the candidate hair designer and the candidate hair designer by displaying at least one of the treatment cost, treatment area, and treatment available date and time of the displayed candidate hair designer.
[0014] In another embodiment of the hair designer matching method, the image synthesis algorithm generates a latent space in which similar hairstyles are distributed in a neighboring space by inverting semantic features of hairstyles from an actual image including a plurality of hairstyles using an encoder, a generator receives a latent vector in the latent space as an input and generates a fake image, a discriminator receives the fake image and a real image as input and calculates a loss related to the difference between them, the generator learns to generate a fake image similar to the real image based on the loss, and the discriminator learns to determine whether the loss is within a threshold based on the loss, thereby training the hair model.
[0015] In another embodiment of the hair designer matching method, the classifier includes a first classifier that determines whether the fake image and the real image have the same face, and a second classifier that determines whether the fake image and the real image have the same hairstyle, and losses calculated through the first classifier and the second classifier are provided to the generator to simultaneously conduct face and hairstyle learning, and the first classifier may be trained based on multiple face photos of the same person, and the second hairstyle classifier may be trained based on multiple hairstyle photos of the same hairstyle.
[0016] Meanwhile, a computer-readable recording medium having recorded thereon a program for causing a computer to execute the image synthesis method and hair designer matching method described above will be provided below. [Effects of the Invention]
[0017] The embodiments of the present specification utilize deep learning technology to generate a composite image in which a user's actual photo has been changed to a desired hairstyle. In particular, by masking the hair area and providing a hair model that has been trained for each face and hairstyle, it is possible to obtain a change in hairstyle only while preserving unique external features. Furthermore, by introducing image synthesis technology into a platform that connects users and hair designers, it is possible to induce hair designer matching based on the user's changed hairstyle. [Brief explanation of the drawings]
[0018] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present specification, provide embodiments of the present specification and, together with the detailed description, explain the technical features of the present specification.
[0019] FIG. 1 is a diagram showing the basic idea of the image synthesis method proposed in the embodiment of this specification.
[0020] FIG. 2 is a diagram showing the basic structure of GANs (Generative Adversarial Networks).
[0021] FIG. 3 is a diagram showing an overview of the image synthesis process proposed in the embodiment of this specification.
[0022] FIG. 4 is a flow chart illustrating a method for combining images according to an embodiment of the present disclosure.
[0023] FIG. 5 is a diagram illustrating the configuration of a generator and a segmenter for image synthesis according to an embodiment of the present specification.
[0024] FIG. 6 is a diagram for explaining a hair model learning process according to an embodiment of the present specification.
[0025] FIG. 7 is a flow chart illustrating a method for matching hair designers based on image synthesis according to another embodiment of the present disclosure.
[0026] 8A to 12 are diagrams illustrating an example of a processing flow of an application that realizes a hair designer matching method according to another embodiment of the present specification.
[0027] FIG. 13 is a block diagram showing a hair designer matching system according to another embodiment of the present specification.
[0028] <Explanation of symbols>
[0029] 10: Hair designer (hair designer terminal)
[0030] 20: User (user terminal)
[0031] 30: Matching System
[0032] 31: Communications Department
[0033] 32: Processor
[0034] 33: Memory DETAILED DESCRIPTION OF THE INVENTION
[0035] Hereinafter, embodiments of the present specification will be described in detail with reference to the drawings. However, in the following description and the accompanying drawings, detailed descriptions of known functions or configurations that may obscure the gist of the embodiments will be omitted. Furthermore, throughout the specification, the term "comprises" or "includes" a certain element does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.
[0036] Furthermore, terms such as "first" and "second" may be used to describe various components, but the components should not be limited by these terms. These terms may be used to distinguish one component from another. For example, a first component may be referred to as a "second component," and similarly, a second component may be referred to as a "first component," without departing from the scope of the present invention.
[0037] The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the specification. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this application, the terms "comprise" or "comprises" are intended to specify the presence of embodied features, numbers, steps, operations, components, parts, or combinations thereof, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which this specification pertains. Terms as defined in commonly used dictionaries should be interpreted as meanings consistent with the meanings they have in the context of the relevant art, and should not be interpreted as idealized or overly formal unless expressly defined herein.
[0039] FIG. 1 illustrates the basic idea of an image synthesis method proposed in an embodiment of the present specification. The goal is to generate a synthetic image C from an original image A by referencing a target image B. In this case, original image A may be an actual photo of a user, and target image B may be a photo in which the person (user) in original image A has a different hairstyle. The ultimately generated synthetic image C may be a photo that reflects only the hairstyle characteristics contained in target image B of the person in original image A. To achieve this, it is necessary to replace information about the hairstyle in original image A with information about the hairstyle in target image B. Despite this change in hairstyle, the user's personal characteristics in original image A should be maintained to achieve a natural image synthesis.
[0040] FIG. 2 illustrates the basic structure of a generative adversarial network (GAN). A GAN is a generative model in which two neural networks—a generator 210 that learns probability distributions and a discriminator 230 that discriminates between different sets—train through competition with each other. The generator 210 creates fake examples and trains them to fool the discriminator as much as possible, while the discriminator 230 trains them to discriminate between fake examples presented by the generator 210 and real examples as accurately as possible. By training the generator 210 in an adversarial manner to fool the discriminator 230, the GAN can generate highly similar imitation examples (real-like fakes) to real examples through the process in which the two neural networks develop in an adversarial manner. These characteristics have led to GANs being recognized as suitable for image generation or synthesis.
[0041] However, some problems have been discovered with image synthesis technology that utilizes GANs.
[0042] As illustrated in FIG. 2, the input values input to the generator 210 are random numbers. When random noise is given, it is difficult to generate a desired image from the original photograph when synthesizing an image using a user's actual photograph. Therefore, the random noise is not simply randomly generated image information, but rather it is necessary to extract certain features from the original photograph and configure input values before the generator 210 so as to project the corresponding features.
[0043] Second, a drawback was discovered in that it reflected all overall characteristics contained in the image. For example, considering the goal presented in Figure 1, the intention was to transform only the hairstyle, but the problem arose that skin tone, makeup, and other facial features were also reflected in the synthetic image, resulting in a synthetic image that was slightly different from the original image. As a result, it was necessary to address the weakness of the image being transformed by incorporating unnecessary features that the user did not perceive as their own face.
[0044] Third, there is a problem that the characteristics of the context image (e.g., hairstyle) cannot be maintained depending on the proportion of the target image during image synthesis. In other words, it is necessary to appropriately control the characteristics of the images involved in synthesis according to the purpose.
[0045] The embodiments of this specification, which were discovered from the above problem recognition, propose a technical means that uses a user's actual photograph as an input value, and can perform synthesis by concentrating on only the intended features to be changed from the features of various regions contained in the image while preserving the features of the target image.
[0046] FIG. 3 is a diagram showing an overview of the image synthesis process proposed in the embodiment of this specification.
[0047] When an original image 310, which is a photograph of a user, is input, preprocessing (cropping 320) and alignment processes are performed in consideration of synthesis performance. For example, if the original image 310 is a full-body photograph, is biased to a certain area within the photographic area, or is mixed with many different objects, it is preferable to crop and align only the face and hair area to be centered from the perspective of hairstyle synthesis, which is the target of this embodiment.
[0048] Then, a mask 330, which is pre-designed to identify only the hair region in the image, is input and the image is masked. At the same time, a hair image 340 to be converted is input. At this time, the hair image 340 is assumed to have a different hairstyle from the original image 310, and a target hairstyle to be changed by the user can be input.
[0049] Here, a composite image 360 based on a pre-trained hair model 350 can be generated and output from the masked image and the hair image 340 to be changed. Here, how the deep learning model is trained and used for image synthesis will be described in detail below with reference to FIGS. 5 and 6.
[0050] 4 is a flowchart illustrating a method for compositing images according to an embodiment of the present disclosure. From an implementation perspective, an image compositing device including at least one processor can execute the process defined by each step of FIG. 4, and software including instructions according to each step can be run via the processor.
[0051] In step S410, the image synthesis device learns a hair model having a GAN (Generative Adversarial Networks) structure using multiple learning data related to hairstyles.
[0052] First, existing data on Korean hairstyle images was collected as training data for hair model training. In the process of implementing this embodiment, a total of 500,000 photos were used: 440,000 photos immediately after cosmetic treatment, 10,000 photos of hairstyles with tied up or up, and 50,000 photos of everyday styles. These photos were labeled and segmented using the same standards and can be used as a single dataset or tailored to the purpose of the service. All datasets were 100% augmented through augmentation. The importance of datasets that collect relevant data to perform specific tasks cannot be overemphasized. In particular, the individual data types that make up a dataset, the format of those types, and the quality of the data have a significant impact on AI learning and prediction performance. Therefore, the datasets proposed in this embodiment will be specifically presented below.
[0053] [Table 1]
[0054] The data types in Table 1 are explained below.
[0055] 1) The hair salon uniform dataset can provide beautiful hairstyles immediately after treatment in image format (extension png), Excel file (extension csv), and JSON (JavaScript Object Notation) format. JSON is a character-based standard format for expressing structured data in JavaScript object grammar, and is used when transferring data from web applications. Since exif data can have different schemas depending on the exif tag version, it is recommended to provide it as json rather than csv.
[0056] 2) The hair salon longtail dataset provides hairstyles with a low probability of being beautiful immediately after the procedure in image format (with a png extension), Excel file (with a CSV extension), and JSON format. Long-tail data is necessary for training AI models, but this data is not always readily available. The term "long tail," which is rooted in statistics, refers to the phenomenon in which a large number of unlikely events are distributed along one side of a statistical distribution. The long tail also has a significant impact on the design and operation of AI systems. Existing AI systems are particularly vulnerable to long-tail data. This is because long-tail data, which has a low probability of occurrence, is often not included in the AI training data, which requires a large amount of data.
[0057] 3) The daily hairstyle dataset is for people who have not been to the hairdresser for more than two weeks, so the style is difficult to distinguish at a glance, and the photo background and lighting are diverse and noisy.The dataset can be provided in image format (extension png), Excel file (extension CSV), and JSON format.
[0058] 4) The special hairstyle dataset can be provided in image format (with extension png), Excel file (with extension CSV), or JSON format, of hairstyles that are not done in hair salons but are maintained by many people (e.g., tied-up hair, hair loss, very long hair, etc.).
[0059] The most important consideration when designing a dataset is data balance. Data must be designed to be evenly distributed according to appropriate classification criteria, minimizing data bias expected during learning. In this embodiment, the dataset was constructed by including data on the long tail of hairstyles that are actually frequently ordered, so that both trend and even distribution can be achieved simultaneously.
[0060] The newly collected data, which is hairstyle images collected in this embodiment, is collected by hair shops and hair designers, who are the application areas of the technology, by taking before and after photos of customers, and has the same schema information as the existing data (Korean hairstyle images). An example of the file structure for the newly collected data is as follows.
[0061] The "Annotation.csv" file can have the structure shown in Table 2 below.
[0062] [Table 2]
[0063] Annotation is the process of adding metadata, such as object or image categories, used to describe the original data to a dataset in the form of tags. In other words, it is the process of adding annotations to the source data so that artificial intelligence can understand the content of the data. The description data can be expressed in various forms depending on the purpose of the function. In this case, it is in CSV format. Hairstyle tile name, hairstyle type, hair length, hair color, bangs, degree of hair loss, side hairstyle, age, representative 2D front shot, left / right angle, up / down angle, color, Garma type, gender, special hairstyle classification, segment RGB average, etc. in CSV format. can be provided.
[0064] The "Meta-Annotation.csv" file may have the structure shown in Table 3 below.
[0065] [Table 3]
[0066] Metadata is structured data about data, i.e., data that describes other data. It is data that is attached to content according to certain rules to efficiently find and use the information you are looking for among a large amount of information. Metadata refers to information that follows any data, i.e., structured information, to analyze, classify, and add additional information. In terms of data, this refers to labeling to explain the data. Labeling adds object information, or metadata, when recognizing objects in an image. This allows you to provide the path to the picture file for the hairstyle, the shooting location, photographer, shooting date, hair and face segment coordinates, resolution, shooting equipment, and more in CSV format.
[0067] The "optional-Annotation.csv" file can have the structure shown in Table 4 below.
[0068] [Table 4]
[0069] Optional annotations can provide additional information about the hair, such as the shooting set, hair thickness, water repellency, whether the hair has curls, and damage level, in CSV format.
[0070] The "exifData.csv" file may have a structure as shown in Table 5 below.
[0071] [Table 5]
[0072] Table 5 provides the path where the data is stored in CSV format.
[0073] As mentioned above, to solve the problem of not being able to adjust detailed features during image synthesis, the present embodiment introduces an inversion process that generates noise reflecting features to be projected onto an image generated from an actual photograph while preserving semantic knowledge, which is the feature of the target image. That is, an encoder is provided that converts an image (actual photograph) into noise to enable image-to-image conversion. The encoder generates a latent vector that reflects the image features and can perform various functions, such as converting the pose or facial expression of the image or generating an averaged image by interpolating two images. In this embodiment, the goal is to derive a latent vector that focuses on features related to hairstyle.
[0074] FIG. 5 is a diagram illustrating the configuration of a generator and a segmenter for image synthesis according to an embodiment of the present disclosure, and more specifically illustrates the process of learning the hair model (S410) of FIG.
[0075] A generator 510 receives input of latent vectors in a latent space and generates a fake image. Furthermore, discriminators 531 and 533 receive input of the fake image and a real image and calculate a loss related to the difference between them. The generator 510 learns to generate fake images similar to real images based on the loss, and the discriminators 531 and 533 learn to determine whether the loss is within a threshold based on the loss.
[0076] Unlike typical GAN technology that includes one classifier, the embodiment of the present specification includes at least two classifiers 531 and 533. When changing a hairstyle using a conventional GAN with the goal of changing a hairstyle, a problem occurs in that the user's facial shape also changes. In this embodiment, the classifier does not simply determine how similar a synthesized fake photo is to the real one, but also classifies and determines the face and hairstyle separately within the photo. To achieve this, two types of classifiers are used: one classifier determines whether the face (person) in the predicted photo is the same as the current user's face, and the other classifier determines whether the hairstyle in the predicted photo is the same as the target hairstyle.
[0077] 5, two classifiers are shown, a first classifier 531 for determining whether the fake image and the real image have the same face, and a second classifier 533 for determining whether the fake image and the real image have the same hairstyle. Then, the losses calculated by the first classifier 531 and the second classifier 533 are provided to the generator 510 to simultaneously guide learning of the face and hairstyle.
[0078] Meanwhile, since the first classifier 531 should be trained based on multiple face photos of the same person, it can receive multiple photos of the same person (e.g., Person 1_Photo 1, Person 1_Photo 2, Person 2_Photo 1, Person 2_Photo 2, ...) as a training data set. Furthermore, since the second classifier 533 should be trained based on multiple hairstyle photos of the same hairstyle, it can receive multiple photos of the same hairstyle (e.g., Target Hair 1_Photo 1, Target Hair 1_Photo 2, Same Hair 2_Photo 1, Same Hair 2_Photo 2, ...) as a training data set.
[0079] A careful examination of the two types of classifiers 531 and 533 described above reveals that they require training to determine the identity of faces and the identity of hairstyles, respectively. Therefore, the training data required for the image synthesis device to train a hair model in step S410 of Figure 4 requires not only images related to hairstyles but also images related to faces. For this reason, the training data may include image data including a face region for face training, image data including a hair region for hairstyle training, and data in which the hair region has been masked.
[0080] FIG. 6 is a diagram for explaining a hair model learning process according to an embodiment of the present specification, showing learning using an encoder 610 and a decoder 630.
[0081] First, to extract features from an actual photo, a photo is input to the encoder 610, which outputs an encoded feature. Of course, the input photo must be preprocessed with a photo related to hairstyle before it can learn a hair model for the target hairstyle. The decoder 630 then receives the input of the feature and operates to infer the original photo again. When this series of processes is performed multiple times for various photos, the encoded features of photos with similar hairstyles are displayed as adjacent points in the latent space, while photos with different hairstyles result in the encoded features being far apart in the latent space.
[0082] In summary, the process of learning a hair model can generate a latent space in which similar hairstyles are distributed in a neighboring space by inverting semantic features of hairstyles from an actual image containing multiple hairstyles using the encoder 610. This process can solve the problem of the conventional GAN technology, in which random noise makes it impossible to project the features of an actual photo.
[0083] Here, after the learning of the encoder 610 is completed, the encoded feature contains information about hairstyle regardless of which photo is input, so depending on the convenience of implementation, only the corresponding feature may be provided to the GAN generator.
[0084] Once the hair model has been trained according to the above process, the remaining configuration of this embodiment will be described by returning to FIG.
[0085] In step S430, the image synthesis device receives an image of the user and a hair image including the new hairstyle, where the image of the user may be a real photograph in which various features of the user's appearance are desired to be preserved.
[0086] In step S450, the image combiner masks the image of the user using a mask of hair regions. Various features related to appearance are preserved, but the domain of transformation is controlled so that only the hairstyle is changed.
[0087] In step S470, the image synthesis device generates a synthetic image based on the masked user image and the hair image using the trained hair model. The previously trained hair model includes one generator and two classifiers, and in particular, the generators are simultaneously trained through a first classifier that determines whether the face is the same and a second classifier that determines whether the hairstyle is the same. Therefore, the synthetic image generated by the hair model proposed in this embodiment can naturally reflect only the hairstyle while preserving features other than the target hairstyle (e.g., skin color or makeup) in the original image (the user's actual photo).
[0088] Below, we will introduce platform application technology that utilizes the hairstyle image synthesis method described above.
[0089] 7 is a flowchart illustrating a method for matching a hair designer based on image synthesis according to another embodiment of the present disclosure. From an implementation perspective, a matching system including at least one processor can execute the processing steps defined by each step of FIG. 7, and software including instructions related to each step can be driven by the processor. The processing steps related to image synthesis have been described in detail above with reference to FIGS. 4 to 6, so only an outline thereof will be provided here to avoid repetition.
[0090] In step S710, the matching system receives an input of a user's image. For example, a user who wants to change their hairstyle can provide the user's image to the matching system by taking a real photo of themselves.
[0091] In step S730, the matching system pre-trains a hair model to be used in the image synthesis algorithm. Alternatively, the matching system may be provided with only the results (hair model) learned through another physically separate device.
[0092] In step S750, the matching system sets a desired hairstyle input by a user and generates a composite image corresponding to the hairstyle from the user's image using an image synthesis algorithm. Here, the image synthesis algorithm may learn a hair model with a Generative Adversarial Networks (GAN) structure using a large amount of learning data related to hairstyles, receive input of the user's image and a hair image including a new hairstyle, mask the user's image using a mask for a hair region, and generate a composite image based on the masked user's image and the hair image using the learned hair model.
[0093] In addition, the image synthesis algorithm uses an encoder to invert semantic features of hairstyles from an actual image containing a number of hairstyles, thereby generating a latent space in which similar hairstyles are distributed in an adjacent space; a generator receives a latent vector in the latent space as an input and generates a fake image; a discriminator receives the fake image and a real image as input and calculates a loss related to the difference between them; the generator learns to generate a fake image similar to the real image based on the loss; and the discriminator learns to determine whether the loss is within a threshold based on the loss, thereby training the hair model.
[0094] Furthermore, the classifier includes a first classifier that determines whether the fake image and the real image have the same face, and a second classifier that determines whether the fake image and the real image have the same hairstyle, and the losses calculated through the first classifier and the second classifier are provided to the generator to simultaneously guide learning of faces and hairstyles, and it is preferable that the first classifier is trained based on multiple face photos of the same person, and the second classifier is trained based on multiple hairstyle photos of the same hairstyle.
[0095] In step S770, the matching system recommends a hair designer corresponding to the hairstyle of the composite image generated in step S750. To this end, the matching system can be realized as a collaboration platform that connects hair shops or hair designers working at hair shops with users. That is, multiple hair designers can be registered in the matching system, and a hair designer that meets the user's needs can be recommended based on the hair designer's available treatments and various treatment conditions. When a user selects a recommended hair designer, a convenient function can be provided that allows treatment reservations and payments to be processed within a single platform.
[0096] In summary, in step S770 of recommending a hair designer, at least one or more candidate hair designers can be displayed taking into consideration at least one of the treatment fields and careers of the multiple hair designers. Furthermore, by displaying at least one of the treatment cost, treatment area, and treatment availability date and time of the displayed candidate hair designer, it is possible to encourage the user and the candidate hair designer to make a treatment reservation.
[0097] 8A to 12 are diagrams illustrating an example of a processing flow of an application that realizes a hair designer matching method according to another embodiment of the present specification.
[0098] 8A and 8B show the user interface of the matching application. 8A shows an example of a user interface. First, in FIG. 8A, a user takes a photo of themselves and displays it on the screen so that they can select various aspects of their current hairstyle they would like to change. For example, hair length, wave, hair style, or hair color may be displayed as options. Second, in FIG. 8B, a user loads their own photo from a storage device on a device (e.g., a smartphone) and selects a target photo of another person (e.g., a celebrity) and displays it as the target photo. Next, they can select to predict the results and view a composite image created from their own photo and the target photo. This is a predicted image of the results of the hairstyle treatment.
[0099] Figure 9 shows a composite image displayed on the screen according to the items selected through the user interface of Figure 8A or Figure 8B. When compared with the original photo (user image) of Figure 8A or Figure 8B, the composite image of Figure 9 retains all the features of the same person while showing a very natural image with different hair lengths and waves. The user can then select the designer search button on the screen of Figure 9 to proceed with the matching service of the matching system (platform).
[0100] In Figure 10, hair designers who can be matched according to the user's criteria are displayed. These hair designers are those who can perform the treatment for the composite image created earlier, and results can be displayed with additional search criteria added as needed. For example, only results filtered by additional criteria such as career range or popularity according to the user's preferences can be viewed. At this time, the user can select one hair designer and proceed to the details screen.
[0101] Fig. 11 shows the services that the selected hair designer can provide. The user can select at least one of the services that the selected hair designer can provide and proceed to the reservation screen of Fig. 12.
[0102] In Figure 12, the available treatment times of the selected hair stylist are shown. If necessary, as shown in the example, a single screen including various hair stylist searched previously can be displayed to guide the user to other selections. The user can then complete the reservation by specifying the available treatment time. If necessary, the user may proceed to a payment screen to provide various options for advance payment.
[0103] FIG. 13 is a block diagram showing a hair designer matching system according to another embodiment of the present specification, in which the matching method of FIG. 7 is reconstructed from the viewpoint of hardware configuration.
[0104] The hair designer 10 can be a terminal carried by the hair designer or a reservation terminal at a hair shop, and is connected to the matching system 30 via a network.
[0105] A user 20 is connected to the matching system 30 via a network using a terminal or PC owned by the user.
[0106] The matching system 30 includes a communication unit 31 for connecting with the hair designer 10 and the user 20 via a network, and serves as an intermediary for matching and booking hair shops for the user. The matching system 30 receives a matching request from the user 20, and can load or store matching software including commands defining a series of processing steps for processing the request into a memory 33. The matching system 30 also includes a processor 32 for executing the matching software loaded or stored in the memory 33.
[0107] The matching software includes commands to receive an input of a user's image, set a desired hairstyle input by the user 20, generate a composite image corresponding to the hairstyle from the user's image using an image synthesis algorithm, and recommend a hair designer 10 corresponding to the hairstyle of the generated composite image. Here, the image synthesis algorithm is defined to train a hair model with a Generative Adversarial Networks (GAN) structure using a large amount of learning data related to hairstyles, receive an input of a user's image and a hair image including a new hairstyle, mask the user's image using a mask for a hair region, and generate a composite image based on the masked user's image and the hair image using the trained hair model.
[0108] The matching system proposed in Figure 13 can accumulate personalized data using photo data acquired from customers, and can also accumulate a large amount of learning data on hairstyles using photos of treatment results entered by hair designers. In this case, designers can achieve the goal of being exposed to customers by actively providing their photos to the matching system for marketing purposes, and from the perspective of the matching system, it can be an opportunity to obtain high-quality learning data.
[0109] Meanwhile, the embodiments of the present specification may be embodied as computer-readable codes on a computer-readable recording medium, which may include any type of recording device that stores data that can be read by a computer system.
[0110] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. Furthermore, the computer-readable recording media may be distributed among computer systems connected to a network, and computer-readable code may be stored and executed in a distributed manner. Functional programs, codes, and code segments for implementing the embodiments may be easily construed by programmers skilled in the art to which this specification pertains.
[0111] From the above, the present specification has been carefully reviewed, focusing on various embodiments thereof. Those skilled in the art will understand that the various embodiments can be realized in modified forms without departing from the essential characteristics of the present specification. Therefore, the disclosed embodiments should be considered from an illustrative rather than a restrictive perspective. The scope of the present specification is defined by the claims, not the foregoing description, and all differences within the scope of the claims should be construed as being included in the present specification. [Industrial Applicability]
[0112] According to the embodiments of the present specification, a composite image can be generated from a user's actual photograph in a desired hairstyle by utilizing deep learning technology. In particular, by masking the hair area and providing a hair model that has been trained for each face and hairstyle, it is possible to obtain a change in hairstyle only while preserving unique appearance features. Furthermore, by introducing image synthesis technology into a platform that connects users and hair designers, it is possible to lead to hair designer matching based on the user's changed hairstyle.
Claims
1. 1. A method for compositing an image in an image compositing device including at least one processor, comprising: The image synthesis device learns a hair model having a Generative Adversarial Networks (GAN) structure using a plurality of learning data including image data including a face region for face learning, image data including a hair region for hairstyle learning, and data in which the hair region is masked; the image synthesis device receiving an input of a hair image including a user's image and a new hairstyle; the image combiner masking the image of the user with a mask of hair regions; an image synthesis device using the learned hair model to generate a synthetic image based on the masked image of the user and the hair image; The step of training the hair model includes: A generator generates fake images, a discriminator receives the fake images and real images as input, and distinguishes between them. The two learn by competing with each other. The divider is a first classifier that receives a plurality of facial photographs of the same person as input to a training dataset and trains the training dataset to determine whether the fake image and the real image are the same face; The image synthesis method includes a second classifier that receives a plurality of hairstyle photos of the same hairstyle as a training data set and determines whether the fake image and the real image have the same hairstyle by learning, thereby classifying and determining the face and hairstyle separately in the photo.
2. The step of training the hair model includes: A generator receives a latent vector in a latent space and generates a fake image; a discriminator receiving the fake image and the real image and calculating a loss related to the difference therebetween; the generator learns to generate fake images that resemble real images based on the loss; The image synthesis method of claim 1 , wherein the segmenter learns based on the loss to determine whether the loss is within a threshold.
3. The step of training the hair model includes:
3. The image synthesis method of claim 2, further comprising: generating a latent space in which similar hairstyles are distributed in a neighboring space by inverting semantic features of hairstyles from an actual image containing multiple hairstyles using an encoder.
4. The divider is The image synthesis method of claim 2 , further comprising providing the generator with losses calculated through the first segmenter and the second segmenter, respectively, to simultaneously guide learning of faces and hairstyles.
5. The learning data is The first dataset provides an image of the hairstyle immediately after the treatment, and The second dataset is long-tail data that provides hairstyle images that are unlikely to be used immediately after a treatment. A third dataset providing everyday hairstyle images in which hairstyles are not readily classified; A fourth dataset provides hairstyle images that are not performed but are maintained by a large number of people, The image synthesis method according to claim 1 , wherein the data set is constructed to include data on a portion of long tail hairstyles that are actually popular.
6. 1. A method for matching a hair designer based on image synthesis, wherein a matching system including at least one processor comprises: a matching system receiving an input of a user's image; The matching system sets a desired hairstyle input by a user, and generates a composite image corresponding to the hairstyle from the user's image using an image synthesis algorithm; The matching system recommends a hair designer corresponding to the hairstyle of the generated composite image, The image compositing algorithm comprises: A hair model having a Generative Adversarial Network (GAN) structure is trained using a plurality of training data including image data including a face region for face training, image data including a hair region for hairstyle training, and data in which the hair region is masked; a user's image and a hair image including a new hairstyle are input, and the user's image is masked using a hair region mask; and a composite image is generated based on the masked user's image and the hair image using the trained hair model; The hair model is generated by a generator to generate fake images, The discriminator receives the fake and real images as inputs, distinguishes the differences between them, and the two are trained while competing with each other. The divider is a first classifier that receives a plurality of facial photographs of the same person as input to a training dataset and learns the input, and determines whether the fake image and the real image are the same face; A hair designer matching method including a second classifier that determines whether the fake image and the real image have the same hairstyle by inputting a plurality of hairstyle photos of the same hairstyle into a training dataset and learning the data, thereby classifying and determining the face and hairstyle separately in the photo.
7. The step of recommending a hair designer includes: The hair designer matching method according to claim 6, further comprising displaying at least one or more hair designer candidates in consideration of at least one of the fields of practice and careers of the plurality of hair designers.
8. The step of recommending a hair designer includes: The hair designer matching method according to claim 7, further comprising a step of guiding the user to make a treatment reservation between the candidate hair designer and the candidate hair designer by displaying at least one of the treatment cost, treatment location, and treatment available date and time of the displayed candidate hair designer.
9. The image compositing algorithm comprises: The encoder inverts the semantic features of hairstyles from an actual image containing multiple hairstyles to generate a latent space in which similar hairstyles are distributed in adjacent spaces. The generator receives the latent vector in the latent space and generates a fake image.
7. The hair designer matching method of claim 6, wherein a discriminator receives input of the fake image and the real image and calculates a loss relating to the difference therebetween, the generator learns to generate a fake image similar to the real image based on the loss, and the discriminator learns to determine whether the loss is within a threshold based on the loss, and the losses calculated through the first discriminator and the second discriminator are provided to the generator to simultaneously guide learning for the face and hairstyle, thereby learning the hair model.
10. The learning data is The first dataset provides an image of the hairstyle immediately after the treatment, and The second dataset is long-tail data that provides hairstyle images that are unlikely to be used immediately after a treatment. A third dataset providing everyday hairstyle images in which hairstyles are not readily classified; A fourth dataset provides hairstyle images that are not performed but are maintained by a large number of people, The hair designer matching method according to claim 6, wherein the data set is constructed so as to compile data on a portion where hairstyles that are actually frequently ordered are long-tailed.
Citation Information
Patent Citations
Image generation system and image generation method using the same
JP2021190062A
Changing the appearance of the hair
JP2022527370A
Method, apparatus, and system for providing virtual fitting service
WO2022039450A1
Try-on with reverse gans
WO2022184858A1