Method for operating information processing device, information processing device, and program

The information processing device uses machine learning to enhance the accuracy and efficiency of selecting makeup products by classifying features of face images, addressing the limitations of subjective user judgment in determining similarity.

JP2025131424APending Publication Date: 2025-09-09KOSE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024029160
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The selection of makeup products based on user preferences is often inaccurate and inefficient due to reliance on subjective user judgment and intuition in determining similarity between desired and sample images.

Method used

An information processing device employs machine learning to generate and classify features of face images, using a shape extraction unit and a color extraction unit to accurately match the desired image with similar samples, reducing reliance on user subjectivity.

Benefits of technology

This approach enables more accurate and efficient selection of makeup products by classifying features of face images, improving the matching process and reducing dependence on user intuition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131424000001_ABST
    Figure 2025131424000001_ABST
Patent Text Reader

Abstract

To execute selection of a sample that resembles an image desired by a user with higher accuracy and efficiently.SOLUTION: A method for operating an information processing device includes: a first step of, using a first feature extracted from a first subject image by an extraction unit and a second feature extracted from a second subject image by another extraction unit, generating a third subject image having the first and second features by a generation unit; and a second step of causing the extraction unit to perform machine learning so as to classify a third feature extracted from the third subject image by the extraction unit and the first feature into the same group.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an operation method of an information processing device, an information processing device, and a program. [Background technology]

[0002] In the beauty industry, various techniques have been proposed for providing users with product recommendation information when selecting makeup products to express a desired makeup look (for example, Patent Documents 1 and 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2023-81971 [Patent Document 2] International Publication No. 2022 / 176279 Summary of the Invention [Problem to be solved by the invention]

[0004] When a user selects a makeup product, the user may select a sample that closely resembles the desired image from among samples of face images with makeup applied, and then select a product to be applied to the selected sample. In this selection process, the determination of the similarity between the desired image and the sample largely depends on the user's subjectivity and intuition, and there is room for improvement in accuracy and efficiency.

[0005] In view of the above, the following discloses an operation method of an information processing device that enables more accurate and efficient selection of a sample that approximates a user's desired image. [Means for solving the problem]

[0006] In order to solve the above problem, the operating method of the information processing device in the present disclosure includes a first step of generating a third object image having a first feature extracted by an extraction unit from a first object image and a second feature extracted by another extraction unit from a second object image, using the first and second features; and a second step of causing the extraction unit to perform machine learning to classify the third feature extracted by the extraction unit from the third object image and the first feature into the same group.

[0007] In addition, the information processing device of the present disclosure includes a memory unit that stores a model in which a generation unit generates a third object image having the first and second features using a first feature extracted by an extraction unit from a first object image and a second feature extracted by another extraction unit from a second object image, and a control unit that causes the extraction unit to perform machine learning to classify the third feature extracted by the extraction unit from the third object image and the first feature into the same group.

[0008] Furthermore, the program in the present disclosure causes an information processing device to execute a first step of generating a third object image having the first and second features by a generation unit using a first feature extracted from a first object image by an extraction unit and a second feature extracted from a second object image by another extraction unit, and a second step of causing the extraction unit to perform machine learning so as to classify the third feature extracted from the third object image by the extraction unit and the first feature into the same group. [Effects of the Invention]

[0009] According to the operation method of the information processing device and the like of the present disclosure, it becomes possible to select a sample that is close to the image desired by the user with higher accuracy and efficiency. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 illustrates an example of the configuration of an information processing system. [Figure 2]FIG. 2 is a diagram illustrating an example of the configuration of an image transformation model. [Figure 3] FIG. 10 is a flowchart illustrating an example of an operation procedure of the server device. [Figure 4] FIG. 10 is a flowchart illustrating an example of an operation procedure of the server device. [Figure 5] 10A and 10B are diagrams illustrating a face image used for adjusting an image transformation model. [Figure 6] FIG. 10 is a flowchart illustrating an example of an operation procedure of the server device. [Figure 7] FIG. 10 is a flowchart illustrating an example of an operation procedure of the server device. [Figure 8A] FIG. 10 is a diagram illustrating the shape characteristics of a face image. [Figure 8B] FIG. 10 is a diagram illustrating color features of a facial image. [Figure 9] FIG. 10 is a flowchart illustrating an example of an operation procedure of a server device according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described.

[0012] [System Configuration] FIG. 1 is a diagram illustrating an example configuration of an embodiment of the present invention. The information processing system 1 includes a server device 10 and a terminal device 12 connected to each other via a network 11 so as to be able to communicate with each other. In the information processing system 1, the server device 10 performs information processing, including machine learning, using various information sent from the terminal device 12. The terminal device 12 is, for example, one or more personal computers. The personal computer may include a tablet terminal device, a smartphone, etc. The server device 10 corresponds to the "information processing device" in this embodiment. The server device 10 is, for example, one or more server computers. When the server device 10 is a single server computer, the server device 10 may be multiple server computers that cooperate to execute the operations of this embodiment and provide a cloud service. The network 11 is, for example, a local area network (LAN), the Internet, an ad hoc network, a metropolitan area network (MAN), a mobile communication network, or other networks, or any combination thereof.

[0013] The server device 10 acquires a subject image from the terminal device 12, performs machine learning using the subject image, and executes a step of generating an image transformation model 108 (hereinafter referred to as a model generation step). The subject image is a captured image obtained by capturing an image of a person's face as a subject. The image transformation model 108 is an image transformation model that adds makeup information of a captured image of the face of the same person or another person with makeup applied (hereinafter referred to as a reference face image) to a captured image of a face (an image including the entire face of a person as seen from the front, hereinafter referred to as an original face image). As shown in FIG. 2, the image transformation model 108 includes functional modules such as an extraction unit 21 (hereinafter referred to as a shape extraction unit) that extracts shape features of various parts of the face from the original face image Isrc, an extraction unit 22 (hereinafter referred to as a color extraction unit) that extracts color features of various parts of the face from the reference face image Iref, and a generation unit 23 that uses the shape features extracted from the original face image Isrc and the color features extracted from the reference face image Iref to generate a face image (hereinafter referred to as a transformed face image) Itransfer that has the shape features of the original face image Isrc and the color features of the reference face image Iref. Here, the transformed face image Itransfer is a face image obtained by adding the color features of the reference face image Iref, i.e., makeup information, to the original face image Isrc. The shape extraction unit 21 or the color extraction unit 22 each has an Lp norm layer (p is an arbitrary positive number) including a scale parameter in its output portion. The outputs of the shape extraction unit 21 and the color extraction unit 22 are added and multiplied and input to the generation unit 23, which generates a transformed facial image Itransfer based on this input.

[0014] The server device 10 also executes a step of adjusting the image transformation model 108 (hereinafter referred to as a model adjustment step) using an image obtained by adding makeup information to the original facial image Isrc in a predetermined procedure (hereinafter referred to as a secondary facial image) so that the attributes of the original facial image Isrc are not lost when the original facial image Isrc is transformed by the image transformation model 108. Specifically, the server device 10 executes an algorithm processing step of adding makeup information to the original facial image Isrc in a predetermined procedure and transforming it into a secondary facial image. The server device 10 also transforms the original facial image Isrc into a transformed facial image Itransfer using the secondary facial image as a reference facial image Iref, using the image transformation model 108 generated by machine learning the process of transforming the original facial image Isrc by adding makeup information included in the secondary facial image. The server device 10 then adjusts the image transformation model 108 by reducing loss of the original facial image Isrc in the transformed facial image Itransfer.

[0015] Furthermore, the server device 10 executes a step (hereinafter referred to as an image transformation step) of generating a transformed face image Itransfer from the original face image Isrc and the reference face image Iref using the image transformation model 108. For example, the server device 10 uses shape features extracted from the original face image Isrc by the shape extraction unit 21 and color features extracted from the reference face image Iref by the color extraction unit 22 to generate, by the generation unit 23, a transformed face image Itransfer having the shape features of the original face image Isrc and the color features of the reference face image Iref.

[0016] The server device 10 then executes a step of causing the shape extraction unit 21 to undergo machine learning so as to classify shape features extracted by the shape extraction unit 21 from the original face image Isrc and the transformed face image Itransfer into the same group, and shape features extracted by the shape extraction unit 21 from the reference face image Iref and the transformed face image Itransfer into different groups. Additionally, the server device 10 may execute a step of causing the color extraction unit 22 to undergo machine learning so as to classify color features extracted by the color extraction unit 22 from the reference face image Iref and the transformed face image Itransfer into the same group, and color features extracted by the color extraction unit 22 from the original face image Isrc and the transformed face image Itransfer into different groups. Hereinafter, this machine learning step by the shape extraction unit 21 and the color extraction unit 22 will be referred to as a feature design step. The machine learning in the feature design step is machine learning without teacher data.

[0017] According to this embodiment, the server device 10 trains the shape extraction unit 21 to classify shape features extracted by the shape extraction unit 21 from the original face image Isrc and the transformed face image Itransfer, which are likely to have similar shape features, into the same group. This improves the accuracy with which the shape extraction unit 21 extracts similar shape features from different subject images. Furthermore, the server device 10 trains the color extraction unit 22 to train the reference face image Iref and the transformed face image Itransfer, which are likely to have similar color features, into the same group. This improves the accuracy with which the color extraction unit 22 extracts similar color features from different subject images. The shape extraction unit 21 and the color extraction unit 22 that have undergone this type of machine learning can be used to select sample face images whose shape or color is similar to that contained in a user's desired image, as described below. This makes it possible to more accurately and efficiently select sample face images that are similar to a user's desired image.

[0018] [Configuration example of server device 10] The server device 10 includes a communication unit 101, a storage unit 102, a control unit 103, an input unit 105, and an output unit 106. When the server device 10 is configured with two or more server computers, these components are appropriately arranged in the two or more server computers.

[0019] The communication unit 101 includes one or more communication interfaces. The communication interface is, for example, a LAN interface. The communication unit 101 receives information used in the operation of the server device 10 and transmits information obtained by the operation of the server device 10. The server device 10 is connected to a network 11 by the communication unit 101 and communicates information with a terminal device 12 via the network 11.

[0020] The storage unit 102 includes, for example, one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of them, that function as a main storage device, an auxiliary storage device, or a cache memory. The semiconductor memory is, for example, a random access memory (RAM) or a read-only memory (ROM). The RAM is, for example, a static RAM (SRAM) or a dynamic RAM (DRAM). The ROM is, for example, an electrically erasable programmable read-only memory (EEPROM). The storage unit 102 stores information used in the operation of the control unit 103 and information obtained by the operation of the control unit 103. The storage unit 102 stores an image transformation model 108 generated by the control unit 103 based on information sent from the terminal device 12.

[0021] The control unit 103 includes one or more processors, one or more dedicated circuits, or a combination thereof. The processor is, for example, a general-purpose processor such as a central processing unit (CPU), or a dedicated processor such as a graphics processing unit (GPU) specialized for a specific process. The dedicated circuit is, for example, a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The control unit 103 executes information processing related to the operation of the server device 10 while controlling each unit of the server device 10. The shape extraction unit 21, color extraction unit 22, and generation unit 23, which are functional modules of the image transformation model 108, are configured by the control unit 103, which executes procedures related to the image transformation model 108.

[0022] The functions of the server device 10 are realized by a processor included in the control unit 103 executing a control program. The control program is a program for causing the processor to function as the control unit 103. Alternatively, some or all of the functions of the server device 10 may be realized by a dedicated circuit included in the control unit 103. Alternatively, the control program may be stored in a non-transitory recording / storage medium readable by the control unit 103, and read by the control unit 103 from the medium.

[0023] The input unit 105 includes one or more input interfaces. The input interfaces are, for example, physical keys, capacitance keys, a pointing device, a touch screen integrated with a display, or a microphone that accepts voice input. The input unit 105 accepts an operation to input information used in the operation of the server device 10 and sends the input information to the control unit 103.

[0024] The output unit 106 includes one or more output interfaces. The output interface is, for example, a display or a speaker. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display. The output unit 106 outputs information obtained by the operation of the server device 10.

[0025] [Configuration example of terminal device 12] The terminal device 12 includes a communication unit 121 , a storage unit 122 , a control unit 123 , an input unit 125 , and an output unit 126 .

[0026] The communication unit 121 includes a communication module compatible with wired or wireless LAN standards, a module compatible with mobile communication standards such as LTE, 4G, 5G, etc. The terminal device 12 is connected to the network 11 by the communication unit 121 via a nearby router device or a mobile communication base station, and performs information communication with the server device 10, etc. via the network 11.

[0027] The storage unit 122 includes one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of these. The semiconductor memories are, for example, RAM or ROM. The RAM is, for example, SRAM or DRAM. The ROM is, for example, EEPROM. The storage unit 122 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 122 stores information used in the operation of the control unit 123 and information obtained by the operation of the control unit 123.

[0028] The control unit 123 has, for example, one or more general-purpose processors such as a CPU, an MPU (Micro Processing Unit), etc., or one or more dedicated processors such as a GPU specialized for a specific process. Alternatively, the control unit 123 may have one or more dedicated circuits such as an FPGA, an ASIC, etc. The control unit 123 performs overall control of the operation of the terminal device 12 by operating according to a control / processing program or operating according to an operating procedure implemented as a circuit. The control unit 123 then transmits and receives various information to and from the server device 10, etc. via the communication unit 121, and performs the operation according to this embodiment.

[0029] The functions of the terminal device 12 are realized by a processor included in the control unit 123 executing a control program. The control program is a program for causing the processor to function as the control unit 123. Alternatively, some or all of the functions of the terminal device 12 may be realized by a dedicated circuit included in the control unit 123. Alternatively, the control program may be stored in a non-transitory recording / storage medium readable by the control unit 123, and read by the control unit 123 from the medium.

[0030] The input unit 125 includes one or more input interfaces. The input interfaces include, for example, physical keys, capacitive keys, a pointing device, and a touch screen integrated with a display. The input interfaces also include a microphone for receiving voice input and a camera for capturing captured images. The input interfaces may also include a scanner or camera for scanning image codes, and an IC card reader. The input unit 125 receives an operation for inputting information used in the operation of the control unit 123 and sends the input information to the control unit 123. The input unit 125 also sends images captured by the camera to the control unit 123.

[0031] The output unit 126 includes one or more output interfaces. The output interfaces include, for example, a display and a speaker. The display is, for example, an LCD or an organic EL display. The output unit 126 outputs information obtained by the operation of the control unit 123.

[0032] [Model generation process] 3 is a flowchart illustrating an example of the operation of the server device 10 related to the generation of the image transformation model 108. Each step is executed by the control unit 103.

[0033] In step S20, the control unit 103 acquires face images necessary for machine learning to generate an image transformation model. The face images include an original face image Isrc to be transformed and a reference face image Iref of a face with makeup applied. The original face image Isrc is generated by capturing an image of a person's face. The reference face image Iref is generated by capturing an image of the person's face with makeup applied. The image of the person's face is captured by, for example, the terminal device 12. For example, the control unit 103 receives multiple original face images Isrc and multiple reference face images Iref sent from the terminal device 12 via the communication unit 101 and stores them in the storage unit 102. The control unit 103 may acquire multiple original face images Isrc and reference face images Iref from open data.

[0034] In step S22, the control unit 103 performs machine learning. The control unit 103 performs deep learning using, for example, GAN. The control unit 103 has modules (e.g., a shape extraction unit 21, a color extraction unit 22, and a generation unit 23) corresponding to a generator that adds makeup information of a reference face image to an original face image, and a module corresponding to a classifier that classifies the transformed face image Itransfer generated by the generator from the original face image Isrc. The control unit 103 generates an image transformation model 108 by training the generator and the classifier in an adversarial manner. The control unit 103 stores the generated image transformation model 108 in the storage unit 102.

[0035] [Model adjustment process] Fig. 4 is a flowchart illustrating an example of the operation of the server device 10 related to the adjustment of the image transformation model 108. Each step is executed by the control unit 103. Fig. 5 is a diagram illustrating a face image used for adjusting the image transformation model 108. The procedure of Fig. 4 will be described with reference to Fig. 5.

[0036] In step S30, the control unit 103 acquires an original face image Isrc for executing the algorithm processing step. The original face image Isrc is generated by, for example, the terminal device 12 capturing an image of a person's face. For example, the control unit 103 receives the original face image Isrc sent from the terminal device 12 via the communication unit 101 and stores it in the storage unit 102. The control unit 103 may acquire multiple original face images Isrc from open data.

[0037] In step S32, the control unit 103 executes an algorithm processing step. The control unit 103 assigns predetermined makeup information MU to the original face image Isrc and converts the original face image Isrc into a reference face image Isyn. The makeup information MU includes one or more of hue, brightness, and saturation to be assigned to areas in the original face image Isrc, including the eyes, nose, cheeks, and lips (hereinafter referred to as makeup areas). The makeup information MU and the procedure for assigning the makeup information are image processing procedures that are arbitrarily set in advance. The hue, brightness, and saturation to be assigned to the makeup areas may be arbitrarily determined quantitatively, or color information already assigned to a reference face image may be extracted and applied. The makeup information MU may be input by an operator at the terminal device 12 and sent from the terminal device 12 to the server device 10. The control unit 103 executes image processing to apply eyeliner, nose shadow, cheek color, lip color, or the like of a color arbitrarily set in the makeup information MU to the eyes, nose bridge, cheeks, or lips of the face image Isrc.

[0038] In step S34, the control unit 103 executes a makeup conversion process. The control unit 103 inputs the original face image Isrc as a face image to be converted into the image conversion model 108, together with a reference face image Iref having desired makeup information and linked to the reference face image Isyn. The reference face image Iref may be a reference face image Isyn simulated by the user to match the user's own image, or may be arbitrarily selected from pre-prepared reference face images Iref of other people and linked to the reference face image Isyn. The reference image Iref includes, for example, face images with various types of makeup, such as elegant, cool, and trendy, and face images of real people. The image transformation model 108 extracts shape features such as facial shape, three-dimensionality, and surface condition from the original face image Isrc and makeup information, i.e., color features, from the reference face image Iref using a shape extraction unit 21 and a color extraction unit 22, respectively, and generates a transformed face image Itransfer by adding the makeup information of the reference face image Iref to the original face image Isrc using a generation unit 23.

[0039] In step S36, the control unit 103 executes an adjustment step. When the transformed face image Itransfer is used as a parameter of a loss function L, the control unit 103 adjusts the parameters of the image transformation model 108 so as to minimize the value of the loss function L. The loss function L includes, for example, one or more of Adversarial loss, Makeup loss, Perceptual loss, and Feature matching loss.

[0040] Adversarial loss is a loss function for training a generator to deceive a classifier in a GAN. The control unit 103 uses a classifier to classify whether the transformed face image Itransfer is the result of a makeup transformation process or the original face image Isrc, and uses the result to train the generator to deceive the classifier. In this way, parameters of the image transformation model 108 are adjusted so that the loss of the original face image Isrc in the transformed face image Itransfer is reduced. The classifier may include a global classifier and a local classifier. The global classifier uses the entire face image to classify whether the transformed face image Itransfer is the result of a makeup transformation process or the original face image Isrc. The local classifier uses makeup areas to classify whether the transformed face image Itransfer is the result of a makeup transformation process or the original face image Isrc. Using both a global classifier and a local classifier makes it possible to improve the learning accuracy of the generator.

[0041] Makeup loss is a loss function for training the generator regarding color distribution. The control unit 103 generates a pseudo-face image having the same color distribution as the reference face image Iref by histogram matching based on the reference face image Iref. The histogram matching may be performed on the color distribution of the entire face image or on the color distribution of the makeup-applied areas. The control unit 103 then causes a classifier to distinguish whether the pseudo-face image or the transformed face image Itransfer is the result of the histogram matching or makeup conversion process or the original face image Isrc, and trains the generator to deceive the classifier using the result. In this way, the parameters of the image conversion model 108 are adjusted so that the loss of the original face image Isrc in the transformed face image Itransfer is reduced.

[0042] Perceptual loss is a loss function for training the generator regarding the contours of a facial image. The control unit 103 acquires a facial image during conversion from the intermediate layer when the generator converts the reference facial image Iref into the converted facial image Itransfer, and generates an edge image by extracting the edges of each part of the image, such as the eyes, nose, and mouth. The control unit 103 then causes a classifier to distinguish whether the edges of the edge image are edges of the edge image resulting from the makeup conversion process or edges of the original facial image Isrc, and uses the result to train the generator to deceive the classifier. In deep learning, edge information of the image to be converted is extracted as a feature in the intermediate layer. By performing training using the edge image in the intermediate layer, the parameters of the image conversion model 108 are adjusted to reduce the loss related to the edges of the original facial image Isrc in the converted facial image Itransfer. In other words, the parameters of the image conversion model 108 are adjusted so that the facial contours in the converted facial image Itransfer match the facial contours in the original facial image Isrc.

[0043] The feature matching loss is a loss function related to features such as the shape or color of a face image. The control unit 103 adjusts the parameters of the image transformation model 108 so as to minimize the following formula 1 or formula 2, which includes cosine similarity. TIFF2025131424000002.tif22169 (In Equation 1 and Equation 2, TIFF2025131424000003.tif7164 is the parameter of the feature generator for extracting the shape and color of the face image. By doing so, the parameters of the image transformation model 108 are adjusted so that loss of shape of the original face image Isrc in the transformed face image Itransfer is reduced, or so that loss of color of the original face image Isrc in the transformed face image Itransfer is reduced.

[0044] [Image conversion process] 6 is a flowchart for explaining an example of the operation of the server device 10 relating to the image transformation process using the image transformation model 108. Each step is executed by the control unit 103. The procedure of FIG. 6 will be explained with reference to FIG. 2.

[0045] In step S60, the control unit 103 acquires an original face image Isrc and a reference face image Iref for executing the image conversion process. The control unit 103 receives the original face image Isrc and the reference face image Iref sent from the terminal device 12 via the communication unit 101 and stores them in the storage unit 102. The control unit 103 may acquire multiple original face images Isrc and reference face images Iref from open data.

[0046] In step S61, the control unit 103 causes the shape extraction unit 21 to extract shape features from the original face image Isrc, and the color extraction unit 22 to extract color features from the reference face image Iref.

[0047] In step S62, the control unit 103 uses the shape features extracted from the original face image Isrc and the color features extracted from the reference face image Iref to generate a transformed face image Itransfer having the shape features of the original face image Isrc and the color features of the reference face image Iref using the generation unit 23.

[0048] In step S63, the control unit 103 stores the original face image Isrc, the reference face image Iref, and the generated transformed face image Itransfer in the storage unit .

[0049] [Feature design process] Fig. 7 is a flowchart for explaining an example of the operation of the server device 10 relating to the feature design process. Each step is executed by the control unit 103. Figs. 8A and 8B are diagrams for schematically explaining the feature design process. The procedure of Fig. 7 will be explained with reference to Figs. 8A and 8B.

[0050] In step S70, the control unit 103 reads out and acquires from the storage unit 102 the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer that have been processed in the image transformation step.

[0051] In step S71, the control unit 103 extracts shape features from each of the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer using the shape extraction unit 21. As shown in Fig. 8A, the shape features of each of the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer are extracted as vector components 81, 82, and 83 in an N-dimensional (N is an integer equal to or greater than 2) vector space 80.

[0052] In step S72, the control unit 103 causes the shape extraction unit 21 to perform machine learning by adjusting hyperparameters of the shape extraction unit 21 so as to group shape features extracted from the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer. Specifically, the control unit 103 classifies the shape features of the original face image Isrc and the transformed face image Itransfer into a single group by decreasing the similarity D1 in the vector space 80 of the corresponding vector components 81 and 83, and classifies the shape features of the original face image Isrc and the reference face image Iref into different groups by increasing the similarity D2 in the vector space 80a of the corresponding vector components 81 and 82. The increase or decrease in the similarity in the vector space 80a of the vector components corresponding to each feature is performed by, for example, random oblivion, hypersphere expansion, or the like.

[0053] Random oblivion is a method that reduces the similarity D1 between the vector components 81 and 83 corresponding to the shape features of the original face image Isrc and the transformed face image Itransfer with an arbitrary probability P (where 0 < P < 1), and increases the similarity D2 between the vector components 81 and 83 corresponding to the shape features of the original face image Isrc and the reference face image Iref with a probability of (1 - P). The probability P can be set, for example, by the user sending information from the terminal device 12 to the server device 10 in advance. Here, the similarities D1 and D2 are represented by the cosine similarities of the following equations 3 and 4, respectively. TIFF2025131424000004.tif21170 (In equations 3 and 4, TIFF2025131424000005.tif9149 is a parameter of the feature generator for extracting the shape of the face image)

[0054] Hypersphere expansion is a method that reduces the similarity D1 between the vector components 81 and 83 corresponding to the shape features of the original face image Isrc and the transformed face image Itransfer with an arbitrary probability P (where 0 < P < 1), and pseudo-increases the similarity D2 between the vector components 81 and 83 by increasing the radius of the hypersphere in which the vector components 81 and 83 corresponding to the shape features of the original face image Isrc and the reference face image Iref are included with a probability of (1 - P). The probability P can be set, for example, by the user sending information from the terminal device 12 to the server device 10 in advance. Here, the similarity D1 is represented by the cosine similarity of the above equation 3. Also, the loss function of the similarity D2 is represented by the following equation 5, and the parameters of the shape extraction unit 21 are adjusted so as to minimize this loss function. After the parameter adjustment, the value of the radius of the hypersphere is returned to the original value. TIFF2025131424000006.tif11169 (In equation 5, TIFF2025131424000007.tif9157 is a parameter of the feature generator for extracting the shape of the face image)

[0055] In step S73, the control unit 103 extracts color features from the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer by the color extraction unit 22. As shown in FIG. 8B, the shape features of the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer are extracted as vector components 85, 86, and 87 in an N-dimensional (N is an integer of 2 or more) vector space 80b.

[0056] In step S74, the control unit 103 adjusts the hyperparameters of the color extraction unit 22 to group the color features extracted from the original face image Isrc, the reference face image Iref, and the transformed face image Itransfer, causing the color extraction unit 22 to perform machine learning. Specifically, the control unit 103 classifies the color features of the reference face image Iref and the transformed face image Itransfer into a single group by reducing the similarity D3 in the vector space 80b of the corresponding vector components 86, 87, and classifies the color features of the original face image Isrc and the reference face image Iref into different groups by increasing the similarity D4 in the vector space 80b of the corresponding vector components 85, 86. The increase and decrease of the similarity in the vector space 80b of the vector components corresponding to each feature are performed by methods such as Random oblivion, Hypersphere expansion, etc.

[0057] In Random oblivion, the control unit 103 decreases the similarity D3 between the vector components 86, 87 corresponding to the color features of the original face image Isrc and the reference face image Iref with an arbitrary probability P (where 0 < P < 1), and increases the similarity D4 between the vector components corresponding to the color features of the original face image Isrc and the reference face image Iref with a probability (1 - P). Here, the similarities D3 and D4 are represented by the cosine similarities of the following equations 6 and 7, respectively. TIFF2025131424000008.tif21166 (In equations 6 and 7, TIFF2025131424000009.tif8153 is a parameter of the feature generator for extracting makeup information of the face image)

[0058] In hypersphere expansion, the control unit 103 decreases the similarity D3 between the vector components 86 and 87 corresponding to the color features of the original face image Isrc and the reference face image Iref with an arbitrary probability P (where 0 < P < 1), and increases the radius of the hypersphere including the vector components 85 and 86 corresponding to the color features of the original face image Isrc and the reference face image Iref with a probability (1 - P), thereby pseudo-increasing the similarity D4 between the vector components 85 and 86. Here, the similarity D3 is represented by the cosine similarity of the above formula 6. Further, the loss function of the similarity D4 is represented by the following formula 8, and the parameters of the color extraction unit 22 are adjusted so as to minimize this loss function. Note that after the parameter adjustment, the value of the radius of the hypersphere is returned to the original value. TIFF2025131424000010.tif10169 (in formula 8, TIFF2025131424000011.tif8150 is a parameter of the feature quantity generator for extracting the makeup information of the face image)

[0059] According to the present embodiment, it is not necessary to extract the feature quantity of the makeup information in the face image and create the accompanying teacher data. When creating a model of the feature quantity generator for application to calculations such as cluster analysis, teacher data is required, and a great deal of work cost is required for creating the teacher data. However, according to the present embodiment, learning can be executed only with the input data, so that the work cost can be reduced.

[0060] [Example] The shape extraction unit 21 and color extraction unit 22, which have undergone machine learning according to the above-described procedure, output, for example, vector components in a vector space corresponding to shape and color features, respectively. For example, if the shape extraction unit 21 extracts shape features as vector components from multiple different face images and derives the similarity between each component, it becomes possible to determine the similarity according to the similarity. In other words, the smaller the similarity, the more similar the shape features of the face images can be determined. Also, if the color extraction unit 22 extracts color features as vector components from multiple different face images and derives the similarity between each component, it becomes possible to determine the similarity according to the similarity. In other words, the smaller the similarity, the more similar the color features of the face images can be determined.

[0061] 9 is a sequence diagram for explaining an example of the operation of the information processing system 1 in the example. The procedure in Fig. 9 shows the procedure of the operation by the control unit 103 of the server device 10 having the shape extraction unit 21 and the color extraction unit 22 that have undergone machine learning according to the procedure of this embodiment.

[0062] In step S90, the control unit 103 acquires a user face image and a plurality of sample face images. The user face image is a captured image of the user's face. The sample face image is a captured image of the face of a person other than the user before or after makeup is applied. The control unit 103 acquires the user face image and the plurality of sample face images sent from the terminal device 12, for example. Alternatively, the control unit 103 may acquire the plurality of sample face images from open data.

[0063] In step S91, the control unit 103 extracts shape or color features of the user face image and each of the plurality of sample face images using the shape extraction unit 21 and the color extraction unit 22. The shape or color features of each face image are output as vector components and stored in the storage unit 102 in association with the identification information of each face image.

[0064] In step S92, the control unit 103 derives a similarity of each sample face image to the user face image. The similarity is derived, for example, as a similarity of shape features of each sample face image to shape features of the user face image, or a similarity of color features of each sample face image to color features of the user face image.

[0065] In step S93, the control unit 103 outputs information based on the similarity. The information based on the similarity is, for example, information obtained by scoring the similarity of the shape features of each sample face image to the shape features of the user face image, and sample face images corresponding to each score. Alternatively, the information based on the similarity is, for example, information obtained by scoring the similarity of the color features of each sample face image to the color features of the user face image, and sample face images corresponding to each score. The sample face images may be sorted in ascending or descending order of score. The control unit 103 may output any number of top-ranked sample images input by the user from the terminal device 12. The information based on the similarity is, for example, sent from the server device 10 to the terminal device 12 and displayed on the terminal device 12.

[0066] According to the above-described procedure, the user can grasp sample face images that are similar to the user's facial shape or color characteristics according to the degree of similarity. This reduces the degree of dependence on the user's subjectivity and intuition, and makes it possible to more accurately and efficiently select samples that are close to the user's desired image.

[0067] This embodiment can be applied not only to adding makeup information to a facial image but also to determining similarity between images. For example, this embodiment can be applied to identifying and extracting sample images of hairstyles similar to an image of a hairstyle specified by a user, or identifying and extracting sample images of paintings similar to an image of a painting specified by a user. Even in such cases, it is possible to determine similar images with higher accuracy and efficiency than when a user makes a subjective or intuitive judgment.

[0068] In the above description, the server device 10 corresponds to the "information processing device." However, the server device 10 and the terminal device 12 may cooperate to configure the "information processing device," or the terminal device 12 may correspond to the "information processing device."

[0069] In the above-described embodiment, the processing / control program that defines the operation of the terminal device 12 may be stored in the memory unit 102 of the server device 10 or in the memory unit of another server device, and may be downloaded to the terminal device 12 via the network 11, or may be stored in a computer-readable non-transitory recording / storage medium and read by the terminal device 12 from the medium.

[0070] Although the embodiments have been described above based on the drawings and examples, it should be noted that those skilled in the art can easily make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each means, step, etc. can be rearranged so as not to be logically inconsistent, and multiple means, steps, etc. can be combined or divided into one. [Explanation of symbols]

[0071] 10: Server device 11: Network 12: Terminal device 101, 121: Communications Department 102, 122: Storage section 103, 123: control unit 105, 125: Input section 106, 126: Output section 108: Image transformation model 21: Shape extraction part 22: Color extraction section 23: Generation part Isrc: Original face image Iref: Reference face image iTransfer: Converted face image L: Loss function

Claims

1. A method for operating an information processing device, comprising: a first step of generating, by a generating unit, a third object image having the first and second features, using a first feature extracted from a first object image by an extracting unit and a second feature extracted from a second object image by another extracting unit; a second step of causing the extraction unit to perform machine learning so as to classify the third feature extracted by the extraction unit from the third object image and the first feature into the same group; A method of operation having the following steps:

2. In claim 1, the second step further includes causing the extraction unit to perform machine learning to classify a fourth feature extracted by the extraction unit from the second object image and the first feature into different groups. How it works.

3. In claim 1, the second step includes a step of classifying the first and third features into the same group by reducing a distance between the first and third features in a vector space, and classifying the second and fourth features into different groups by increasing a distance between the second feature and a fourth feature extracted from the second object image by the extraction unit in the vector space. How it works.

4. In claim 2, the second step includes a step of decreasing the distance between the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1) and increasing the distance between the fourth feature and the first feature in the vector space with a probability (1−P). How it works.

5. In claim 3, the second step includes a step of decreasing each radius of the hypersphere of the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1), or increasing each radius of the hypersphere of the fourth feature and the first feature in the vector space with a probability (1-P). How it works.

6. In claim 1, the first characteristic is one of shape and color, and the second characteristic is the other of shape and color; How it works.

7. a storage unit that stores a model for generating a third object image having the first and second features by a generating unit using a first feature extracted from a first object image by an extracting unit and a second feature extracted from a second object image by another extracting unit; a control unit that causes the extraction unit to perform machine learning so as to classify the third feature extracted by the extraction unit from the third object image and the first feature into the same group; An information processing device having the above.

8. In claim 7, The control unit further causes the extraction unit to perform machine learning to classify a fourth feature extracted from the second object image by the extraction unit and the first feature into different groups. Information processing device.

9. In claim 7, the control unit classifies the first and third features into the same group by reducing a distance in the vector space between the first and third features, and classifies the second and fourth features into different groups by increasing a distance in the vector space between the second feature and a fourth feature extracted from the second object image by the extraction unit. Information processing device.

10. In claim 8, the control unit reduces the distance between the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1), and increases the distance between the fourth feature and the first feature in the vector space with a probability (1−P); Information processing device.

11. In claim 9, the control unit decreases each radius of the hypersphere of the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1), or increases each radius of the hypersphere of the fourth feature and the first feature in the vector space with a probability (1−P). Information processing device.

12. In claim 7, the first characteristic is one of shape and color, and the second characteristic is the other of shape and color; Information processing device.

13. In the information processing device, a first step of generating, by a generating unit, a third object image having the first and second features, using a first feature extracted from a first object image by an extracting unit and a second feature extracted from a second object image by another extracting unit; a second step of causing the extraction unit to perform machine learning so as to classify the third feature extracted by the extraction unit from the third object image and the first feature into the same group; A program that executes the following.

14. In claim 13, the second step further includes causing the extraction unit to perform machine learning to classify a fourth feature extracted by the extraction unit from the second object image and the first feature into different groups. Program.

15. In claim 13, the second step includes a step of classifying the first and third features into the same group by reducing a distance between the first and third features in a vector space, and classifying the second and fourth features into different groups by increasing a distance between the second feature and a fourth feature extracted from the second object image by the extraction unit in the vector space. program.

16. In claim 14, the second step includes a step of decreasing the distance between the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1) and increasing the distance between the fourth feature and the first feature in the vector space with a probability (1−P). program.

17. In claim 15, the second step includes a step of decreasing each radius of the hypersphere of the third feature and the first feature in the vector space with a predetermined probability P (where 0<P<1), or increasing each radius of the hypersphere of the fourth feature and the first feature in the vector space with a probability (1-P). program.

18. In claim 13, the first characteristic is one of shape and color, and the second characteristic is the other of shape and color; program.

Citation Information

Patent Citations

  • Systems and methods for providing personalized product recommendations using deep learning

    JP2023081971A

  • Information processing device, information processing method, and program

    WO2022176279A1