Clothing try-on simulation system and program

The system addresses the issue of incongruity in clothing try-on simulations by aligning facial features and line of sight between user and model images, resulting in a realistic and comfortable try-on experience.

JP7713733B2Active Publication Date: 2025-07-28PRISMATEC CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023063778
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-07-28
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

Existing clothing try-on simulation systems fail to adequately eliminate the sense of incongruity when synthesizing user and model images, particularly due to discrepancies in face orientation, distance, and line of sight.

Method used

A system and program that utilize feature point extraction, orientation detection, and image synthesis to align and correct face orientations and distances between user and model images, adjusting transparency and line of sight to minimize discrepancies.

Benefits of technology

Generates a composite image that effectively reduces the sense of incongruity by aligning facial features and line of sight, providing a realistic clothing try-on simulation without discomfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713733000001
    Figure 0007713733000001
  • Figure 0007713733000002
    Figure 0007713733000002
  • Figure 0007713733000003
    Figure 0007713733000003
Patent Text Reader

Abstract

To provide a system that enables trial fitting simulation without any discomfort on screen.SOLUTION: A program 2 configuring a system 1 is installed on a user terminal 3 of a user U, connected to an SNS 4 or the like via a network, and accessible to images uploaded on a model terminal 7 of a model M. When the user U wishes to try on clothes, the user U selects a user image 8 and searches for images of the clothes that the user U wants to try on. When a model image 10 is selected in the system 1, display control means 12 extracts facial feature points of the user image 8 and the model image 10, obtains roll, pitch, and yaw values from a reference posture of the face, corrects the facial orientation of the user image 8, performs correction for adjusting the distance between feature points 8b and 10b, also corrects the line of sight, and synthesizes the image of a facial area of the user image 8 onto the model image 10. This correction enables generating a synthetic image without any discomfort, and enables highly accurate trial fitting simulation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system and a program that can easily perform a try-on simulation of clothes desired by a user on so-called SNS (Social Networking Service), flea market sites, auction sites, or EC sites on the Internet.

Background Art

[0002] In recent years, uploading an image of wearing clothes purchased on SNS has been carried out. In addition, sites such as so-called flea market sites and auction sites where sellers offer clothes or accessories (such as clothes) that they no longer need, or clothes that the sellers want to sell, are widely used. Also, in electronic commerce such as EC sites, clothes and the like are circulated as major trading items.

[0003] When a user wants to obtain clothes or desires to view clothes like window shopping, the user often hopes to try them on. Patent Document 1 proposes a system in which a user can obtain their own try-on image using a smartphone or a personal computer, and the system provides a try-on state according to the user's body shape.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] The system described in Patent Document 1 acquires an image of a user and parameters indicating the user's body shape, acquires identification information of clothing that the user wishes to try on, acquires clothing body shape parameters corresponding to the size of the clothing from the clothing identification information, acquires a model image showing the state of a model of a body shape that matches the clothing body shape parameters wearing the clothing, displays on the user terminal a model image worn by a model having body shape parameters close to the user's body shape parameters, and displays a composite image obtained by synthesizing the model image and the user image.

[0006] In the system in this Patent Document 1, a user image obtained by photographing the user and a model image worn by a model having parameters close to the user's body shape are displayed on the user terminal, and by synthesizing the user image and the model image, it is possible to view an image as if the user had tried on the clothing.

[0007] On the other hand, in the system described in Patent Document 1, when synthesizing the user image and the model image, a depth map using distance information is acquired for the user image, the parameters of the user's body shape are estimated, and a model image having body shape parameters close to the estimated body shape parameters is selected for synthesis. Thus, in Patent Document 1, a method is adopted to reduce the sense of incongruity when creating a composite image by selecting images with close body shape parameters for the user image and the model image.

[0008] However, regarding the synthesis of the user image and the model image, the face is only used for detecting the body directions of the user and the model, and no means for eliminating the sense of incongruity when the face and the torso are synthesized is disclosed.

[0009] Patent Document 2 discloses a simulation device for performing on-screen clothing fitting for clothing such as wedding dresses in a hotel or a ceremony hall. In the device in Patent Document 2, only the face image is cut out from the user image along the contour of the face, moved to the position of the face of the model in the model image to create a composite image, and adjustments such as the position, size, hue, brightness, and shading of the face image are performed.

[0010] Patent Document 3 discloses a simulation method for synthesizing an image of a mannequin dressed in clothes and a user image. The method includes photographing a mannequin with feature points representing body shape displayed thereon, photographing the mannequin dressed in clothes, photographing the user's image, and posing the user's image in the same pose as the mannequin.

[0011] In this Patent Document 3, in the embodiment, only the face part is synthesized with the mannequin image. However, there is also a description to the effect that for the parts where the skin is exposed in a human image such as the neck, hands, and feet, they may be cut out from the human image and synthesized with the mannequin image and smoothed.

[0012] Thus, in the conventional system, with regard to the simulation of clothes, various contrivances are made in synthesizing the user image and the model image. However, when actually generating a synthesized image, it sometimes becomes an image with a sense of incongruity.

[0013] An object of the present invention is to provide a fitting simulation system and a program capable of generating a synthesized image that does not cause a sense of incongruity when generating a synthesized image by synthesizing a user image and a model image.

Means for Solving the Problems

[0014] To achieve the above object, the fitting simulation system of the present invention includes a user image input means for inputting a user image obtained by photographing a user, a model image input means for inputting a model image in which a model is wearing clothes, and a display control means for displaying the user image and the model image on the screen of a display device. The display control means includes a feature point extraction unit for extracting feature points in the face regions of the user image and the model image, an orientation detection unit for detecting the orientations of the faces of the user and the model from the feature points of the user image and the model image, a correction unit for correcting the difference in the distance between the feature points and the difference in the face orientation in the user image and the model image within a predetermined range, and an image synthesis unit for synthesizing the face portion of the user image corrected by the correction unit and the model image and displaying a synthesized image on the display device.

[0015] When the inventors of the present application examined the cause of the discomfort in the conventional system, it was found that it occurred due to a slight difference in the orientation of the user's face and the orientation of the model's face when the images in the face regions of the user image and the model image were replaced. Therefore, in the fitting simulation system of the present invention, by extracting the feature points in the face regions of the user image and the model image and making the difference in the orientation of the face of the user image and the orientation of the face of the model image within a predetermined range, it is possible to generate a synthesized image that does not cause discomfort when synthesizing the user image and the model image to generate a synthesized image.

[0016] In the fitting simulation system of the present invention, the orientation detection unit may detect the difference in the orientation of the faces of the user and the model by obtaining the respective values of roll, pitch, and yaw from the reference posture of the face. In this way, by detecting the orientations of the faces of the user image and the model image by obtaining the respective values of roll, pitch, and yaw from the reference posture of the face, it is possible to generate a synthesized image without discomfort.

[0017] Also, in the fitting simulation system of the present invention, the correction unit may be configured such that the sum of the squared differences in the distances between a plurality of the feature points in the user image and the model image is minimized, so that the difference in the distances between the feature points in the user image and the model image is within a predetermined range. With this configuration, the accuracy of eliminating the sense of incongruity in the composite image can be improved.

[0018] Also, in the fitting simulation system of the present invention, the image composite unit may extract the hair portion image of the user image, detect the transparency of the hair portion image, and composite the transparency of the hair portion of the user in the composite image as the transparency in the user image with the image of the background of the hair portion. According to this configuration, the sense of incongruity in the hair portion of the composite image can be reduced.

[0019] Also, in the fitting simulation system of the present invention, the image composite unit may detect the directions of the lines of sight in the user image and the model image, and when the lines of sight in both images are reversed left and right, perform an inversion process of reversing the line of sight of the user image. The sense of incongruity that occurs in the composite image may be caused by the difference in the lines of sight between the user image and the model image, but by performing the inversion process on the line of sight of the user image, the sense of incongruity due to the difference in the lines of sight can be eliminated.

[0020] Also, the fitting simulation program of the present invention is a program for operating a computer as each of the fitting simulation systems. This program may be installed and used in a device incorporating a computer such as a smartphone, or may be installed in a computer such as a server and used with a smartphone or the like.

Advantages of the Invention

[0021] According to the present invention, with the above-described respective configurations, when generating a composite image by compositing a user image and a model image, it is possible to generate a composite image that does not cause a sense of incongruity.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0023] Next, a fitting simulation system and a fitting simulation program, which are examples of embodiments of the present invention, will be described with reference to FIGS. 1 to 6. The fitting simulation system 1 of this embodiment (hereinafter simply referred to as "system 1") is realized by software constituting the program 2 and hardware that installs and executes the program, as shown in FIG. 1.

[0024] The program 2 is installed in the user terminal 3 owned by the user U or the server S, and can be respectively connected via a network constructed by the Internet, a telephone line, etc. to other sites such as the SNS 4, the flea market site 5, the EC site 6, or the auction. Also, from the model terminal 7 owned by the model M, an image of the model M can be uploaded to the SNS 4 or the like.

[0025] The server S is composed of hardware (computer) including a computing device such as a CPU, a RAM, a ROM, various storage devices such as a hard disk or an SSD, various interfaces, etc., and software installed and executed on the hardware. The server S includes not only the case where it is constructed in one place but also the case where it is constructed dispersedly like a so-called cloud computing system.

[0026] As the functional configuration of Program 2, it includes a user image input means 9 for inputting a user image 8 of the user U, a model image input means 11 for inputting a model image 10 of a person wearing clothes, a display control means 12 for displaying the user image 8 and the model image 10 on the screen of a display device such as the user terminal 3, and a communication means 13 for communicating with an external terminal, database, etc.

[0027] The system 1 of the present embodiment is executed by a program called an app that is executed on a user terminal 3 such as a smartphone owned by the user U. The user terminal 3 includes personal computers such as tablet terminals and notebook PCs in addition to smartphones. Further, the program of the present embodiment may be stored in the server S, and the program may be executed using a program such as a web browser on the user terminal 3.

[0028] The user image input means 9 is a functional unit that displays the user image 8 stored in the user terminal 3 or the user image 8 stored in various storage services (online image storage services) on the user terminal 3 as a display device. The user U can select an image of himself / herself for which he / she wants to perform a try-on simulation by this user image input means 9.

[0029] The model image input means 11 is a functional unit that causes the user terminal 3 to display the model image 10 uploaded by the model M on SNS, blogs, etc. The user U can search for an image of a model wearing their favorite clothes from a large number of images uploaded on SNS, etc., and select a model image 10 that they want to try on and display it on the user terminal 3. In addition, the model image input means 11 analyzes the tendency of the clothes selected by the user U by the program 2, and can also search for recommended images for the user U from information such as the user U's profile and display them on the user terminal 3.

[0030] The display control means 12 is a functional unit that realizes a simulation of trying on the clothes worn by the model M by synthesizing the image of the face region (for example, the region combining the face, hair, and neck) of the user image 8 with the model image 10. As shown in FIG. 2, the display control means 12 includes a feature point extraction unit 14 that extracts the feature points 8b, 10b in the face region 8a of the user image 8 and the face region 10a of the model image 10, and an orientation detection unit 15 that detects the orientation of each face from the feature points 8b, 10b in the face regions 8a, 10a of the user U and the model M.

[0031] In addition, the display control means 12 calculates the distance between the feature points 8b in the user image 8, calculates the distance between the feature points 10b in the model image 10, and includes a correction unit 16 that corrects the user image 8 so that the difference in the distance between the corresponding feature points in the user image 8 and the model image 10 is within a predetermined range.

[0032] Furthermore, the display control means 12 includes an image synthesis unit 17 that performs image synthesis by replacing the image of the face region 8a of the user image 8 corrected by the correction unit 16 with the image of the face region 10a of the model image 10.

[0033] Next, the operation of the system 1 in this embodiment will be described with reference to FIGS. 1 to 6. As shown in FIG. 1, the user U uses the user terminal 3 to start the application provided by the system 1 in this embodiment. When the application is started and login processing such as personal authentication of the user U is performed, it is possible to select a user image 8 for performing a try-on simulation.

[0034] For the selection of the user image 8, it is possible to select an image stored in the user terminal 3 or an image uploaded by the user himself / herself to SNS or the like. FIG. 3(A) is a diagram showing a state in which a plurality of images are displayed on the user terminal 3. When one of the images is selected, the selected user image 8 is displayed as shown in FIG. 3(B).

[0035] The user U taps (or clicks) the registration button 18 displayed on the screen of the user terminal 3 to register the selected user image 8 in the system 1 as the user image 8 to be used in the simulation.

[0036] Next, the user U searches for clothes that he / she wants to try on. The search can be performed on the SNS 4, the flea market site 5, or the EC site 6 as shown in FIG. 1. Specifically, as shown in FIG. 3(A), a plurality of images uploaded to the SNS 4 are tapped, and one favorite model image 10 is selected. Note that the SNS 4, the flea market site 5, or the EC site 6 may provide a service or site that collects only the images of the users of the system 1 in this embodiment.

[0037] When the model image 10 is selected, the model image 10 is displayed on the screen of the user terminal 3, and a simulation button 19 is displayed on a part of the model image 10. Here, when the user U taps the simulation button 19, the processing by the display control means 12 is started.

[0038] In the display control means 12, first, the feature point extraction unit 14 extracts feature points in the face regions 10a of the user image 8 and the model image 10. Examples of the feature points include the centers of the pupils of both eyes, the midpoint between both eyes, the apex of the nose, the horizontal contours of both eyes, the tip of the chin, both ends of the mouth, the boundary between the face contour and the neck, the apex of the head, and a part of the hair contour. These feature points are selected as the points necessary for detecting the face orientation. FIGS. 5(A) and (B) show the state in which the feature points 8b and 10b are extracted in the face region 8a of the user image 8 and the face region 10a of the model image 10 by the feature point extraction unit 14.

[0039] Next, the orientation detection unit 15 detects the orientation of the face of the user U in the user image 8 and the orientation of the face of the model M in the model image 10. In the orientation detection unit 15, not only the inclination detected from each feature point but also each value of roll, pitch, and yaw from the reference pose of the face is obtained for detection.

[0040] The reference pose of the face can be, for example, a state in which the line connecting the center of the forehead, the apex of the nose, and the tip of the chin of the face is vertical, the line connecting both eyes is horizontal, and the line of sight is straight ahead.

[0041] In the orientation detection unit 15, for this reference pose, the respective values of how much the line connecting both eyes is inclined from the horizontal line (roll), how much the line of sight is inclined vertically (pitch), and how much the line of sight is inclined horizontally (yaw) are detected to detect the face orientation.

[0042] The detection of the face orientation in this orientation detection unit 15 can be performed by so-called AI processing using a learned model trained with images of various face orientations of the user U as teacher data for the user image 8. Alternatively, in addition to methods such as digitizing the body shape of the user U with a 3D scanner or the like for detection, a gyro sensor or the like may be attached to the user's head in advance, and various orientation photos and the detection values of the gyro sensor may be digitized for detection.

[0043] Also, the orientation detection unit 15 also obtains the values of roll, pitch, and yaw from the reference pose of the face of the model M in the model image 10, and detects the orientation of the face region 10a in the model image 10.

[0044] Next, the correction unit 16 corrects the model image 10 so that the difference in the distances between the respective feature points in the user image 8 and the model image 10 is within a predetermined range. Specifically, regarding the difference in the distances between the respective feature points in the user image 8 and the model image 10, the sum of the squared values of the differences in the distances between the plurality of said feature points in the user image 8 and the model image 10 is minimized, thereby making the difference in the distances between the feature points in the user image 8 and the model image 10 within a predetermined range.

[0045] For example, regarding the distance between the tip of the nose and the tip of the chin, the distance between the midpoint of both eyes and the tip of the chin, the distance between the horizontal contours of both eyes, etc., subtract the distance in the model image 10 from the distance in the user image 8, square the result, and add each value. Adjust the scale of the user image 8 or the model image 10 so that the added value is minimized.

[0046] Next, the image synthesis unit 17 synthesizes the face portion of the user image 8 corrected by the correction unit 16 and the model image 10. At this time, the image synthesis unit 17 extracts the hair portion image of the user image 8 and detects the transparency of the hair portion image. For example, the transparency is low in the part where the hair is dense, and the transparency is high in the part where the hair is not dense.

[0047] When the image synthesis unit 17 synthesizes the face portion of the user image 8 with the model image 10, it adjusts the transparency at the time of image pasting according to the transparency of the hair portion. As a result, where the transparency of the hair portion image is high, the background image can be seen according to the transparency of the hair portion, so the sense of incongruity of the synthesized image is eliminated.

[0048] In addition, the image synthesis unit 17 detects the direction of the user U's line of sight in the user image 8 and the direction of the model M's line of sight in the model image 10. When the lines of sight in both images are reversed left and right, an inversion process may be performed to invert the line of sight of the user image 8 left and right.

[0049] FIG. 5(A) shows the user image 8, but the line of sight of the user U is slightly directed to the left. On the other hand, looking at the model image 10 in FIG. 5(B), the line of sight of the model M is slightly directed to the right. In such a case, the image synthesis unit 17 performs an inversion process of inverting the position of the black pupil in the white eye in the eye portion of the user image 8. By this inversion process, the line of sight of the user image 8 when synthesized into the model image 10 becomes the same line of sight as the model image 10, and the discomfort of the synthesized image is eliminated.

[0050] Also, at this time, the face and hair portions of the model M in the model image 10 can be cut out and synthesized with the face and hair portions of the user U in the user image 8. At this time, when the cut-out portion is large, a redrawing process may be performed by matching the background of the blank portion to the surroundings.

[0051] In addition, when synthesizing the face region of the user image 8, a gap may occur between the face and the neck. In this case, a so-called filling process of extending the neck region to the boundary portion of the face region may be performed. Also, when the hue and saturation of the surroundings of the face region of the user image 8 and the model image 10 are different, the hue and saturation of the face region of the user image 8 may be adjusted to match the surroundings.

[0052] By performing the above-described processes, even when the image of the face region of the user image 8 is synthesized into the model image 10, a synthesized image without discomfort can be generated. The synthesized image generated by the system of this embodiment becomes, for example, the image shown in FIG. 6.

[0053] The composite image 20 shown in FIG. 6 is such that the image of the face region of the user image 8 is adjusted to the face orientation (roll, pitch, yaw) in the model image 10, the distance between the feature points of the user U is adjusted to be close to the distance between the feature points of the model M, and the line of sight is also adjusted. Therefore, there is no sense of discomfort, and it can be made into an image as if the user U is wearing the clothes in the model image 10.

[0054] As a result, even when the user U cannot actually try on clothes, the user U can use their own smartphone as the user terminal 3 to perform a try-on simulation.

[0055] Here, when the user U taps the simulation button 19 shown in FIG. 6, information regarding the clothes and accessories in the model image 10 can be obtained. If the user likes them, they can connect to the SNS 4, flea market site 5, or EC site 6 to purchase these clothes and accessories.

[0056] In the above embodiment, an example of using a smartphone as the user terminal 3 has been described. However, as in Patent Document 1, an apparatus that photographs the user U life-size and acquires numerical values such as an image, height, and weight may be used, or equipment such as a 3D scanner may be used to acquire an image and dimensional data of the user U.

[0057] Also, in the above embodiment, the face region is defined as the region including the face, hair, and neck. However, only the face and hair may be defined as the face region.

Explanation of Reference Numerals

[0058] 1... Try-on simulation system 2... Program (app) 3... User terminal 4... SNS 5... Flea market site 6... EC site 7... Model terminal 8... User image 8a... Face region 8b... Feature points 9… User image input means 10… Model image 10a… Face area 10b… Feature points 11… Model image input means 12… Display control means 13… Communication means 14… Feature point extraction unit 15… Orientation detection unit 16… Correction unit 17… Image synthesis unit 18… Registration button 19… Simulation button 20… Composite image U… User M… Model S… Server

Claims

1. User image input means for inputting a user image obtained by photographing a user; Model image input means for inputting a model image in which a model is wearing clothes; Display control means for displaying the user image and the model image on the screen of a display device, The display control means includes: A feature point extraction unit that extracts feature points in the face regions of the user image and the model image; An orientation detection unit that detects the orientations of the faces of the user and the model from the feature points of the user image and the model image; A correction unit that corrects the difference in the distances between the feature points in the user image and the model image and the difference in the face orientations within a predetermined range; An image synthesis unit that synthesizes the face portion of the user image corrected by the correction unit and the model image and causes the display device to display a synthesized image, The image synthesis unit extracts a hair portion image of the user image, detects the transparency of the hair portion image, and synthesizes the transparency of the user's hair portion in the synthesized image as the transparency in the user image with the image of the background of the hair portion, characterized by a try-on simulation system.

2. User image input means for inputting a user image obtained by photographing a user; Model image input means for inputting a model image in which a model is wearing clothes; Display control means for displaying the user image and the model image on the screen of a display device, The display control means includes: A feature point extraction unit that extracts feature points in the face regions of the user image and the model image; An orientation detection unit that detects the orientations of the faces of the user and the model from the feature points of the user image and the model image; A correction unit that corrects the difference in the distances between the feature points in the user image and the model image and the difference in the face orientations within a predetermined range; An image synthesis unit that synthesizes the face portion of the user image corrected by the correction unit and the model image and causes the display device to display a synthesized image, The image synthesis unit detects the directions of the lines of sight in the user image and the model image, and when the lines of sight in both images are reversed left and right, performs an inversion process of inverting the line of sight of the user image, characterized by a try-on simulation system.

3. The try-on simulation system according to claim 1 or 2, The orientation detection unit detects the difference in the face orientations of the user and the model by obtaining the roll, pitch, and yaw values from the reference posture of the face. A fitting simulation system characterized by this.

4. The fitting simulation system according to claim 1 or 2, The correction unit makes the sum of the squared differences in the distances between the plurality of feature points in the user image and the model image minimum, thereby setting the difference in the distances between the feature points in the user image and the model image within a predetermined range. A fitting simulation system characterized by this.

5. A fitting simulation program for operating a computer as the fitting simulation system according to claim 1 or 2.

Citation Information

Patent Citations

  • Clothes fitting simulation method

    JP1997106419A

  • Try-on simulation system

    JP2001134745A

  • Image-processing program, computer-readable recording medium recording the program, image processor and image processing method

    JP2009064086A

  • Image display method, and image display device

    JP2013168969A

  • Sight line conversion device, sight line conversion method, and program

    JP2015149016A