Image generation program, image generation system, and image generation method

The image generation system addresses the limitation of requiring a predetermined dressing item by using 3D models and learning models to generate accurate virtual try-on images for subjects not wearing dedicated items, enabling versatile and real-time virtual try-on.

JP2025105087APending Publication Date: 2025-07-10THE UNIV OF TOKYO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023223385
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing image processing systems require the subject to wear a predetermined dressing item for virtual try-on, limiting their applicability when the subject is not wearing such an item.

Method used

An image generation system that includes an input image acquisition unit, a 3D model generation unit, a reference image generation unit, an image generation model, and an output image generation unit to create virtual wearable item images based on a target person's pose, using learning models and 3D models to generate accurate virtual try-on images even when the subject is not wearing a dedicated item.

Benefits of technology

Enables virtual try-on of wearable items in various images, including those where the subject is not wearing a dedicated item, with high accuracy and real-time capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105087000001_ABST
    Figure 2025105087000001_ABST
Patent Text Reader

Abstract

To realize virtual fitting of a wearing article in various images.SOLUTION: An image generation program allows a computer to realize: an input image acquisition section for acquiring an input image of an object person; an object person model generation section for generating an object person model by inputting an input image to a three-dimensional model; a reference image generation section for generating a reference image of a virtual reference ornament in a three-dimensional shape corresponding to a pose of an object person in an input image on the basis of an object person model; an object image generation section for inputting a reference image to an image generation model that is allowed to learn multiple learning reference images in a shape corresponding to each of multiple poses and multiple learning object images in a shape corresponding to each of multiple poses in association with each other, and generating an object image of an object ornament in a three-dimensional shape corresponding to a pose shown by the reference image; and an output image generation section for generating an output image of an object person who attaches an object ornament on the basis of an input image and an object image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image generation program, an image generation system, and an image generation method.

Background Art

[0002] Conventionally, a technique for virtually dressing a subject's image has been known.

[0003] For example, in the image processing apparatus described in Patent Document 1, an image of a reference dressing item worn by a wearing subject in a predetermined posture is used as input data, and an image of a trial dressing item worn by the wearing subject in a posture common to the predetermined posture is used as output data. A trained model trained by machine learning using the associated learning data is provided with an image generation unit that inputs an image of a user wearing a reference dressing item and generates an image of a user trying on a trial dressing item.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the image processing apparatus described in Patent Document 1, it is premised that the wearing subject (the subject) is wearing a predetermined dressing item, that is, a dedicated dressing item for virtual try-on. Therefore, when the subject is not wearing a predetermined dressing item, it is impossible to virtually dress the subject.

[0006] Therefore, an object of the present invention is to realize virtual dressing in various images.

Means for Solving the Problems

[0007] An image generation program according to an aspect of the present invention causes a computer to include: an input image acquisition unit that acquires an input image of a target person; a 3D model generation unit that generates a 3D model of a predetermined object from an image of the predetermined object, and inputs the input image to generate a target person model that is a 3D model of the target person; a reference image generation unit that generates a reference image of a virtual reference wearable item having a 3D shape corresponding to the pose of the target person in the input image based on the target person model; an image generation model that is learned by associating a plurality of learning reference images of reference wearable items having shapes corresponding to each of a plurality of poses, and a plurality of learning target images of target wearable items having shapes corresponding to each of the plurality of poses, and inputs the reference image to generate a target image of the target wearable item having a 3D shape corresponding to the pose indicated by the reference image; and an output image generation unit that generates an output image of the target person wearing the target wearable item based on the input image and the target image.

Advantages of the Invention

[0008] According to the present invention, virtual wearable item wearing can be realized in various images.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0010] With reference to the accompanying drawings, preferred embodiments of the present invention will be described. FIG. 1 is a diagram showing an overview of processing in an image generation system 100 which is an embodiment of the present invention.

[0011] The image generation system 100 is an information processing system realized by an image generation program. The image generation system 100 is an information processing system that inputs an image of a subject to a learning model and generates an image of the subject virtually wearing a wearable (so-called trying on).

[0012] Here, the wearable includes, for example, items worn on the body of the subject such as clothes, hats, glasses, shoes, accessories, etc. In this embodiment, the case where the wearable is clothes will be described.

[0013] First, the image generation system 100 inputs the acquired input image to a 3D model (S101) and generates a subject model (S102). The image generation system 100 generates a reference image based on the subject model and the reference wearable model (S103).

[0014] Subsequently, the image generation system 100 inputs a reference image into an image generation model (S104) and generates a target image (S105).

[0015] Then, the image generation system 100 generates a mask image based on the input image (S106), inputs the target image and the mask image into a gap filling model (S107), and generates an output image (S108). The image generation system 100 outputs the output image.

[0016] Note that in this embodiment, unless otherwise specified, the "image" includes still images and moving images.

[0017] FIG. 2 is a diagram showing the configuration of an image generation system 100 according to an embodiment of the present invention. The image generation system 100 is communicably connected to a user device 200 and a photographing device 300 via a network such as the Internet. Details of the image generation system 100 will be described later.

[0018] The user device 200 is an information processing device used by a user, such as a computer, a smartphone, a tablet terminal, or a personal computer.

[0019] The user accesses the image generation system 100 using the user device 200 and provides, for example, an input image or an input video to the image generation system 100. Also, the user accesses the image generation system 100 using the user device 200 and acquires, for example, an output image or an output video.

[0020] The user device 200 may acquire an input image photographed by the photographing device 300 described later and output the input image to the image generation system 100.

[0021] The photographing device 300 is a device that photographs a subject and outputs the photographed image to an external information processing system. The photographing device 300 can output, for example, the photographed image as an input image or an input video to the image generation system 100.

[0022] The image generation system 100 may acquire an input image or an input video from the user device 200, or may acquire an input image or an input video from the imaging device 300.

[0023] In addition, in FIG. 2, one user device 200 and one imaging device 300 are shown respectively, but there may be a plurality of user devices 200 and imaging devices 300 respectively.

[0024] Subsequently, the details of the image generation system 100 will be described. The image generation system 100 includes a storage unit 110, a learning unit 120, an input image acquisition unit 130, a subject model generation unit 140, a reference image generation unit 150, a target image generation unit 160, a mask generation unit 170, an output image generation unit 180, and an output unit 190. Each unit shown in FIG. 2 can be realized, for example, by using a storage area or by the processor executing a program stored in the storage area.

[0025] The storage unit 110 stores information processed in the image generation system 100. The storage unit 110 can store, for example, a learning reference image, an image generation model, an input image, a subject model, a reference image, a target image, a mask image, an output image, and a gap completion model, which will be described later.

[0026] The learning unit 120 acquires a learning reference image and a learning target image used for learning the image generation model, and stores them in the storage unit 110. In addition, the learning unit 120 generates an image generation model based on the learning reference image and the learning target image, and stores the generated image generation model in the storage unit 110. Note that at least one of the learning reference image, the learning target image, and the image generation model may be stored in an external information processing system.

[0027] The learning unit 120 acquires, for example, a learning reference image and a learning target image from an external information processing system. At this time, the learning unit 120 can acquire the learning reference image and the learning target image, for example, based on the operation of the administrator of the image generation system 100.

[0028] The learning target images include, for example, a plurality of images obtained by photographing a plurality of poses of a predetermined mannequin wearing the target wearable from a plurality of angles.

[0029] The target wearable is, for example, a wearable to be virtually worn by a target person. The target wearable may be, for example, a wearable that is on sale or scheduled to be sold. That is, the target person can virtually wear (so-called try on) the target wearable through the image generation system 100 and consider purchasing the target wearable.

[0030] The mannequin is, for example, a humanoid mannequin. The mannequin may be, for example, a mannequin of a part where the wearable is worn. That is, for example, when the wearable is a shirt, the mannequin may be, for example, a mannequin of a human chest. The mannequin is configured such that a predetermined part (for example, a joint) can be driven with a plurality of degrees of freedom and can take various poses.

[0031] The learning target images are images obtained by photographing a plurality of poses of a predetermined mannequin wearing the target wearable from a plurality of angles before generating the image generation model. That is, the learning target images are images that have photographed the shape and appearance of the target wearable in each of the plurality of poses of the mannequin.

[0032] The learning reference images include, for example, images of virtual reference wearables corresponding to a plurality of poses of the mannequin.

[0033] The images of the reference wearables are, for example, images of wearables (reference wearables) that serve as a reference when the target image generation unit 160 described later generates target images. The reference wearable is, for example, a virtual wearable having the same or similar shape as the target wearable. Here, the shape may be an external shape. That is, when the target wearable is a short-sleeved shirt, the reference wearable has the shape of a short-sleeved shirt.

[0034] Each of the learning reference images is generated, for example, based on each of the learning target images. That is, each of the learning reference images is an image of a reference wearing item having a shape corresponding to the shape of the target wearing item indicated by each of the learning target images. Thereby, each of the learning reference images can be associated with each of the learning target images.

[0035] The reference wearing item and the target wearing item may correspond one-to-one. Thereby, the target image generation unit 160 described later can generate a target image with high accuracy.

[0036] Further, the reference wearing item may correspond to a plurality of target wearing items. That is, when the reference wearing item is a short-sleeved shirt, the reference wearing item may correspond to, for example, target wearing items of a short-sleeved shirt without a collar and a short-sleeved shirt with a collar. Thereby, the image generation system 100 can realize virtual wearing of a plurality of target wearing items based on an image of one reference wearing item, and the versatility of the image generation system 100 is improved.

[0037] Further, the learning reference image may be, for example, an image in which a predetermined color pattern is mapped for each predetermined region of the reference wearing item. Here, the mapping may be, for example, a process of dividing the reference wearing item into a plurality of regions in a mesh shape and associating the plurality of regions with a predetermined color pattern. Thereby, compared with the case of generating a target image using a learning reference image that does not match the color pattern or generating a target image using UV coordinates that are texture coordinates, the shape of the wearing item can be expressed more accurately. More specifically, for example, the shape of the wearing item can be accurately expressed due to the distortion of the shape for each predetermined region, which serves as a guide when generating the target image. Therefore, a more accurate virtual wearing can be realized.

[0038] The image generation model is a learning model that inputs a reference image, which will be described later, and generates a target image, which will also be described later. That is, the image generation model is a learning model trained by associating a plurality of learning reference images of reference wearing items in shapes corresponding to each of a plurality of poses with a plurality of learning target images of target wearing items in shapes corresponding to each of the plurality of poses.

[0039] More specifically, as will be described later, the reference image is an image of the reference wearing item when the subject takes a predetermined pose. And the image generation model is a learning model that takes the reference image as an input and outputs an image of the target wearing item when the subject takes the predetermined pose as the target image.

[0040] The image generation model is, for example, a learning model trained using machine learning techniques. The algorithm used for training the image generation model is not limited to this.

[0041] Note that the image generation model may be a model generated for each target wearing item. Thereby, the image generation model can generate a target image with high accuracy.

[0042] Note that the image generation system 100 may use, for example, an existing image generation model. Here, the existing image generation model may be a learning model stored in an external information processing system or a learning model stored in the image generation system 100.

[0043] Also, the image generation model may be an image generation model trained by further associating a predetermined color pattern mapped for each predetermined region of the reference wearing item. Thereby, since the shape of the wearing item can be expressed with higher accuracy, a more accurate virtual wearing can be realized.

[0044] FIG. 3A is a diagram showing an example of a learning reference image. As shown in FIG. 3A, the learning reference image is, for example, an image of a virtual reference wearable item, and is a plurality of images corresponding to a plurality of poses in which a predetermined color pattern is mapped for each predetermined region.

[0045] FIG. 3B is a diagram showing an example of a learning target image. As shown in FIG. 3B, the learning target image is, for example, a plurality of images of a target wearable item corresponding to a plurality of poses.

[0046] The input image acquisition unit 130 acquires an input image of a target person and stores the acquired input image in the storage unit 110.

[0047] Here, the target person is the person photographed in the input image. In the input image, the target person may or may not be wearing a predetermined wearable item. The input image includes at least the region where the target person virtually wears the wearable item in the target image described later. For example, when the wearable item is a shirt, the region where the wearable item is virtually worn is, for example, the chest.

[0048] The target person model generation unit 140 inputs the input image into a 3D model generation unit that generates a 3D model of a predetermined object from an image of the predetermined object, generates a target person model that is a 3D model of the target person, and stores the generated target person model in the storage unit 110.

[0049] Here, the predetermined object may be the target person. That is, the 3D model generation unit is, for example, a learning model that can generate a 3D model of the target person from an image including the target person as a predetermined object.

[0050] The image generation system 100 can use, for example, an existing 3D model generation unit. Here, the existing 3D model generation unit may be a learning model stored in an external information processing system or a learning model stored in the image generation system 100.

[0051] The subject model is a three-dimensional model of the subject in the pose in the input image. That is, the subject model is, for example, a three-dimensional model corresponding to the pose of the subject in the input image, which is reproduced on a predetermined information processing device.

[0052] The reference image generation unit 150 generates a reference image of a virtual reference wearable item having a three-dimensional shape corresponding to the pose of the subject in the input image based on the subject model, and stores the generated reference image in the storage unit 110.

[0053] Here, the reference image is an image of the reference wearable item corresponding to a predetermined pose of the subject in the input image. That is, when the wearable item is a shirt and the subject is raising a hand in the input image, the reference image is an image of the reference wearable item corresponding to the pose of raising a hand.

[0054] At this time, the reference image generation unit 150 may, for example, deform a reference wearable item model, which is a three-dimensional model of the reference wearable item set in advance, so as to correspond to the pose indicated by the subject model, and generate a reference image based on the deformed reference wearable item model.

[0055] The process of deforming the reference clothing model to correspond to the pose indicated by the subject model is realized by existing technologies. The process may be, for example, a process of adjusting various parameters (such as the angle of the sleeve, etc.) included in the reference clothing model based on various parameters included in the subject model. Here, the deformation of the reference clothing model may be realized, for example, by changing the coordinates v of the vertices constituting the reference clothing model. In this case, each vertex is associated with the nearest body part (such as the torso, upper arm, lower arm, etc.). First, in the reference pose, calculate the relative coordinates (T0) of the vertex (v0) of the reference clothing model as seen from the nearest body part (b0) (that is, for example, v0 = T0(b0)). When the pose of the human body changes, based on the pose (b) of the nearest body part in the changed pose and the relative coordinates (T0) therefrom, calculate the absolute coordinates of the vertices of the reference clothing model (that is, for example, v = T0(b)).

[0056] The process of generating a reference image based on the deformed reference clothing model may be, for example, a process of projecting the deformed reference clothing model, which is a 3D model, onto a predetermined plane (such as the plane where the viewpoint is located) to generate a 2D reference image. The process may be, for example, a process of obtaining a 2D image by projecting each vertex v constituting the reference clothing model onto the 2D drawing screen. More specifically, the coordinates of the vertex v after projection are obtained, for example, as the intersection of the straight line ev and p, where the viewpoint is e and the drawing screen is p.

[0057] Also, the reference image generation unit 150 may generate a reference image with a color pattern mapped thereon. At this time, the reference clothing model may be a 3D model with a predetermined color pattern mapped for each predetermined region of the reference clothing. Thereby, since the shape of the clothing can be expressed more accurately, a more accurate virtual wearing can be realized.

[0058] The target image generation unit 160 inputs the reference image into the image generation model, generates a target image of the target wearing item in a three-dimensional shape corresponding to the pose shown in the reference image, and stores the generated target image in the storage unit 110.

[0059] The target image is an image of the target wearing item corresponding to a predetermined pose of the target person in the input image. That is, when the wearing item is a shirt and the target person is raising their hand in the input image, the target image is an image of the target wearing item corresponding to the pose of raising the hand.

[0060] Further, the target image generation unit 160 may input a reference image with a color pattern mapped thereto into the image generation model to generate a target image. Thereby, since the shape of the wearing item can be expressed more accurately, a more accurate virtual wearing can be realized.

[0061] FIG. 4A is a diagram showing an example of an input image. In the input image shown in FIG. 4A, the target person is facing forward.

[0062] FIG. 4B is a diagram showing an example of a target person model. The target person model shown in FIG. 4B is, for example, a three-dimensional model corresponding to the pose of the target person in the input image shown in FIG. 4A.

[0063] FIG. 4C is a diagram showing an example of a reference image. The reference image shown in FIG. 4C shows an image of a reference wearing item (in this case, a shirt) corresponding to the pose of the target person model shown in FIG. 4B.

[0064] FIG. 4D is a diagram showing an example of a target image. The target image shown in FIG. 4D is, for example, a target image generated based on the reference image shown in FIG. 4C, and shows an image of a target wearing item (in this case, a shirt) corresponding to the pose of the target person model shown in FIG. 4B.

[0065] The mask generation unit 170 generates a mask image obtained by masking the area of the wearing item worn by the target person based on the input image, and stores the generated mask image in the storage unit 110.

[0066] The generation process of the mask image by the mask generation unit 170 is realized by existing technologies. This process may be, for example, a process in which a learning model learned by machine learning recognizes the wearing item area in the input image and masks the recognized wearing item area.

[0067] The output image generation unit 180 generates an output image of the target person wearing the target wearing item based on the input image and the target image, and stores the generated output image in the storage unit 110.

[0068] Further, the output image generation unit 180 may generate an output image by synthesizing the target image in the masked area based on the mask image.

[0069] Further, the output image generation unit 180 inputs the output image generated by synthesis to a gap compensation model that compensates for the gap portion included in the input image and outputs a predetermined image, and generates an output image in which the gap between the target image and the mask image is compensated.

[0070] Here, the gap compensation model is, for example, a learning model learned based on an image including a gap portion, a mask image of the gap portion, and an image in which the gap portion is compensated, and is a learning model that inputs an image including a gap portion and a mask image of the gap portion and generates an image in which the gap portion is compensated.

[0071] The gap compensation model may be, for example, an existing learning model. The gap compensation model may be a learning model stored in an external information processing system, or may be a learning model stored in the image generation system 100.

[0072] FIG. 5A is a diagram showing an example of a mask image. As shown in FIG. 5A, the area of the wearing item (in this case, a shirt) in the target image shown in FIG. 4A is masked.

[0073] FIG. 5B is a diagram showing an example of an output image. The output image shown in FIG. 5B shows an example of an output image generated without using a gap completion model. As shown in FIG. 5B, the subject in the input image shown in FIG. 4A is virtually wearing the target clothing item in the pose shown in the input image shown in FIG. 4A.

[0074] FIG. 5C shows an example of a gap image between a target image and a mask image. The gap image is, for example, the difference between the target image and the mask image. As shown in FIG. 5C, gaps can be seen in the sleeves and hems of the shirt in the target image and the mask image.

[0075] FIG. 5D is a diagram showing an example of an output image. The output image shown in FIG. 5D shows an example of an output image generated using a gap completion model. As shown in FIG. 5D, the subject in the input image shown in FIG. 4A is virtually wearing the target clothing item in the pose shown in the input image shown in FIG. 4A. Also, as shown in FIG. 5D, the gap portion shown in FIG. 5C is complemented by the gap completion model.

[0076] In this way, even if the subject does not wear a dedicated clothing item for virtual try-on, the image generation system 100 can generate a target image of virtual try-on. That is, the image generation system 100 can realize virtual try-on in various images including an image in which the subject does not wear a dedicated clothing item.

[0077] The output unit 190 outputs the output image. The output unit 190 can output the output image to, for example, the user device 200.

[0078] The image generation system 100 can acquire an input image and generate and output an output image in real time.

[0079] Specifically, the input image acquisition unit 130 sequentially acquires input images, and the output image generation unit 180 sequentially generates output images corresponding to the input images sequentially acquired.

[0080] Accordingly, for example, in a state where the subject stands in front of the imaging device 300, an output image in which the target wearing item is virtually worn on the subject can be generated and output in real time.

[0081] In addition, the image generation system 100 can generate and output a video in which the target wearing item is virtually worn on the subject imaged in a video format.

[0082] Specifically, the image generation system 100 further includes a video acquisition unit and a video output unit. The video acquisition unit acquires an input video in which the subject is imaged. The input image acquisition unit 130 acquires each of a plurality of frames included in the input video as an input image. The output image generation unit 180 generates an output image corresponding to each of the plurality of acquired frames. Then, the video output unit outputs an output video based on each of the generated output images.

[0083] At this time, the video may be a video that has been finished being shot. Accordingly, the image generation system 100 can virtually change the wearing item of the subject in the video that has been finished being shot. Also, the video may be a video that is being continuously shot. Accordingly, for example, in a state where the subject stands in front of the imaging device 300, the image generation system 100 can generate and output an output image in which the target wearing item is virtually worn on the subject in real time.

[0084] FIG. 6 is a flowchart showing an example of processing in the image generation system 100. The flowchart shown in FIG. 6 shows an example of the generation process of the image generation model in the image generation system 100.

[0085] Specifically, the learning unit 120 acquires a learning reference image and a learning target image (S601), and learns based on the learning reference image and the learning target image to generate an image generation model (S602).

[0086] FIG. 7 is a flowchart showing an example of processing in the image generation system 100. The flowchart shown in FIG. 7 shows an example of the process of generating an output image in the image generation system 100.

[0087] First, the input image acquisition unit 130 acquires an input image (S701). The subject model generation unit 140 inputs the input image into a three-dimensional model to generate a subject model (S702). The reference image generation unit 150 generates a reference image based on the subject model (S703). The target image generation unit 160 generates a target image based on the reference image (S704).

[0088] The mask generation unit 170 generates a mask image based on the input image (S705). The output image generation unit 180 generates an output image based on the target image and the mask image, and the output unit 190 outputs the output image (S706).

[0089] Next, with reference to FIG. 8, an example of the hardware configuration when the image generation system 100 is realized by a computer 800 will be described. FIG. 8 is a diagram showing an example of the hardware configuration of the computer 800.

[0090] As shown in FIG. 8, the computer 800 includes, for example, a processor 801, a memory 802, a storage device 803, an input I / F unit 804, a data I / F unit 805, a communication I / F unit 806, and a display device 807.

[0091] The computer 800 may be, for example, a server computer, a personal computer (e.g., desktop, laptop, tablet, etc.), a media computer platform (e.g., cable, satellite set-top box, digital video recorder, etc.), a handheld computer device (e.g., PDA, email client, etc.), or other types of computers, or a communication platform.

[0092] The processor 801 is a control unit that controls various processes in the computer 800 by executing programs stored in the memory 802.

[0093] The memory 802 is a storage medium such as a RAM (Random Access Memory). The memory 802 temporarily stores program codes of programs executed by the processor 801 and data required during program execution.

[0094] The storage device 803 is a non-volatile storage medium such as a hard disk drive (HDD) or a flash memory. The storage device 803 stores an operating system and various programs for realizing the above-described components.

[0095] The input I / F unit 804 is a device for receiving inputs from a user. The input I / F unit 804 is, for example, a keyboard, a mouse, a touch panel, various sensors, a wearable device, etc. The input I / F unit 804 may be connected to the computer 800 via an interface such as a USB (Universal Serial Bus).

[0096] The data I / F unit 805 is a device for inputting data from outside the computer 800. The data I / F unit 805 is, for example, a drive device for reading data stored in various storage media. The data I / F unit 805 may be provided outside the computer 800. When the data I / F unit 805 is provided outside the computer 800, the data I / F unit 805 is connected to the computer 800 via an interface such as a USB.

[0097] The communication I / F unit 806 is a device for performing data communication with a device external to the computer 800 via a network such as the Internet, either wired or wirelessly. The communication I / F unit 806 may be provided outside the computer 800. When the communication I / F unit 806 is provided outside the computer 800, the communication I / F unit 806 is connected to the computer 800 via an interface such as USB, for example.

[0098] The display device 807 is a device for displaying various kinds of information. The display device 807 is, for example, a liquid crystal display, an organic EL (Electro-Luminescence) display, a display of a wearable device, or the like. The display device 807 may be provided outside the computer 800. When the display device 807 is provided outside the computer 800, the display device 807 is connected to the computer 800 via a display cable or the like, for example. Also, when a touch panel is adopted as the input I / F unit 804, the display device 807 may be configured integrally with the input I / F unit 804.

[0099] As described above, one embodiment of the present invention has been explained. The image generation system 100 can acquire an input image of a target person, input the input image into a 3D model to generate a target person model, generate a reference image based on the target person model, input the reference image into an image generation model to generate a target image, and generate an output image of the target person wearing the target wearing item based on the input image and the target image. Thereby, even in the case of various images, particularly, for example, an image in which the target person is not wearing a predetermined wearing item, virtual wearing of the wearing item on the target person can be realized.

[0100] In addition, the image generation system 100 can generate a reference image with a color pattern mapped thereto based on an image generation model that has been further learned by associating a predetermined color pattern with each predetermined region of the reference clothing item, and input the reference image with the color pattern mapped thereto into the image generation model to generate a target image. Thereby, the image generation system 100 can more accurately realize virtual clothing item wearing for the target person.

[0101] In addition, the image generation system 100 can deform a pre-set reference clothing item model so as to correspond to the pose shown by the target person model, and generate a reference image based on the deformed reference clothing item model. Thereby, the image generation system 100 can more accurately generate a target image.

[0102] In addition, the image generation system 100 can generate a mask image based on the input image, and further generate an output image by synthesizing the target image in the masked region based on the mask image. Thereby, the image generation system 100 can realize virtual clothing item wearing for the target person.

[0103] In addition, the image generation system 100 can input the output image generated by synthesizing the target image in the masked region into the gap compensation model to generate an output image with the gap between the target image and the mask image compensated. Thereby, the image generation system 100 can generate a more accurate output image.

[0104] In addition, the image generation system 100 can sequentially acquire input images and sequentially generate output images corresponding to the sequentially acquired input images. Thereby, the image generation system 100 can realize virtual clothing item wearing in real time.

[0105] In addition, the image generation system 100 can acquire an input video, obtain each of a plurality of frames included in the input video as an input image, generate an output image corresponding to each of the plurality of acquired frames, and output an output video based on each of the generated output images. Thereby, the image generation system 100 can realize virtual wearing of wearing items for an object photographed as a video.

[0106] Note that this embodiment is for facilitating the understanding of the present invention and is not for limiting and interpreting the present invention. The present invention can be changed / improved without departing from its gist, and equivalents thereof are also included in the present invention.

[0107] In the present invention, the "part" does not simply mean a physical means, and also includes a case where the function of the "part" is realized by software. Also, even if the function of one "part" or device is realized by two or more physical means, devices, or software, or the functions of two or more "parts" or devices are realized by one physical means, device, or software, it is also acceptable.

Explanation of Reference Numerals

[0108] 100 Image generation system, 110 Storage unit, 120 Learning unit, 130 Input image acquisition unit, 140 Subject model generation unit, 150 Reference image generation unit, 160 Target image generation unit, 170 Mask generation unit, 180 Output image generation unit, 190 Output unit, 200 User device, 300 Photographing device

Claims

1. A computer, an input image acquisition unit that acquires an input image of a target person, a target person model generation unit that inputs the input image into a 3D model generation unit that generates a 3D model of the predetermined object from an image of the predetermined object taken, and generates a target person model that is a 3D model of the target person, a reference image generation unit that generates a reference image of a virtual reference wearable in a 3D shape corresponding to the pose of the target person in the input image based on the target person model, an object image generation unit that inputs the reference image into an image generation model that is learned by associating a plurality of learning reference images of reference wearables in shapes corresponding to each of a plurality of poses and a plurality of learning target images of target wearables in shapes corresponding to each of the plurality of poses, and generates an object image of the target wearable in a 3D shape corresponding to the pose indicated by the reference image, an output image generation unit that generates an output image of the target person wearing the target wearable based on the input image and the target image, An image generation program for realizing the above.

2. Each of the plurality of learning reference images is an image in which a predetermined color pattern is mapped for each predetermined region of the reference wearable, The image generation model is an image generation model that is further learned by associating the color pattern, The reference image generation unit generates the reference image in which the color pattern is mapped, The object image generation unit inputs the reference image in which the color pattern is mapped into the image generation model to generate the object image. The image generation program according to claim 1.

3. The reference image generation unit deforms a reference wearable model, which is a 3D model of the reference wearable set in advance, so as to correspond to the pose indicated by the target person model, and generates the reference image based on the deformed reference wearable model. The image generation program according to claim 1 or 2.

4. The reference wearable model is a 3D model in which a predetermined color pattern is mapped for each predetermined region of the reference wearable, The reference image generation unit generates the reference image in which the color pattern is mapped. The image generation program according to claim 3.

5. The computer further includes a mask generation unit that generates a mask image obtained by masking a region of a wearable item worn by the subject based on the input image. The output image generation unit generates the output image by synthesizing the target image with the masked region based on the mask image. The image generation program according to claim 1 or 2.

6. The output image generation unit inputs the output image generated by the synthesis into a gap filling model that fills a gap portion included in the input image and outputs a predetermined image, and generates the output image in which the gap between the target image and the mask image is filled. The image generation program according to claim 5.

7. The input image acquisition unit sequentially acquires the input images. The output image generation unit sequentially generates the output images corresponding to the input images acquired in the order. The image generation program according to claim 1 or 2.

8. In the computer, a video acquisition unit that acquires an input video in which the subject is photographed; a video output unit that outputs an output video based on the output image; are further implemented, The input image acquisition unit acquires each of a plurality of frames included in the input video as the input image. The output image generation unit generates the output images corresponding to each of the plurality of acquired frames, respectively. The video output unit outputs the output video based on each of the generated output images. The image generation program according to claim 1 or 2.

9. an input image acquisition unit that acquires an input image of a subject; a subject model generation unit that inputs the input image into a three-dimensionalization model that generates a three-dimensional model of a predetermined object from an image of the predetermined object photographed, and generates a subject model that is a three-dimensional model of the subject; a reference image generation unit that generates a reference image of a virtual reference wearable item having a three-dimensional shape corresponding to the pose of the subject in the input image based on the subject model; an object image generation unit that inputs the reference image into an image generation model learned by associating a plurality of learning reference images of reference wearable items having shapes corresponding to each of a plurality of poses and a plurality of learning target images of target wearable items having shapes corresponding to each of the plurality of poses, and generates an object image of the target wearable item having a three-dimensional shape corresponding to the pose indicated by the reference image. An output image generation unit that generates an output image of the subject wearing the target wearable item based on the input image and the target image; An image generation system. **Claim 10** A computer: Obtains an input image of a subject; Inputs the input image into a 3D model generation unit that generates a 3D model of a predetermined object from an image of the predetermined object that has been photographed, and generates a subject model that is a 3D model of the subject; Generates a reference image of a virtual reference wearable item having a 3D shape corresponding to the pose of the subject in the input image based on the subject model; Inputs the reference image into an image generation model that has been learned by associating a plurality of learning reference images of reference wearable items having shapes corresponding to each of a plurality of poses, and a plurality of learning target images of target wearable items having shapes corresponding to each of the plurality of poses, and generates a target image of the target wearable item having a 3D shape corresponding to the pose indicated by the reference image; Generates an output image of the subject wearing the target wearable item based on the input image and the target image; An image generation method.

Citation Information

Patent Citations

  • Image processing device, learned model, image collection device, image processing method, and image processing program

    WO2020171237A1