System

The system preprocesses and integrates apparel item photos with model images using generative AI to create realistic wearing look images and videos, addressing the inefficiencies in current online shopping systems and enhancing the shopping experience.

JP2026022447APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123964
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Current online shopping systems lack the ability to efficiently and cost-effectively provide high-quality visual representations of apparel items, particularly in terms of how they look and fit, due to the time and resources required for generating model images and videos.

Method used

A system that preprocesses apparel item photos, integrates them with model images using generative AI models, generates natural-looking worn-look images and videos, and applies user-specified backgrounds to create realistic visual content.

Benefits of technology

Enables efficient generation of high-quality wearing look images and model videos, allowing consumers to visualize apparel items in a realistic manner, improving online shopping experiences and reducing the risk of returns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022447000001_ABST
    Figure 2026022447000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a picture of an apparel item; means for receiving a model image; means for receiving a background image; means for preprocessing the received picture of the apparel item; a generative AI model for generating a wearing-look image using the preprocessed picture of the apparel item and the model image; means for generating a model animation based on the generated wearing-look image; means for applying the background image to generate a final wearing-look image and model animation; and means for providing the generated wearing-look image and model animation to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] A lack of ways for consumers to visualize products more clearly is a problem in online shopping. The apparel industry, in particular, needs to clearly represent the actual look and fit of products, but current technology cannot adequately meet this need. Furthermore, conventional technology requires significant cost and time to prepare numerous model images and videos, so a more efficient method is needed. [Means for solving the problem]

[0005] To solve the above problems, the present invention proposes the following means: A means is provided for receiving photographs of apparel items and preprocessing them. Preprocessing includes adjusting resolution, removing noise, and automatically cropping the background. Furthermore, a means is provided for generating natural-looking worn-look images through a generative AI model using the preprocessed apparel item photographs and model images. A means is also provided for setting an animation framework based on the generated worn-look images and integrating multiple frames to generate a model video. Finally, a background image is applied to generate worn-look images and model videos that resemble a real environment, which are then provided to the user. This makes it possible to provide high-quality visual content efficiently and at low cost.

[0006] "Apparel items" is a general term for fashion-related products such as clothing and accessories.

[0007] A "photograph" refers to a still image that records visual information about an object.

[0008] "Model image" refers to an image or data of a model wearing an apparel item, and serves as a reference when generating a wearing look.

[0009] "Background image" refers to an image or data of a virtual background on which a model wearing an apparel item is set.

[0010] "Preprocessing" refers to the process of adjusting the resolution of the received photographs of the apparel items, removing noise, and cutting out the background.

[0011] "Generative AI model" refers to an algorithm and its model that uses artificial intelligence techniques to generate specific outputs from input data.

[0012] A "wearing look image" refers to an image that simulates how an apparel item would look when worn by a model.

[0013] "Model video" refers to a video generated based on a wearing look image, depicting a model moving while wearing an apparel item.

[0014] An "animation framework" defines the basic frame structure and movement sequences used to generate moving images.

[0015] "Means for providing to users" refers to a method or process for providing the generated worn look images and model videos to users in an accessible format. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system for generating wear look images and model videos from photos of apparel items, and uses AI technology to generate them using photos, background images, and model images uploaded by users. Specifically, the process is carried out in the following steps.

[0038] System program processing description

[0039] Upload a photo

[0040] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to the server. For example, a user selects a photo of a jacket and uploads a background image of a city and an image of Model A.

[0041] Photo pre-processing

[0042] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out other unnecessary background elements, and adjusts the resolution accordingly.

[0043] Wearing look image generation

[0044] The server passes the preprocessed apparel item photos and model images to a generative AI model. The generative AI model integrates the apparel item and model's body data to generate seamless, natural-looking images of the item being worn. For example, the generative AI model integrates a photo of a jacket with an image of Model A to generate an image of the model wearing the jacket.

[0045] Model video generation

[0046] The server creates an animation framework based on the generated images of the clothing look, and adds movement using a generative AI model. For example, it generates a video of model A wearing the jacket and spinning or walking.

[0047] Applying a Background Image

[0048] The server applies a user-specified background image to the generated wearing look images and model video. For example, it creates a video of model A walking around wearing the jacket against a cityscape.

[0049] Providing results

[0050] The server generates the final worn-look images and model videos, saves them as files, and generates and provides a download link to the user. The user can access the link to download the generated images and videos. For example, the user can click the provided link to obtain images and videos of Model A wearing the jacket.

[0051] Through these processes, the system of the present invention can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to consumers.

[0052] The processing flow will be explained below.

[0053] Step 1:

[0054] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0055] Step 2:

[0056] The server performs pre-processing on the uploaded apparel item photos according to their purpose. Specifically, it adjusts the image resolution to ensure appropriate image quality. It also performs noise reduction to improve image clarity. It also performs automatic background cropping to extract the outline of the apparel item. For example, the server identifies the outline of a jacket and crops out unnecessary background.

[0057] Step 3:

[0058] The server inputs the preprocessed photos of the apparel items and the model images into a generative AI model. This generative AI model integrates the apparel items with the model's posture and body data to generate natural, seamless images of how the items look when worn. For example, the generative AI model combines a photo of a jacket with an image of Model A to generate an image of Model A wearing the jacket.

[0059] Step 4:

[0060] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets frames that simulate the images and movements of the model as seen from all directions. This allows for greater reproducibility of the animation.

[0061] Step 5:

[0062] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0063] Step 6:

[0064] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create smooth, natural videos.

[0065] Step 7:

[0066] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0067] Step 8:

[0068] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0069] Through the above steps, the server can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to users.

[0070] Example 1

[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0072] Conventional apparel fitting simulation systems often make it difficult for users to specifically check how they will look when wearing apparel. For example, there is a growing need to visually understand how an item will look in dynamic situations, not just still images, but systems that offer such functionality are limited. Furthermore, automated, highly accurate processing is required for image quality and background application, but current systems have difficulty achieving satisfactory results. Therefore, the challenge is to develop a system that can generate high-quality wearing look images and videos using photos of apparel items uploaded by users and provide them to users.

[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0074] In this invention, the server includes means for receiving images of apparel items, means for receiving images of people, means for receiving background images, a generative AI model for generating worn look images using preprocessed apparel item images and images of people, means for generating animated videos based on the generated worn look images, means for generating final worn look images and animated videos by applying the background image, and means for providing the generated worn look images and animated videos to users, thereby enabling users to easily perform high-quality try-on simulations.

[0075] "Apparel item images" refer to digital images of fashion items such as clothing and accessories.

[0076] "Human Image" means a digital image of a model or user that is used to simulate trying on apparel items.

[0077] A "background image" is a digital image of a particular scene or location specified by the user, which is combined with an image of a person wearing an apparel item.

[0078] "Pre-processing" refers to a series of operations performed on uploaded apparel item images, including resolution adjustment, noise removal, and automatic background cropping.

[0079] "Generative AI models" are algorithms and systems that use machine learning techniques to generate new images and videos.

[0080] A "wearing look image" is a digital image created by combining images of a person wearing an apparel item.

[0081] An "animated video" is a moving digital image created by integrating a series of frames.

[0082] "User" refers to a person who uses the system to perform a try-on simulation and receives images and videos.

[0083] The present invention provides a system for generating high-quality images and animation videos of apparel items from images of the apparel items, and providing the images and animation videos to users. The system operates based on the interaction between a server, a terminal, and a user.

[0084] First, the user uploads the apparel item image, background image, and model image to the server from a terminal, such as a commonly used personal computer or smartphone.

[0085] The server then performs preprocessing on the received apparel item images. This preprocessing includes resolution adjustment, noise removal, and automatic background cropping. To perform these processes, software libraries such as OpenCV, PIL (Python Imaging Library), and Mask R-CNN are used. Specifically, OpenCV is used to resize the images for resolution adjustment, and Gaussian Blur and Median Filter are applied for noise removal. For automatic background cropping, Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0086] The preprocessed apparel item images are passed to the generative AI model by the server. This generative AI model uses algorithms such as DALL-E and Stable Diffusion to generate a wearing look image based on the preprocessed apparel item images and model images. As a concrete example, the following prompt sentence is input to the generative AI model:

[0087] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0088] After the wearable look image is generated, the server generates an animated video based on this image. Using animation software such as Blender or Adobe After Effects, the model's movements (e.g., rotation or walking) are defined, and the generative AI model adds realistic animation based on those movements. As a specific example, the following prompt sentence is input to the generative AI model:

[0089] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0090] Finally, the server applies the background image specified by the user to complete the look images and animation video. To apply the background image, video editing software such as Adobe Premiere Pro is used to composite the background onto the model images and videos wearing the apparel items.

[0091] The completed wearable look images and animation videos are stored by the server and provided to the user. Specifically, they are stored in a cloud storage service (e.g., an S3 bucket) and an access link is provided to the user. The user can download the generated images and videos by accessing the link.

[0092] As a result, the system of the present invention provides an environment in which users can visually check their apparel items in a high-quality and realistic manner.

[0093] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0094] Step 1:

[0095] The user uploads images of apparel items, background images, and model images from their device to the server. Specifically, the user selects images of clothes, accessories, etc. from their local disk using, for example, a PC or smartphone through the system interface and sends them to the server. This allows the server to receive input data for processing these images.

[0096] Input: Apparel item image, background image, model image

[0097] Output: Image data uploaded to the server

[0098] Step 2:

[0099] The server performs preprocessing on the received apparel item images. First, resolution adjustment is performed. The server uses OpenCV to resize the image resolution to an appropriate size. Next, noise reduction is performed. Techniques such as Gaussian Blur and Median Filter are used to remove noise from the image. Finally, automatic background cropping is performed. The deep learning model Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0100] Input: Uploaded apparel item image

[0101] Output: Preprocessed apparel item images

[0102] Step 3:

[0103] The server passes the preprocessed apparel item images and model images to a generative AI model to generate a worn-look image, using generative AI models such as DALL-E and Stable Diffusion. The server inputs a prompt statement into the generative AI model to generate a natural, seamless composite image.

[0104] Input: Preprocessed apparel item images, model images

[0105] Output: Images of the outfit worn

[0106] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0107] Step 4:

[0108] The server generates an animated video based on the generated images of the outfits. Using animation software such as Blender or Adobe After Effects, the model's movements are defined. For example, an animation of Model A rotating can be defined, and then a generative AI model can be used to create an animated video with detailed movements added based on that movement.

[0109] Input: Image of the look you're wearing

[0110] Output: Animation video

[0111] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0112] Step 5:

[0113] The server composites the background image with the animation video, and then the user can use video editing software such as Adobe Premiere Pro to apply the background image specified by the user to the generated animation video, making the video more realistic and polished.

[0114] Input: Animation video, background image

[0115] Output: Final wear look video

[0116] Step 6:

[0117] The server saves the final generated wearable look images and animation videos in cloud storage and provides users with download links. For example, the server may save the files in an AWS S3 bucket and notify the user of the URL via email or dashboard. Users can download the images and videos by clicking the provided link.

[0118] Input: Final wear look images and animation video

[0119] Output: Provide download link to user

[0120] (Application example 1)

[0121] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0122] In conventional virtual and physical stores, consumers have limited means to visually check how apparel items look when worn, making it difficult to fully understand how the items will fit and coordinate with each other when actually worn. Furthermore, even in online shopping experiences, customers tend to feel unsure about product selection, and there is a high risk of returning items after purchase.

[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0124] In this invention, the server includes means for receiving photos of apparel items, means for receiving model images, means for receiving background images, means for pre-processing the received photos of the apparel items, a generative AI model for generating worn-look images using the pre-processed photos of the apparel items and the model images, means for generating model videos based on the generated worn-look images, means for generating final worn-look images and model videos by applying a background image, means for providing the generated worn-look images and model videos to a user, interface means for executing each of the above means using a smartphone, means for applying a background image to the worn-look images and model videos to simulate displays in a physical store or a virtual store, and means for generating prompts for the generative AI model. This allows consumers to concretely and visually see what the apparel items will look like when actually worn, improving the online purchasing experience and reducing the risk of returns after purchase.

[0125] "Apparel item photos" refers to images of fashion-related products such as clothing and accessories.

[0126] "Model Image" refers to an image showing a person wearing an apparel item.

[0127] "Background image" refers to an image such as a landscape or scene that appears behind an apparel item or model image.

[0128] "Preprocessing" refers to performing processes such as adjusting the image resolution, removing noise, and cutting out the background.

[0129] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate new images and videos from input images.

[0130] "Worn look image" refers to an image in which the apparel item is naturally integrated into the model and appears to be worn.

[0131] "Model video" refers to a video showing a model moving and posing based on an image of the look worn.

[0132] "Means for providing to users" refers to a method for providing the generated images and videos in a form that allows users to access them.

[0133] "Interface means" refers to an interface that allows a user to perform various operations through a smartphone.

[0134] "Means to simulate" refers to a method for applying a background image to replicate the display in a physical or virtual store environment.

[0135] "Means for generating prompt sentences" refers to a method for automatically generating input sentences for a generative AI model.

[0136] The present invention is a system for generating a wear look image and a model video from a photograph of an apparel item, and aims to improve the shopping experience of consumers in real stores and virtual stores using a smartphone. The system of the present invention is configured as follows.

[0137] Upload a photo

[0138] Users use their smartphones to upload photos of apparel items, model images, and background images to the system, and a user interface is provided to allow for intuitive operation.

[0139] Photo pre-processing

[0140] The server preprocesses the uploaded photos of apparel items. Specifically, it uses software such as OpenCV to adjust the resolution, remove noise, and crop the background. For example, it adjusts the uploaded photo of a jacket to the appropriate resolution, crops the background, and removes noise.

[0141] Wearing look image generation

[0142] The preprocessed photos of the apparel items and the model images are input into a generative AI model to generate a worn look image. The generative AI model uses a machine learning model such as StyleGAN, which naturally integrates the apparel items into the model, resulting in an image that looks as if the item is being worn.

[0143] Model video generation

[0144] Based on the generated images of the clothing look, a video of the model is generated. The server uses the generative AI model to generate a video of the model moving and posing. For example, a video of the model spinning or walking while wearing the jacket is generated.

[0145] Applying a Background Image

[0146] A user-specified background image is applied to the wearable look images and model video. For example, a video of a model walking around wearing a jacket against a cityscape background is generated. The background image is then composited using libraries such as OpenCV.

[0147] Providing results

[0148] The server generates the final wearable look images and model videos and provides them to the user. Specifically, the server saves the generated images and videos as files and provides the user with a download link. The user can access the link to download the generated content.

[0149] Interface methods and prompt generation

[0150] Furthermore, the present invention provides an interface means that allows users to perform each operation using a smartphone. Through a smartphone application, users can intuitively upload content, and prompt sentences for the generative AI model are automatically generated. Examples of prompt sentences include:

[0151] Input image: path / to / apparel.jpg

[0152] Background image: path / to / background.jpg

[0153] Model image: path / to / model.jpg

[0154] Generated content: Images and videos of a model wearing the jacket in the street

[0155] By passing this prompt to the generative AI model, users can easily obtain high-quality images of the outfits being worn and videos of the models.

[0156] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0157] Step 1:

[0158] Users use their smartphones to upload photos of apparel items, model images, and background images to the system. This allows users to intuitively select various images through the application interface and send them to the server. The inputs are photos of apparel items, model images, and background images, which are then provided to the server.

[0159] Step 2:

[0160] The server preprocesses the received photos of apparel items using software such as OpenCV to adjust the resolution (resizing the input image to 256x256 pixels), remove noise (using non-local means), and crop the background (using edge detection and masking). The output is an image of the apparel item after preprocessing.

[0161] Step 3:

[0162] The server inputs the preprocessed apparel item photos and model images into a generative AI model. Specifically, it uses a generative AI model such as StyleGAN to generate a worn-look image in which the apparel item is naturally integrated into the model. In this process, the inputs are the preprocessed apparel item images and model images, and the output is a worn-look image.

[0163] Step 4:

[0164] The server generates a model video based on the generated wearable look images. Using a generative AI model, it generates a video of the model moving and posing. The input is a wearable look image, and the generated multiple frames are integrated to output a video.

[0165] Step 5:

[0166] The server applies a background image to the generated wearing-look images and model videos. Specifically, it uses libraries such as OpenCV to synthesize the background image with the wearing-look images and videos. The input is the wearing-look image and background image, and the output is the final wearing-look image and video with the background applied.

[0167] Step 6:

[0168] The server generates the final worn-look images and model videos and provides them to the user. Specifically, it saves the generated images and videos as files and generates a download link to provide to the user. The user can download the generated content by accessing the link. The input is the final worn-look images and videos, and the output is the download link.

[0169] Step 7:

[0170] The server generates prompt sentences and automatically generates input sentences for the generative AI model. These prompt sentences automate the processing flow of the entire system, helping users easily obtain high-quality content. For example, a prompt sentence might be generated as follows: "Input image: path / to / apparel.jpg, background image: path / to / background.jpg, model image: path / to / model.jpg, generated content: images and videos of a model wearing a jacket in the city." The inputs are an image of an apparel item, a model image, and a background image, and the output is the prompt sentence.

[0171] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0172] The present invention is a system that generates wearing look images and model videos from photographs of apparel items, and provides personalized content to each user by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system is explained below, along with specific examples.

[0173] System program processing description

[0174] Upload a photo

[0175] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0176] Applying the Emotion Engine

[0177] The server uses an emotion engine to analyze the user's facial expression data and emotion-related data received from the user's device. This allows the server to recognize the user's current emotional state. For example, if the user is smiling, it is recognized as a positive emotion.

[0178] Photo pre-processing

[0179] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out unnecessary background, and adjusts the resolution accordingly.

[0180] Emotion-Based Adjustment

[0181] The server then changes the color and design of the apparel item based on the user's recognized emotion, adjusting the generative AI model to match the user's emotion. For example, if the user is in an active mood, a bright-colored jacket might be suggested.

[0182] Wearing look image generation

[0183] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0184] Model video generation

[0185] The server creates an animation framework based on the generated wearing look images. Specifically, it sets frames that simulate the image and movement of the model as seen from all directions. This improves the reproducibility of the video.

[0186] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0187] Applying a Background Image

[0188] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0189] Providing results

[0190] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0191] Through these processes, the system of the present invention can generate personalized, high-quality wearing look images and model videos based on the user's emotions and provide them to consumers.

[0192] The processing flow will be explained below.

[0193] Step 1:

[0194] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0195] Step 2:

[0196] The server receives facial expression and voice data from the user's device and analyzes this data using an emotion engine. For example, if the user's facial expression is smiling, the server recognizes that the user has a positive emotion.

[0197] Step 3:

[0198] The server performs the necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, in a photo of a jacket, the server identifies the jacket's outline and crops out any unnecessary background.

[0199] Step 4:

[0200] The server adjusts the color and design of the apparel item based on the user's perceived emotion. For example, if the user is perceived as being in an active mood, the server suggests a brightly colored jacket.

[0201] Step 5:

[0202] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model then integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0203] Step 6:

[0204] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets the frames necessary to simulate the model's movements from all angles, improving the reproducibility of the animation.

[0205] Step 7:

[0206] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A wearing a jacket and spinning or walking in each frame.

[0207] Step 8:

[0208] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create a smooth, natural video.

[0209] Step 9:

[0210] The server analyzes the user's emotional data and selects an appropriate background image. For example, if the user is relaxing, the server will select a background of a park or natural scenery.

[0211] Step 10:

[0212] The server then synthesizes the generated look images and model video with an appropriate background image. For example, a video of model A wearing a jacket against a park landscape can be created.

[0213] Step 11:

[0214] The server saves the final look images and model video as files and generates a download link to provide to the user. The user can download the final result using the link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0215] Through these steps, the system of the present invention makes it possible to generate and provide to consumers high-quality wearing look images and model videos that are personalized based on the user's emotions.

[0216] Example 2

[0217] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0218] Conventional apparel fitting systems do not personalize the experience based on the user's emotions, and only provide standard images and videos, making it difficult to provide content tailored to the needs and emotions of individual users. Furthermore, there are technical challenges in obtaining natural, high-quality output when processing photos of apparel items and generating model videos.

[0219] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0220] In this invention, the server includes means for receiving images of apparel items, means for receiving images of models, means for receiving images of backgrounds, means for preprocessing the received images of apparel items, a generative AI model for generating worn-look images using the preprocessed images of apparel items and images of models, means for analyzing and recognizing user emotion data, means for adjusting the color and design of the apparel items based on the user emotion data, means for generating videos of the models based on the generated images of worn looks, means for generating final images of worn looks and videos of the models by applying the background images, and means for providing the generated images of worn looks and videos of the models to the user. This enables the generation of personalized, high-quality images of worn looks and videos of the models based on the user's emotions.

[0221] "Apparel item images" are photographs or digital images of fashion items such as clothing and accessories.

[0222] "Model Image" means a photograph or digital image of a human model used to simulate trying on an apparel item.

[0223] "Background image" refers to visual materials such as photographs or illustrations that show the background onto which digital characters or apparel items are projected.

[0224] "Preprocessing" refers to processing operations such as adjusting resolution, removing noise, and removing background that are performed on image data.

[0225] A "generative AI model" is an artificial intelligence algorithm that uses deep learning techniques to generate images and videos, creating new images and videos based on specific input data.

[0226] "User emotion data" is information about the user's psychological state obtained through facial expression analysis, biometric signals, and the like.

[0227] A "wearing look image" is a still image that simulates a model wearing an apparel item.

[0228] "Model video" refers to a moving image that simulates a model moving while wearing an apparel item.

[0229] "Means of providing" refers to the process or functionality of distributing the generated digital content to users in a downloadable format.

[0230] This clearly defines the technical elements included in the claims and clarifies their scope of application.

[0231] The present invention is a system for generating images of appearances of apparel items and model videos from photographs of apparel items, and by combining this with an emotion engine that recognizes the emotions of users, the system provides personalized content to individual users. Specific embodiments of the system are described below.

[0232] The user uses a device to upload images of apparel items, background images, and model images to the server. The user opens the device's browser or a dedicated application and clicks the upload button to select and send each image file. For example, the user selects a photo of a jacket, a background image of a city, and a photo of Model A, and uploads them to the server.

[0233] The server analyzes the facial expression data and emotion-related data received from the user's device using an emotion engine. The user's facial expression data is captured by a camera and sent to the server in real time. The server uses an emotion engine (e.g., a deep learning-based facial expression recognition engine) to analyze the facial expression data and recognize the user's emotional state (e.g., smiling, surprised, etc.).

[0234] The server pre-processes the uploaded apparel item images by adjusting the resolution using an image processing library (e.g., OpenCV), applying a noise reduction algorithm, and performing image segmentation to identify the contours of the apparel items and remove unwanted background.

[0235] The server then adjusts the color and design of the apparel item based on the user's emotional data recognized by the emotion engine. The server uses generative AI models (e.g., GANs) to select bright colors (e.g., yellow or red) when the user is in an active mood.

[0236] After this preprocessing and emotion adjustment step, the server inputs the preprocessed apparel item images and the model's image into a generative AI model to generate a worn-look image. The deep learning-based generative AI model integrates the apparel item and model's body data to generate a natural and seamless worn-look image.

[0237] A model video is generated based on the generated images of the outfits. The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generated AI model adds movements (e.g., rotation, walking) to each frame, and then combines multiple frames to create a model video.

[0238] To apply the background image, the server uses an image synthesis algorithm. It integrates the background image with the look image and model video, and adjusts the color and shadows to make the background and model look natural. For example, it creates a video of Model A walking in a jacket against the backdrop of a cityscape.

[0239] Finally, the server stores the generated worn look images and model videos in cloud storage and generates a download link to provide them to the user. The user can download the final images and videos using the provided link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0240] Specific examples

[0241] Prompt Sentence Examples

[0242] "Upload a photo of the jacket. Use a cityscape as the background image and combine it with an image of Model A to generate images and videos of the suggested look."

[0243] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0244] Step 1:

[0245] Input: The user selects an apparel item image, a background image, and a model image via their device.

[0246] Specific operation: The user opens a browser or a dedicated application on their device, clicks the upload button, and selects each image file. For example, they select a photo of a jacket, a background image of a city, and an image of Model A. After selection, those files are uploaded to the server.

[0247] Output: All selected images are sent to the server.

[0248] Step 2:

[0249] Input: Facial expression data and emotion-related data sent from the user's device.

[0250] Specific operation: The user captures facial expression data in real time using the device's camera. The user's facial expression is recognized as a state such as a smile or surprise. This data is sent to the server.

[0251] Data processing / calculation: The server uses an emotion engine to analyze the received facial expression data and recognize the user's emotional state. For example, a deep learning-based facial expression recognition engine is used.

[0252] Output: User's current emotional state data.

[0253] Step 3:

[0254] Input: An image of an apparel item.

[0255] What it does: The server uses an image processing library (e.g., OpenCV) to adjust the resolution and apply noise reduction algorithms. It also performs image segmentation to identify the contours of apparel items and cut out unwanted background. For example, it automatically traces the outline of a jacket and removes the background.

[0256] Data processing / calculation: Resolution optimization, noise removal, and background removal.

[0257] Output: Preprocessed apparel item images.

[0258] Step 4:

[0259] Input: Recognized user emotional state data, preprocessed apparel item images.

[0260] Specific operation: The server adjusts the color and design of the apparel item according to the user's emotional state as recognized by the emotion engine. For example, if the user is in a lively mood, the color will be changed to a bright color (e.g., yellow or red).

[0261] Data manipulation / calculation: Use the customization engine to make color and design changes.

[0262] Output: Images of apparel items adjusted based on emotion.

[0263] Step 5:

[0264] Input: Preprocessed and adjusted apparel item images, model images.

[0265] How it works: The server uses a generative AI model (e.g., GANs) to combine images of apparel items with images of models to generate natural, seamless images of the items being worn. For example, it uses deep learning technology to generate an image of Model A wearing a jacket.

[0266] Data processing / computation: Image synthesis and generation.

[0267] Output: Generated worn look images.

[0268] Step 6:

[0269] Input: Generated worn look images.

[0270] How it works: The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generative AI model adds movements (e.g., rotating, walking) for each frame to create the video.

[0271] Data processing / computation: Video frame generation and integration.

[0272] Output: The generated model video.

[0273] Step 7:

[0274] Input: Background image, generated look images and model video.

[0275] How it works: The server uses an image synthesis algorithm to apply a background image to each frame, adjusting colors and shadows to make the background and model look natural. For example, the server creates a video of Model A walking in a jacket against a cityscape.

[0276] Data processing / calculation: background image synthesis and color adjustment.

[0277] Output: Final worn-look images and model video.

[0278] Step 8:

[0279] Input: Final wear-look images and model video.

[0280] Specific operation: The server saves the generated content to cloud storage and generates a download link to provide to the user. The user can use the provided link to download the final images and videos. For example, the user accesses the download link to obtain images and videos of Model A wearing the jacket.

[0281] Output: The content file that the user gets via the download link.

[0282] (Application example 2)

[0283] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0284] Modern online shopping sites have a problem in that users cannot actually try on apparel products when purchasing them. This can lead to problems such as anxiety when purchasing products and an increased rate of returns. Furthermore, personalized try-on experiences based on each user's emotions and preferences are rarely offered. The present invention aims to solve this problem by analyzing a user's emotions and providing a personalized virtual try-on experience based on the results.

[0285] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving photographs of apparel items, means for receiving model images, means for receiving background images, means for receiving emotion-related data, an emotion engine for analyzing the received emotion-related data, means for preprocessing the received photographs of apparel items, a generative AI model for generating worn look images adjusted based on the user's emotion using the preprocessed photographs of the apparel items and the model images, means for generating model videos based on the generated worn look images, means for generating final worn look images and model videos by applying the background image, and means for providing the generated worn look images and model videos to the user. This allows the user to enjoy a high-quality virtual try-on experience personalized to their emotions.

[0286] "Photos of apparel items" are image data of specific clothing, accessories, etc. uploaded by users.

[0287] A "model image" is a photograph of the user or a photograph of a person selected by the user, and is image data used to simulate the appearance of the user wearing an apparel item.

[0288] A "background image" is image data of a background selected by the user, and is used as the background of a wearing look image or a model video.

[0289] "Emotion-related data" is data obtained from the user's facial expressions and the like, and is used to analyze the user's current emotional state.

[0290] An "emotion engine" is a software or hardware component for analyzing emotion-related data and recognizing a user's emotional state.

[0291] "Pre-processing" refers to image processing operations such as resolution adjustment, noise removal, and background removal that are performed to prepare photographs of apparel items for use.

[0292] A "generative AI model" is an algorithm or system that uses AI technology to integrate pre-processed photographs of apparel items with model images to generate personalized worn-look images.

[0293] A "wear look image" is an image generated by a generative AI model that simulates an apparel item being worn by a model.

[0294] "Model video" is video data that uses multiple frames to recreate the appearance of a model wearing an apparel item, based on a wearing look image.

[0295] The "means for providing to the user" refers to an interface or link generation function that allows the user to view or download the generated wearing look images and model videos.

[0296] The system for implementing this invention comprises a series of processes that allow a user to have a virtual try-on experience, and hardware and software components for realizing this process.

[0297] The central part of the system is the server, which includes:

[0298] 1. Receiving photos and data:

[0299] Users use devices (such as smartphones or HMDs) to upload photos of apparel items, background images, model images, and emotion-related data to the server. The emotion-related data is necessary for analyzing the user's facial expressions.

[0300] 2. Emotion analysis:

[0301] The server then uses the emotion engine to analyze the user's emotional state based on the emotion-related data it receives. This analysis uses software such as machine learning models and emotion recognition algorithms. For example, if the user is smiling with satisfaction, their emotional state is recognized as positive.

[0302] 3. Image preprocessing:

[0303] The server preprocesses the received photos of apparel items by adjusting the resolution, removing noise, and cropping the background using image processing libraries such as OpenCV. By removing unnecessary background, the outlines of the items become clearer.

[0304] 4. Emotional customization:

[0305] The server adjusts the color and design of the apparel item based on the analyzed user's emotion. For example, if the user is in a lively mood, apparel items with bright colors and lively designs are selected.

[0306] 5. Generation of worn look images and model videos:

[0307] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. This generative AI model uses algorithms such as GAN (Generative Adversarial Networks). Based on the generated images, a model video is then generated. This process uses a technique to create a video from multiple frames.

[0308] 6. Applying a background image:

[0309] The server then combines the generated clothing look images and model video with a background image specified by the user. For example, it adjusts the background to make the model appear natural walking while wearing the apparel items against a cityscape.

[0310] 7. Providing Results:

[0311] The server finally generates a download link to provide the worn look images and model videos to the user, who can use the link to obtain the results.

[0312] Specific examples

[0313] For example, a user using an HMD provides the following prompt to the system:

[0314] Prompt statement:

[0315] The user uploaded a photo of themselves smiling with a satisfied expression and chose a beautiful beach background. The task was to crop the photo of the red dress, adjust it to the optimal resolution, and generate a video on the HMD of the model wearing the red dress and walking against the beach background.

[0316] Based on this prompt, the system executes each of the above steps to provide a virtual try-on experience tailored to the user's emotions, allowing the user to enjoy a personalized try-on experience for apparel items tailored to their emotions.

[0317] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0318] Step 1:

[0319] The user uploads a photo of the apparel item, a background image, a model image, and the user's facial expression from the device to the server. At this time, the device sends this data to the server as a single request. The input is the photo of the apparel item, the background image, the model image, and facial expression data, and the output is the state in which the data has been sent to the server.

[0320] Step 2:

[0321] The server inputs the received facial expression data into the emotion engine to analyze the user's emotional state. This emotion engine uses a machine learning model to perform real-time facial expression analysis. The input is the user's facial expression data, and the output is the analyzed emotional state (positive, negative, neutral, etc.).

[0322] Step 3:

[0323] The server preprocesses the photos of the apparel items. This preprocessing includes adjusting the image resolution, removing noise, and automatically cropping the background. Specifically, OpenCV is used to perform these operations. The input is the photos of the apparel items, and the output is the preprocessed photos.

[0324] Step 4:

[0325] The server adjusts the color and design of the preprocessed apparel item based on the emotion analysis results. For example, if the user's emotion is positive, it adjusts the color to a brighter color. The input is a photo of the preprocessed apparel item and the user's emotional state, and the output is a photo of the apparel item adjusted according to the emotion.

[0326] Step 5:

[0327] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. The generative AI model uses GAN (Generative Adversarial Networks). The input is the preprocessed apparel item photos and model images, and the output is a worn look image.

[0328] Step 6:

[0329] The server generates a model video using the generated wearing-look images. In this process, the server generates multiple frames and integrates them to construct a video. The input is the wearing-look images, and the output is the model video.

[0330] Step 7:

[0331] The server composites the generated wearing-look images and model videos with a background image specified by the user. Specifically, the background image is applied to each frame to achieve a natural look. The input is a wearing-look image, a model video, and a background image, and the output is the final wearing-look image and model video with the background composited.

[0332] Step 8:

[0333] The server generates a download link for providing the final worn-look image and model video to the user and provides the link to the user. The input is the final worn-look image and model video, and the output is the download link generated and available for the user to access.

[0334] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0335] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0336] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0337] [Second embodiment]

[0338] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0339] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0340] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0341] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0342] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0343] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0344] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0345] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0346] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0347] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0348] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0349] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0350] This invention is a system for generating wear look images and model videos from photos of apparel items, and uses AI technology to generate them using photos, background images, and model images uploaded by users. Specifically, the process is carried out in the following steps.

[0351] System program processing description

[0352] Upload a photo

[0353] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to the server. For example, a user selects a photo of a jacket and uploads a background image of a city and an image of Model A.

[0354] Photo pre-processing

[0355] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out other unnecessary background elements, and adjusts the resolution accordingly.

[0356] Wearing look image generation

[0357] The server passes the preprocessed apparel item photos and model images to a generative AI model. The generative AI model integrates the apparel item and model's body data to generate seamless, natural-looking images of the item being worn. For example, the generative AI model integrates a photo of a jacket with an image of Model A to generate an image of the model wearing the jacket.

[0358] Model video generation

[0359] The server creates an animation framework based on the generated images of the clothing look, and adds movement using a generative AI model. For example, it generates a video of model A wearing the jacket and spinning or walking.

[0360] Applying a Background Image

[0361] The server applies a user-specified background image to the generated wearing look images and model video. For example, it creates a video of model A walking around wearing the jacket against a cityscape.

[0362] Providing results

[0363] The server generates the final worn-look images and model videos, saves them as files, and generates and provides a download link to the user. The user can access the link to download the generated images and videos. For example, the user can click the provided link to obtain images and videos of Model A wearing the jacket.

[0364] Through these processes, the system of the present invention can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to consumers.

[0365] The processing flow will be explained below.

[0366] Step 1:

[0367] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0368] Step 2:

[0369] The server performs pre-processing on the uploaded apparel item photos according to their purpose. Specifically, it adjusts the image resolution to ensure appropriate image quality. It also performs noise reduction to improve image clarity. It also performs automatic background cropping to extract the outline of the apparel item. For example, the server identifies the outline of a jacket and crops out unnecessary background.

[0370] Step 3:

[0371] The server inputs the preprocessed photos of the apparel items and the model images into a generative AI model. This generative AI model integrates the apparel items with the model's posture and body data to generate natural, seamless images of how the items look when worn. For example, the generative AI model combines a photo of a jacket with an image of Model A to generate an image of Model A wearing the jacket.

[0372] Step 4:

[0373] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets frames that simulate the images and movements of the model as seen from all directions. This allows for greater reproducibility of the animation.

[0374] Step 5:

[0375] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0376] Step 6:

[0377] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create smooth, natural videos.

[0378] Step 7:

[0379] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0380] Step 8:

[0381] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0382] Through the above steps, the server can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to users.

[0383] Example 1

[0384] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0385] Conventional apparel fitting simulation systems often make it difficult for users to specifically check how they will look when wearing apparel. For example, there is a growing need to visually understand how an item will look in dynamic situations, not just still images, but systems that offer such functionality are limited. Furthermore, automated, highly accurate processing is required for image quality and background application, but current systems have difficulty achieving satisfactory results. Therefore, the challenge is to develop a system that can generate high-quality wearing look images and videos using photos of apparel items uploaded by users and provide them to users.

[0386] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0387] In this invention, the server includes means for receiving images of apparel items, means for receiving images of people, means for receiving background images, a generative AI model for generating worn look images using preprocessed apparel item images and images of people, means for generating animated videos based on the generated worn look images, means for generating final worn look images and animated videos by applying the background image, and means for providing the generated worn look images and animated videos to users, thereby enabling users to easily perform high-quality try-on simulations.

[0388] "Apparel item images" refer to digital images of fashion items such as clothing and accessories.

[0389] "Human Image" means a digital image of a model or user that is used to simulate trying on apparel items.

[0390] A "background image" is a digital image of a particular scene or location specified by the user, which is combined with an image of a person wearing an apparel item.

[0391] "Pre-processing" refers to a series of operations performed on uploaded apparel item images, including resolution adjustment, noise removal, and automatic background cropping.

[0392] "Generative AI models" are algorithms and systems that use machine learning techniques to generate new images and videos.

[0393] A "wearing look image" is a digital image created by combining images of a person wearing an apparel item.

[0394] An "animated video" is a moving digital image created by integrating a series of frames.

[0395] "User" refers to a person who uses the system to perform a try-on simulation and receives images and videos.

[0396] The present invention provides a system for generating high-quality images and animation videos of apparel items from images of the apparel items, and providing the images and animation videos to users. The system operates based on the interaction between a server, a terminal, and a user.

[0397] First, the user uploads the apparel item image, background image, and model image to the server from a terminal, such as a commonly used personal computer or smartphone.

[0398] The server then performs preprocessing on the received apparel item images. This preprocessing includes resolution adjustment, noise removal, and automatic background cropping. To perform these processes, software libraries such as OpenCV, PIL (Python Imaging Library), and Mask R-CNN are used. Specifically, OpenCV is used to resize the images for resolution adjustment, and Gaussian Blur and Median Filter are applied for noise removal. For automatic background cropping, Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0399] The preprocessed apparel item images are passed to the generative AI model by the server. This generative AI model uses algorithms such as DALL-E and Stable Diffusion to generate a wearing look image based on the preprocessed apparel item images and model images. As a concrete example, the following prompt sentence is input to the generative AI model:

[0400] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0401] After the wearable look image is generated, the server generates an animated video based on this image. Using animation software such as Blender or Adobe After Effects, the model's movements (e.g., rotation or walking) are defined, and the generative AI model adds realistic animation based on those movements. As a specific example, the following prompt sentence is input to the generative AI model:

[0402] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0403] Finally, the server applies the background image specified by the user to complete the look images and animation video. To apply the background image, video editing software such as Adobe Premiere Pro is used to composite the background onto the model images and videos wearing the apparel items.

[0404] The completed wearable look images and animation videos are stored by the server and provided to the user. Specifically, they are stored in a cloud storage service (e.g., an S3 bucket) and an access link is provided to the user. The user can download the generated images and videos by accessing the link.

[0405] As a result, the system of the present invention provides an environment in which users can visually check their apparel items in a high-quality and realistic manner.

[0406] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0407] Step 1:

[0408] The user uploads images of apparel items, background images, and model images from their device to the server. Specifically, the user selects images of clothes, accessories, etc. from their local disk using, for example, a PC or smartphone through the system interface and sends them to the server. This allows the server to receive input data for processing these images.

[0409] Input: Apparel item image, background image, model image

[0410] Output: Image data uploaded to the server

[0411] Step 2:

[0412] The server performs preprocessing on the received apparel item images. First, resolution adjustment is performed. The server uses OpenCV to resize the image resolution to an appropriate size. Next, noise reduction is performed. Techniques such as Gaussian Blur and Median Filter are used to remove noise from the image. Finally, automatic background cropping is performed. The deep learning model Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0413] Input: Uploaded apparel item image

[0414] Output: Preprocessed apparel item images

[0415] Step 3:

[0416] The server passes the preprocessed apparel item images and model images to a generative AI model to generate a worn-look image, using generative AI models such as DALL-E and Stable Diffusion. The server inputs a prompt statement into the generative AI model to generate a natural, seamless composite image.

[0417] Input: Preprocessed apparel item images, model images

[0418] Output: Images of the outfit worn

[0419] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0420] Step 4:

[0421] The server generates an animated video based on the generated images of the outfits. Using animation software such as Blender or Adobe After Effects, the model's movements are defined. For example, an animation of Model A rotating can be defined, and then a generative AI model can be used to create an animated video with detailed movements added based on that movement.

[0422] Input: Image of the look you're wearing

[0423] Output: Animation video

[0424] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0425] Step 5:

[0426] The server composites the background image with the animation video, and then the user can use video editing software such as Adobe Premiere Pro to apply the background image specified by the user to the generated animation video, making the video more realistic and polished.

[0427] Input: Animation video, background image

[0428] Output: Final wear look video

[0429] Step 6:

[0430] The server saves the final generated wearable look images and animation videos in cloud storage and provides users with download links. For example, the server may save the files in an AWS S3 bucket and notify the user of the URL via email or dashboard. Users can download the images and videos by clicking the provided link.

[0431] Input: Final wear look images and animation video

[0432] Output: Provide download link to user

[0433] (Application example 1)

[0434] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0435] In conventional virtual and physical stores, consumers have limited means to visually check how apparel items look when worn, making it difficult to fully understand how the items will fit and coordinate with each other when actually worn. Furthermore, even in online shopping experiences, customers tend to feel unsure about product selection, and there is a high risk of returning items after purchase.

[0436] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0437] In this invention, the server includes means for receiving photos of apparel items, means for receiving model images, means for receiving background images, means for pre-processing the received photos of the apparel items, a generative AI model for generating worn-look images using the pre-processed photos of the apparel items and the model images, means for generating model videos based on the generated worn-look images, means for generating final worn-look images and model videos by applying a background image, means for providing the generated worn-look images and model videos to a user, interface means for executing each of the above means using a smartphone, means for applying a background image to the worn-look images and model videos to simulate displays in a physical store or a virtual store, and means for generating prompts for the generative AI model. This allows consumers to concretely and visually see what the apparel items will look like when actually worn, improving the online purchasing experience and reducing the risk of returns after purchase.

[0438] "Apparel item photos" refers to images of fashion-related products such as clothing and accessories.

[0439] "Model Image" refers to an image showing a person wearing an apparel item.

[0440] "Background image" refers to an image such as a landscape or scene that appears behind an apparel item or model image.

[0441] "Preprocessing" refers to performing processes such as adjusting the image resolution, removing noise, and cutting out the background.

[0442] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate new images and videos from input images.

[0443] "Worn look image" refers to an image in which the apparel item is naturally integrated into the model and appears to be worn.

[0444] "Model video" refers to a video showing a model moving and posing based on an image of the look worn.

[0445] "Means for providing to users" refers to a method for providing the generated images and videos in a form that allows users to access them.

[0446] "Interface means" refers to an interface that allows a user to perform various operations through a smartphone.

[0447] "Means to simulate" refers to a method for applying a background image to replicate the display in a physical or virtual store environment.

[0448] "Means for generating prompt sentences" refers to a method for automatically generating input sentences for a generative AI model.

[0449] The present invention is a system for generating a wear look image and a model video from a photograph of an apparel item, and aims to improve the shopping experience of consumers in real stores and virtual stores using a smartphone. The system of the present invention is configured as follows.

[0450] Upload a photo

[0451] Users use their smartphones to upload photos of apparel items, model images, and background images to the system, and a user interface is provided to allow for intuitive operation.

[0452] Photo pre-processing

[0453] The server preprocesses the uploaded photos of apparel items. Specifically, it uses software such as OpenCV to adjust the resolution, remove noise, and crop the background. For example, it adjusts the uploaded photo of a jacket to the appropriate resolution, crops the background, and removes noise.

[0454] Wearing look image generation

[0455] The preprocessed photos of the apparel items and the model images are input into a generative AI model to generate a worn look image. The generative AI model uses a machine learning model such as StyleGAN, which naturally integrates the apparel items into the model, resulting in an image that looks as if the item is being worn.

[0456] Model video generation

[0457] Based on the generated images of the clothing look, a video of the model is generated. The server uses the generative AI model to generate a video of the model moving and posing. For example, a video of the model spinning or walking while wearing the jacket is generated.

[0458] Applying a Background Image

[0459] A user-specified background image is applied to the wearable look images and model video. For example, a video of a model walking around wearing a jacket against a cityscape background is generated. The background image is then composited using libraries such as OpenCV.

[0460] Providing results

[0461] The server generates the final wearable look images and model videos and provides them to the user. Specifically, the server saves the generated images and videos as files and provides the user with a download link. The user can access the link to download the generated content.

[0462] Interface methods and prompt generation

[0463] Furthermore, the present invention provides an interface means that allows users to perform each operation using a smartphone. Through a smartphone application, users can intuitively upload content, and prompt sentences for the generative AI model are automatically generated. Examples of prompt sentences include:

[0464] Input image: path / to / apparel.jpg

[0465] Background image: path / to / background.jpg

[0466] Model image: path / to / model.jpg

[0467] Generated content: Images and videos of a model wearing the jacket in the street

[0468] By passing this prompt to the generative AI model, users can easily obtain high-quality images of the outfits being worn and videos of the models.

[0469] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0470] Step 1:

[0471] Users use their smartphones to upload photos of apparel items, model images, and background images to the system. This allows users to intuitively select various images through the application interface and send them to the server. The inputs are photos of apparel items, model images, and background images, which are then provided to the server.

[0472] Step 2:

[0473] The server preprocesses the received photos of apparel items using software such as OpenCV to adjust the resolution (resizing the input image to 256x256 pixels), remove noise (using non-local means), and crop the background (using edge detection and masking). The output is an image of the apparel item after preprocessing.

[0474] Step 3:

[0475] The server inputs the preprocessed apparel item photos and model images into a generative AI model. Specifically, it uses a generative AI model such as StyleGAN to generate a worn-look image in which the apparel item is naturally integrated into the model. In this process, the inputs are the preprocessed apparel item images and model images, and the output is a worn-look image.

[0476] Step 4:

[0477] The server generates a model video based on the generated wearable look images. Using a generative AI model, it generates a video of the model moving and posing. The input is a wearable look image, and the generated multiple frames are integrated to output a video.

[0478] Step 5:

[0479] The server applies a background image to the generated wearing-look images and model videos. Specifically, it uses libraries such as OpenCV to synthesize the background image with the wearing-look images and videos. The input is the wearing-look image and background image, and the output is the final wearing-look image and video with the background applied.

[0480] Step 6:

[0481] The server generates the final worn-look images and model videos and provides them to the user. Specifically, it saves the generated images and videos as files and generates a download link to provide to the user. The user can download the generated content by accessing the link. The input is the final worn-look images and videos, and the output is the download link.

[0482] Step 7:

[0483] The server generates prompt sentences and automatically generates input sentences for the generative AI model. These prompt sentences automate the processing flow of the entire system, helping users easily obtain high-quality content. For example, a prompt sentence might be generated as follows: "Input image: path / to / apparel.jpg, background image: path / to / background.jpg, model image: path / to / model.jpg, generated content: images and videos of a model wearing a jacket in the city." The inputs are an image of an apparel item, a model image, and a background image, and the output is the prompt sentence.

[0484] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0485] The present invention is a system that generates wearing look images and model videos from photographs of apparel items, and provides personalized content to each user by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system is explained below, along with specific examples.

[0486] System program processing description

[0487] Upload a photo

[0488] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0489] Applying the Emotion Engine

[0490] The server uses an emotion engine to analyze the user's facial expression data and emotion-related data received from the user's device. This allows the server to recognize the user's current emotional state. For example, if the user is smiling, it is recognized as a positive emotion.

[0491] Photo pre-processing

[0492] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out unnecessary background, and adjusts the resolution accordingly.

[0493] Emotion-Based Adjustment

[0494] The server then changes the color and design of the apparel item based on the user's recognized emotion, adjusting the generative AI model to match the user's emotion. For example, if the user is in an active mood, a bright-colored jacket might be suggested.

[0495] Wearing look image generation

[0496] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0497] Model video generation

[0498] The server creates an animation framework based on the generated wearing look images. Specifically, it sets frames that simulate the image and movement of the model as seen from all directions. This improves the reproducibility of the video.

[0499] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0500] Applying a Background Image

[0501] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0502] Providing results

[0503] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0504] Through these processes, the system of the present invention can generate personalized, high-quality wearing look images and model videos based on the user's emotions and provide them to consumers.

[0505] The processing flow will be explained below.

[0506] Step 1:

[0507] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0508] Step 2:

[0509] The server receives facial expression and voice data from the user's device and analyzes this data using an emotion engine. For example, if the user's facial expression is smiling, the server recognizes that the user has a positive emotion.

[0510] Step 3:

[0511] The server performs the necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, in a photo of a jacket, the server identifies the jacket's outline and crops out any unnecessary background.

[0512] Step 4:

[0513] The server adjusts the color and design of the apparel item based on the user's perceived emotion. For example, if the user is perceived as being in an active mood, the server suggests a brightly colored jacket.

[0514] Step 5:

[0515] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model then integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0516] Step 6:

[0517] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets the frames necessary to simulate the model's movements from all angles, improving the reproducibility of the animation.

[0518] Step 7:

[0519] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A wearing a jacket and spinning or walking in each frame.

[0520] Step 8:

[0521] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create a smooth, natural video.

[0522] Step 9:

[0523] The server analyzes the user's emotional data and selects an appropriate background image. For example, if the user is relaxing, the server will select a background of a park or natural scenery.

[0524] Step 10:

[0525] The server then synthesizes the generated look images and model video with an appropriate background image. For example, a video of model A wearing a jacket against a park landscape can be created.

[0526] Step 11:

[0527] The server saves the final look images and model video as files and generates a download link to provide to the user. The user can download the final result using the link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0528] Through these steps, the system of the present invention makes it possible to generate and provide to consumers high-quality wearing look images and model videos that are personalized based on the user's emotions.

[0529] Example 2

[0530] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0531] Conventional apparel fitting systems do not personalize the experience based on the user's emotions, and only provide standard images and videos, making it difficult to provide content tailored to the needs and emotions of individual users. Furthermore, there are technical challenges in obtaining natural, high-quality output when processing photos of apparel items and generating model videos.

[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0533] In this invention, the server includes means for receiving images of apparel items, means for receiving images of models, means for receiving images of backgrounds, means for preprocessing the received images of apparel items, a generative AI model for generating worn-look images using the preprocessed images of apparel items and images of models, means for analyzing and recognizing user emotion data, means for adjusting the color and design of the apparel items based on the user emotion data, means for generating videos of the models based on the generated images of worn looks, means for generating final images of worn looks and videos of the models by applying the background images, and means for providing the generated images of worn looks and videos of the models to the user. This enables the generation of personalized, high-quality images of worn looks and videos of the models based on the user's emotions.

[0534] "Apparel item images" are photographs or digital images of fashion items such as clothing and accessories.

[0535] "Model Image" means a photograph or digital image of a human model used to simulate trying on an apparel item.

[0536] "Background image" refers to visual materials such as photographs or illustrations that show the background onto which digital characters or apparel items are projected.

[0537] "Preprocessing" refers to processing operations such as adjusting resolution, removing noise, and removing background that are performed on image data.

[0538] A "generative AI model" is an artificial intelligence algorithm that uses deep learning techniques to generate images and videos, creating new images and videos based on specific input data.

[0539] "User emotion data" is information about the user's psychological state obtained through facial expression analysis, biometric signals, and the like.

[0540] A "wearing look image" is a still image that simulates a model wearing an apparel item.

[0541] "Model video" refers to a moving image that simulates a model moving while wearing an apparel item.

[0542] "Means of providing" refers to the process or functionality of distributing the generated digital content to users in a downloadable format.

[0543] This clearly defines the technical elements included in the claims and clarifies their scope of application.

[0544] The present invention is a system for generating images of appearances of apparel items and model videos from photographs of apparel items, and by combining this with an emotion engine that recognizes the emotions of users, the system provides personalized content to individual users. Specific embodiments of the system are described below.

[0545] The user uses a device to upload images of apparel items, background images, and model images to the server. The user opens the device's browser or a dedicated application and clicks the upload button to select and send each image file. For example, the user selects a photo of a jacket, a background image of a city, and a photo of Model A, and uploads them to the server.

[0546] The server analyzes the facial expression data and emotion-related data received from the user's device using an emotion engine. The user's facial expression data is captured by a camera and sent to the server in real time. The server uses an emotion engine (e.g., a deep learning-based facial expression recognition engine) to analyze the facial expression data and recognize the user's emotional state (e.g., smiling, surprised, etc.).

[0547] The server pre-processes the uploaded apparel item images by adjusting the resolution using an image processing library (e.g., OpenCV), applying a noise reduction algorithm, and performing image segmentation to identify the contours of the apparel items and remove unwanted background.

[0548] The server then adjusts the color and design of the apparel item based on the user's emotional data recognized by the emotion engine. The server uses generative AI models (e.g., GANs) to select bright colors (e.g., yellow or red) when the user is in an active mood.

[0549] After this preprocessing and emotion adjustment step, the server inputs the preprocessed apparel item images and the model's image into a generative AI model to generate a worn-look image. The deep learning-based generative AI model integrates the apparel item and model's body data to generate a natural and seamless worn-look image.

[0550] A model video is generated based on the generated images of the outfits. The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generated AI model adds movements (e.g., rotation, walking) to each frame, and then combines multiple frames to create a model video.

[0551] To apply the background image, the server uses an image synthesis algorithm. It integrates the background image with the look image and model video, and adjusts the color and shadows to make the background and model look natural. For example, it creates a video of Model A walking in a jacket against the backdrop of a cityscape.

[0552] Finally, the server stores the generated worn look images and model videos in cloud storage and generates a download link to provide them to the user. The user can download the final images and videos using the provided link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0553] Specific examples

[0554] Prompt Sentence Examples

[0555] "Upload a photo of the jacket. Use a cityscape as the background image and combine it with an image of Model A to generate images and videos of the suggested look."

[0556] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0557] Step 1:

[0558] Input: The user selects an apparel item image, a background image, and a model image via their device.

[0559] Specific operation: The user opens a browser or a dedicated application on their device, clicks the upload button, and selects each image file. For example, they select a photo of a jacket, a background image of a city, and an image of Model A. After selection, those files are uploaded to the server.

[0560] Output: All selected images are sent to the server.

[0561] Step 2:

[0562] Input: Facial expression data and emotion-related data sent from the user's device.

[0563] Specific operation: The user captures facial expression data in real time using the device's camera. The user's facial expression is recognized as a state such as a smile or surprise. This data is sent to the server.

[0564] Data processing / calculation: The server uses an emotion engine to analyze the received facial expression data and recognize the user's emotional state. For example, a deep learning-based facial expression recognition engine is used.

[0565] Output: User's current emotional state data.

[0566] Step 3:

[0567] Input: An image of an apparel item.

[0568] What it does: The server uses an image processing library (e.g., OpenCV) to adjust the resolution and apply noise reduction algorithms. It also performs image segmentation to identify the contours of apparel items and cut out unwanted background. For example, it automatically traces the outline of a jacket and removes the background.

[0569] Data processing / calculation: Resolution optimization, noise removal, and background removal.

[0570] Output: Preprocessed apparel item images.

[0571] Step 4:

[0572] Input: Recognized user emotional state data, preprocessed apparel item images.

[0573] Specific operation: The server adjusts the color and design of the apparel item according to the user's emotional state as recognized by the emotion engine. For example, if the user is in a lively mood, the color will be changed to a bright color (e.g., yellow or red).

[0574] Data manipulation / calculation: Use the customization engine to make color and design changes.

[0575] Output: Images of apparel items adjusted based on emotion.

[0576] Step 5:

[0577] Input: Preprocessed and adjusted apparel item images, model images.

[0578] How it works: The server uses a generative AI model (e.g., GANs) to combine images of apparel items with images of models to generate natural, seamless images of the items being worn. For example, it uses deep learning technology to generate an image of Model A wearing a jacket.

[0579] Data processing / computation: Image synthesis and generation.

[0580] Output: Generated worn look images.

[0581] Step 6:

[0582] Input: Generated worn look images.

[0583] How it works: The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generative AI model adds movements (e.g., rotating, walking) for each frame to create the video.

[0584] Data processing / computation: Video frame generation and integration.

[0585] Output: The generated model video.

[0586] Step 7:

[0587] Input: Background image, generated look images and model video.

[0588] How it works: The server uses an image synthesis algorithm to apply a background image to each frame, adjusting colors and shadows to make the background and model look natural. For example, the server creates a video of Model A walking in a jacket against a cityscape.

[0589] Data processing / calculation: background image synthesis and color adjustment.

[0590] Output: Final worn-look images and model video.

[0591] Step 8:

[0592] Input: Final wear-look images and model video.

[0593] Specific operation: The server saves the generated content to cloud storage and generates a download link to provide to the user. The user can use the provided link to download the final images and videos. For example, the user accesses the download link to obtain images and videos of Model A wearing the jacket.

[0594] Output: The content file that the user gets via the download link.

[0595] (Application example 2)

[0596] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0597] Modern online shopping sites have a problem in that users cannot actually try on apparel products when purchasing them. This can lead to problems such as anxiety when purchasing products and an increased rate of returns. Furthermore, personalized try-on experiences based on each user's emotions and preferences are rarely offered. The present invention aims to solve this problem by analyzing a user's emotions and providing a personalized virtual try-on experience based on the results.

[0598] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving photographs of apparel items, means for receiving model images, means for receiving background images, means for receiving emotion-related data, an emotion engine for analyzing the received emotion-related data, means for preprocessing the received photographs of apparel items, a generative AI model for generating worn look images adjusted based on the user's emotion using the preprocessed photographs of the apparel items and the model images, means for generating model videos based on the generated worn look images, means for generating final worn look images and model videos by applying the background image, and means for providing the generated worn look images and model videos to the user. This allows the user to enjoy a high-quality virtual try-on experience personalized to their emotions.

[0599] "Photos of apparel items" are image data of specific clothing, accessories, etc. uploaded by users.

[0600] A "model image" is a photograph of the user or a photograph of a person selected by the user, and is image data used to simulate the appearance of the user wearing an apparel item.

[0601] A "background image" is image data of a background selected by the user, and is used as the background of a wearing look image or a model video.

[0602] "Emotion-related data" is data obtained from the user's facial expressions and the like, and is used to analyze the user's current emotional state.

[0603] An "emotion engine" is a software or hardware component for analyzing emotion-related data and recognizing a user's emotional state.

[0604] "Pre-processing" refers to image processing operations such as resolution adjustment, noise removal, and background removal that are performed to prepare photographs of apparel items for use.

[0605] A "generative AI model" is an algorithm or system that uses AI technology to integrate pre-processed photographs of apparel items with model images to generate personalized worn-look images.

[0606] A "wear look image" is an image generated by a generative AI model that simulates an apparel item being worn by a model.

[0607] "Model video" is video data that uses multiple frames to recreate the appearance of a model wearing an apparel item, based on a wearing look image.

[0608] The "means for providing to the user" refers to an interface or link generation function that allows the user to view or download the generated wearing look images and model videos.

[0609] The system for implementing this invention comprises a series of processes that allow a user to have a virtual try-on experience, and hardware and software components for realizing this process.

[0610] The central part of the system is the server, which includes:

[0611] 1. Receiving photos and data:

[0612] Users use devices (such as smartphones or HMDs) to upload photos of apparel items, background images, model images, and emotion-related data to the server. The emotion-related data is necessary for analyzing the user's facial expressions.

[0613] 2. Emotion analysis:

[0614] The server then uses the emotion engine to analyze the user's emotional state based on the emotion-related data it receives. This analysis uses software such as machine learning models and emotion recognition algorithms. For example, if the user is smiling with satisfaction, their emotional state is recognized as positive.

[0615] 3. Image preprocessing:

[0616] The server preprocesses the received photos of apparel items by adjusting the resolution, removing noise, and cropping the background using image processing libraries such as OpenCV. By removing unnecessary background, the outlines of the items become clearer.

[0617] 4. Emotional customization:

[0618] The server adjusts the color and design of the apparel item based on the analyzed user's emotion. For example, if the user is in a lively mood, apparel items with bright colors and lively designs are selected.

[0619] 5. Generation of worn look images and model videos:

[0620] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. This generative AI model uses algorithms such as GAN (Generative Adversarial Networks). Based on the generated images, a model video is then generated. This process uses a technique to create a video from multiple frames.

[0621] 6. Applying a background image:

[0622] The server then combines the generated clothing look images and model video with a background image specified by the user. For example, it adjusts the background to make the model appear natural walking while wearing the apparel items against a cityscape.

[0623] 7. Providing Results:

[0624] The server finally generates a download link to provide the worn look images and model videos to the user, who can use the link to obtain the results.

[0625] Specific examples

[0626] For example, a user using an HMD provides the following prompt to the system:

[0627] Prompt statement:

[0628] The user uploaded a photo of themselves smiling with a satisfied expression and chose a beautiful beach background. The task was to crop the photo of the red dress, adjust it to the optimal resolution, and generate a video on the HMD of the model wearing the red dress and walking against the beach background.

[0629] Based on this prompt, the system executes each of the above steps to provide a virtual try-on experience tailored to the user's emotions, allowing the user to enjoy a personalized try-on experience for apparel items tailored to their emotions.

[0630] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0631] Step 1:

[0632] The user uploads a photo of the apparel item, a background image, a model image, and the user's facial expression from the device to the server. At this time, the device sends this data to the server as a single request. The input is the photo of the apparel item, the background image, the model image, and facial expression data, and the output is the state in which the data has been sent to the server.

[0633] Step 2:

[0634] The server inputs the received facial expression data into the emotion engine to analyze the user's emotional state. This emotion engine uses a machine learning model to perform real-time facial expression analysis. The input is the user's facial expression data, and the output is the analyzed emotional state (positive, negative, neutral, etc.).

[0635] Step 3:

[0636] The server preprocesses the photos of the apparel items. This preprocessing includes adjusting the image resolution, removing noise, and automatically cropping the background. Specifically, OpenCV is used to perform these operations. The input is the photos of the apparel items, and the output is the preprocessed photos.

[0637] Step 4:

[0638] The server adjusts the color and design of the preprocessed apparel item based on the emotion analysis results. For example, if the user's emotion is positive, it adjusts the color to a brighter color. The input is a photo of the preprocessed apparel item and the user's emotional state, and the output is a photo of the apparel item adjusted according to the emotion.

[0639] Step 5:

[0640] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. The generative AI model uses GAN (Generative Adversarial Networks). The input is the preprocessed apparel item photos and model images, and the output is a worn look image.

[0641] Step 6:

[0642] The server generates a model video using the generated wearing-look images. In this process, the server generates multiple frames and integrates them to construct a video. The input is the wearing-look images, and the output is the model video.

[0643] Step 7:

[0644] The server composites the generated wearing-look images and model videos with a background image specified by the user. Specifically, the background image is applied to each frame to achieve a natural look. The input is a wearing-look image, a model video, and a background image, and the output is the final wearing-look image and model video with the background composited.

[0645] Step 8:

[0646] The server generates a download link for providing the final worn-look image and model video to the user and provides the link to the user. The input is the final worn-look image and model video, and the output is the download link generated and available for the user to access.

[0647] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0648] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0649] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0650] [Third embodiment]

[0651] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0652] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0653] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0654] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0655] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0656] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0657] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0658] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0659] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0660] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0661] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0662] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0663] This invention is a system for generating wear look images and model videos from photos of apparel items, and uses AI technology to generate them using photos, background images, and model images uploaded by users. Specifically, the process is carried out in the following steps.

[0664] System program processing description

[0665] Upload a photo

[0666] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to the server. For example, a user selects a photo of a jacket and uploads a background image of a city and an image of Model A.

[0667] Photo pre-processing

[0668] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out other unnecessary background elements, and adjusts the resolution accordingly.

[0669] Wearing look image generation

[0670] The server passes the preprocessed apparel item photos and model images to a generative AI model. The generative AI model integrates the apparel item and model's body data to generate seamless, natural-looking images of the item being worn. For example, the generative AI model integrates a photo of a jacket with an image of Model A to generate an image of the model wearing the jacket.

[0671] Model video generation

[0672] The server creates an animation framework based on the generated images of the clothing look, and adds movement using a generative AI model. For example, it generates a video of model A wearing the jacket and spinning or walking.

[0673] Applying a Background Image

[0674] The server applies a user-specified background image to the generated wearing look images and model video. For example, it creates a video of model A walking around wearing the jacket against a cityscape.

[0675] Providing results

[0676] The server generates the final worn-look images and model videos, saves them as files, and generates and provides a download link to the user. The user can access the link to download the generated images and videos. For example, the user can click the provided link to obtain images and videos of Model A wearing the jacket.

[0677] Through these processes, the system of the present invention can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to consumers.

[0678] The processing flow will be explained below.

[0679] Step 1:

[0680] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0681] Step 2:

[0682] The server performs pre-processing on the uploaded apparel item photos according to their purpose. Specifically, it adjusts the image resolution to ensure appropriate image quality. It also performs noise reduction to improve image clarity. It also performs automatic background cropping to extract the outline of the apparel item. For example, the server identifies the outline of a jacket and crops out unnecessary background.

[0683] Step 3:

[0684] The server inputs the preprocessed photos of the apparel items and the model images into a generative AI model. This generative AI model integrates the apparel items with the model's posture and body data to generate natural, seamless images of how the items look when worn. For example, the generative AI model combines a photo of a jacket with an image of Model A to generate an image of Model A wearing the jacket.

[0685] Step 4:

[0686] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets frames that simulate the images and movements of the model as seen from all directions. This allows for greater reproducibility of the animation.

[0687] Step 5:

[0688] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0689] Step 6:

[0690] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create smooth, natural videos.

[0691] Step 7:

[0692] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0693] Step 8:

[0694] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0695] Through the above steps, the server can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to users.

[0696] Example 1

[0697] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0698] Conventional apparel fitting simulation systems often make it difficult for users to specifically check how they will look when wearing apparel. For example, there is a growing need to visually understand how an item will look in dynamic situations, not just still images, but systems that offer such functionality are limited. Furthermore, automated, highly accurate processing is required for image quality and background application, but current systems have difficulty achieving satisfactory results. Therefore, the challenge is to develop a system that can generate high-quality wearing look images and videos using photos of apparel items uploaded by users and provide them to users.

[0699] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0700] In this invention, the server includes means for receiving images of apparel items, means for receiving images of people, means for receiving background images, a generative AI model for generating worn look images using preprocessed apparel item images and images of people, means for generating animated videos based on the generated worn look images, means for generating final worn look images and animated videos by applying the background image, and means for providing the generated worn look images and animated videos to users, thereby enabling users to easily perform high-quality try-on simulations.

[0701] "Apparel item images" refer to digital images of fashion items such as clothing and accessories.

[0702] "Human Image" means a digital image of a model or user that is used to simulate trying on apparel items.

[0703] A "background image" is a digital image of a particular scene or location specified by the user, which is combined with an image of a person wearing an apparel item.

[0704] "Pre-processing" refers to a series of operations performed on uploaded apparel item images, including resolution adjustment, noise removal, and automatic background cropping.

[0705] "Generative AI models" are algorithms and systems that use machine learning techniques to generate new images and videos.

[0706] A "wearing look image" is a digital image created by combining images of a person wearing an apparel item.

[0707] An "animated video" is a moving digital image created by integrating a series of frames.

[0708] "User" refers to a person who uses the system to perform a try-on simulation and receives images and videos.

[0709] The present invention provides a system for generating high-quality images and animation videos of apparel items from images of the apparel items, and providing the images and animation videos to users. The system operates based on the interaction between a server, a terminal, and a user.

[0710] First, the user uploads the apparel item image, background image, and model image to the server from a terminal, such as a commonly used personal computer or smartphone.

[0711] The server then performs preprocessing on the received apparel item images. This preprocessing includes resolution adjustment, noise removal, and automatic background cropping. To perform these processes, software libraries such as OpenCV, PIL (Python Imaging Library), and Mask R-CNN are used. Specifically, OpenCV is used to resize the images for resolution adjustment, and Gaussian Blur and Median Filter are applied for noise removal. For automatic background cropping, Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0712] The preprocessed apparel item images are passed to the generative AI model by the server. This generative AI model uses algorithms such as DALL-E and Stable Diffusion to generate a wearing look image based on the preprocessed apparel item images and model images. As a concrete example, the following prompt sentence is input to the generative AI model:

[0713] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0714] After the wearable look image is generated, the server generates an animated video based on this image. Using animation software such as Blender or Adobe After Effects, the model's movements (e.g., rotation or walking) are defined, and the generative AI model adds realistic animation based on those movements. As a specific example, the following prompt sentence is input to the generative AI model:

[0715] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0716] Finally, the server applies the background image specified by the user to complete the look images and animation video. To apply the background image, video editing software such as Adobe Premiere Pro is used to composite the background onto the model images and videos wearing the apparel items.

[0717] The completed wearable look images and animation videos are stored by the server and provided to the user. Specifically, they are stored in a cloud storage service (e.g., an S3 bucket) and an access link is provided to the user. The user can download the generated images and videos by accessing the link.

[0718] As a result, the system of the present invention provides an environment in which users can visually check their apparel items in a high-quality and realistic manner.

[0719] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0720] Step 1:

[0721] The user uploads images of apparel items, background images, and model images from their device to the server. Specifically, the user selects images of clothes, accessories, etc. from their local disk using, for example, a PC or smartphone through the system interface and sends them to the server. This allows the server to receive input data for processing these images.

[0722] Input: Apparel item image, background image, model image

[0723] Output: Image data uploaded to the server

[0724] Step 2:

[0725] The server performs preprocessing on the received apparel item images. First, resolution adjustment is performed. The server uses OpenCV to resize the image resolution to an appropriate size. Next, noise reduction is performed. Techniques such as Gaussian Blur and Median Filter are used to remove noise from the image. Finally, automatic background cropping is performed. The deep learning model Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[0726] Input: Uploaded apparel item image

[0727] Output: Preprocessed apparel item images

[0728] Step 3:

[0729] The server passes the preprocessed apparel item images and model images to a generative AI model to generate a worn-look image, using generative AI models such as DALL-E and Stable Diffusion. The server inputs a prompt statement into the generative AI model to generate a natural, seamless composite image.

[0730] Input: Preprocessed apparel item images, model images

[0731] Output: Images of the outfit worn

[0732] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[0733] Step 4:

[0734] The server generates an animated video based on the generated images of the outfits. Using animation software such as Blender or Adobe After Effects, the model's movements are defined. For example, an animation of Model A rotating can be defined, and then a generative AI model can be used to create an animated video with detailed movements added based on that movement.

[0735] Input: Image of the look you're wearing

[0736] Output: Animation video

[0737] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[0738] Step 5:

[0739] The server composites the background image with the animation video, and then the user can use video editing software such as Adobe Premiere Pro to apply the background image specified by the user to the generated animation video, making the video more realistic and polished.

[0740] Input: Animation video, background image

[0741] Output: Final wear look video

[0742] Step 6:

[0743] The server saves the final generated wearable look images and animation videos in cloud storage and provides users with download links. For example, the server may save the files in an AWS S3 bucket and notify the user of the URL via email or dashboard. Users can download the images and videos by clicking the provided link.

[0744] Input: Final wear look images and animation video

[0745] Output: Provide download link to user

[0746] (Application example 1)

[0747] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0748] In conventional virtual and physical stores, consumers have limited means to visually check how apparel items look when worn, making it difficult to fully understand how the items will fit and coordinate with each other when actually worn. Furthermore, even in online shopping experiences, customers tend to feel unsure about product selection, and there is a high risk of returning items after purchase.

[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0750] In this invention, the server includes means for receiving photos of apparel items, means for receiving model images, means for receiving background images, means for pre-processing the received photos of the apparel items, a generative AI model for generating worn-look images using the pre-processed photos of the apparel items and the model images, means for generating model videos based on the generated worn-look images, means for generating final worn-look images and model videos by applying a background image, means for providing the generated worn-look images and model videos to a user, interface means for executing each of the above means using a smartphone, means for applying a background image to the worn-look images and model videos to simulate displays in a physical store or a virtual store, and means for generating prompts for the generative AI model. This allows consumers to concretely and visually see what the apparel items will look like when actually worn, improving the online purchasing experience and reducing the risk of returns after purchase.

[0751] "Apparel item photos" refers to images of fashion-related products such as clothing and accessories.

[0752] "Model Image" refers to an image showing a person wearing an apparel item.

[0753] "Background image" refers to an image such as a landscape or scene that appears behind an apparel item or model image.

[0754] "Preprocessing" refers to performing processes such as adjusting the image resolution, removing noise, and cutting out the background.

[0755] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate new images and videos from input images.

[0756] "Worn look image" refers to an image in which the apparel item is naturally integrated into the model and appears to be worn.

[0757] "Model video" refers to a video showing a model moving and posing based on an image of the look worn.

[0758] "Means for providing to users" refers to a method for providing the generated images and videos in a form that allows users to access them.

[0759] "Interface means" refers to an interface that allows a user to perform various operations through a smartphone.

[0760] "Means to simulate" refers to a method for applying a background image to replicate the display in a physical or virtual store environment.

[0761] "Means for generating prompt sentences" refers to a method for automatically generating input sentences for a generative AI model.

[0762] The present invention is a system for generating a wear look image and a model video from a photograph of an apparel item, and aims to improve the shopping experience of consumers in real stores and virtual stores using a smartphone. The system of the present invention is configured as follows.

[0763] Upload a photo

[0764] Users use their smartphones to upload photos of apparel items, model images, and background images to the system, and a user interface is provided to allow for intuitive operation.

[0765] Photo pre-processing

[0766] The server preprocesses the uploaded photos of apparel items. Specifically, it uses software such as OpenCV to adjust the resolution, remove noise, and crop the background. For example, it adjusts the uploaded photo of a jacket to the appropriate resolution, crops the background, and removes noise.

[0767] Wearing look image generation

[0768] The preprocessed photos of the apparel items and the model images are input into a generative AI model to generate a worn look image. The generative AI model uses a machine learning model such as StyleGAN, which naturally integrates the apparel items into the model, resulting in an image that looks as if the item is being worn.

[0769] Model video generation

[0770] Based on the generated images of the clothing look, a video of the model is generated. The server uses the generative AI model to generate a video of the model moving and posing. For example, a video of the model spinning or walking while wearing the jacket is generated.

[0771] Applying a Background Image

[0772] A user-specified background image is applied to the wearable look images and model video. For example, a video of a model walking around wearing a jacket against a cityscape background is generated. The background image is then composited using libraries such as OpenCV.

[0773] Providing results

[0774] The server generates the final wearable look images and model videos and provides them to the user. Specifically, the server saves the generated images and videos as files and provides the user with a download link. The user can access the link to download the generated content.

[0775] Interface methods and prompt generation

[0776] Furthermore, the present invention provides an interface means that allows users to perform each operation using a smartphone. Through a smartphone application, users can intuitively upload content, and prompt sentences for the generative AI model are automatically generated. Examples of prompt sentences include:

[0777] Input image: path / to / apparel.jpg

[0778] Background image: path / to / background.jpg

[0779] Model image: path / to / model.jpg

[0780] Generated content: Images and videos of a model wearing the jacket in the street

[0781] By passing this prompt to the generative AI model, users can easily obtain high-quality images of the outfits being worn and videos of the models.

[0782] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0783] Step 1:

[0784] Users use their smartphones to upload photos of apparel items, model images, and background images to the system. This allows users to intuitively select various images through the application interface and send them to the server. The inputs are photos of apparel items, model images, and background images, which are then provided to the server.

[0785] Step 2:

[0786] The server preprocesses the received photos of apparel items using software such as OpenCV to adjust the resolution (resizing the input image to 256x256 pixels), remove noise (using non-local means), and crop the background (using edge detection and masking). The output is an image of the apparel item after preprocessing.

[0787] Step 3:

[0788] The server inputs the preprocessed apparel item photos and model images into a generative AI model. Specifically, it uses a generative AI model such as StyleGAN to generate a worn-look image in which the apparel item is naturally integrated into the model. In this process, the inputs are the preprocessed apparel item images and model images, and the output is a worn-look image.

[0789] Step 4:

[0790] The server generates a model video based on the generated wearable look images. Using a generative AI model, it generates a video of the model moving and posing. The input is a wearable look image, and the generated multiple frames are integrated to output a video.

[0791] Step 5:

[0792] The server applies a background image to the generated wearing-look images and model videos. Specifically, it uses libraries such as OpenCV to synthesize the background image with the wearing-look images and videos. The input is the wearing-look image and background image, and the output is the final wearing-look image and video with the background applied.

[0793] Step 6:

[0794] The server generates the final worn-look images and model videos and provides them to the user. Specifically, it saves the generated images and videos as files and generates a download link to provide to the user. The user can download the generated content by accessing the link. The input is the final worn-look images and videos, and the output is the download link.

[0795] Step 7:

[0796] The server generates prompt sentences and automatically generates input sentences for the generative AI model. These prompt sentences automate the processing flow of the entire system, helping users easily obtain high-quality content. For example, a prompt sentence might be generated as follows: "Input image: path / to / apparel.jpg, background image: path / to / background.jpg, model image: path / to / model.jpg, generated content: images and videos of a model wearing a jacket in the city." The inputs are an image of an apparel item, a model image, and a background image, and the output is the prompt sentence.

[0797] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0798] The present invention is a system that generates wearing look images and model videos from photographs of apparel items, and provides personalized content to each user by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system is explained below, along with specific examples.

[0799] System program processing description

[0800] Upload a photo

[0801] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0802] Applying the Emotion Engine

[0803] The server uses an emotion engine to analyze the user's facial expression data and emotion-related data received from the user's device. This allows the server to recognize the user's current emotional state. For example, if the user is smiling, it is recognized as a positive emotion.

[0804] Photo pre-processing

[0805] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out unnecessary background, and adjusts the resolution accordingly.

[0806] Emotion-Based Adjustment

[0807] The server then changes the color and design of the apparel item based on the user's recognized emotion, adjusting the generative AI model to match the user's emotion. For example, if the user is in an active mood, a bright-colored jacket might be suggested.

[0808] Wearing look image generation

[0809] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0810] Model video generation

[0811] The server creates an animation framework based on the generated wearing look images. Specifically, it sets frames that simulate the image and movement of the model as seen from all directions. This improves the reproducibility of the video.

[0812] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[0813] Applying a Background Image

[0814] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[0815] Providing results

[0816] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0817] Through these processes, the system of the present invention can generate personalized, high-quality wearing look images and model videos based on the user's emotions and provide them to consumers.

[0818] The processing flow will be explained below.

[0819] Step 1:

[0820] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0821] Step 2:

[0822] The server receives facial expression and voice data from the user's device and analyzes this data using an emotion engine. For example, if the user's facial expression is smiling, the server recognizes that the user has a positive emotion.

[0823] Step 3:

[0824] The server performs the necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, in a photo of a jacket, the server identifies the jacket's outline and crops out any unnecessary background.

[0825] Step 4:

[0826] The server adjusts the color and design of the apparel item based on the user's perceived emotion. For example, if the user is perceived as being in an active mood, the server suggests a brightly colored jacket.

[0827] Step 5:

[0828] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model then integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[0829] Step 6:

[0830] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets the frames necessary to simulate the model's movements from all angles, improving the reproducibility of the animation.

[0831] Step 7:

[0832] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A wearing a jacket and spinning or walking in each frame.

[0833] Step 8:

[0834] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create a smooth, natural video.

[0835] Step 9:

[0836] The server analyzes the user's emotional data and selects an appropriate background image. For example, if the user is relaxing, the server will select a background of a park or natural scenery.

[0837] Step 10:

[0838] The server then synthesizes the generated look images and model video with an appropriate background image. For example, a video of model A wearing a jacket against a park landscape can be created.

[0839] Step 11:

[0840] The server saves the final look images and model video as files and generates a download link to provide to the user. The user can download the final result using the link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0841] Through these steps, the system of the present invention makes it possible to generate and provide to consumers high-quality wearing look images and model videos that are personalized based on the user's emotions.

[0842] Example 2

[0843] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0844] Conventional apparel fitting systems do not personalize the experience based on the user's emotions, and only provide standard images and videos, making it difficult to provide content tailored to the needs and emotions of individual users. Furthermore, there are technical challenges in obtaining natural, high-quality output when processing photos of apparel items and generating model videos.

[0845] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0846] In this invention, the server includes means for receiving images of apparel items, means for receiving images of models, means for receiving images of backgrounds, means for preprocessing the received images of apparel items, a generative AI model for generating worn-look images using the preprocessed images of apparel items and images of models, means for analyzing and recognizing user emotion data, means for adjusting the color and design of the apparel items based on the user emotion data, means for generating videos of the models based on the generated images of worn looks, means for generating final images of worn looks and videos of the models by applying the background images, and means for providing the generated images of worn looks and videos of the models to the user. This enables the generation of personalized, high-quality images of worn looks and videos of the models based on the user's emotions.

[0847] "Apparel item images" are photographs or digital images of fashion items such as clothing and accessories.

[0848] "Model Image" means a photograph or digital image of a human model used to simulate trying on an apparel item.

[0849] "Background image" refers to visual materials such as photographs or illustrations that show the background onto which digital characters or apparel items are projected.

[0850] "Preprocessing" refers to processing operations such as adjusting resolution, removing noise, and removing background that are performed on image data.

[0851] A "generative AI model" is an artificial intelligence algorithm that uses deep learning techniques to generate images and videos, creating new images and videos based on specific input data.

[0852] "User emotion data" is information about the user's psychological state obtained through facial expression analysis, biometric signals, and the like.

[0853] A "wearing look image" is a still image that simulates a model wearing an apparel item.

[0854] "Model video" refers to a moving image that simulates a model moving while wearing an apparel item.

[0855] "Means of providing" refers to the process or functionality of distributing the generated digital content to users in a downloadable format.

[0856] This clearly defines the technical elements included in the claims and clarifies their scope of application.

[0857] The present invention is a system for generating images of appearances of apparel items and model videos from photographs of apparel items, and by combining this with an emotion engine that recognizes the emotions of users, the system provides personalized content to individual users. Specific embodiments of the system are described below.

[0858] The user uses a device to upload images of apparel items, background images, and model images to the server. The user opens the device's browser or a dedicated application and clicks the upload button to select and send each image file. For example, the user selects a photo of a jacket, a background image of a city, and a photo of Model A, and uploads them to the server.

[0859] The server analyzes the facial expression data and emotion-related data received from the user's device using an emotion engine. The user's facial expression data is captured by a camera and sent to the server in real time. The server uses an emotion engine (e.g., a deep learning-based facial expression recognition engine) to analyze the facial expression data and recognize the user's emotional state (e.g., smiling, surprised, etc.).

[0860] The server pre-processes the uploaded apparel item images by adjusting the resolution using an image processing library (e.g., OpenCV), applying a noise reduction algorithm, and performing image segmentation to identify the contours of the apparel items and remove unwanted background.

[0861] The server then adjusts the color and design of the apparel item based on the user's emotional data recognized by the emotion engine. The server uses generative AI models (e.g., GANs) to select bright colors (e.g., yellow or red) when the user is in an active mood.

[0862] After this preprocessing and emotion adjustment step, the server inputs the preprocessed apparel item images and the model's image into a generative AI model to generate a worn-look image. The deep learning-based generative AI model integrates the apparel item and model's body data to generate a natural and seamless worn-look image.

[0863] A model video is generated based on the generated images of the outfits. The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generated AI model adds movements (e.g., rotation, walking) to each frame, and then combines multiple frames to create a model video.

[0864] To apply the background image, the server uses an image synthesis algorithm. It integrates the background image with the look image and model video, and adjusts the color and shadows to make the background and model look natural. For example, it creates a video of Model A walking in a jacket against the backdrop of a cityscape.

[0865] Finally, the server stores the generated worn look images and model videos in cloud storage and generates a download link to provide them to the user. The user can download the final images and videos using the provided link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[0866] Specific examples

[0867] Prompt Sentence Examples

[0868] "Upload a photo of the jacket. Use a cityscape as the background image and combine it with an image of Model A to generate images and videos of the suggested look."

[0869] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0870] Step 1:

[0871] Input: The user selects an apparel item image, a background image, and a model image via their device.

[0872] Specific operation: The user opens a browser or a dedicated application on their device, clicks the upload button, and selects each image file. For example, they select a photo of a jacket, a background image of a city, and an image of Model A. After selection, those files are uploaded to the server.

[0873] Output: All selected images are sent to the server.

[0874] Step 2:

[0875] Input: Facial expression data and emotion-related data sent from the user's device.

[0876] Specific operation: The user captures facial expression data in real time using the device's camera. The user's facial expression is recognized as a state such as a smile or surprise. This data is sent to the server.

[0877] Data processing / calculation: The server uses an emotion engine to analyze the received facial expression data and recognize the user's emotional state. For example, a deep learning-based facial expression recognition engine is used.

[0878] Output: User's current emotional state data.

[0879] Step 3:

[0880] Input: An image of an apparel item.

[0881] What it does: The server uses an image processing library (e.g., OpenCV) to adjust the resolution and apply noise reduction algorithms. It also performs image segmentation to identify the contours of apparel items and cut out unwanted background. For example, it automatically traces the outline of a jacket and removes the background.

[0882] Data processing / calculation: Resolution optimization, noise removal, and background removal.

[0883] Output: Preprocessed apparel item images.

[0884] Step 4:

[0885] Input: Recognized user emotional state data, preprocessed apparel item images.

[0886] Specific operation: The server adjusts the color and design of the apparel item according to the user's emotional state as recognized by the emotion engine. For example, if the user is in a lively mood, the color will be changed to a bright color (e.g., yellow or red).

[0887] Data manipulation / calculation: Use the customization engine to make color and design changes.

[0888] Output: Images of apparel items adjusted based on emotion.

[0889] Step 5:

[0890] Input: Preprocessed and adjusted apparel item images, model images.

[0891] How it works: The server uses a generative AI model (e.g., GANs) to combine images of apparel items with images of models to generate natural, seamless images of the items being worn. For example, it uses deep learning technology to generate an image of Model A wearing a jacket.

[0892] Data processing / computation: Image synthesis and generation.

[0893] Output: Generated worn look images.

[0894] Step 6:

[0895] Input: Generated worn look images.

[0896] How it works: The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generative AI model adds movements (e.g., rotating, walking) for each frame to create the video.

[0897] Data processing / computation: Video frame generation and integration.

[0898] Output: The generated model video.

[0899] Step 7:

[0900] Input: Background image, generated look images and model video.

[0901] How it works: The server uses an image synthesis algorithm to apply a background image to each frame, adjusting colors and shadows to make the background and model look natural. For example, the server creates a video of Model A walking in a jacket against a cityscape.

[0902] Data processing / calculation: background image synthesis and color adjustment.

[0903] Output: Final worn-look images and model video.

[0904] Step 8:

[0905] Input: Final wear-look images and model video.

[0906] Specific operation: The server saves the generated content to cloud storage and generates a download link to provide to the user. The user can use the provided link to download the final images and videos. For example, the user accesses the download link to obtain images and videos of Model A wearing the jacket.

[0907] Output: The content file that the user gets via the download link.

[0908] (Application example 2)

[0909] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0910] Modern online shopping sites have a problem in that users cannot actually try on apparel products when purchasing them. This can lead to problems such as anxiety when purchasing products and an increased rate of returns. Furthermore, personalized try-on experiences based on each user's emotions and preferences are rarely offered. The present invention aims to solve this problem by analyzing a user's emotions and providing a personalized virtual try-on experience based on the results.

[0911] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving photographs of apparel items, means for receiving model images, means for receiving background images, means for receiving emotion-related data, an emotion engine for analyzing the received emotion-related data, means for preprocessing the received photographs of apparel items, a generative AI model for generating worn look images adjusted based on the user's emotion using the preprocessed photographs of the apparel items and the model images, means for generating model videos based on the generated worn look images, means for generating final worn look images and model videos by applying the background image, and means for providing the generated worn look images and model videos to the user. This allows the user to enjoy a high-quality virtual try-on experience personalized to their emotions.

[0912] "Photos of apparel items" are image data of specific clothing, accessories, etc. uploaded by users.

[0913] A "model image" is a photograph of the user or a photograph of a person selected by the user, and is image data used to simulate the appearance of the user wearing an apparel item.

[0914] A "background image" is image data of a background selected by the user, and is used as the background of a wearing look image or a model video.

[0915] "Emotion-related data" is data obtained from the user's facial expressions and the like, and is used to analyze the user's current emotional state.

[0916] An "emotion engine" is a software or hardware component for analyzing emotion-related data and recognizing a user's emotional state.

[0917] "Pre-processing" refers to image processing operations such as resolution adjustment, noise removal, and background removal that are performed to prepare photographs of apparel items for use.

[0918] A "generative AI model" is an algorithm or system that uses AI technology to integrate pre-processed photographs of apparel items with model images to generate personalized worn-look images.

[0919] A "wear look image" is an image generated by a generative AI model that simulates an apparel item being worn by a model.

[0920] "Model video" is video data that uses multiple frames to recreate the appearance of a model wearing an apparel item, based on a wearing look image.

[0921] The "means for providing to the user" refers to an interface or link generation function that allows the user to view or download the generated wearing look images and model videos.

[0922] The system for implementing this invention comprises a series of processes that allow a user to have a virtual try-on experience, and hardware and software components for realizing this process.

[0923] The central part of the system is the server, which includes:

[0924] 1. Receiving photos and data:

[0925] Users use devices (such as smartphones or HMDs) to upload photos of apparel items, background images, model images, and emotion-related data to the server. The emotion-related data is necessary for analyzing the user's facial expressions.

[0926] 2. Emotion analysis:

[0927] The server then uses the emotion engine to analyze the user's emotional state based on the emotion-related data it receives. This analysis uses software such as machine learning models and emotion recognition algorithms. For example, if the user is smiling with satisfaction, their emotional state is recognized as positive.

[0928] 3. Image preprocessing:

[0929] The server preprocesses the received photos of apparel items by adjusting the resolution, removing noise, and cropping the background using image processing libraries such as OpenCV. By removing unnecessary background, the outlines of the items become clearer.

[0930] 4. Emotional customization:

[0931] The server adjusts the color and design of the apparel item based on the analyzed user's emotion. For example, if the user is in a lively mood, apparel items with bright colors and lively designs are selected.

[0932] 5. Generation of worn look images and model videos:

[0933] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. This generative AI model uses algorithms such as GAN (Generative Adversarial Networks). Based on the generated images, a model video is then generated. This process uses a technique to create a video from multiple frames.

[0934] 6. Applying a background image:

[0935] The server then combines the generated clothing look images and model video with a background image specified by the user. For example, it adjusts the background to make the model appear natural walking while wearing the apparel items against a cityscape.

[0936] 7. Providing Results:

[0937] The server finally generates a download link to provide the worn look images and model videos to the user, who can use the link to obtain the results.

[0938] Specific examples

[0939] For example, a user using an HMD provides the following prompt to the system:

[0940] Prompt statement:

[0941] The user uploaded a photo of themselves smiling with a satisfied expression and chose a beautiful beach background. The task was to crop the photo of the red dress, adjust it to the optimal resolution, and generate a video on the HMD of the model wearing the red dress and walking against the beach background.

[0942] Based on this prompt, the system executes each of the above steps to provide a virtual try-on experience tailored to the user's emotions, allowing the user to enjoy a personalized try-on experience for apparel items tailored to their emotions.

[0943] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0944] Step 1:

[0945] The user uploads a photo of the apparel item, a background image, a model image, and the user's facial expression from the device to the server. At this time, the device sends this data to the server as a single request. The input is the photo of the apparel item, the background image, the model image, and facial expression data, and the output is the state in which the data has been sent to the server.

[0946] Step 2:

[0947] The server inputs the received facial expression data into the emotion engine to analyze the user's emotional state. This emotion engine uses a machine learning model to perform real-time facial expression analysis. The input is the user's facial expression data, and the output is the analyzed emotional state (positive, negative, neutral, etc.).

[0948] Step 3:

[0949] The server preprocesses the photos of the apparel items. This preprocessing includes adjusting the image resolution, removing noise, and automatically cropping the background. Specifically, OpenCV is used to perform these operations. The input is the photos of the apparel items, and the output is the preprocessed photos.

[0950] Step 4:

[0951] The server adjusts the color and design of the preprocessed apparel item based on the emotion analysis results. For example, if the user's emotion is positive, it adjusts the color to a brighter color. The input is a photo of the preprocessed apparel item and the user's emotional state, and the output is a photo of the apparel item adjusted according to the emotion.

[0952] Step 5:

[0953] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. The generative AI model uses GAN (Generative Adversarial Networks). The input is the preprocessed apparel item photos and model images, and the output is a worn look image.

[0954] Step 6:

[0955] The server generates a model video using the generated wearing-look images. In this process, the server generates multiple frames and integrates them to construct a video. The input is the wearing-look images, and the output is the model video.

[0956] Step 7:

[0957] The server composites the generated wearing-look images and model videos with a background image specified by the user. Specifically, the background image is applied to each frame to achieve a natural look. The input is a wearing-look image, a model video, and a background image, and the output is the final wearing-look image and model video with the background composited.

[0958] Step 8:

[0959] The server generates a download link for providing the final worn-look image and model video to the user and provides the link to the user. The input is the final worn-look image and model video, and the output is the download link generated and available for the user to access.

[0960] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0961] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0962] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0963] [Fourth embodiment]

[0964] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0965] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0966] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0967] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0968] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0969] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0970] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0971] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0972] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0973] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0974] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0975] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0976] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0977] This invention is a system for generating wear look images and model videos from photos of apparel items, and uses AI technology to generate them using photos, background images, and model images uploaded by users. Specifically, the process is carried out in the following steps.

[0978] System program processing description

[0979] Upload a photo

[0980] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to the server. For example, a user selects a photo of a jacket and uploads a background image of a city and an image of Model A.

[0981] Photo pre-processing

[0982] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out other unnecessary background elements, and adjusts the resolution accordingly.

[0983] Wearing look image generation

[0984] The server passes the preprocessed apparel item photos and model images to a generative AI model. The generative AI model integrates the apparel item and model's body data to generate seamless, natural-looking images of the item being worn. For example, the generative AI model integrates a photo of a jacket with an image of Model A to generate an image of the model wearing the jacket.

[0985] Model video generation

[0986] The server creates an animation framework based on the generated images of the clothing look, and adds movement using a generative AI model. For example, it generates a video of model A wearing the jacket and spinning or walking.

[0987] Applying a Background Image

[0988] The server applies a user-specified background image to the generated wearing look images and model video. For example, it creates a video of model A walking around wearing the jacket against a cityscape.

[0989] Providing results

[0990] The server generates the final worn-look images and model videos, saves them as files, and generates and provides a download link to the user. The user can access the link to download the generated images and videos. For example, the user can click the provided link to obtain images and videos of Model A wearing the jacket.

[0991] Through these processes, the system of the present invention can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to consumers.

[0992] The processing flow will be explained below.

[0993] Step 1:

[0994] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[0995] Step 2:

[0996] The server performs pre-processing on the uploaded apparel item photos according to their purpose. Specifically, it adjusts the image resolution to ensure appropriate image quality. It also performs noise reduction to improve image clarity. It also performs automatic background cropping to extract the outline of the apparel item. For example, the server identifies the outline of a jacket and crops out unnecessary background.

[0997] Step 3:

[0998] The server inputs the preprocessed photos of the apparel items and the model images into a generative AI model. This generative AI model integrates the apparel items with the model's posture and body data to generate natural, seamless images of how the items look when worn. For example, the generative AI model combines a photo of a jacket with an image of Model A to generate an image of Model A wearing the jacket.

[0999] Step 4:

[1000] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets frames that simulate the images and movements of the model as seen from all directions. This allows for greater reproducibility of the animation.

[1001] Step 5:

[1002] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[1003] Step 6:

[1004] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create smooth, natural videos.

[1005] Step 7:

[1006] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[1007] Step 8:

[1008] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[1009] Through the above steps, the server can efficiently generate high-quality wearing look images and model videos from photographs of apparel items and provide them to users.

[1010] Example 1

[1011] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1012] Conventional apparel fitting simulation systems often make it difficult for users to specifically check how they will look when wearing apparel. For example, there is a growing need to visually understand how an item will look in dynamic situations, not just still images, but systems that offer such functionality are limited. Furthermore, automated, highly accurate processing is required for image quality and background application, but current systems have difficulty achieving satisfactory results. Therefore, the challenge is to develop a system that can generate high-quality wearing look images and videos using photos of apparel items uploaded by users and provide them to users.

[1013] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1014] In this invention, the server includes means for receiving images of apparel items, means for receiving images of people, means for receiving background images, a generative AI model for generating worn look images using preprocessed apparel item images and images of people, means for generating animated videos based on the generated worn look images, means for generating final worn look images and animated videos by applying the background image, and means for providing the generated worn look images and animated videos to users, thereby enabling users to easily perform high-quality try-on simulations.

[1015] "Apparel item images" refer to digital images of fashion items such as clothing and accessories.

[1016] "Human Image" means a digital image of a model or user that is used to simulate trying on apparel items.

[1017] A "background image" is a digital image of a particular scene or location specified by the user, which is combined with an image of a person wearing an apparel item.

[1018] "Pre-processing" refers to a series of operations performed on uploaded apparel item images, including resolution adjustment, noise removal, and automatic background cropping.

[1019] "Generative AI models" are algorithms and systems that use machine learning techniques to generate new images and videos.

[1020] A "wearing look image" is a digital image created by combining images of a person wearing an apparel item.

[1021] An "animated video" is a moving digital image created by integrating a series of frames.

[1022] "User" refers to a person who uses the system to perform a try-on simulation and receives images and videos.

[1023] The present invention provides a system for generating high-quality images and animation videos of apparel items from images of the apparel items, and providing the images and animation videos to users. The system operates based on the interaction between a server, a terminal, and a user.

[1024] First, the user uploads the apparel item image, background image, and model image to the server from a terminal, such as a commonly used personal computer or smartphone.

[1025] The server then performs preprocessing on the received apparel item images. This preprocessing includes resolution adjustment, noise removal, and automatic background cropping. To perform these processes, software libraries such as OpenCV, PIL (Python Imaging Library), and Mask R-CNN are used. Specifically, OpenCV is used to resize the images for resolution adjustment, and Gaussian Blur and Median Filter are applied for noise removal. For automatic background cropping, Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[1026] The preprocessed apparel item images are passed to the generative AI model by the server. This generative AI model uses algorithms such as DALL-E and Stable Diffusion to generate a wearing look image based on the preprocessed apparel item images and model images. As a concrete example, the following prompt sentence is input to the generative AI model:

[1027] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[1028] After the wearable look image is generated, the server generates an animated video based on this image. Using animation software such as Blender or Adobe After Effects, the model's movements (e.g., rotation or walking) are defined, and the generative AI model adds realistic animation based on those movements. As a specific example, the following prompt sentence is input to the generative AI model:

[1029] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[1030] Finally, the server applies the background image specified by the user to complete the look images and animation video. To apply the background image, video editing software such as Adobe Premiere Pro is used to composite the background onto the model images and videos wearing the apparel items.

[1031] The completed wearable look images and animation videos are stored by the server and provided to the user. Specifically, they are stored in a cloud storage service (e.g., an S3 bucket) and an access link is provided to the user. The user can download the generated images and videos by accessing the link.

[1032] As a result, the system of the present invention provides an environment in which users can visually check their apparel items in a high-quality and realistic manner.

[1033] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1034] Step 1:

[1035] The user uploads images of apparel items, background images, and model images from their device to the server. Specifically, the user selects images of clothes, accessories, etc. from their local disk using, for example, a PC or smartphone through the system interface and sends them to the server. This allows the server to receive input data for processing these images.

[1036] Input: Apparel item image, background image, model image

[1037] Output: Image data uploaded to the server

[1038] Step 2:

[1039] The server performs preprocessing on the received apparel item images. First, resolution adjustment is performed. The server uses OpenCV to resize the image resolution to an appropriate size. Next, noise reduction is performed. Techniques such as Gaussian Blur and Median Filter are used to remove noise from the image. Finally, automatic background cropping is performed. The deep learning model Mask R-CNN is used to identify the contours of the apparel items and remove unnecessary background areas.

[1040] Input: Uploaded apparel item image

[1041] Output: Preprocessed apparel item images

[1042] Step 3:

[1043] The server passes the preprocessed apparel item images and model images to a generative AI model to generate a worn-look image, using generative AI models such as DALL-E and Stable Diffusion. The server inputs a prompt statement into the generative AI model to generate a natural, seamless composite image.

[1044] Input: Preprocessed apparel item images, model images

[1045] Output: Images of the outfit worn

[1046] Prompt: "Use this jacket image (URL of jacket image) and this model image (URL of model image) to generate an image of a model wearing the jacket."

[1047] Step 4:

[1048] The server generates an animated video based on the generated images of the outfits. Using animation software such as Blender or Adobe After Effects, the model's movements are defined. For example, an animation of Model A rotating can be defined, and then a generative AI model can be used to create an animated video with detailed movements added based on that movement.

[1049] Input: Image of the look you're wearing

[1050] Output: Animation video

[1051] Prompt: "Create a video of the model walking while wearing the jacket, based on the generated model images."

[1052] Step 5:

[1053] The server composites the background image with the animation video, and then the user can use video editing software such as Adobe Premiere Pro to apply the background image specified by the user to the generated animation video, making the video more realistic and polished.

[1054] Input: Animation video, background image

[1055] Output: Final wear look video

[1056] Step 6:

[1057] The server saves the final generated wearable look images and animation videos in cloud storage and provides users with download links. For example, the server may save the files in an AWS S3 bucket and notify the user of the URL via email or dashboard. Users can download the images and videos by clicking the provided link.

[1058] Input: Final wear look images and animation video

[1059] Output: Provide download link to user

[1060] (Application example 1)

[1061] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1062] In conventional virtual and physical stores, consumers have limited means to visually check how apparel items look when worn, making it difficult to fully understand how the items will fit and coordinate with each other when actually worn. Furthermore, even in online shopping experiences, customers tend to feel unsure about product selection, and there is a high risk of returning items after purchase.

[1063] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1064] In this invention, the server includes means for receiving photos of apparel items, means for receiving model images, means for receiving background images, means for pre-processing the received photos of the apparel items, a generative AI model for generating worn-look images using the pre-processed photos of the apparel items and the model images, means for generating model videos based on the generated worn-look images, means for generating final worn-look images and model videos by applying a background image, means for providing the generated worn-look images and model videos to a user, interface means for executing each of the above means using a smartphone, means for applying a background image to the worn-look images and model videos to simulate displays in a physical store or a virtual store, and means for generating prompts for the generative AI model. This allows consumers to concretely and visually see what the apparel items will look like when actually worn, improving the online purchasing experience and reducing the risk of returns after purchase.

[1065] "Apparel item photos" refers to images of fashion-related products such as clothing and accessories.

[1066] "Model Image" refers to an image showing a person wearing an apparel item.

[1067] "Background image" refers to an image such as a landscape or scene that appears behind an apparel item or model image.

[1068] "Preprocessing" refers to performing processes such as adjusting the image resolution, removing noise, and cutting out the background.

[1069] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate new images and videos from input images.

[1070] "Worn look image" refers to an image in which the apparel item is naturally integrated into the model and appears to be worn.

[1071] "Model video" refers to a video showing a model moving and posing based on an image of the look worn.

[1072] "Means for providing to users" refers to a method for providing the generated images and videos in a form that allows users to access them.

[1073] "Interface means" refers to an interface that allows a user to perform various operations through a smartphone.

[1074] "Means to simulate" refers to a method for applying a background image to replicate the display in a physical or virtual store environment.

[1075] "Means for generating prompt sentences" refers to a method for automatically generating input sentences for a generative AI model.

[1076] The present invention is a system for generating a wear look image and a model video from a photograph of an apparel item, and aims to improve the shopping experience of consumers in real stores and virtual stores using a smartphone. The system of the present invention is configured as follows.

[1077] Upload a photo

[1078] Users use their smartphones to upload photos of apparel items, model images, and background images to the system, and a user interface is provided to allow for intuitive operation.

[1079] Photo pre-processing

[1080] The server preprocesses the uploaded photos of apparel items. Specifically, it uses software such as OpenCV to adjust the resolution, remove noise, and crop the background. For example, it adjusts the uploaded photo of a jacket to the appropriate resolution, crops the background, and removes noise.

[1081] Wearing look image generation

[1082] The preprocessed photos of the apparel items and the model images are input into a generative AI model to generate a worn look image. The generative AI model uses a machine learning model such as StyleGAN, which naturally integrates the apparel items into the model, resulting in an image that looks as if the item is being worn.

[1083] Model video generation

[1084] Based on the generated images of the clothing look, a video of the model is generated. The server uses the generative AI model to generate a video of the model moving and posing. For example, a video of the model spinning or walking while wearing the jacket is generated.

[1085] Applying a Background Image

[1086] A user-specified background image is applied to the wearable look images and model video. For example, a video of a model walking around wearing a jacket against a cityscape background is generated. The background image is then composited using libraries such as OpenCV.

[1087] Providing results

[1088] The server generates the final wearable look images and model videos and provides them to the user. Specifically, the server saves the generated images and videos as files and provides the user with a download link. The user can access the link to download the generated content.

[1089] Interface methods and prompt generation

[1090] Furthermore, the present invention provides an interface means that allows users to perform each operation using a smartphone. Through a smartphone application, users can intuitively upload content, and prompt sentences for the generative AI model are automatically generated. Examples of prompt sentences include:

[1091] Input image: path / to / apparel.jpg

[1092] Background image: path / to / background.jpg

[1093] Model image: path / to / model.jpg

[1094] Generated content: Images and videos of a model wearing the jacket in the street

[1095] By passing this prompt to the generative AI model, users can easily obtain high-quality images of the outfits being worn and videos of the models.

[1096] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1097] Step 1:

[1098] Users use their smartphones to upload photos of apparel items, model images, and background images to the system. This allows users to intuitively select various images through the application interface and send them to the server. The inputs are photos of apparel items, model images, and background images, which are then provided to the server.

[1099] Step 2:

[1100] The server preprocesses the received photos of apparel items using software such as OpenCV to adjust the resolution (resizing the input image to 256x256 pixels), remove noise (using non-local means), and crop the background (using edge detection and masking). The output is an image of the apparel item after preprocessing.

[1101] Step 3:

[1102] The server inputs the preprocessed apparel item photos and model images into a generative AI model. Specifically, it uses a generative AI model such as StyleGAN to generate a worn-look image in which the apparel item is naturally integrated into the model. In this process, the inputs are the preprocessed apparel item images and model images, and the output is a worn-look image.

[1103] Step 4:

[1104] The server generates a model video based on the generated wearable look images. Using a generative AI model, it generates a video of the model moving and posing. The input is a wearable look image, and the generated multiple frames are integrated to output a video.

[1105] Step 5:

[1106] The server applies a background image to the generated wearing-look images and model videos. Specifically, it uses libraries such as OpenCV to synthesize the background image with the wearing-look images and videos. The input is the wearing-look image and background image, and the output is the final wearing-look image and video with the background applied.

[1107] Step 6:

[1108] The server generates the final worn-look images and model videos and provides them to the user. Specifically, it saves the generated images and videos as files and generates a download link to provide to the user. The user can download the generated content by accessing the link. The input is the final worn-look images and videos, and the output is the download link.

[1109] Step 7:

[1110] The server generates prompt sentences and automatically generates input sentences for the generative AI model. These prompt sentences automate the processing flow of the entire system, helping users easily obtain high-quality content. For example, a prompt sentence might be generated as follows: "Input image: path / to / apparel.jpg, background image: path / to / background.jpg, model image: path / to / model.jpg, generated content: images and videos of a model wearing a jacket in the city." The inputs are an image of an apparel item, a model image, and a background image, and the output is the prompt sentence.

[1111] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1112] The present invention is a system that generates wearing look images and model videos from photographs of apparel items, and provides personalized content to each user by combining it with an emotion engine that recognizes the user's emotions. Specific program processing of the system is explained below, along with specific examples.

[1113] System program processing description

[1114] Upload a photo

[1115] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[1116] Applying the Emotion Engine

[1117] The server uses an emotion engine to analyze the user's facial expression data and emotion-related data received from the user's device. This allows the server to recognize the user's current emotional state. For example, if the user is smiling, it is recognized as a positive emotion.

[1118] Photo pre-processing

[1119] The server performs necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, the server identifies the outline of the jacket, crops out unnecessary background, and adjusts the resolution accordingly.

[1120] Emotion-Based Adjustment

[1121] The server then changes the color and design of the apparel item based on the user's recognized emotion, adjusting the generative AI model to match the user's emotion. For example, if the user is in an active mood, a bright-colored jacket might be suggested.

[1122] Wearing look image generation

[1123] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[1124] Model video generation

[1125] The server creates an animation framework based on the generated wearing look images. Specifically, it sets frames that simulate the image and movement of the model as seen from all directions. This improves the reproducibility of the video.

[1126] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A turning around and walking while wearing the jacket for each frame.

[1127] Applying a Background Image

[1128] The server applies the user-specified background image to the generated wearing look images and model video. Specifically, the background image is composited into each frame to achieve a natural look. For example, a video of model A walking in a jacket against a cityscape in the background can be completed.

[1129] Providing results

[1130] The server saves the final look images and model videos as files and generates a download link to provide them to users. Users can use the link to download the final results. For example, a user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[1131] Through these processes, the system of the present invention can generate personalized, high-quality wearing look images and model videos based on the user's emotions and provide them to consumers.

[1132] The processing flow will be explained below.

[1133] Step 1:

[1134] A user uploads a photo of an apparel item, a background image, and a model image from a terminal to a server. For example, a user selects and uploads a photo of a jacket, a background image of a city, and an image of Model A.

[1135] Step 2:

[1136] The server receives facial expression and voice data from the user's device and analyzes this data using an emotion engine. For example, if the user's facial expression is smiling, the server recognizes that the user has a positive emotion.

[1137] Step 3:

[1138] The server performs the necessary pre-processing on the uploaded apparel item photos, such as adjusting the image resolution, removing noise, and automatically cropping the background. For example, in a photo of a jacket, the server identifies the jacket's outline and crops out any unnecessary background.

[1139] Step 4:

[1140] The server adjusts the color and design of the apparel item based on the user's perceived emotion. For example, if the user is perceived as being in an active mood, the server suggests a brightly colored jacket.

[1141] Step 5:

[1142] The server inputs the preprocessed apparel item photos and model images into a generative AI model. The generative AI model then integrates the apparel item and model's body data to generate natural, seamless images of the garment being worn. For example, the generative AI model combines a photo of a jacket with an image of model A to generate an image of model A wearing the jacket.

[1143] Step 6:

[1144] The server creates an animation framework based on the images of the outfits worn by the model. Specifically, it sets the frames necessary to simulate the model's movements from all angles, improving the reproducibility of the animation.

[1145] Step 7:

[1146] The server uses the generative AI model to add movement to each animation frame. For example, the generative AI model simulates Model A wearing a jacket and spinning or walking in each frame.

[1147] Step 8:

[1148] The server then integrates the generated frames to generate a series of model videos, ensuring continuity between frames to create a smooth, natural video.

[1149] Step 9:

[1150] The server analyzes the user's emotional data and selects an appropriate background image. For example, if the user is relaxing, the server will select a background of a park or natural scenery.

[1151] Step 10:

[1152] The server then synthesizes the generated look images and model video with an appropriate background image. For example, a video of model A wearing a jacket against a park landscape can be created.

[1153] Step 11:

[1154] The server saves the final look images and model video as files and generates a download link to provide to the user. The user can download the final result using the link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[1155] Through these steps, the system of the present invention makes it possible to generate and provide to consumers high-quality wearing look images and model videos that are personalized based on the user's emotions.

[1156] Example 2

[1157] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1158] Conventional apparel fitting systems do not personalize the experience based on the user's emotions, and only provide standard images and videos, making it difficult to provide content tailored to the needs and emotions of individual users. Furthermore, there are technical challenges in obtaining natural, high-quality output when processing photos of apparel items and generating model videos.

[1159] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1160] In this invention, the server includes means for receiving images of apparel items, means for receiving images of models, means for receiving images of backgrounds, means for preprocessing the received images of apparel items, a generative AI model for generating worn-look images using the preprocessed images of apparel items and images of models, means for analyzing and recognizing user emotion data, means for adjusting the color and design of the apparel items based on the user emotion data, means for generating videos of the models based on the generated images of worn looks, means for generating final images of worn looks and videos of the models by applying the background images, and means for providing the generated images of worn looks and videos of the models to the user. This enables the generation of personalized, high-quality images of worn looks and videos of the models based on the user's emotions.

[1161] "Apparel item images" are photographs or digital images of fashion items such as clothing and accessories.

[1162] "Model Image" means a photograph or digital image of a human model used to simulate trying on an apparel item.

[1163] "Background image" refers to visual materials such as photographs or illustrations that show the background onto which digital characters or apparel items are projected.

[1164] "Preprocessing" refers to processing operations such as adjusting resolution, removing noise, and removing background that are performed on image data.

[1165] A "generative AI model" is an artificial intelligence algorithm that uses deep learning techniques to generate images and videos, creating new images and videos based on specific input data.

[1166] "User emotion data" is information about the user's psychological state obtained through facial expression analysis, biometric signals, and the like.

[1167] A "wearing look image" is a still image that simulates a model wearing an apparel item.

[1168] "Model video" refers to a moving image that simulates a model moving while wearing an apparel item.

[1169] "Means of providing" refers to the process or functionality of distributing the generated digital content to users in a downloadable format.

[1170] This clearly defines the technical elements included in the claims and clarifies their scope of application.

[1171] The present invention is a system for generating images of appearances of apparel items and model videos from photographs of apparel items, and by combining this with an emotion engine that recognizes the emotions of users, the system provides personalized content to individual users. Specific embodiments of the system are described below.

[1172] The user uses a device to upload images of apparel items, background images, and model images to the server. The user opens the device's browser or a dedicated application and clicks the upload button to select and send each image file. For example, the user selects a photo of a jacket, a background image of a city, and a photo of Model A, and uploads them to the server.

[1173] The server analyzes the facial expression data and emotion-related data received from the user's device using an emotion engine. The user's facial expression data is captured by a camera and sent to the server in real time. The server uses an emotion engine (e.g., a deep learning-based facial expression recognition engine) to analyze the facial expression data and recognize the user's emotional state (e.g., smiling, surprised, etc.).

[1174] The server pre-processes the uploaded apparel item images by adjusting the resolution using an image processing library (e.g., OpenCV), applying a noise reduction algorithm, and performing image segmentation to identify the contours of the apparel items and remove unwanted background.

[1175] The server then adjusts the color and design of the apparel item based on the user's emotional data recognized by the emotion engine. The server uses generative AI models (e.g., GANs) to select bright colors (e.g., yellow or red) when the user is in an active mood.

[1176] After this preprocessing and emotion adjustment step, the server inputs the preprocessed apparel item images and the model's image into a generative AI model to generate a worn-look image. The deep learning-based generative AI model integrates the apparel item and model's body data to generate a natural and seamless worn-look image.

[1177] A model video is generated based on the generated images of the outfits. The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generated AI model adds movements (e.g., rotation, walking) to each frame, and then combines multiple frames to create a model video.

[1178] To apply the background image, the server uses an image synthesis algorithm. It integrates the background image with the look image and model video, and adjusts the color and shadows to make the background and model look natural. For example, it creates a video of Model A walking in a jacket against the backdrop of a cityscape.

[1179] Finally, the server stores the generated worn look images and model videos in cloud storage and generates a download link to provide them to the user. The user can download the final images and videos using the provided link. For example, the user accesses the provided link to obtain images and videos of Model A wearing the jacket.

[1180] Specific examples

[1181] Prompt Sentence Examples

[1182] "Upload a photo of the jacket. Use a cityscape as the background image and combine it with an image of Model A to generate images and videos of the suggested look."

[1183] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1184] Step 1:

[1185] Input: The user selects an apparel item image, a background image, and a model image via their device.

[1186] Specific operation: The user opens a browser or a dedicated application on their device, clicks the upload button, and selects each image file. For example, they select a photo of a jacket, a background image of a city, and an image of Model A. After selection, those files are uploaded to the server.

[1187] Output: All selected images are sent to the server.

[1188] Step 2:

[1189] Input: Facial expression data and emotion-related data sent from the user's device.

[1190] Specific operation: The user captures facial expression data in real time using the device's camera. The user's facial expression is recognized as a state such as a smile or surprise. This data is sent to the server.

[1191] Data processing / calculation: The server uses an emotion engine to analyze the received facial expression data and recognize the user's emotional state. For example, a deep learning-based facial expression recognition engine is used.

[1192] Output: User's current emotional state data.

[1193] Step 3:

[1194] Input: An image of an apparel item.

[1195] What it does: The server uses an image processing library (e.g., OpenCV) to adjust the resolution and apply noise reduction algorithms. It also performs image segmentation to identify the contours of apparel items and cut out unwanted background. For example, it automatically traces the outline of a jacket and removes the background.

[1196] Data processing / calculation: Resolution optimization, noise removal, and background removal.

[1197] Output: Preprocessed apparel item images.

[1198] Step 4:

[1199] Input: Recognized user emotional state data, preprocessed apparel item images.

[1200] Specific operation: The server adjusts the color and design of the apparel item according to the user's emotional state as recognized by the emotion engine. For example, if the user is in a lively mood, the color will be changed to a bright color (e.g., yellow or red).

[1201] Data manipulation / calculation: Use the customization engine to make color and design changes.

[1202] Output: Images of apparel items adjusted based on emotion.

[1203] Step 5:

[1204] Input: Preprocessed and adjusted apparel item images, model images.

[1205] How it works: The server uses a generative AI model (e.g., GANs) to combine images of apparel items with images of models to generate natural, seamless images of the items being worn. For example, it uses deep learning technology to generate an image of Model A wearing a jacket.

[1206] Data processing / computation: Image synthesis and generation.

[1207] Output: Generated worn look images.

[1208] Step 6:

[1209] Input: Generated worn look images.

[1210] How it works: The server uses 3D animation software (e.g., Blender) to simulate the model's movements from all directions. The generative AI model adds movements (e.g., rotating, walking) for each frame to create the video.

[1211] Data processing / computation: Video frame generation and integration.

[1212] Output: The generated model video.

[1213] Step 7:

[1214] Input: Background image, generated look images and model video.

[1215] How it works: The server uses an image synthesis algorithm to apply a background image to each frame, adjusting colors and shadows to make the background and model look natural. For example, the server creates a video of Model A walking in a jacket against a cityscape.

[1216] Data processing / calculation: background image synthesis and color adjustment.

[1217] Output: Final worn-look images and model video.

[1218] Step 8:

[1219] Input: Final wear-look images and model video.

[1220] Specific operation: The server saves the generated content to cloud storage and generates a download link to provide to the user. The user can use the provided link to download the final images and videos. For example, the user accesses the download link to obtain images and videos of Model A wearing the jacket.

[1221] Output: The content file that the user gets via the download link.

[1222] (Application example 2)

[1223] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1224] Modern online shopping sites have a problem in that users cannot actually try on apparel products when purchasing them. This can lead to problems such as anxiety when purchasing products and an increased rate of returns. Furthermore, personalized try-on experiences based on each user's emotions and preferences are rarely offered. The present invention aims to solve this problem by analyzing a user's emotions and providing a personalized virtual try-on experience based on the results.

[1225] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving photographs of apparel items, means for receiving model images, means for receiving background images, means for receiving emotion-related data, an emotion engine for analyzing the received emotion-related data, means for preprocessing the received photographs of apparel items, a generative AI model for generating worn look images adjusted based on the user's emotion using the preprocessed photographs of the apparel items and the model images, means for generating model videos based on the generated worn look images, means for generating final worn look images and model videos by applying the background image, and means for providing the generated worn look images and model videos to the user. This allows the user to enjoy a high-quality virtual try-on experience personalized to their emotions.

[1226] "Photos of apparel items" are image data of specific clothing, accessories, etc. uploaded by users.

[1227] A "model image" is a photograph of the user or a photograph of a person selected by the user, and is image data used to simulate the appearance of the user wearing an apparel item.

[1228] A "background image" is image data of a background selected by the user, and is used as the background of a wearing look image or a model video.

[1229] "Emotion-related data" is data obtained from the user's facial expressions and the like, and is used to analyze the user's current emotional state.

[1230] An "emotion engine" is a software or hardware component for analyzing emotion-related data and recognizing a user's emotional state.

[1231] "Pre-processing" refers to image processing operations such as resolution adjustment, noise removal, and background removal that are performed to prepare photographs of apparel items for use.

[1232] A "generative AI model" is an algorithm or system that uses AI technology to integrate pre-processed photographs of apparel items with model images to generate personalized worn-look images.

[1233] A "wear look image" is an image generated by a generative AI model that simulates an apparel item being worn by a model.

[1234] "Model video" is video data that uses multiple frames to recreate the appearance of a model wearing an apparel item, based on a wearing look image.

[1235] The "means for providing to the user" refers to an interface or link generation function that allows the user to view or download the generated wearing look images and model videos.

[1236] The system for implementing this invention comprises a series of processes that allow a user to have a virtual try-on experience, and hardware and software components for realizing this process.

[1237] The central part of the system is the server, which includes:

[1238] 1. Receiving photos and data:

[1239] Users use devices (such as smartphones or HMDs) to upload photos of apparel items, background images, model images, and emotion-related data to the server. The emotion-related data is necessary for analyzing the user's facial expressions.

[1240] 2. Emotion analysis:

[1241] The server then uses the emotion engine to analyze the user's emotional state based on the emotion-related data it receives. This analysis uses software such as machine learning models and emotion recognition algorithms. For example, if the user is smiling with satisfaction, their emotional state is recognized as positive.

[1242] 3. Image preprocessing:

[1243] The server preprocesses the received photos of apparel items by adjusting the resolution, removing noise, and cropping the background using image processing libraries such as OpenCV. By removing unnecessary background, the outlines of the items become clearer.

[1244] 4. Emotional customization:

[1245] The server adjusts the color and design of the apparel item based on the analyzed user's emotion. For example, if the user is in a lively mood, apparel items with bright colors and lively designs are selected.

[1246] 5. Generation of worn look images and model videos:

[1247] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. This generative AI model uses algorithms such as GAN (Generative Adversarial Networks). Based on the generated images, a model video is then generated. This process uses a technique to create a video from multiple frames.

[1248] 6. Applying a background image:

[1249] The server then combines the generated clothing look images and model video with a background image specified by the user. For example, it adjusts the background to make the model appear natural walking while wearing the apparel items against a cityscape.

[1250] 7. Providing Results:

[1251] The server finally generates a download link to provide the worn look images and model videos to the user, who can use the link to obtain the results.

[1252] Specific examples

[1253] For example, a user using an HMD provides the following prompt to the system:

[1254] Prompt statement:

[1255] The user uploaded a photo of themselves smiling with a satisfied expression and chose a beautiful beach background. The task was to crop the photo of the red dress, adjust it to the optimal resolution, and generate a video on the HMD of the model wearing the red dress and walking against the beach background.

[1256] Based on this prompt, the system executes each of the above steps to provide a virtual try-on experience tailored to the user's emotions, allowing the user to enjoy a personalized try-on experience for apparel items tailored to their emotions.

[1257] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1258] Step 1:

[1259] The user uploads a photo of the apparel item, a background image, a model image, and the user's facial expression from the device to the server. At this time, the device sends this data to the server as a single request. The input is the photo of the apparel item, the background image, the model image, and facial expression data, and the output is the state in which the data has been sent to the server.

[1260] Step 2:

[1261] The server inputs the received facial expression data into the emotion engine to analyze the user's emotional state. This emotion engine uses a machine learning model to perform real-time facial expression analysis. The input is the user's facial expression data, and the output is the analyzed emotional state (positive, negative, neutral, etc.).

[1262] Step 3:

[1263] The server preprocesses the photos of the apparel items. This preprocessing includes adjusting the image resolution, removing noise, and automatically cropping the background. Specifically, OpenCV is used to perform these operations. The input is the photos of the apparel items, and the output is the preprocessed photos.

[1264] Step 4:

[1265] The server adjusts the color and design of the preprocessed apparel item based on the emotion analysis results. For example, if the user's emotion is positive, it adjusts the color to a brighter color. The input is a photo of the preprocessed apparel item and the user's emotional state, and the output is a photo of the apparel item adjusted according to the emotion.

[1266] Step 5:

[1267] The server inputs the preprocessed apparel item photos and model images into a generative AI model to generate a worn look image. The generative AI model uses GAN (Generative Adversarial Networks). The input is the preprocessed apparel item photos and model images, and the output is a worn look image.

[1268] Step 6:

[1269] The server generates a model video using the generated wearing-look images. In this process, the server generates multiple frames and integrates them to construct a video. The input is the wearing-look images, and the output is the model video.

[1270] Step 7:

[1271] The server composites the generated wearing-look images and model videos with a background image specified by the user. Specifically, the background image is applied to each frame to achieve a natural look. The input is a wearing-look image, a model video, and a background image, and the output is the final wearing-look image and model video with the background composited.

[1272] Step 8:

[1273] The server generates a download link for providing the final worn-look image and model video to the user and provides the link to the user. The input is the final worn-look image and model video, and the output is the download link generated and available for the user to access.

[1274] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1275] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1276] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1277] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1278] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1279] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1280] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1281] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1282] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1283] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1284] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1285] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1286] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1287] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1288] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1289] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1290] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1291] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1292] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1293] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1294] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1295] The following is further disclosed regarding the above embodiment.

[1296] (Claim 1)

[1297] a means for receiving photographs of the apparel items;

[1298] a means for receiving a model image;

[1299] means for receiving a background image;

[1300] means for pre-processing the photographs of the received apparel items;

[1301] a generative AI model for generating a wearing look image using the preprocessed apparel item photograph and a model image;

[1302] A means for generating a model video based on the generated wearing look image;

[1303] means for applying a background image to generate a final worn look image and model video;

[1304] The system includes a means for providing the generated worn look images and model videos to a user.

[1305] (Claim 2)

[1306] 10. The system of claim 1, further comprising means for adjusting the resolution of the preprocessed photograph of the apparel item, means for removing noise, and means for removing background from the image.

[1307] (Claim 3)

[1308] The system of claim 1, further comprising means for generating multiple frames using a generative AI model and integrating the frames to construct a model video.

[1309] "Example 1"

[1310] (Claim 1)

[1311] a means for receiving an image of the apparel item;

[1312] a means for receiving an image of a person;

[1313] a means for receiving a background image;

[1314] means for pre-processing the received images of the apparel items;

[1315] A generative AI model for generating a wearing look image using preprocessed apparel item images and a person image;

[1316] A means for generating an animation video based on the generated wearing look image;

[1317] means for applying a background image to generate a final worn look image and animation video;

[1318] The system includes a means for providing the generated worn look images and animation videos to a user.

[1319] (Claim 2)

[1320] 10. The system of claim 1, further comprising: means for adjusting the resolution of the preprocessed apparel item image; means for removing noise; and means for removing background from the image.

[1321] (Claim 3)

[1322] The system of claim 1, further comprising means for generating multiple frames using a generative AI model and integrating the frames to create an animated video.

[1323] "Application Example 1"

[1324] (Claim 1)

[1325] a means for receiving photographs of the apparel items;

[1326] a means for receiving a model image;

[1327] means for receiving a background image;

[1328] means for pre-processing the photographs of the received apparel items;

[1329] a generative AI model for generating a wearing look image using the preprocessed apparel item photograph and a model image;

[1330] A means for generating a model video based on the generated wearing look image;

[1331] means for applying a background image to generate a final worn look image and model video;

[1332] A means for providing the generated wearing look images and model videos to a user;

[1333] an interface means for executing each of the means using a smartphone;

[1334] A means for applying background images to the images of the outfits and the model videos to simulate displays in physical stores and virtual stores;

[1335] A system including means for generating a prompt sentence for a generative AI model.

[1336] (Claim 2)

[1337] 10. The system of claim 1, further comprising: means for adjusting the resolution of the preprocessed photograph of the apparel item; means for removing noise; means for removing background from the image; and means for integrating the preprocessed model image.

[1338] (Claim 3)

[1339] The system of claim 1 includes a means for generating multiple frames using a generative AI model and integrating the frames to create a model video, and a means for displaying the generated model video on a smartphone.

[1340] "Example 2: Combining Emotion Engines"

[1341] (Claim 1)

[1342] a means for receiving an image of the apparel item;

[1343] a means for receiving images of the model;

[1344] a means for receiving a background image;

[1345] means for pre-processing the received images of the apparel items;

[1346] a generative AI model for generating a wearing look image using the preprocessed apparel item image and the model image;

[1347] means for analyzing and recognizing user emotion data;

[1348] means for adjusting the color and design of the apparel item based on the user's emotional data;

[1349] A means for generating a video of a model based on the generated wearing look image;

[1350] means for applying background images to generate final worn look images and model videos;

[1351] The system includes a means for providing the generated worn-look images and model videos to a user.

[1352] (Claim 2)

[1353] 10. The system of claim 1, further comprising: means for adjusting the resolution of the preprocessed apparel item image; means for removing noise; and means for removing background from the image.

[1354] (Claim 3)

[1355] 2. The system of claim 1, further comprising means for generating a plurality of frames using a generative AI model and integrating the frames to construct a video of the model.

[1356] "Application example 2 when combining emotion engines"

[1357] (Claim 1)

[1358] a means for receiving photographs of the apparel items;

[1359] a means for receiving a model image;

[1360] means for receiving a background image;

[1361] means for receiving emotion-related data;

[1362] an emotion engine that analyzes the received emotion-related data;

[1363] means for pre-processing the photographs of the received apparel items;

[1364] a generative AI model for generating a wearing look image adjusted based on a user's emotion using the preprocessed apparel item photos and model images;

[1365] A means for generating a model video based on the generated wearing look image;

[1366] means for applying a background image to generate a final worn look image and model video;

[1367] The system includes a means for providing the generated worn look images and model videos to a user.

[1368] (Claim 2)

[1369] 10. The system of claim 1, further comprising: means for adjusting the resolution of the preprocessed photograph of the apparel item; means for removing noise; means for removing background from the image; and means for changing the color or design of the apparel item based on a user's emotion.

[1370] (Claim 3)

[1371] The system of claim 1 further comprising means for generating multiple frames using a generative AI model, integrating the frames to construct a model video, and synthesizing a background image. [Explanation of symbols]

[1372] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for receiving photographs of the apparel items; a means for receiving a model image; means for receiving a background image; means for pre-processing the photographs of the received apparel items; a generative AI model for generating a wearing look image using the preprocessed apparel item photograph and a model image; A means for generating a model video based on the generated wearing look image; means for applying a background image to generate a final worn look image and model video; The system includes a means for providing the generated worn look images and model videos to a user.

2. 10. The system of claim 1, further comprising means for adjusting the resolution of the pre-processed photograph of the apparel item, means for removing noise, and means for removing background from the image.

3. The system of claim 1, further comprising means for generating a plurality of frames using a generative AI model and integrating the frames to construct a model video.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A