System
A system that generates a 3D model based on user input for virtual try-on simulations addresses the uncertainty of online clothing purchases, enhancing the shopping experience by providing fit and appearance feedback.
Patent Information
- Application Number
- JP2024119079
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Users often feel unsure about the size, fit, or appearance of clothing when purchasing online, leading to dissatisfaction and returns due to the inability to physically try on products.
A system that receives personal information and images from users, generates a 3D model, performs a try-on simulation, and provides feedback on fit and appearance, allowing virtual try-on before purchase.
Enables users to check the fit and appearance of clothing virtually, reducing anxiety and stress associated with online shopping and minimizing returns.
Smart Images

Figure 2026018018000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When purchasing clothing on conventional e-commerce sites, users often felt unsure about their product selection because they could not actually touch and try on the product. A major issue was not knowing the size, fit, or actual appearance of the product. As a result, users often found that the product they purchased did not meet their expectations, leading to returns. The present invention aims to provide a system that solves these problems. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system that includes: means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user; means for processing the received personal information and image to generate a 3D model of the user; means for running a try-on simulation based on clothing data and the generated 3D model; and means for providing the user with the results of the try-on simulation and their impressions. Using this system, users can check the fit and appearance of clothes before purchasing without actually trying them on.
[0006] "User" means an individual who uses the System to try on or purchase a Product.
[0007] "Personal information" refers to information including physical characteristics necessary for the try-on simulation, such as the user's height, weight, body type, hairstyle, etc.
[0008] "Image" refers to visual data necessary for performing a try-on simulation, such as a full-body photo of the user.
[0009] "Receiving" refers to the system obtaining data sent by a user.
[0010] "Processing" refers to analyzing the received data and converting it into a format suitable for the try-on simulation.
[0011] "3D model" refers to digital data that is reproduced in three dimensions based on the user's physical characteristics.
[0012] "Clothing data" refers to information such as the texture, color, and design of the fabric of the clothing that is the subject of the try-on simulation.
[0013] "Try-on simulation" refers to applying clothing data to a 3D model of the user to digitally recreate the look and fit of the clothing when tried on.
[0014] "Results" refers to the images and feedback generated by the try-on simulation.
[0015] "Opinions" refers to fit, appearance, and other evaluation information that is automatically generated based on the results of a try-on simulation. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system performs a try-on simulation based on personal information such as height, weight, figure, hairstyle, etc. provided by the user, as well as an image, and provides the results.
[0038] System Overview
[0039] The system mainly includes four main components:
[0040] 1. Data entry phase
[0041] 2. Data analysis phase
[0042] 3. Feedback Phase
[0043] 4. Purchasing Phase (Optional)
[0044] The specific processing of the program for each phase will be explained below.
[0045] Data Entry Phase
[0046] 1. User Actions
[0047] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[0048] 2. Processing performed by the server
[0049] The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. It also performs preprocessing on image data, such as cropping and resizing.
[0050] Data analysis phase
[0051] 1. Processing performed by the server: Image preprocessing
[0052] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[0053] 2. Processing performed by the server: 3D model generation
[0054] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[0055] 3. Processing performed by the server: Acquiring clothing data
[0056] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[0057] 4. Processing performed by the server: Try-on simulation
[0058] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0059] Feedback Phase
[0060] 1. Server processing: Try-on images and impression generation
[0061] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[0062] 2. Server action: Sending feedback
[0063] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[0064] 3. User Action: Feedback Check
[0065] The user can then check the try-on images and their impressions and decide whether to purchase. If they like the item, they can proceed with the purchase.
[0066] Specific examples
[0067] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, integrates it with the data for the "blue T-shirt" that the user selected to try on, and performs a try-on simulation. Finally, the generated try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user.
[0068] In this way, users can choose products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress that comes with purchasing clothing on an e-commerce site.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] User process: The user accesses an e-commerce site or a dedicated app from their own device. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[0072] Step 2:
[0073] Processing performed by the device: The entered personal information and image are sent to the server. By pressing the send button, the data is sent to the specified API endpoint.
[0074] Step 3:
[0075] Processing by the server: The server converts the received personal information and images into an appropriate format and stores them in a database. The images are pre-processed for analysis.
[0076] Step 4:
[0077] Server processing: Preprocesses the uploaded image by removing the background and extracting only the user's body. At this stage, the image is resized and a noise reduction filter is applied.
[0078] Step 5:
[0079] Server processing: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. This information is used to run a 3D modeling algorithm to create a three-dimensional digital model of the user.
[0080] Step 6:
[0081] Processing performed by the server: Retrieves data on the clothing item selected by the user from the database, including information on the clothing item's fabric texture, design, color, etc.
[0082] Step 7:
[0083] Server processing: The acquired clothing data is applied to the user's 3D model, and a fitting simulation is performed. A physics engine is used to simulate the fit, drape, wrinkles, etc. of the clothing.
[0084] Step 8:
[0085] Server processing: Analyzes visual information and automatically generates impressions along with the post-try-on images generated as a result of the simulation, creating text feedback on fit and appearance.
[0086] Step 9:
[0087] Server process: The generated try-on images and feedback are sent to the user's device. The data is formatted for the user interface and returned as an API response.
[0088] Step 10:
[0089] User action: Check the submitted try-on images and comments. Display the feedback on the e-commerce site or in the dedicated app and decide whether to purchase the product.
[0090] Step 11 (Optional):
[0091] What the user does: If they like the item, they press the "Purchase" button to complete the purchase. They can also try on other items if they wish.
[0092] This process allows users to simulate the fit and appearance of clothing through virtual try-on without actually trying it on. This system is expected to significantly reduce the anxiety and stress felt when purchasing clothing on an e-commerce site.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] Online shopping, especially when purchasing clothing, can cause anxiety and stress when users choose products without actually trying them on. Specifically, it is difficult for users to check in advance how the size, fit, and visual appearance of the clothing they choose will actually look like. Another issue is the time and effort required to visit a physical store to try on items.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a three-dimensional model of the user, means for running a try-on simulation based on clothing data and the generated three-dimensional model, means for providing the user with the results of the try-on simulation and feedback, means for sending the generated feedback to the user's device, and means for the user to check the feedback on their own device. This allows users to select products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress of online shopping.
[0098] "User" refers to an end user who uses the system to provide their personal information and images and simulate trying on clothing items.
[0099] "Personal information" refers to attribute information such as height, weight, body type, hairstyle, etc. provided by the user.
[0100] "Images" refers to photographs or videos of a user's entire body that they upload to the system.
[0101] "3D model" refers to a 3D virtual human model generated based on a user's personal information and image.
[0102] "Clothing data" refers to data including information such as the texture, color, and design of the clothing fabric, which is necessary for performing a fitting simulation.
[0103] "Try-on simulation" refers to the process of applying clothing data to a three-dimensional model of the user and virtually recreating the fit and appearance of the clothing using a physics engine or other means.
[0104] "Opinions" refer to evaluation comments about fit and appearance generated based on the results of the try-on simulation.
[0105] "Terminal" refers to an electronic device such as a smartphone or computer that a user uses to access the system.
[0106] "Feedback" refers to providing information to the user, including the results of the try-on simulation and the generated impressions.
[0107] "Generative AI model" refers to an artificial intelligence algorithm or program that automatically generates impressions about fit and appearance based on the results of a fitting simulation.
[0108] MODE FOR CARRYING OUT THE INVENTION
[0109] The present invention provides a system that allows users to virtually try on clothing when purchasing it online, and check the fit and appearance of the clothing after trying it on. This system is primarily built using the following hardware and software:
[0110] Hardware used
[0111] 1. Server: A high-performance server capable of processing and storing large amounts of data.
[0112] 2. Device: A device that can connect to the internet, such as a smartphone, tablet, or PC.
[0113] Software used
[0114] 1. Database: Database software (e.g., MySQL, PostgreSQL) for storing personal information, 3D models, and clothing data.
[0115] 2. Image processing software: Software for cropping, resizing, and noise removal of images (e.g., OpenCV).
[0116] 3. 3D modeling algorithms: Software used to generate a 3D model of the user (e.g. Blender API).
[0117] 4. Try-on simulation software: Software that applies clothing data to a three-dimensional model to simulate trying on the garment (e.g., Unity).
[0118] 5. Generative AI model: An AI model (e.g., a GPT-based model) that automatically generates impressions on fit and appearance based on the results of a try-on simulation.
[0119] System overview and specific processing
[0120] Data Entry Phase
[0121] Users access a dedicated app or e-commerce site using their own device, enter personal information such as their height, weight, body type, and hairstyle, along with a full-body image, and submit it. The server converts the received data into an appropriate format and stores it in a database. The image data undergoes pre-processing such as cropping, resizing, and noise removal.
[0122] Data analysis phase
[0123] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It also resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information to generate a 3D model, which also incorporates the preprocessed image information. It then retrieves data for the selected clothing item from a database and converts it into a format optimized for simulation.
[0124] Try-on simulation phase
[0125] The server applies the clothing data to the generated 3D model of the user and performs a fitting simulation using a physics engine, thereby reproducing the fit and visual appearance of the clothing and generating the final after-fit image.
[0126] Feedback Phase
[0127] The server uses a generative AI model to automatically generate feedback on the fit and appearance of the clothing based on the generated try-on images. The try-on images and feedback are then sent to the user's device, where the user can review the feedback and decide whether to purchase the clothing.
[0128] Specific examples
[0129] For example, if a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair, they can take a full-body photo with their smartphone and upload it to the system. The server generates a three-dimensional model of the user based on the received information, obtains data on the "blue T-shirt" selected for trying on, and performs a try-on simulation. Finally, the generated try-on image and a user's feedback, such as "It fits perfectly and the color is nice," are sent to the user's device.
[0130] Prompt Sentence Examples
[0131] "The user is 170cm tall, weighs 65kg, has a slim build, and has short hair. Please simulate trying on the 'blue T-shirt' selected by this user and generate impressions about the fit and appearance."
[0132] By building a system like this, users can check the fit and appearance of clothing through virtual try-ons, allowing them to enjoy online shopping with peace of mind.
[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0134] Step 1:
[0135] User Action: Data Entry
[0136] Users access a dedicated app or e-commerce site using a device such as a smartphone or PC, where they enter personal information such as their height, weight, body type, and hairstyle, and take and upload a full-body image. The input data is then sent to the server.
[0137] Input: User's personal information (height, weight, body type, hairstyle) and full-body image
[0138] Output: Formatted personal information and image data sent to the server
[0139] Step 2:
[0140] Server operations: receiving and storing data
[0141] The server receives the personal information and image data sent by the user, converts the received data into an appropriate format (e.g., JSON format), and stores it in a database.
[0142] Input: Personal information and image data sent by the user
[0143] Output: Formatted personal information and image data stored in a database
[0144] Step 3:
[0145] Processing performed by the server: Image preprocessing
[0146] The server retrieves the image stored in the database, first cuts out the background and extracts only the user's body, then resizes the image to the appropriate size and applies a noise reduction filter to improve the image quality.
[0147] Input: Image data stored in a database
[0148] Output: Preprocessed image data
[0149] Step 4:
[0150] Server processing: 3D model generation
[0151] The server runs a 3D modeling algorithm based on the user's personal information (height, weight, body type) and preprocessed image data to generate a 3D model of the user, using the Blender API, for example.
[0152] Input: Personal information such as height, weight, and body shape, and preprocessed image data
[0153] Output: Generated 3D model of the user
[0154] Step 5:
[0155] Server process: Clothing data acquisition
[0156] The server retrieves data about the clothing item selected by the user from a database and converts it into a format optimized for simulation, including the texture, color, and design of the fabric.
[0157] Input: The identity of the clothing item selected by the user
[0158] Output: Optimized clothing data
[0159] Step 6:
[0160] Processing performed by the server: Try-on simulation
[0161] The server then applies the acquired clothing data to the generated 3D model of the user. The fitting simulation is performed using Unity's physics engine to reproduce the fit and visual appearance of the clothing. Finally, an image of the user after trying it on is generated.
[0162] Input: Generated 3D model and optimized clothing data
[0163] Output: Image after trying on
[0164] Step 7:
[0165] Server processing: Impression generation
[0166] The server then uses a generative AI model, such as a GPT-based model, to automatically generate feedback on the fit and appearance of the garment based on the generated try-on images.
[0167] Input: Image after trying on
[0168] Output: Generated impressions (text format)
[0169] Step 8:
[0170] Server action: Send feedback
[0171] The server then sends the generated post-try-on images and feedback to the user's device, allowing the user to view the feedback on their own device.
[0172] Input: Image after try-on and generated impressions
[0173] Output: Feedback sent to the user's device
[0174] Step 9:
[0175] User Action: Check Feedback
[0176] The user checks the images and feedback received on their device after trying on the item. Based on the feedback, the user decides whether to purchase the item, and if they do proceed with the purchase, they complete it on the e-commerce site.
[0177] Input: Try-on images and impressions received on the device
[0178] Output: Purchase decision and checkout
[0179] (Application example 1)
[0180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0181] On conventional e-commerce sites, when users purchase clothing, they are unable to try it on, which can lead to concerns about fit and appearance. Furthermore, users are often unable to actually try on clothing before purchasing, which often results in the hassle and expense of returning the product. There is a need to improve this situation and provide a system that allows users to comfortably purchase clothing.
[0182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0183] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the results of the try-on simulation and impressions, means for providing feedback using the generated AI model, and means for outputting the feedback content as prompt text. This allows users to select products through a virtual try-on experience without actually trying them on, reducing anxiety and stress and making it possible to reduce the effort and cost of returning items.
[0184] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[0185] "Image" refers to photographs or video data that show the user's entire body.
[0186] A "3D model" is a three-dimensional virtual model generated based on personal information and image data, which mimics the user's body.
[0187] "Clothing data" is information about the texture, color, design, etc. of the fabric of the clothing to be tried on.
[0188] "Try-on simulation" is a process in which clothing data is applied to a 3D model of the user, allowing them to virtually try on the clothes.
[0189] A "generative AI model" is an algorithm that uses artificial intelligence to automatically generate feedback and impressions.
[0190] "Feedback" refers to comments and evaluations provided to users by the generative AI model based on the results of the fitting simulation.
[0191] A "prompt" is a document of instructions or guidance that a generative AI model outputs to a user.
[0192] A description will be given of an embodiment of the present invention. A system for implementing the present invention includes the following main phases: a data input phase, a data analysis phase, a feedback phase, and a purchasing phase (optional).
[0193] Data Entry Phase
[0194] Users access an e-commerce site or a dedicated app from their smartphone or PC, enter their personal information such as height, weight, body type, hairstyle, and a full-body image, and then submit the data. The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. The server also performs preprocessing on the image data, such as cropping and resizing.
[0195] Data analysis phase
[0196] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It then resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body shape, to generate a 3D model of the user. This 3D model also incorporates the preprocessed image information.
[0197] The server also retrieves data (fabric texture, color, design) from the database for the clothing item the user selects to try on. Based on this data, the server runs a fitting simulation and uses a physics engine to reproduce the fit and visual appearance of the clothing.
[0198] Feedback Phase
[0199] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated post-try-on images. This feedback is output as prompt text and sent to the user's device, allowing the user to check the fitting simulation results and their feedback on their own device.
[0200] Hardware and software used
[0201] 1. Hardware: Smartphone, personal computer (PC), server, head-mounted display (HMD)
[0202] 2. Software: Flask (web framework), OpenCV (image processing library), machine learning libraries (such as scikit-learn), generative AI models (for natural language generation)
[0203] Specific examples
[0204] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. When the user takes a full-body photo with their smartphone and uploads it to the system, the server generates a 3D model of the user based on the received information. The server then combines this with data on the "blue T-shirt" the user selected to try on, and performs a try-on simulation. The final try-on image and user feedback, such as "The fit is perfect and the color is nice," are sent to the user's device. In this way, the user can virtually choose products without actually trying them on.
[0205] Prompt Sentence Examples
[0206] "If the user is 170cm tall, weighs 65kg, has a slim build, and has short hair, generate a fitting simulation of the blue T-shirt that this user wants to try on."
[0207] This allows users to purchase clothing while reducing anxiety and stress.
[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0209] Step 1:
[0210] Users access an e-commerce site or dedicated app using their smartphone or PC, enter personal information such as their height, weight, body type, hairstyle, and a full-body image, and then submit it. Specifically, after entering their personal information and image, the user presses the send button to send the data to the server. Input is via a text input form and an image upload function, and output is data in JSON format.
[0211] Step 2:
[0212] The server receives the JSON data sent by the user and analyzes each field. The received personal information and image data are saved in a database. At this time, the image data is preprocessed by cropping and resizing it and converting it into an appropriate format. Specifically, the image data is cropped using OpenCV and a noise reduction filter is applied. The output is the preprocessed image data and the analyzed personal information.
[0213] Step 3:
[0214] The server analyzes the preprocessed image, removes the background, and extracts only the user's body. It then resizes the image to an appropriate size and applies the noise reduction filter again. This step produces an image that extracts only the user's key physical characteristics. The input is the preprocessed image, and the output is an image of the user's body with the background removed.
[0215] Step 4:
[0216] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user. This model also incorporates preprocessed image information. Specifically, the 3D modeling software is used to recreate the user's three-dimensional body profile. The input is the analyzed personal information and preprocessed images, and the output is the user's 3D model.
[0217] Step 5:
[0218] The server retrieves data (fabric texture, color, design) of the clothing item selected by the user from the database. Specifically, it executes a database query to retrieve optimized clothing data. The input is the clothing item ID selected by the user, and the output is the corresponding clothing item data.
[0219] Step 6:
[0220] The server applies clothing data to the generated 3D model and performs a fitting simulation. This simulation uses a physics engine to reproduce the fit and visual appearance of the clothing, and generates a final image of the clothing being tried on. Specifically, the server uses simulation software to perform a dynamic simulation of the clothing. The input is the 3D model and clothing data, and the output is a simulated image of the clothing being tried on.
[0221] Step 7:
[0222] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated try-on images. This feedback is generated through analysis by a natural language processing model. The input is the try-on image, and the output is the feedback content.
[0223] Step 8:
[0224] The server formats the feedback content as a prompt sentence and sends it to the user's device. Specifically, the server obtains the feedback content from the generative AI model and formats it into a prompt sentence format. The input is the generated feedback, and the output is feedback as a prompt sentence.
[0225] Step 9:
[0226] The user checks the try-on simulation results and their impressions sent to their own device. After checking, they can make a purchase decision. Specifically, they check the try-on images on their smartphone or PC screen and read the feedback. The input is the feedback and try-on images sent from the server, and the output is the user's purchasing decision.
[0227] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0228] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[0229] System Overview
[0230] The system mainly includes five main components:
[0231] 1. Data entry phase
[0232] 2. Data analysis phase
[0233] 3. Feedback Phase
[0234] 4. Emotion Recognition Phase
[0235] 5. Purchasing Phase (Optional)
[0236] The specific processing of the program for each phase will be explained below.
[0237] Data Entry Phase
[0238] 1. User Actions
[0239] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[0240] 2. Processing performed by the device
[0241] The device sends the entered personal information and image to the server. By pressing the send button, the data is sent to the specified API endpoint.
[0242] Data analysis phase
[0243] 1. Processing performed by the server: Image preprocessing
[0244] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[0245] 2. Processing performed by the server: 3D model generation
[0246] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[0247] 3. Processing performed by the server: Acquiring clothing data
[0248] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[0249] 4. Processing performed by the server: Try-on simulation
[0250] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0251] Feedback Phase
[0252] 1. Server processing: Try-on images and impression generation
[0253] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[0254] 2. Server action: Sending feedback
[0255] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[0256] Emotion Recognition Phase
[0257] 1. Processing performed by the server: Execution of the emotion engine
[0258] The server analyzes the facial expressions and voices of the user while checking the fitting results and uses an emotion engine to recognize the user's emotions, such as joy, surprise, or dissatisfaction, based on facial expressions captured by a camera and voice recorded by a microphone.
[0259] 2. Server processing: Emotion-based feedback adjustment
[0260] The system adjusts the feedback provided based on the perceived emotions. For example, if the user is happy, it will emphasize the positive feedback and encourage them to purchase the product. On the other hand, if negative emotions are detected, it will adjust the feedback by suggesting a different product.
[0261] Purchasing Phase (Optional)
[0262] User actions
[0263] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[0264] Specific examples
[0265] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[0266] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[0267] The processing flow will be explained below.
[0268] Step 1:
[0269] User processing: The user accesses an e-commerce site or a dedicated app from their own device and logs in. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[0270] Step 2:
[0271] Processing performed by the device: Sends the entered personal information and image to the server. By pressing the send button, the entered data and image are sent to the specified API endpoint.
[0272] Step 3:
[0273] Server processing: Receives the transmitted data, converts it into an appropriate format, and stores it in the database. For images, preprocessing involves cropping, resizing, and noise removal.
[0274] Step 4:
[0275] Processing performed by the server: The background is removed from the preprocessed image data, and only the user's body is extracted. This processing allows the necessary features to be extracted effectively.
[0276] Step 5:
[0277] Processing performed by the server: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. Specifically, it runs a 3D modeling algorithm to create a three-dimensional digital model of the user.
[0278] Step 6:
[0279] Processing performed by the server: Retrieves data on the clothing item to be tried on from the database, including information on the clothing's fabric texture, design, color, etc.
[0280] Step 7:
[0281] Server processing: Applying clothing data to the user's 3D model and running a fitting simulation using a physics engine. This simulates the fit and visual appearance of the clothing in a realistic way.
[0282] Step 8:
[0283] Server processing: Based on the simulation results, automatically generate images and impressions after trying on the garment. Create text feedback including evaluations of fit and appearance.
[0284] Step 9:
[0285] Server process: The generated try-on images and feedback are sent to the user's device. The feedback is formatted and returned as an API response.
[0286] Step 10:
[0287] User action: Check the submitted try-on images and feedback and decide whether to purchase the product. Based on the feedback, the user can choose to either press the purchase button or try on other products.
[0288] Step 11:
[0289] Processing performed by the server: While the user is checking the feedback, the emotion engine analyzes the user's facial expressions and voice to recognize their emotions. Based on data acquired through the camera and microphone, emotions such as joy, surprise, and dissatisfaction are determined.
[0290] Step 12:
[0291] Processing performed by the server: The content of the feedback is adjusted based on the user's recognized emotions. For example, if a positive emotion is detected, a message to encourage purchase is displayed. If a negative emotion is detected, reference information or suggestions for other products are provided.
[0292] Step 13:
[0293] What happens to the user: They review the tailored feedback and make a final purchase decision. They can either click the buy button to proceed with the purchase or try on other items.
[0294] This process allows users to check the fit and appearance of clothing through virtual try-on, and the emotion engine provides more personalized feedback, reducing anxiety when purchasing clothing on e-commerce sites and increasing user motivation.
[0295] Example 2
[0296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0297] When purchasing clothing on traditional e-commerce sites, users are unable to try on items, which can often lead to anxiety when making a purchase. Furthermore, a simple virtual try-on simulation makes it difficult to accurately grasp how users will feel, and it is unable to provide emotional feedback, which fails to increase purchase motivation. Furthermore, it is difficult to provide personalized feedback based on each user's individual body type and preferences.
[0298] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0299] In this invention, the server includes: means for receiving personal information such as height, weight, body type, and hairstyle, as well as an image, from the user; means for processing the received personal information and image to generate a three-dimensional model of the user; and means for running a try-on simulation based on clothing information and the generated three-dimensional model. This allows the user to check the fit and appearance of clothing through a virtual try-on without actually visiting a store. The server also includes means for providing the user with the results of the try-on simulation and their impressions, means for recognizing the user's emotions, and means for adjusting feedback based on the recognized emotions. This allows for feedback that reflects the user's emotional state, enabling a more personalized shopping experience.
[0300] "Personal information" refers to data that indicates individual characteristics of a user, such as height, weight, body type, hairstyle, etc.
[0301] A "three-dimensional model" is a digital replica of a user's body shape and features, generated based on the user's personal information and image.
[0302] "Try-on simulation" is a simulation technology that applies clothing information to a generated three-dimensional model to reproduce the visual appearance and fit of the garment as if it were actually being tried on.
[0303] "Clothing information" refers to data that includes characteristics of clothing, such as texture, color, and design, and that is applied to a three-dimensional model.
[0304] "Feedback" refers to information and evaluations provided to users, including results of try-on simulations and impressions.
[0305] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state (happiness, dissatisfaction, etc.) from their facial expressions and voice.
[0306] "Feedback modulation" means changing the content of feedback based on perceived emotions.
[0307] This invention is a system that supports virtual try-on of clothing when users purchase it on an e-commerce website. The system receives personal information (height, weight, body type, hairstyle, etc.) and images from the user, generates a three-dimensional model based on the information, and then performs a try-on simulation based on the clothing information, providing the results and user feedback. It also recognizes the user's emotions and adjusts the feedback based on those emotions, providing a more personalized shopping experience.
[0308] Overall system flow
[0309] Data Entry Phase
[0310] The user accesses the e-commerce site or a dedicated app using a device (smartphone or PC), enters personal information such as their height, weight, body type, hairstyle, etc., along with a full-body image, and presses the send button. The device then sends this information to the server. The transmission is encrypted using SSL / TLS.
[0311] Data analysis phase
[0312] The server first preprocesses the received image by using an algorithm (e.g., GrabCut) to remove the background and extract only the user's body. It also resizes the image if necessary and applies a noise reduction filter (e.g., a Gaussian filter).
[0313] Next, the server runs a 3D modeling algorithm (e.g., Blender's API) based on the user's personal information such as height, weight, and body type to generate a 3D model of the user.
[0314] The server then retrieves the specified clothing information (texture, color, design, etc.) from the database, applies the clothing information to the user's 3D model, and performs a fitting simulation using a physics engine (e.g., PhysX). The fitting simulation reproduces the fit and visual appearance of the clothing, and generates a final try-on image.
[0315] Feedback Phase
[0316] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback on the fit and appearance of the garment based on the post-try-on images. The generated try-on images and feedback are then formatted and sent to the user's device, where the user can view the results.
[0317] Emotion Recognition Phase
[0318] The server analyzes the user's facial expressions and voice in real time while checking the fitting results. Using facial expression data captured by the camera and voice data obtained by the microphone, an emotion recognition engine (such as Microsoft Azure's Emotion API) is used to recognize the user's emotional state. Based on the recognized emotion, the server adjusts the content of the feedback. For example, if the user is pleased, it will emphasize positive feedback to encourage the purchase.
[0319] Purchasing Phase (Optional)
[0320] The user checks the feedback and if they like the product, they press the purchase button to complete the purchase. They are then prompted to enter their credit card information and shipping address, and once these are entered, the order is confirmed.
[0321] Specific examples
[0322] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a three-dimensional model of the user based on the received information, retrieves information about a "blue shirt" from the database, and simulates trying it on. A physics engine is activated to simulate how the shirt will fit the user, and generates an image of the user after trying it on. The generated image and a comment such as "It fits perfectly and the color is nice" are sent to the user.
[0323] Next, the camera captures the user's facial expressions as they check the fitting results, and if the emotion engine recognizes their joy, it provides positive feedback such as, "This shirt looks great on you!"
[0324] Prompt Sentence Examples
[0325] "The user is 170cm tall, 65kg, slim, with short hair. He is trying on a blue shirt. Generate the results of the fitting simulation."
[0326] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0327] Step 1:
[0328] A user accesses an e-commerce site or a dedicated app from their own device (smartphone or PC) and enters their personal information (height, weight, body type, hairstyle, etc.) and a full-body image. The entered data is validated within the device to ensure that all required information has been entered. The input data is often in JSON format.
[0329] Input: Personal information and full-body image entered by the user
[0330] Output: Verified personal information and image data
[0331] Step 2:
[0332] The device sends the entered personal information and image to the API endpoint. The data is encrypted with SSL / TLS and sent securely to the server.
[0333] Input: Verified personal information and image data
[0334] Output: Data sent to the API endpoint
[0335] Step 3:
[0336] The server analyzes the received image data and performs image preprocessing. Specifically, it applies the GrabCut algorithm to remove the background and extract the user's body parts. Then, it resizes the image and removes noise using a Gaussian filter as needed.
[0337] Input: Image data sent
[0338] Output: Image data of the body part with the background removed
[0339] Step 4:
[0340] Based on the personal information received, the server runs a 3D modeling algorithm (e.g., Blender API) to generate a 3D model of the user. This step also integrates data obtained from image pre-processing.
[0341] Input: Personal information and preprocessed image data
[0342] Output: 3D model of the user
[0343] Step 5:
[0344] The server retrieves information about the clothing item the user has selected to try on from the database, including the clothing's texture, color, design, etc. This data is then optimized for the simulation.
[0345] Input: Clothing identification information
[0346] Output: Optimized clothing data
[0347] Step 6:
[0348] The server applies the optimized clothing data to the 3D model and performs a fitting simulation using a physics engine (e.g., PhysX), generating a fitting image that reproduces the fit and appearance of the clothing.
[0349] Input: 3D model and optimized clothing data
[0350] Output: Image after trying on
[0351] Step 7:
[0352] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback about the fit and appearance of the garment based on post-try-on images, which are then written in natural language and formatted as feedback to the user.
[0353] Input: Image after trying on
[0354] Output: Feedback on the generated fit and appearance
[0355] Step 8:
[0356] The server sends the generated post-try-on images and feedback to the user's device. The feedback consists of images and textual feedback.
[0357] Input: Images and impressions after trying on
[0358] Output: Feedback sent to the user's device
[0359] Step 9:
[0360] The user terminal displays the fitting results, and the user checks them. While the user is checking, the camera and microphone are used to capture facial expressions and voice data.
[0361] Input: Feedback
[0362] Output: Captured facial and voice data
[0363] Step 10:
[0364] The server sends the captured facial and voice data to an emotion recognition engine (such as Microsoft Azure's Emotion API) to recognize the user's emotions. The recognized emotions are analyzed in real time.
[0365] Input: Captured facial and voice data
[0366] Output: Recognized emotion data
[0367] Step 11:
[0368] The server adjusts the feedback based on the recognized emotion. For example, if the user is happy, it will emphasize the positive feedback and encourage them to buy the product. On the other hand, if a negative emotion is detected, it will adjust the feedback by suggesting a different product.
[0369] Input: Recognized emotion data
[0370] Output: Regulated Feedback
[0371] Step 12:
[0372] If the user is motivated to purchase after reviewing the adjusted feedback, they press the purchase button to complete the purchase. At this point, they are prompted to enter credit card information and a shipping address, and the order is confirmed once the entered data is sent to the server.
[0373] Input: Tailored feedback and checkout information
[0374] Output: Purchase confirmation notice
[0375] (Application example 2)
[0376] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0377] Traditional e-commerce sites face the challenge of making it difficult for users to check the size and fit of clothing when purchasing. As a result, products are often disappointing when they arrive, leading to frequent returns and exchanges. Another problem is the lack of feedback to encourage users to purchase. In particular, the lack of feedback that reflects users' emotions limits the effectiveness of these sites in promoting purchases.
[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0379] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the user with the results of the try-on simulation and their impressions, and means for analyzing the user's emotions and adjusting the feedback content based on the recognized emotions. This allows the user to virtually try on clothing that fits their body, and by providing feedback based on their emotions, it is possible to achieve a greater purchasing promotion effect.
[0380] "User" refers to an individual who uses the System to virtually try on clothing items.
[0381] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[0382] "Image" refers to photographic data capturing the user's entire body.
[0383] "3D model" refers to a three-dimensional virtual shape of a user that is generated based on the received personal information and image.
[0384] "Clothing data" refers to data that includes information such as the texture, color, and design of clothing fabrics.
[0385] "Try-on simulation" refers to the process of applying clothing data to a generated 3D model to perform a virtual try-on.
[0386] "Try-on simulation results" refers to the visual results and feedback obtained after running a try-on simulation.
[0387] "Opinion" refers to an evaluation provided on fit and appearance based on the results of a try-on simulation.
[0388] "Analyzing emotions" refers to the process of recognizing emotions from the user's facial expressions and voice.
[0389] "Adjusting feedback content" refers to the process of changing the content of the feedback provided based on perceived emotions.
[0390] MODE FOR CARRYING OUT THE INVENTION
[0391] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[0392] Data Entry Phase
[0393] First, the user accesses an e-commerce site or a dedicated app using their own device (smartphone or PC) and uploads a full-body image along with personal information such as their height, weight, body type, and hairstyle. They enter this information through the app's interface and press the send button. The device then sends the entered personal information and image to the server. The data is sent to a specified API endpoint.
[0394] Data analysis phase
[0395] The server performs the following processes based on the received personal information and images:
[0396] 1. Image preprocessing: Use OpenCV to cut out the background and extract only the user's body. Then, resize the image to the appropriate size and apply a noise reduction filter.
[0397] 2. 3D model generation: A 3D modeling algorithm using TensorFlow generates a 3D model of the user, incorporating preprocessed image information.
[0398] 3. Clothing data acquisition: Data (fabric texture, color, design) of the clothing item selected by the user to try on is acquired from a database (e.g., AWS DynamoDB). The acquired data is optimized for simulation.
[0399] 4. Try-on simulation: Using a physics engine, clothing data is applied to a 3D model of the user to simulate a try-on, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0400] Feedback Phase
[0401] The server uses the AI model to automatically generate feedback about the fit and appearance of the garment based on the generated post-try-on images. These feedback provide feedback as if the user had actually tried the garment on. The server then formats the try-on images and feedback and sends them to the user's device.
[0402] Emotion Recognition Phase
[0403] The server analyzes the user's facial expressions and voice while checking the fitting results, and uses an emotion engine to recognize the user's emotions. For example, it can recognize happiness, surprise, dissatisfaction, etc. from facial expressions captured by a camera and voice captured by a microphone. It then adjusts the feedback it provides based on the recognized emotions. If the user is happy, it emphasizes positive feedback and encourages the user to purchase the product. On the other hand, if negative emotions are detected, it makes adjustments such as suggesting a different product.
[0404] Purchasing Phase (Optional)
[0405] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[0406] Specific examples
[0407] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[0408] Prompt Sentence Examples
[0409] User A: Height 170cm, weight 65kg, slim build, short hair. Considering buying a blue T-shirt. Please generate a try-on simulation result and feedback.
[0410] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[0411] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0412] Step 1:
[0413] Users use their own devices to upload personal information such as height, weight, body type, hairstyle, and a full-body image. The entered data is sent from the device to the server.
[0414] Step 2:
[0415] The server performs preprocessing on the received image. Specifically, it uses OpenCV to cut out the background of the image, remove noise, and extract only the user's body. This process outputs the user's outline data.
[0416] Step 3:
[0417] The server runs a 3D modeling algorithm using TensorFlow based on the preprocessed images and personal information to generate a 3D model of the user. The input data here is the user's physical features and contour data, and the output data is the user's 3D model.
[0418] Step 4:
[0419] The server retrieves clothing data from a database (e.g., AWS DynamoDB), including the texture, color, and design of the clothing fabric, and outputs the data in a format optimized for simulation.
[0420] Step 5:
[0421] The server uses a physics engine to simulate the user trying on the clothes based on their 3D model and clothing data. As a result of the simulation, visual image data of the clothes after trying them on is generated.
[0422] Step 6:
[0423] The server uses an AI model (NLP model) to generate automatically generated impressions about fit and appearance based on the image data after trying on. The input is the image data after trying on, and the output is a textual impression about fit and appearance.
[0424] Step 7:
[0425] The server formats the image data after trying on the clothes and the created feedback, and sends it to the user's device. The input is the image data after trying on the clothes and the feedback, and the output is feedback data in a format that the user can check.
[0426] Step 8:
[0427] When the user checks the fitting results, the device's camera captures the user's facial expression and the device's microphone captures voice data, which is then sent to the server in real time.
[0428] Step 9:
[0429] The server analyzes the user's facial expressions and voice data using an emotion engine, recognizes the user's emotions (happiness, surprise, dissatisfaction, etc.), and outputs this emotional data.
[0430] Step 10:
[0431] The server adjusts the feedback based on the recognized emotion. For example, if a positive emotion is detected, it outputs positive feedback to encourage the purchase of a product, and if a negative emotion is detected, it suggests a different product.
[0432] Step 11:
[0433] The user decides whether to purchase the product based on the provided feedback. If the user decides to purchase the product, he or she presses the purchase button, the purchase procedure is executed on the server, and the purchase data is output.
[0434] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0435] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0436] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0437] [Second embodiment]
[0438] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0439] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0440] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0441] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0442] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0444] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0445] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0446] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0447] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0448] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0449] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0450] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system performs a try-on simulation based on personal information such as height, weight, figure, hairstyle, etc. provided by the user, as well as an image, and provides the results.
[0451] System Overview
[0452] The system mainly includes four main components:
[0453] 1. Data entry phase
[0454] 2. Data analysis phase
[0455] 3. Feedback Phase
[0456] 4. Purchasing Phase (Optional)
[0457] The specific processing of the program for each phase will be explained below.
[0458] Data Entry Phase
[0459] 1. User Actions
[0460] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[0461] 2. Processing performed by the server
[0462] The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. It also performs preprocessing on image data, such as cropping and resizing.
[0463] Data analysis phase
[0464] 1. Processing performed by the server: Image preprocessing
[0465] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[0466] 2. Processing performed by the server: 3D model generation
[0467] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[0468] 3. Processing performed by the server: Acquiring clothing data
[0469] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[0470] 4. Processing performed by the server: Try-on simulation
[0471] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0472] Feedback Phase
[0473] 1. Server processing: Try-on images and impression generation
[0474] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[0475] 2. Server action: Sending feedback
[0476] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[0477] 3. User Action: Feedback Check
[0478] The user can then check the try-on images and their impressions and decide whether to purchase. If they like the item, they can proceed with the purchase.
[0479] Specific examples
[0480] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, integrates it with the data for the "blue T-shirt" that the user selected to try on, and performs a try-on simulation. Finally, the generated try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user.
[0481] In this way, users can choose products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress that comes with purchasing clothing on an e-commerce site.
[0482] The processing flow will be explained below.
[0483] Step 1:
[0484] User process: The user accesses an e-commerce site or a dedicated app from their own device. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[0485] Step 2:
[0486] Processing performed by the device: The entered personal information and image are sent to the server. By pressing the send button, the data is sent to the specified API endpoint.
[0487] Step 3:
[0488] Processing by the server: The server converts the received personal information and images into an appropriate format and stores them in a database. The images are pre-processed for analysis.
[0489] Step 4:
[0490] Server processing: Preprocesses the uploaded image by removing the background and extracting only the user's body. At this stage, the image is resized and a noise reduction filter is applied.
[0491] Step 5:
[0492] Server processing: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. This information is used to run a 3D modeling algorithm to create a three-dimensional digital model of the user.
[0493] Step 6:
[0494] Processing performed by the server: Retrieves data on the clothing item selected by the user from the database, including information on the clothing item's fabric texture, design, color, etc.
[0495] Step 7:
[0496] Server processing: The acquired clothing data is applied to the user's 3D model, and a fitting simulation is performed. A physics engine is used to simulate the fit, drape, wrinkles, etc. of the clothing.
[0497] Step 8:
[0498] Server processing: Analyzes visual information and automatically generates impressions along with the post-try-on images generated as a result of the simulation, creating text feedback on fit and appearance.
[0499] Step 9:
[0500] Server process: The generated try-on images and feedback are sent to the user's device. The data is formatted for the user interface and returned as an API response.
[0501] Step 10:
[0502] User action: Check the submitted try-on images and comments. Display the feedback on the e-commerce site or in the dedicated app and decide whether to purchase the product.
[0503] Step 11 (Optional):
[0504] What the user does: If they like the item, they press the "Purchase" button to complete the purchase. They can also try on other items if they wish.
[0505] This process allows users to simulate the fit and appearance of clothing through virtual try-on without actually trying it on. This system is expected to significantly reduce the anxiety and stress felt when purchasing clothing on an e-commerce site.
[0506] Example 1
[0507] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0508] Online shopping, especially when purchasing clothing, can cause anxiety and stress when users choose products without actually trying them on. Specifically, it is difficult for users to check in advance how the size, fit, and visual appearance of the clothing they choose will actually look like. Another issue is the time and effort required to visit a physical store to try on items.
[0509] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0510] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a three-dimensional model of the user, means for running a try-on simulation based on clothing data and the generated three-dimensional model, means for providing the user with the results of the try-on simulation and feedback, means for sending the generated feedback to the user's device, and means for the user to check the feedback on their own device. This allows users to select products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress of online shopping.
[0511] "User" refers to an end user who uses the system to provide their personal information and images and simulate trying on clothing items.
[0512] "Personal information" refers to attribute information such as height, weight, body type, hairstyle, etc. provided by the user.
[0513] "Images" refers to photographs or videos of a user's entire body that they upload to the system.
[0514] "3D model" refers to a 3D virtual human model generated based on a user's personal information and image.
[0515] "Clothing data" refers to data including information such as the texture, color, and design of the clothing fabric, which is necessary for performing a fitting simulation.
[0516] "Try-on simulation" refers to the process of applying clothing data to a three-dimensional model of the user and virtually recreating the fit and appearance of the clothing using a physics engine or other means.
[0517] "Opinions" refer to evaluation comments about fit and appearance generated based on the results of the try-on simulation.
[0518] "Terminal" refers to an electronic device such as a smartphone or computer that a user uses to access the system.
[0519] "Feedback" refers to providing information to the user, including the results of the try-on simulation and the generated impressions.
[0520] "Generative AI model" refers to an artificial intelligence algorithm or program that automatically generates impressions about fit and appearance based on the results of a fitting simulation.
[0521] MODE FOR CARRYING OUT THE INVENTION
[0522] The present invention provides a system that allows users to virtually try on clothing when purchasing it online, and check the fit and appearance of the clothing after trying it on. This system is primarily built using the following hardware and software:
[0523] Hardware used
[0524] 1. Server: A high-performance server capable of processing and storing large amounts of data.
[0525] 2. Device: A device that can connect to the internet, such as a smartphone, tablet, or PC.
[0526] Software used
[0527] 1. Database: Database software (e.g., MySQL, PostgreSQL) for storing personal information, 3D models, and clothing data.
[0528] 2. Image processing software: Software for cropping, resizing, and noise removal of images (e.g., OpenCV).
[0529] 3. 3D modeling algorithms: Software used to generate a 3D model of the user (e.g. Blender API).
[0530] 4. Try-on simulation software: Software that applies clothing data to a three-dimensional model to simulate trying on the garment (e.g., Unity).
[0531] 5. Generative AI model: An AI model (e.g., a GPT-based model) that automatically generates impressions on fit and appearance based on the results of a try-on simulation.
[0532] System overview and specific processing
[0533] Data Entry Phase
[0534] Users access a dedicated app or e-commerce site using their own device, enter personal information such as their height, weight, body type, and hairstyle, along with a full-body image, and submit it. The server converts the received data into an appropriate format and stores it in a database. The image data undergoes pre-processing such as cropping, resizing, and noise removal.
[0535] Data analysis phase
[0536] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It also resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information to generate a 3D model, which also incorporates the preprocessed image information. It then retrieves data for the selected clothing item from a database and converts it into a format optimized for simulation.
[0537] Try-on simulation phase
[0538] The server applies the clothing data to the generated 3D model of the user and performs a fitting simulation using a physics engine, thereby reproducing the fit and visual appearance of the clothing and generating the final after-fit image.
[0539] Feedback Phase
[0540] The server uses a generative AI model to automatically generate feedback on the fit and appearance of the clothing based on the generated try-on images. The try-on images and feedback are then sent to the user's device, where the user can review the feedback and decide whether to purchase the clothing.
[0541] Specific examples
[0542] For example, if a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair, they can take a full-body photo with their smartphone and upload it to the system. The server generates a three-dimensional model of the user based on the received information, obtains data on the "blue T-shirt" selected for trying on, and performs a try-on simulation. Finally, the generated try-on image and a user's feedback, such as "It fits perfectly and the color is nice," are sent to the user's device.
[0543] Prompt Sentence Examples
[0544] "The user is 170cm tall, weighs 65kg, has a slim build, and has short hair. Please simulate trying on the 'blue T-shirt' selected by this user and generate impressions about the fit and appearance."
[0545] By building a system like this, users can check the fit and appearance of clothing through virtual try-ons, allowing them to enjoy online shopping with peace of mind.
[0546] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0547] Step 1:
[0548] User Action: Data Entry
[0549] Users access a dedicated app or e-commerce site using a device such as a smartphone or PC, where they enter personal information such as their height, weight, body type, and hairstyle, and take and upload a full-body image. The input data is then sent to the server.
[0550] Input: User's personal information (height, weight, body type, hairstyle) and full-body image
[0551] Output: Formatted personal information and image data sent to the server
[0552] Step 2:
[0553] Server operations: receiving and storing data
[0554] The server receives the personal information and image data sent by the user, converts the received data into an appropriate format (e.g., JSON format), and stores it in a database.
[0555] Input: Personal information and image data sent by the user
[0556] Output: Formatted personal information and image data stored in a database
[0557] Step 3:
[0558] Processing performed by the server: Image preprocessing
[0559] The server retrieves the image stored in the database, first cuts out the background and extracts only the user's body, then resizes the image to the appropriate size and applies a noise reduction filter to improve the image quality.
[0560] Input: Image data stored in a database
[0561] Output: Preprocessed image data
[0562] Step 4:
[0563] Server processing: 3D model generation
[0564] The server runs a 3D modeling algorithm based on the user's personal information (height, weight, body type) and preprocessed image data to generate a 3D model of the user, using the Blender API, for example.
[0565] Input: Personal information such as height, weight, and body shape, and preprocessed image data
[0566] Output: Generated 3D model of the user
[0567] Step 5:
[0568] Server process: Clothing data acquisition
[0569] The server retrieves data about the clothing item selected by the user from a database and converts it into a format optimized for simulation, including the texture, color, and design of the fabric.
[0570] Input: The identity of the clothing item selected by the user
[0571] Output: Optimized clothing data
[0572] Step 6:
[0573] Processing performed by the server: Try-on simulation
[0574] The server then applies the acquired clothing data to the generated 3D model of the user. The fitting simulation is performed using Unity's physics engine to reproduce the fit and visual appearance of the clothing. Finally, an image of the user after trying it on is generated.
[0575] Input: Generated 3D model and optimized clothing data
[0576] Output: Image after trying on
[0577] Step 7:
[0578] Server processing: Impression generation
[0579] The server then uses a generative AI model, such as a GPT-based model, to automatically generate feedback on the fit and appearance of the garment based on the generated try-on images.
[0580] Input: Image after trying on
[0581] Output: Generated impressions (text format)
[0582] Step 8:
[0583] Server action: Send feedback
[0584] The server then sends the generated post-try-on images and feedback to the user's device, allowing the user to view the feedback on their own device.
[0585] Input: Image after try-on and generated impressions
[0586] Output: Feedback sent to the user's device
[0587] Step 9:
[0588] User Action: Check Feedback
[0589] The user checks the images and feedback received on their device after trying on the item. Based on the feedback, the user decides whether to purchase the item, and if they do proceed with the purchase, they complete it on the e-commerce site.
[0590] Input: Try-on images and impressions received on the device
[0591] Output: Purchase decision and checkout
[0592] (Application example 1)
[0593] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0594] On conventional e-commerce sites, when users purchase clothing, they are unable to try it on, which can lead to concerns about fit and appearance. Furthermore, users are often unable to actually try on clothing before purchasing, which often results in the hassle and expense of returning the product. There is a need to improve this situation and provide a system that allows users to comfortably purchase clothing.
[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0596] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the results of the try-on simulation and impressions, means for providing feedback using the generated AI model, and means for outputting the feedback content as prompt text. This allows users to select products through a virtual try-on experience without actually trying them on, reducing anxiety and stress and making it possible to reduce the effort and cost of returning items.
[0597] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[0598] "Image" refers to photographs or video data that show the user's entire body.
[0599] A "3D model" is a three-dimensional virtual model generated based on personal information and image data, which mimics the user's body.
[0600] "Clothing data" is information about the texture, color, design, etc. of the fabric of the clothing to be tried on.
[0601] "Try-on simulation" is a process in which clothing data is applied to a 3D model of the user, allowing them to virtually try on the clothes.
[0602] A "generative AI model" is an algorithm that uses artificial intelligence to automatically generate feedback and impressions.
[0603] "Feedback" refers to comments and evaluations provided to users by the generative AI model based on the results of the fitting simulation.
[0604] A "prompt" is a document of instructions or guidance that a generative AI model outputs to a user.
[0605] A description will be given of an embodiment of the present invention. A system for implementing the present invention includes the following main phases: a data input phase, a data analysis phase, a feedback phase, and a purchasing phase (optional).
[0606] Data Entry Phase
[0607] Users access an e-commerce site or a dedicated app from their smartphone or PC, enter their personal information such as height, weight, body type, hairstyle, and a full-body image, and then submit the data. The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. The server also performs preprocessing on the image data, such as cropping and resizing.
[0608] Data analysis phase
[0609] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It then resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body shape, to generate a 3D model of the user. This 3D model also incorporates the preprocessed image information.
[0610] The server also retrieves data (fabric texture, color, design) from the database for the clothing item the user selects to try on. Based on this data, the server runs a fitting simulation and uses a physics engine to reproduce the fit and visual appearance of the clothing.
[0611] Feedback Phase
[0612] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated post-try-on images. This feedback is output as prompt text and sent to the user's device, allowing the user to check the fitting simulation results and their feedback on their own device.
[0613] Hardware and software used
[0614] 1. Hardware: Smartphone, personal computer (PC), server, head-mounted display (HMD)
[0615] 2. Software: Flask (web framework), OpenCV (image processing library), machine learning libraries (such as scikit-learn), generative AI models (for natural language generation)
[0616] Specific examples
[0617] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. When the user takes a full-body photo with their smartphone and uploads it to the system, the server generates a 3D model of the user based on the received information. The server then combines this with data on the "blue T-shirt" the user selected to try on, and performs a try-on simulation. The final try-on image and user feedback, such as "The fit is perfect and the color is nice," are sent to the user's device. In this way, the user can virtually choose products without actually trying them on.
[0618] Prompt Sentence Examples
[0619] "If the user is 170cm tall, weighs 65kg, has a slim build, and has short hair, generate a fitting simulation of the blue T-shirt that this user wants to try on."
[0620] This allows users to purchase clothing while reducing anxiety and stress.
[0621] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0622] Step 1:
[0623] Users access an e-commerce site or dedicated app using their smartphone or PC, enter personal information such as their height, weight, body type, hairstyle, and a full-body image, and then submit it. Specifically, after entering their personal information and image, the user presses the send button to send the data to the server. Input is via a text input form and an image upload function, and output is data in JSON format.
[0624] Step 2:
[0625] The server receives the JSON data sent by the user and analyzes each field. The received personal information and image data are saved in a database. At this time, the image data is preprocessed by cropping and resizing it and converting it into an appropriate format. Specifically, the image data is cropped using OpenCV and a noise reduction filter is applied. The output is the preprocessed image data and the analyzed personal information.
[0626] Step 3:
[0627] The server analyzes the preprocessed image, removes the background, and extracts only the user's body. It then resizes the image to an appropriate size and applies the noise reduction filter again. This step produces an image that extracts only the user's key physical characteristics. The input is the preprocessed image, and the output is an image of the user's body with the background removed.
[0628] Step 4:
[0629] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user. This model also incorporates preprocessed image information. Specifically, the 3D modeling software is used to recreate the user's three-dimensional body profile. The input is the analyzed personal information and preprocessed images, and the output is the user's 3D model.
[0630] Step 5:
[0631] The server retrieves data (fabric texture, color, design) of the clothing item selected by the user from the database. Specifically, it executes a database query to retrieve optimized clothing data. The input is the clothing item ID selected by the user, and the output is the corresponding clothing item data.
[0632] Step 6:
[0633] The server applies clothing data to the generated 3D model and performs a fitting simulation. This simulation uses a physics engine to reproduce the fit and visual appearance of the clothing, and generates a final image of the clothing being tried on. Specifically, the server uses simulation software to perform a dynamic simulation of the clothing. The input is the 3D model and clothing data, and the output is a simulated image of the clothing being tried on.
[0634] Step 7:
[0635] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated try-on images. This feedback is generated through analysis by a natural language processing model. The input is the try-on image, and the output is the feedback content.
[0636] Step 8:
[0637] The server formats the feedback content as a prompt sentence and sends it to the user's device. Specifically, the server obtains the feedback content from the generative AI model and formats it into a prompt sentence format. The input is the generated feedback, and the output is feedback as a prompt sentence.
[0638] Step 9:
[0639] The user checks the try-on simulation results and their impressions sent to their own device. After checking, they can make a purchase decision. Specifically, they check the try-on images on their smartphone or PC screen and read the feedback. The input is the feedback and try-on images sent from the server, and the output is the user's purchasing decision.
[0640] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0641] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[0642] System Overview
[0643] The system mainly includes five main components:
[0644] 1. Data entry phase
[0645] 2. Data analysis phase
[0646] 3. Feedback Phase
[0647] 4. Emotion Recognition Phase
[0648] 5. Purchasing Phase (Optional)
[0649] The specific processing of the program for each phase will be explained below.
[0650] Data Entry Phase
[0651] 1. User Actions
[0652] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[0653] 2. Processing performed by the device
[0654] The device sends the entered personal information and image to the server. By pressing the send button, the data is sent to the specified API endpoint.
[0655] Data analysis phase
[0656] 1. Processing performed by the server: Image preprocessing
[0657] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[0658] 2. Processing performed by the server: 3D model generation
[0659] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[0660] 3. Processing performed by the server: Acquiring clothing data
[0661] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[0662] 4. Processing performed by the server: Try-on simulation
[0663] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0664] Feedback Phase
[0665] 1. Server processing: Try-on images and impression generation
[0666] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[0667] 2. Server action: Sending feedback
[0668] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[0669] Emotion Recognition Phase
[0670] 1. Processing performed by the server: Execution of the emotion engine
[0671] The server analyzes the facial expressions and voices of the user while checking the fitting results and uses an emotion engine to recognize the user's emotions, such as joy, surprise, or dissatisfaction, based on facial expressions captured by a camera and voice recorded by a microphone.
[0672] 2. Server processing: Emotion-based feedback adjustment
[0673] The system adjusts the feedback provided based on the perceived emotions. For example, if the user is happy, it will emphasize the positive feedback and encourage them to purchase the product. On the other hand, if negative emotions are detected, it will adjust the feedback by suggesting a different product.
[0674] Purchasing Phase (Optional)
[0675] User actions
[0676] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[0677] Specific examples
[0678] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[0679] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[0680] The processing flow will be explained below.
[0681] Step 1:
[0682] User processing: The user accesses an e-commerce site or a dedicated app from their own device and logs in. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[0683] Step 2:
[0684] Processing performed by the device: Sends the entered personal information and image to the server. By pressing the send button, the entered data and image are sent to the specified API endpoint.
[0685] Step 3:
[0686] Server processing: Receives the transmitted data, converts it into an appropriate format, and stores it in the database. For images, preprocessing involves cropping, resizing, and noise removal.
[0687] Step 4:
[0688] Processing performed by the server: The background is removed from the preprocessed image data, and only the user's body is extracted. This processing allows the necessary features to be extracted effectively.
[0689] Step 5:
[0690] Processing performed by the server: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. Specifically, it runs a 3D modeling algorithm to create a three-dimensional digital model of the user.
[0691] Step 6:
[0692] Processing performed by the server: Retrieves data on the clothing item to be tried on from the database, including information on the clothing's fabric texture, design, color, etc.
[0693] Step 7:
[0694] Server processing: Applying clothing data to the user's 3D model and running a fitting simulation using a physics engine. This simulates the fit and visual appearance of the clothing in a realistic way.
[0695] Step 8:
[0696] Server processing: Based on the simulation results, automatically generate images and impressions after trying on the garment. Create text feedback including evaluations of fit and appearance.
[0697] Step 9:
[0698] Server process: The generated try-on images and feedback are sent to the user's device. The feedback is formatted and returned as an API response.
[0699] Step 10:
[0700] User action: Check the submitted try-on images and feedback and decide whether to purchase the product. Based on the feedback, the user can choose to either press the purchase button or try on other products.
[0701] Step 11:
[0702] Processing performed by the server: While the user is checking the feedback, the emotion engine analyzes the user's facial expressions and voice to recognize their emotions. Based on data acquired through the camera and microphone, emotions such as joy, surprise, and dissatisfaction are determined.
[0703] Step 12:
[0704] Processing performed by the server: The content of the feedback is adjusted based on the user's recognized emotions. For example, if a positive emotion is detected, a message to encourage purchase is displayed. If a negative emotion is detected, reference information or suggestions for other products are provided.
[0705] Step 13:
[0706] What happens to the user: They review the tailored feedback and make a final purchase decision. They can either click the buy button to proceed with the purchase or try on other items.
[0707] This process allows users to check the fit and appearance of clothing through virtual try-on, and the emotion engine provides more personalized feedback, reducing anxiety when purchasing clothing on e-commerce sites and increasing user motivation.
[0708] Example 2
[0709] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0710] When purchasing clothing on traditional e-commerce sites, users are unable to try on items, which can often lead to anxiety when making a purchase. Furthermore, a simple virtual try-on simulation makes it difficult to accurately grasp how users will feel, and it is unable to provide emotional feedback, which fails to increase purchase motivation. Furthermore, it is difficult to provide personalized feedback based on each user's individual body type and preferences.
[0711] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0712] In this invention, the server includes: means for receiving personal information such as height, weight, body type, and hairstyle, as well as an image, from the user; means for processing the received personal information and image to generate a three-dimensional model of the user; and means for running a try-on simulation based on clothing information and the generated three-dimensional model. This allows the user to check the fit and appearance of clothing through a virtual try-on without actually visiting a store. The server also includes means for providing the user with the results of the try-on simulation and their impressions, means for recognizing the user's emotions, and means for adjusting feedback based on the recognized emotions. This allows for feedback that reflects the user's emotional state, enabling a more personalized shopping experience.
[0713] "Personal information" refers to data that indicates individual characteristics of a user, such as height, weight, body type, hairstyle, etc.
[0714] A "three-dimensional model" is a digital replica of a user's body shape and features, generated based on the user's personal information and image.
[0715] "Try-on simulation" is a simulation technology that applies clothing information to a generated three-dimensional model to reproduce the visual appearance and fit of the garment as if it were actually being tried on.
[0716] "Clothing information" refers to data that includes characteristics of clothing, such as texture, color, and design, and that is applied to a three-dimensional model.
[0717] "Feedback" refers to information and evaluations provided to users, including results of try-on simulations and impressions.
[0718] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state (happiness, dissatisfaction, etc.) from their facial expressions and voice.
[0719] "Feedback modulation" means changing the content of feedback based on perceived emotions.
[0720] This invention is a system that supports virtual try-on of clothing when users purchase it on an e-commerce website. The system receives personal information (height, weight, body type, hairstyle, etc.) and images from the user, generates a three-dimensional model based on the information, and then performs a try-on simulation based on the clothing information, providing the results and user feedback. It also recognizes the user's emotions and adjusts the feedback based on those emotions, providing a more personalized shopping experience.
[0721] Overall system flow
[0722] Data Entry Phase
[0723] The user accesses the e-commerce site or a dedicated app using a device (smartphone or PC), enters personal information such as their height, weight, body type, hairstyle, etc., along with a full-body image, and presses the send button. The device then sends this information to the server. The transmission is encrypted using SSL / TLS.
[0724] Data analysis phase
[0725] The server first preprocesses the received image by using an algorithm (e.g., GrabCut) to remove the background and extract only the user's body. It also resizes the image if necessary and applies a noise reduction filter (e.g., a Gaussian filter).
[0726] Next, the server runs a 3D modeling algorithm (e.g., Blender's API) based on the user's personal information such as height, weight, and body type to generate a 3D model of the user.
[0727] The server then retrieves the specified clothing information (texture, color, design, etc.) from the database, applies the clothing information to the user's 3D model, and performs a fitting simulation using a physics engine (e.g., PhysX). The fitting simulation reproduces the fit and visual appearance of the clothing, and generates a final try-on image.
[0728] Feedback Phase
[0729] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback on the fit and appearance of the garment based on the post-try-on images. The generated try-on images and feedback are then formatted and sent to the user's device, where the user can view the results.
[0730] Emotion Recognition Phase
[0731] The server analyzes the user's facial expressions and voice in real time while checking the fitting results. Using facial expression data captured by the camera and voice data obtained by the microphone, an emotion recognition engine (such as Microsoft Azure's Emotion API) is used to recognize the user's emotional state. Based on the recognized emotion, the server adjusts the content of the feedback. For example, if the user is pleased, it will emphasize positive feedback to encourage the purchase.
[0732] Purchasing Phase (Optional)
[0733] The user checks the feedback and if they like the product, they press the purchase button to complete the purchase. They are then prompted to enter their credit card information and shipping address, and once these are entered, the order is confirmed.
[0734] Specific examples
[0735] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a three-dimensional model of the user based on the received information, retrieves information about a "blue shirt" from the database, and simulates trying it on. A physics engine is activated to simulate how the shirt will fit the user, and generates an image of the user after trying it on. The generated image and a comment such as "It fits perfectly and the color is nice" are sent to the user.
[0736] Next, the camera captures the user's facial expressions as they check the fitting results, and if the emotion engine recognizes their joy, it provides positive feedback such as, "This shirt looks great on you!"
[0737] Prompt Sentence Examples
[0738] "The user is 170cm tall, 65kg, slim, with short hair. He is trying on a blue shirt. Generate the results of the fitting simulation."
[0739] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0740] Step 1:
[0741] A user accesses an e-commerce site or a dedicated app from their own device (smartphone or PC) and enters their personal information (height, weight, body type, hairstyle, etc.) and a full-body image. The entered data is validated within the device to ensure that all required information has been entered. The input data is often in JSON format.
[0742] Input: Personal information and full-body image entered by the user
[0743] Output: Verified personal information and image data
[0744] Step 2:
[0745] The device sends the entered personal information and image to the API endpoint. The data is encrypted with SSL / TLS and sent securely to the server.
[0746] Input: Verified personal information and image data
[0747] Output: Data sent to the API endpoint
[0748] Step 3:
[0749] The server analyzes the received image data and performs image preprocessing. Specifically, it applies the GrabCut algorithm to remove the background and extract the user's body parts. Then, it resizes the image and removes noise using a Gaussian filter as needed.
[0750] Input: Image data sent
[0751] Output: Image data of the body part with the background removed
[0752] Step 4:
[0753] Based on the personal information received, the server runs a 3D modeling algorithm (e.g., Blender API) to generate a 3D model of the user. This step also integrates data obtained from image pre-processing.
[0754] Input: Personal information and preprocessed image data
[0755] Output: 3D model of the user
[0756] Step 5:
[0757] The server retrieves information about the clothing item the user has selected to try on from the database, including the clothing's texture, color, design, etc. This data is then optimized for the simulation.
[0758] Input: Clothing identification information
[0759] Output: Optimized clothing data
[0760] Step 6:
[0761] The server applies the optimized clothing data to the 3D model and performs a fitting simulation using a physics engine (e.g., PhysX), generating a fitting image that reproduces the fit and appearance of the clothing.
[0762] Input: 3D model and optimized clothing data
[0763] Output: Image after trying on
[0764] Step 7:
[0765] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback about the fit and appearance of the garment based on post-try-on images, which are then written in natural language and formatted as feedback to the user.
[0766] Input: Image after trying on
[0767] Output: Feedback on the generated fit and appearance
[0768] Step 8:
[0769] The server sends the generated post-try-on images and feedback to the user's device. The feedback consists of images and textual feedback.
[0770] Input: Images and impressions after trying on
[0771] Output: Feedback sent to the user's device
[0772] Step 9:
[0773] The user terminal displays the fitting results, and the user checks them. While the user is checking, the camera and microphone are used to capture facial expressions and voice data.
[0774] Input: Feedback
[0775] Output: Captured facial and voice data
[0776] Step 10:
[0777] The server sends the captured facial and voice data to an emotion recognition engine (such as Microsoft Azure's Emotion API) to recognize the user's emotions. The recognized emotions are analyzed in real time.
[0778] Input: Captured facial and voice data
[0779] Output: Recognized emotion data
[0780] Step 11:
[0781] The server adjusts the feedback based on the recognized emotion. For example, if the user is happy, it will emphasize the positive feedback and encourage them to buy the product. On the other hand, if a negative emotion is detected, it will adjust the feedback by suggesting a different product.
[0782] Input: Recognized emotion data
[0783] Output: Regulated Feedback
[0784] Step 12:
[0785] If the user is motivated to purchase after reviewing the adjusted feedback, they press the purchase button to complete the purchase. At this point, they are prompted to enter credit card information and a shipping address, and the order is confirmed once the entered data is sent to the server.
[0786] Input: Tailored feedback and checkout information
[0787] Output: Purchase confirmation notice
[0788] (Application example 2)
[0789] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0790] Traditional e-commerce sites face the challenge of making it difficult for users to check the size and fit of clothing when purchasing. As a result, products are often disappointing when they arrive, leading to frequent returns and exchanges. Another problem is the lack of feedback to encourage users to purchase. In particular, the lack of feedback that reflects users' emotions limits the effectiveness of these sites in promoting purchases.
[0791] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0792] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the user with the results of the try-on simulation and their impressions, and means for analyzing the user's emotions and adjusting the feedback content based on the recognized emotions. This allows the user to virtually try on clothing that fits their body, and by providing feedback based on their emotions, it is possible to achieve a greater purchasing promotion effect.
[0793] "User" refers to an individual who uses the System to virtually try on clothing items.
[0794] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[0795] "Image" refers to photographic data capturing the user's entire body.
[0796] "3D model" refers to a three-dimensional virtual shape of a user that is generated based on the received personal information and image.
[0797] "Clothing data" refers to data that includes information such as the texture, color, and design of clothing fabrics.
[0798] "Try-on simulation" refers to the process of applying clothing data to a generated 3D model to perform a virtual try-on.
[0799] "Try-on simulation results" refers to the visual results and feedback obtained after running a try-on simulation.
[0800] "Opinion" refers to an evaluation provided on fit and appearance based on the results of a try-on simulation.
[0801] "Analyzing emotions" refers to the process of recognizing emotions from the user's facial expressions and voice.
[0802] "Adjusting feedback content" refers to the process of changing the content of the feedback provided based on perceived emotions.
[0803] MODE FOR CARRYING OUT THE INVENTION
[0804] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[0805] Data Entry Phase
[0806] First, the user accesses an e-commerce site or a dedicated app using their own device (smartphone or PC) and uploads a full-body image along with personal information such as their height, weight, body type, and hairstyle. They enter this information through the app's interface and press the send button. The device then sends the entered personal information and image to the server. The data is sent to a specified API endpoint.
[0807] Data analysis phase
[0808] The server performs the following processes based on the received personal information and images:
[0809] 1. Image preprocessing: Use OpenCV to cut out the background and extract only the user's body. Then, resize the image to the appropriate size and apply a noise reduction filter.
[0810] 2. 3D model generation: A 3D modeling algorithm using TensorFlow generates a 3D model of the user, incorporating preprocessed image information.
[0811] 3. Clothing data acquisition: Data (fabric texture, color, design) of the clothing item selected by the user to try on is acquired from a database (e.g., AWS DynamoDB). The acquired data is optimized for simulation.
[0812] 4. Try-on simulation: Using a physics engine, clothing data is applied to a 3D model of the user to simulate a try-on, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0813] Feedback Phase
[0814] The server uses the AI model to automatically generate feedback about the fit and appearance of the garment based on the generated post-try-on images. These feedback provide feedback as if the user had actually tried the garment on. The server then formats the try-on images and feedback and sends them to the user's device.
[0815] Emotion Recognition Phase
[0816] The server analyzes the user's facial expressions and voice while checking the fitting results, and uses an emotion engine to recognize the user's emotions. For example, it can recognize happiness, surprise, dissatisfaction, etc. from facial expressions captured by a camera and voice captured by a microphone. It then adjusts the feedback it provides based on the recognized emotions. If the user is happy, it emphasizes positive feedback and encourages the user to purchase the product. On the other hand, if negative emotions are detected, it makes adjustments such as suggesting a different product.
[0817] Purchasing Phase (Optional)
[0818] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[0819] Specific examples
[0820] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[0821] Prompt Sentence Examples
[0822] User A: Height 170cm, weight 65kg, slim build, short hair. Considering buying a blue T-shirt. Please generate a try-on simulation result and feedback.
[0823] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[0824] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0825] Step 1:
[0826] Users use their own devices to upload personal information such as height, weight, body type, hairstyle, and a full-body image. The entered data is sent from the device to the server.
[0827] Step 2:
[0828] The server performs preprocessing on the received image. Specifically, it uses OpenCV to cut out the background of the image, remove noise, and extract only the user's body. This process outputs the user's outline data.
[0829] Step 3:
[0830] The server runs a 3D modeling algorithm using TensorFlow based on the preprocessed images and personal information to generate a 3D model of the user. The input data here is the user's physical features and contour data, and the output data is the user's 3D model.
[0831] Step 4:
[0832] The server retrieves clothing data from a database (e.g., AWS DynamoDB), including the texture, color, and design of the clothing fabric, and outputs the data in a format optimized for simulation.
[0833] Step 5:
[0834] The server uses a physics engine to simulate the user trying on the clothes based on their 3D model and clothing data. As a result of the simulation, visual image data of the clothes after trying them on is generated.
[0835] Step 6:
[0836] The server uses an AI model (NLP model) to generate automatically generated impressions about fit and appearance based on the image data after trying on. The input is the image data after trying on, and the output is a textual impression about fit and appearance.
[0837] Step 7:
[0838] The server formats the image data after trying on the clothes and the created feedback, and sends it to the user's device. The input is the image data after trying on the clothes and the feedback, and the output is feedback data in a format that the user can check.
[0839] Step 8:
[0840] When the user checks the fitting results, the device's camera captures the user's facial expression and the device's microphone captures voice data, which is then sent to the server in real time.
[0841] Step 9:
[0842] The server analyzes the user's facial expressions and voice data using an emotion engine, recognizes the user's emotions (happiness, surprise, dissatisfaction, etc.), and outputs this emotional data.
[0843] Step 10:
[0844] The server adjusts the feedback based on the recognized emotion. For example, if a positive emotion is detected, it outputs positive feedback to encourage the purchase of a product, and if a negative emotion is detected, it suggests a different product.
[0845] Step 11:
[0846] The user decides whether to purchase the product based on the provided feedback. If the user decides to purchase the product, he or she presses the purchase button, the purchase procedure is executed on the server, and the purchase data is output.
[0847] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0848] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0849] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0850] [Third embodiment]
[0851] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0852] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0853] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0854] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0855] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0856] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0857] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0858] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0859] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0860] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0861] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0862] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0863] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system performs a try-on simulation based on personal information such as height, weight, figure, hairstyle, etc. provided by the user, as well as an image, and provides the results.
[0864] System Overview
[0865] The system mainly includes four main components:
[0866] 1. Data entry phase
[0867] 2. Data analysis phase
[0868] 3. Feedback Phase
[0869] 4. Purchasing Phase (Optional)
[0870] The specific processing of the program for each phase will be explained below.
[0871] Data Entry Phase
[0872] 1. User Actions
[0873] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[0874] 2. Processing performed by the server
[0875] The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. It also performs preprocessing on image data, such as cropping and resizing.
[0876] Data analysis phase
[0877] 1. Processing performed by the server: Image preprocessing
[0878] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[0879] 2. Processing performed by the server: 3D model generation
[0880] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[0881] 3. Processing performed by the server: Acquiring clothing data
[0882] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[0883] 4. Processing performed by the server: Try-on simulation
[0884] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[0885] Feedback Phase
[0886] 1. Server processing: Try-on images and impression generation
[0887] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[0888] 2. Server action: Sending feedback
[0889] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[0890] 3. User Action: Feedback Check
[0891] The user can then check the try-on images and their impressions and decide whether to purchase. If they like the item, they can proceed with the purchase.
[0892] Specific examples
[0893] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, integrates it with the data for the "blue T-shirt" that the user selected to try on, and performs a try-on simulation. Finally, the generated try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user.
[0894] In this way, users can choose products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress that comes with purchasing clothing on an e-commerce site.
[0895] The processing flow will be explained below.
[0896] Step 1:
[0897] User process: The user accesses an e-commerce site or a dedicated app from their own device. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[0898] Step 2:
[0899] Processing performed by the device: The entered personal information and image are sent to the server. By pressing the send button, the data is sent to the specified API endpoint.
[0900] Step 3:
[0901] Processing by the server: The server converts the received personal information and images into an appropriate format and stores them in a database. The images are pre-processed for analysis.
[0902] Step 4:
[0903] Server processing: Preprocesses the uploaded image by removing the background and extracting only the user's body. At this stage, the image is resized and a noise reduction filter is applied.
[0904] Step 5:
[0905] Server processing: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. This information is used to run a 3D modeling algorithm to create a three-dimensional digital model of the user.
[0906] Step 6:
[0907] Processing performed by the server: Retrieves data on the clothing item selected by the user from the database, including information on the clothing item's fabric texture, design, color, etc.
[0908] Step 7:
[0909] Server processing: The acquired clothing data is applied to the user's 3D model, and a fitting simulation is performed. A physics engine is used to simulate the fit, drape, wrinkles, etc. of the clothing.
[0910] Step 8:
[0911] Server processing: Analyzes visual information and automatically generates impressions along with the post-try-on images generated as a result of the simulation, creating text feedback on fit and appearance.
[0912] Step 9:
[0913] Server process: The generated try-on images and feedback are sent to the user's device. The data is formatted for the user interface and returned as an API response.
[0914] Step 10:
[0915] User action: Check the submitted try-on images and comments. Display the feedback on the e-commerce site or in the dedicated app and decide whether to purchase the product.
[0916] Step 11 (Optional):
[0917] What the user does: If they like the item, they press the "Purchase" button to complete the purchase. They can also try on other items if they wish.
[0918] This process allows users to simulate the fit and appearance of clothing through virtual try-on without actually trying it on. This system is expected to significantly reduce the anxiety and stress felt when purchasing clothing on an e-commerce site.
[0919] Example 1
[0920] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0921] Online shopping, especially when purchasing clothing, can cause anxiety and stress when users choose products without actually trying them on. Specifically, it is difficult for users to check in advance how the size, fit, and visual appearance of the clothing they choose will actually look like. Another issue is the time and effort required to visit a physical store to try on items.
[0922] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0923] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a three-dimensional model of the user, means for running a try-on simulation based on clothing data and the generated three-dimensional model, means for providing the user with the results of the try-on simulation and feedback, means for sending the generated feedback to the user's device, and means for the user to check the feedback on their own device. This allows users to select products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress of online shopping.
[0924] "User" refers to an end user who uses the system to provide their personal information and images and simulate trying on clothing items.
[0925] "Personal information" refers to attribute information such as height, weight, body type, hairstyle, etc. provided by the user.
[0926] "Images" refers to photographs or videos of a user's entire body that they upload to the system.
[0927] "3D model" refers to a 3D virtual human model generated based on a user's personal information and image.
[0928] "Clothing data" refers to data including information such as the texture, color, and design of the clothing fabric, which is necessary for performing a fitting simulation.
[0929] "Try-on simulation" refers to the process of applying clothing data to a three-dimensional model of the user and virtually recreating the fit and appearance of the clothing using a physics engine or other means.
[0930] "Opinions" refer to evaluation comments about fit and appearance generated based on the results of the try-on simulation.
[0931] "Terminal" refers to an electronic device such as a smartphone or computer that a user uses to access the system.
[0932] "Feedback" refers to providing information to the user, including the results of the try-on simulation and the generated impressions.
[0933] "Generative AI model" refers to an artificial intelligence algorithm or program that automatically generates impressions about fit and appearance based on the results of a fitting simulation.
[0934] MODE FOR CARRYING OUT THE INVENTION
[0935] The present invention provides a system that allows users to virtually try on clothing when purchasing it online, and check the fit and appearance of the clothing after trying it on. This system is primarily built using the following hardware and software:
[0936] Hardware used
[0937] 1. Server: A high-performance server capable of processing and storing large amounts of data.
[0938] 2. Device: A device that can connect to the internet, such as a smartphone, tablet, or PC.
[0939] Software used
[0940] 1. Database: Database software (e.g., MySQL, PostgreSQL) for storing personal information, 3D models, and clothing data.
[0941] 2. Image processing software: Software for cropping, resizing, and noise removal of images (e.g., OpenCV).
[0942] 3. 3D modeling algorithms: Software used to generate a 3D model of the user (e.g. Blender API).
[0943] 4. Try-on simulation software: Software that applies clothing data to a three-dimensional model to simulate trying on the garment (e.g., Unity).
[0944] 5. Generative AI model: An AI model (e.g., a GPT-based model) that automatically generates impressions on fit and appearance based on the results of a try-on simulation.
[0945] System overview and specific processing
[0946] Data Entry Phase
[0947] Users access a dedicated app or e-commerce site using their own device, enter personal information such as their height, weight, body type, and hairstyle, along with a full-body image, and submit it. The server converts the received data into an appropriate format and stores it in a database. The image data undergoes pre-processing such as cropping, resizing, and noise removal.
[0948] Data analysis phase
[0949] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It also resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information to generate a 3D model, which also incorporates the preprocessed image information. It then retrieves data for the selected clothing item from a database and converts it into a format optimized for simulation.
[0950] Try-on simulation phase
[0951] The server applies the clothing data to the generated 3D model of the user and performs a fitting simulation using a physics engine, thereby reproducing the fit and visual appearance of the clothing and generating the final after-fit image.
[0952] Feedback Phase
[0953] The server uses a generative AI model to automatically generate feedback on the fit and appearance of the clothing based on the generated try-on images. The try-on images and feedback are then sent to the user's device, where the user can review the feedback and decide whether to purchase the clothing.
[0954] Specific examples
[0955] For example, if a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair, they can take a full-body photo with their smartphone and upload it to the system. The server generates a three-dimensional model of the user based on the received information, obtains data on the "blue T-shirt" selected for trying on, and performs a try-on simulation. Finally, the generated try-on image and a user's feedback, such as "It fits perfectly and the color is nice," are sent to the user's device.
[0956] Prompt Sentence Examples
[0957] "The user is 170cm tall, weighs 65kg, has a slim build, and has short hair. Please simulate trying on the 'blue T-shirt' selected by this user and generate impressions about the fit and appearance."
[0958] By building a system like this, users can check the fit and appearance of clothing through virtual try-ons, allowing them to enjoy online shopping with peace of mind.
[0959] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0960] Step 1:
[0961] User Action: Data Entry
[0962] Users access a dedicated app or e-commerce site using a device such as a smartphone or PC, where they enter personal information such as their height, weight, body type, and hairstyle, and take and upload a full-body image. The input data is then sent to the server.
[0963] Input: User's personal information (height, weight, body type, hairstyle) and full-body image
[0964] Output: Formatted personal information and image data sent to the server
[0965] Step 2:
[0966] Server operations: receiving and storing data
[0967] The server receives the personal information and image data sent by the user, converts the received data into an appropriate format (e.g., JSON format), and stores it in a database.
[0968] Input: Personal information and image data sent by the user
[0969] Output: Formatted personal information and image data stored in a database
[0970] Step 3:
[0971] Processing performed by the server: Image preprocessing
[0972] The server retrieves the image stored in the database, first cuts out the background and extracts only the user's body, then resizes the image to the appropriate size and applies a noise reduction filter to improve the image quality.
[0973] Input: Image data stored in a database
[0974] Output: Preprocessed image data
[0975] Step 4:
[0976] Server processing: 3D model generation
[0977] The server runs a 3D modeling algorithm based on the user's personal information (height, weight, body type) and preprocessed image data to generate a 3D model of the user, using the Blender API, for example.
[0978] Input: Personal information such as height, weight, and body shape, and preprocessed image data
[0979] Output: Generated 3D model of the user
[0980] Step 5:
[0981] Server process: Clothing data acquisition
[0982] The server retrieves data about the clothing item selected by the user from a database and converts it into a format optimized for simulation, including the texture, color, and design of the fabric.
[0983] Input: The identity of the clothing item selected by the user
[0984] Output: Optimized clothing data
[0985] Step 6:
[0986] Processing performed by the server: Try-on simulation
[0987] The server then applies the acquired clothing data to the generated 3D model of the user. The fitting simulation is performed using Unity's physics engine to reproduce the fit and visual appearance of the clothing. Finally, an image of the user after trying it on is generated.
[0988] Input: Generated 3D model and optimized clothing data
[0989] Output: Image after trying on
[0990] Step 7:
[0991] Server processing: Impression generation
[0992] The server then uses a generative AI model, such as a GPT-based model, to automatically generate feedback on the fit and appearance of the garment based on the generated try-on images.
[0993] Input: Image after trying on
[0994] Output: Generated impressions (text format)
[0995] Step 8:
[0996] Server action: Send feedback
[0997] The server then sends the generated post-try-on images and feedback to the user's device, allowing the user to view the feedback on their own device.
[0998] Input: Image after try-on and generated impressions
[0999] Output: Feedback sent to the user's device
[1000] Step 9:
[1001] User Action: Check Feedback
[1002] The user checks the images and feedback received on their device after trying on the item. Based on the feedback, the user decides whether to purchase the item, and if they do proceed with the purchase, they complete it on the e-commerce site.
[1003] Input: Try-on images and impressions received on the device
[1004] Output: Purchase decision and checkout
[1005] (Application example 1)
[1006] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1007] On conventional e-commerce sites, when users purchase clothing, they are unable to try it on, which can lead to concerns about fit and appearance. Furthermore, users are often unable to actually try on clothing before purchasing, which often results in the hassle and expense of returning the product. There is a need to improve this situation and provide a system that allows users to comfortably purchase clothing.
[1008] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1009] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the results of the try-on simulation and impressions, means for providing feedback using the generated AI model, and means for outputting the feedback content as prompt text. This allows users to select products through a virtual try-on experience without actually trying them on, reducing anxiety and stress and making it possible to reduce the effort and cost of returning items.
[1010] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[1011] "Image" refers to photographs or video data that show the user's entire body.
[1012] A "3D model" is a three-dimensional virtual model generated based on personal information and image data, which mimics the user's body.
[1013] "Clothing data" is information about the texture, color, design, etc. of the fabric of the clothing to be tried on.
[1014] "Try-on simulation" is a process in which clothing data is applied to a 3D model of the user, allowing them to virtually try on the clothes.
[1015] A "generative AI model" is an algorithm that uses artificial intelligence to automatically generate feedback and impressions.
[1016] "Feedback" refers to comments and evaluations provided to users by the generative AI model based on the results of the fitting simulation.
[1017] A "prompt" is a document of instructions or guidance that a generative AI model outputs to a user.
[1018] A description will be given of an embodiment of the present invention. A system for implementing the present invention includes the following main phases: a data input phase, a data analysis phase, a feedback phase, and a purchasing phase (optional).
[1019] Data Entry Phase
[1020] Users access an e-commerce site or a dedicated app from their smartphone or PC, enter their personal information such as height, weight, body type, hairstyle, and a full-body image, and then submit the data. The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. The server also performs preprocessing on the image data, such as cropping and resizing.
[1021] Data analysis phase
[1022] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It then resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body shape, to generate a 3D model of the user. This 3D model also incorporates the preprocessed image information.
[1023] The server also retrieves data (fabric texture, color, design) from the database for the clothing item the user selects to try on. Based on this data, the server runs a fitting simulation and uses a physics engine to reproduce the fit and visual appearance of the clothing.
[1024] Feedback Phase
[1025] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated post-try-on images. This feedback is output as prompt text and sent to the user's device, allowing the user to check the fitting simulation results and their feedback on their own device.
[1026] Hardware and software used
[1027] 1. Hardware: Smartphone, personal computer (PC), server, head-mounted display (HMD)
[1028] 2. Software: Flask (web framework), OpenCV (image processing library), machine learning libraries (such as scikit-learn), generative AI models (for natural language generation)
[1029] Specific examples
[1030] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. When the user takes a full-body photo with their smartphone and uploads it to the system, the server generates a 3D model of the user based on the received information. The server then combines this with data on the "blue T-shirt" the user selected to try on, and performs a try-on simulation. The final try-on image and user feedback, such as "The fit is perfect and the color is nice," are sent to the user's device. In this way, the user can virtually choose products without actually trying them on.
[1031] Prompt Sentence Examples
[1032] "If the user is 170cm tall, weighs 65kg, has a slim build, and has short hair, generate a fitting simulation of the blue T-shirt that this user wants to try on."
[1033] This allows users to purchase clothing while reducing anxiety and stress.
[1034] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1035] Step 1:
[1036] Users access an e-commerce site or dedicated app using their smartphone or PC, enter personal information such as their height, weight, body type, hairstyle, and a full-body image, and then submit it. Specifically, after entering their personal information and image, the user presses the send button to send the data to the server. Input is via a text input form and an image upload function, and output is data in JSON format.
[1037] Step 2:
[1038] The server receives the JSON data sent by the user and analyzes each field. The received personal information and image data are saved in a database. At this time, the image data is preprocessed by cropping and resizing it and converting it into an appropriate format. Specifically, the image data is cropped using OpenCV and a noise reduction filter is applied. The output is the preprocessed image data and the analyzed personal information.
[1039] Step 3:
[1040] The server analyzes the preprocessed image, removes the background, and extracts only the user's body. It then resizes the image to an appropriate size and applies the noise reduction filter again. This step produces an image that extracts only the user's key physical characteristics. The input is the preprocessed image, and the output is an image of the user's body with the background removed.
[1041] Step 4:
[1042] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user. This model also incorporates preprocessed image information. Specifically, the 3D modeling software is used to recreate the user's three-dimensional body profile. The input is the analyzed personal information and preprocessed images, and the output is the user's 3D model.
[1043] Step 5:
[1044] The server retrieves data (fabric texture, color, design) of the clothing item selected by the user from the database. Specifically, it executes a database query to retrieve optimized clothing data. The input is the clothing item ID selected by the user, and the output is the corresponding clothing item data.
[1045] Step 6:
[1046] The server applies clothing data to the generated 3D model and performs a fitting simulation. This simulation uses a physics engine to reproduce the fit and visual appearance of the clothing, and generates a final image of the clothing being tried on. Specifically, the server uses simulation software to perform a dynamic simulation of the clothing. The input is the 3D model and clothing data, and the output is a simulated image of the clothing being tried on.
[1047] Step 7:
[1048] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated try-on images. This feedback is generated through analysis by a natural language processing model. The input is the try-on image, and the output is the feedback content.
[1049] Step 8:
[1050] The server formats the feedback content as a prompt sentence and sends it to the user's device. Specifically, the server obtains the feedback content from the generative AI model and formats it into a prompt sentence format. The input is the generated feedback, and the output is feedback as a prompt sentence.
[1051] Step 9:
[1052] The user checks the try-on simulation results and their impressions sent to their own device. After checking, they can make a purchase decision. Specifically, they check the try-on images on their smartphone or PC screen and read the feedback. The input is the feedback and try-on images sent from the server, and the output is the user's purchasing decision.
[1053] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1054] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[1055] System Overview
[1056] The system mainly includes five main components:
[1057] 1. Data entry phase
[1058] 2. Data analysis phase
[1059] 3. Feedback Phase
[1060] 4. Emotion Recognition Phase
[1061] 5. Purchasing Phase (Optional)
[1062] The specific processing of the program for each phase will be explained below.
[1063] Data Entry Phase
[1064] 1. User Actions
[1065] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[1066] 2. Processing performed by the device
[1067] The device sends the entered personal information and image to the server. By pressing the send button, the data is sent to the specified API endpoint.
[1068] Data analysis phase
[1069] 1. Processing performed by the server: Image preprocessing
[1070] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[1071] 2. Processing performed by the server: 3D model generation
[1072] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[1073] 3. Processing performed by the server: Acquiring clothing data
[1074] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[1075] 4. Processing performed by the server: Try-on simulation
[1076] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[1077] Feedback Phase
[1078] 1. Server processing: Try-on images and impression generation
[1079] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[1080] 2. Server action: Sending feedback
[1081] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[1082] Emotion Recognition Phase
[1083] 1. Processing performed by the server: Execution of the emotion engine
[1084] The server analyzes the facial expressions and voices of the user while checking the fitting results and uses an emotion engine to recognize the user's emotions, such as joy, surprise, or dissatisfaction, based on facial expressions captured by a camera and voice recorded by a microphone.
[1085] 2. Server processing: Emotion-based feedback adjustment
[1086] The system adjusts the feedback provided based on the perceived emotions. For example, if the user is happy, it will emphasize the positive feedback and encourage them to purchase the product. On the other hand, if negative emotions are detected, it will adjust the feedback by suggesting a different product.
[1087] Purchasing Phase (Optional)
[1088] User actions
[1089] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[1090] Specific examples
[1091] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[1092] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[1093] The processing flow will be explained below.
[1094] Step 1:
[1095] User processing: The user accesses an e-commerce site or a dedicated app from their own device and logs in. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[1096] Step 2:
[1097] Processing performed by the device: Sends the entered personal information and image to the server. By pressing the send button, the entered data and image are sent to the specified API endpoint.
[1098] Step 3:
[1099] Server processing: Receives the transmitted data, converts it into an appropriate format, and stores it in the database. For images, preprocessing involves cropping, resizing, and noise removal.
[1100] Step 4:
[1101] Processing performed by the server: The background is removed from the preprocessed image data, and only the user's body is extracted. This processing allows the necessary features to be extracted effectively.
[1102] Step 5:
[1103] Processing performed by the server: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. Specifically, it runs a 3D modeling algorithm to create a three-dimensional digital model of the user.
[1104] Step 6:
[1105] Processing performed by the server: Retrieves data on the clothing item to be tried on from the database, including information on the clothing's fabric texture, design, color, etc.
[1106] Step 7:
[1107] Server processing: Applying clothing data to the user's 3D model and running a fitting simulation using a physics engine. This simulates the fit and visual appearance of the clothing in a realistic way.
[1108] Step 8:
[1109] Server processing: Based on the simulation results, automatically generate images and impressions after trying on the garment. Create text feedback including evaluations of fit and appearance.
[1110] Step 9:
[1111] Server process: The generated try-on images and feedback are sent to the user's device. The feedback is formatted and returned as an API response.
[1112] Step 10:
[1113] User action: Check the submitted try-on images and feedback and decide whether to purchase the product. Based on the feedback, the user can choose to either press the purchase button or try on other products.
[1114] Step 11:
[1115] Processing performed by the server: While the user is checking the feedback, the emotion engine analyzes the user's facial expressions and voice to recognize their emotions. Based on data acquired through the camera and microphone, emotions such as joy, surprise, and dissatisfaction are determined.
[1116] Step 12:
[1117] Processing performed by the server: The content of the feedback is adjusted based on the user's recognized emotions. For example, if a positive emotion is detected, a message to encourage purchase is displayed. If a negative emotion is detected, reference information or suggestions for other products are provided.
[1118] Step 13:
[1119] What happens to the user: They review the tailored feedback and make a final purchase decision. They can either click the buy button to proceed with the purchase or try on other items.
[1120] This process allows users to check the fit and appearance of clothing through virtual try-on, and the emotion engine provides more personalized feedback, reducing anxiety when purchasing clothing on e-commerce sites and increasing user motivation.
[1121] Example 2
[1122] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1123] When purchasing clothing on traditional e-commerce sites, users are unable to try on items, which can often lead to anxiety when making a purchase. Furthermore, a simple virtual try-on simulation makes it difficult to accurately grasp how users will feel, and it is unable to provide emotional feedback, which fails to increase purchase motivation. Furthermore, it is difficult to provide personalized feedback based on each user's individual body type and preferences.
[1124] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1125] In this invention, the server includes: means for receiving personal information such as height, weight, body type, and hairstyle, as well as an image, from the user; means for processing the received personal information and image to generate a three-dimensional model of the user; and means for running a try-on simulation based on clothing information and the generated three-dimensional model. This allows the user to check the fit and appearance of clothing through a virtual try-on without actually visiting a store. The server also includes means for providing the user with the results of the try-on simulation and their impressions, means for recognizing the user's emotions, and means for adjusting feedback based on the recognized emotions. This allows for feedback that reflects the user's emotional state, enabling a more personalized shopping experience.
[1126] "Personal information" refers to data that indicates individual characteristics of a user, such as height, weight, body type, hairstyle, etc.
[1127] A "three-dimensional model" is a digital replica of a user's body shape and features, generated based on the user's personal information and image.
[1128] "Try-on simulation" is a simulation technology that applies clothing information to a generated three-dimensional model to reproduce the visual appearance and fit of the garment as if it were actually being tried on.
[1129] "Clothing information" refers to data that includes characteristics of clothing, such as texture, color, and design, and that is applied to a three-dimensional model.
[1130] "Feedback" refers to information and evaluations provided to users, including results of try-on simulations and impressions.
[1131] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state (happiness, dissatisfaction, etc.) from their facial expressions and voice.
[1132] "Feedback modulation" means changing the content of feedback based on perceived emotions.
[1133] This invention is a system that supports virtual try-on of clothing when users purchase it on an e-commerce website. The system receives personal information (height, weight, body type, hairstyle, etc.) and images from the user, generates a three-dimensional model based on the information, and then performs a try-on simulation based on the clothing information, providing the results and user feedback. It also recognizes the user's emotions and adjusts the feedback based on those emotions, providing a more personalized shopping experience.
[1134] Overall system flow
[1135] Data Entry Phase
[1136] The user accesses the e-commerce site or a dedicated app using a device (smartphone or PC), enters personal information such as their height, weight, body type, hairstyle, etc., along with a full-body image, and presses the send button. The device then sends this information to the server. The transmission is encrypted using SSL / TLS.
[1137] Data analysis phase
[1138] The server first preprocesses the received image by using an algorithm (e.g., GrabCut) to remove the background and extract only the user's body. It also resizes the image if necessary and applies a noise reduction filter (e.g., a Gaussian filter).
[1139] Next, the server runs a 3D modeling algorithm (e.g., Blender's API) based on the user's personal information such as height, weight, and body type to generate a 3D model of the user.
[1140] The server then retrieves the specified clothing information (texture, color, design, etc.) from the database, applies the clothing information to the user's 3D model, and performs a fitting simulation using a physics engine (e.g., PhysX). The fitting simulation reproduces the fit and visual appearance of the clothing, and generates a final try-on image.
[1141] Feedback Phase
[1142] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback on the fit and appearance of the garment based on the post-try-on images. The generated try-on images and feedback are then formatted and sent to the user's device, where the user can view the results.
[1143] Emotion Recognition Phase
[1144] The server analyzes the user's facial expressions and voice in real time while checking the fitting results. Using facial expression data captured by the camera and voice data obtained by the microphone, an emotion recognition engine (such as Microsoft Azure's Emotion API) is used to recognize the user's emotional state. Based on the recognized emotion, the server adjusts the content of the feedback. For example, if the user is pleased, it will emphasize positive feedback to encourage the purchase.
[1145] Purchasing Phase (Optional)
[1146] The user checks the feedback and if they like the product, they press the purchase button to complete the purchase. They are then prompted to enter their credit card information and shipping address, and once these are entered, the order is confirmed.
[1147] Specific examples
[1148] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a three-dimensional model of the user based on the received information, retrieves information about a "blue shirt" from the database, and simulates trying it on. A physics engine is activated to simulate how the shirt will fit the user, and generates an image of the user after trying it on. The generated image and a comment such as "It fits perfectly and the color is nice" are sent to the user.
[1149] Next, the camera captures the user's facial expressions as they check the fitting results, and if the emotion engine recognizes their joy, it provides positive feedback such as, "This shirt looks great on you!"
[1150] Prompt Sentence Examples
[1151] "The user is 170cm tall, 65kg, slim, with short hair. He is trying on a blue shirt. Generate the results of the fitting simulation."
[1152] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1153] Step 1:
[1154] A user accesses an e-commerce site or a dedicated app from their own device (smartphone or PC) and enters their personal information (height, weight, body type, hairstyle, etc.) and a full-body image. The entered data is validated within the device to ensure that all required information has been entered. The input data is often in JSON format.
[1155] Input: Personal information and full-body image entered by the user
[1156] Output: Verified personal information and image data
[1157] Step 2:
[1158] The device sends the entered personal information and image to the API endpoint. The data is encrypted with SSL / TLS and sent securely to the server.
[1159] Input: Verified personal information and image data
[1160] Output: Data sent to the API endpoint
[1161] Step 3:
[1162] The server analyzes the received image data and performs image preprocessing. Specifically, it applies the GrabCut algorithm to remove the background and extract the user's body parts. Then, it resizes the image and removes noise using a Gaussian filter as needed.
[1163] Input: Image data sent
[1164] Output: Image data of the body part with the background removed
[1165] Step 4:
[1166] Based on the personal information received, the server runs a 3D modeling algorithm (e.g., Blender API) to generate a 3D model of the user. This step also integrates data obtained from image pre-processing.
[1167] Input: Personal information and preprocessed image data
[1168] Output: 3D model of the user
[1169] Step 5:
[1170] The server retrieves information about the clothing item the user has selected to try on from the database, including the clothing's texture, color, design, etc. This data is then optimized for the simulation.
[1171] Input: Clothing identification information
[1172] Output: Optimized clothing data
[1173] Step 6:
[1174] The server applies the optimized clothing data to the 3D model and performs a fitting simulation using a physics engine (e.g., PhysX), generating a fitting image that reproduces the fit and appearance of the clothing.
[1175] Input: 3D model and optimized clothing data
[1176] Output: Image after trying on
[1177] Step 7:
[1178] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback about the fit and appearance of the garment based on post-try-on images, which are then written in natural language and formatted as feedback to the user.
[1179] Input: Image after trying on
[1180] Output: Feedback on the generated fit and appearance
[1181] Step 8:
[1182] The server sends the generated post-try-on images and feedback to the user's device. The feedback consists of images and textual feedback.
[1183] Input: Images and impressions after trying on
[1184] Output: Feedback sent to the user's device
[1185] Step 9:
[1186] The user terminal displays the fitting results, and the user checks them. While the user is checking, the camera and microphone are used to capture facial expressions and voice data.
[1187] Input: Feedback
[1188] Output: Captured facial and voice data
[1189] Step 10:
[1190] The server sends the captured facial and voice data to an emotion recognition engine (such as Microsoft Azure's Emotion API) to recognize the user's emotions. The recognized emotions are analyzed in real time.
[1191] Input: Captured facial and voice data
[1192] Output: Recognized emotion data
[1193] Step 11:
[1194] The server adjusts the feedback based on the recognized emotion. For example, if the user is happy, it will emphasize the positive feedback and encourage them to buy the product. On the other hand, if a negative emotion is detected, it will adjust the feedback by suggesting a different product.
[1195] Input: Recognized emotion data
[1196] Output: Regulated Feedback
[1197] Step 12:
[1198] If the user is motivated to purchase after reviewing the adjusted feedback, they press the purchase button to complete the purchase. At this point, they are prompted to enter credit card information and a shipping address, and the order is confirmed once the entered data is sent to the server.
[1199] Input: Tailored feedback and checkout information
[1200] Output: Purchase confirmation notice
[1201] (Application example 2)
[1202] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1203] Traditional e-commerce sites face the challenge of making it difficult for users to check the size and fit of clothing when purchasing. As a result, products are often disappointing when they arrive, leading to frequent returns and exchanges. Another problem is the lack of feedback to encourage users to purchase. In particular, the lack of feedback that reflects users' emotions limits the effectiveness of these sites in promoting purchases.
[1204] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1205] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the user with the results of the try-on simulation and their impressions, and means for analyzing the user's emotions and adjusting the feedback content based on the recognized emotions. This allows the user to virtually try on clothing that fits their body, and by providing feedback based on their emotions, it is possible to achieve a greater purchasing promotion effect.
[1206] "User" refers to an individual who uses the System to virtually try on clothing items.
[1207] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[1208] "Image" refers to photographic data capturing the user's entire body.
[1209] "3D model" refers to a three-dimensional virtual shape of a user that is generated based on the received personal information and image.
[1210] "Clothing data" refers to data that includes information such as the texture, color, and design of clothing fabrics.
[1211] "Try-on simulation" refers to the process of applying clothing data to a generated 3D model to perform a virtual try-on.
[1212] "Try-on simulation results" refers to the visual results and feedback obtained after running a try-on simulation.
[1213] "Opinion" refers to an evaluation provided on fit and appearance based on the results of a try-on simulation.
[1214] "Analyzing emotions" refers to the process of recognizing emotions from the user's facial expressions and voice.
[1215] "Adjusting feedback content" refers to the process of changing the content of the feedback provided based on perceived emotions.
[1216] MODE FOR CARRYING OUT THE INVENTION
[1217] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[1218] Data Entry Phase
[1219] First, the user accesses an e-commerce site or a dedicated app using their own device (smartphone or PC) and uploads a full-body image along with personal information such as their height, weight, body type, and hairstyle. They enter this information through the app's interface and press the send button. The device then sends the entered personal information and image to the server. The data is sent to a specified API endpoint.
[1220] Data analysis phase
[1221] The server performs the following processes based on the received personal information and images:
[1222] 1. Image preprocessing: Use OpenCV to cut out the background and extract only the user's body. Then, resize the image to the appropriate size and apply a noise reduction filter.
[1223] 2. 3D model generation: A 3D modeling algorithm using TensorFlow generates a 3D model of the user, incorporating preprocessed image information.
[1224] 3. Clothing data acquisition: Data (fabric texture, color, design) of the clothing item selected by the user to try on is acquired from a database (e.g., AWS DynamoDB). The acquired data is optimized for simulation.
[1225] 4. Try-on simulation: Using a physics engine, clothing data is applied to a 3D model of the user to simulate a try-on, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[1226] Feedback Phase
[1227] The server uses the AI model to automatically generate feedback about the fit and appearance of the garment based on the generated post-try-on images. These feedback provide feedback as if the user had actually tried the garment on. The server then formats the try-on images and feedback and sends them to the user's device.
[1228] Emotion Recognition Phase
[1229] The server analyzes the user's facial expressions and voice while checking the fitting results, and uses an emotion engine to recognize the user's emotions. For example, it can recognize happiness, surprise, dissatisfaction, etc. from facial expressions captured by a camera and voice captured by a microphone. It then adjusts the feedback it provides based on the recognized emotions. If the user is happy, it emphasizes positive feedback and encourages the user to purchase the product. On the other hand, if negative emotions are detected, it makes adjustments such as suggesting a different product.
[1230] Purchasing Phase (Optional)
[1231] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[1232] Specific examples
[1233] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[1234] Prompt Sentence Examples
[1235] User A: Height 170cm, weight 65kg, slim build, short hair. Considering buying a blue T-shirt. Please generate a try-on simulation result and feedback.
[1236] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[1237] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1238] Step 1:
[1239] Users use their own devices to upload personal information such as height, weight, body type, hairstyle, and a full-body image. The entered data is sent from the device to the server.
[1240] Step 2:
[1241] The server performs preprocessing on the received image. Specifically, it uses OpenCV to cut out the background of the image, remove noise, and extract only the user's body. This process outputs the user's outline data.
[1242] Step 3:
[1243] The server runs a 3D modeling algorithm using TensorFlow based on the preprocessed images and personal information to generate a 3D model of the user. The input data here is the user's physical features and contour data, and the output data is the user's 3D model.
[1244] Step 4:
[1245] The server retrieves clothing data from a database (e.g., AWS DynamoDB), including the texture, color, and design of the clothing fabric, and outputs the data in a format optimized for simulation.
[1246] Step 5:
[1247] The server uses a physics engine to simulate the user trying on the clothes based on their 3D model and clothing data. As a result of the simulation, visual image data of the clothes after trying them on is generated.
[1248] Step 6:
[1249] The server uses an AI model (NLP model) to generate automatically generated impressions about fit and appearance based on the image data after trying on. The input is the image data after trying on, and the output is a textual impression about fit and appearance.
[1250] Step 7:
[1251] The server formats the image data after trying on the clothes and the created feedback, and sends it to the user's device. The input is the image data after trying on the clothes and the feedback, and the output is feedback data in a format that the user can check.
[1252] Step 8:
[1253] When the user checks the fitting results, the device's camera captures the user's facial expression and the device's microphone captures voice data, which is then sent to the server in real time.
[1254] Step 9:
[1255] The server analyzes the user's facial expressions and voice data using an emotion engine, recognizes the user's emotions (happiness, surprise, dissatisfaction, etc.), and outputs this emotional data.
[1256] Step 10:
[1257] The server adjusts the feedback based on the recognized emotion. For example, if a positive emotion is detected, it outputs positive feedback to encourage the purchase of a product, and if a negative emotion is detected, it suggests a different product.
[1258] Step 11:
[1259] The user decides whether to purchase the product based on the provided feedback. If the user decides to purchase the product, he or she presses the purchase button, the purchase procedure is executed on the server, and the purchase data is output.
[1260] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1261] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1262] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1263] [Fourth embodiment]
[1264] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1265] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1266] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1267] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1268] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1269] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1270] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1271] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1272] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1273] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1274] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1275] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1276] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1277] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system performs a try-on simulation based on personal information such as height, weight, figure, hairstyle, etc. provided by the user, as well as an image, and provides the results.
[1278] System Overview
[1279] The system mainly includes four main components:
[1280] 1. Data entry phase
[1281] 2. Data analysis phase
[1282] 3. Feedback Phase
[1283] 4. Purchasing Phase (Optional)
[1284] The specific processing of the program for each phase will be explained below.
[1285] Data Entry Phase
[1286] 1. User Actions
[1287] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[1288] 2. Processing performed by the server
[1289] The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. It also performs preprocessing on image data, such as cropping and resizing.
[1290] Data analysis phase
[1291] 1. Processing performed by the server: Image preprocessing
[1292] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[1293] 2. Processing performed by the server: 3D model generation
[1294] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[1295] 3. Processing performed by the server: Acquiring clothing data
[1296] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[1297] 4. Processing performed by the server: Try-on simulation
[1298] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[1299] Feedback Phase
[1300] 1. Server processing: Try-on images and impression generation
[1301] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[1302] 2. Server action: Sending feedback
[1303] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[1304] 3. User Action: Feedback Check
[1305] The user can then check the try-on images and their impressions and decide whether to purchase. If they like the item, they can proceed with the purchase.
[1306] Specific examples
[1307] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, integrates it with the data for the "blue T-shirt" that the user selected to try on, and performs a try-on simulation. Finally, the generated try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user.
[1308] In this way, users can choose products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress that comes with purchasing clothing on an e-commerce site.
[1309] The processing flow will be explained below.
[1310] Step 1:
[1311] User process: The user accesses an e-commerce site or a dedicated app from their own device. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[1312] Step 2:
[1313] Processing performed by the device: The entered personal information and image are sent to the server. By pressing the send button, the data is sent to the specified API endpoint.
[1314] Step 3:
[1315] Processing by the server: The server converts the received personal information and images into an appropriate format and stores them in a database. The images are pre-processed for analysis.
[1316] Step 4:
[1317] Server processing: Preprocesses the uploaded image by removing the background and extracting only the user's body. At this stage, the image is resized and a noise reduction filter is applied.
[1318] Step 5:
[1319] Server processing: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. This information is used to run a 3D modeling algorithm to create a three-dimensional digital model of the user.
[1320] Step 6:
[1321] Processing performed by the server: Retrieves data on the clothing item selected by the user from the database, including information on the clothing item's fabric texture, design, color, etc.
[1322] Step 7:
[1323] Server processing: The acquired clothing data is applied to the user's 3D model, and a fitting simulation is performed. A physics engine is used to simulate the fit, drape, wrinkles, etc. of the clothing.
[1324] Step 8:
[1325] Server processing: Analyzes visual information and automatically generates impressions along with the post-try-on images generated as a result of the simulation, creating text feedback on fit and appearance.
[1326] Step 9:
[1327] Server process: The generated try-on images and feedback are sent to the user's device. The data is formatted for the user interface and returned as an API response.
[1328] Step 10:
[1329] User action: Check the submitted try-on images and comments. Display the feedback on the e-commerce site or in the dedicated app and decide whether to purchase the product.
[1330] Step 11 (Optional):
[1331] What the user does: If they like the item, they press the "Purchase" button to complete the purchase. They can also try on other items if they wish.
[1332] This process allows users to simulate the fit and appearance of clothing through virtual try-on without actually trying it on. This system is expected to significantly reduce the anxiety and stress felt when purchasing clothing on an e-commerce site.
[1333] Example 1
[1334] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1335] Online shopping, especially when purchasing clothing, can cause anxiety and stress when users choose products without actually trying them on. Specifically, it is difficult for users to check in advance how the size, fit, and visual appearance of the clothing they choose will actually look like. Another issue is the time and effort required to visit a physical store to try on items.
[1336] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1337] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a three-dimensional model of the user, means for running a try-on simulation based on clothing data and the generated three-dimensional model, means for providing the user with the results of the try-on simulation and feedback, means for sending the generated feedback to the user's device, and means for the user to check the feedback on their own device. This allows users to select products through a virtual try-on experience without actually trying them on, reducing the anxiety and stress of online shopping.
[1338] "User" refers to an end user who uses the system to provide their personal information and images and simulate trying on clothing items.
[1339] "Personal information" refers to attribute information such as height, weight, body type, hairstyle, etc. provided by the user.
[1340] "Images" refers to photographs or videos of a user's entire body that they upload to the system.
[1341] "3D model" refers to a 3D virtual human model generated based on a user's personal information and image.
[1342] "Clothing data" refers to data including information such as the texture, color, and design of the clothing fabric, which is necessary for performing a fitting simulation.
[1343] "Try-on simulation" refers to the process of applying clothing data to a three-dimensional model of the user and virtually recreating the fit and appearance of the clothing using a physics engine or other means.
[1344] "Opinions" refer to evaluation comments about fit and appearance generated based on the results of the try-on simulation.
[1345] "Terminal" refers to an electronic device such as a smartphone or computer that a user uses to access the system.
[1346] "Feedback" refers to providing information to the user, including the results of the try-on simulation and the generated impressions.
[1347] "Generative AI model" refers to an artificial intelligence algorithm or program that automatically generates impressions about fit and appearance based on the results of a fitting simulation.
[1348] MODE FOR CARRYING OUT THE INVENTION
[1349] The present invention provides a system that allows users to virtually try on clothing when purchasing it online, and check the fit and appearance of the clothing after trying it on. This system is primarily built using the following hardware and software:
[1350] Hardware used
[1351] 1. Server: A high-performance server capable of processing and storing large amounts of data.
[1352] 2. Device: A device that can connect to the internet, such as a smartphone, tablet, or PC.
[1353] Software used
[1354] 1. Database: Database software (e.g., MySQL, PostgreSQL) for storing personal information, 3D models, and clothing data.
[1355] 2. Image processing software: Software for cropping, resizing, and noise removal of images (e.g., OpenCV).
[1356] 3. 3D modeling algorithms: Software used to generate a 3D model of the user (e.g. Blender API).
[1357] 4. Try-on simulation software: Software that applies clothing data to a three-dimensional model to simulate trying on the garment (e.g., Unity).
[1358] 5. Generative AI model: An AI model (e.g., a GPT-based model) that automatically generates impressions on fit and appearance based on the results of a try-on simulation.
[1359] System overview and specific processing
[1360] Data Entry Phase
[1361] Users access a dedicated app or e-commerce site using their own device, enter personal information such as their height, weight, body type, and hairstyle, along with a full-body image, and submit it. The server converts the received data into an appropriate format and stores it in a database. The image data undergoes pre-processing such as cropping, resizing, and noise removal.
[1362] Data analysis phase
[1363] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It also resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information to generate a 3D model, which also incorporates the preprocessed image information. It then retrieves data for the selected clothing item from a database and converts it into a format optimized for simulation.
[1364] Try-on simulation phase
[1365] The server applies the clothing data to the generated 3D model of the user and performs a fitting simulation using a physics engine, thereby reproducing the fit and visual appearance of the clothing and generating the final after-fit image.
[1366] Feedback Phase
[1367] The server uses a generative AI model to automatically generate feedback on the fit and appearance of the clothing based on the generated try-on images. The try-on images and feedback are then sent to the user's device, where the user can review the feedback and decide whether to purchase the clothing.
[1368] Specific examples
[1369] For example, if a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair, they can take a full-body photo with their smartphone and upload it to the system. The server generates a three-dimensional model of the user based on the received information, obtains data on the "blue T-shirt" selected for trying on, and performs a try-on simulation. Finally, the generated try-on image and a user's feedback, such as "It fits perfectly and the color is nice," are sent to the user's device.
[1370] Prompt Sentence Examples
[1371] "The user is 170cm tall, weighs 65kg, has a slim build, and has short hair. Please simulate trying on the 'blue T-shirt' selected by this user and generate impressions about the fit and appearance."
[1372] By building a system like this, users can check the fit and appearance of clothing through virtual try-ons, allowing them to enjoy online shopping with peace of mind.
[1373] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1374] Step 1:
[1375] User Action: Data Entry
[1376] Users access a dedicated app or e-commerce site using a device such as a smartphone or PC, where they enter personal information such as their height, weight, body type, and hairstyle, and take and upload a full-body image. The input data is then sent to the server.
[1377] Input: User's personal information (height, weight, body type, hairstyle) and full-body image
[1378] Output: Formatted personal information and image data sent to the server
[1379] Step 2:
[1380] Server operations: receiving and storing data
[1381] The server receives the personal information and image data sent by the user, converts the received data into an appropriate format (e.g., JSON format), and stores it in a database.
[1382] Input: Personal information and image data sent by the user
[1383] Output: Formatted personal information and image data stored in a database
[1384] Step 3:
[1385] Processing performed by the server: Image preprocessing
[1386] The server retrieves the image stored in the database, first cuts out the background and extracts only the user's body, then resizes the image to the appropriate size and applies a noise reduction filter to improve the image quality.
[1387] Input: Image data stored in a database
[1388] Output: Preprocessed image data
[1389] Step 4:
[1390] Server processing: 3D model generation
[1391] The server runs a 3D modeling algorithm based on the user's personal information (height, weight, body type) and preprocessed image data to generate a 3D model of the user, using the Blender API, for example.
[1392] Input: Personal information such as height, weight, and body shape, and preprocessed image data
[1393] Output: Generated 3D model of the user
[1394] Step 5:
[1395] Server process: Clothing data acquisition
[1396] The server retrieves data about the clothing item selected by the user from a database and converts it into a format optimized for simulation, including the texture, color, and design of the fabric.
[1397] Input: The identity of the clothing item selected by the user
[1398] Output: Optimized clothing data
[1399] Step 6:
[1400] Processing performed by the server: Try-on simulation
[1401] The server then applies the acquired clothing data to the generated 3D model of the user. The fitting simulation is performed using Unity's physics engine to reproduce the fit and visual appearance of the clothing. Finally, an image of the user after trying it on is generated.
[1402] Input: Generated 3D model and optimized clothing data
[1403] Output: Image after trying on
[1404] Step 7:
[1405] Server processing: Impression generation
[1406] The server then uses a generative AI model, such as a GPT-based model, to automatically generate feedback on the fit and appearance of the garment based on the generated try-on images.
[1407] Input: Image after trying on
[1408] Output: Generated impressions (text format)
[1409] Step 8:
[1410] Server action: Send feedback
[1411] The server then sends the generated post-try-on images and feedback to the user's device, allowing the user to view the feedback on their own device.
[1412] Input: Image after try-on and generated impressions
[1413] Output: Feedback sent to the user's device
[1414] Step 9:
[1415] User Action: Check Feedback
[1416] The user checks the images and feedback received on their device after trying on the item. Based on the feedback, the user decides whether to purchase the item, and if they do proceed with the purchase, they complete it on the e-commerce site.
[1417] Input: Try-on images and impressions received on the device
[1418] Output: Purchase decision and checkout
[1419] (Application example 1)
[1420] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1421] On conventional e-commerce sites, when users purchase clothing, they are unable to try it on, which can lead to concerns about fit and appearance. Furthermore, users are often unable to actually try on clothing before purchasing, which often results in the hassle and expense of returning the product. There is a need to improve this situation and provide a system that allows users to comfortably purchase clothing.
[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1423] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the results of the try-on simulation and impressions, means for providing feedback using the generated AI model, and means for outputting the feedback content as prompt text. This allows users to select products through a virtual try-on experience without actually trying them on, reducing anxiety and stress and making it possible to reduce the effort and cost of returning items.
[1424] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[1425] "Image" refers to photographs or video data that show the user's entire body.
[1426] A "3D model" is a three-dimensional virtual model generated based on personal information and image data, which mimics the user's body.
[1427] "Clothing data" is information about the texture, color, design, etc. of the fabric of the clothing to be tried on.
[1428] "Try-on simulation" is a process in which clothing data is applied to a 3D model of the user, allowing them to virtually try on the clothes.
[1429] A "generative AI model" is an algorithm that uses artificial intelligence to automatically generate feedback and impressions.
[1430] "Feedback" refers to comments and evaluations provided to users by the generative AI model based on the results of the fitting simulation.
[1431] A "prompt" is a document of instructions or guidance that a generative AI model outputs to a user.
[1432] A description will be given of an embodiment of the present invention. A system for implementing the present invention includes the following main phases: a data input phase, a data analysis phase, a feedback phase, and a purchasing phase (optional).
[1433] Data Entry Phase
[1434] Users access an e-commerce site or a dedicated app from their smartphone or PC, enter their personal information such as height, weight, body type, hairstyle, and a full-body image, and then submit the data. The server receives the data sent by the user, converts it into an appropriate format (e.g., JSON format), and stores it in a database. The server also performs preprocessing on the image data, such as cropping and resizing.
[1435] Data analysis phase
[1436] The server analyzes the uploaded image, removes the background, and extracts only the user's body. It then resizes the image to the appropriate size and applies a noise reduction filter. It then runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body shape, to generate a 3D model of the user. This 3D model also incorporates the preprocessed image information.
[1437] The server also retrieves data (fabric texture, color, design) from the database for the clothing item the user selects to try on. Based on this data, the server runs a fitting simulation and uses a physics engine to reproduce the fit and visual appearance of the clothing.
[1438] Feedback Phase
[1439] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated post-try-on images. This feedback is output as prompt text and sent to the user's device, allowing the user to check the fitting simulation results and their feedback on their own device.
[1440] Hardware and software used
[1441] 1. Hardware: Smartphone, personal computer (PC), server, head-mounted display (HMD)
[1442] 2. Software: Flask (web framework), OpenCV (image processing library), machine learning libraries (such as scikit-learn), generative AI models (for natural language generation)
[1443] Specific examples
[1444] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. When the user takes a full-body photo with their smartphone and uploads it to the system, the server generates a 3D model of the user based on the received information. The server then combines this with data on the "blue T-shirt" the user selected to try on, and performs a try-on simulation. The final try-on image and user feedback, such as "The fit is perfect and the color is nice," are sent to the user's device. In this way, the user can virtually choose products without actually trying them on.
[1445] Prompt Sentence Examples
[1446] "If the user is 170cm tall, weighs 65kg, has a slim build, and has short hair, generate a fitting simulation of the blue T-shirt that this user wants to try on."
[1447] This allows users to purchase clothing while reducing anxiety and stress.
[1448] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1449] Step 1:
[1450] Users access an e-commerce site or dedicated app using their smartphone or PC, enter personal information such as their height, weight, body type, hairstyle, and a full-body image, and then submit it. Specifically, after entering their personal information and image, the user presses the send button to send the data to the server. Input is via a text input form and an image upload function, and output is data in JSON format.
[1451] Step 2:
[1452] The server receives the JSON data sent by the user and analyzes each field. The received personal information and image data are saved in a database. At this time, the image data is preprocessed by cropping and resizing it and converting it into an appropriate format. Specifically, the image data is cropped using OpenCV and a noise reduction filter is applied. The output is the preprocessed image data and the analyzed personal information.
[1453] Step 3:
[1454] The server analyzes the preprocessed image, removes the background, and extracts only the user's body. It then resizes the image to an appropriate size and applies the noise reduction filter again. This step produces an image that extracts only the user's key physical characteristics. The input is the preprocessed image, and the output is an image of the user's body with the background removed.
[1455] Step 4:
[1456] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user. This model also incorporates preprocessed image information. Specifically, the 3D modeling software is used to recreate the user's three-dimensional body profile. The input is the analyzed personal information and preprocessed images, and the output is the user's 3D model.
[1457] Step 5:
[1458] The server retrieves data (fabric texture, color, design) of the clothing item selected by the user from the database. Specifically, it executes a database query to retrieve optimized clothing data. The input is the clothing item ID selected by the user, and the output is the corresponding clothing item data.
[1459] Step 6:
[1460] The server applies clothing data to the generated 3D model and performs a fitting simulation. This simulation uses a physics engine to reproduce the fit and visual appearance of the clothing, and generates a final image of the clothing being tried on. Specifically, the server uses simulation software to perform a dynamic simulation of the clothing. The input is the 3D model and clothing data, and the output is a simulated image of the clothing being tried on.
[1461] Step 7:
[1462] The server uses a generative AI model to automatically generate feedback on fit and appearance based on the generated try-on images. This feedback is generated through analysis by a natural language processing model. The input is the try-on image, and the output is the feedback content.
[1463] Step 8:
[1464] The server formats the feedback content as a prompt sentence and sends it to the user's device. Specifically, the server obtains the feedback content from the generative AI model and formats it into a prompt sentence format. The input is the generated feedback, and the output is feedback as a prompt sentence.
[1465] Step 9:
[1466] The user checks the try-on simulation results and their impressions sent to their own device. After checking, they can make a purchase decision. Specifically, they check the try-on images on their smartphone or PC screen and read the feedback. The input is the feedback and try-on images sent from the server, and the output is the user's purchasing decision.
[1467] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1468] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[1469] System Overview
[1470] The system mainly includes five main components:
[1471] 1. Data entry phase
[1472] 2. Data analysis phase
[1473] 3. Feedback Phase
[1474] 4. Emotion Recognition Phase
[1475] 5. Purchasing Phase (Optional)
[1476] The specific processing of the program for each phase will be explained below.
[1477] Data Entry Phase
[1478] 1. User Actions
[1479] Users access the e-commerce site or a dedicated app using their own device (smartphone or PC), upload personal information such as their height, weight, body type, hairstyle, and a full-body image, enter this information through the app interface, and press the send button.
[1480] 2. Processing performed by the device
[1481] The device sends the entered personal information and image to the server. By pressing the send button, the data is sent to the specified API endpoint.
[1482] Data analysis phase
[1483] 1. Processing performed by the server: Image preprocessing
[1484] The server analyzes the uploaded image, cuts out the background, extracts only the user's body, resizes the image to the appropriate size, and applies a noise reduction filter.
[1485] 2. Processing performed by the server: 3D model generation
[1486] The server runs a 3D modeling algorithm based on the user's personal information, such as height, weight, and body type, to generate a 3D model of the user, which also incorporates pre-processed image information.
[1487] 3. Processing performed by the server: Acquiring clothing data
[1488] The data (texture, color, design) of the clothing item that the user selects to try on is retrieved from the database. The retrieved data is in a format optimized for simulation.
[1489] 4. Processing performed by the server: Try-on simulation
[1490] The server applies the clothing data to the user's 3D model and uses a physics engine to simulate a fitting, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[1491] Feedback Phase
[1492] 1. Server processing: Try-on images and impression generation
[1493] Based on the generated try-on images, the AI model automatically generates feedback on fit and appearance, providing users with the same experience as if they were trying the garment on in real life.
[1494] 2. Server action: Sending feedback
[1495] The try-on images and feedback are formatted and sent to the user's device, allowing the user to view the feedback on their own device.
[1496] Emotion Recognition Phase
[1497] 1. Processing performed by the server: Execution of the emotion engine
[1498] The server analyzes the facial expressions and voices of the user while checking the fitting results and uses an emotion engine to recognize the user's emotions, such as joy, surprise, or dissatisfaction, based on facial expressions captured by a camera and voice recorded by a microphone.
[1499] 2. Server processing: Emotion-based feedback adjustment
[1500] The system adjusts the feedback provided based on the perceived emotions. For example, if the user is happy, it will emphasize the positive feedback and encourage them to purchase the product. On the other hand, if negative emotions are detected, it will adjust the feedback by suggesting a different product.
[1501] Purchasing Phase (Optional)
[1502] User actions
[1503] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[1504] Specific examples
[1505] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[1506] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[1507] The processing flow will be explained below.
[1508] Step 1:
[1509] User processing: The user accesses an e-commerce site or a dedicated app from their own device and logs in. After logging in, they enter personal information such as height, weight, body type, and hairstyle, and take and upload a full-body image.
[1510] Step 2:
[1511] Processing performed by the device: Sends the entered personal information and image to the server. By pressing the send button, the entered data and image are sent to the specified API endpoint.
[1512] Step 3:
[1513] Server processing: Receives the transmitted data, converts it into an appropriate format, and stores it in the database. For images, preprocessing involves cropping, resizing, and noise removal.
[1514] Step 4:
[1515] Processing performed by the server: The background is removed from the preprocessed image data, and only the user's body is extracted. This processing allows the necessary features to be extracted effectively.
[1516] Step 5:
[1517] Processing performed by the server: Generates a 3D model based on the user's height, weight, body type, and hairstyle information. Specifically, it runs a 3D modeling algorithm to create a three-dimensional digital model of the user.
[1518] Step 6:
[1519] Processing performed by the server: Retrieves data on the clothing item to be tried on from the database, including information on the clothing's fabric texture, design, color, etc.
[1520] Step 7:
[1521] Server processing: Applying clothing data to the user's 3D model and running a fitting simulation using a physics engine. This simulates the fit and visual appearance of the clothing in a realistic way.
[1522] Step 8:
[1523] Server processing: Based on the simulation results, automatically generate images and impressions after trying on the garment. Create text feedback including evaluations of fit and appearance.
[1524] Step 9:
[1525] Server process: The generated try-on images and feedback are sent to the user's device. The feedback is formatted and returned as an API response.
[1526] Step 10:
[1527] User action: Check the submitted try-on images and feedback and decide whether to purchase the product. Based on the feedback, the user can choose to either press the purchase button or try on other products.
[1528] Step 11:
[1529] Processing performed by the server: While the user is checking the feedback, the emotion engine analyzes the user's facial expressions and voice to recognize their emotions. Based on data acquired through the camera and microphone, emotions such as joy, surprise, and dissatisfaction are determined.
[1530] Step 12:
[1531] Processing performed by the server: The content of the feedback is adjusted based on the user's recognized emotions. For example, if a positive emotion is detected, a message to encourage purchase is displayed. If a negative emotion is detected, reference information or suggestions for other products are provided.
[1532] Step 13:
[1533] What happens to the user: They review the tailored feedback and make a final purchase decision. They can either click the buy button to proceed with the purchase or try on other items.
[1534] This process allows users to check the fit and appearance of clothing through virtual try-on, and the emotion engine provides more personalized feedback, reducing anxiety when purchasing clothing on e-commerce sites and increasing user motivation.
[1535] Example 2
[1536] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1537] When purchasing clothing on traditional e-commerce sites, users are unable to try on items, which can often lead to anxiety when making a purchase. Furthermore, a simple virtual try-on simulation makes it difficult to accurately grasp how users will feel, and it is unable to provide emotional feedback, which fails to increase purchase motivation. Furthermore, it is difficult to provide personalized feedback based on each user's individual body type and preferences.
[1538] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1539] In this invention, the server includes: means for receiving personal information such as height, weight, body type, and hairstyle, as well as an image, from the user; means for processing the received personal information and image to generate a three-dimensional model of the user; and means for running a try-on simulation based on clothing information and the generated three-dimensional model. This allows the user to check the fit and appearance of clothing through a virtual try-on without actually visiting a store. The server also includes means for providing the user with the results of the try-on simulation and their impressions, means for recognizing the user's emotions, and means for adjusting feedback based on the recognized emotions. This allows for feedback that reflects the user's emotional state, enabling a more personalized shopping experience.
[1540] "Personal information" refers to data that indicates individual characteristics of a user, such as height, weight, body type, hairstyle, etc.
[1541] A "three-dimensional model" is a digital replica of a user's body shape and features, generated based on the user's personal information and image.
[1542] "Try-on simulation" is a simulation technology that applies clothing information to a generated three-dimensional model to reproduce the visual appearance and fit of the garment as if it were actually being tried on.
[1543] "Clothing information" refers to data that includes characteristics of clothing, such as texture, color, and design, and that is applied to a three-dimensional model.
[1544] "Feedback" refers to information and evaluations provided to users, including results of try-on simulations and impressions.
[1545] "Emotion recognition" is a technology that analyzes and recognizes a user's emotional state (happiness, dissatisfaction, etc.) from their facial expressions and voice.
[1546] "Feedback modulation" means changing the content of feedback based on perceived emotions.
[1547] This invention is a system that supports virtual try-on of clothing when users purchase it on an e-commerce website. The system receives personal information (height, weight, body type, hairstyle, etc.) and images from the user, generates a three-dimensional model based on the information, and then performs a try-on simulation based on the clothing information, providing the results and user feedback. It also recognizes the user's emotions and adjusts the feedback based on those emotions, providing a more personalized shopping experience.
[1548] Overall system flow
[1549] Data Entry Phase
[1550] The user accesses the e-commerce site or a dedicated app using a device (smartphone or PC), enters personal information such as their height, weight, body type, hairstyle, etc., along with a full-body image, and presses the send button. The device then sends this information to the server. The transmission is encrypted using SSL / TLS.
[1551] Data analysis phase
[1552] The server first preprocesses the received image by using an algorithm (e.g., GrabCut) to remove the background and extract only the user's body. It also resizes the image if necessary and applies a noise reduction filter (e.g., a Gaussian filter).
[1553] Next, the server runs a 3D modeling algorithm (e.g., Blender's API) based on the user's personal information such as height, weight, and body type to generate a 3D model of the user.
[1554] The server then retrieves the specified clothing information (texture, color, design, etc.) from the database, applies the clothing information to the user's 3D model, and performs a fitting simulation using a physics engine (e.g., PhysX). The fitting simulation reproduces the fit and visual appearance of the clothing, and generates a final try-on image.
[1555] Feedback Phase
[1556] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback on the fit and appearance of the garment based on the post-try-on images. The generated try-on images and feedback are then formatted and sent to the user's device, where the user can view the results.
[1557] Emotion Recognition Phase
[1558] The server analyzes the user's facial expressions and voice in real time while checking the fitting results. Using facial expression data captured by the camera and voice data obtained by the microphone, an emotion recognition engine (such as Microsoft Azure's Emotion API) is used to recognize the user's emotional state. Based on the recognized emotion, the server adjusts the content of the feedback. For example, if the user is pleased, it will emphasize positive feedback to encourage the purchase.
[1559] Purchasing Phase (Optional)
[1560] The user checks the feedback and if they like the product, they press the purchase button to complete the purchase. They are then prompted to enter their credit card information and shipping address, and once these are entered, the order is confirmed.
[1561] Specific examples
[1562] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a three-dimensional model of the user based on the received information, retrieves information about a "blue shirt" from the database, and simulates trying it on. A physics engine is activated to simulate how the shirt will fit the user, and generates an image of the user after trying it on. The generated image and a comment such as "It fits perfectly and the color is nice" are sent to the user.
[1563] Next, the camera captures the user's facial expressions as they check the fitting results, and if the emotion engine recognizes their joy, it provides positive feedback such as, "This shirt looks great on you!"
[1564] Prompt Sentence Examples
[1565] "The user is 170cm tall, 65kg, slim, with short hair. He is trying on a blue shirt. Generate the results of the fitting simulation."
[1566] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1567] Step 1:
[1568] A user accesses an e-commerce site or a dedicated app from their own device (smartphone or PC) and enters their personal information (height, weight, body type, hairstyle, etc.) and a full-body image. The entered data is validated within the device to ensure that all required information has been entered. The input data is often in JSON format.
[1569] Input: Personal information and full-body image entered by the user
[1570] Output: Verified personal information and image data
[1571] Step 2:
[1572] The device sends the entered personal information and image to the API endpoint. The data is encrypted with SSL / TLS and sent securely to the server.
[1573] Input: Verified personal information and image data
[1574] Output: Data sent to the API endpoint
[1575] Step 3:
[1576] The server analyzes the received image data and performs image preprocessing. Specifically, it applies the GrabCut algorithm to remove the background and extract the user's body parts. Then, it resizes the image and removes noise using a Gaussian filter as needed.
[1577] Input: Image data sent
[1578] Output: Image data of the body part with the background removed
[1579] Step 4:
[1580] Based on the personal information received, the server runs a 3D modeling algorithm (e.g., Blender API) to generate a 3D model of the user. This step also integrates data obtained from image pre-processing.
[1581] Input: Personal information and preprocessed image data
[1582] Output: 3D model of the user
[1583] Step 5:
[1584] The server retrieves information about the clothing item the user has selected to try on from the database, including the clothing's texture, color, design, etc. This data is then optimized for the simulation.
[1585] Input: Clothing identification information
[1586] Output: Optimized clothing data
[1587] Step 6:
[1588] The server applies the optimized clothing data to the 3D model and performs a fitting simulation using a physics engine (e.g., PhysX), generating a fitting image that reproduces the fit and appearance of the clothing.
[1589] Input: 3D model and optimized clothing data
[1590] Output: Image after trying on
[1591] Step 7:
[1592] The server uses a generative AI model (e.g., GPT-3) to automatically generate feedback about the fit and appearance of the garment based on post-try-on images, which are then written in natural language and formatted as feedback to the user.
[1593] Input: Image after trying on
[1594] Output: Feedback on the generated fit and appearance
[1595] Step 8:
[1596] The server sends the generated post-try-on images and feedback to the user's device. The feedback consists of images and textual feedback.
[1597] Input: Images and impressions after trying on
[1598] Output: Feedback sent to the user's device
[1599] Step 9:
[1600] The user terminal displays the fitting results, and the user checks them. While the user is checking, the camera and microphone are used to capture facial expressions and voice data.
[1601] Input: Feedback
[1602] Output: Captured facial and voice data
[1603] Step 10:
[1604] The server sends the captured facial and voice data to an emotion recognition engine (such as Microsoft Azure's Emotion API) to recognize the user's emotions. The recognized emotions are analyzed in real time.
[1605] Input: Captured facial and voice data
[1606] Output: Recognized emotion data
[1607] Step 11:
[1608] The server adjusts the feedback based on the recognized emotion. For example, if the user is happy, it will emphasize the positive feedback and encourage them to buy the product. On the other hand, if a negative emotion is detected, it will adjust the feedback by suggesting a different product.
[1609] Input: Recognized emotion data
[1610] Output: Regulated Feedback
[1611] Step 12:
[1612] If the user is motivated to purchase after reviewing the adjusted feedback, they press the purchase button to complete the purchase. At this point, they are prompted to enter credit card information and a shipping address, and the order is confirmed once the entered data is sent to the server.
[1613] Input: Tailored feedback and checkout information
[1614] Output: Purchase confirmation notice
[1615] (Application example 2)
[1616] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1617] Traditional e-commerce sites face the challenge of making it difficult for users to check the size and fit of clothing when purchasing. As a result, products are often disappointing when they arrive, leading to frequent returns and exchanges. Another problem is the lack of feedback to encourage users to purchase. In particular, the lack of feedback that reflects users' emotions limits the effectiveness of these sites in promoting purchases.
[1618] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1619] In this invention, the server includes means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from the user, means for processing the received personal information and image and generating a 3D model of the user, means for running a try-on simulation based on clothing data and the generated 3D model, means for providing the user with the results of the try-on simulation and their impressions, and means for analyzing the user's emotions and adjusting the feedback content based on the recognized emotions. This allows the user to virtually try on clothing that fits their body, and by providing feedback based on their emotions, it is possible to achieve a greater purchasing promotion effect.
[1620] "User" refers to an individual who uses the System to virtually try on clothing items.
[1621] "Personal information" refers to information about a user's physical characteristics, such as height, weight, body type, and hairstyle.
[1622] "Image" refers to photographic data capturing the user's entire body.
[1623] "3D model" refers to a three-dimensional virtual shape of a user that is generated based on the received personal information and image.
[1624] "Clothing data" refers to data that includes information such as the texture, color, and design of clothing fabrics.
[1625] "Try-on simulation" refers to the process of applying clothing data to a generated 3D model to perform a virtual try-on.
[1626] "Try-on simulation results" refers to the visual results and feedback obtained after running a try-on simulation.
[1627] "Opinion" refers to an evaluation provided on fit and appearance based on the results of a try-on simulation.
[1628] "Analyzing emotions" refers to the process of recognizing emotions from the user's facial expressions and voice.
[1629] "Adjusting feedback content" refers to the process of changing the content of the feedback provided based on perceived emotions.
[1630] MODE FOR CARRYING OUT THE INVENTION
[1631] This invention provides a system that supports virtual try-on of clothing when users purchase it on an e-commerce site. This system not only performs a try-on simulation based on personal information such as height, weight, figure, and hairstyle provided by the user, as well as an image, but also recognizes the user's emotions regarding the try-on results by combining it with an emotion engine, improving feedback.
[1632] Data Entry Phase
[1633] First, the user accesses an e-commerce site or a dedicated app using their own device (smartphone or PC) and uploads a full-body image along with personal information such as their height, weight, body type, and hairstyle. They enter this information through the app's interface and press the send button. The device then sends the entered personal information and image to the server. The data is sent to a specified API endpoint.
[1634] Data analysis phase
[1635] The server performs the following processes based on the received personal information and images:
[1636] 1. Image preprocessing: Use OpenCV to cut out the background and extract only the user's body. Then, resize the image to the appropriate size and apply a noise reduction filter.
[1637] 2. 3D model generation: A 3D modeling algorithm using TensorFlow generates a 3D model of the user, incorporating preprocessed image information.
[1638] 3. Clothing data acquisition: Data (fabric texture, color, design) of the clothing item selected by the user to try on is acquired from a database (e.g., AWS DynamoDB). The acquired data is optimized for simulation.
[1639] 4. Try-on simulation: Using a physics engine, clothing data is applied to a 3D model of the user to simulate a try-on, recreating the fit and visual appearance of the clothing and generating the final try-on image.
[1640] Feedback Phase
[1641] The server uses the AI model to automatically generate feedback about the fit and appearance of the garment based on the generated post-try-on images. These feedback provide feedback as if the user had actually tried the garment on. The server then formats the try-on images and feedback and sends them to the user's device.
[1642] Emotion Recognition Phase
[1643] The server analyzes the user's facial expressions and voice while checking the fitting results, and uses an emotion engine to recognize the user's emotions. For example, it can recognize happiness, surprise, dissatisfaction, etc. from facial expressions captured by a camera and voice captured by a microphone. It then adjusts the feedback it provides based on the recognized emotions. If the user is happy, it emphasizes positive feedback and encourages the user to purchase the product. On the other hand, if negative emotions are detected, it makes adjustments such as suggesting a different product.
[1644] Purchasing Phase (Optional)
[1645] Users can check the feedback and if they like the product, they can press the purchase button to complete the purchase process. If they do not purchase, they can try on other products.
[1646] Specific examples
[1647] For example, suppose a user is 170 cm tall, weighs 65 kg, has a slim build, and has short hair. The user takes a full-body photo with their smartphone and uploads it to the system. The server generates a 3D model of the user based on the received information, and integrates it with the data for the "blue T-shirt" selected for trying on to simulate a try-on. The try-on image and a comment such as "The fit is perfect and the color is nice" are sent to the user. Furthermore, the camera captures the user's facial expression when checking the try-on results, and if the emotion engine recognizes the user's joy, it provides positive feedback to encourage purchase.
[1648] Prompt Sentence Examples
[1649] User A: Height 170cm, weight 65kg, slim build, short hair. Considering buying a blue T-shirt. Please generate a try-on simulation result and feedback.
[1650] In this way, users can choose products through virtual try-on without visiting a physical store, and emotion recognition provides a more personalized purchasing experience, while e-commerce site operators can expect higher customer satisfaction and purchase rates.
[1651] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1652] Step 1:
[1653] Users use their own devices to upload personal information such as height, weight, body type, hairstyle, and a full-body image. The entered data is sent from the device to the server.
[1654] Step 2:
[1655] The server performs preprocessing on the received image. Specifically, it uses OpenCV to cut out the background of the image, remove noise, and extract only the user's body. This process outputs the user's outline data.
[1656] Step 3:
[1657] The server runs a 3D modeling algorithm using TensorFlow based on the preprocessed images and personal information to generate a 3D model of the user. The input data here is the user's physical features and contour data, and the output data is the user's 3D model.
[1658] Step 4:
[1659] The server retrieves clothing data from a database (e.g., AWS DynamoDB), including the texture, color, and design of the clothing fabric, and outputs the data in a format optimized for simulation.
[1660] Step 5:
[1661] The server uses a physics engine to simulate the user trying on the clothes based on their 3D model and clothing data. As a result of the simulation, visual image data of the clothes after trying them on is generated.
[1662] Step 6:
[1663] The server uses an AI model (NLP model) to generate automatically generated impressions about fit and appearance based on the image data after trying on. The input is the image data after trying on, and the output is a textual impression about fit and appearance.
[1664] Step 7:
[1665] The server formats the image data after trying on the clothes and the created feedback, and sends it to the user's device. The input is the image data after trying on the clothes and the feedback, and the output is feedback data in a format that the user can check.
[1666] Step 8:
[1667] When the user checks the fitting results, the device's camera captures the user's facial expression and the device's microphone captures voice data, which is then sent to the server in real time.
[1668] Step 9:
[1669] The server analyzes the user's facial expressions and voice data using an emotion engine, recognizes the user's emotions (happiness, surprise, dissatisfaction, etc.), and outputs this emotional data.
[1670] Step 10:
[1671] The server adjusts the feedback based on the recognized emotion. For example, if a positive emotion is detected, it outputs positive feedback to encourage the purchase of a product, and if a negative emotion is detected, it suggests a different product.
[1672] Step 11:
[1673] The user decides whether to purchase the product based on the provided feedback. If the user decides to purchase the product, he or she presses the purchase button, the purchase procedure is executed on the server, and the purchase data is output.
[1674] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1675] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1676] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1677] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1678] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1679] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1680] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1681] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1682] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1683] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1684] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1685] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1686] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1687] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1688] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1689] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1690] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1691] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1692] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1693] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1694] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1695] The following is further disclosed regarding the above embodiment.
[1696] (Claim 1)
[1697] A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user;
[1698] means for processing the received personal information and images to generate a 3D model of the user;
[1699] A means for performing a fitting simulation based on the clothing data and the generated 3D model;
[1700] The system includes a means for providing a user with the results of a try-on simulation and their impressions.
[1701] (Claim 2)
[1702] 10. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
[1703] (Claim 3)
[1704] 2. The system according to claim 1, further comprising means for automatically generating impressions regarding fit and appearance based on the results of the fitting simulation.
[1705] "Example 1"
[1706] (Claim 1)
[1707] A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user;
[1708] means for processing the received personal information and images to generate a three-dimensional model of the user;
[1709] a means for executing a fitting simulation based on the clothing data and the generated three-dimensional model;
[1710] A means for providing a user with a fitting simulation result and an impression;
[1711] means for transmitting the generated impressions to a user's terminal;
[1712] A system that includes a means for users to view feedback on their own devices.
[1713] (Claim 2)
[1714] 10. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
[1715] (Claim 3)
[1716] 10. The system of claim 1, further comprising means for using a generative AI model to automatically generate impressions regarding fit and appearance based on the results of the fitting simulation.
[1717] "Application Example 1"
[1718] (Claim 1)
[1719] A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user;
[1720] means for processing the received personal information and images to generate a 3D model of the user;
[1721] A means for performing a fitting simulation based on the clothing data and the generated 3D model;
[1722] A means for providing a user with a fitting simulation result and an impression;
[1723] a means of providing feedback using a generative AI model; and
[1724] A means of outputting feedback as a prompt
[1725] A system including:
[1726] (Claim 2)
[1727] 10. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
[1728] (Claim 3)
[1729] 2. The system according to claim 1, further comprising means for automatically generating impressions regarding fit and appearance based on the results of the fitting simulation.
[1730] "Example 2: Combining Emotion Engines"
[1731] (Claim 1)
[1732] A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user;
[1733] means for processing the received personal information and images to generate a three-dimensional model of the user;
[1734] means for executing a try-on simulation based on the clothing information and the generated three-dimensional model;
[1735] A means for providing a user with a fitting simulation result and an impression;
[1736] a means for recognizing a user's emotion;
[1737] a means for adjusting feedback based on perceived emotions;
[1738] A system including:
[1739] (Claim 2)
[1740] 10. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
[1741] (Claim 3)
[1742] 2. The system according to claim 1, further comprising means for analyzing facial expressions and voices of the user while checking the fitting results and executing an emotion engine.
[1743] (Claim 4)
[1744] 2. The system according to claim 1, further comprising means for automatically generating impressions regarding fit and appearance based on the results of the fitting simulation.
[1745] "Application example 2 when combining emotion engines"
[1746] (Claim 1)
[1747] A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user;
[1748] means for processing the received personal information and images to generate a 3D model of the user;
[1749] A means for performing a fitting simulation based on the clothing data and the generated 3D model;
[1750] A means for providing a user with a fitting simulation result and an impression;
[1751] The system includes a means for analyzing a user's emotions and adjusting feedback content based on the recognized emotions.
[1752] (Claim 2)
[1753] 10. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
[1754] (Claim 3)
[1755] 2. The system according to claim 1, further comprising means for automatically generating impressions regarding fit and appearance based on the results of the fitting simulation. [Explanation of symbols]
[1756] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving personal information such as height, weight, body type, hairstyle, etc. and an image from a user; means for processing the received personal information and images to generate a 3D model of the user; A means for performing a fitting simulation based on the clothing data and the generated 3D model; The system includes a means for providing a user with the results of a try-on simulation and their impressions.
2. 2. The system of claim 1, further comprising means for pre-processing the received image to remove background and extract only the user's body.
3. The system according to claim 1, further comprising means for automatically generating impressions regarding fit and appearance based on the results of the fitting simulation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A