system
A system generates composite images of users wearing selected clothing, facilitating efficient and convenient clothing selection and fitting by combining facial images with clothing databases, reducing the need for physical try-ons and reservations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Users face a time-consuming and laborious process when selecting suitable clothing for events or ceremonies, making it difficult to efficiently choose from a large number of items and requiring significant effort for actual fitting.
A system that generates a composite image by combining a user's facial image with selected clothing from a database, allowing virtual try-on, reservation, and purchase/rental procedures, thereby streamlining the selection and fitting process.
Enables efficient and convenient selection and try-on of clothing, reducing user burden by allowing digital try-on and reservation without physical visits, and improving the overall fitting experience.
Smart Images

Figure 2026085786000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, when choosing special clothing for events or ceremonies, users are forced into a time-consuming and laborious fitting process. In this process, it is difficult to efficiently select a suitable item from a large number of clothing items, and the time and effort required for actual fitting pose a significant burden. Therefore, there is a need for a system that enables users to efficiently select clothing suitable for themselves from a variety of clothing items and smooths the process up to the fitting stage.
Means for Solving the Problems
[0005] This invention provides a system that enables digital costume try-on by generating a composite image based on the user's facial image. Specifically, the user uploads their own image data, and a composite image is generated by combining that data with a costume selected from a clothing database. The generated image is presented to the user, and the system supports the user in making a try-on reservation and purchasing / renting the selected costume. This allows for efficient costume selection and try-on procedures within a limited time, thereby reducing the burden on the user.
[0006] "User" is a concept that refers to an individual who uses the system to go through the process of selecting and trying on clothing.
[0007] "Image data" refers to visual information obtained from users for the purpose of generating composite images, and is digital data that is particularly used for facial images.
[0008] A "clothing database" refers to a storage system that accumulates information and visual data on various types of clothing, and is referenced when generating composite images.
[0009] A "composite image" refers to a digital image generated by combining the user's image data with the selected clothing, and is used by users to virtually try on clothing.
[0010] "Fitting reservation" refers to the process of securing a time and place for a user to actually try on the clothing they have selected.
[0011] "Purchase / rental procedures" refer to the commercial transactions and related procedures necessary for users to obtain the costumes they have selected. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0018] In the following embodiments, a communication I / F (Interface) with a reference numeral is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention is a system that enables users to efficiently select suitable clothing and smoothly proceed through the fitting process. This system primarily functions through the coordinated operation of three elements: a server, a terminal, and the user.
[0034] First, the user accesses the system, completes the necessary registration process, and then uploads a picture of their face. This image is used as the foundation for providing a digital costume try-on experience. The terminal receives the image uploaded by the user and sends it to the server.
[0035] The server processes the received facial image data, selects various outfits from a clothing database, and generates a composite image by combining the outfit with the user's face. This composite image allows the user to virtually try on many different outfits.
[0036] The generated composite image is provided to the user in a viewable format. The user can select an outfit that they think suits them through the composite image. Based on this result, the server saves the selected outfit information and manages the reservation process for actual try-ons, as well as the purchase or rental process for the outfit, if desired.
[0037] As a concrete example, let's consider the case of choosing a furisode (long-sleeved kimono) for a coming-of-age ceremony. In this case, the user wants to try on furisode in order to attend the ceremony. First, the user uploads a facial image to the system. The server generates composite images of the user wearing various furisode based on the user's facial image data and presents them to the user. The user then selects a furisode they like and makes a reservation to try it on. If necessary, they can make a payment on the spot and complete the furisode rental procedure.
[0038] This system allows users to easily find the outfit that suits them best and efficiently go through the entire process from trying it on to purchasing or renting it.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The user accesses the system and first registers an account. They enter the required information (name, email address, password, etc.) to create an account. The terminal collects this information and sends the account data to the server. The server stores the received information in its database and creates the account.
[0042] Step 2:
[0043] After registering an account, the user uploads a facial image. This process utilizes an image selection tool. The device sends the uploaded image data to the server. The server receives the image data and prepares it for image processing.
[0044] Step 3:
[0045] The server inputs the received facial image data into an AI model and generates a composite image by combining it with clothing selected from a clothing database. The AI model performs high-precision image processing to digitally composite various clothing onto the user's face.
[0046] Step 4:
[0047] The server sends the generated composite images to the terminal and presents them to the user in a viewable format. The user reviews these composite images and selects the outfit they like. The selection process is performed on the terminal, and the selection results are sent back to the server.
[0048] Step 5:
[0049] The server stores the selected costume information and provides an interface for making fitting reservations and purchasing / renting items. Users select a fitting date and time and make a reservation. The server processes this information and coordinates the schedule with partner sales / rental businesses.
[0050] Step 6:
[0051] Once the reservation is complete, the server sends a confirmation notice to the user. If the user wishes to purchase or rent, the server provides a secure payment environment and completes the transaction. The user receives confirmation that the process is complete and can then finish using the service.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] The process of selecting and trying on clothing requires time and effort for users to try on many options. Physical try-ons require travel to stores and making appointments, which is inconvenient. Furthermore, the limited variety of clothing available for try-on makes it difficult to make the best clothing choice.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes means for acquiring and storing the user's image information, means for analyzing visual features based on the acquired image information to select clothing information from a data set and generating a composite image using an image synthesis model, and means for presenting the generated composite image to the user and presenting selectable clothing information. As a result, the user can choose from a variety of clothing without physically trying them on, enabling them to make the optimal clothing selection while significantly reducing time and effort.
[0057] "User image information" refers to visual data that individuals provide to the system, and this data is used to analyze the characteristics of the users.
[0058] "Visual features" refer to information about appearance and shape extracted from image data, and represent the individual characteristics of the user necessary for generating a composite image.
[0059] A "data set" is a database containing diverse clothing information, used to select the most suitable outfit based on the user's image information.
[0060] "Costume information" refers to data related to clothing, including detailed information such as the design, color, and material of the garments.
[0061] An "image synthesis model" refers to an algorithm or program used to generate a realistic composite image by combining image information and clothing information acquired using AI technology.
[0062] A "composite image" is a simulated image created based on the user's image information and selected clothing information, visually representing the try-on state.
[0063] "To present" refers to the act of showing generated information or images to the user through a user interface, in order to enable them to make choices.
[0064] "Try-on procedure" refers to the process of making reservations and arrangements for users to actually try on the outfits they have selected.
[0065] "Commercial transaction procedures" refer to the series of activities involved in processing payments and contracts necessary when a user purchases or rents an outfit they have selected.
[0066] The following describes embodiments for carrying out this system invention.
[0067] The present invention provides a system for users to efficiently select suitable clothing digitally. This system primarily consists of three elements: the user, the terminal, and the server.
[0068] Users first access the system using their device and register by entering the required information. This includes providing basic personal information. Afterward, users use their device's camera function to take a picture of their face and upload it to the system. This face image serves as the basis for generating a composite image of the costume.
[0069] The device is equipped with the ability to receive facial images uploaded by users and quickly transmit them to the server. This communication process is supported by a high-speed and stable internet connection.
[0070] The server analyzes the facial image data received from the user using an "image processing API." Here, a specific technology called the "FaceMatch API" is used to extract the visual features of the face. Based on this, the server utilizes a "generative AI model" to select the most suitable outfit from a clothing database. The selected outfit is then combined with the user's facial image using a "DressAI model" to generate a composite image that provides a realistic try-on experience.
[0071] The device receives a composite image sent from the server and displays it to the user. This image allows the user to compare various outfits without actually trying them on. This process is performed using the device's high-resolution display, making it easy for the user to view the composite image.
[0072] As a concrete example, in the case of a user choosing a furisode (long-sleeved kimono) for their coming-of-age ceremony, the user makes a reservation to try on furisode for the ceremony. The user uploads a facial image, and after processing with an "image processing API" and a "generative AI model," receives a composite image of themselves wearing various furisode. Based on this image, the user selects the most suitable furisode and makes a reservation to try it on from their device. Furthermore, when purchasing or renting, credit card payment can also be made via the device.
[0073] In this way, users can intuitively select costumes even from remote locations and efficiently proceed with costume selection, purchase, or rental procedures.
[0074] An example of a prompt is: "Use the user's facial image to generate a composite image of a kimono that can be tried on. Limit the suggestions to styles suitable for coming-of-age ceremonies and propose designs that are compatible with the user's facial features." This prompt is used to instruct the generation AI model when generating the composite image.
[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0076] Step 1:
[0077] Users access the system using their devices and complete the registration process. During this process, they enter personal information such as their name and email address. This information is sent from the device to the server and stored in the server's database. This creates the user's account.
[0078] Step 2:
[0079] The user takes a picture of their face using the device's camera and uploads it to the system. The device receives this face image and sends it to the server. At this point, the input is the user's face image, and the output is data for storage on the server.
[0080] Step 3:
[0081] The server analyzes the received facial image data using an "image processing API." Specifically, it uses the "FaceMatch API" to extract the visual features of the face. The input to this process is the user's facial image, and the output is the extracted feature data. This clarifies the user's unique visual features.
[0082] Step 4:
[0083] The server selects suitable clothing information from the clothing database based on the extracted visual features. It executes a database query to list clothing that matches the facial features. The input for this step is visual feature data, and the output is the selected clothing information.
[0084] Step 5:
[0085] The server uses a "generative AI model" to combine selected clothing information with the user's facial image to generate a realistic composite image. The AI model matches the clothing to the face, visually recreating the try-on experience. The input for this step is clothing information and a facial image, and the output is a composite image.
[0086] Step 6:
[0087] The terminal receives a composite image sent from the server and presents it to the user. The image is displayed on the terminal's screen, and the user can view it and select their desired outfit. The input is a composite image from the server, and the output is a visual display for the user to see.
[0088] Step 7:
[0089] Users select an outfit from a collection of composite images displayed via their terminal and make a fitting reservation. Purchase and rental procedures are also possible. Selection and transaction information is transmitted from the terminal to a server and stored in a management system. Inputs are the user's selection and transaction information, while outputs are reservation confirmation and payment confirmation.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] In online shopping, users often feel anxious about purchasing clothing because it's difficult to try things on in person. Furthermore, there's a lack of efficient ways to try on clothes and make a purchase decision without visiting a physical store. In this situation, there's a need to provide a system that allows users to conveniently and accurately select and purchase clothing in a digital environment.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining the acquired image data with clothing selected from a clothing database, and means for providing the generated composite image to the user and allowing them to select their preferred clothing. This enables the user to efficiently and easily select the clothing best suited to them within a digital virtual store and to consistently carry out procedures from trying on to purchasing or renting.
[0095] "Means for acquiring user image data" refers to the method by which the system receives photographs and image information from the user and transmits it to the processor.
[0096] "A means of generating a composite image by combining clothing selected from a clothing database based on acquired image data" refers to a method of creating a virtual image that makes it appear as if the user is wearing the clothing, using the user's image and clothing information in the database.
[0097] "A means of providing users with generated composite images and allowing them to select their preferred outfit" refers to a method that displays composite try-on images to the user and enables them to choose their preferred outfit from among them.
[0098] "A means of managing and providing fitting reservations or purchase / rental procedures based on selected costume information within a digital virtual store" refers to a method of aggregating information regarding reservations, purchases, or rentals related to selected costumes and efficiently carrying them out online.
[0099] "Performing facial synthesis processing and providing digital try-on via smart devices" refers to the process of processing a user's facial image and providing a virtual try-on experience through the user's smart device.
[0100] "Optimizing the selection of suggested outfits using a generative AI model and generating and presenting prompt sentences" refers to a process that utilizes artificial intelligence technology to suggest outfits that match the user's preferences and outputs related instruction sentences based on the results.
[0101] The embodiment of the invention is based on a system that enables users to have a digital try-on experience in their daily lives using smart devices. This system mainly consists of three elements: a server, a terminal, and the user.
[0102] First, the user uploads their facial image to the system using their device. The device is a smartphone or tablet computer, and its role is to send the image data to the server.
[0103] The servers are built on cloud services such as AWS® and Google® Cloud Platform. The servers use OpenCV to process the received facial images and, based on the obtained data, use pre-trained generative AI models with TENSORFLOW® and PyTorch to generate a composite image combining the user's face and clothing. This composite image functions as a predictive image of what the user would look like wearing the created clothing.
[0104] Furthermore, by providing the generated composite images, the server accumulates user preference data and uses a suggestion algorithm to select even more suitable outfits. These optimized suggestions are then presented to the user again. During this process, the generative AI model generates prompt messages to guide the user to the next action.
[0105] Users can see their avatar trying on outfits through a virtual display. For example, they can easily proceed by receiving prompts such as, "Please upload a user face image," "Next, please select the category of outfit to try on: casual, formal, traditional, etc.," and "Select your favorite outfit based on the generated try-on images and proceed with the purchase process."
[0106] This system allows users to try on various outfits from the comfort of their homes, without having to physically go to a store, and efficiently select their favorites. This digital try-on experience offers a new way for users to enjoy online shopping more in line with their lifestyle.
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The user takes a picture of their face using a smart device and uploads it to the system. The system receives the face image as input and sends this image to the server as output.
[0110] Step 2:
[0111] The server uses OpenCV to perform image analysis based on the user's face image received. The input is the user's face image, and the output is the extracted facial features after analysis. Specific operations include facial landmark detection and image preprocessing (resolution adjustment and noise reduction).
[0112] Step 3:
[0113] The server uses TensorFlow to generate a synthesizer image of the clothing based on the analyzed facial features and clothing database. The input is the analyzed facial feature data and clothing database, and the output is the synthesized try-on image. This step involves the integration of digital data by applying clothing to facial features.
[0114] Step 4:
[0115] The server sends the generated composite image to the terminal and displays it to the user. The input is the composite fitting image data, and the output is displayed on the user's device screen. This display allows the user to have a virtual fitting experience.
[0116] Step 5:
[0117] The user views a composite image displayed on their device and selects their preferred outfit. The input is the visual data of the try-on image, and the output is the digital information of the selected outfit. This selection is made through the application's user interface (UI).
[0118] Step 6:
[0119] The server collects information about the outfit selected by the user and generates prompt messages to guide the user's next action. The input is the data of the selected outfit, and the output is the generated prompt message. Here, an AI model is used to generate the prompt message and suggest the next action.
[0120] Step 7:
[0121] Based on the selected outfit, the server manages the fitting reservation or purchase / rental process. Inputs are the user's selection information and registration information, while outputs are reservation confirmations and rental completion statuses. This process is achieved by the backend system updating customer management information.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] This invention relates to a system that acquires user image data, generates composite images for digitally trying on clothing, and further recognizes the user's emotions to optimize clothing selection. This system improves the accuracy of suggestions by understanding the user's preferences through the integration of an emotion engine.
[0124] The system works as follows: First, the user accesses the system and creates an account. Then, the user uploads their facial image using a device. The device retrieves the image data and sends it to the server.
[0125] The server processes the facial image while simultaneously selecting various outfits from a clothing database and generating a composite image by combining them with the user's facial image. In addition, when the composite image is displayed, the emotion engine evaluates the user's response by analyzing the user's facial expressions in real time through the camera on the terminal or a separately connected device.
[0126] The emotion engine captures subtle changes in the user's facial expressions to infer their emotional state, and based on the emotion data obtained during the suggestion of synthesized images, it provides more suitable clothing options for the user. To achieve this, it combines past preference history and emotion data, and builds a feedback system that learns the user's preference history to improve the accuracy of future suggestions.
[0127] For example, when a user is choosing a kimono for their coming-of-age ceremony, if the emotion engine recognizes feelings of approval or surprise, it can prioritize presenting more vibrant and stylish kimono options based on that emotional state. This system makes the virtual try-on experience through synthesized images even smoother, enabling users to proceed quickly and efficiently with everything from making a fitting reservation to purchasing or renting.
[0128] In this way, the present invention realizes highly accurate clothing selection support that reflects the user's emotions, thereby improving the comfort and satisfaction of the fitting process.
[0129] The following describes the processing flow.
[0130] Step 1:
[0131] The user accesses the system and registers an account. They enter the necessary information (name, email address, password, etc.) to create an account. The device sends this information to the server and saves the account data.
[0132] Step 2:
[0133] The user selects a photo using their device to upload a facial image. The device sends the facial image data to the server. The server receives this image information and begins image processing.
[0134] Step 3:
[0135] The server generates a composite image by combining the acquired facial image data with clothing selected from a clothing database. Furthermore, it constructs the composite image to facilitate analysis by the emotion engine.
[0136] Step 4:
[0137] The device displays the generated composite image to the user. At the same time, it activates a function that uses the device's built-in camera to capture the user's facial expressions in real time.
[0138] Step 5:
[0139] The emotion engine analyzes the user's captured facial expression data to identify emotions such as anger and joy. Based on the emotional state, the server filters recommended items to suggest the most suitable outfit for the user.
[0140] Step 6:
[0141] The user selects their preferred outfit from the suggested options. The device sends the selection information to the server, which then uses that information to initiate the process of making a fitting reservation or purchasing / renting the outfit.
[0142] Step 7:
[0143] The server coordinates with the sales / rental provider for the selected costume and schedules the fitting. Once the fitting reservation is complete, the server sends a confirmation email or notification to the user.
[0144] Step 8:
[0145] The server uses feedback from the emotion engine to learn from the acquired emotion data and preference history, and updates the database to improve the accuracy of future suggestions. This update makes it possible to provide more personalized suggestions on subsequent uses.
[0146] (Example 2)
[0147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0148] In conventional costume fitting systems, user image data is handled statically, making it difficult to provide dynamic costume suggestions based on individual user preferences. Furthermore, the inability to make suggestions that take into account the user's emotional state resulted in a limited user experience and difficulty in improving satisfaction.
[0149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0150] In this invention, the server includes means for acquiring the user's digital data, means for generating composite data by combining items selected from a database based on the acquired digital data, and means for analyzing the user's facial expressions, estimating their emotional state, and optimizing the items to be suggested. This makes it possible to suggest the optimal clothing that reflects the user's individual emotions and preferences in real time.
[0151] "User digital data" refers to information such as digital images and facial photographs that the system acquires from users.
[0152] A "database" is a collection of information containing diverse item information and options, which a system uses to select items to suggest to users.
[0153] "Synthetic data" refers to digital data generated by combining acquired user digital data with selected item information from a database.
[0154] "Analyzing facial expressions" refers to the process of analyzing the characteristics and movements of a user's face and using that information to infer their emotions and psychological state.
[0155] "Emotional state" refers to the user's current emotional response and psychological state, and is information inferred through facial expression analysis.
[0156] "Optimizing products" refers to the process of adjusting the selection of products offered to the user to be the most appropriate, based on the user's preferences and emotional state.
[0157] To implement this invention, it is necessary to build a system in which a server, terminal, and user work together. The specific operation method is described below.
[0158] Users first access the system using a terminal. The terminal is equipped with a camera, allowing users to take photos of their own face or select and upload existing digital data. This digital data is encrypted and sent from the terminal to the server.
[0159] The server digitally processes the received facial data. Specifically, it extracts facial feature data using an image processing library. Next, the server retrieves data on various items from a database and generates composite data by combining it with the user's facial data. This process uses a generative AI model, primarily employing techniques to achieve style transfer and realistic object synthesis.
[0160] Furthermore, the device's camera captures the user's facial expressions in real time. The server analyzes this facial data to estimate the user's emotional state. Analysis using a deep learning-based emotion recognition model enables optimal product recommendations based on the user's emotions.
[0161] For example, when a user selects an outfit for their coming-of-age ceremony, the server generates composite data and displays it to the user via their device. By analyzing the user's facial expressions and selection history, the system can suggest outfits that match the user's preferences.
[0162] An example of a prompt message is: "I would like to try on a virtual kimono for my coming-of-age ceremony. There are several styles of kimonos available; please display the best option based on the user's preferences."
[0163] In this way, users can obtain an optimal fitting experience tailored to their individual needs, allowing them to make more informed decisions regarding purchases and procedures.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] The user uses a terminal to access the system and create a new account. The terminal then sends account data to the server based on the user's input. The input consists of the user's basic information, and the output is the registered account data.
[0167] Step 2:
[0168] The user uses a device to capture or select their own facial data and upload it. The device compresses and encrypts the image data to standardize the format before sending it to the server. Here, the input is the user's facial image data, and the output is the facial data securely transmitted to the server.
[0169] Step 3:
[0170] The server extracts feature points from the received facial data using an image processing library. This involves using a facial recognition algorithm to extract basic facial features such as the positions of the eyes, nose, and mouth. The input is facial image data, and the output is facial feature point data.
[0171] Step 4:
[0172] The server retrieves item data tailored to the user from the database. Here, a generative AI model is used to combine facial feature point data with item data to generate realistic synthetic data. At this stage, the inputs are facial feature point data and item data, and the output is the synthetic data.
[0173] Step 5:
[0174] The device's camera captures the user's facial expressions in real time and sends them to the server. The server receives this data and analyzes the emotional state using an emotion recognition model. The input is real-time facial expression data, and the output is the analyzed emotional state data.
[0175] Step 6:
[0176] The server uses emotional state data to select the most suitable item from synthesized data and outputs it to the user. In this case, the input is emotional state data and synthesized data, and the output is the most suitable item suggestion according to the emotion.
[0177] Step 7:
[0178] The user reviews the suggested items on the terminal and selects the desired items. The terminal sends the selected item information to the server and supports the reservation or purchase process. The input is the user's item selection, and the output is the selection data sent to the server and a notification that the process is complete.
[0179] (Application Example 2)
[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0181] Traditional clothing selection and fitting processes required customers to physically try on garments, which presented problems due to the time and effort involved. Furthermore, online shopping carried the risk of post-purchase dissatisfaction because customers hadn't actually tried on the clothes. Additionally, the inability to adequately consider individual customer preferences and feelings resulted in low customer satisfaction with the selection process.
[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0183] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining selected clothing from a clothing database based on the acquired image data, and means for analyzing the user's facial expression data in real time to recognize emotions and optimize the selected clothing. This makes it possible for customers to efficiently select the optimal clothing based on their emotions digitally without having to physically try on clothes.
[0184] "User image data" refers to image information acquired for the purpose of identifying an individual, and is particularly used for the purpose of selecting clothing.
[0185] A "costume database" is a source of information that stores digital information on various costumes and is used to select appropriate costumes based on the user's image data.
[0186] A "composite image" is an image generated by combining the user's image data and clothing data, enabling digital try-on.
[0187] "Real-time analysis" refers to a technical process that processes input data immediately and provides the results almost instantly.
[0188] "Recognizing emotions" refers to the process of analyzing a user's facial expression data to infer their psychological state and preferences.
[0189] "Optimizing selected outfits" means adjusting the selection process to suggest the most appropriate outfit based on the user's emotions and past preference history.
[0190] The system that realizes this invention allows users to digitally try on clothes by standing in front of a smart mirror in a store. The system mainly functions as follows:
[0191] The server acquires the user's facial image from a camera built into the smart mirror. The camera instantly captures the user's face and body in high resolution and sends the data to the server. The server processes the acquired image data using image processing software such as OpenCV or TensorFlow, and combines it with multiple selected clothing data from a clothing database to generate a composite image. This allows the user to virtually try on and check their appearance on the mirror.
[0192] Furthermore, the server uses AI technologies for emotion recognition, such as Azure Cognitive Services and AWS Rekognition, to analyze real-time facial expression data acquired through the camera installed in the smart mirror. This method recognizes the user's emotions and predicts whether the selected outfit matches the user's preferences. Based on the results, the emotion engine optimizes outfit suggestions, presenting the user with more suitable options.
[0193] For example, if a user choosing a kimono for their coming-of-age ceremony stands in front of a mirror and smiles while virtually trying on a red and gold kimono, the system can recognize this expression as joy and present this kimono as a preferred option. An example of a prompt message used as input to the generating AI model might be: "A woman in her 20s is digitally trying on a vibrant and stylish kimono for her coming-of-age ceremony. She has a happy expression as she looks at the red and gold kimono."
[0194] This system allows users to enjoy a highly personalized try-on experience and select their ideal outfit while saving time and effort.
[0195] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0196] Step 1:
[0197] The user stands in front of a smart mirror in the store. The smart mirror's built-in camera captures the user's face and body in high resolution. The input data is the user's image data, which is then sent to the server.
[0198] Step 2:
[0199] The server receives the acquired image data and processes it using image processing software such as OpenCV or TensorFlow. A face recognition algorithm extracts facial feature points, and the server outputs search criteria for the clothing database. Based on these criteria, it selects appropriate clothing data.
[0200] Step 3:
[0201] The server combines the selected costume data with the user's image data to generate a composite image. This process utilizes CGI to output a visual representation that makes it appear as if the user is wearing the costume. The composite image is then mirrored back and provided to the user.
[0202] Step 4:
[0203] The smart mirror, which is the terminal device, displays a composite image. At the same time, the mirror's built-in camera continuously acquires the user's facial expression data in real time. This facial expression data, as input, is sent to a server for emotion recognition.
[0204] Step 5:
[0205] The server uses emotion recognition AI such as Azure Cognitive Services and AWS Rekognition to analyze facial expression data. The analysis results are generated as output data indicating the user's emotional state. Based on this information, the emotion engine selects the optimal outfit and ultimately presents the user with the most suitable options.
[0206] Step 6:
[0207] The user accepts the sentiment analysis and outfit suggestions and makes a final selection. This selection is recorded on the server and used as feedback data to improve the accuracy of suggestions in the future.
[0208] Step 7:
[0209] The server sends the necessary information for the purchase or reservation process to the user's terminal based on the costume selected by the user. The output is provided to the user as a purchase confirmation or reservation completion notification.
[0210] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0211] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0212] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0213] [Second Embodiment]
[0214] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0215] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0216] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0217] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0218] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0219] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0220] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0221] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0222] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0223] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0224] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0225] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0226] This invention is a system that enables users to efficiently select suitable clothing and smoothly proceed through the fitting process. This system primarily functions through the coordinated operation of three elements: a server, a terminal, and the user.
[0227] First, the user accesses the system, completes the necessary registration process, and then uploads a picture of their face. This image is used as the foundation for providing a digital costume try-on experience. The terminal receives the image uploaded by the user and sends it to the server.
[0228] The server processes the received facial image data, selects various outfits from a clothing database, and generates a composite image by combining the outfit with the user's face. This composite image allows the user to virtually try on many different outfits.
[0229] The generated composite image is provided to the user in a viewable format. The user can select an outfit that they think suits them through the composite image. Based on this result, the server saves the selected outfit information and manages the reservation process for actual try-ons, as well as the purchase or rental process for the outfit, if desired.
[0230] As a concrete example, let's consider the case of choosing a furisode (long-sleeved kimono) for a coming-of-age ceremony. In this case, the user wants to try on furisode in order to attend the ceremony. First, the user uploads a facial image to the system. The server generates composite images of the user wearing various furisode based on the user's facial image data and presents them to the user. The user then selects a furisode they like and makes a reservation to try it on. If necessary, they can make a payment on the spot and complete the furisode rental procedure.
[0231] This system allows users to easily find the outfit that suits them best and efficiently go through the entire process from trying it on to purchasing or renting it.
[0232] The following describes the processing flow.
[0233] Step 1:
[0234] The user accesses the system and first registers an account. They enter the required information (name, email address, password, etc.) to create an account. The terminal collects this information and sends the account data to the server. The server stores the received information in its database and creates the account.
[0235] Step 2:
[0236] After registering an account, the user uploads a facial image. This process utilizes an image selection tool. The device sends the uploaded image data to the server. The server receives the image data and prepares it for image processing.
[0237] Step 3:
[0238] The server inputs the received facial image data into an AI model and generates a composite image by combining it with clothing selected from a clothing database. The AI model performs high-precision image processing to digitally composite various clothing onto the user's face.
[0239] Step 4:
[0240] The server sends the generated composite images to the terminal and presents them to the user in a viewable format. The user reviews these composite images and selects their favorite outfit. The selection process is performed on the terminal, and the selection results are sent back to the server.
[0241] Step 5:
[0242] The server stores the selected costume information and provides an interface for making fitting reservations and purchasing / renting items. Users select a fitting date and time and make a reservation. The server processes this information and coordinates the schedule with partner sales / rental businesses.
[0243] Step 6:
[0244] Once the reservation is complete, the server sends a confirmation notice to the user. If the user wishes to purchase or rent, the server provides a secure payment environment and completes the transaction. The user receives confirmation that the process is complete and can then finish using the service.
[0245] (Example 1)
[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0247] The process of selecting and trying on clothing requires time and effort for users to try on many options. Physical try-ons require travel to stores and making appointments, which is inconvenient. Furthermore, the limited variety of clothing available for try-on makes it difficult to make the best clothing choice.
[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0249] In this invention, the server includes means for acquiring and storing the user's image information, means for analyzing visual features based on the acquired image information to select clothing information from a data set and generating a composite image using an image synthesis model, and means for presenting the generated composite image to the user and presenting selectable clothing information. As a result, the user can choose from a variety of clothing without physically trying them on, enabling them to make the optimal clothing selection while significantly reducing time and effort.
[0250] "User image information" refers to visual data that individuals provide to the system, and this data is used to analyze the characteristics of the users.
[0251] "Visual features" refer to information about appearance and shape extracted from image data, and represent the individual characteristics of the user necessary for generating a composite image.
[0252] A "data set" is a database containing diverse clothing information, used to select the most suitable outfit based on the user's image information.
[0253] "Costume information" refers to data related to clothing, including detailed information such as the design, color, and material of the garments.
[0254] An "image synthesis model" refers to an algorithm or program used to generate a realistic composite image by combining image information and clothing information acquired using AI technology.
[0255] A "composite image" is a simulated image created based on the user's image information and selected clothing information, visually representing the try-on state.
[0256] "To present" refers to the act of showing generated information or images to the user through a user interface, in order to enable them to make choices.
[0257] "Try-on procedure" refers to the process of making reservations and arrangements for users to actually try on the outfits they have selected.
[0258] "Commercial transaction procedures" refer to the series of activities involved in processing payments and contracts necessary when a user purchases or rents an outfit they have selected.
[0259] The following describes embodiments for carrying out this system invention.
[0260] The present invention provides a system for users to efficiently select suitable clothing digitally. This system primarily consists of three elements: the user, the terminal, and the server.
[0261] Users first access the system using their device and register by entering the required information. This includes providing basic personal information. Afterward, users use their device's camera function to take a picture of their face and upload it to the system. This face image serves as the basis for generating a composite image of the costume.
[0262] The device is equipped with the ability to receive facial images uploaded by users and quickly transmit them to the server. This communication process is supported by a high-speed and stable internet connection.
[0263] The server analyzes the facial image data received from the user using an "image processing API." Here, a specific technology called the "FaceMatch API" is used to extract the visual features of the face. Based on this, the server utilizes a "generative AI model" to select the most suitable outfit from a clothing database. The selected outfit is then combined with the user's facial image using a "DressAI model" to generate a composite image that provides a realistic try-on experience.
[0264] The device receives a composite image sent from the server and displays it to the user. This image allows the user to compare various outfits without actually trying them on. This process is performed using the device's high-resolution display, making it easy for the user to view the composite image.
[0265] As a concrete example, in the case of a user choosing a furisode (long-sleeved kimono) for their coming-of-age ceremony, the user makes a reservation to try on furisode for the ceremony. The user uploads a facial image, and after processing with an "image processing API" and a "generative AI model," receives a composite image of themselves wearing various furisode. Based on this image, the user selects the most suitable furisode and makes a reservation to try it on from their device. Furthermore, when purchasing or renting, credit card payment can also be made via the device.
[0266] In this way, users can intuitively select costumes even from remote locations and efficiently proceed with costume selection, purchase, or rental procedures.
[0267] An example of a prompt is: "Use the user's facial image to generate a composite image of a kimono that can be tried on. Limit the suggestions to styles suitable for coming-of-age ceremonies and propose designs that are compatible with the user's facial features." This prompt is used to instruct the generation AI model when generating the composite image.
[0268] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0269] Step 1:
[0270] Users access the system using their devices and complete the registration process. During this process, they enter personal information such as their name and email address. This information is sent from the device to the server and stored in the server's database. This creates the user's account.
[0271] Step 2:
[0272] The user takes a picture of their face using the device's camera and uploads it to the system. The device receives this face image and sends it to the server. At this point, the input is the user's face image, and the output is data for storage on the server.
[0273] Step 3:
[0274] The server analyzes the received facial image data using an "image processing API." Specifically, it uses the "FaceMatch API" to extract the visual features of the face. The input to this process is the user's facial image, and the output is the extracted feature data. This clarifies the user's unique visual features.
[0275] Step 4:
[0276] The server selects suitable clothing information from the clothing database based on the extracted visual features. It executes a database query to list clothing that matches the facial features. The input for this step is visual feature data, and the output is the selected clothing information.
[0277] Step 5:
[0278] The server uses a "generative AI model" to combine selected clothing information with the user's facial image to generate a realistic composite image. The AI model matches the clothing to the face, visually recreating the try-on experience. The input for this step is clothing information and a facial image, and the output is a composite image.
[0279] Step 6:
[0280] The terminal receives the synthesized image transmitted from the server and presents it to the user. The terminal displays the image on its display, and the user can view it and select the desired clothing. The input is the synthesized image from the server, and the output is the visual display for showing to the user.
[0281] Step 7:
[0282] The user selects the clothing selected from the synthesized images presented through the terminal and makes a reservation for a fitting. Also, procedures for purchase or rental are possible. The selection information and transaction information are transmitted from the terminal to the server and stored in the management system. The input is the user's selection information and transaction information, and the output is the reservation confirmation and payment confirmation.
[0283] (Application Example 1)
[0284] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0285] In online shopping, when a user selects clothing suitable for themselves, it is difficult to actually try it on, and they often feel不安 about purchasing. Also, there is a lack of means to efficiently try on and decide to purchase without going to a store. In such a situation, it is necessary to provide a system that allows users to conveniently and highly accurately select and purchase clothing in a digital environment.
[0286] The specific processing by the specific processing unit of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0287] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining the acquired image data with clothing selected from a clothing database, and means for providing the generated composite image to the user and allowing them to select their preferred clothing. This enables the user to efficiently and easily select the clothing best suited to them within a digital virtual store and to consistently carry out procedures from trying on to purchasing or renting.
[0288] "Means for acquiring user image data" refers to the method by which the system receives photographs and image information from the user and transmits it to the processor.
[0289] "A means of generating a composite image by combining clothing selected from a clothing database based on acquired image data" refers to a method of creating a virtual image that makes it appear as if the user is wearing the clothing, using the user's image and clothing information in the database.
[0290] "A means of providing users with generated composite images and allowing them to select their preferred outfit" refers to a method that displays composite try-on images to the user and enables them to choose their preferred outfit from among them.
[0291] "A means of managing and providing fitting reservations or purchase / rental procedures based on selected costume information within a digital virtual store" refers to a method of aggregating information regarding reservations, purchases, or rentals related to selected costumes and efficiently carrying them out online.
[0292] "Performing facial synthesis processing and providing digital try-on via smart devices" refers to the process of processing a user's facial image and providing a virtual try-on experience through the user's smart device.
[0293] "Optimizing the selection of suggested outfits using a generative AI model and generating and presenting prompt sentences" refers to a process that utilizes artificial intelligence technology to suggest outfits that match the user's preferences and outputs related instruction sentences based on the results.
[0294] The embodiment of the invention is based on a system that enables users to have a digital try-on experience in their daily lives using smart devices. This system mainly consists of three elements: a server, a terminal, and the user.
[0295] First, the user uploads their facial image to the system using their device. The device is a smartphone or tablet computer, and its role is to send the image data to the server.
[0296] The servers are built on cloud services such as AWS and Google Cloud Platform. The servers use OpenCV to process the received facial images and, based on the obtained data, utilize TensorFlow, PyTorch, and other tools to generate a composite image combining the user's face and clothing using a pre-trained generative AI model. This composite image functions as a predictive image of what the person would look like wearing the created clothing.
[0297] Furthermore, by providing the generated composite images, the server accumulates user preference data and uses a suggestion algorithm to select even more suitable outfits. These optimized suggestions are then presented to the user again. During this process, the generative AI model generates prompt messages to guide the user to the next action.
[0298] Users can see their avatar trying on outfits through a virtual display. For example, they can easily proceed by receiving prompts such as, "Please upload a user face image," "Next, please select the category of outfit to try on: casual, formal, traditional, etc.," and "Select your favorite outfit based on the generated try-on images and proceed with the purchase process."
[0299] With this system, users can try on various outfits and efficiently select the ones they like without physically going to a store and while staying at home. This digital try-on experience provides a new way for users to enjoy online shopping according to their own lifestyle.
[0300] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0301] Step 1:
[0302] The user uses a smart device to take a picture of their face and upload it to the system. The terminal receives the face image as input and sends this image to the server as output.
[0303] Step 2:
[0304] Based on the received face image of the user, the server performs image analysis using OpenCV. The input is the user's face image, and the output is to extract the analyzed facial features. Specific operations include face landmark detection and preprocessing of the image (adjusting resolution and removing noise).
[0305] Step 3:
[0306] Based on the analyzed facial features and the clothing database, the server uses the generated AI model in TensorFlow to generate a composite image of the clothing. The input is the analyzed facial feature data and the clothing database, and the output is the synthesized try-on image. In this step, the integration of digital data, namely the application of clothing to facial features, is performed.
[0307] Step 4:
[0308] The server sends the generated composite image to the terminal and displays it to the user. The input is the composite try-on image data, and the output is displayed on the user's device screen. Through this display, the user can experience virtual try-on.
[0309] Step 5:
[0310] The user views a composite image displayed on their device and selects their preferred outfit. The input is the visual data of the try-on image, and the output is the digital information of the selected outfit. This selection is made through the application's user interface (UI).
[0311] Step 6:
[0312] The server collects information about the outfit selected by the user and generates prompt messages to guide the user's next action. The input is the data of the selected outfit, and the output is the generated prompt message. Here, an AI model is used to generate the prompt message and suggest the next action.
[0313] Step 7:
[0314] Based on the selected outfit, the server manages the fitting reservation or purchase / rental process. Inputs are the user's selection information and registration information, while outputs are reservation confirmations and rental completion statuses. This process is achieved by the backend system updating customer management information.
[0315] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0316] This invention relates to a system that acquires user image data, generates composite images for digitally trying on clothing, and further recognizes the user's emotions to optimize clothing selection. This system improves the accuracy of suggestions by understanding the user's preferences through the integration of an emotion engine.
[0317] The system works as follows: First, the user accesses the system and creates an account. Then, the user uploads their facial image using a device. The device retrieves the image data and sends it to the server.
[0318] The server processes the facial image while simultaneously selecting various outfits from a clothing database and generating a composite image by combining them with the user's facial image. In addition, when the composite image is displayed, the emotion engine evaluates the user's response by analyzing the user's facial expressions in real time through the camera on the terminal or a separately connected device.
[0319] The emotion engine captures subtle changes in the user's facial expressions to infer their emotional state, and based on the emotion data obtained during the suggestion of synthesized images, it provides more suitable clothing options for the user. To achieve this, it combines past preference history and emotion data, and builds a feedback system that learns the user's preference history to improve the accuracy of future suggestions.
[0320] For example, when a user is choosing a kimono for their coming-of-age ceremony, if the emotion engine recognizes feelings of approval or surprise, it can prioritize presenting more vibrant and stylish kimono options based on that emotional state. This system makes the virtual try-on experience through synthesized images even smoother, enabling users to proceed quickly and efficiently with everything from making a fitting reservation to purchasing or renting.
[0321] In this way, the present invention realizes highly accurate clothing selection support that reflects the user's emotions, thereby improving the comfort and satisfaction of the fitting process.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The user accesses the system and registers an account. They enter the necessary information (name, email address, password, etc.) to create an account. The device sends this information to the server and saves the account data.
[0325] Step 2:
[0326] The user selects a photo using their device to upload a facial image. The device sends the facial image data to the server. The server receives this image information and begins image processing.
[0327] Step 3:
[0328] The server generates a composite image by combining the acquired facial image data with clothing selected from a clothing database. Furthermore, it constructs the composite image to facilitate analysis by the emotion engine.
[0329] Step 4:
[0330] The device displays the generated composite image to the user. At the same time, it activates a function that uses the device's built-in camera to capture the user's facial expressions in real time.
[0331] Step 5:
[0332] The emotion engine analyzes the user's captured facial expression data to identify emotions such as anger and joy. Based on the emotional state, the server filters recommended items to suggest the most suitable outfit for the user.
[0333] Step 6:
[0334] The user selects their preferred outfit from the suggested options. The device sends the selection information to the server, which then uses that information to initiate the process of making a fitting reservation or purchasing / renting the outfit.
[0335] Step 7:
[0336] The server coordinates with the sales / rental provider for the selected costume and schedules the fitting. Once the fitting reservation is complete, the server sends a confirmation email or notification to the user.
[0337] Step 8:
[0338] The server uses feedback from the emotion engine to learn from the acquired emotion data and preference history, and updates the database to improve the accuracy of future suggestions. This update makes it possible to provide more personalized suggestions on subsequent uses.
[0339] (Example 2)
[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0341] In conventional costume fitting systems, user image data is handled statically, making it difficult to provide dynamic costume suggestions based on individual user preferences. Furthermore, the inability to make suggestions that take into account the user's emotional state resulted in a limited user experience and difficulty in improving satisfaction.
[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0343] In this invention, the server includes means for acquiring the user's digital data, means for generating composite data by combining items selected from a database based on the acquired digital data, and means for analyzing the user's facial expressions, estimating their emotional state, and optimizing the items to be suggested. This makes it possible to suggest the optimal clothing that reflects the user's individual emotions and preferences in real time.
[0344] "User digital data" refers to information such as digital images and facial photographs that the system acquires from users.
[0345] A "database" is a collection of information containing diverse item information and options, which a system uses to select items to suggest to users.
[0346] "Synthetic data" refers to digital data generated by combining acquired user digital data with selected item information from a database.
[0347] "Analyzing facial expressions" refers to the process of analyzing the characteristics and movements of a user's face and using that information to infer their emotions and psychological state.
[0348] "Emotional state" refers to the user's current emotional response and psychological state, and is information inferred through facial expression analysis.
[0349] "Optimizing products" refers to the process of adjusting the selection of products offered to the user to be the most appropriate, based on the user's preferences and emotional state.
[0350] To implement this invention, it is necessary to build a system in which a server, terminal, and user work together. The specific operation method is described below.
[0351] Users first access the system using a terminal. The terminal is equipped with a camera, allowing users to take photos of their own face or select and upload existing digital data. This digital data is encrypted and sent from the terminal to the server.
[0352] The server digitally processes the received facial data. Specifically, it extracts facial feature data using an image processing library. Next, the server retrieves data on various items from a database and generates composite data by combining it with the user's facial data. This process uses a generative AI model, primarily employing techniques to achieve style transfer and realistic object synthesis.
[0353] Furthermore, the device's camera captures the user's facial expressions in real time. The server analyzes this facial data to estimate the user's emotional state. Analysis using a deep learning-based emotion recognition model enables optimal product recommendations based on the user's emotions.
[0354] For example, when a user selects an outfit for their coming-of-age ceremony, the server generates composite data and displays it to the user via their device. By analyzing the user's facial expressions and selection history, the system can suggest outfits that match the user's preferences.
[0355] An example of a prompt message is: "I would like to try on a virtual kimono for my coming-of-age ceremony. There are several styles of kimonos available; please display the best option based on the user's preferences."
[0356] In this way, users can obtain an optimal fitting experience tailored to their individual needs, allowing them to make more informed decisions regarding purchases and procedures.
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] The user uses a terminal to access the system and create a new account. The terminal then sends account data to the server based on the user's input. The input consists of the user's basic information, and the output is the registered account data.
[0360] Step 2:
[0361] The user uses a device to capture or select their own facial data and upload it. The device compresses and encrypts the image data to standardize the format before sending it to the server. Here, the input is the user's facial image data, and the output is the facial data securely transmitted to the server.
[0362] Step 3:
[0363] The server extracts feature points from the received facial data using an image processing library. This involves using a facial recognition algorithm to extract basic facial features such as the positions of the eyes, nose, and mouth. The input is facial image data, and the output is facial feature point data.
[0364] Step 4:
[0365] The server retrieves item data tailored to the user from the database. Here, a generative AI model is used to combine facial feature point data with item data to generate realistic synthetic data. At this stage, the inputs are facial feature point data and item data, and the output is the synthetic data.
[0366] Step 5:
[0367] The device's camera captures the user's facial expressions in real time and sends them to the server. The server receives this data and analyzes the emotional state using an emotion recognition model. The input is real-time facial expression data, and the output is the analyzed emotional state data.
[0368] Step 6:
[0369] The server uses emotional state data to select the most suitable item from synthesized data and outputs it to the user. In this case, the input is emotional state data and synthesized data, and the output is the most suitable item suggestion according to the emotion.
[0370] Step 7:
[0371] The user reviews the suggested items on the terminal and selects the desired items. The terminal sends the selected item information to the server and supports the reservation or purchase process. The input is the user's item selection, and the output is the selection data sent to the server and a notification that the process is complete.
[0372] (Application Example 2)
[0373] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0374] Traditional clothing selection and fitting processes required customers to physically try on garments, which presented problems due to the time and effort involved. Furthermore, online shopping carried the risk of post-purchase dissatisfaction because customers hadn't actually tried on the clothes. Additionally, the inability to adequately consider individual customer preferences and feelings resulted in low customer satisfaction with the selection process.
[0375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0376] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining selected clothing from a clothing database based on the acquired image data, and means for analyzing the user's facial expression data in real time to recognize emotions and optimize the selected clothing. This makes it possible for customers to efficiently select the optimal clothing based on their emotions digitally without having to physically try on clothes.
[0377] "User image data" refers to image information acquired for the purpose of identifying an individual, and is particularly used for the purpose of selecting clothing.
[0378] A "costume database" is a source of information that stores digital information on various costumes and is used to select appropriate costumes based on the user's image data.
[0379] A "composite image" is an image generated by combining the user's image data and clothing data, enabling digital try-on.
[0380] "Real-time analysis" refers to a technical process that processes input data immediately and provides the results almost instantly.
[0381] "Recognizing emotions" refers to the process of analyzing a user's facial expression data to infer their psychological state and preferences.
[0382] "Optimizing selected outfits" means adjusting the selection process to suggest the most appropriate outfit based on the user's emotions and past preference history.
[0383] The system that realizes this invention allows users to digitally try on clothes by standing in front of a smart mirror in a store. The system mainly functions as follows:
[0384] The server acquires the user's facial image from a camera built into the smart mirror. The camera instantly captures the user's face and body in high resolution and sends the data to the server. The server processes the acquired image data using image processing software such as OpenCV or TensorFlow, and combines it with multiple selected clothing data from a clothing database to generate a composite image. This allows the user to virtually try on and check their appearance on the mirror.
[0385] Furthermore, the server uses AI technologies for emotion recognition, such as Azure Cognitive Services and AWS Rekognition, to analyze real-time facial expression data acquired through the camera installed in the smart mirror. This method recognizes the user's emotions and predicts whether the selected outfit matches the user's preferences. Based on the results, the emotion engine optimizes outfit suggestions, presenting the user with more suitable options.
[0386] For example, if a user choosing a kimono for their coming-of-age ceremony stands in front of a mirror and smiles while virtually trying on a red and gold kimono, the system can recognize this expression as joy and present this kimono as a preferred option. An example of a prompt message used as input to the generating AI model might be: "A woman in her 20s is digitally trying on a vibrant and stylish kimono for her coming-of-age ceremony. She has a happy expression as she looks at the red and gold kimono."
[0387] This system allows users to enjoy a highly personalized try-on experience and select their ideal outfit while saving time and effort.
[0388] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0389] Step 1:
[0390] The user stands in front of a smart mirror in the store. The smart mirror's built-in camera captures the user's face and body in high resolution. The input data is the user's image data, which is then sent to the server.
[0391] Step 2:
[0392] The server receives the acquired image data and processes it using image processing software such as OpenCV or TensorFlow. A face recognition algorithm extracts facial feature points, and the server outputs search criteria for the clothing database. Based on these criteria, it selects appropriate clothing data.
[0393] Step 3:
[0394] The server combines the selected costume data with the user's image data to generate a composite image. This process utilizes CGI to output a visual representation that makes it appear as if the user is wearing the costume. The composite image is then mirrored back and provided to the user.
[0395] Step 4:
[0396] The smart mirror, which is the terminal device, displays a composite image. At the same time, the mirror's built-in camera continuously acquires the user's facial expression data in real time. This facial expression data, as input, is sent to a server for emotion recognition.
[0397] Step 5:
[0398] The server uses emotion recognition AI such as Azure Cognitive Services and AWS Rekognition to analyze facial expression data. The analysis results are generated as output data indicating the user's emotional state. Based on this information, the emotion engine selects the optimal outfit and ultimately presents the user with the most suitable options.
[0399] Step 6:
[0400] The user accepts the sentiment analysis and outfit suggestions and makes a final selection. This selection is recorded on the server and used as feedback data to improve the accuracy of suggestions in the future.
[0401] Step 7:
[0402] The server sends the necessary information for the purchase or reservation process to the user's terminal based on the costume selected by the user. The output is provided to the user as a purchase confirmation or reservation completion notification.
[0403] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0404] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0405] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0406] [Third Embodiment]
[0407] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0408] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0409] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0410] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0411] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0412] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0413] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0414] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0415] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0416] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0417] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0418] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0419] This invention is a system that enables users to efficiently select suitable clothing and smoothly proceed through the fitting process. This system primarily functions through the coordinated operation of three elements: a server, a terminal, and the user.
[0420] First, the user accesses the system, completes the necessary registration process, and then uploads a picture of their face. This image is used as the foundation for providing a digital costume try-on experience. The terminal receives the image uploaded by the user and sends it to the server.
[0421] The server processes the received facial image data, selects various outfits from a clothing database, and generates a composite image by combining the outfit with the user's face. This composite image allows the user to virtually try on many different outfits.
[0422] The generated composite image is provided to the user in a viewable format. The user can select an outfit that they think suits them through the composite image. Based on this result, the server saves the selected outfit information and manages the reservation process for actual try-ons, as well as the purchase or rental process for the outfit, if desired.
[0423] As a concrete example, let's consider the case of choosing a furisode (long-sleeved kimono) for a coming-of-age ceremony. In this case, the user wants to try on furisode in order to attend the ceremony. First, the user uploads a facial image to the system. The server generates composite images of the user wearing various furisode based on the user's facial image data and presents them to the user. The user then selects a furisode they like and makes a reservation to try it on. If necessary, they can make a payment on the spot and complete the furisode rental procedure.
[0424] This system allows users to easily find the outfit that suits them best and efficiently go through the entire process from trying it on to purchasing or renting it.
[0425] The following describes the processing flow.
[0426] Step 1:
[0427] The user accesses the system and first registers an account. They enter the required information (name, email address, password, etc.) to create an account. The terminal collects this information and sends the account data to the server. The server stores the received information in its database and creates the account.
[0428] Step 2:
[0429] After registering an account, the user uploads a facial image. This process utilizes an image selection tool. The device sends the uploaded image data to the server. The server receives the image data and prepares it for image processing.
[0430] Step 3:
[0431] The server inputs the received facial image data into an AI model and generates a composite image by combining it with clothing selected from a clothing database. The AI model performs high-precision image processing to digitally composite various clothing onto the user's face.
[0432] Step 4:
[0433] The server sends the generated composite images to the terminal and presents them to the user in a viewable format. The user reviews these composite images and selects their favorite outfit. The selection process is performed on the terminal, and the selection results are sent back to the server.
[0434] Step 5:
[0435] The server stores the selected costume information and provides an interface for making fitting reservations and purchasing / renting items. Users select a fitting date and time and make a reservation. The server processes this information and coordinates the schedule with partner sales / rental businesses.
[0436] Step 6:
[0437] Once the reservation is complete, the server sends a confirmation notice to the user. If the user wishes to purchase or rent, the server provides a secure payment environment and completes the transaction. The user receives confirmation that the process is complete and can then finish using the service.
[0438] (Example 1)
[0439] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0440] The process of selecting and trying on clothing requires time and effort for users to try on many options. Physical try-ons require travel to stores and making appointments, which is inconvenient. Furthermore, the limited variety of clothing available for try-on makes it difficult to make the best clothing choice.
[0441] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0442] In this invention, the server includes means for acquiring and storing the user's image information, means for analyzing visual features based on the acquired image information to select clothing information from a data set and generating a composite image using an image synthesis model, and means for presenting the generated composite image to the user and presenting selectable clothing information. As a result, the user can choose from a variety of clothing without physically trying them on, enabling them to make the optimal clothing selection while significantly reducing time and effort.
[0443] "User image information" refers to visual data that individuals provide to the system, and this data is used to analyze the characteristics of the users.
[0444] "Visual features" refer to information about appearance and shape extracted from image data, and represent the individual characteristics of the user necessary for generating a composite image.
[0445] A "data set" is a database containing diverse clothing information, used to select the most suitable outfit based on the user's image information.
[0446] "Costume information" refers to data related to clothing, including detailed information such as the design, color, and material of the garments.
[0447] An "image synthesis model" refers to an algorithm or program used to generate a realistic composite image by combining image information and clothing information acquired using AI technology.
[0448] A "composite image" is a simulated image created based on the user's image information and selected clothing information, visually representing the try-on state.
[0449] "To present" refers to the act of showing generated information or images to the user through a user interface, in order to enable them to make choices.
[0450] "Try-on procedure" refers to the process of making reservations and arrangements for users to actually try on the outfits they have selected.
[0451] "Commercial transaction procedures" refer to the series of activities involved in processing payments and contracts necessary when a user purchases or rents an outfit they have selected.
[0452] The following describes embodiments for carrying out this system invention.
[0453] The present invention provides a system for users to efficiently select suitable clothing digitally. This system primarily consists of three elements: the user, the terminal, and the server.
[0454] Users first access the system using their device and register by entering the required information. This includes providing basic personal information. Afterward, users use their device's camera function to take a picture of their face and upload it to the system. This face image serves as the basis for generating a composite image of the costume.
[0455] The device is equipped with the ability to receive facial images uploaded by users and quickly transmit them to the server. This communication process is supported by a high-speed and stable internet connection.
[0456] The server analyzes the facial image data received from the user using an "image processing API." Here, a specific technology called the "FaceMatch API" is used to extract the visual features of the face. Based on this, the server utilizes a "generative AI model" to select the most suitable outfit from a clothing database. The selected outfit is then combined with the user's facial image using a "DressAI model" to generate a composite image that provides a realistic try-on experience.
[0457] The device receives a composite image sent from the server and displays it to the user. This image allows the user to compare various outfits without actually trying them on. This process is performed using the device's high-resolution display, making it easy for the user to view the composite image.
[0458] As a concrete example, in the case of a user choosing a furisode (long-sleeved kimono) for their coming-of-age ceremony, the user makes a reservation to try on furisode for the ceremony. The user uploads a facial image, and after processing with an "image processing API" and a "generative AI model," receives a composite image of themselves wearing various furisode. Based on this image, the user selects the most suitable furisode and makes a reservation to try it on from their device. Furthermore, when purchasing or renting, credit card payment can also be made via the device.
[0459] In this way, users can intuitively select costumes even from remote locations and efficiently proceed with costume selection, purchase, or rental procedures.
[0460] An example of a prompt is: "Use the user's facial image to generate a composite image of a kimono that can be tried on. Limit the suggestions to styles suitable for coming-of-age ceremonies and propose designs that are compatible with the user's facial features." This prompt is used to instruct the generation AI model when generating the composite image.
[0461] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0462] Step 1:
[0463] Users access the system using their devices and complete the registration process. During this process, they enter personal information such as their name and email address. This information is sent from the device to the server and stored in the server's database. This creates the user's account.
[0464] Step 2:
[0465] The user takes a picture of their face using the device's camera and uploads it to the system. The device receives this face image and sends it to the server. At this point, the input is the user's face image, and the output is data for storage on the server.
[0466] Step 3:
[0467] The server analyzes the received facial image data using an "image processing API." Specifically, it uses the "FaceMatch API" to extract the visual features of the face. The input to this process is the user's facial image, and the output is the extracted feature data. This clarifies the user's unique visual features.
[0468] Step 4:
[0469] The server selects suitable clothing information from the clothing database based on the extracted visual features. It executes a database query to list clothing that matches the facial features. The input for this step is visual feature data, and the output is the selected clothing information.
[0470] Step 5:
[0471] The server uses a "generative AI model" to combine selected clothing information with the user's facial image to generate a realistic composite image. The AI model matches the clothing to the face, visually recreating the try-on experience. The input for this step is clothing information and a facial image, and the output is a composite image.
[0472] Step 6:
[0473] The terminal receives a composite image sent from the server and presents it to the user. The image is displayed on the terminal's screen, and the user can view it and select their desired outfit. The input is a composite image from the server, and the output is a visual display for the user to see.
[0474] Step 7:
[0475] Users select an outfit from a collection of composite images displayed via their terminal and make a fitting reservation. Purchase and rental procedures are also possible. Selection and transaction information is transmitted from the terminal to a server and stored in a management system. Inputs are the user's selection and transaction information, while outputs are reservation confirmation and payment confirmation.
[0476] (Application Example 1)
[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0478] In online shopping, users often feel anxious about purchasing clothing because it's difficult to try things on in person. Furthermore, there's a lack of efficient ways to try on clothes and make a purchase decision without visiting a physical store. In this situation, there's a need to provide a system that allows users to conveniently and accurately select and purchase clothing in a digital environment.
[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0480] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining the acquired image data with clothing selected from a clothing database, and means for providing the generated composite image to the user and allowing them to select their preferred clothing. This enables the user to efficiently and easily select the clothing best suited to them within a digital virtual store and to consistently carry out procedures from trying on to purchasing or renting.
[0481] "Means for acquiring user image data" refers to the method by which the system receives photographs and image information from the user and transmits it to the processor.
[0482] "A means of generating a composite image by combining clothing selected from a clothing database based on acquired image data" refers to a method of creating a virtual image that makes it appear as if the user is wearing the clothing, using the user's image and clothing information in the database.
[0483] "A means of providing users with generated composite images and allowing them to select their preferred outfit" refers to a method that displays composite try-on images to the user and enables them to choose their preferred outfit from among them.
[0484] "A means of managing and providing fitting reservations or purchase / rental procedures based on selected costume information within a digital virtual store" refers to a method of aggregating information regarding reservations, purchases, or rentals related to selected costumes and efficiently carrying them out online.
[0485] "Performing facial synthesis processing and providing digital try-on via smart devices" refers to the process of processing a user's facial image and providing a virtual try-on experience through the user's smart device.
[0486] "Optimizing the selection of suggested outfits using a generative AI model and generating and presenting prompt sentences" refers to a process that utilizes artificial intelligence technology to suggest outfits that match the user's preferences and outputs related instruction sentences based on the results.
[0487] The embodiment of the invention is based on a system that enables users to have a digital try-on experience in their daily lives using smart devices. This system mainly consists of three elements: a server, a terminal, and the user.
[0488] First, the user uploads their facial image to the system using their device. The device is a smartphone or tablet computer, and its role is to send the image data to the server.
[0489] The servers are built on cloud services such as AWS and Google Cloud Platform. The servers use OpenCV to process the received facial images and, based on the obtained data, utilize TensorFlow, PyTorch, and other tools to generate a composite image combining the user's face and clothing using a pre-trained generative AI model. This composite image functions as a predictive image of what the person would look like wearing the created clothing.
[0490] Furthermore, by providing the generated composite images, the server accumulates user preference data and uses a suggestion algorithm to select even more suitable outfits. These optimized suggestions are then presented to the user again. During this process, the generative AI model generates prompt messages to guide the user to the next action.
[0491] Users can see their avatar trying on outfits through a virtual display. For example, they can easily proceed by receiving prompts such as, "Please upload a user face image," "Next, please select the category of outfit to try on: casual, formal, traditional, etc.," and "Select your favorite outfit based on the generated try-on images and proceed with the purchase process."
[0492] This system allows users to try on various outfits from the comfort of their homes, without having to physically go to a store, and efficiently select their favorites. This digital try-on experience offers a new way for users to enjoy online shopping more in line with their lifestyle.
[0493] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0494] Step 1:
[0495] The user takes a picture of their face using a smart device and uploads it to the system. The system receives the face image as input and sends this image to the server as output.
[0496] Step 2:
[0497] The server uses OpenCV to perform image analysis based on the user's face image received. The input is the user's face image, and the output is the extracted facial features after analysis. Specific operations include facial landmark detection and image preprocessing (resolution adjustment and noise reduction).
[0498] Step 3:
[0499] The server uses TensorFlow to generate a synthesizer image of the clothing based on the analyzed facial features and clothing database. The input is the analyzed facial feature data and clothing database, and the output is the synthesized try-on image. This step involves the integration of digital data by applying clothing to facial features.
[0500] Step 4:
[0501] The server sends the generated composite image to the terminal and displays it to the user. The input is the composite fitting image data, and the output is displayed on the user's device screen. This display allows the user to have a virtual fitting experience.
[0502] Step 5:
[0503] The user views a composite image displayed on their device and selects their preferred outfit. The input is the visual data of the try-on image, and the output is the digital information of the selected outfit. This selection is made through the application's user interface (UI).
[0504] Step 6:
[0505] The server collects information about the outfit selected by the user and generates prompt messages to guide the user's next action. The input is the data of the selected outfit, and the output is the generated prompt message. Here, an AI model is used to generate the prompt message and suggest the next action.
[0506] Step 7:
[0507] Based on the selected outfit, the server manages the fitting reservation or purchase / rental process. Inputs are the user's selection information and registration information, while outputs are reservation confirmations and rental completion statuses. This process is achieved by the backend system updating customer management information.
[0508] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0509] This invention relates to a system that acquires user image data, generates composite images for digitally trying on clothing, and further recognizes the user's emotions to optimize clothing selection. This system improves the accuracy of suggestions by understanding the user's preferences through the integration of an emotion engine.
[0510] The system works as follows: First, the user accesses the system and creates an account. Then, the user uploads their facial image using a device. The device retrieves the image data and sends it to the server.
[0511] The server processes the facial image while simultaneously selecting various outfits from a clothing database and generating a composite image by combining them with the user's facial image. In addition, when the composite image is displayed, the emotion engine evaluates the user's response by analyzing the user's facial expressions in real time through the camera on the terminal or a separately connected device.
[0512] The emotion engine captures subtle changes in the user's facial expressions to infer their emotional state, and based on the emotion data obtained during the suggestion of synthesized images, it provides more suitable clothing options for the user. To achieve this, it combines past preference history and emotion data, and builds a feedback system that learns the user's preference history to improve the accuracy of future suggestions.
[0513] For example, when a user is choosing a kimono for their coming-of-age ceremony, if the emotion engine recognizes feelings of approval or surprise, it can prioritize presenting more vibrant and stylish kimono options based on that emotional state. This system makes the virtual try-on experience through synthesized images even smoother, enabling users to proceed quickly and efficiently with everything from making a fitting reservation to purchasing or renting.
[0514] In this way, the present invention realizes highly accurate clothing selection support that reflects the user's emotions, thereby improving the comfort and satisfaction of the fitting process.
[0515] The following describes the processing flow.
[0516] Step 1:
[0517] The user accesses the system and registers an account. They enter the necessary information (name, email address, password, etc.) to create an account. The terminal sends this information to the server and saves the account data.
[0518] Step 2:
[0519] The user selects a photo using their device to upload a facial image. The device sends the facial image data to the server. The server receives this image information and begins image processing.
[0520] Step 3:
[0521] The server generates a composite image by combining the acquired facial image data with clothing selected from a clothing database. Furthermore, it constructs the composite image to facilitate analysis by the emotion engine.
[0522] Step 4:
[0523] The device displays the generated composite image to the user. At the same time, it activates a function that uses the device's built-in camera to capture the user's facial expressions in real time.
[0524] Step 5:
[0525] The emotion engine analyzes the user's captured facial expression data to identify emotions such as anger and joy. Based on the emotional state, the server filters recommended items to suggest the most suitable outfit for the user.
[0526] Step 6:
[0527] The user selects their preferred outfit from the suggested options. The device sends the selection information to the server, which then uses that information to initiate the process of making a fitting reservation or purchasing / renting the outfit.
[0528] Step 7:
[0529] The server coordinates with the sales / rental provider for the selected costume and schedules the fitting. Once the fitting reservation is complete, the server sends a confirmation email or notification to the user.
[0530] Step 8:
[0531] The server uses feedback from the emotion engine to learn from the acquired emotion data and preference history, and updates the database to improve the accuracy of future suggestions. This update makes it possible to provide more personalized suggestions on subsequent uses.
[0532] (Example 2)
[0533] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0534] In conventional costume fitting systems, user image data is handled statically, making it difficult to provide dynamic costume suggestions based on individual user preferences. Furthermore, the inability to make suggestions that take into account the user's emotional state resulted in a limited user experience and difficulty in improving satisfaction.
[0535] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0536] In this invention, the server includes means for acquiring the user's digital data, means for generating composite data by combining items selected from a database based on the acquired digital data, and means for analyzing the user's facial expressions, estimating their emotional state, and optimizing the items to be suggested. This makes it possible to suggest the optimal clothing that reflects the user's individual emotions and preferences in real time.
[0537] "User digital data" refers to information such as digital images and facial photographs that the system acquires from users.
[0538] A "database" is a collection of information containing diverse item information and options, which a system uses to select items to suggest to users.
[0539] "Synthetic data" refers to digital data generated by combining acquired user digital data with selected item information from a database.
[0540] "Analyzing facial expressions" refers to the process of analyzing the characteristics and movements of a user's face and using that information to infer their emotions and psychological state.
[0541] "Emotional state" refers to the user's current emotional response and psychological state, and is information inferred through facial expression analysis.
[0542] "Optimizing products" refers to the process of adjusting the selection of products offered to the user to be the most appropriate, based on the user's preferences and emotional state.
[0543] To implement this invention, it is necessary to build a system in which a server, terminal, and user work together. The specific operation method is described below.
[0544] Users first access the system using a terminal. The terminal is equipped with a camera, allowing users to take photos of their own face or select and upload existing digital data. This digital data is encrypted and sent from the terminal to the server.
[0545] The server digitally processes the received facial data. Specifically, it extracts facial feature data using an image processing library. Next, the server retrieves data on various items from a database and generates composite data by combining it with the user's facial data. This process uses a generative AI model, primarily employing techniques to achieve style transfer and realistic object synthesis.
[0546] Furthermore, the device's camera captures the user's facial expressions in real time. The server analyzes this facial data to estimate the user's emotional state. Analysis using a deep learning-based emotion recognition model enables optimal product recommendations based on the user's emotions.
[0547] For example, when a user selects an outfit for their coming-of-age ceremony, the server generates composite data and displays it to the user via their device. By analyzing the user's facial expressions and selection history, the system can suggest outfits that match the user's preferences.
[0548] An example of a prompt message is: "I would like to try on a virtual kimono for my coming-of-age ceremony. There are several styles of kimonos available; please display the best option based on the user's preferences."
[0549] In this way, users can obtain an optimal fitting experience tailored to their individual needs, allowing them to make more informed decisions regarding purchases and procedures.
[0550] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0551] Step 1:
[0552] The user uses a terminal to access the system and create a new account. The terminal then sends account data to the server based on the user's input. The input consists of the user's basic information, and the output is the registered account data.
[0553] Step 2:
[0554] The user uses a device to capture or select their own facial data and upload it. The device compresses and encrypts the image data to standardize the format before sending it to the server. Here, the input is the user's facial image data, and the output is the facial data securely transmitted to the server.
[0555] Step 3:
[0556] The server extracts feature points from the received facial data using an image processing library. This involves using a facial recognition algorithm to extract basic facial features such as the positions of the eyes, nose, and mouth. The input is facial image data, and the output is facial feature point data.
[0557] Step 4:
[0558] The server retrieves item data tailored to the user from the database. Here, a generative AI model is used to combine facial feature point data with item data to generate realistic synthetic data. At this stage, the inputs are facial feature point data and item data, and the output is the synthetic data.
[0559] Step 5:
[0560] The device's camera captures the user's facial expressions in real time and sends them to the server. The server receives this data and analyzes the emotional state using an emotion recognition model. The input is real-time facial expression data, and the output is the analyzed emotional state data.
[0561] Step 6:
[0562] The server uses emotional state data to select the most suitable item from synthesized data and outputs it to the user. In this case, the input is emotional state data and synthesized data, and the output is the most suitable item suggestion according to the emotion.
[0563] Step 7:
[0564] The user reviews the suggested items on the terminal and selects the desired items. The terminal sends the selected item information to the server and supports the reservation or purchase process. The input is the user's item selection, and the output is the selection data sent to the server and a notification that the process is complete.
[0565] (Application Example 2)
[0566] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0567] Traditional clothing selection and fitting processes required customers to physically try on garments, which presented problems due to the time and effort involved. Furthermore, online shopping carried the risk of post-purchase dissatisfaction because customers hadn't actually tried on the clothes. Additionally, the inability to adequately consider individual customer preferences and feelings resulted in low customer satisfaction with the selection process.
[0568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0569] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining selected clothing from a clothing database based on the acquired image data, and means for analyzing the user's facial expression data in real time to recognize emotions and optimize the selected clothing. This makes it possible for customers to efficiently select the optimal clothing based on their emotions digitally without having to physically try on clothes.
[0570] "User image data" refers to image information acquired for the purpose of identifying an individual, and is particularly used for the purpose of selecting clothing.
[0571] A "costume database" is a source of information that stores digital information on various costumes and is used to select appropriate costumes based on the user's image data.
[0572] A "composite image" is an image generated by combining the user's image data and clothing data, enabling digital try-on.
[0573] "Real-time analysis" refers to a technical process that processes input data immediately and provides the results almost instantly.
[0574] "Recognizing emotions" refers to the process of analyzing a user's facial expression data to infer their psychological state and preferences.
[0575] "Optimizing selected outfits" means adjusting the selection process to suggest the most appropriate outfits based on the user's emotions and past preference history.
[0576] The system that realizes this invention allows users to digitally try on clothes by standing in front of a smart mirror in a store. The system mainly functions as follows:
[0577] The server acquires the user's facial image from a camera built into the smart mirror. The camera instantly captures the user's face and body in high resolution and sends the data to the server. The server processes the acquired image data using image processing software such as OpenCV or TensorFlow, and combines it with multiple selected clothing data from a clothing database to generate a composite image. This allows the user to virtually try on and check their appearance on the mirror.
[0578] Furthermore, the server uses AI technologies for emotion recognition, such as Azure Cognitive Services and AWS Rekognition, to analyze real-time facial expression data acquired through the camera installed in the smart mirror. This method recognizes the user's emotions and predicts whether the selected outfit matches the user's preferences. Based on the results, the emotion engine optimizes outfit suggestions, presenting the user with more suitable options.
[0579] For example, if a user choosing a kimono for their coming-of-age ceremony stands in front of a mirror and smiles while virtually trying on a red and gold kimono, the system can recognize this expression as joy and present this kimono as a preferred option. An example of a prompt message used as input to the generating AI model might be: "A woman in her 20s is digitally trying on a vibrant and stylish kimono for her coming-of-age ceremony. She has a happy expression as she looks at the red and gold kimono."
[0580] This system allows users to enjoy a highly personalized try-on experience and select their ideal outfit while saving time and effort.
[0581] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0582] Step 1:
[0583] The user stands in front of a smart mirror in the store. The smart mirror's built-in camera captures the user's face and body in high resolution. The input data is the user's image data, which is then sent to the server.
[0584] Step 2:
[0585] The server receives the acquired image data and processes it using image processing software such as OpenCV or TensorFlow. A face recognition algorithm extracts facial feature points, and the server outputs search criteria for the clothing database. Based on these criteria, it selects appropriate clothing data.
[0586] Step 3:
[0587] The server combines the selected costume data with the user's image data to generate a composite image. This process utilizes CGI to output a visual representation that makes it appear as if the user is wearing the costume. The composite image is then mirrored back and provided to the user.
[0588] Step 4:
[0589] The smart mirror, which is the terminal device, displays a composite image. At the same time, the mirror's built-in camera continuously acquires the user's facial expression data in real time. This facial expression data, as input, is sent to a server for emotion recognition.
[0590] Step 5:
[0591] The server uses emotion recognition AI such as Azure Cognitive Services and AWS Rekognition to analyze facial expression data. The analysis results are generated as output data indicating the user's emotional state. Based on this information, the emotion engine selects the optimal outfit and ultimately presents the user with the most suitable options.
[0592] Step 6:
[0593] The user accepts the sentiment analysis and outfit suggestions and makes a final selection. This selection is recorded on the server and used as feedback data to improve the accuracy of suggestions in the future.
[0594] Step 7:
[0595] The server sends the necessary information for the purchase or reservation process to the user's terminal based on the costume selected by the user. The output is provided to the user as a purchase confirmation or reservation completion notification.
[0596] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0597] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0598] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0599] [Fourth Embodiment]
[0600] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0601] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0602] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0603] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0604] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0605] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0606] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0607] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0608] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0609] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0610] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0611] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0612] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0613] This invention is a system that enables users to efficiently select suitable clothing and smoothly proceed through the fitting process. This system primarily functions through the coordinated operation of three elements: a server, a terminal, and the user.
[0614] First, the user accesses the system, completes the necessary registration process, and then uploads a picture of their face. This image is used as the foundation for providing a digital costume try-on experience. The terminal receives the image uploaded by the user and sends it to the server.
[0615] The server processes the received facial image data, selects various outfits from a clothing database, and generates a composite image by combining the outfit with the user's face. This composite image allows the user to virtually try on many different outfits.
[0616] The generated composite image is provided to the user in a viewable format. The user can select an outfit that they think suits them through the composite image. Based on this result, the server saves the selected outfit information and manages the reservation process for actual try-ons, as well as the purchase or rental process for the outfit, if desired.
[0617] As a concrete example, let's consider the case of choosing a furisode (long-sleeved kimono) for a coming-of-age ceremony. In this case, the user wants to try on furisode in order to attend the ceremony. First, the user uploads a facial image to the system. The server generates composite images of the user wearing various furisode based on the user's facial image data and presents them to the user. The user then selects a furisode they like and makes a reservation to try it on. If necessary, they can make a payment on the spot and complete the furisode rental procedure.
[0618] This system allows users to easily find the outfit that suits them best and efficiently go through the entire process from trying it on to purchasing or renting it.
[0619] The following describes the processing flow.
[0620] Step 1:
[0621] The user accesses the system and first registers an account. They enter the required information (name, email address, password, etc.) to create an account. The terminal collects this information and sends the account data to the server. The server stores the received information in its database and creates the account.
[0622] Step 2:
[0623] After registering an account, the user uploads a facial image. This process utilizes an image selection tool. The device sends the uploaded image data to the server. The server receives the image data and prepares it for image processing.
[0624] Step 3:
[0625] The server inputs the received facial image data into an AI model and generates a composite image by combining it with clothing selected from a clothing database. The AI model performs high-precision image processing to digitally composite various clothing onto the user's face.
[0626] Step 4:
[0627] The server sends the generated composite images to the terminal and presents them to the user in a viewable format. The user reviews these composite images and selects their favorite outfit. The selection process is performed on the terminal, and the selection results are sent back to the server.
[0628] Step 5:
[0629] The server stores the selected costume information and provides an interface for making fitting reservations and purchasing / renting items. Users select a fitting date and time and make a reservation. The server processes this information and coordinates the schedule with partner sales / rental businesses.
[0630] Step 6:
[0631] Once the reservation is complete, the server sends a confirmation notice to the user. If the user wishes to purchase or rent, the server provides a secure payment environment and completes the transaction. The user receives confirmation that the process is complete and can then finish using the service.
[0632] (Example 1)
[0633] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0634] The process of selecting and trying on clothing requires time and effort for users to try on many options. Physical try-ons require travel to stores and making appointments, which is inconvenient. Furthermore, the limited variety of clothing available for try-on makes it difficult to make the best clothing choice.
[0635] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0636] In this invention, the server includes means for acquiring and storing the user's image information, means for analyzing visual features based on the acquired image information to select clothing information from a data set and generating a composite image using an image synthesis model, and means for presenting the generated composite image to the user and presenting selectable clothing information. As a result, the user can choose from a variety of clothing without physically trying them on, enabling them to make the optimal clothing selection while significantly reducing time and effort.
[0637] "User image information" refers to visual data that individuals provide to the system, and this data is used to analyze the characteristics of the users.
[0638] "Visual features" refer to information about appearance and shape extracted from image data, and represent the individual characteristics of the user necessary for generating a composite image.
[0639] A "data set" is a database containing diverse clothing information, used to select the most suitable outfit based on the user's image information.
[0640] "Costume information" refers to data related to clothing, including detailed information such as the design, color, and material of the garments.
[0641] An "image synthesis model" refers to an algorithm or program used to generate a realistic composite image by combining image information and clothing information acquired using AI technology.
[0642] A "composite image" is a simulated image created based on the user's image information and selected clothing information, visually representing the try-on state.
[0643] "To present" refers to the act of showing generated information or images to the user through a user interface, in order to enable them to make choices.
[0644] "Try-on procedure" refers to the process of making reservations and arrangements for users to actually try on the outfits they have selected.
[0645] "Commercial transaction procedures" refer to the series of activities involved in processing payments and contracts necessary when a user purchases or rents an outfit they have selected.
[0646] The following describes embodiments for carrying out this system invention.
[0647] The present invention provides a system for users to efficiently select suitable clothing digitally. This system primarily consists of three elements: the user, the terminal, and the server.
[0648] Users first access the system using their device and register by entering the required information. This includes providing basic personal information. Afterward, users use their device's camera function to take a picture of their face and upload it to the system. This face image serves as the basis for generating a composite image of the costume.
[0649] The device is equipped with the ability to receive facial images uploaded by users and quickly transmit them to the server. This communication process is supported by a high-speed and stable internet connection.
[0650] The server analyzes the facial image data received from the user using an "image processing API." Here, a specific technology called the "FaceMatch API" is used to extract the visual features of the face. Based on this, the server utilizes a "generative AI model" to select the most suitable outfit from a clothing database. The selected outfit is then combined with the user's facial image using a "DressAI model" to generate a composite image that provides a realistic try-on experience.
[0651] The device receives a composite image sent from the server and displays it to the user. This image allows the user to compare various outfits without actually trying them on. This process is performed using the device's high-resolution display, making it easy for the user to view the composite image.
[0652] As a concrete example, in the case of a user choosing a furisode (long-sleeved kimono) for their coming-of-age ceremony, the user makes a reservation to try on furisode for the ceremony. The user uploads a facial image, and after processing with an "image processing API" and a "generative AI model," receives a composite image of themselves wearing various furisode. Based on this image, the user selects the most suitable furisode and makes a reservation to try it on from their device. Furthermore, when purchasing or renting, credit card payment can also be made via the device.
[0653] In this way, users can intuitively select costumes even from remote locations and efficiently proceed with costume selection, purchase, or rental procedures.
[0654] An example of a prompt is: "Use the user's facial image to generate a composite image of a kimono that can be tried on. Limit the suggestions to styles suitable for coming-of-age ceremonies and propose designs that are compatible with the user's facial features." This prompt is used to instruct the generation AI model when generating the composite image.
[0655] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0656] Step 1:
[0657] Users access the system using their devices and complete the registration process. During this process, they enter personal information such as their name and email address. This information is sent from the device to the server and stored in the server's database. This creates the user's account.
[0658] Step 2:
[0659] The user takes a picture of their face using the device's camera and uploads it to the system. The device receives this face image and sends it to the server. At this point, the input is the user's face image, and the output is data for storage on the server.
[0660] Step 3:
[0661] The server analyzes the received facial image data using an "image processing API." Specifically, it uses the "FaceMatch API" to extract the visual features of the face. The input to this process is the user's facial image, and the output is the extracted feature data. This clarifies the user's unique visual features.
[0662] Step 4:
[0663] The server selects suitable clothing information from the clothing database based on the extracted visual features. It executes a database query to list clothing that matches the facial features. The input for this step is visual feature data, and the output is the selected clothing information.
[0664] Step 5:
[0665] The server uses a "generative AI model" to combine selected clothing information with the user's facial image to generate a realistic composite image. The AI model matches the clothing to the face, visually recreating the try-on experience. The input for this step is clothing information and a facial image, and the output is a composite image.
[0666] Step 6:
[0667] The terminal receives a composite image sent from the server and presents it to the user. The image is displayed on the terminal's screen, and the user can view it and select their desired outfit. The input is a composite image from the server, and the output is a visual display for the user to see.
[0668] Step 7:
[0669] Users select an outfit from a collection of composite images displayed via their terminal and make a fitting reservation. Purchase and rental procedures are also possible. Selection and transaction information is transmitted from the terminal to a server and stored in a management system. Inputs are the user's selection and transaction information, while outputs are reservation confirmation and payment confirmation.
[0670] (Application Example 1)
[0671] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0672] In online shopping, users often feel anxious about purchasing clothing because it's difficult to try things on in person. Furthermore, there's a lack of efficient ways to try on clothes and make a purchase decision without visiting a physical store. In this situation, there's a need to provide a system that allows users to conveniently and accurately select and purchase clothing in a digital environment.
[0673] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0674] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining the acquired image data with clothing selected from a clothing database, and means for providing the generated composite image to the user and allowing them to select their preferred clothing. This enables the user to efficiently and easily select the clothing best suited to them within a digital virtual store and to consistently carry out procedures from trying on to purchasing or renting.
[0675] "Means for acquiring user image data" refers to the method by which the system receives photographs and image information from the user and transmits it to the processor.
[0676] "A means of generating a composite image by combining clothing selected from a clothing database based on acquired image data" refers to a method of creating a virtual image that makes it appear as if the user is wearing the clothing, using the user's image and clothing information in the database.
[0677] "A means of providing users with generated composite images and allowing them to select their preferred outfit" refers to a method that displays composite try-on images to the user and enables them to choose their preferred outfit from among them.
[0678] "A means of managing and providing fitting reservations or purchase / rental procedures based on selected costume information within a digital virtual store" refers to a method of aggregating information regarding reservations, purchases, or rentals related to selected costumes and efficiently carrying them out online.
[0679] "Performing facial synthesis processing and providing digital try-on via smart devices" refers to the process of processing a user's facial image and providing a virtual try-on experience through the user's smart device.
[0680] "Optimizing the selection of suggested outfits using a generative AI model and generating and presenting prompt sentences" refers to a process that utilizes artificial intelligence technology to suggest outfits that match the user's preferences and outputs related instruction sentences based on the results.
[0681] The embodiment of the invention is based on a system that enables users to have a digital try-on experience in their daily lives using smart devices. This system mainly consists of three elements: a server, a terminal, and the user.
[0682] First, the user uploads their facial image to the system using their device. The device is a smartphone or tablet computer, and its role is to send the image data to the server.
[0683] The servers are built on cloud services such as AWS and Google Cloud Platform. The servers use OpenCV to process the received facial images and, based on the obtained data, utilize TensorFlow, PyTorch, and other tools to generate a composite image combining the user's face and clothing using a pre-trained generative AI model. This composite image functions as a predictive image of what the person would look like wearing the created clothing.
[0684] Furthermore, by providing the generated composite images, the server accumulates user preference data and uses a suggestion algorithm to select even more suitable outfits. These optimized suggestions are then presented to the user again. During this process, the generative AI model generates prompt messages to guide the user to the next action.
[0685] Users can see their avatar trying on outfits through a virtual display. For example, they can easily proceed by receiving prompts such as, "Please upload a user face image," "Next, please select the category of outfit to try on: casual, formal, traditional, etc.," and "Select your favorite outfit based on the generated try-on images and proceed with the purchase process."
[0686] This system allows users to try on various outfits from the comfort of their homes, without having to physically go to a store, and efficiently select their favorites. This digital try-on experience offers a new way for users to enjoy online shopping more in line with their lifestyle.
[0687] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0688] Step 1:
[0689] The user takes a picture of their face using a smart device and uploads it to the system. The system receives the face image as input and sends this image to the server as output.
[0690] Step 2:
[0691] The server uses OpenCV to perform image analysis based on the user's face image received. The input is the user's face image, and the output is the extracted facial features after analysis. Specific operations include facial landmark detection and image preprocessing (resolution adjustment and noise reduction).
[0692] Step 3:
[0693] The server uses TensorFlow to generate a synthesizer image of the clothing based on the analyzed facial features and clothing database. The input is the analyzed facial feature data and clothing database, and the output is the synthesized try-on image. This step involves the integration of digital data by applying clothing to facial features.
[0694] Step 4:
[0695] The server sends the generated composite image to the terminal and displays it to the user. The input is the composite fitting image data, and the output is displayed on the user's device screen. This display allows the user to have a virtual fitting experience.
[0696] Step 5:
[0697] The user views a composite image displayed on their device and selects their preferred outfit. The input is the visual data of the try-on image, and the output is the digital information of the selected outfit. This selection is made through the application's user interface (UI).
[0698] Step 6:
[0699] The server collects information about the outfit selected by the user and generates prompt messages to guide the user's next action. The input is the data of the selected outfit, and the output is the generated prompt message. Here, an AI model is used to generate the prompt message and suggest the next action.
[0700] Step 7:
[0701] Based on the selected outfit, the server manages the fitting reservation or purchase / rental process. Inputs are the user's selection information and registration information, while outputs are reservation confirmations and rental completion statuses. This process is achieved by the backend system updating customer management information.
[0702] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0703] This invention relates to a system that acquires user image data, generates composite images for digitally trying on clothing, and further recognizes the user's emotions to optimize clothing selection. This system improves the accuracy of suggestions by understanding the user's preferences through the integration of an emotion engine.
[0704] The system works as follows: First, the user accesses the system and creates an account. Then, the user uploads their facial image using a device. The device retrieves the image data and sends it to the server.
[0705] The server processes the facial image while simultaneously selecting various outfits from a clothing database and generating a composite image by combining them with the user's facial image. In addition, when the composite image is displayed, the emotion engine evaluates the user's response by analyzing the user's facial expressions in real time through the camera on the terminal or a separately connected device.
[0706] The emotion engine captures subtle changes in the user's facial expressions to infer their emotional state, and based on the emotion data obtained during the suggestion of synthesized images, it provides more suitable clothing options for the user. To achieve this, it combines past preference history and emotion data, and builds a feedback system that learns the user's preference history to improve the accuracy of future suggestions.
[0707] For example, when a user is choosing a kimono for their coming-of-age ceremony, if the emotion engine recognizes feelings of approval or surprise, it can prioritize presenting more vibrant and stylish kimono options based on that emotional state. This system makes the virtual try-on experience through synthesized images even smoother, enabling users to proceed quickly and efficiently with everything from making a fitting reservation to purchasing or renting.
[0708] In this way, the present invention realizes highly accurate clothing selection support that reflects the user's emotions, thereby improving the comfort and satisfaction of the fitting process.
[0709] The following describes the processing flow.
[0710] Step 1:
[0711] The user accesses the system and registers an account. They enter the necessary information (name, email address, password, etc.) to create an account. The terminal sends this information to the server and saves the account data.
[0712] Step 2:
[0713] The user selects a photo using their device to upload a facial image. The device sends the facial image data to the server. The server receives this image information and begins image processing.
[0714] Step 3:
[0715] The server generates a composite image by combining the acquired facial image data with clothing selected from a clothing database. Furthermore, it constructs the composite image to facilitate analysis by the emotion engine.
[0716] Step 4:
[0717] The device displays the generated composite image to the user. At the same time, it activates a function that uses the device's built-in camera to capture the user's facial expressions in real time.
[0718] Step 5:
[0719] The emotion engine analyzes the user's captured facial expression data to identify emotions such as anger and joy. Based on the emotional state, the server filters recommended items to suggest the most suitable outfit for the user.
[0720] Step 6:
[0721] The user selects their preferred outfit from the suggested options. The device sends the selection information to the server, which then uses that information to initiate the process of making a fitting reservation or purchasing / renting the outfit.
[0722] Step 7:
[0723] The server coordinates with the sales / rental provider for the selected costume and schedules the fitting. Once the fitting reservation is complete, the server sends a confirmation email or notification to the user.
[0724] Step 8:
[0725] The server uses feedback from the emotion engine to learn from the acquired emotion data and preference history, and updates the database to improve the accuracy of future suggestions. This update makes it possible to provide more personalized suggestions on subsequent uses.
[0726] (Example 2)
[0727] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0728] In conventional costume fitting systems, user image data is handled statically, making it difficult to provide dynamic costume suggestions based on individual user preferences. Furthermore, the inability to make suggestions that take into account the user's emotional state resulted in a limited user experience and difficulty in improving satisfaction.
[0729] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0730] In this invention, the server includes means for acquiring the user's digital data, means for generating composite data by combining items selected from a database based on the acquired digital data, and means for analyzing the user's facial expressions, estimating their emotional state, and optimizing the items to be suggested. This makes it possible to suggest the optimal clothing that reflects the user's individual emotions and preferences in real time.
[0731] "User digital data" refers to information such as digital images and facial photographs that the system acquires from users.
[0732] A "database" is a collection of information containing diverse item information and options, which a system uses to select items to suggest to users.
[0733] "Synthetic data" refers to digital data generated by combining acquired user digital data with selected item information from a database.
[0734] "Analyzing facial expressions" refers to the process of analyzing the characteristics and movements of a user's face and using that information to infer their emotions and psychological state.
[0735] "Emotional state" refers to the user's current emotional response and psychological state, and is information inferred through facial expression analysis.
[0736] "Optimizing products" refers to the process of adjusting the selection of products offered to the user to be the most appropriate, based on the user's preferences and emotional state.
[0737] To implement this invention, it is necessary to build a system in which a server, terminal, and user work together. The specific operation method is described below.
[0738] Users first access the system using a terminal. The terminal is equipped with a camera, allowing users to take photos of their own face or select and upload existing digital data. This digital data is encrypted and sent from the terminal to the server.
[0739] The server digitally processes the received facial data. Specifically, it extracts facial feature data using an image processing library. Next, the server retrieves data on various items from a database and generates composite data by combining it with the user's facial data. This process uses a generative AI model, primarily employing techniques to achieve style transfer and realistic object synthesis.
[0740] Furthermore, the device's camera captures the user's facial expressions in real time. The server analyzes this facial data to estimate the user's emotional state. Analysis using a deep learning-based emotion recognition model enables optimal product recommendations based on the user's emotions.
[0741] For example, when a user selects an outfit for their coming-of-age ceremony, the server generates composite data and displays it to the user via their device. By analyzing the user's facial expressions and selection history, the system can suggest outfits that match the user's preferences.
[0742] An example of a prompt message is: "I would like to try on a virtual kimono for my coming-of-age ceremony. There are several styles of kimonos available; please display the best option based on the user's preferences."
[0743] In this way, users can obtain an optimal fitting experience tailored to their individual needs, allowing them to make more informed decisions regarding purchases and procedures.
[0744] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0745] Step 1:
[0746] The user uses a terminal to access the system and create a new account. The terminal then sends account data to the server based on the user's input. The input consists of the user's basic information, and the output is the registered account data.
[0747] Step 2:
[0748] The user uses a device to capture or select their own facial data and upload it. The device compresses and encrypts the image data to standardize the format before sending it to the server. Here, the input is the user's facial image data, and the output is the facial data securely transmitted to the server.
[0749] Step 3:
[0750] The server extracts feature points from the received facial data using an image processing library. This involves using a facial recognition algorithm to extract basic facial features such as the positions of the eyes, nose, and mouth. The input is facial image data, and the output is facial feature point data.
[0751] Step 4:
[0752] The server retrieves item data tailored to the user from the database. Here, a generative AI model is used to combine facial feature point data with item data to generate realistic synthetic data. At this stage, the inputs are facial feature point data and item data, and the output is the synthetic data.
[0753] Step 5:
[0754] The device's camera captures the user's facial expressions in real time and sends them to the server. The server receives this data and analyzes the emotional state using an emotion recognition model. The input is real-time facial expression data, and the output is the analyzed emotional state data.
[0755] Step 6:
[0756] The server uses emotional state data to select the most suitable item from synthesized data and outputs it to the user. In this case, the input is emotional state data and synthesized data, and the output is the most suitable item suggestion according to the emotion.
[0757] Step 7:
[0758] The user reviews the suggested items on the terminal and selects the desired items. The terminal sends the selected item information to the server and supports the reservation or purchase process. The input is the user's item selection, and the output is the selection data sent to the server and a notification that the process is complete.
[0759] (Application Example 2)
[0760] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] Traditional clothing selection and fitting processes required customers to physically try on garments, which presented problems due to the time and effort involved. Furthermore, online shopping carried the risk of post-purchase dissatisfaction because customers hadn't actually tried on the clothes. Additionally, the inability to adequately consider individual customer preferences and feelings resulted in low customer satisfaction with the selection process.
[0762] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0763] In this invention, the server includes means for acquiring user image data, means for generating a composite image by combining selected clothing from a clothing database based on the acquired image data, and means for analyzing the user's facial expression data in real time to recognize emotions and optimize the selected clothing. This makes it possible for customers to efficiently select the optimal clothing based on their emotions digitally without having to physically try on clothes.
[0764] "User image data" refers to image information acquired for the purpose of identifying an individual, and is particularly used for the purpose of selecting clothing.
[0765] A "costume database" is a source of information that stores digital information on various costumes and is used to select appropriate costumes based on the user's image data.
[0766] A "composite image" is an image generated by combining the user's image data and clothing data, enabling digital try-on.
[0767] "Real-time analysis" refers to a technical process that processes input data immediately and provides the results almost instantly.
[0768] "Recognizing emotions" refers to the process of analyzing a user's facial expression data to infer their psychological state and preferences.
[0769] "Optimizing selected outfits" means adjusting the selection process to suggest the most appropriate outfits based on the user's emotions and past preference history.
[0770] The system that realizes this invention allows users to digitally try on clothes by standing in front of a smart mirror in a store. The system mainly functions as follows:
[0771] The server acquires the user's facial image from a camera built into the smart mirror. The camera instantly captures the user's face and body in high resolution and sends the data to the server. The server processes the acquired image data using image processing software such as OpenCV or TensorFlow, and combines it with multiple selected clothing data from a clothing database to generate a composite image. This allows the user to virtually try on and check their appearance on the mirror.
[0772] Furthermore, the server uses AI technologies for emotion recognition, such as Azure Cognitive Services and AWS Rekognition, to analyze real-time facial expression data acquired through the camera installed in the smart mirror. This method recognizes the user's emotions and predicts whether the selected outfit matches the user's preferences. Based on the results, the emotion engine optimizes outfit suggestions, presenting the user with more suitable options.
[0773] For example, if a user choosing a kimono for their coming-of-age ceremony stands in front of a mirror and smiles while virtually trying on a red and gold kimono, the system can recognize this expression as joy and present this kimono as a preferred option. An example of a prompt message used as input to the generating AI model might be: "A woman in her 20s is digitally trying on a vibrant and stylish kimono for her coming-of-age ceremony. She has a happy expression as she looks at the red and gold kimono."
[0774] This system allows users to enjoy a highly personalized try-on experience and select their ideal outfit while saving time and effort.
[0775] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0776] Step 1:
[0777] The user stands in front of a smart mirror in the store. The smart mirror's built-in camera captures the user's face and body in high resolution. The input data is the user's image data, which is then sent to the server.
[0778] Step 2:
[0779] The server receives the acquired image data and processes it using image processing software such as OpenCV or TensorFlow. A face recognition algorithm extracts facial feature points, and the server outputs search criteria for the clothing database. Based on these criteria, it selects appropriate clothing data.
[0780] Step 3:
[0781] The server combines the selected costume data with the user's image data to generate a composite image. This process utilizes CGI to output a visual representation that makes it appear as if the user is wearing the costume. The composite image is then mirrored back and provided to the user.
[0782] Step 4:
[0783] The smart mirror, which is the terminal device, displays a composite image. At the same time, the mirror's built-in camera continuously acquires the user's facial expression data in real time. This facial expression data, as input, is sent to a server for emotion recognition.
[0784] Step 5:
[0785] The server uses emotion recognition AI such as Azure Cognitive Services and AWS Rekognition to analyze facial expression data. The analysis results are generated as output data indicating the user's emotional state. Based on this information, the emotion engine selects the optimal outfit and ultimately presents the user with the most suitable options.
[0786] Step 6:
[0787] The user accepts the sentiment analysis and outfit suggestions and makes a final selection. This selection is recorded on the server and used as feedback data to improve the accuracy of suggestions in the future.
[0788] Step 7:
[0789] The server sends the necessary information for the purchase or reservation process to the user's terminal based on the costume selected by the user. The output is provided to the user as a purchase confirmation or reservation completion notification.
[0790] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0791] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0792] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0793] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0794] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0795] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0796] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0797] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0798] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0799] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0800] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0801] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0802] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0803] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0804] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0805] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0806] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0807] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0808] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0809] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0810] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0811] The following is further disclosed regarding the embodiments described above.
[0812] (Claim 1)
[0813] Means for obtaining user image data,
[0814] A means for generating a composite image by combining clothing selected from a clothing database based on acquired image data,
[0815] A means of providing the user with a generated composite image and allowing them to select their preferred outfit,
[0816] A means of managing the reservation of a fitting or the purchase / rental procedure based on the information of the selected costume,
[0817] A system that includes this.
[0818] (Claim 2)
[0819] The system according to claim 1, wherein a facial image is used as the user's image data and facial synthesis processing is performed.
[0820] (Claim 3)
[0821] The system according to claim 1, wherein the system learns the user's preferences through the provision of generated composite images and optimizes the selection of suggested costumes.
[0822] "Example 1"
[0823] (Claim 1)
[0824] A means of acquiring and saving the user's image information,
[0825] A means for analyzing visual features based on acquired image information, selecting costume information from a data set, and generating a composite image using an image synthesis model,
[0826] A means of presenting the generated composite image to the user and showing selectable costume information,
[0827] A means of managing individual fitting and commercial transaction procedures based on the costume information selected by the user,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, wherein a facial image is used as the user's image information, and a visual and clothing synthesis process is performed.
[0831] (Claim 3)
[0832] The system according to claim 1, wherein the system collects the user's preference trends through the presentation of generated composite images and enhances the selection of recommended clothing information.
[0833] "Application Example 1"
[0834] (Claim 1)
[0835] Means for obtaining user image data,
[0836] A means for generating a composite image by combining clothing selected from a clothing database based on acquired image data,
[0837] A means of providing the user with a generated composite image and allowing them to select their preferred outfit,
[0838] Based on the information of the selected costume, the system manages the reservation of a fitting or the purchase / rental process, and provides this service within a digital virtual store.
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, wherein a facial image is used as the user's image data, facial synthesis processing is performed, and digital try-on is provided via a smart device.
[0842] (Claim 3)
[0843] The system according to claim 1, wherein the system learns the user's preferences through the provision of generated composite images, optimizes the selection of suggested clothing using a generation AI model, and generates and presents prompt sentences.
[0844] "Example 2 of combining an emotion engine"
[0845] (Claim 1)
[0846] Means of acquiring users' digital data,
[0847] A means for generating composite data by combining items selected from a database based on acquired digital data,
[0848] A means of providing the generated synthetic data to the user and allowing them to select their preferred items,
[0849] A means of managing reservations or procedures based on information about selected items,
[0850] A means for analyzing the user's facial expressions, estimating their emotional state, and optimizing the suggested items,
[0851] A system that includes this.
[0852] (Claim 2)
[0853] The system according to claim 1, wherein facial data is used as the user's digital data and a synthesis process is performed.
[0854] (Claim 3)
[0855] The system according to claim 1, wherein the system learns the user's preferences through the provision of generated synthetic data and dynamically changes the selection of suggested items.
[0856] "Application example 2 when combining with an emotional engine"
[0857] (Claim 1)
[0858] Means for obtaining user image data,
[0859] A means for generating a composite image by combining costumes selected from a costume database based on acquired image data,
[0860] A means of providing the user with a generated composite image and allowing them to select their preferred outfit,
[0861] A method for analyzing user facial expression data in real time to recognize emotions and optimize the selection of clothing,
[0862] A means of managing the process of booking a fitting or purchasing / renting an outfit based on optimized outfit information,
[0863] A system that includes this.
[0864] (Claim 2)
[0865] The system according to claim 1, wherein a facial image is used as the user's image data and facial synthesis processing is performed.
[0866] (Claim 3)
[0867] The system according to claim 1, wherein the system learns the user's preferences and emotions based on facial expressions through the provision of generated composite images, and optimizes the selection of suggested clothing with higher accuracy. [Explanation of symbols]
[0868] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for obtaining user image data, A means for generating a composite image by combining clothing selected from a clothing database based on acquired image data, A means of providing the user with a generated composite image and allowing them to select their preferred outfit, A means of managing the reservation of a fitting or the purchase / rental procedure based on the information of the selected costume, A system that includes this.
2. The system according to claim 1, wherein a facial image is used as the user's image data and a facial synthesis process is performed.
3. The system according to claim 1, wherein the system learns the user's preferences through the provision of generated composite images and optimizes the selection of suggested costumes.