System
The system addresses online shopping anxieties by using a smart mirror with body data capture, voice commands, and generative AI to simulate trying on clothes, enhancing the shopping experience and satisfaction.
Patent Information
- Application Number
- JP2024122726
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Users face challenges in online shopping, particularly with clothes, as they cannot try them on before purchasing, leading to issues with size and fit, and lack the real-life experience, causing anxiety.
A system utilizing a smart mirror with a camera for body data capture, microphone for voice commands, display for selection menus, generative AI for try-on images, and communication for seamless purchases, combined with authentication and online shop APIs, to simulate trying on clothes and facilitate purchases.
Provides a realistic online shopping experience similar to a physical store, reducing anxiety and improving user satisfaction by allowing users to try on clothes virtually and complete purchases efficiently.
Smart Images

Figure 2026021044000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With traditional online shopping, users often purchase clothes without actually trying them on, which can lead to problems such as the wrong size or clothes not matching what they imagined. This problem is particularly serious for users who do not have a physical store of a particular brand nearby. Furthermore, online shopping makes it difficult to get the same kind of real-life experience as trying on clothes before purchasing, which can lead to anxiety about purchasing. There is a need to solve these issues and provide an environment where users can purchase clothes online with peace of mind. [Means for solving the problem]
[0005] The present invention solves these problems by providing a system that includes a camera means for acquiring a user's body type data, a microphone means for recognizing the user's voice instructions, a display means for displaying a clothing selection menu based on the user's voice instructions, a communication means for sending information about the clothing selected by the user to a generation AI, a generation AI means for generating a fitting image of the user wearing the selected clothing, a display means for displaying the generated fitting image on a mirror, and a communication means for completing the purchase process based on the user's voice instructions.Furthermore, by adding an authentication means for authenticating the user and a communication means for acquiring clothing information through an online shop's API, a seamless shopping experience is realized, allowing users to try on clothes with an atmosphere similar to that of a real store and make a purchase with peace of mind.
[0006] "User" means any person who uses the System to shop online.
[0007] "Body data" is information about the user's physique and dimensions, obtained through the camera.
[0008] "Camera Means" refers to a device for capturing a user's body shape data.
[0009] "Voice Commands" means verbal commands or requests given by the User through a microphone.
[0010] "Microphone Means" refers to a device for recognizing a user's voice instructions.
[0011] "Display means" refers to a device used to display information, menus, try-on images, etc. to users.
[0012] "Communication means" refers to technology for sending and receiving data, and has the function of sending information about the clothing selected by the user to the generation AI or server.
[0013] "Generative AI" refers to artificial intelligence technology that generates try-on images of the clothes a user selects.
[0014] "Generative AI means" refers to a system or device that uses generative AI to generate images of users trying on clothes.
[0015] "Authentication Method" means a system or process capable of verifying a user's identity.
[0016] "Online Shop API" means the application program interface for accessing the information and services of the Online Shop.
[0017] "Checkout" means the process a User goes through to purchase the selected Product. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention provides an online shopping support system using a smart full-length mirror, which allows users to experience the experience of trying on clothes in real life. Specific embodiments of the present invention will be described below.
[0040] 1. Initial Setup
[0041] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, and speaker and verifies that they are working properly. The camera captures the user's body shape data, the microphone receives the user's voice instructions, and the speaker provides voice responses.
[0042] 2. User Identification and Login
[0043] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0044] 3. Clothing choices
[0045] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0046] 4. Try-on simulation
[0047] When a user selects an item, the device sends the clothing information to the AI generator, which then generates an image of the selected clothing based on the user's body type data. This image is displayed in real time on the device's mirror, making it appear as if the user is actually trying it on.
[0048] 5. Purchase Procedure
[0049] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0050] Specific examples
[0051] For example, let's say a user wants to try on a denim jacket. When the user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about that denim jacket is sent to the generation AI, which generates a fitting image adapted to the user's body type. This fitting image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user wants to purchase it, the device recognizes the command and starts the purchase process by saying, "I want to buy this jacket."
[0052] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, thereby eliminating concerns about online shopping and increasing user satisfaction.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0056] Step 2:
[0057] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0058] Step 3:
[0059] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0060] Step 4:
[0061] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0062] Step 5:
[0063] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0064] Step 6:
[0065] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0066] Step 7:
[0067] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0068] Step 8:
[0069] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0070] Step 9:
[0071] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0072] Step 10:
[0073] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0074] Step 11:
[0075] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0076] Step 12:
[0077] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0078] Step 13:
[0079] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] With traditional online shopping, customers could not actually try on the products, which meant they could not check the size or fit, which often led to anxiety when making a purchase. Furthermore, it was not possible to provide optimal suggestions for individual users, so diverse needs could not be met. Furthermore, the purchasing process was complicated, making it difficult to improve the user experience.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes an image acquisition means for acquiring a user's body type data, a voice recognition means for recognizing the user's voice instructions, a display means for displaying a product selection menu based on the user's voice instructions, a communication means for sending information about the product selected by the user to a generative AI model, a generative AI means for generating try-on images of the user wearing the selected product, a display means for displaying the generated try-on images on a display device, and a communication means for carrying out the purchase procedure based on the user's voice instructions. This allows the user to have an experience that feels like they are actually trying on the product, eliminating the anxiety of online shopping and enabling individually optimized suggestions, making the purchase procedure smoother.
[0085] "Image acquisition means" refers to devices or technologies for acquiring a user's body shape data, and specifically includes cameras and 3D scanners.
[0086] "Voice recognition means" refers to devices or technology for recognizing a user's voice instructions, and specifically includes a microphone and voice recognition software.
[0087] "Display means" refers to devices or technologies for displaying a user interface, and specifically includes displays and touch screens.
[0088] "Communication means" refers to devices and technologies for transmitting and receiving data, and specifically includes internet connections and wireless communication technologies.
[0089] A "generative AI model" is an artificial intelligence model that generates try-on images based on the user's body type data and selected product information.
[0090] "Generative AI means" refers to technology or equipment for generating try-on images from specified data using a generative AI model.
[0091] The "display device" refers to a device for visually presenting the generated try-on images to the user, and specifically includes a smart mirror or display.
[0092] "Authentication means" refers to devices or technologies for authenticating users, and specifically includes facial recognition systems and login systems.
[0093] A "database" is a system for organizing and storing information, and specifically includes a system for storing information about products in an online shop.
[0094] The present invention is an online shopping support system that provides users with an experience similar to trying on clothes in a physical store. This system acquires the user's body type data, selects products based on voice instructions, generates try-on images using a generative AI model, and performs the entire process from purchase to purchase. Specific embodiments are described below.
[0095] The server prepares a database to manage user account information, past purchase history, and product data, enabling personalized recommendations to be made to users. The database used can be a commercial or open-source relational database management system such as MySQL or PostgreSQL.
[0096] The device is a smart mirror and is configured as follows: the camera is used to acquire the user's body shape data, and can be, for example, a high-performance webcam or a 3D scanner; the microphone is used to receive the user's voice instructions, and is preferably, for example, a directional microphone or a microphone with noise-canceling capabilities; and the speaker is used to provide voice feedback.
[0097] When a user stands in front of the smart mirror, the device recognizes the user's face through the camera. The facial recognition technology uses libraries and frameworks such as OpenCV and TensorFlow. Once authentication is complete, the device displays a login screen, allowing the user to enter login information via voice commands or the touch panel.
[0098] When a user says, "I want to choose clothes," the device's microphone captures the voice command and analyzes it using the Google Speech-to-Text API. Based on the analysis results, the device displays a selection menu, allowing the user to select a specific category or brand. Once the user confirms their selection, the server retrieves product information from the online store's database and sends it to the device.
[0099] Next, the user selects the product they want to try on, and the device sends that product information to a generative AI model. The generative AI model uses OpenAI's generative AI technology, for example. This model generates a fitting image based on the user's body data. The generated fitting image is displayed in real time on the smart mirror, allowing the user to experience the experience of actually trying on the product.
[0100] For example, if a user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about the denim jacket is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is instantly displayed on the smart mirror, allowing the user to check the denim jacket as if they were trying it on. When the user says, "I want to buy this jacket," the device recognizes the voice command and begins the purchase process. The server uses the online shop API to process the payment and sends a confirmation message to the device.
[0101] As described above, the online shopping support system of the present invention allows users to enjoy a realistic try-on experience from the comfort of their own home, thereby eliminating concerns about online shopping and contributing to improved user satisfaction.
[0102] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0103] Step 1:
[0104] Database Setup
[0105] The server prepares a database and manages user account information, past purchase history, and product data. Database management systems such as MySQL and PostgreSQL are used for this. Specifically, the server creates tables for registering new users and for storing detailed product information.
[0106] Input: User information, purchase history, product data
[0107] Data processing: storing data in a database
[0108] Output: Database completed, data stored
[0109] Step 2:
[0110] Initial Hardware Setup
[0111] The device's camera, microphone, and speaker are initially configured. At this time, the camera resolution and frame rate are set, the microphone sensitivity is adjusted, and the speaker volume is set. For the camera, a high-performance webcam or 3D scanner is used, and for the microphone, a directional microphone with noise-canceling functions is used. For the speaker, a speaker that can output clear audio is selected.
[0112] Input: Hardware configuration information
[0113] Data processing: Applying setting parameters
[0114] Output: Camera, microphone, and speaker initial settings completed
[0115] Step 3:
[0116] User facial recognition and login
[0117] When a user stands in front of the smart mirror, the device's camera recognizes the user's face. Facial recognition technology uses OpenCV and TensorFlow. If facial recognition is successful, the device displays a login screen and the user enters their login information. The server receives this and authenticates it against the database. If authentication is successful, the home screen is displayed.
[0118] Input: Face image, login information
[0119] Data processing: facial recognition, login information authentication
[0120] Output: Authentication result, home screen display
[0121] Step 4:
[0122] Voice command recognition and product selection
[0123] When a user says, "I want to choose some clothes," the device's microphone captures this voice command and converts the voice data into text using the Google Speech-to-Text API. Based on this text data, the device displays a product selection menu. When the user selects a specific category or brand, this information is sent to the server. The server retrieves a list of corresponding products through the online shop's database and sends it to the device. The device then displays the retrieved item list to the user.
[0124] Input: Voice data, text data, product category
[0125] Data processing: voice analysis, item list acquisition
[0126] Output: Product selection menu, item list display
[0127] Step 5:
[0128] Try-on simulation
[0129] When a user selects an item, the device sends the product information and the user's body shape data to a generative AI model, which uses OpenAI technology to generate images of the selected item being tried on. The generated images are then displayed in real time on a smart mirror, giving the user the experience of actually trying the item on.
[0130] Input: Product information, body type data
[0131] Data processing: Generative AI generates try-on images
[0132] Output: Show try-on image
[0133] Step 6:
[0134] Purchase procedure
[0135] When a user says, "I want to buy this outfit," the device's microphone captures this voice command and converts it into text data using the Google Speech-to-Text API. Based on this text data, the device generates a purchase request and sends it to the server. The server processes the payment via the online shop API, and once the order is complete, a confirmation message is sent to the device and displayed to the user.
[0136] Input: Voice data, purchase request
[0137] Data processing: voice analysis, payment processing
[0138] Output: Display purchase confirmation message
[0139] (Application example 1)
[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0141] Currently, when purchasing clothes in a physical store, customers must actually try on the clothes in a fitting room. However, this method takes time, reducing user convenience. Furthermore, due to the impact of COVID-19, many consumers are concerned about using fitting rooms. Therefore, there is a need for a method that allows customers to easily and quickly simulate trying on clothes. Furthermore, there is a need for a system that can provide personalized suggestions based on the user's past purchase history, providing a more satisfying shopping experience.
[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0143] In this invention, the server includes a photographing device that captures the user's body shape data, a voice input device that recognizes the user's voice instructions, a display device that displays a clothing selection menu based on the user's voice instructions, a communication device that transmits information about the user's selected clothing to a generative AI model, a generative AI model device that generates a try-on image of the user wearing the selected clothing, a display device that displays the generated try-on image on a reflective surface, a communication device that completes the purchase process based on the user's voice instructions, a recognition device installed in a physical store that recognizes the user's face, and a database device that makes personalized suggestions based on past purchase history. This allows users to quickly check out and purchase clothing through a try-on simulation without actually trying on the clothing in a physical store. Furthermore, suggestions based on past purchase history also increase user satisfaction.
[0144] "Photographing means" refers to a device for acquiring the user's body shape data, and includes an image acquisition device such as a camera.
[0145] "Voice input means" refers to a device for recognizing a user's voice instructions, and includes a voice recording device such as a microphone.
[0146] "Display means" refers to a device that visually presents information to a user, including a display or screen.
[0147] "Communication means" refers to a network connection device for sending and receiving data, including an internet connection and Wi-Fi.
[0148] "Generative AI model means" refers to AI technology that generates try-on images of a user wearing the clothing selected by the user, and includes a generative AI model.
[0149] "Reflective surface" refers to a mirror or smart mirror that allows users to see themselves.
[0150] "Recognition means" refers to devices installed in physical stores that recognize users' faces, including facial recognition cameras and facial recognition software.
[0151] "Database means" refers to a system that manages data on users' past purchase history and personalized suggestions.
[0152] "Online Shop API" refers to the application programming interface for obtaining clothing information.
[0153] This invention provides a system that allows users to try on clothes without actually trying them on by using a smart full-length mirror in a physical store. The system is composed of the following hardware and software:
[0154] 1. Camera as a photography tool:
[0155] The camera is used to capture the user's body shape data. When the user stands in front of the smart mirror, the camera automatically captures the user's image and generates body shape data.
[0156] 2. Microphone as a means of voice input:
[0157] The microphone is used to recognize the user's voice instructions: when the user selects clothes or makes purchases by voice, the microphone captures the voice and inputs it into the system.
[0158] 3. Display as a means of presentation:
[0159] The display allows users to select clothing and check images of the clothing they are trying on based on voice instructions. Images of the selected clothing and images of the clothing being tried on based on the generative AI model are then displayed.
[0160] 4. Network connectivity as a means of communication:
[0161] The communication means is used to send information about the clothes selected by the user to the generative AI model and receive the generated try-on images, which is achieved via an internet connection or Wi-Fi.
[0162] 5. Generative AI model means:
[0163] The generative AI model generates try-on images based on the user's body data and the selected clothing. The generated try-on images are displayed on the screen in real time, allowing the user to see how the clothes will look when tried on.
[0164] 6. Smart mirrors as reflective surfaces:
[0165] The smart mirror is used to visually present the generated fitting images to the user. It functions as a normal mirror but has a built-in display that displays the fitting images.
[0166] 7. Facial recognition cameras as a means of recognition:
[0167] Facial recognition cameras are used to recognize users' faces in physical stores. When a user stands in front of a smart mirror, the camera recognizes their face and provides user information to the system.
[0168] 8. Database Means:
[0169] The database is used to manage users' past purchase history and provide personalized recommendations, and has the function of suggesting the most suitable clothing based on the user's purchase history.
[0170] As a concrete example, a user stands in front of a smart full-length mirror, and the system recognizes the user's face using a facial recognition camera. When the user gives a voice command such as "I want to see a red dress," the microphone recognizes the voice and a selection menu of red dresses appears on the display. When the user selects a specific dress, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. The generated try-on image is displayed on the smart mirror, allowing the user to check how the dress will look when tried on. Finally, when the user gives a voice command such as "I want to buy this dress," the purchase process is carried out via a communication means.
[0171] The following text can be used as an example of a prompt to input to the generative AI model:
[0172] Prompt: "Given a user's body data and an image of a red dress, generate an image of the user trying on the red dress. The user is a size medium."
[0173] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0174] Step 1:
[0175] The device will perform the initial settings
[0176] The device will perform the initial setup of the camera, microphone, and speaker, and check that they are working properly. Specifically, the device will acquire the user's body shape data from the camera, and prepare the microphone as a voice input method to receive voice instructions. After the setup is complete, the device will check that these devices are working properly.
[0177] Step 2:
[0178] Recognize the user's face and log them in
[0179] When a user stands in front of the smart mirror, the device's camera captures the user's face and uses facial recognition technology to identify the user. The facial image data is taken as input, and the user ID is identified as output. Based on the identified user ID, the server retrieves the user's account information and sends it to the device, which then displays an individually personalized interface.
[0180] Step 3:
[0181] Displaying a clothing selection menu based on the user's voice commands
[0182] When a user voices the instruction "I want to choose some clothes," the device's microphone recognizes this instruction and captures the voice data as input. Voice recognition technology is used to analyze the user's intention. Based on the analysis results, the server retrieves the appropriate product data via the online shop API and sends it to the device. A selection menu is then displayed on the screen as output.
[0183] Step 4:
[0184] Send user-selected clothing information to a generative AI model
[0185] When a user selects a specific item, the device sends detailed information about the selected clothing to the server. The server then sends this information to the generative AI model, which generates a prompt such as, "Please generate a try-on image using the user's body data and an image of the selected clothing." The selected clothing information is taken as input, and the prompt is sent as output to the generative AI model.
[0186] Step 5:
[0187] Try-on images are generated and displayed using a generative AI model
[0188] The generative AI model generates try-on images based on the user's body type data and selected clothing information. It receives body type data and clothing information as input and generates try-on images as output. These images are sent to the device via the server, and the device displays the generated try-on images on the smart mirror.
[0189] Step 6:
[0190] Proceed with purchases based on user voice commands
[0191] When a user gives a voice command such as "I want to buy this outfit," the device's microphone recognizes the voice and captures the voice data as input. The device analyzes the command and sends a purchase request to the server. The server confirms the purchase information via the online shop API and processes the payment. A purchase confirmation message is generated as output and sent to the device.
[0192] By following the steps above, users can experience trying on clothes without actually trying them on, and can easily complete the purchase process through the display on the smart mirror.
[0193] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0194] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions, providing users with an experience of trying on clothes through the mirror. Specific embodiments of the present invention are described below.
[0195] 1. Initial Setup
[0196] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly. The camera captures the user's body shape data and facial expressions, the microphone receives the user's voice instructions, and the speaker responds with voice. The emotion engine recognizes emotions from the user's facial expressions.
[0197] 2. User Identification and Login
[0198] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0199] 3. Clothing choices
[0200] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0201] 4. Emotion Recognition and Suggestion
[0202] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0203] 5. Try-on simulation
[0204] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0205] 6. Purchase Procedure
[0206] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0207] Specific examples
[0208] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0209] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. Furthermore, the use of an emotion engine makes it possible to make suggestions more suited to the user and adjust visual effects, further improving the shopping experience.
[0210] The processing flow will be explained below.
[0211] Step 1:
[0212] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0213] Step 2:
[0214] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0215] Step 3:
[0216] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0217] Step 4:
[0218] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0219] Step 5:
[0220] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0221] Step 6:
[0222] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0223] Step 7:
[0224] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0225] Step 8:
[0226] The emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions, which are then sent to the server.
[0227] Step 9:
[0228] The server generates additional clothing suggestions based on the recognized user emotions and sends a suggestion message to the terminal.
[0229] Step 10:
[0230] The device will display a suggestion message on the screen and the user will confirm it.
[0231] Step 11:
[0232] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0233] Step 12:
[0234] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0235] Step 13:
[0236] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0237] Step 14:
[0238] If the emotion engine recognizes joy or satisfaction from the user's facial expression, it adjusts the visual effects of the try-on images (e.g., background color and lighting effects).
[0239] Step 15:
[0240] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0241] Step 16:
[0242] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0243] Step 17:
[0244] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0245] Example 2
[0246] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0247] In conventional online shopping, users select products without trying them on, which can lead to a decrease in purchasing motivation and an increase in returns after purchase. Furthermore, users cannot see how the product will look when actually worn, which raises concerns about a decrease in satisfaction. Furthermore, suggestions do not take emotions into account, and visual effects are not well-adjusted, leaving a need for an improved user experience.
[0248] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0249] In this invention, the server includes a photographing means for capturing the user's body type data and facial expression, a voice input means for recognizing the user's voice instructions, a display means for displaying a clothing selection menu based on the user's voice instructions, an emotion engine for recognizing the user's emotions, a communication means for sending information about the clothing selected by the user to a generative AI model, a generation AI means for generating try-on images of the user wearing the selected clothing based on the user's body type data, a display means for displaying the generated try-on images on a mirror, an adjustment means for adjusting the visual effects of the try-on images based on the user's emotions, and a communication means for completing the purchase process based on the user's voice instructions. This allows the user to check the appearance of products as if they were trying them on, increasing their desire to purchase and enabling a satisfying online shopping experience.
[0250] "Photographing means for acquiring the user's body type data and facial expression" refers to means for capturing the user's body type data and facial expression using photographing equipment such as a camera, and acquiring them as digital data.
[0251] "A voice input means for recognizing a user's voice instructions" refers to a means for collecting a user's voice instructions using a voice input device such as a microphone, and analyzing the content of the instructions using voice recognition technology.
[0252] "Display means for displaying a clothing selection menu based on the user's voice instructions" refers to means for visually displaying a clothing selection menu based on the content of the user's voice instructions using a display or the like.
[0253] The "emotion engine that recognizes user emotions" is a software and hardware component that analyzes emotions from the user's facial expressions, tone of voice, etc., and recognizes the user's current emotions.
[0254] "Communication means for transmitting information about clothing selected by the user to the generative AI model" refers to a means for transmitting data about clothing selected by the user to the generative AI model via network communication.
[0255] "Generative AI means for generating try-on images of a user wearing selected clothing based on the user's body type data" refers to a generative AI technology used to generate try-on images of a user wearing selected clothing using the user's body type data as input.
[0256] The "display means for displaying the generated try-on images on a mirror" refers to a means for using a mirror display or the like to display the generated try-on images in real time.
[0257] The "adjustment means for adjusting the visual effects of the try-on images based on the user's emotions" is a means for dynamically adjusting the visual effects of the try-on images, such as the background color and lighting, according to the recognized user's emotions.
[0258] "Communication means for carrying out purchase procedures based on the user's voice instructions" refers to a communication means for transmitting information to a server when a user expresses their intention to purchase by voice, and for carrying out the online purchase procedure.
[0259] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions. This system provides users with an experience of trying on clothes through the mirror, enhancing online shopping satisfaction. Specific embodiments of the system are described below.
[0260] The system includes the following hardware and software components:
[0261] Camera: Captures the user's body data and facial expressions.
[0262] Microphone: Inputs the user's voice commands.
[0263] Speaker: Outputs voice responses from the system.
[0264] Emotion engine: Analyzes and recognizes emotions from the user's facial expressions. For example, Haarcascades or deep learning-based emotion recognition models are used.
[0265] Generative AI model: Generates images of the user trying on the selected clothing based on their body shape data. For example, GANs (generative artificial network) or deep learning-based image generation models are used.
[0266] Mirror display: Displays the generated fitting image in real time.
[0267] Communication method: Sending and receiving data via the Internet.
[0268] 1. Initial Setup
[0269] The server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. The device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly.
[0270] 2. User Identification and Login
[0271] When a user stands in front of the smart mirror, the device's camera recognizes the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access a personalized interface.
[0272] 3. Clothing choices
[0273] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0274] 4. Emotion Recognition and Suggestion
[0275] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0276] 5. Try-on simulation
[0277] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0278] 6. Purchase Procedure
[0279] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0280] Specific examples
[0281] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0282] In this way, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. In addition, the use of an emotion engine makes it possible to make more user-friendly suggestions and adjust visual effects, further improving the shopping experience.
[0283] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0284] Step 1:
[0285] The server sets up a database to manage user account information, past purchase history, and product data. The input is user information and product metadata. It then uses a database management system (e.g., MySQL or PostgreSQL) to store this data and organizes the user's individual data. The output is an organized database that allows for personalized offers.
[0286] Step 2:
[0287] The device initializes the camera, microphone, speaker, and emotion engine. The input is the setting parameters of each device. Specifically, it calibrates the camera, adjusts the microphone sensitivity, sets the speaker volume, and initializes the emotion engine. The output is a properly functioning hardware device.
[0288] Step 3:
[0289] When a user stands in front of the device, the device uses the camera to recognize the user's face and displays a login screen. The input is a face image captured by the camera. A facial recognition algorithm (e.g., OpenCV or Dlib) is used to identify the user. The output is a login screen where the user enters their login information.
[0290] Step 4:
[0291] When a user enters their login information, the server authenticates it and displays the user's home screen. The input is a user ID and password. The authentication information is retrieved from a database and verified. The output is the authentication result and a personalized home screen.
[0292] Step 5:
[0293] When a user issues a voice command such as "I want to choose clothes," the device captures this with a microphone and converts it into text using a speech recognition engine. The input is the user's voice command. The speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes it and generates a text command. The output is the generated text command, and a selection menu is displayed based on it.
[0294] Step 6:
[0295] When a user selects a specific category or brand, the server uses the online shop API to retrieve the corresponding item list and sends it to the terminal. The input is the category or brand selected by the user. A request is sent to the online shop API to retrieve product data. The output is the retrieved item list.
[0296] Step 7:
[0297] When a user selects a specific item from the item list, the device sends information about that item to the generative AI model. The input is detailed information about the selected item. The generative AI model processes the data to generate a try-on image. The output is try-on image data.
[0298] Step 8:
[0299] A generative AI model generates try-on images based on the user's body data, and the device displays these images in real time on the mirror display. The input is the user's body data and selected item information. The generative AI model (e.g., GANs) executes the image generation process. The output is an image that looks as if the user is actually trying on the clothing.
[0300] Step 9:
[0301] While trying on clothes, the emotion engine analyzes the user's facial expressions through the camera and recognizes the current emotion. The input is the user's facial expression data captured by the camera. The emotion recognition algorithm analyzes the data. The output is the recognized emotion information.
[0302] Step 10:
[0303] Based on the recognized emotion, the server suggests additional outfits that suit the user and adjusts the visual effects. The input is the recognized emotion information and past purchase history. The suggestion algorithm generates an appropriate product list. The output is additional suggested outfits and adjusted visual effects.
[0304] Step 11:
[0305] When a user wishes to make a purchase, they issue a voice command such as "I would like to purchase this clothing." The device captures this and sends a purchase request to the server. The input is a text command generated by the voice command. The voice recognition engine analyzes the text and generates a purchase request. The output is the purchase request data.
[0306] Step 12:
[0307] The server verifies the purchase information and processes the payment. The input is the purchase request and the user's payment information. The payment process is performed via the online shop API. The output is a purchase confirmation message, which is sent to the terminal and displayed to the user.
[0308] (Application example 2)
[0309] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0310] In today's online shopping environment, the inability to actually try on clothes is a major source of anxiety for users and a hurdle to purchasing. Furthermore, the lack of personalized suggestions that take emotions into account when selecting products also limits the shopping experience. The objective of this invention is to solve these problems and provide users with a more satisfying shopping experience.
[0311] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0312] In this invention, the system includes optical means for acquiring a user's body shape data, voice input means for recognizing the user's voice instructions, display means for displaying a clothing selection menu based on the user's voice instructions, camera means for acquiring the user's facial expression data, emotion analysis means for recognizing the user's emotions, communication means for transmitting information about the user's selected clothing to a generative AI model, generative AI means for generating try-on images of the user wearing the selected clothing, optical display means for displaying the generated try-on images, suggestion means for suggesting additional clothing based on the user's emotion recognition, and communication means for completing the purchase process based on the user's voice instructions. This allows the user to have an online experience similar to trying on actual clothing. Furthermore, personalized suggestions based on emotion analysis are possible, increasing user satisfaction.
[0313] "Optical means" refers to a device for acquiring the user's body shape data, and includes a camera and a sensor.
[0314] "Voice input means" refers to a device that uses a microphone or voice recognition technology to recognize a user's voice instructions and input them into the system.
[0315] "Display means" refers to a display or screen for providing visual information to the user, and displays a clothing selection menu based on the user's voice instructions.
[0316] The "camera means" is a device used to acquire facial expression data of the user, such as a high-resolution camera.
[0317] "Emotion analysis means" refers to software or algorithms that recognize emotions based on a user's facial expression data.
[0318] "Communication means" refers to the network technology and interface used to send information about the clothing selected by the user to the generative AI model and to complete the purchase process.
[0319] The "generative AI means" is a system that includes an artificial intelligence model or algorithm for generating try-on images of the user wearing the selected clothing.
[0320] "Optical display means" refers to a display or screen for displaying the generated try-on image, including the display of smart glasses.
[0321] "Suggestion mechanism" refers to an algorithm or system for providing additional clothing suggestions based on the user's emotion recognition.
[0322] This invention aims to enhance the online and brick-and-mortar shopping experience by combining a system with an emotion analysis engine that recognizes user emotions. The system acquires the user's body shape data and facial expression data, and then simulates trying on clothes and carrying out the purchase process based on that data.
[0323] First, the user puts on the smart glasses and stands in front of the system. The system uses optical means (high-resolution camera) to acquire the user's body shape data and camera means to acquire the user's facial expression data. This allows the system to collect basic data about the user and proceed to the next process.
[0324] Next, the user gives a voice command using the voice input means (microphone). The voice input means recognizes this and displays a clothing selection menu on the display means (the display of the smart glasses) based on the voice command. When the user selects a specific clothing item, the information is sent to the generative AI model via the communication means.
[0325] The generation AI means generates a try-on image of the clothing when worn by the user based on the received body shape data and clothing information. This generated try-on image is displayed in real time on the optical display means (the display of the smart glasses). At this time, the emotion analysis means recognizes emotions based on the user's facial expressions and detects whether the user looks happy or unhappy. Based on the results of the emotion analysis, the system uses the suggestion means to suggest additional clothing items.
[0326] For example, if a user voices the command "I want to look at denim jackets," the voice input means recognizes this and a category menu is displayed. When the user selects a specific denim jacket, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. This try-on image is displayed in real time on the smart glasses display, allowing the user to see the item as if they were actually trying it on. If the user is smiling, the emotion analysis means detects this and the suggestion means suggests additional denim jackets of similar style and color.
[0327] An example of a specific prompt would be, "Generate a fitting image of a denim jacket based on the user's body type data. The user's height is 170 cm, weight is 65 kg, shoulder width is 45 cm, and preferred style is casual." This prompt is sent to the generation AI model, which then generates an appropriate fitting image.
[0328] The system utilizes the Microsoft Azure Emotion API for emotion analysis and the Google Speech-to-Text API for voice command recognition, and uses advanced artificial intelligence technologies such as GPT-4 for generative AI models to respond to users in real time, enabling users to enjoy a smooth and satisfying shopping experience in-store and online.
[0329] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0330] Step 1:
[0331] The user puts on the smart glasses and stands in front of the system. The terminal (smart glasses) acquires the user's body shape data using optical means (high-resolution camera). Specifically, the camera captures the user's full-body image, which is then analyzed using image processing software to acquire data such as height, weight, and shoulder width. The input is the "user's full-body image," and the output is the "user's body shape data."
[0332] Step 2:
[0333] The user issues voice instructions using the voice input means (microphone). The device (smart glasses) captures the user's voice instructions through the voice input means. The voice data is converted into text by voice recognition software (Google Speech-to-Text API). The input is "voice data" and the output is "text data." Based on the voice instructions, the device displays a clothing selection menu on the display means (the display of the smart glasses).
[0334] Step 3:
[0335] The user selects a specific piece of clothing from the displayed menu. The device (smart glasses) recognizes the user's selection using a touch interface or eye tracking. Information about the selected clothing is sent to the server via communication means. The input is "user selection information" and the output is "transmission of clothing information."
[0336] Step 4:
[0337] The server sends the received clothing information and the user's body data to the generative AI model. The generative AI model generates a try-on image of the selected clothing based on the user's body data. Specifically, it uses a prompt to give instructions to the AI model. An example of a prompt is: "Based on the user's body data, please generate an image of a denim jacket being tried on. The user is 170 cm tall, weighs 65 kg, has a shoulder width of 45 cm, and prefers a casual style." The input is "body data and prompt," and the output is "a try-on image."
[0338] Step 5:
[0339] The try-on images generated by the generation AI means are sent to the terminal (smart glasses) via the communication means. The terminal displays the try-on images in real time using the optical display means (the display of the smart glasses). The input is the "try-on image" and the output is the "display of the try-on image."
[0340] Step 6:
[0341] The device (smart glasses) acquires the user's facial expression data using a camera. It analyzes the user's emotions using an emotion analysis tool (Microsoft Azure Emotion API). In this process, the user's facial expression data is input and emotion data is output as the analysis result. The input is "facial expression data" and the output is "emotion data."
[0342] Step 7:
[0343] The server uses the suggestion means to suggest additional clothing based on the emotion data. Specifically, if the user shows a happy emotion, it automatically selects clothing of a similar style and color and suggests it to the user. The input is "emotion data" and the output is "additional clothing suggestions." The terminal displays the suggested additional clothing on its display.
[0344] Step 8:
[0345] The user issues a voice command such as "I would like to purchase this clothing." The device captures this voice command using a voice input means and converts it into text using voice recognition software. This text data is sent to the server via a communication means. The input is "voice data" and the output is "text data of the purchase command."
[0346] Step 9:
[0347] The server receives the purchase instruction and carries out the purchase procedure via the online shop's API. Specifically, it sends the purchase information to the API and processes the payment. The input is "text data of the purchase instruction" and the output is "purchase procedure completed." Once the procedure is complete, a confirmation message is sent to the terminal to notify the user.
[0348] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0349] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0350] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0351] [Second embodiment]
[0352] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0353] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0354] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0355] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0356] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0357] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0358] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0359] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0360] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0361] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0362] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0363] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0364] The present invention provides an online shopping support system using a smart full-length mirror, which allows users to experience the experience of trying on clothes in real life. Specific embodiments of the present invention will be described below.
[0365] 1. Initial Setup
[0366] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, and speaker and verifies that they are working properly. The camera captures the user's body shape data, the microphone receives the user's voice instructions, and the speaker provides voice responses.
[0367] 2. User Identification and Login
[0368] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0369] 3. Clothing choices
[0370] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0371] 4. Try-on simulation
[0372] When a user selects an item, the device sends the clothing information to the AI generator, which then generates an image of the selected clothing based on the user's body type data. This image is displayed in real time on the device's mirror, making it appear as if the user is actually trying it on.
[0373] 5. Purchase Procedure
[0374] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0375] Specific examples
[0376] For example, let's say a user wants to try on a denim jacket. When the user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about that denim jacket is sent to the generation AI, which generates a fitting image adapted to the user's body type. This fitting image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user wants to purchase it, the device recognizes the command and starts the purchase process by saying, "I want to buy this jacket."
[0377] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, thereby eliminating concerns about online shopping and increasing user satisfaction.
[0378] The processing flow will be explained below.
[0379] Step 1:
[0380] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0381] Step 2:
[0382] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0383] Step 3:
[0384] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0385] Step 4:
[0386] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0387] Step 5:
[0388] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0389] Step 6:
[0390] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0391] Step 7:
[0392] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0393] Step 8:
[0394] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0395] Step 9:
[0396] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0397] Step 10:
[0398] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0399] Step 11:
[0400] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0401] Step 12:
[0402] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0403] Step 13:
[0404] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0405] Example 1
[0406] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0407] With traditional online shopping, customers could not actually try on the products, which meant they could not check the size or fit, which often led to anxiety when making a purchase. Furthermore, it was not possible to provide optimal suggestions for individual users, so diverse needs could not be met. Furthermore, the purchasing process was complicated, making it difficult to improve the user experience.
[0408] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0409] In this invention, the server includes an image acquisition means for acquiring a user's body type data, a voice recognition means for recognizing the user's voice instructions, a display means for displaying a product selection menu based on the user's voice instructions, a communication means for sending information about the product selected by the user to a generative AI model, a generative AI means for generating try-on images of the user wearing the selected product, a display means for displaying the generated try-on images on a display device, and a communication means for carrying out the purchase procedure based on the user's voice instructions. This allows the user to have an experience that feels like they are actually trying on the product, eliminating the anxiety of online shopping and enabling individually optimized suggestions, making the purchase procedure smoother.
[0410] "Image acquisition means" refers to devices or technologies for acquiring a user's body shape data, and specifically includes cameras and 3D scanners.
[0411] "Voice recognition means" refers to devices or technology for recognizing a user's voice instructions, and specifically includes a microphone and voice recognition software.
[0412] "Display means" refers to devices or technologies for displaying a user interface, and specifically includes displays and touch screens.
[0413] "Communication means" refers to devices and technologies for transmitting and receiving data, and specifically includes internet connections and wireless communication technologies.
[0414] A "generative AI model" is an artificial intelligence model that generates try-on images based on the user's body type data and selected product information.
[0415] "Generative AI means" refers to technology or equipment for generating try-on images from specified data using a generative AI model.
[0416] The "display device" refers to a device for visually presenting the generated try-on images to the user, and specifically includes a smart mirror or display.
[0417] "Authentication means" refers to devices or technologies for authenticating users, and specifically includes facial recognition systems and login systems.
[0418] A "database" is a system for organizing and storing information, and specifically includes a system for storing information about products in an online shop.
[0419] The present invention is an online shopping support system that provides users with an experience similar to trying on clothes in a physical store. This system acquires the user's body type data, selects products based on voice instructions, generates try-on images using a generative AI model, and performs the entire process from purchase to purchase. Specific embodiments are described below.
[0420] The server prepares a database to manage user account information, past purchase history, and product data, enabling personalized recommendations to be made to users. The database used can be a commercial or open-source relational database management system such as MySQL or PostgreSQL.
[0421] The device is a smart mirror and is configured as follows: the camera is used to acquire the user's body shape data, and can be, for example, a high-performance webcam or a 3D scanner; the microphone is used to receive the user's voice instructions, and is preferably, for example, a directional microphone or a microphone with noise-canceling capabilities; and the speaker is used to provide voice feedback.
[0422] When a user stands in front of the smart mirror, the device recognizes the user's face through the camera. The facial recognition technology uses libraries and frameworks such as OpenCV and TensorFlow. Once authentication is complete, the device displays a login screen, allowing the user to enter login information via voice commands or the touch panel.
[0423] When a user says, "I want to choose clothes," the device's microphone captures the voice command and analyzes it using the Google Speech-to-Text API. Based on the analysis results, the device displays a selection menu, allowing the user to select a specific category or brand. Once the user confirms their selection, the server retrieves product information from the online store's database and sends it to the device.
[0424] Next, the user selects the product they want to try on, and the device sends that product information to a generative AI model. The generative AI model uses OpenAI's generative AI technology, for example. This model generates a fitting image based on the user's body data. The generated fitting image is displayed in real time on the smart mirror, allowing the user to experience the experience of actually trying on the product.
[0425] For example, if a user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about the denim jacket is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is instantly displayed on the smart mirror, allowing the user to check the denim jacket as if they were trying it on. When the user says, "I want to buy this jacket," the device recognizes the voice command and begins the purchase process. The server uses the online shop API to process the payment and sends a confirmation message to the device.
[0426] As described above, the online shopping support system of the present invention allows users to enjoy a realistic try-on experience from the comfort of their own home, thereby eliminating concerns about online shopping and contributing to improved user satisfaction.
[0427] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0428] Step 1:
[0429] Database Setup
[0430] The server prepares a database and manages user account information, past purchase history, and product data. Database management systems such as MySQL and PostgreSQL are used for this. Specifically, the server creates tables for registering new users and for storing detailed product information.
[0431] Input: User information, purchase history, product data
[0432] Data processing: storing data in a database
[0433] Output: Database completed, data stored
[0434] Step 2:
[0435] Initial Hardware Setup
[0436] The device's camera, microphone, and speaker are initially configured. At this time, the camera resolution and frame rate are set, the microphone sensitivity is adjusted, and the speaker volume is set. For the camera, a high-performance webcam or 3D scanner is used, and for the microphone, a directional microphone with noise-canceling functions is used. For the speaker, a speaker that can output clear audio is selected.
[0437] Input: Hardware configuration information
[0438] Data processing: Applying setting parameters
[0439] Output: Camera, microphone, and speaker initial settings completed
[0440] Step 3:
[0441] User facial recognition and login
[0442] When a user stands in front of the smart mirror, the device's camera recognizes the user's face. Facial recognition technology uses OpenCV and TensorFlow. If facial recognition is successful, the device displays a login screen and the user enters their login information. The server receives this and authenticates it against the database. If authentication is successful, the home screen is displayed.
[0443] Input: Face image, login information
[0444] Data processing: facial recognition, login information authentication
[0445] Output: Authentication result, home screen display
[0446] Step 4:
[0447] Voice command recognition and product selection
[0448] When a user says, "I want to choose some clothes," the device's microphone captures this voice command and converts the voice data into text using the Google Speech-to-Text API. Based on this text data, the device displays a product selection menu. When the user selects a specific category or brand, this information is sent to the server. The server retrieves a list of corresponding products through the online shop's database and sends it to the device. The device then displays the retrieved item list to the user.
[0449] Input: Voice data, text data, product category
[0450] Data processing: voice analysis, item list acquisition
[0451] Output: Product selection menu, item list display
[0452] Step 5:
[0453] Try-on simulation
[0454] When a user selects an item, the device sends the product information and the user's body shape data to a generative AI model, which uses OpenAI technology to generate images of the selected item being tried on. The generated images are then displayed in real time on a smart mirror, giving the user the experience of actually trying the item on.
[0455] Input: Product information, body type data
[0456] Data processing: Generative AI generates try-on images
[0457] Output: Show try-on image
[0458] Step 6:
[0459] Purchase procedure
[0460] When a user says, "I want to buy this outfit," the device's microphone captures this voice command and converts it into text data using the Google Speech-to-Text API. Based on this text data, the device generates a purchase request and sends it to the server. The server processes the payment via the online shop API, and once the order is complete, a confirmation message is sent to the device and displayed to the user.
[0461] Input: Voice data, purchase request
[0462] Data processing: voice analysis, payment processing
[0463] Output: Display purchase confirmation message
[0464] (Application example 1)
[0465] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0466] Currently, when purchasing clothes in a physical store, customers must actually try on the clothes in a fitting room. However, this method takes time, reducing user convenience. Furthermore, due to the impact of COVID-19, many consumers are concerned about using fitting rooms. Therefore, there is a need for a method that allows customers to easily and quickly simulate trying on clothes. Furthermore, there is a need for a system that can provide personalized suggestions based on the user's past purchase history, providing a more satisfying shopping experience.
[0467] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0468] In this invention, the server includes a photographing device that captures the user's body shape data, a voice input device that recognizes the user's voice instructions, a display device that displays a clothing selection menu based on the user's voice instructions, a communication device that transmits information about the user's selected clothing to a generative AI model, a generative AI model device that generates a try-on image of the user wearing the selected clothing, a display device that displays the generated try-on image on a reflective surface, a communication device that completes the purchase process based on the user's voice instructions, a recognition device installed in a physical store that recognizes the user's face, and a database device that makes personalized suggestions based on past purchase history. This allows users to quickly check out and purchase clothing through a try-on simulation without actually trying on the clothing in a physical store. Furthermore, suggestions based on past purchase history also increase user satisfaction.
[0469] "Photographing means" refers to a device for acquiring the user's body shape data, and includes an image acquisition device such as a camera.
[0470] "Voice input means" refers to a device for recognizing a user's voice instructions, and includes a voice recording device such as a microphone.
[0471] "Display means" refers to a device that visually presents information to a user, including a display or screen.
[0472] "Communication means" refers to a network connection device for sending and receiving data, including an internet connection and Wi-Fi.
[0473] "Generative AI model means" refers to AI technology that generates try-on images of a user wearing the clothing selected by the user, and includes a generative AI model.
[0474] "Reflective surface" refers to a mirror or smart mirror that allows users to see themselves.
[0475] "Recognition means" refers to devices installed in physical stores that recognize users' faces, including facial recognition cameras and facial recognition software.
[0476] "Database means" refers to a system that manages data on users' past purchase history and personalized suggestions.
[0477] "Online Shop API" refers to the application programming interface for obtaining clothing information.
[0478] This invention provides a system that allows users to try on clothes without actually trying them on by using a smart full-length mirror in a physical store. The system is composed of the following hardware and software:
[0479] 1. Camera as a photography tool:
[0480] The camera is used to capture the user's body shape data. When the user stands in front of the smart mirror, the camera automatically captures the user's image and generates body shape data.
[0481] 2. Microphone as a means of voice input:
[0482] The microphone is used to recognize the user's voice instructions: when the user selects clothes or makes purchases by voice, the microphone captures the voice and inputs it into the system.
[0483] 3. Display as a means of presentation:
[0484] The display allows users to select clothing and check images of the clothing they are trying on based on voice instructions. Images of the selected clothing and images of the clothing being tried on based on the generative AI model are then displayed.
[0485] 4. Network connectivity as a means of communication:
[0486] The communication means is used to send information about the clothes selected by the user to the generative AI model and receive the generated try-on images, which is achieved via an internet connection or Wi-Fi.
[0487] 5. Generative AI model means:
[0488] The generative AI model generates try-on images based on the user's body data and the selected clothing. The generated try-on images are displayed on the screen in real time, allowing the user to see how the clothes will look when tried on.
[0489] 6. Smart mirrors as reflective surfaces:
[0490] The smart mirror is used to visually present the generated fitting images to the user. It functions as a normal mirror but has a built-in display that displays the fitting images.
[0491] 7. Facial recognition cameras as a means of recognition:
[0492] Facial recognition cameras are used to recognize users' faces in physical stores. When a user stands in front of a smart mirror, the camera recognizes their face and provides user information to the system.
[0493] 8. Database Means:
[0494] The database is used to manage users' past purchase history and provide personalized recommendations, and has the function of suggesting the most suitable clothing based on the user's purchase history.
[0495] As a concrete example, a user stands in front of a smart full-length mirror, and the system recognizes the user's face using a facial recognition camera. When the user gives a voice command such as "I want to see a red dress," the microphone recognizes the voice and a selection menu of red dresses appears on the display. When the user selects a specific dress, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. The generated try-on image is displayed on the smart mirror, allowing the user to check how the dress will look when tried on. Finally, when the user gives a voice command such as "I want to buy this dress," the purchase process is carried out via a communication means.
[0496] The following text can be used as an example of a prompt to input to the generative AI model:
[0497] Prompt: "Given a user's body data and an image of a red dress, generate an image of the user trying on the red dress. The user is a size medium."
[0498] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0499] Step 1:
[0500] The device will perform the initial settings
[0501] The device will perform the initial setup of the camera, microphone, and speaker, and check that they are working properly. Specifically, the device will acquire the user's body shape data from the camera, and prepare the microphone as a voice input method to receive voice instructions. After the setup is complete, the device will check that these devices are working properly.
[0502] Step 2:
[0503] Recognize the user's face and log them in
[0504] When a user stands in front of the smart mirror, the device's camera captures the user's face and uses facial recognition technology to identify the user. The facial image data is taken as input, and the user ID is identified as output. Based on the identified user ID, the server retrieves the user's account information and sends it to the device, which then displays an individually personalized interface.
[0505] Step 3:
[0506] Displaying a clothing selection menu based on the user's voice commands
[0507] When a user voices the instruction "I want to choose some clothes," the device's microphone recognizes this instruction and captures the voice data as input. Voice recognition technology is used to analyze the user's intention. Based on the analysis results, the server retrieves the appropriate product data via the online shop API and sends it to the device. A selection menu is then displayed on the screen as output.
[0508] Step 4:
[0509] Send user-selected clothing information to a generative AI model
[0510] When a user selects a specific item, the device sends detailed information about the selected clothing to the server. The server then sends this information to the generative AI model, which generates a prompt such as, "Please generate a try-on image using the user's body data and an image of the selected clothing." The selected clothing information is taken as input, and the prompt is sent as output to the generative AI model.
[0511] Step 5:
[0512] Try-on images are generated and displayed using a generative AI model
[0513] The generative AI model generates try-on images based on the user's body type data and selected clothing information. It receives body type data and clothing information as input and generates try-on images as output. These images are sent to the device via the server, and the device displays the generated try-on images on the smart mirror.
[0514] Step 6:
[0515] Proceed with purchases based on user voice commands
[0516] When a user gives a voice command such as "I want to buy this outfit," the device's microphone recognizes the voice and captures the voice data as input. The device analyzes the command and sends a purchase request to the server. The server confirms the purchase information via the online shop API and processes the payment. A purchase confirmation message is generated as output and sent to the device.
[0517] By following the steps above, users can experience trying on clothes without actually trying them on, and can easily complete the purchase process through the display on the smart mirror.
[0518] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0519] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions, providing users with an experience of trying on clothes through the mirror. Specific embodiments of the present invention are described below.
[0520] 1. Initial Setup
[0521] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly. The camera captures the user's body shape data and facial expressions, the microphone receives the user's voice instructions, and the speaker responds with voice. The emotion engine recognizes emotions from the user's facial expressions.
[0522] 2. User Identification and Login
[0523] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0524] 3. Clothing choices
[0525] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0526] 4. Emotion Recognition and Suggestion
[0527] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0528] 5. Try-on simulation
[0529] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0530] 6. Purchase Procedure
[0531] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0532] Specific examples
[0533] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0534] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. Furthermore, the use of an emotion engine makes it possible to make suggestions more suited to the user and adjust visual effects, further improving the shopping experience.
[0535] The processing flow will be explained below.
[0536] Step 1:
[0537] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0538] Step 2:
[0539] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0540] Step 3:
[0541] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0542] Step 4:
[0543] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0544] Step 5:
[0545] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0546] Step 6:
[0547] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0548] Step 7:
[0549] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0550] Step 8:
[0551] The emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions, which are then sent to the server.
[0552] Step 9:
[0553] The server generates additional clothing suggestions based on the recognized user emotions and sends a suggestion message to the terminal.
[0554] Step 10:
[0555] The device will display a suggestion message on the screen and the user will confirm it.
[0556] Step 11:
[0557] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0558] Step 12:
[0559] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0560] Step 13:
[0561] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0562] Step 14:
[0563] If the emotion engine recognizes joy or satisfaction from the user's facial expression, it adjusts the visual effects of the try-on images (e.g., background color and lighting effects).
[0564] Step 15:
[0565] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0566] Step 16:
[0567] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0568] Step 17:
[0569] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0570] Example 2
[0571] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0572] In conventional online shopping, users select products without trying them on, which can lead to a decrease in purchasing motivation and an increase in returns after purchase. Furthermore, users cannot see how the product will look when actually worn, which raises concerns about a decrease in satisfaction. Furthermore, suggestions do not take emotions into account, and visual effects are not well-adjusted, leaving a need for an improved user experience.
[0573] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0574] In this invention, the server includes a photographing means for capturing the user's body type data and facial expression, a voice input means for recognizing the user's voice instructions, a display means for displaying a clothing selection menu based on the user's voice instructions, an emotion engine for recognizing the user's emotions, a communication means for sending information about the clothing selected by the user to a generative AI model, a generation AI means for generating try-on images of the user wearing the selected clothing based on the user's body type data, a display means for displaying the generated try-on images on a mirror, an adjustment means for adjusting the visual effects of the try-on images based on the user's emotions, and a communication means for completing the purchase process based on the user's voice instructions. This allows the user to check the appearance of products as if they were trying them on, increasing their desire to purchase and enabling a satisfying online shopping experience.
[0575] "Photographing means for acquiring the user's body type data and facial expression" refers to means for capturing the user's body type data and facial expression using photographing equipment such as a camera, and acquiring them as digital data.
[0576] "A voice input means for recognizing a user's voice instructions" refers to a means for collecting a user's voice instructions using a voice input device such as a microphone, and analyzing the content of the instructions using voice recognition technology.
[0577] "Display means for displaying a clothing selection menu based on the user's voice instructions" refers to means for visually displaying a clothing selection menu based on the content of the user's voice instructions using a display or the like.
[0578] The "emotion engine that recognizes user emotions" is a software and hardware component that analyzes emotions from the user's facial expressions, tone of voice, etc., and recognizes the user's current emotions.
[0579] "Communication means for transmitting information about clothing selected by the user to the generative AI model" refers to a means for transmitting data about clothing selected by the user to the generative AI model via network communication.
[0580] "Generative AI means for generating try-on images of a user wearing selected clothing based on the user's body type data" refers to a generative AI technology used to generate try-on images of a user wearing selected clothing using the user's body type data as input.
[0581] The "display means for displaying the generated try-on images on a mirror" refers to a means for using a mirror display or the like to display the generated try-on images in real time.
[0582] The "adjustment means for adjusting the visual effects of the try-on images based on the user's emotions" is a means for dynamically adjusting the visual effects of the try-on images, such as the background color and lighting, according to the recognized user's emotions.
[0583] "Communication means for carrying out purchase procedures based on the user's voice instructions" refers to a communication means for transmitting information to a server when a user expresses their intention to purchase by voice, and for carrying out the online purchase procedure.
[0584] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions. This system provides users with an experience of trying on clothes through the mirror, enhancing online shopping satisfaction. Specific embodiments of the system are described below.
[0585] The system includes the following hardware and software components:
[0586] Camera: Captures the user's body data and facial expressions.
[0587] Microphone: Inputs the user's voice commands.
[0588] Speaker: Outputs voice responses from the system.
[0589] Emotion engine: Analyzes and recognizes emotions from the user's facial expressions. For example, Haarcascades or deep learning-based emotion recognition models are used.
[0590] Generative AI model: Generates images of the user trying on the selected clothing based on their body shape data. For example, GANs (generative artificial network) or deep learning-based image generation models are used.
[0591] Mirror display: Displays the generated fitting image in real time.
[0592] Communication method: Sending and receiving data via the Internet.
[0593] 1. Initial Setup
[0594] The server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. The device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly.
[0595] 2. User Identification and Login
[0596] When a user stands in front of the smart mirror, the device's camera recognizes the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access a personalized interface.
[0597] 3. Clothing choices
[0598] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0599] 4. Emotion Recognition and Suggestion
[0600] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0601] 5. Try-on simulation
[0602] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0603] 6. Purchase Procedure
[0604] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0605] Specific examples
[0606] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0607] In this way, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. In addition, the use of an emotion engine makes it possible to make more user-friendly suggestions and adjust visual effects, further improving the shopping experience.
[0608] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0609] Step 1:
[0610] The server sets up a database to manage user account information, past purchase history, and product data. The input is user information and product metadata. It then uses a database management system (e.g., MySQL or PostgreSQL) to store this data and organizes the user's individual data. The output is an organized database that allows for personalized offers.
[0611] Step 2:
[0612] The device initializes the camera, microphone, speaker, and emotion engine. The input is the setting parameters of each device. Specifically, it calibrates the camera, adjusts the microphone sensitivity, sets the speaker volume, and initializes the emotion engine. The output is a properly functioning hardware device.
[0613] Step 3:
[0614] When a user stands in front of the device, the device uses the camera to recognize the user's face and displays a login screen. The input is a face image captured by the camera. A facial recognition algorithm (e.g., OpenCV or Dlib) is used to identify the user. The output is a login screen where the user enters their login information.
[0615] Step 4:
[0616] When a user enters their login information, the server authenticates it and displays the user's home screen. The input is a user ID and password. The authentication information is retrieved from a database and verified. The output is the authentication result and a personalized home screen.
[0617] Step 5:
[0618] When a user issues a voice command such as "I want to choose clothes," the device captures this with a microphone and converts it into text using a speech recognition engine. The input is the user's voice command. The speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes it and generates a text command. The output is the generated text command, and a selection menu is displayed based on it.
[0619] Step 6:
[0620] When a user selects a specific category or brand, the server uses the online shop API to retrieve the corresponding item list and sends it to the terminal. The input is the category or brand selected by the user. A request is sent to the online shop API to retrieve product data. The output is the retrieved item list.
[0621] Step 7:
[0622] When a user selects a specific item from the item list, the device sends information about that item to the generative AI model. The input is detailed information about the selected item. The generative AI model processes the data to generate a try-on image. The output is try-on image data.
[0623] Step 8:
[0624] A generative AI model generates try-on images based on the user's body data, and the device displays these images in real time on the mirror display. The input is the user's body data and selected item information. The generative AI model (e.g., GANs) executes the image generation process. The output is an image that looks as if the user is actually trying on the clothing.
[0625] Step 9:
[0626] While trying on clothes, the emotion engine analyzes the user's facial expressions through the camera and recognizes the current emotion. The input is the user's facial expression data captured by the camera. The emotion recognition algorithm analyzes the data. The output is the recognized emotion information.
[0627] Step 10:
[0628] Based on the recognized emotion, the server suggests additional outfits that suit the user and adjusts the visual effects. The input is the recognized emotion information and past purchase history. The suggestion algorithm generates an appropriate product list. The output is additional suggested outfits and adjusted visual effects.
[0629] Step 11:
[0630] When a user wishes to make a purchase, they issue a voice command such as "I would like to purchase this clothing." The device captures this and sends a purchase request to the server. The input is a text command generated by the voice command. The voice recognition engine analyzes the text and generates a purchase request. The output is the purchase request data.
[0631] Step 12:
[0632] The server verifies the purchase information and processes the payment. The input is the purchase request and the user's payment information. The payment process is performed via the online shop API. The output is a purchase confirmation message, which is sent to the terminal and displayed to the user.
[0633] (Application example 2)
[0634] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0635] In today's online shopping environment, the inability to actually try on clothes is a major source of anxiety for users and a hurdle to purchasing. Furthermore, the lack of personalized suggestions that take emotions into account when selecting products also limits the shopping experience. The objective of this invention is to solve these problems and provide users with a more satisfying shopping experience.
[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0637] In this invention, the system includes optical means for acquiring a user's body shape data, voice input means for recognizing the user's voice instructions, display means for displaying a clothing selection menu based on the user's voice instructions, camera means for acquiring the user's facial expression data, emotion analysis means for recognizing the user's emotions, communication means for transmitting information about the user's selected clothing to a generative AI model, generative AI means for generating try-on images of the user wearing the selected clothing, optical display means for displaying the generated try-on images, suggestion means for suggesting additional clothing based on the user's emotion recognition, and communication means for completing the purchase process based on the user's voice instructions. This allows the user to have an online experience similar to trying on actual clothing. Furthermore, personalized suggestions based on emotion analysis are possible, increasing user satisfaction.
[0638] "Optical means" refers to a device for acquiring the user's body shape data, and includes a camera and a sensor.
[0639] "Voice input means" refers to a device that uses a microphone or voice recognition technology to recognize a user's voice instructions and input them into the system.
[0640] "Display means" refers to a display or screen for providing visual information to the user, and displays a clothing selection menu based on the user's voice instructions.
[0641] The "camera means" is a device used to acquire facial expression data of the user, such as a high-resolution camera.
[0642] "Emotion analysis means" refers to software or algorithms that recognize emotions based on a user's facial expression data.
[0643] "Communication means" refers to the network technology and interface used to send information about the clothing selected by the user to the generative AI model and to complete the purchase process.
[0644] The "generative AI means" is a system that includes an artificial intelligence model or algorithm for generating try-on images of the user wearing the selected clothing.
[0645] "Optical display means" refers to a display or screen for displaying the generated try-on image, including the display of smart glasses.
[0646] "Suggestion mechanism" refers to an algorithm or system for providing additional clothing suggestions based on the user's emotion recognition.
[0647] This invention aims to enhance the online and brick-and-mortar shopping experience by combining a system with an emotion analysis engine that recognizes user emotions. The system acquires the user's body shape data and facial expression data, and then simulates trying on clothes and carrying out the purchase process based on that data.
[0648] First, the user puts on the smart glasses and stands in front of the system. The system uses optical means (high-resolution camera) to acquire the user's body shape data and camera means to acquire the user's facial expression data. This allows the system to collect basic data about the user and proceed to the next process.
[0649] Next, the user gives a voice command using the voice input means (microphone). The voice input means recognizes this and displays a clothing selection menu on the display means (the display of the smart glasses) based on the voice command. When the user selects a specific clothing item, the information is sent to the generative AI model via the communication means.
[0650] The generation AI means generates a try-on image of the clothing when worn by the user based on the received body shape data and clothing information. This generated try-on image is displayed in real time on the optical display means (the display of the smart glasses). At this time, the emotion analysis means recognizes emotions based on the user's facial expressions and detects whether the user looks happy or unhappy. Based on the results of the emotion analysis, the system uses the suggestion means to suggest additional clothing items.
[0651] For example, if a user voices the command "I want to look at denim jackets," the voice input means recognizes this and a category menu is displayed. When the user selects a specific denim jacket, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. This try-on image is displayed in real time on the smart glasses display, allowing the user to see the item as if they were actually trying it on. If the user is smiling, the emotion analysis means detects this and the suggestion means suggests additional denim jackets of similar style and color.
[0652] An example of a specific prompt would be, "Generate a fitting image of a denim jacket based on the user's body type data. The user's height is 170 cm, weight is 65 kg, shoulder width is 45 cm, and preferred style is casual." This prompt is sent to the generation AI model, which then generates an appropriate fitting image.
[0653] The system utilizes the Microsoft Azure Emotion API for emotion analysis and the Google Speech-to-Text API for voice command recognition, and uses advanced artificial intelligence technologies such as GPT-4 for generative AI models to respond to users in real time, enabling users to enjoy a smooth and satisfying shopping experience in-store and online.
[0654] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0655] Step 1:
[0656] The user puts on the smart glasses and stands in front of the system. The terminal (smart glasses) acquires the user's body shape data using optical means (high-resolution camera). Specifically, the camera captures the user's full-body image, which is then analyzed using image processing software to acquire data such as height, weight, and shoulder width. The input is the "user's full-body image," and the output is the "user's body shape data."
[0657] Step 2:
[0658] The user issues voice instructions using the voice input means (microphone). The device (smart glasses) captures the user's voice instructions through the voice input means. The voice data is converted into text by voice recognition software (Google Speech-to-Text API). The input is "voice data" and the output is "text data." Based on the voice instructions, the device displays a clothing selection menu on the display means (the display of the smart glasses).
[0659] Step 3:
[0660] The user selects a specific piece of clothing from the displayed menu. The device (smart glasses) recognizes the user's selection using a touch interface or eye tracking. Information about the selected clothing is sent to the server via communication means. The input is "user selection information" and the output is "transmission of clothing information."
[0661] Step 4:
[0662] The server sends the received clothing information and the user's body data to the generative AI model. The generative AI model generates a try-on image of the selected clothing based on the user's body data. Specifically, it uses a prompt to give instructions to the AI model. An example of a prompt is: "Based on the user's body data, please generate an image of a denim jacket being tried on. The user is 170 cm tall, weighs 65 kg, has a shoulder width of 45 cm, and prefers a casual style." The input is "body data and prompt," and the output is "a try-on image."
[0663] Step 5:
[0664] The try-on images generated by the generation AI means are sent to the terminal (smart glasses) via the communication means. The terminal displays the try-on images in real time using the optical display means (the display of the smart glasses). The input is the "try-on image" and the output is the "display of the try-on image."
[0665] Step 6:
[0666] The device (smart glasses) acquires the user's facial expression data using a camera. It analyzes the user's emotions using an emotion analysis tool (Microsoft Azure Emotion API). In this process, the user's facial expression data is input and emotion data is output as the analysis result. The input is "facial expression data" and the output is "emotion data."
[0667] Step 7:
[0668] The server uses the suggestion means to suggest additional clothing based on the emotion data. Specifically, if the user shows a happy emotion, it automatically selects clothing of a similar style and color and suggests it to the user. The input is "emotion data" and the output is "additional clothing suggestions." The terminal displays the suggested additional clothing on its display.
[0669] Step 8:
[0670] The user issues a voice command such as "I would like to purchase this clothing." The device captures this voice command using a voice input means and converts it into text using voice recognition software. This text data is sent to the server via a communication means. The input is "voice data" and the output is "text data of the purchase command."
[0671] Step 9:
[0672] The server receives the purchase instruction and carries out the purchase procedure via the online shop's API. Specifically, it sends the purchase information to the API and processes the payment. The input is "text data of the purchase instruction" and the output is "purchase procedure completed." Once the procedure is complete, a confirmation message is sent to the terminal to notify the user.
[0673] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0674] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0675] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0676] [Third embodiment]
[0677] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0678] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0679] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0680] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0681] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0682] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0683] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0684] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0685] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0686] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0687] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0688] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0689] The present invention provides an online shopping support system using a smart full-length mirror, which allows users to experience the experience of trying on clothes in real life. Specific embodiments of the present invention will be described below.
[0690] 1. Initial Setup
[0691] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, and speaker and verifies that they are working properly. The camera captures the user's body shape data, the microphone receives the user's voice instructions, and the speaker provides voice responses.
[0692] 2. User Identification and Login
[0693] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0694] 3. Clothing choices
[0695] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0696] 4. Try-on simulation
[0697] When a user selects an item, the device sends the clothing information to the AI generator, which then generates an image of the selected clothing based on the user's body type data. This image is displayed in real time on the device's mirror, making it appear as if the user is actually trying it on.
[0698] 5. Purchase Procedure
[0699] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0700] Specific examples
[0701] For example, let's say a user wants to try on a denim jacket. When the user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about that denim jacket is sent to the generation AI, which generates a fitting image adapted to the user's body type. This fitting image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user wants to purchase it, the device recognizes the command and starts the purchase process by saying, "I want to buy this jacket."
[0702] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, thereby eliminating concerns about online shopping and increasing user satisfaction.
[0703] The processing flow will be explained below.
[0704] Step 1:
[0705] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0706] Step 2:
[0707] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0708] Step 3:
[0709] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0710] Step 4:
[0711] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0712] Step 5:
[0713] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0714] Step 6:
[0715] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0716] Step 7:
[0717] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0718] Step 8:
[0719] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0720] Step 9:
[0721] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0722] Step 10:
[0723] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0724] Step 11:
[0725] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0726] Step 12:
[0727] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0728] Step 13:
[0729] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0730] Example 1
[0731] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0732] With traditional online shopping, customers could not actually try on the products, which meant they could not check the size or fit, which often led to anxiety when making a purchase. Furthermore, it was not possible to provide optimal suggestions for individual users, so diverse needs could not be met. Furthermore, the purchasing process was complicated, making it difficult to improve the user experience.
[0733] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0734] In this invention, the server includes an image acquisition means for acquiring a user's body type data, a voice recognition means for recognizing the user's voice instructions, a display means for displaying a product selection menu based on the user's voice instructions, a communication means for sending information about the product selected by the user to a generative AI model, a generative AI means for generating try-on images of the user wearing the selected product, a display means for displaying the generated try-on images on a display device, and a communication means for carrying out the purchase procedure based on the user's voice instructions. This allows the user to have an experience that feels like they are actually trying on the product, eliminating the anxiety of online shopping and enabling individually optimized suggestions, making the purchase procedure smoother.
[0735] "Image acquisition means" refers to devices or technologies for acquiring a user's body shape data, and specifically includes cameras and 3D scanners.
[0736] "Voice recognition means" refers to devices or technology for recognizing a user's voice instructions, and specifically includes a microphone and voice recognition software.
[0737] "Display means" refers to devices or technologies for displaying a user interface, and specifically includes displays and touch screens.
[0738] "Communication means" refers to devices and technologies for transmitting and receiving data, and specifically includes internet connections and wireless communication technologies.
[0739] A "generative AI model" is an artificial intelligence model that generates try-on images based on the user's body type data and selected product information.
[0740] "Generative AI means" refers to technology or equipment for generating try-on images from specified data using a generative AI model.
[0741] The "display device" refers to a device for visually presenting the generated try-on images to the user, and specifically includes a smart mirror or display.
[0742] "Authentication means" refers to devices or technologies for authenticating users, and specifically includes facial recognition systems and login systems.
[0743] A "database" is a system for organizing and storing information, and specifically includes a system for storing information about products in an online shop.
[0744] The present invention is an online shopping support system that provides users with an experience similar to trying on clothes in a physical store. This system acquires the user's body type data, selects products based on voice instructions, generates try-on images using a generative AI model, and performs the entire process from purchase to purchase. Specific embodiments are described below.
[0745] The server prepares a database to manage user account information, past purchase history, and product data, enabling personalized recommendations to be made to users. The database used can be a commercial or open-source relational database management system such as MySQL or PostgreSQL.
[0746] The device is a smart mirror and is configured as follows: the camera is used to acquire the user's body shape data, and can be, for example, a high-performance webcam or a 3D scanner; the microphone is used to receive the user's voice instructions, and is preferably, for example, a directional microphone or a microphone with noise-canceling capabilities; and the speaker is used to provide voice feedback.
[0747] When a user stands in front of the smart mirror, the device recognizes the user's face through the camera. The facial recognition technology uses libraries and frameworks such as OpenCV and TensorFlow. Once authentication is complete, the device displays a login screen, allowing the user to enter login information via voice commands or the touch panel.
[0748] When a user says, "I want to choose clothes," the device's microphone captures the voice command and analyzes it using the Google Speech-to-Text API. Based on the analysis results, the device displays a selection menu, allowing the user to select a specific category or brand. Once the user confirms their selection, the server retrieves product information from the online store's database and sends it to the device.
[0749] Next, the user selects the product they want to try on, and the device sends that product information to a generative AI model. The generative AI model uses OpenAI's generative AI technology, for example. This model generates a fitting image based on the user's body data. The generated fitting image is displayed in real time on the smart mirror, allowing the user to experience the experience of actually trying on the product.
[0750] For example, if a user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about the denim jacket is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is instantly displayed on the smart mirror, allowing the user to check the denim jacket as if they were trying it on. When the user says, "I want to buy this jacket," the device recognizes the voice command and begins the purchase process. The server uses the online shop API to process the payment and sends a confirmation message to the device.
[0751] As described above, the online shopping support system of the present invention allows users to enjoy a realistic try-on experience from the comfort of their own home, thereby eliminating concerns about online shopping and contributing to improved user satisfaction.
[0752] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0753] Step 1:
[0754] Database Setup
[0755] The server prepares a database and manages user account information, past purchase history, and product data. Database management systems such as MySQL and PostgreSQL are used for this. Specifically, the server creates tables for registering new users and for storing detailed product information.
[0756] Input: User information, purchase history, product data
[0757] Data processing: storing data in a database
[0758] Output: Database completed, data stored
[0759] Step 2:
[0760] Initial Hardware Setup
[0761] The device's camera, microphone, and speaker are initially configured. At this time, the camera resolution and frame rate are set, the microphone sensitivity is adjusted, and the speaker volume is set. For the camera, a high-performance webcam or 3D scanner is used, and for the microphone, a directional microphone with noise-canceling functions is used. For the speaker, a speaker that can output clear audio is selected.
[0762] Input: Hardware configuration information
[0763] Data processing: Applying setting parameters
[0764] Output: Camera, microphone, and speaker initial settings completed
[0765] Step 3:
[0766] User facial recognition and login
[0767] When a user stands in front of the smart mirror, the device's camera recognizes the user's face. Facial recognition technology uses OpenCV and TensorFlow. If facial recognition is successful, the device displays a login screen and the user enters their login information. The server receives this and authenticates it against the database. If authentication is successful, the home screen is displayed.
[0768] Input: Face image, login information
[0769] Data processing: facial recognition, login information authentication
[0770] Output: Authentication result, home screen display
[0771] Step 4:
[0772] Voice command recognition and product selection
[0773] When a user says, "I want to choose some clothes," the device's microphone captures this voice command and converts the voice data into text using the Google Speech-to-Text API. Based on this text data, the device displays a product selection menu. When the user selects a specific category or brand, this information is sent to the server. The server retrieves a list of corresponding products through the online shop's database and sends it to the device. The device then displays the retrieved item list to the user.
[0774] Input: Voice data, text data, product category
[0775] Data processing: voice analysis, item list acquisition
[0776] Output: Product selection menu, item list display
[0777] Step 5:
[0778] Try-on simulation
[0779] When a user selects an item, the device sends the product information and the user's body shape data to a generative AI model, which uses OpenAI technology to generate images of the selected item being tried on. The generated images are then displayed in real time on a smart mirror, giving the user the experience of actually trying the item on.
[0780] Input: Product information, body type data
[0781] Data processing: Generative AI generates try-on images
[0782] Output: Show try-on image
[0783] Step 6:
[0784] Purchase procedure
[0785] When a user says, "I want to buy this outfit," the device's microphone captures this voice command and converts it into text data using the Google Speech-to-Text API. Based on this text data, the device generates a purchase request and sends it to the server. The server processes the payment via the online shop API, and once the order is complete, a confirmation message is sent to the device and displayed to the user.
[0786] Input: Voice data, purchase request
[0787] Data processing: voice analysis, payment processing
[0788] Output: Display purchase confirmation message
[0789] (Application example 1)
[0790] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0791] Currently, when purchasing clothes in a physical store, customers must actually try on the clothes in a fitting room. However, this method takes time, reducing user convenience. Furthermore, due to the impact of COVID-19, many consumers are concerned about using fitting rooms. Therefore, there is a need for a method that allows customers to easily and quickly simulate trying on clothes. Furthermore, there is a need for a system that can provide personalized suggestions based on the user's past purchase history, providing a more satisfying shopping experience.
[0792] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0793] In this invention, the server includes a photographing device that captures the user's body shape data, a voice input device that recognizes the user's voice instructions, a display device that displays a clothing selection menu based on the user's voice instructions, a communication device that transmits information about the user's selected clothing to a generative AI model, a generative AI model device that generates a try-on image of the user wearing the selected clothing, a display device that displays the generated try-on image on a reflective surface, a communication device that completes the purchase process based on the user's voice instructions, a recognition device installed in a physical store that recognizes the user's face, and a database device that makes personalized suggestions based on past purchase history. This allows users to quickly check out and purchase clothing through a try-on simulation without actually trying on the clothing in a physical store. Furthermore, suggestions based on past purchase history also increase user satisfaction.
[0794] "Photographing means" refers to a device for acquiring the user's body shape data, and includes an image acquisition device such as a camera.
[0795] "Voice input means" refers to a device for recognizing a user's voice instructions, and includes a voice recording device such as a microphone.
[0796] "Display means" refers to a device that visually presents information to a user, including a display or screen.
[0797] "Communication means" refers to a network connection device for sending and receiving data, including an internet connection and Wi-Fi.
[0798] "Generative AI model means" refers to AI technology that generates try-on images of a user wearing the clothing selected by the user, and includes a generative AI model.
[0799] "Reflective surface" refers to a mirror or smart mirror that allows users to see themselves.
[0800] "Recognition means" refers to devices installed in physical stores that recognize users' faces, including facial recognition cameras and facial recognition software.
[0801] "Database means" refers to a system that manages data on users' past purchase history and personalized suggestions.
[0802] "Online Shop API" refers to the application programming interface for obtaining clothing information.
[0803] This invention provides a system that allows users to try on clothes without actually trying them on by using a smart full-length mirror in a physical store. The system is composed of the following hardware and software:
[0804] 1. Camera as a photography tool:
[0805] The camera is used to capture the user's body shape data. When the user stands in front of the smart mirror, the camera automatically captures the user's image and generates body shape data.
[0806] 2. Microphone as a means of voice input:
[0807] The microphone is used to recognize the user's voice instructions: when the user selects clothes or makes purchases by voice, the microphone captures the voice and inputs it into the system.
[0808] 3. Display as a means of presentation:
[0809] The display allows users to select clothing and check images of the clothing they are trying on based on voice instructions. Images of the selected clothing and images of the clothing being tried on based on the generative AI model are then displayed.
[0810] 4. Network connectivity as a means of communication:
[0811] The communication means is used to send information about the clothes selected by the user to the generative AI model and receive the generated try-on images, which is achieved via an internet connection or Wi-Fi.
[0812] 5. Generative AI model means:
[0813] The generative AI model generates try-on images based on the user's body data and the selected clothing. The generated try-on images are displayed on the screen in real time, allowing the user to see how the clothes will look when tried on.
[0814] 6. Smart mirrors as reflective surfaces:
[0815] The smart mirror is used to visually present the generated fitting images to the user. It functions as a normal mirror but has a built-in display that displays the fitting images.
[0816] 7. Facial recognition cameras as a means of recognition:
[0817] Facial recognition cameras are used to recognize users' faces in physical stores. When a user stands in front of a smart mirror, the camera recognizes their face and provides user information to the system.
[0818] 8. Database Means:
[0819] The database is used to manage users' past purchase history and provide personalized recommendations, and has the function of suggesting the most suitable clothing based on the user's purchase history.
[0820] As a concrete example, a user stands in front of a smart full-length mirror, and the system recognizes the user's face using a facial recognition camera. When the user gives a voice command such as "I want to see a red dress," the microphone recognizes the voice and a selection menu of red dresses appears on the display. When the user selects a specific dress, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. The generated try-on image is displayed on the smart mirror, allowing the user to check how the dress will look when tried on. Finally, when the user gives a voice command such as "I want to buy this dress," the purchase process is carried out via a communication means.
[0821] The following text can be used as an example of a prompt to input to the generative AI model:
[0822] Prompt: "Given a user's body data and an image of a red dress, generate an image of the user trying on the red dress. The user is a size medium."
[0823] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0824] Step 1:
[0825] The device will perform the initial settings
[0826] The device will perform the initial setup of the camera, microphone, and speaker, and check that they are working properly. Specifically, the device will acquire the user's body shape data from the camera, and prepare the microphone as a voice input method to receive voice instructions. After the setup is complete, the device will check that these devices are working properly.
[0827] Step 2:
[0828] Recognize the user's face and log them in
[0829] When a user stands in front of the smart mirror, the device's camera captures the user's face and uses facial recognition technology to identify the user. The facial image data is taken as input, and the user ID is identified as output. Based on the identified user ID, the server retrieves the user's account information and sends it to the device, which then displays an individually personalized interface.
[0830] Step 3:
[0831] Displaying a clothing selection menu based on the user's voice commands
[0832] When a user voices the instruction "I want to choose some clothes," the device's microphone recognizes this instruction and captures the voice data as input. Voice recognition technology is used to analyze the user's intention. Based on the analysis results, the server retrieves the appropriate product data via the online shop API and sends it to the device. A selection menu is then displayed on the screen as output.
[0833] Step 4:
[0834] Send user-selected clothing information to a generative AI model
[0835] When a user selects a specific item, the device sends detailed information about the selected clothing to the server. The server then sends this information to the generative AI model, which generates a prompt such as, "Please generate a try-on image using the user's body data and an image of the selected clothing." The selected clothing information is taken as input, and the prompt is sent as output to the generative AI model.
[0836] Step 5:
[0837] Try-on images are generated and displayed using a generative AI model
[0838] The generative AI model generates try-on images based on the user's body type data and selected clothing information. It receives body type data and clothing information as input and generates try-on images as output. These images are sent to the device via the server, and the device displays the generated try-on images on the smart mirror.
[0839] Step 6:
[0840] Proceed with purchases based on user voice commands
[0841] When a user gives a voice command such as "I want to buy this outfit," the device's microphone recognizes the voice and captures the voice data as input. The device analyzes the command and sends a purchase request to the server. The server confirms the purchase information via the online shop API and processes the payment. A purchase confirmation message is generated as output and sent to the device.
[0842] By following the steps above, users can experience trying on clothes without actually trying them on, and can easily complete the purchase process through the display on the smart mirror.
[0843] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0844] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions, providing users with an experience of trying on clothes through the mirror. Specific embodiments of the present invention are described below.
[0845] 1. Initial Setup
[0846] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly. The camera captures the user's body shape data and facial expressions, the microphone receives the user's voice instructions, and the speaker responds with voice. The emotion engine recognizes emotions from the user's facial expressions.
[0847] 2. User Identification and Login
[0848] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[0849] 3. Clothing choices
[0850] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0851] 4. Emotion Recognition and Suggestion
[0852] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0853] 5. Try-on simulation
[0854] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0855] 6. Purchase Procedure
[0856] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0857] Specific examples
[0858] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0859] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. Furthermore, the use of an emotion engine makes it possible to make suggestions more suited to the user and adjust visual effects, further improving the shopping experience.
[0860] The processing flow will be explained below.
[0861] Step 1:
[0862] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[0863] Step 2:
[0864] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[0865] Step 3:
[0866] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[0867] Step 4:
[0868] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[0869] Step 5:
[0870] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[0871] Step 6:
[0872] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[0873] Step 7:
[0874] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[0875] Step 8:
[0876] The emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions, which are then sent to the server.
[0877] Step 9:
[0878] The server generates additional clothing suggestions based on the recognized user emotions and sends a suggestion message to the terminal.
[0879] Step 10:
[0880] The device will display a suggestion message on the screen and the user will confirm it.
[0881] Step 11:
[0882] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[0883] Step 12:
[0884] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[0885] Step 13:
[0886] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[0887] Step 14:
[0888] If the emotion engine recognizes joy or satisfaction from the user's facial expression, it adjusts the visual effects of the try-on images (e.g., background color and lighting effects).
[0889] Step 15:
[0890] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[0891] Step 16:
[0892] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[0893] Step 17:
[0894] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[0895] Example 2
[0896] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0897] In conventional online shopping, users select products without trying them on, which can lead to a decrease in purchasing motivation and an increase in returns after purchase. Furthermore, users cannot see how the product will look when actually worn, which raises concerns about a decrease in satisfaction. Furthermore, suggestions do not take emotions into account, and visual effects are not well-adjusted, leaving a need for an improved user experience.
[0898] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0899] In this invention, the server includes a photographing means for capturing the user's body type data and facial expression, a voice input means for recognizing the user's voice instructions, a display means for displaying a clothing selection menu based on the user's voice instructions, an emotion engine for recognizing the user's emotions, a communication means for sending information about the clothing selected by the user to a generative AI model, a generation AI means for generating try-on images of the user wearing the selected clothing based on the user's body type data, a display means for displaying the generated try-on images on a mirror, an adjustment means for adjusting the visual effects of the try-on images based on the user's emotions, and a communication means for completing the purchase process based on the user's voice instructions. This allows the user to check the appearance of products as if they were trying them on, increasing their desire to purchase and enabling a satisfying online shopping experience.
[0900] "Photographing means for acquiring the user's body type data and facial expression" refers to means for capturing the user's body type data and facial expression using photographing equipment such as a camera, and acquiring them as digital data.
[0901] "A voice input means for recognizing a user's voice instructions" refers to a means for collecting a user's voice instructions using a voice input device such as a microphone, and analyzing the content of the instructions using voice recognition technology.
[0902] "Display means for displaying a clothing selection menu based on the user's voice instructions" refers to means for visually displaying a clothing selection menu based on the content of the user's voice instructions using a display or the like.
[0903] The "emotion engine that recognizes user emotions" is a software and hardware component that analyzes emotions from the user's facial expressions, tone of voice, etc., and recognizes the user's current emotions.
[0904] "Communication means for transmitting information about clothing selected by the user to the generative AI model" refers to a means for transmitting data about clothing selected by the user to the generative AI model via network communication.
[0905] "Generative AI means for generating try-on images of a user wearing selected clothing based on the user's body type data" refers to a generative AI technology used to generate try-on images of a user wearing selected clothing using the user's body type data as input.
[0906] The "display means for displaying the generated try-on images on a mirror" refers to a means for using a mirror display or the like to display the generated try-on images in real time.
[0907] The "adjustment means for adjusting the visual effects of the try-on images based on the user's emotions" is a means for dynamically adjusting the visual effects of the try-on images, such as the background color and lighting, according to the recognized user's emotions.
[0908] "Communication means for carrying out purchase procedures based on the user's voice instructions" refers to a communication means for transmitting information to a server when a user expresses their intention to purchase by voice, and for carrying out the online purchase procedure.
[0909] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions. This system provides users with an experience of trying on clothes through the mirror, enhancing online shopping satisfaction. Specific embodiments of the system are described below.
[0910] The system includes the following hardware and software components:
[0911] Camera: Captures the user's body data and facial expressions.
[0912] Microphone: Inputs the user's voice commands.
[0913] Speaker: Outputs voice responses from the system.
[0914] Emotion engine: Analyzes and recognizes emotions from the user's facial expressions. For example, Haarcascades or deep learning-based emotion recognition models are used.
[0915] Generative AI model: Generates images of the user trying on the selected clothing based on their body shape data. For example, GANs (generative artificial network) or deep learning-based image generation models are used.
[0916] Mirror display: Displays the generated fitting image in real time.
[0917] Communication method: Sending and receiving data via the Internet.
[0918] 1. Initial Setup
[0919] The server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. The device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly.
[0920] 2. User Identification and Login
[0921] When a user stands in front of the smart mirror, the device's camera recognizes the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access a personalized interface.
[0922] 3. Clothing choices
[0923] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[0924] 4. Emotion Recognition and Suggestion
[0925] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[0926] 5. Try-on simulation
[0927] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[0928] 6. Purchase Procedure
[0929] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[0930] Specific examples
[0931] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[0932] In this way, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. In addition, the use of an emotion engine makes it possible to make more user-friendly suggestions and adjust visual effects, further improving the shopping experience.
[0933] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0934] Step 1:
[0935] The server sets up a database to manage user account information, past purchase history, and product data. The input is user information and product metadata. It then uses a database management system (e.g., MySQL or PostgreSQL) to store this data and organizes the user's individual data. The output is an organized database that allows for personalized offers.
[0936] Step 2:
[0937] The device initializes the camera, microphone, speaker, and emotion engine. The input is the setting parameters of each device. Specifically, it calibrates the camera, adjusts the microphone sensitivity, sets the speaker volume, and initializes the emotion engine. The output is a properly functioning hardware device.
[0938] Step 3:
[0939] When a user stands in front of the device, the device uses the camera to recognize the user's face and displays a login screen. The input is a face image captured by the camera. A facial recognition algorithm (e.g., OpenCV or Dlib) is used to identify the user. The output is a login screen where the user enters their login information.
[0940] Step 4:
[0941] When a user enters their login information, the server authenticates it and displays the user's home screen. The input is a user ID and password. The authentication information is retrieved from a database and verified. The output is the authentication result and a personalized home screen.
[0942] Step 5:
[0943] When a user issues a voice command such as "I want to choose clothes," the device captures this with a microphone and converts it into text using a speech recognition engine. The input is the user's voice command. The speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes it and generates a text command. The output is the generated text command, and a selection menu is displayed based on it.
[0944] Step 6:
[0945] When a user selects a specific category or brand, the server uses the online shop API to retrieve the corresponding item list and sends it to the terminal. The input is the category or brand selected by the user. A request is sent to the online shop API to retrieve product data. The output is the retrieved item list.
[0946] Step 7:
[0947] When a user selects a specific item from the item list, the device sends information about that item to the generative AI model. The input is detailed information about the selected item. The generative AI model processes the data to generate a try-on image. The output is try-on image data.
[0948] Step 8:
[0949] A generative AI model generates try-on images based on the user's body data, and the device displays these images in real time on the mirror display. The input is the user's body data and selected item information. The generative AI model (e.g., GANs) executes the image generation process. The output is an image that looks as if the user is actually trying on the clothing.
[0950] Step 9:
[0951] While trying on clothes, the emotion engine analyzes the user's facial expressions through the camera and recognizes the current emotion. The input is the user's facial expression data captured by the camera. The emotion recognition algorithm analyzes the data. The output is the recognized emotion information.
[0952] Step 10:
[0953] Based on the recognized emotion, the server suggests additional outfits that suit the user and adjusts the visual effects. The input is the recognized emotion information and past purchase history. The suggestion algorithm generates an appropriate product list. The output is additional suggested outfits and adjusted visual effects.
[0954] Step 11:
[0955] When a user wishes to make a purchase, they issue a voice command such as "I would like to purchase this clothing." The device captures this and sends a purchase request to the server. The input is a text command generated by the voice command. The voice recognition engine analyzes the text and generates a purchase request. The output is the purchase request data.
[0956] Step 12:
[0957] The server verifies the purchase information and processes the payment. The input is the purchase request and the user's payment information. The payment process is performed via the online shop API. The output is a purchase confirmation message, which is sent to the terminal and displayed to the user.
[0958] (Application example 2)
[0959] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0960] In today's online shopping environment, the inability to actually try on clothes is a major source of anxiety for users and a hurdle to purchasing. Furthermore, the lack of personalized suggestions that take emotions into account when selecting products also limits the shopping experience. The objective of this invention is to solve these problems and provide users with a more satisfying shopping experience.
[0961] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0962] In this invention, the system includes optical means for acquiring a user's body shape data, voice input means for recognizing the user's voice instructions, display means for displaying a clothing selection menu based on the user's voice instructions, camera means for acquiring the user's facial expression data, emotion analysis means for recognizing the user's emotions, communication means for transmitting information about the user's selected clothing to a generative AI model, generative AI means for generating try-on images of the user wearing the selected clothing, optical display means for displaying the generated try-on images, suggestion means for suggesting additional clothing based on the user's emotion recognition, and communication means for completing the purchase process based on the user's voice instructions. This allows the user to have an online experience similar to trying on actual clothing. Furthermore, personalized suggestions based on emotion analysis are possible, increasing user satisfaction.
[0963] "Optical means" refers to a device for acquiring the user's body shape data, and includes a camera and a sensor.
[0964] "Voice input means" refers to a device that uses a microphone or voice recognition technology to recognize a user's voice instructions and input them into the system.
[0965] "Display means" refers to a display or screen for providing visual information to the user, and displays a clothing selection menu based on the user's voice instructions.
[0966] The "camera means" is a device used to acquire facial expression data of the user, such as a high-resolution camera.
[0967] "Emotion analysis means" refers to software or algorithms that recognize emotions based on a user's facial expression data.
[0968] "Communication means" refers to the network technology and interface used to send information about the clothing selected by the user to the generative AI model and to complete the purchase process.
[0969] The "generative AI means" is a system that includes an artificial intelligence model or algorithm for generating try-on images of the user wearing the selected clothing.
[0970] "Optical display means" refers to a display or screen for displaying the generated try-on image, including the display of smart glasses.
[0971] "Suggestion mechanism" refers to an algorithm or system for providing additional clothing suggestions based on the user's emotion recognition.
[0972] This invention aims to enhance the online and brick-and-mortar shopping experience by combining a system with an emotion analysis engine that recognizes user emotions. The system acquires the user's body shape data and facial expression data, and then simulates trying on clothes and carrying out the purchase process based on that data.
[0973] First, the user puts on the smart glasses and stands in front of the system. The system uses optical means (high-resolution camera) to acquire the user's body shape data and camera means to acquire the user's facial expression data. This allows the system to collect basic data about the user and proceed to the next process.
[0974] Next, the user gives a voice command using the voice input means (microphone). The voice input means recognizes this and displays a clothing selection menu on the display means (the display of the smart glasses) based on the voice command. When the user selects a specific clothing item, the information is sent to the generative AI model via the communication means.
[0975] The generation AI means generates a try-on image of the clothing when worn by the user based on the received body shape data and clothing information. This generated try-on image is displayed in real time on the optical display means (the display of the smart glasses). At this time, the emotion analysis means recognizes emotions based on the user's facial expressions and detects whether the user looks happy or unhappy. Based on the results of the emotion analysis, the system uses the suggestion means to suggest additional clothing items.
[0976] For example, if a user voices the command "I want to look at denim jackets," the voice input means recognizes this and a category menu is displayed. When the user selects a specific denim jacket, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. This try-on image is displayed in real time on the smart glasses display, allowing the user to see the item as if they were actually trying it on. If the user is smiling, the emotion analysis means detects this and the suggestion means suggests additional denim jackets of similar style and color.
[0977] An example of a specific prompt would be, "Generate a fitting image of a denim jacket based on the user's body type data. The user's height is 170 cm, weight is 65 kg, shoulder width is 45 cm, and preferred style is casual." This prompt is sent to the generation AI model, which then generates an appropriate fitting image.
[0978] The system utilizes the Microsoft Azure Emotion API for emotion analysis and the Google Speech-to-Text API for voice command recognition, and uses advanced artificial intelligence technologies such as GPT-4 for generative AI models to respond to users in real time, enabling users to enjoy a smooth and satisfying shopping experience in-store and online.
[0979] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0980] Step 1:
[0981] The user puts on the smart glasses and stands in front of the system. The terminal (smart glasses) acquires the user's body shape data using optical means (high-resolution camera). Specifically, the camera captures the user's full-body image, which is then analyzed using image processing software to acquire data such as height, weight, and shoulder width. The input is the "user's full-body image," and the output is the "user's body shape data."
[0982] Step 2:
[0983] The user issues voice instructions using the voice input means (microphone). The device (smart glasses) captures the user's voice instructions through the voice input means. The voice data is converted into text by voice recognition software (Google Speech-to-Text API). The input is "voice data" and the output is "text data." Based on the voice instructions, the device displays a clothing selection menu on the display means (the display of the smart glasses).
[0984] Step 3:
[0985] The user selects a specific piece of clothing from the displayed menu. The device (smart glasses) recognizes the user's selection using a touch interface or eye tracking. Information about the selected clothing is sent to the server via communication means. The input is "user selection information" and the output is "transmission of clothing information."
[0986] Step 4:
[0987] The server sends the received clothing information and the user's body data to the generative AI model. The generative AI model generates a try-on image of the selected clothing based on the user's body data. Specifically, it uses a prompt to give instructions to the AI model. An example of a prompt is: "Based on the user's body data, please generate an image of a denim jacket being tried on. The user is 170 cm tall, weighs 65 kg, has a shoulder width of 45 cm, and prefers a casual style." The input is "body data and prompt," and the output is "a try-on image."
[0988] Step 5:
[0989] The try-on images generated by the generation AI means are sent to the terminal (smart glasses) via the communication means. The terminal displays the try-on images in real time using the optical display means (the display of the smart glasses). The input is the "try-on image" and the output is the "display of the try-on image."
[0990] Step 6:
[0991] The device (smart glasses) acquires the user's facial expression data using a camera. It analyzes the user's emotions using an emotion analysis tool (Microsoft Azure Emotion API). In this process, the user's facial expression data is input and emotion data is output as the analysis result. The input is "facial expression data" and the output is "emotion data."
[0992] Step 7:
[0993] The server uses the suggestion means to suggest additional clothing based on the emotion data. Specifically, if the user shows a happy emotion, it automatically selects clothing of a similar style and color and suggests it to the user. The input is "emotion data" and the output is "additional clothing suggestions." The terminal displays the suggested additional clothing on its display.
[0994] Step 8:
[0995] The user issues a voice command such as "I would like to purchase this clothing." The device captures this voice command using a voice input means and converts it into text using voice recognition software. This text data is sent to the server via a communication means. The input is "voice data" and the output is "text data of the purchase command."
[0996] Step 9:
[0997] The server receives the purchase instruction and carries out the purchase procedure via the online shop's API. Specifically, it sends the purchase information to the API and processes the payment. The input is "text data of the purchase instruction" and the output is "purchase procedure completed." Once the procedure is complete, a confirmation message is sent to the terminal to notify the user.
[0998] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0999] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1000] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1001] [Fourth embodiment]
[1002] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1003] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1004] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1005] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1006] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1007] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1008] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1009] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1010] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1011] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1012] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1013] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1014] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1015] The present invention provides an online shopping support system using a smart full-length mirror, which allows users to experience the experience of trying on clothes in real life. Specific embodiments of the present invention will be described below.
[1016] 1. Initial Setup
[1017] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, and speaker and verifies that they are working properly. The camera captures the user's body shape data, the microphone receives the user's voice instructions, and the speaker provides voice responses.
[1018] 2. User Identification and Login
[1019] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[1020] 3. Clothing choices
[1021] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[1022] 4. Try-on simulation
[1023] When a user selects an item, the device sends the clothing information to the AI generator, which then generates an image of the selected clothing based on the user's body type data. This image is displayed in real time on the device's mirror, making it appear as if the user is actually trying it on.
[1024] 5. Purchase Procedure
[1025] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[1026] Specific examples
[1027] For example, let's say a user wants to try on a denim jacket. When the user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about that denim jacket is sent to the generation AI, which generates a fitting image adapted to the user's body type. This fitting image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user wants to purchase it, the device recognizes the command and starts the purchase process by saying, "I want to buy this jacket."
[1028] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, thereby eliminating concerns about online shopping and increasing user satisfaction.
[1029] The processing flow will be explained below.
[1030] Step 1:
[1031] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[1032] Step 2:
[1033] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[1034] Step 3:
[1035] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[1036] Step 4:
[1037] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[1038] Step 5:
[1039] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[1040] Step 6:
[1041] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[1042] Step 7:
[1043] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[1044] Step 8:
[1045] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[1046] Step 9:
[1047] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[1048] Step 10:
[1049] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[1050] Step 11:
[1051] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[1052] Step 12:
[1053] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[1054] Step 13:
[1055] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[1056] Example 1
[1057] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1058] With traditional online shopping, customers could not actually try on the products, which meant they could not check the size or fit, which often led to anxiety when making a purchase. Furthermore, it was not possible to provide optimal suggestions for individual users, so diverse needs could not be met. Furthermore, the purchasing process was complicated, making it difficult to improve the user experience.
[1059] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1060] In this invention, the server includes an image acquisition means for acquiring a user's body type data, a voice recognition means for recognizing the user's voice instructions, a display means for displaying a product selection menu based on the user's voice instructions, a communication means for sending information about the product selected by the user to a generative AI model, a generative AI means for generating try-on images of the user wearing the selected product, a display means for displaying the generated try-on images on a display device, and a communication means for carrying out the purchase procedure based on the user's voice instructions. This allows the user to have an experience that feels like they are actually trying on the product, eliminating the anxiety of online shopping and enabling individually optimized suggestions, making the purchase procedure smoother.
[1061] "Image acquisition means" refers to devices or technologies for acquiring a user's body shape data, and specifically includes cameras and 3D scanners.
[1062] "Voice recognition means" refers to devices or technology for recognizing a user's voice instructions, and specifically includes a microphone and voice recognition software.
[1063] "Display means" refers to devices or technologies for displaying a user interface, and specifically includes displays and touch screens.
[1064] "Communication means" refers to devices and technologies for transmitting and receiving data, and specifically includes internet connections and wireless communication technologies.
[1065] A "generative AI model" is an artificial intelligence model that generates try-on images based on the user's body type data and selected product information.
[1066] "Generative AI means" refers to technology or equipment for generating try-on images from specified data using a generative AI model.
[1067] The "display device" refers to a device for visually presenting the generated try-on images to the user, and specifically includes a smart mirror or display.
[1068] "Authentication means" refers to devices or technologies for authenticating users, and specifically includes facial recognition systems and login systems.
[1069] A "database" is a system for organizing and storing information, and specifically includes a system for storing information about products in an online shop.
[1070] The present invention is an online shopping support system that provides users with an experience similar to trying on clothes in a physical store. This system acquires the user's body type data, selects products based on voice instructions, generates try-on images using a generative AI model, and performs the entire process from purchase to purchase. Specific embodiments are described below.
[1071] The server prepares a database to manage user account information, past purchase history, and product data, enabling personalized recommendations to be made to users. The database used can be a commercial or open-source relational database management system such as MySQL or PostgreSQL.
[1072] The device is a smart mirror and is configured as follows: the camera is used to acquire the user's body shape data, and can be, for example, a high-performance webcam or a 3D scanner; the microphone is used to receive the user's voice instructions, and is preferably, for example, a directional microphone or a microphone with noise-canceling capabilities; and the speaker is used to provide voice feedback.
[1073] When a user stands in front of the smart mirror, the device recognizes the user's face through the camera. The facial recognition technology uses libraries and frameworks such as OpenCV and TensorFlow. Once authentication is complete, the device displays a login screen, allowing the user to enter login information via voice commands or the touch panel.
[1074] When a user says, "I want to choose clothes," the device's microphone captures the voice command and analyzes it using the Google Speech-to-Text API. Based on the analysis results, the device displays a selection menu, allowing the user to select a specific category or brand. Once the user confirms their selection, the server retrieves product information from the online store's database and sends it to the device.
[1075] Next, the user selects the product they want to try on, and the device sends that product information to a generative AI model. The generative AI model uses OpenAI's generative AI technology, for example. This model generates a fitting image based on the user's body data. The generated fitting image is displayed in real time on the smart mirror, allowing the user to experience the experience of actually trying on the product.
[1076] For example, if a user says, "I want to look at denim jackets," the device recognizes this voice command and displays a category of denim jackets. When the user selects a specific denim jacket, the information about the denim jacket is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is instantly displayed on the smart mirror, allowing the user to check the denim jacket as if they were trying it on. When the user says, "I want to buy this jacket," the device recognizes the voice command and begins the purchase process. The server uses the online shop API to process the payment and sends a confirmation message to the device.
[1077] As described above, the online shopping support system of the present invention allows users to enjoy a realistic try-on experience from the comfort of their own home, thereby eliminating concerns about online shopping and contributing to improved user satisfaction.
[1078] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1079] Step 1:
[1080] Database Setup
[1081] The server prepares a database and manages user account information, past purchase history, and product data. Database management systems such as MySQL and PostgreSQL are used for this. Specifically, the server creates tables for registering new users and for storing detailed product information.
[1082] Input: User information, purchase history, product data
[1083] Data processing: storing data in a database
[1084] Output: Database completed, data stored
[1085] Step 2:
[1086] Initial Hardware Setup
[1087] The device's camera, microphone, and speaker are initially configured. At this time, the camera resolution and frame rate are set, the microphone sensitivity is adjusted, and the speaker volume is set. For the camera, a high-performance webcam or 3D scanner is used, and for the microphone, a directional microphone with noise-canceling functions is used. For the speaker, a speaker that can output clear audio is selected.
[1088] Input: Hardware configuration information
[1089] Data processing: Applying setting parameters
[1090] Output: Camera, microphone, and speaker initial settings completed
[1091] Step 3:
[1092] User facial recognition and login
[1093] When a user stands in front of the smart mirror, the device's camera recognizes the user's face. Facial recognition technology uses OpenCV and TensorFlow. If facial recognition is successful, the device displays a login screen and the user enters their login information. The server receives this and authenticates it against the database. If authentication is successful, the home screen is displayed.
[1094] Input: Face image, login information
[1095] Data processing: facial recognition, login information authentication
[1096] Output: Authentication result, home screen display
[1097] Step 4:
[1098] Voice command recognition and product selection
[1099] When a user says, "I want to choose some clothes," the device's microphone captures this voice command and converts the voice data into text using the Google Speech-to-Text API. Based on this text data, the device displays a product selection menu. When the user selects a specific category or brand, this information is sent to the server. The server retrieves a list of corresponding products through the online shop's database and sends it to the device. The device then displays the retrieved item list to the user.
[1100] Input: Voice data, text data, product category
[1101] Data processing: voice analysis, item list acquisition
[1102] Output: Product selection menu, item list display
[1103] Step 5:
[1104] Try-on simulation
[1105] When a user selects an item, the device sends the product information and the user's body shape data to a generative AI model, which uses OpenAI technology to generate images of the selected item being tried on. The generated images are then displayed in real time on a smart mirror, giving the user the experience of actually trying the item on.
[1106] Input: Product information, body type data
[1107] Data processing: Generative AI generates try-on images
[1108] Output: Show try-on image
[1109] Step 6:
[1110] Purchase procedure
[1111] When a user says, "I want to buy this outfit," the device's microphone captures this voice command and converts it into text data using the Google Speech-to-Text API. Based on this text data, the device generates a purchase request and sends it to the server. The server processes the payment via the online shop API, and once the order is complete, a confirmation message is sent to the device and displayed to the user.
[1112] Input: Voice data, purchase request
[1113] Data processing: voice analysis, payment processing
[1114] Output: Display purchase confirmation message
[1115] (Application example 1)
[1116] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1117] Currently, when purchasing clothes in a physical store, customers must actually try on the clothes in a fitting room. However, this method takes time, reducing user convenience. Furthermore, due to the impact of COVID-19, many consumers are concerned about using fitting rooms. Therefore, there is a need for a method that allows customers to easily and quickly simulate trying on clothes. Furthermore, there is a need for a system that can provide personalized suggestions based on the user's past purchase history, providing a more satisfying shopping experience.
[1118] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1119] In this invention, the server includes a photographing device that captures the user's body shape data, a voice input device that recognizes the user's voice instructions, a display device that displays a clothing selection menu based on the user's voice instructions, a communication device that transmits information about the user's selected clothing to a generative AI model, a generative AI model device that generates a try-on image of the user wearing the selected clothing, a display device that displays the generated try-on image on a reflective surface, a communication device that completes the purchase process based on the user's voice instructions, a recognition device installed in a physical store that recognizes the user's face, and a database device that makes personalized suggestions based on past purchase history. This allows users to quickly check out and purchase clothing through a try-on simulation without actually trying on the clothing in a physical store. Furthermore, suggestions based on past purchase history also increase user satisfaction.
[1120] "Photographing means" refers to a device for acquiring the user's body shape data, and includes an image acquisition device such as a camera.
[1121] "Voice input means" refers to a device for recognizing a user's voice instructions, and includes a voice recording device such as a microphone.
[1122] "Display means" refers to a device that visually presents information to a user, including a display or screen.
[1123] "Communication means" refers to a network connection device for sending and receiving data, including an internet connection and Wi-Fi.
[1124] "Generative AI model means" refers to AI technology that generates try-on images of a user wearing the clothing selected by the user, and includes a generative AI model.
[1125] "Reflective surface" refers to a mirror or smart mirror that allows users to see themselves.
[1126] "Recognition means" refers to devices installed in physical stores that recognize users' faces, including facial recognition cameras and facial recognition software.
[1127] "Database means" refers to a system that manages data on users' past purchase history and personalized suggestions.
[1128] "Online Shop API" refers to the application programming interface for obtaining clothing information.
[1129] This invention provides a system that allows users to try on clothes without actually trying them on by using a smart full-length mirror in a physical store. The system is composed of the following hardware and software:
[1130] 1. Camera as a photography tool:
[1131] The camera is used to capture the user's body shape data. When the user stands in front of the smart mirror, the camera automatically captures the user's image and generates body shape data.
[1132] 2. Microphone as a means of voice input:
[1133] The microphone is used to recognize the user's voice instructions: when the user selects clothes or makes purchases by voice, the microphone captures the voice and inputs it into the system.
[1134] 3. Display as a means of presentation:
[1135] The display allows users to select clothing and check images of the clothing they are trying on based on voice instructions. Images of the selected clothing and images of the clothing being tried on based on the generative AI model are then displayed.
[1136] 4. Network connectivity as a means of communication:
[1137] The communication means is used to send information about the clothes selected by the user to the generative AI model and receive the generated try-on images, which is achieved via an internet connection or Wi-Fi.
[1138] 5. Generative AI model means:
[1139] The generative AI model generates try-on images based on the user's body data and the selected clothing. The generated try-on images are displayed on the screen in real time, allowing the user to see how the clothes will look when tried on.
[1140] 6. Smart mirrors as reflective surfaces:
[1141] The smart mirror is used to visually present the generated fitting images to the user. It functions as a normal mirror but has a built-in display that displays the fitting images.
[1142] 7. Facial recognition cameras as a means of recognition:
[1143] Facial recognition cameras are used to recognize users' faces in physical stores. When a user stands in front of a smart mirror, the camera recognizes their face and provides user information to the system.
[1144] 8. Database Means:
[1145] The database is used to manage users' past purchase history and provide personalized recommendations, and has the function of suggesting the most suitable clothing based on the user's purchase history.
[1146] As a concrete example, a user stands in front of a smart full-length mirror, and the system recognizes the user's face using a facial recognition camera. When the user gives a voice command such as "I want to see a red dress," the microphone recognizes the voice and a selection menu of red dresses appears on the display. When the user selects a specific dress, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. The generated try-on image is displayed on the smart mirror, allowing the user to check how the dress will look when tried on. Finally, when the user gives a voice command such as "I want to buy this dress," the purchase process is carried out via a communication means.
[1147] The following text can be used as an example of a prompt to input to the generative AI model:
[1148] Prompt: "Given a user's body data and an image of a red dress, generate an image of the user trying on the red dress. The user is a size medium."
[1149] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1150] Step 1:
[1151] The device will perform the initial settings
[1152] The device will perform the initial setup of the camera, microphone, and speaker, and check that they are working properly. Specifically, the device will acquire the user's body shape data from the camera, and prepare the microphone as a voice input method to receive voice instructions. After the setup is complete, the device will check that these devices are working properly.
[1153] Step 2:
[1154] Recognize the user's face and log them in
[1155] When a user stands in front of the smart mirror, the device's camera captures the user's face and uses facial recognition technology to identify the user. The facial image data is taken as input, and the user ID is identified as output. Based on the identified user ID, the server retrieves the user's account information and sends it to the device, which then displays an individually personalized interface.
[1156] Step 3:
[1157] Displaying a clothing selection menu based on the user's voice commands
[1158] When a user voices the instruction "I want to choose some clothes," the device's microphone recognizes this instruction and captures the voice data as input. Voice recognition technology is used to analyze the user's intention. Based on the analysis results, the server retrieves the appropriate product data via the online shop API and sends it to the device. A selection menu is then displayed on the screen as output.
[1159] Step 4:
[1160] Send user-selected clothing information to a generative AI model
[1161] When a user selects a specific item, the device sends detailed information about the selected clothing to the server. The server then sends this information to the generative AI model, which generates a prompt such as, "Please generate a try-on image using the user's body data and an image of the selected clothing." The selected clothing information is taken as input, and the prompt is sent as output to the generative AI model.
[1162] Step 5:
[1163] Try-on images are generated and displayed using a generative AI model
[1164] The generative AI model generates try-on images based on the user's body type data and selected clothing information. It receives body type data and clothing information as input and generates try-on images as output. These images are sent to the device via the server, and the device displays the generated try-on images on the smart mirror.
[1165] Step 6:
[1166] Proceed with purchases based on user voice commands
[1167] When a user gives a voice command such as "I want to buy this outfit," the device's microphone recognizes the voice and captures the voice data as input. The device analyzes the command and sends a purchase request to the server. The server confirms the purchase information via the online shop API and processes the payment. A purchase confirmation message is generated as output and sent to the device.
[1168] By following the steps above, users can experience trying on clothes without actually trying them on, and can easily complete the purchase process through the display on the smart mirror.
[1169] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1170] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions, providing users with an experience of trying on clothes through the mirror. Specific embodiments of the present invention are described below.
[1171] 1. Initial Setup
[1172] First, the server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. Next, the device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly. The camera captures the user's body shape data and facial expressions, the microphone receives the user's voice instructions, and the speaker responds with voice. The emotion engine recognizes emotions from the user's facial expressions.
[1173] 2. User Identification and Login
[1174] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access their personalized interface.
[1175] 3. Clothing choices
[1176] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[1177] 4. Emotion Recognition and Suggestion
[1178] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[1179] 5. Try-on simulation
[1180] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[1181] 6. Purchase Procedure
[1182] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[1183] Specific examples
[1184] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[1185] As described above, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. Furthermore, the use of an emotion engine makes it possible to make suggestions more suited to the user and adjust visual effects, further improving the shopping experience.
[1186] The processing flow will be explained below.
[1187] Step 1:
[1188] When a user stands in front of the smart mirror, the device uses the camera to recognize the user's face, and if facial recognition is successful, the login screen is displayed.
[1189] Step 2:
[1190] The user enters their login information and is authenticated. The information is sent to the server, which then authenticates them. If authentication is successful, the home screen is displayed.
[1191] Step 3:
[1192] The user issues a voice command such as "I want to choose clothes." The microphone recognizes and analyzes the voice command.
[1193] Step 4:
[1194] The device will display a menu of choices based on your voice commands, including options for specific categories and brands.
[1195] Step 5:
[1196] The user selects a specific category (e.g., denim jackets) using a touchscreen or voice input.
[1197] Step 6:
[1198] The device sends the selected category information to the server, which then uses the online shop API to retrieve the corresponding item list.
[1199] Step 7:
[1200] The retrieved item list is sent to the device, which displays it to the user, allowing the user to select an item.
[1201] Step 8:
[1202] The emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions, which are then sent to the server.
[1203] Step 9:
[1204] The server generates additional clothing suggestions based on the recognized user emotions and sends a suggestion message to the terminal.
[1205] Step 10:
[1206] The device will display a suggestion message on the screen and the user will confirm it.
[1207] Step 11:
[1208] The user selects a specific item, and this selection information is sent to the generation AI via the server.
[1209] Step 12:
[1210] The generation AI generates an image of the item being tried on, and this generated image is sent back to the device.
[1211] Step 13:
[1212] The device displays the generated fitting image on the mirror, allowing the user to see the fitting results in real time.
[1213] Step 14:
[1214] If the emotion engine recognizes joy or satisfaction from the user's facial expression, it adjusts the visual effects of the try-on images (e.g., background color and lighting effects).
[1215] Step 15:
[1216] If a user issues a voice command such as "I want to buy this outfit," the device will recognize the voice command and send a purchase request to the server.
[1217] Step 16:
[1218] The server verifies the purchase request and processes the payment using the online store API, sending the necessary payment information.
[1219] Step 17:
[1220] Once the order is complete, the server sends a confirmation message to the device, which displays a confirmation screen to let the user know the purchase is complete.
[1221] Example 2
[1222] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1223] In conventional online shopping, users select products without trying them on, which can lead to a decrease in purchasing motivation and an increase in returns after purchase. Furthermore, users cannot see how the product will look when actually worn, which raises concerns about a decrease in satisfaction. Furthermore, suggestions do not take emotions into account, and visual effects are not well-adjusted, leaving a need for an improved user experience.
[1224] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1225] In this invention, the server includes a photographing means for capturing the user's body type data and facial expression, a voice input means for recognizing the user's voice instructions, a display means for displaying a clothing selection menu based on the user's voice instructions, an emotion engine for recognizing the user's emotions, a communication means for sending information about the clothing selected by the user to a generative AI model, a generation AI means for generating try-on images of the user wearing the selected clothing based on the user's body type data, a display means for displaying the generated try-on images on a mirror, an adjustment means for adjusting the visual effects of the try-on images based on the user's emotions, and a communication means for completing the purchase process based on the user's voice instructions. This allows the user to check the appearance of products as if they were trying them on, increasing their desire to purchase and enabling a satisfying online shopping experience.
[1226] "Photographing means for acquiring the user's body type data and facial expression" refers to means for capturing the user's body type data and facial expression using photographing equipment such as a camera, and acquiring them as digital data.
[1227] "A voice input means for recognizing a user's voice instructions" refers to a means for collecting a user's voice instructions using a voice input device such as a microphone, and analyzing the content of the instructions using voice recognition technology.
[1228] "Display means for displaying a clothing selection menu based on the user's voice instructions" refers to means for visually displaying a clothing selection menu based on the content of the user's voice instructions using a display or the like.
[1229] The "emotion engine that recognizes user emotions" is a software and hardware component that analyzes emotions from the user's facial expressions, tone of voice, etc., and recognizes the user's current emotions.
[1230] "Communication means for transmitting information about clothing selected by the user to the generative AI model" refers to a means for transmitting data about clothing selected by the user to the generative AI model via network communication.
[1231] "Generative AI means for generating try-on images of a user wearing selected clothing based on the user's body type data" refers to a generative AI technology used to generate try-on images of a user wearing selected clothing using the user's body type data as input.
[1232] The "display means for displaying the generated try-on images on a mirror" refers to a means for using a mirror display or the like to display the generated try-on images in real time.
[1233] The "adjustment means for adjusting the visual effects of the try-on images based on the user's emotions" is a means for dynamically adjusting the visual effects of the try-on images, such as the background color and lighting, according to the recognized user's emotions.
[1234] "Communication means for carrying out purchase procedures based on the user's voice instructions" refers to a communication means for transmitting information to a server when a user expresses their intention to purchase by voice, and for carrying out the online purchase procedure.
[1235] The present invention is an online shopping support system that uses a smart full-length mirror combined with an emotion engine that recognizes user emotions. This system provides users with an experience of trying on clothes through the mirror, enhancing online shopping satisfaction. Specific embodiments of the system are described below.
[1236] The system includes the following hardware and software components:
[1237] Camera: Captures the user's body data and facial expressions.
[1238] Microphone: Inputs the user's voice commands.
[1239] Speaker: Outputs voice responses from the system.
[1240] Emotion engine: Analyzes and recognizes emotions from the user's facial expressions. For example, Haarcascades or deep learning-based emotion recognition models are used.
[1241] Generative AI model: Generates images of the user trying on the selected clothing based on their body shape data. For example, GANs (generative artificial network) or deep learning-based image generation models are used.
[1242] Mirror display: Displays the generated fitting image in real time.
[1243] Communication method: Sending and receiving data via the Internet.
[1244] 1. Initial Setup
[1245] The server sets up a database and manages the user's account information, past purchase history, and product data. This enables personalized suggestions based on the user's past purchase history. The device (smart mirror) initializes the camera, microphone, speaker, and emotion engine and verifies that they are working properly.
[1246] 2. User Identification and Login
[1247] When a user stands in front of the smart mirror, the device's camera recognizes the user's face and displays a login screen. After the user enters their login information, the server authenticates the login information and displays the home screen, allowing the user to access a personalized interface.
[1248] 3. Clothing choices
[1249] When a user says, "I want to choose some clothes," the device recognizes this voice command through the microphone and displays a selection menu using the display means. When the user selects a specific category or brand, the server retrieves the corresponding item list via the online shop's API and sends it to the device. The device then displays the retrieved item list to the user, allowing them to make a selection.
[1250] 4. Emotion Recognition and Suggestion
[1251] While the user is looking at items, the emotion engine analyzes the user's facial expressions through the camera and recognizes their emotions. Based on the recognized emotions, the server will suggest additional clothing that matches the user's preferences. For example, if the user shows a happy expression, clothing of a similar style and color will be suggested.
[1252] 5. Try-on simulation
[1253] When a user selects a specific item, the device sends the clothing information to the generation AI. The generation AI then generates a try-on image of the selected clothing based on the user's body data. This generated image is displayed in real time in the device's mirror, making it appear as if the user is actually trying it on. Furthermore, the emotion engine recognizes the user's emotions and adjusts the visual effects of the try-on image (e.g., background color and lighting effects).
[1254] 6. Purchase Procedure
[1255] If the user says, "I want to buy this outfit," the device recognizes this voice command and sends a purchase request to the server. The server verifies the purchase information and processes the payment via the online shop API. Once the order is complete, the server sends a confirmation message to the device, which displays it to the user.
[1256] Specific examples
[1257] For example, let's say a user wants to try on a denim jacket. If the user says, "I want to see denim jackets," the device recognizes this voice command and displays a category of denim jackets. If the user selects a specific denim jacket, the denim jacket information is sent to the generation AI, which generates a try-on image adapted to the user's body type. This try-on image is displayed in real time on a smart full-length mirror, allowing the user to see the denim jacket as if they were actually trying it on. If the user shows a happy expression, the emotion engine recognizes this emotion, and the server suggests denim jackets with similar designs and colors. If the user wants to purchase it, the device recognizes the command and begins the purchase process.
[1258] In this way, the present invention provides an online shopping support system that allows users to have an experience similar to trying on clothes in a real store, eliminating the anxiety of online shopping and increasing user satisfaction. In addition, the use of an emotion engine makes it possible to make more user-friendly suggestions and adjust visual effects, further improving the shopping experience.
[1259] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1260] Step 1:
[1261] The server sets up a database to manage user account information, past purchase history, and product data. The input is user information and product metadata. It then uses a database management system (e.g., MySQL or PostgreSQL) to store this data and organizes the user's individual data. The output is an organized database that allows for personalized offers.
[1262] Step 2:
[1263] The device initializes the camera, microphone, speaker, and emotion engine. The input is the setting parameters of each device. Specifically, it calibrates the camera, adjusts the microphone sensitivity, sets the speaker volume, and initializes the emotion engine. The output is a properly functioning hardware device.
[1264] Step 3:
[1265] When a user stands in front of the device, the device uses the camera to recognize the user's face and displays a login screen. The input is a face image captured by the camera. A facial recognition algorithm (e.g., OpenCV or Dlib) is used to identify the user. The output is a login screen where the user enters their login information.
[1266] Step 4:
[1267] When a user enters their login information, the server authenticates it and displays the user's home screen. The input is a user ID and password. The authentication information is retrieved from a database and verified. The output is the authentication result and a personalized home screen.
[1268] Step 5:
[1269] When a user issues a voice command such as "I want to choose clothes," the device captures this with a microphone and converts it into text using a speech recognition engine. The input is the user's voice command. The speech recognition engine (e.g., Google Cloud Speech-to-Text) analyzes it and generates a text command. The output is the generated text command, and a selection menu is displayed based on it.
[1270] Step 6:
[1271] When a user selects a specific category or brand, the server uses the online shop API to retrieve the corresponding item list and sends it to the terminal. The input is the category or brand selected by the user. A request is sent to the online shop API to retrieve product data. The output is the retrieved item list.
[1272] Step 7:
[1273] When a user selects a specific item from the item list, the device sends information about that item to the generative AI model. The input is detailed information about the selected item. The generative AI model processes the data to generate a try-on image. The output is try-on image data.
[1274] Step 8:
[1275] A generative AI model generates try-on images based on the user's body data, and the device displays these images in real time on the mirror display. The input is the user's body data and selected item information. The generative AI model (e.g., GANs) executes the image generation process. The output is an image that looks as if the user is actually trying on the clothing.
[1276] Step 9:
[1277] While trying on clothes, the emotion engine analyzes the user's facial expressions through the camera and recognizes the current emotion. The input is the user's facial expression data captured by the camera. The emotion recognition algorithm analyzes the data. The output is the recognized emotion information.
[1278] Step 10:
[1279] Based on the recognized emotion, the server suggests additional outfits that suit the user and adjusts the visual effects. The input is the recognized emotion information and past purchase history. The suggestion algorithm generates an appropriate product list. The output is additional suggested outfits and adjusted visual effects.
[1280] Step 11:
[1281] When a user wishes to make a purchase, they issue a voice command such as "I would like to purchase this clothing." The device captures this and sends a purchase request to the server. The input is a text command generated by the voice command. The voice recognition engine analyzes the text and generates a purchase request. The output is the purchase request data.
[1282] Step 12:
[1283] The server verifies the purchase information and processes the payment. The input is the purchase request and the user's payment information. The payment process is performed via the online shop API. The output is a purchase confirmation message, which is sent to the terminal and displayed to the user.
[1284] (Application example 2)
[1285] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1286] In today's online shopping environment, the inability to actually try on clothes is a major source of anxiety for users and a hurdle to purchasing. Furthermore, the lack of personalized suggestions that take emotions into account when selecting products also limits the shopping experience. The objective of this invention is to solve these problems and provide users with a more satisfying shopping experience.
[1287] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1288] In this invention, the system includes optical means for acquiring a user's body shape data, voice input means for recognizing the user's voice instructions, display means for displaying a clothing selection menu based on the user's voice instructions, camera means for acquiring the user's facial expression data, emotion analysis means for recognizing the user's emotions, communication means for transmitting information about the user's selected clothing to a generative AI model, generative AI means for generating try-on images of the user wearing the selected clothing, optical display means for displaying the generated try-on images, suggestion means for suggesting additional clothing based on the user's emotion recognition, and communication means for completing the purchase process based on the user's voice instructions. This allows the user to have an online experience similar to trying on actual clothing. Furthermore, personalized suggestions based on emotion analysis are possible, increasing user satisfaction.
[1289] "Optical means" refers to a device for acquiring the user's body shape data, and includes a camera and a sensor.
[1290] "Voice input means" refers to a device that uses a microphone or voice recognition technology to recognize a user's voice instructions and input them into the system.
[1291] "Display means" refers to a display or screen for providing visual information to the user, and displays a clothing selection menu based on the user's voice instructions.
[1292] The "camera means" is a device used to acquire facial expression data of the user, such as a high-resolution camera.
[1293] "Emotion analysis means" refers to software or algorithms that recognize emotions based on a user's facial expression data.
[1294] "Communication means" refers to the network technology and interface used to send information about the clothing selected by the user to the generative AI model and to complete the purchase process.
[1295] The "generative AI means" is a system that includes an artificial intelligence model or algorithm for generating try-on images of the user wearing the selected clothing.
[1296] "Optical display means" refers to a display or screen for displaying the generated try-on image, including the display of smart glasses.
[1297] "Suggestion mechanism" refers to an algorithm or system for providing additional clothing suggestions based on the user's emotion recognition.
[1298] This invention aims to enhance the online and brick-and-mortar shopping experience by combining a system with an emotion analysis engine that recognizes user emotions. The system acquires the user's body shape data and facial expression data, and then simulates trying on clothes and carrying out the purchase process based on that data.
[1299] First, the user puts on the smart glasses and stands in front of the system. The system uses optical means (high-resolution camera) to acquire the user's body shape data and camera means to acquire the user's facial expression data. This allows the system to collect basic data about the user and proceed to the next process.
[1300] Next, the user gives a voice command using the voice input means (microphone). The voice input means recognizes this and displays a clothing selection menu on the display means (the display of the smart glasses) based on the voice command. When the user selects a specific clothing item, the information is sent to the generative AI model via the communication means.
[1301] The generation AI means generates a try-on image of the clothing when worn by the user based on the received body shape data and clothing information. This generated try-on image is displayed in real time on the optical display means (the display of the smart glasses). At this time, the emotion analysis means recognizes emotions based on the user's facial expressions and detects whether the user looks happy or unhappy. Based on the results of the emotion analysis, the system uses the suggestion means to suggest additional clothing items.
[1302] For example, if a user voices the command "I want to look at denim jackets," the voice input means recognizes this and a category menu is displayed. When the user selects a specific denim jacket, that information is sent to a generative AI model, which generates a try-on image based on the user's body type. This try-on image is displayed in real time on the smart glasses display, allowing the user to see the item as if they were actually trying it on. If the user is smiling, the emotion analysis means detects this and the suggestion means suggests additional denim jackets of similar style and color.
[1303] An example of a specific prompt would be, "Generate a fitting image of a denim jacket based on the user's body type data. The user's height is 170 cm, weight is 65 kg, shoulder width is 45 cm, and preferred style is casual." This prompt is sent to the generation AI model, which then generates an appropriate fitting image.
[1304] The system utilizes the Microsoft Azure Emotion API for emotion analysis and the Google Speech-to-Text API for voice command recognition, and uses advanced artificial intelligence technologies such as GPT-4 for generative AI models to respond to users in real time, enabling users to enjoy a smooth and satisfying shopping experience in-store and online.
[1305] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1306] Step 1:
[1307] The user puts on the smart glasses and stands in front of the system. The terminal (smart glasses) acquires the user's body shape data using optical means (high-resolution camera). Specifically, the camera captures the user's full-body image, which is then analyzed using image processing software to acquire data such as height, weight, and shoulder width. The input is the "user's full-body image," and the output is the "user's body shape data."
[1308] Step 2:
[1309] The user issues voice instructions using the voice input means (microphone). The device (smart glasses) captures the user's voice instructions through the voice input means. The voice data is converted into text by voice recognition software (Google Speech-to-Text API). The input is "voice data" and the output is "text data." Based on the voice instructions, the device displays a clothing selection menu on the display means (the display of the smart glasses).
[1310] Step 3:
[1311] The user selects a specific piece of clothing from the displayed menu. The device (smart glasses) recognizes the user's selection using a touch interface or eye tracking. Information about the selected clothing is sent to the server via communication means. The input is "user selection information" and the output is "transmission of clothing information."
[1312] Step 4:
[1313] The server sends the received clothing information and the user's body data to the generative AI model. The generative AI model generates a try-on image of the selected clothing based on the user's body data. Specifically, it uses a prompt to give instructions to the AI model. An example of a prompt is: "Based on the user's body data, please generate an image of a denim jacket being tried on. The user is 170 cm tall, weighs 65 kg, has a shoulder width of 45 cm, and prefers a casual style." The input is "body data and prompt," and the output is "a try-on image."
[1314] Step 5:
[1315] The try-on images generated by the generation AI means are sent to the terminal (smart glasses) via the communication means. The terminal displays the try-on images in real time using the optical display means (the display of the smart glasses). The input is the "try-on image" and the output is the "display of the try-on image."
[1316] Step 6:
[1317] The device (smart glasses) acquires the user's facial expression data using a camera. It analyzes the user's emotions using an emotion analysis tool (Microsoft Azure Emotion API). In this process, the user's facial expression data is input and emotion data is output as the analysis result. The input is "facial expression data" and the output is "emotion data."
[1318] Step 7:
[1319] The server uses the suggestion means to suggest additional clothing based on the emotion data. Specifically, if the user shows a happy emotion, it automatically selects clothing of a similar style and color and suggests it to the user. The input is "emotion data" and the output is "additional clothing suggestions." The terminal displays the suggested additional clothing on its display.
[1320] Step 8:
[1321] The user issues a voice command such as "I would like to purchase this clothing." The device captures this voice command using a voice input means and converts it into text using voice recognition software. This text data is sent to the server via a communication means. The input is "voice data" and the output is "text data of the purchase command."
[1322] Step 9:
[1323] The server receives the purchase instruction and carries out the purchase procedure via the online shop's API. Specifically, it sends the purchase information to the API and processes the payment. The input is "text data of the purchase instruction" and the output is "purchase procedure completed." Once the procedure is complete, a confirmation message is sent to the terminal to notify the user.
[1324] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1325] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1326] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1327] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1328] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1329] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1330] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1331] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1332] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1333] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1334] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1335] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1336] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1337] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1338] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1339] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1340] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1341] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1342] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1343] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1344] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1345] The following is further disclosed regarding the above embodiment.
[1346] (Claim 1)
[1347] a camera means for acquiring body shape data of a user;
[1348] a microphone means for recognizing voice instructions from a user;
[1349] a display means for displaying a clothing selection menu based on a user's voice instruction;
[1350] A communication method to send information about the clothes selected by the user to the generation AI;
[1351] A generation AI means for generating a try-on image of the user wearing the selected clothing;
[1352] a display means for displaying the generated try-on image on a mirror;
[1353] a communication means for carrying out a purchase procedure based on a user's voice instruction;
[1354] A system including:
[1355] (Claim 2)
[1356] 10. The system of claim 1, further comprising an authentication means for authenticating a user.
[1357] (Claim 3)
[1358] The system according to claim 1, further comprising a communication means for acquiring information about clothing through an API of an online shop.
[1359] "Example 1"
[1360] (Claim 1)
[1361] image acquisition means for acquiring body shape data of a user;
[1362] a voice recognition means for recognizing a voice instruction from a user;
[1363] a display means for displaying a product selection menu based on a user's voice instruction;
[1364] A communication means for transmitting information about the product selected by the user to the generation AI model;
[1365] A generation AI means for generating try-on images of the product selected by the user when worn;
[1366] a display means for displaying the generated try-on image on a display device;
[1367] a communication means for carrying out a purchase procedure based on a user's voice instruction;
[1368] A system including:
[1369] (Claim 2)
[1370] 10. The system of claim 1, further comprising an authentication means for authenticating a user.
[1371] (Claim 3)
[1372] 2. The system according to claim 1, further comprising a communication means for acquiring product information through a database of an online shop.
[1373] "Application Example 1"
[1374] (Claim 1)
[1375] A photographing means for acquiring the user's body shape data;
[1376] a voice input means for recognizing voice instructions from a user;
[1377] a display means for displaying a clothing selection menu based on a user's voice instruction;
[1378] A communication means for transmitting information about the clothing selected by the user to the generation AI model;
[1379] A generation AI model means for generating a try-on image of the user wearing the selected clothing;
[1380] a display means for displaying the generated try-on image on a reflective surface;
[1381] a communication means for carrying out a purchase procedure based on a user's voice instruction;
[1382] A recognition device installed in a physical store that recognizes users' faces,
[1383] a database means for providing personalized recommendations based on past purchase history;
[1384] A system including:
[1385] (Claim 2)
[1386] 10. The system of claim 1, further comprising an authentication means for authenticating a user.
[1387] (Claim 3)
[1388] The system according to claim 1, further comprising a communication means for acquiring information about clothing through an API of an online shop.
[1389] "Example 2: Combining Emotion Engines"
[1390] (Claim 1)
[1391] a photographing means for acquiring the user's body shape data and facial expression;
[1392] a voice input means for recognizing voice instructions from a user;
[1393] a display means for displaying a clothing selection menu based on a user's voice instruction;
[1394] An emotion engine that recognizes user emotions;
[1395] A communication means for transmitting information about the clothing selected by the user to the generation AI model;
[1396] A generation AI means for generating try-on images of the user wearing selected clothing based on the user's body type data;
[1397] a display means for displaying the generated try-on image on a mirror;
[1398] an adjusting means for adjusting the visual effect of the try-on image based on the user's emotions;
[1399] a communication means for carrying out a purchase procedure based on a user's voice instruction;
[1400] A system including:
[1401] (Claim 2)
[1402] 10. The system of claim 1, further comprising an authentication means for authenticating a user.
[1403] (Claim 3)
[1404] The system according to claim 1, further comprising a communication means for acquiring information about clothing through an API of an online shop.
[1405] "Application example 2 when combining emotion engines"
[1406] (Claim 1)
[1407] an optical means for acquiring body shape data of a user;
[1408] a voice input means for recognizing voice instructions from a user;
[1409] a display means for displaying a clothing selection menu based on a user's voice instruction;
[1410] a camera means for acquiring facial expression data of a user;
[1411] an emotion analysis means for recognizing the emotion of a user;
[1412] A communication means for transmitting information about the clothing selected by the user to the generation AI model;
[1413] A generation AI means for generating a try-on image of the user wearing the selected clothing;
[1414] an optical display means for displaying the generated try-on image;
[1415] suggestion means for providing additional clothing suggestions based on user emotion recognition;
[1416] a communication means for carrying out a purchase procedure based on a user's voice instruction;
[1417] A system including:
[1418] (Claim 2)
[1419] 10. The system of claim 1, further comprising an authentication means for authenticating a user.
[1420] (Claim 3)
[1421] The system according to claim 1, further comprising a communication means for acquiring information about clothing through an API of an online shop. [Explanation of symbols]
[1422] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a camera means for acquiring body shape data of a user; a microphone means for recognizing voice instructions from a user; a display means for displaying a clothing selection menu based on a user's voice instruction; A communication method to send information about the clothes selected by the user to the generation AI; A generation AI means for generating a try-on image of the user wearing the selected clothing; a display means for displaying the generated try-on image on a mirror; a communication means for carrying out a purchase procedure based on a user's voice instruction; A system including:
2. 10. The system of claim 1, further comprising authentication means for authenticating a user.
3. The system according to claim 1, further comprising a communication means for acquiring information about clothing through an API of an online shop.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A