System
The system addresses challenges in determining color attributes and body type by allowing users to upload images, analyze, suggest styling items, and generate images, facilitating easy purchasing.
Patent Information
- Application Number
- JP2024129321
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-18
AI Technical Summary
Users face challenges in finding color attributes and body type suited to them, difficulty in predicting how styling items will look, and a cumbersome purchasing process for suggested items.
A system that allows users to upload images, analyze color attributes and body type, suggest styling items, generate images incorporating these items, and provide purchase links, while ensuring image quality and using generative AI for visualization.
Enables users to efficiently discover their personal style and easily purchase suitable styling items, improving the online shopping experience.
Smart Images

Figure 2026026900000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Previously, users needed expert advice to find the color attributes and body type that suited them, which was time-consuming and costly. It was also difficult to predict how a styling item would look on the user, requiring them to actually try it on. Furthermore, the purchasing process for the suggested styling items was cumbersome, creating a need for technological solutions to streamline the entire process. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system including a means for users to upload their own images, a means for analyzing the uploaded images and determining the user's color attributes and body type, a means for suggesting styling items suitable for the user based on the results of the determination, a means for generating an image of the user when incorporating the suggested styling items, and a means for displaying the suggestions and the generated image to the user. Furthermore, by providing a means for providing a link to purchase the suggested styling items and a means for checking the quality of the images and selecting images suitable for analysis, the system allows users to efficiently and effectively discover styling that suits them and easily complete the purchasing process.
[0006] "User" refers to a person who uses the system to upload their own image and receive styling suggestions.
[0007] "Image" refers to a photograph or drawing uploaded by a User to the System.
[0008] "Upload" refers to the operation of sending image data from a terminal to a server.
[0009] "Analysis" refers to the process by which the system determines color attributes and skeletal type based on the information in the uploaded image.
[0010] "Color attributes" are characteristics that are classified based on the user's skin tone and color, and are generally divided into cool (blue-based) and warm (yellow-based).
[0011] "Body type" refers to a characteristic that is classified based on the shape and proportions of the user's body, and is categorized as straight, wavy, natural, etc.
[0012] "Styling items" refers to items such as suggested clothing, hair color, makeup products, and colored contact lenses.
[0013] "Suggestion" refers to the process by which the system presents the most suitable styling items to the user based on the results of the analysis.
[0014] "Image image" refers to an image that simulates how the user will look when incorporating the suggested styling item.
[0015] "Display" refers to the process by which the device visually presents analysis results, suggestions, and images to the user.
[0016] "Purchase Link" refers to a link to an online store that guides the user to easily purchase the suggested styling item.
[0017] "Quality check" refers to the process by which the system checks the brightness, resolution, and whether the face is clearly visible in the uploaded image.
[0018] "Server" refers to the central processing unit that receives images uploaded by users, analyzes them, and creates images using generative AI.
[0019] "Terminal" refers to a device used by a user to upload images and receive and display suggestions and images from the system. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (cool or warm skin tones) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like.
[0042] System Program Processing
[0043] 1. Upload a photo
[0044] A user accesses the system's website or application and logs in.
[0045] The user selects their image and clicks the upload button.
[0046] The device sends the photo file to the server.
[0047] 2. Receiving and analyzing photos
[0048] The server receives the uploaded photos.
[0049] The server checks the quality of the received photo to ensure it is suitable for analysis, specifically whether the face is clearly visible and whether the brightness and resolution are sufficient.
[0050] The server inputs photos that pass the quality check into the AI model, which analyzes the user's color attributes and bone structure. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure to determine the user's personal color (cool or warm undertones) and bone structure (straight, wavy, natural).
[0051] 3. Styling suggestion generation
[0052] The server compares the analysis results obtained from the AI model with a database to suggest the most suitable styling items for the user. The suggestions include:
[0053] Clothing color and style
[0054] hair color
[0055] Makeup products
[0056] colored contact lenses
[0057] The server uses these suggestions to generate and store information to provide to the user.
[0058] 4. Image generation
[0059] The server requests the generation AI to generate an image that reflects the proposed styling items.
[0060] The generative AI model creates a new image of the user incorporating each of the suggested elements (clothing, hair color, makeup, colored contact lenses).
[0061] The server receives the generated image and stores it in association with the user data.
[0062] 5. Displaying results and making purchasing suggestions
[0063] The server sends the analysis results, proposals, and generated images to the terminal.
[0064] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0065] Analysis results of the user's personal color and bone structure type
[0066] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0067] Generated image
[0068] Links to purchase recommended items and access to our online store
[0069] The user can check the displayed information and purchase the items they like from the online store.
[0070] Specific examples
[0071] 1. The user uploads their image to the system.
[0072] The device transfers the photo to the server.
[0073] 2. The server uses an AI model to analyze the photo and determine the user's color attribute as "spring warm skin tone" and their bone structure as "wave."
[0074] Based on this, the server will suggest "soft coral" as the best clothing color for the user, "fresh peach-toned blush" as a makeup product, and "warm brown" as colored contact lenses.
[0075] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0076] The terminal displays this to the user.
[0077] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[0078] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0079] In this way, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0080] The processing flow will be explained below.
[0081] Step 1:
[0082] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[0083] Step 2:
[0084] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[0085] Step 3:
[0086] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0087] Step 4:
[0088] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0089] Step 5:
[0090] The server then compares the image analysis results with a database to suggest styling items suitable for the user, such as clothing colors and styles, hair colors, makeup products, and colored contact lenses that match the user's color attributes and body type.
[0091] Step 6:
[0092] The server sends a request to the generative AI model to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0093] Step 7:
[0094] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[0095] Step 8:
[0096] The server sends the analysis results, proposals, and generated images to the terminal.
[0097] Step 9:
[0098] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0099] Analysis results of the user's personal color and bone structure type
[0100] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0101] Generated image
[0102] Links to purchase recommended items and access to our online store
[0103] Step 10:
[0104] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[0105] Example 1
[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0107] In conventional personal styling systems, even if users upload their own images, the system may not accurately determine color attributes or bone structure. It is also difficult for users to visualize how the suggested styling items will look on them, making it difficult to make a purchasing decision. Furthermore, depending on the quality of the image, the analysis results may be inaccurate, making it impossible to provide optimal styling suggestions.
[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0109] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the results of the determination, means for generating an image of the user incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for checking the quality of the images and selecting images suitable for analysis, means for inputting a prompt to a generation AI model for generating an image reflecting the suggested styling items, and means for receiving the generated image and storing it in association with user data. This allows users to accurately grasp their own personal style, easily imagine the suggested styling items, and select and purchase the most suitable products.
[0110] A "user" is a person who accesses the system and receives suggestions for uploading images and styling items.
[0111] "Server" refers to a computer system that receives images uploaded by users, analyzes them, and generates and stores suggestions.
[0112] "Terminal" refers to the device a user uses to access the system, including smartphones and personal computers.
[0113] "Image" refers to a photo file uploaded by a user and used to identify facial features and color attributes.
[0114] "Analysis" refers to the process of determining color attributes and bone structure type based on uploaded images.
[0115] "Color attribute" refers to a personal color determined based on the user's skin tone and hair color, and includes "cool skin" and "warm skin."
[0116] "Body type" is a type classified based on the characteristics of the user's body frame and figure, and includes "straight," "wavy," "natural," and the like.
[0117] "Styling items" are items such as clothes, makeup products, and colored contact lenses that are suggested based on the user's color attributes and body type.
[0118] "Suggestion" refers to the act of providing styling items suitable for the user based on the results of the determination of color attributes and bone structure type.
[0119] An "image image" is a composite image that generates an image of the user wearing the suggested styling item.
[0120] A "generative AI model" is an artificial intelligence model that generates imagery incorporating suggested styling items based on an input prompt.
[0121] A "prompt sentence" is an input sentence to a generative AI model that defines the details of the image to be generated.
[0122] "Quality check" is the process of checking the resolution and brightness of uploaded images and selecting images suitable for analysis.
[0123] "Saving" refers to the act of recording the generated image and proposal content in a database and associating them with the user profile.
[0124] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (blue-based, yellow-based) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like. Specific embodiments are described below.
[0125] First, a user accesses the system's website or application and logs in to upload their own image. After logging in, the user selects their own image and clicks the upload button. At this time, the user's device sends the photo file to the server.
[0126] Next, the server receives the uploaded photo. It performs a quality check on the received photo to ensure it is suitable for analysis. Specifically, it checks whether the face is clearly visible, and whether the brightness and resolution are sufficient. This is done using OpenCV and other image analysis libraries.
[0127] Once the photos pass the quality check, they are fed into an AI model on a server that extracts facial features and analyzes skin tone, hair color, and bone shape to determine the person's personal color (blue-based or yellow-based) and bone type (straight, wavy, natural).
[0128] The server compares the analysis results obtained from the AI model with a database that suggests the best styling items for the user. Suggestions include clothing color and style, hair color, makeup products, colored contact lenses, etc. These suggestions are used to generate and store information to be provided to the user.
[0129] Furthermore, the server requests the generative AI model to generate an image that reflects the suggested styling items. The generative AI model creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, colored contact lenses). An example of a specific prompt is, "Based on this image, please generate an image of styling items (soft coral clothes, fresh peach-toned blush, warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type."
[0130] The generated image is received by the server and stored in association with the user data. Finally, the server sends the analysis results, proposals, and generated image to the terminal, which then displays this information on the user's screen.
[0131] Users can check the displayed information and purchase items they like from the online store. Purchase links are also provided for suggested styling items, making it easy for users to purchase the products.
[0132] As described above, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0134] Step 1: A user visits the system's website or application and logs in.
[0135] Specifically, the user enters their email address and password and clicks the login button, which causes the server to process the user's authentication data and output a login success or failure result.
[0136] Step 2: User selects their image and clicks the upload button.
[0137] Specifically, the user opens the photo library on their device and selects an appropriate image. When the user clicks the upload button, the device inputs the selected image file to be sent to the server. The server receives this request and outputs the image file to a temporary directory.
[0138] Step 3: The server checks the quality of the uploaded photos.
[0139] Specifically, the server uses an image analysis library such as OpenCV to check the image's resolution, brightness, and whether or not it contains a face. The input is the uploaded image file, and image analysis is performed on it. The output is the quality check result (appropriate or inappropriate). If it is determined to be inappropriate, the server notifies the user to upload the image again.
[0140] Step 4: The server inputs photos that pass the quality check into the AI model, which analyzes color attributes and bone structure type.
[0141] Specifically, the server sends the image to the AI model, which extracts facial features. The image must pass a quality check before it is input. The AI model analyzes skin tone, hair color, and bone structure, and outputs a personal color (blue-based or yellow-based) and bone structure type (straight, wavy, or natural).
[0142] Step 5: The server recommends styling items based on the analysis results obtained from the AI model.
[0143] Specifically, the server queries the database for styling items that fit the user's analysis results. The input is the analysis results, which are then compared with the database. The output is styling suggestions such as clothing color and style, hair color, makeup products, and colored contact lenses.
[0144] Step 6: The server generates a prompt sentence for the generated AI model and requests it to generate an image.
[0145] Specifically, the server generates a prompt and sends it to the generative AI model. The input is a styling suggestion, which is converted into text. An example of a prompt is, "Based on this image, please generate an image of styling items (soft coral clothing, fresh peach-toned blush, and warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type." The generative AI model generates an image based on this and sends it back to the server as output.
[0146] Step 7: The server receives the generated image and stores it in association with the user data.
[0147] Specifically, the server receives the image returned from the generative AI model, associates it with the user's profile, and stores it in a database. The input is the image from the generative AI model, and the output is a database update.
[0148] Step 8: The server sends the analysis results, proposals, and generated images to the terminal.
[0149] Specifically, the server compiles the generated information and generates an HTTP response to send to the user's device. The inputs include analysis results, styling suggestions, and images, and these data are sent together as a single response. The output is sent to the device.
[0150] Step 9: The terminal displays the received information on the user screen.
[0151] Specifically, the terminal analyzes the data received from the server and displays it in a format that is easy for the user to view. The input is the data from the server, and the output is displayed on the user interface.
[0152] Step 10: The user reviews the displayed information and purchases the items they like from the online store.
[0153] Specifically, the user clicks on a link to a suggested styling item and is taken to an online store. The input is the user's click, and the output is the browser displaying the online store page. The user then completes the purchase process and orders the product.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] In conventional fashion try-on systems, users had to spend a lot of time and effort finding the perfect styling item for themselves. It was also difficult to receive styling advice without actually trying the items on. Furthermore, the lack of visual information to help users visualize the suggested items made the process of purchasing complicated, hindering the online shopping experience.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, generative AI model means for generating an image of the user's appearance when incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for providing the user with links to purchase items for realizing the suggested styling, means for generating an image using prompt text for the generative AI model, and means for checking the quality of the images and selecting images suitable for analysis. This enables users to efficiently find styling items that are best suited to them and instantly obtain specific images based on visual information, thereby improving the online shopping experience.
[0159] "User" refers to a person who uses this system to upload their own image and receive suggestions for the most suitable styling items.
[0160] "Means for uploading images" refers to the interface or process that allows users to send their own image data to the server.
[0161] "Means for analyzing images to determine a user's color attributes and bone structure type" refers to processes or algorithms that analyze a user's skin tone, hair color, and bone structure based on uploaded images to identify a user's color attributes and bone structure type.
[0162] "Means for suggesting styling items" refers to the process of algorithms and database searches that recommend the most suitable clothes, makeup products, accessories, etc. to users based on the analysis results.
[0163] "Generative AI Model" refers to the artificial intelligence model used to generate an image of what a user would look like when incorporating a suggested styling item.
[0164] The "means for displaying the proposed content and the generated image image" refers to a user interface or display for providing the user with the styling proposal and the generated image image in a visible form.
[0165] "Means for providing links to purchase items" refers to the process of providing a user with links to online stores or shopping lists for purchasing suggested styling items.
[0166] "Means for quality checks and selection of images suitable for analysis" refers to an algorithm that evaluates whether uploaded images are suitable for analysis, checking whether faces are clearly visible and whether the brightness and resolution are sufficient.
[0167] A "prompt" is a sentence or text that provides specific instructions or information to a generative AI model for generating an image.
[0168] A "virtual fitting room" refers to a system or application that allows users to upload their own images, receive suggestions for the most suitable styling items, and virtually try them on.
[0169] This detailed description of the present invention will explain in detail how the components of the system work together to generate and display styling item suggestions and images to the user.
[0170] System Program
[0171] The system allows users to upload their own images, analyzes them, and determines their color attributes and body type. Based on the results, it recommends optimal styling items and generates images incorporating the suggested items. This information is then displayed to the user, and if necessary, a link to purchase the suggested items is provided. It also uses a generative AI model to prompt users when generating specific images.
[0172] What the program does
[0173] 1. Upload a photo
[0174] Users access the system's web application and upload their own photos, which are then sent from devices such as smartphones and PCs to the server.
[0175] 2. Receiving and analyzing photos
[0176] The server checks the quality of the received photos, ensuring that the face is clearly visible and that the brightness and resolution are sufficient. To analyze photos that pass the quality check, an AI model specialized for image analysis is used. An AI model powered by TensorFlow is used to determine the user's color attributes and body type.
[0177] 3. Styling item suggestions
[0178] Based on the results of photo analysis, the system compares the results with a database to suggest the most suitable styling items. These styling items include clothing color and style, makeup, accessories, colored contact lenses, etc. The suggestions are generated on the server side and stored in association with user data.
[0179] 4. Image generation
[0180] The server requests the generative AI model to generate an image that reflects the suggested styling items. Specific prompts are used to provide instructions. For example, a prompt like this might be used: "Upload a photo of the user and determine their color attribute "cool undertones" and bone type "straight." Then, generate an image that incorporates a light blue blouse, natural makeup, and clear blue colored contact lenses."
[0181] 5. Displaying results and making purchasing suggestions
[0182] The server sends the analysis results, recommendations, and generated images to the device, which display them on the user's screen. The user can review the displayed information and purchase the items they like from the online store. A purchase link is provided for each suggested item.
[0183] Specific examples
[0184] When User A uploads a photo of himself / herself and the server receives and analyzes the photo, it determines that User A's color attribute is "cool-toned" and his / her bone structure type is "straight."
[0185] The server suggests the best styling items for user A: a light blue blouse, natural makeup, and clear blue colored contact lenses.
[0186] The server uses the generative AI model to generate an image of user A that reflects the suggested item and sends it to the device.
[0187] User A can check the images generated on their device and purchase items they like from the online store.
[0188] This invention allows users to efficiently find the styling items that are best suited to them and instantly get a concrete image of them, greatly improving the online shopping experience.
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1:
[0191] Users upload their own images.
[0192] A user accesses the system's web application, selects a photo file, and clicks the upload button. The terminal sends the selected photo file to the server. The input data is the user's photo file, and the output is the photo file sent to the server.
[0193] Step 2:
[0194] The server checks the quality of the received photos.
[0195] The server performs a quality check on the received photo before analyzing it. It checks whether the face is clearly visible and whether the brightness and resolution are sufficient. The input data is the user's photo file, and the output is a photo file that has passed the quality check. If the quality check is not passed, the server sends a notification to the user, encouraging them to upload again.
[0196] Step 3:
[0197] The server analyzes the photo and determines the user's color attributes and body type.
[0198] The server inputs photos that pass the quality check into the AI model for image analysis. The AI model extracts the user's facial features and analyzes their skin tone, hair color, and bone structure to determine their color attributes (cool or warm undertones) and bone structure type (straight, wavy, natural). The input data is a photo file that has passed the quality check, and the output is the analysis results of the user's color attributes and bone structure type.
[0199] Step 4:
[0200] The server will suggest styling items.
[0201] The server compares the acquired analysis results with the database and suggests styling items that are best suited to the user. Suggestions include clothing color and style, makeup products, colored contact lenses, etc. The input data is the analysis results of the user's color attributes and body type, and the output is a list of suggested styling items.
[0202] Step 5:
[0203] The server requests the generative AI model to generate an image.
[0204] The server sends a request to the generative AI model using a prompt to generate an image that reflects the suggested styling items. The input data is a list of suggested styling items and the prompt, and the output is the generated image. An example of a prompt is, "Please upload a photo of the user and determine the color attribute 'cool' and bone type 'straight'. Then, generate an image that includes a light blue blouse, natural makeup, and clear blue colored contact lenses."
[0205] Step 6:
[0206] The server displays the generated image to the user.
[0207] The server receives the generated image and stores it in association with the user data. It then transmits the analysis results, suggested styling items, and the generated image to the terminal. The input data is the user data associated with the generated image, and the output is the analysis results, suggested item list, and image displayed to the user.
[0208] Step 7:
[0209] The device will provide a link to purchase the suggested styling item.
[0210] The terminal displays a list of suggested styling items and a link to purchase them to the user. The user can click the link to purchase the suggested items from the online store. The input data is the list of suggested styling items, and the output is a user screen with the purchase link displayed.
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] The system of the present invention allows users to upload their own images, and the system analyzes the images to determine color attributes (cool or warm undertones) and bone structure (straight, wavy, natural), and then suggests the most suitable styling items. It also combines this with an emotion engine that recognizes the user's emotions to provide even more personalized styling suggestions. The system can also generate and display to the user an image of what the suggested styling items would look like when worn. It also provides a link to purchase the suggested styling items.
[0213] System Program Processing
[0214] 1. Upload a photo
[0215] A user accesses the system's website or application and logs in.
[0216] User clicks the image upload button, selects and uploads their own image.
[0217] The terminal transmits the selected image file to the server.
[0218] 2. Receiving and analyzing photos
[0219] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0220] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0221] 3. Emotion Recognition by Emotion Engine
[0222] The server simultaneously uses an emotion engine to analyze emotions from the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[0223] 4. Styling suggestion generation
[0224] The server compares the results of the image analysis and emotion engine with a database to suggest styling items suitable for the user. The suggestions include the following items:
[0225] Clothing color and style
[0226] hair color
[0227] Makeup products
[0228] colored contact lenses
[0229] The server uses these suggestions to generate and store information to provide to the user, which is tailored to take into account feedback on the user's emotional state.
[0230] 5. Image Generation
[0231] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0232] 6. Displaying results and making purchasing suggestions
[0233] The server sends the analysis results, proposals, and generated images to the terminal.
[0234] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0235] Analysis results of the user's personal color and bone structure type
[0236] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0237] Generated image
[0238] Links to purchase recommended items and access to our online store
[0239] Specific examples
[0240] 1. The user uploads their image to the system.
[0241] The device transfers the photo to the server.
[0242] 2. The server uses an AI model and emotion engine to analyze the photo, determine the user's color attributes as "spring warm skin tone" and body type as "wave," and further recognize the emotion of "joy" from the user's facial expression.
[0243] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[0244] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0245] The terminal displays this to the user.
[0246] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[0247] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0248] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[0249] The processing flow will be explained below.
[0250] Step 1:
[0251] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[0252] Step 2:
[0253] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[0254] Step 3:
[0255] The server receives the uploaded images. After receiving them, it checks the quality of the images to ensure that the face is clearly visible and that the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0256] Step 4:
[0257] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0258] Step 5:
[0259] At the same time, the server uses an emotion engine to analyze the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to determine emotions such as "happiness," "sadness," "surprise," and "anger."
[0260] Step 6:
[0261] The server compares the results of image analysis and emotion recognition with a database to suggest the most suitable styling items for the user. Specifically, it selects clothing colors and styles, hair colors, makeup products, colored contact lenses, etc. that correspond to the user's color attributes and body type. The suggestions are adjusted taking into account the user's emotional state.
[0262] Step 7:
[0263] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0264] Step 8:
[0265] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[0266] Step 9:
[0267] The server sends the analysis results, suggestions, and generated images to the device, along with a link to purchase the suggested styling items.
[0268] Step 10:
[0269] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0270] Analysis results of the user's personal color and bone structure type
[0271] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0272] Generated image
[0273] Links to purchase recommended items and access to our online store
[0274] Step 11:
[0275] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[0276] Specific examples
[0277] 1. The user uploads their image to the system.
[0278] The device transfers the photo to the server.
[0279] 2. The server analyzes the photo using an AI model to determine the attributes of "spring warm skin tone" and "wave skin tone." The emotion engine then recognizes "joy" from the facial expressions in the image.
[0280] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[0281] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0282] The terminal displays this to the user.
[0283] 4. The device displays a link to purchase items to help the user achieve the suggested styling.
[0284] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0285] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] Conventional styling suggestion systems can analyze a user's personal color and bone structure to suggest suitable styling items, but they cannot make suggestions that take into account the user's emotional state. As a result, the suggested styling items may not match the user's current emotions, making it difficult to increase user satisfaction. Another issue is that it is difficult for users to specifically imagine the suggested styling, which can reduce their motivation to purchase.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0290] In this invention, the server includes a means for users to upload their own images, a means for analyzing the uploaded images and determining the user's color attributes and body type, a means for analyzing the user's emotions based on the uploaded images, a means for suggesting styling items suitable for the user based on the results of the determination and emotion analysis, a means for generating an image of the user incorporating the suggested styling items, and a means for displaying the suggested content and the generated image to the user. This enables personalized styling suggestions that take into account not only the user's personal color and body type but also their emotional state, thereby increasing user satisfaction. Furthermore, the generated image provides a concrete styling image, thereby increasing the user's desire to purchase.
[0291] "User" refers to a person who uses the system to upload their own image and receive styling suggestions.
[0292] "Image" refers to photographic data uploaded by a user to the system.
[0293] "Color attribute" refers to a personal color (e.g., cool-toned, warm-toned) that is classified based on the user's skin tone and hair color.
[0294] "Body type" refers to a styling type (e.g., straight, wavy, natural) that is classified based on the user's body shape and skeletal characteristics.
[0295] "Emotion" refers to the psychological state (e.g., joy, sadness, surprise, anger) that can be read from the user's facial expression.
[0296] "Styling items" refer to fashion-related items such as clothes, hair color, makeup products, and colored contact lenses that are recommended to users.
[0297] "Image" refers to image data used to generate a new visual for the user that reflects the suggested styling item.
[0298] "Server" refers to the central computing resource that performs the primary data processing and management of the system.
[0299] "Terminal" refers to a device such as a computer or smartphone that a user uses to access the system.
[0300] "Quality check" refers to the inspection process to ensure that uploaded images are suitable for analysis.
[0301] "Analysis" refers to the process of extracting features and recognizing emotions from uploaded images using AI models and emotion engines.
[0302] "Suggestion" refers to the act of providing the user with the most suitable styling items based on the analysis results.
[0303] "Generative AI" refers to artificial intelligence that generates image images that reflect suggested styling items in the user's image.
[0304] A "prompt sentence" refers to an instruction sentence that instructs the generation AI to generate a specific image.
[0305] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and then combines them with an emotion engine to provide personalized styling suggestions. The system can generate and display to the user an image of what the suggested styling items would look like. It also provides a link to purchase the suggested styling items.
[0306] The system uses the following major hardware and software:
[0307] 1. Hardware
[0308] Device: The computer or smartphone that a user uses to access the system.
[0309] Server: The central computing resource that handles the main data processing and management of the system.
[0310] 2. Software
[0311] Website or application: An interface where users can upload images and view analysis results and recommendations
[0312] AI model: a model for image analysis (e.g., a TensorFlow-based model)
[0313] Emotion engine: API for emotion recognition (e.g., Face API by Microsoft)
[0314] Generative AI: A model for generating images that reflect suggested styling items (e.g., DALL-E by OpenAI)
[0315] A specific embodiment of the system will be described below.
[0316] Uploading and receiving photos
[0317] Users access the system's website or application, log in, and then click the image upload button to select and upload their own image. At this time, the device sends the selected image file to the server. The image file is securely transmitted using HTTPS.
[0318] Image quality check and analysis
[0319] The server performs a quality check on the received images to determine whether they are suitable for analysis. For example, it uses OpenCV to check whether the face in the image is clearly visible and whether the resolution and brightness are appropriate. Images that pass the quality check are then analyzed using an AI model (e.g., TensorFlow). The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color attribute (e.g., cool-toned, warm-toned) and bone type (e.g., straight, wavy, natural).
[0320] emotion recognition
[0321] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze the user's emotions from the image. The emotion engine analyzes the user's facial expressions, eye movements, mouth features, etc., and recognizes emotions such as joy, sadness, surprise, and anger. The analysis results and emotion analysis results are then stored in a database.
[0322] Styling suggestions
[0323] The server compares the results of image analysis and emotion analysis with the system's styling database. For example, for a user with "spring warm skin tone," "waves," and the emotion "joy," it would suggest "soft coral" clothing and "fresh peach-toned blush." The server generates suggestions and saves a list of styling items suitable for the user in JSON format.
[0324] Image generation
[0325] The server requests the generation AI (e.g., DALL-E by OpenAI) to generate an image that reflects the suggested styling items. An example of a prompt to be input to the generation AI is "An image of a warm-toned, spring-season woman wearing soft coral clothing." The generation AI model creates a new image for the user that reflects the suggestions. The server retrieves the generated image and saves it by associating it with the user's profile.
[0326] Viewing and purchasing results
[0327] The server sends a dataset containing the analysis results, suggestions, and generated images to the device. The device displays the received information on the user's screen. The user sees:
[0328] Analysis results of the user's personal color and bone structure type
[0329] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0330] Generated image
[0331] Links to purchase recommended items and access to our online store
[0332] Specific examples
[0333] For example, a user uploads an image of themselves to the system. The device sends the uploaded image to the server, which analyzes the image using an AI model and emotion engine. If the analysis results indicate that the user's color attributes are "spring warm skin," their bone structure is "wave," and their emotion is "joy," the server will suggest "soft coral" clothing, "fresh peach-toned blush," and "warm brown" colored contact lenses as the optimal styling for the user. Using generative AI, a new image of the user incorporating these styling items is created and displayed to the user. The user can then purchase their favorite products from the online store.
[0334] In this way, the system of the present invention efficiently discovers the user's personal style, makes styling suggestions that take emotions into consideration, and allows the user to easily purchase specific styling items.
[0335] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0336] Step 1:
[0337] A user accesses a website or application of the system and logs in.
[0338] Input: User login information (email address, password)
[0339] How it works: A user enters their email address and password on the login screen and clicks the login button. The system authenticates the user and starts a login session.
[0340] Output: Successful login notification, user's home screen
[0341] Step 2:
[0342] User clicks the image upload button, selects and uploads their own image.
[0343] Input: User's image file
[0344] How it works: The user clicks the image upload button on the main screen. A file selection dialog opens and the user selects their image file. The device sends the selected image file to the server using an HTTP POST request.
[0345] Output: Notification of image file transfer to server
[0346] Step 3:
[0347] The server receives the uploaded images and performs a quality check.
[0348] Input: Uploaded image file
[0349] How it works: The server receives image files and performs a quality check on the images using a facial recognition algorithm (e.g., OpenCV). Check items include facial clarity, resolution, brightness, etc. Only image files that pass the quality check proceed to the next analysis step.
[0350] Output: Image files that pass the quality check or notification of poor quality
[0351] Step 4:
[0352] The server inputs images that pass the quality check into an AI model (e.g., TensorFlow) and performs image analysis.
[0353] Input: Image files that have passed the quality check
[0354] How it works: The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color profile (cool, warm) and bone type (straight, wavy, natural).
[0355] Output: Analysis results of color attributes and skeletal type
[0356] Step 5:
[0357] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze emotions from the uploaded image and the user's facial expressions.
[0358] Input: Uploaded image file
[0359] Behavior: The emotion engine analyzes the user's facial expressions, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[0360] Output: Emotion analysis results
[0361] Step 6:
[0362] The server suggests styling items suitable for the user based on the results of image analysis and the emotion engine.
[0363] Input: Color attribute and skeletal type analysis results, emotion analysis results
[0364] How it works: The system compares the styling database within the system and selects the styling items that are best suited to the user (e.g., "soft coral" clothing, "fresh peach-toned blush," etc.).
[0365] Output: A list of suggested styling items
[0366] Step 7:
[0367] The server uses a generation AI (e.g., DALL-E) to generate an image that reflects the suggested styling items.
[0368] Input: A list of suggested styling items
[0369] How it works: The user provides the AI with a prompt (e.g., "Image of a warm-toned spring woman wearing a soft coral dress") and requests that it generate an image. The AI then creates a new image for the user that reflects the suggestions.
[0370] Output: Generated image
[0371] Step 8:
[0372] The server sends the analysis results, proposals, and generated images to the terminal.
[0373] Input: Analysis results, proposals, generated images
[0374] How it works: The server sends these datasets in JSON format to the device.
[0375] Output: Notification of completion of dataset transmission to the terminal
[0376] Step 9:
[0377] The terminal displays the received information on the user screen.
[0378] Input: Analysis results, proposals, and generated images sent in JSON format
[0379] How it works: The device analyzes the received data and dynamically renders a user screen using a front-end framework such as Vue.js or React.js. The following information is displayed on the user screen: analysis results, recommendations, generated images, links to purchase recommended items, and links to access the online store.
[0380] Output: Styling suggestions displayed on the user's screen
[0381] Step 10:
[0382] The user clicks on the link for the item of interest and is taken to the specified online store page.
[0383] Input: User clicks
[0384] What it does: When the user clicks, the browser navigates to the product page in the specified online store.
[0385] Output: Online store product page
[0386] (Application example 2)
[0387] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0388] To quickly and accurately respond to today's diversifying customer needs, personalized styling suggestions are required. However, while conventional systems take into account a user's personal color and bone structure, they do not provide personalized styling that takes into account emotional fluctuations. Furthermore, there is a lack of methods for providing real-time suggestions in-store or instantly providing visually easy-to-understand images. The present invention aims to solve these problems and improve the customer experience.
[0389] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload their own image, means for analyzing the uploaded image and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, means for generating an image of the user when incorporating the suggested styling items, means for using an emotion engine that recognizes emotions from the user's facial expression, means for adjusting the suggestions taking into account the emotion recognition results, and means for displaying the suggestions and the generated image to the user. This enables the user to receive personalized styling suggestions in real time according to their emotional state.
[0390] A "user" is an individual or customer who uses the system.
[0391] An "image uploading means" is a method or device by which a user can send their image to the system.
[0392] "Means for analyzing images" refers to a method or device for identifying user information (color attributes and bone structure type) based on uploaded images.
[0393] "Color attribute" indicates the personal color (cool or warm) based on the user's skin tone.
[0394] "Body type" indicates a classification based on the shape of the user's body frame (straight, wavy, natural).
[0395] The "means for suggesting styling items" is a method or device for suggesting clothes, accessories, makeup products, etc. that are suitable for the user.
[0396] The "means for generating an image" is a method or device for creating a new image of the user incorporating the suggested styling item.
[0397] An "emotion engine" is a program or device that recognizes emotions (happiness, sadness, surprise, anger, etc.) from a user's facial expression.
[0398] The "means for adjusting the suggestions taking into account the emotion recognition results" is a method or apparatus for modifying or adjusting styling suggestions based on the emotions recognized by the emotion engine.
[0399] The "means for displaying the generated image" is a method or device for visually presenting new styling suggestions to the user.
[0400] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and uses an emotion engine to recognize emotions from the user's facial expressions. Based on the results of these analyses, the system suggests optimal styling items for the user and generates and displays images using the suggested items. The system aims to visually present the suggestions and the generated images to the user. Links to purchase the suggested styling items are also provided.
[0401] System hardware and software configuration
[0402] This system uses the following hardware and software:
[0403] Hardware:
[0404] Tablet device (e.g. iPad)
[0405] Smart mirrors (e.g., Novera Smart Mirror)
[0406] software:
[0407] AI image analysis engine (e.g. Google Cloud Vision API)
[0408] Emotion recognition engine (e.g. Microsoft Azure Face API)
[0409] Generative AI models (e.g., OpenAI DALL-E)
[0410] Data processing and calculation flow
[0411] Upload a photo
[0412] Users access the system from a tablet or smart mirror and upload their own images, which are then sent to the server.
[0413] Image reception and analysis
[0414] The server receives the uploaded images and performs a quality check. Images that pass the quality check are input into an AI image analysis engine (Google Cloud Vision API) to determine personal color and bone structure type.
[0415] Emotion recognition by emotion engine
[0416] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to analyze the user's facial expressions and recognize emotions such as joy, sadness, surprise, and anger.
[0417] Generate styling suggestions
[0418] The server then uses the results of image analysis and emotion recognition to suggest styling items suitable for the user, including clothing color and style, hair color, makeup products, and colored contact lenses, while also taking into account the user's emotional state.
[0419] Image generation
[0420] The server uses a generative AI model (OpenAI DALL-E) to generate new images for the user incorporating the suggested styling items.
[0421] Displaying results and making purchasing suggestions
[0422] The generated image and the proposed content are sent to the device, which then displays them on the user's screen. The displayed content includes the analysis results of personal color and bone structure type, recommended styling items, the generated image, and a link to purchase the items.
[0423] Specific examples
[0424] Below are some specific examples.
[0425] Example prompt sentence:
[0426] "Using Google Cloud Vision API and Microsoft Azure Face API, we analyze the customer's facial image to determine their personal color. We then analyze their facial expressions to recognize their emotions. Based on the results, we provide fashion styling and product suggestions, and generate new images."
[0427] Specifically, when a user uploads a photo of themselves using a tablet or smart mirror, the server analyzes the photo and determines the user's color attributes as "cool winter" and bone type as "straight," and then uses emotion recognition to identify the emotion of "surprise." Based on this, the server suggests a "sharp black jacket" and "cool-toned makeup products" to the user, and uses a generative AI model to create an image that reflects these suggestions and displays it on the device. The user can also receive a link to purchase the items from an online store based on the suggestions.
[0428] In this way, the system of the present invention allows users to easily receive personalized styling suggestions based on their individual characteristics and emotions, and by visually checking the generated image, the suggestions can be easily understood.
[0429] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0430] Step 1:
[0431] A user uploads their own image using a tablet device or smart mirror. The user opens a dedicated application, clicks the image upload button, and selects their own image file. The input is the user's image file, and the output is that the device sends this image to the server.
[0432] Step 2:
[0433] The server receives the uploaded image. The server performs a quality check on the image to ensure that the face is clearly visible, and that the brightness and resolution are sufficient. Only images that pass this quality check proceed to the next step. The input is the user's image file, and the output is an image that passes the quality check.
[0434] Step 3:
[0435] The server inputs images that pass the quality check into an AI image analysis engine (Google Cloud Vision API) to determine the user's color attributes and bone structure. During this process, the AI analyzes facial features and identifies skin tone, hair color, and bone structure. The input is an image that passes the quality check, and the output is the user's color attributes and bone structure data.
[0436] Step 4:
[0437] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to recognize emotions from the user's facial expressions. This process detects subtle facial features and identifies emotions such as joy, sadness, surprise, and anger. The input is a user image that has passed quality checks, and the output is the user's emotional data.
[0438] Step 5:
[0439] The server compares the results of image analysis and emotion recognition with a database of styling item suggestions. It then suggests the most suitable clothing, hair color, makeup products, and colored contact lenses for the user. The suggestions are adjusted based on the emotion data. The input is the user's color attributes, bone structure, and emotion data, and the output is a list of suggested styling items.
[0440] Step 6:
[0441] The server uses a generative AI model (OpenAI DALL-E) to generate an image that reflects the suggested styling items. In this process, a new image is created by combining styling items with the user's facial photo. The input is a list of suggested styling items and the user's image, and the output is the generated image.
[0442] Step 7:
[0443] The server sends the generated image and proposal details to the terminal, which then displays them on the user's screen. The displayed content includes the analysis results of the personal color and body type, recommended styling items, the generated image, and a link to purchase the items. The input is the generated image and proposal details, and the output is the specific styling proposal information displayed to the user.
[0444] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0445] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0446] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0447] [Second embodiment]
[0448] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0449] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0450] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0451] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0452] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0454] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0455] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0456] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0457] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0458] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0459] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0460] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (cool or warm skin tones) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like.
[0461] System Program Processing
[0462] 1. Upload a photo
[0463] A user accesses the system's website or application and logs in.
[0464] The user selects their image and clicks the upload button.
[0465] The device sends the photo file to the server.
[0466] 2. Receiving and analyzing photos
[0467] The server receives the uploaded photos.
[0468] The server checks the quality of the received photo to ensure it is suitable for analysis, specifically whether the face is clearly visible and whether the brightness and resolution are sufficient.
[0469] The server inputs photos that pass the quality check into the AI model, which analyzes the user's color attributes and bone structure. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure to determine the user's personal color (cool or warm undertones) and bone structure (straight, wavy, natural).
[0470] 3. Styling suggestion generation
[0471] The server compares the analysis results obtained from the AI model with a database to suggest the most suitable styling items for the user. The suggestions include:
[0472] Clothing color and style
[0473] hair color
[0474] Makeup products
[0475] colored contact lenses
[0476] The server uses these suggestions to generate and store information to provide to the user.
[0477] 4. Image generation
[0478] The server requests the generation AI to generate an image that reflects the proposed styling items.
[0479] The generative AI model creates a new image of the user incorporating each of the suggested elements (clothing, hair color, makeup, colored contact lenses).
[0480] The server receives the generated image and stores it in association with the user data.
[0481] 5. Displaying results and making purchasing suggestions
[0482] The server sends the analysis results, proposals, and generated images to the terminal.
[0483] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0484] Analysis results of the user's personal color and bone structure type
[0485] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0486] Generated image
[0487] Links to purchase recommended items and access to our online store
[0488] The user can check the displayed information and purchase the items they like from the online store.
[0489] Specific examples
[0490] 1. The user uploads their image to the system.
[0491] The device transfers the photo to the server.
[0492] 2. The server uses an AI model to analyze the photo and determine the user's color attribute as "spring warm skin tone" and their bone structure as "wave."
[0493] Based on this, the server will suggest "soft coral" as the best clothing color for the user, "fresh peach-toned blush" as a makeup product, and "warm brown" as colored contact lenses.
[0494] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0495] The terminal displays this to the user.
[0496] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[0497] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0498] In this way, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0499] The processing flow will be explained below.
[0500] Step 1:
[0501] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[0502] Step 2:
[0503] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[0504] Step 3:
[0505] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0506] Step 4:
[0507] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0508] Step 5:
[0509] The server then compares the image analysis results with a database to suggest styling items suitable for the user, such as clothing colors and styles, hair colors, makeup products, and colored contact lenses that match the user's color attributes and body type.
[0510] Step 6:
[0511] The server sends a request to the generative AI model to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0512] Step 7:
[0513] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[0514] Step 8:
[0515] The server sends the analysis results, proposals, and generated images to the terminal.
[0516] Step 9:
[0517] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0518] Analysis results of the user's personal color and bone structure type
[0519] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0520] Generated image
[0521] Links to purchase recommended items and access to our online store
[0522] Step 10:
[0523] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[0524] Example 1
[0525] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0526] In conventional personal styling systems, even if users upload their own images, the system may not accurately determine color attributes or bone structure. It is also difficult for users to visualize how the suggested styling items will look on them, making it difficult to make a purchasing decision. Furthermore, depending on the quality of the image, the analysis results may be inaccurate, making it impossible to provide optimal styling suggestions.
[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0528] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the results of the determination, means for generating an image of the user incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for checking the quality of the images and selecting images suitable for analysis, means for inputting a prompt to a generation AI model for generating an image reflecting the suggested styling items, and means for receiving the generated image and storing it in association with user data. This allows users to accurately grasp their own personal style, easily imagine the suggested styling items, and select and purchase the most suitable products.
[0529] A "user" is a person who accesses the system and receives suggestions for uploading images and styling items.
[0530] "Server" refers to a computer system that receives images uploaded by users, analyzes them, and generates and stores suggestions.
[0531] "Terminal" refers to the device a user uses to access the system, including smartphones and personal computers.
[0532] "Image" refers to a photo file uploaded by a user and used to identify facial features and color attributes.
[0533] "Analysis" refers to the process of determining color attributes and bone structure type based on uploaded images.
[0534] "Color attribute" refers to a personal color determined based on the user's skin tone and hair color, and includes "cool skin" and "warm skin."
[0535] "Body type" is a type classified based on the characteristics of the user's body frame and figure, and includes "straight," "wavy," "natural," and the like.
[0536] "Styling items" are items such as clothes, makeup products, and colored contact lenses that are suggested based on the user's color attributes and body type.
[0537] "Suggestion" refers to the act of providing styling items suitable for the user based on the results of the determination of color attributes and bone structure type.
[0538] An "image image" is a composite image that generates an image of the user wearing the suggested styling item.
[0539] A "generative AI model" is an artificial intelligence model that generates imagery incorporating suggested styling items based on an input prompt.
[0540] A "prompt sentence" is an input sentence to a generative AI model that defines the details of the image to be generated.
[0541] "Quality check" is the process of checking the resolution and brightness of uploaded images and selecting images suitable for analysis.
[0542] "Saving" refers to the act of recording the generated image and proposal content in a database and associating them with the user profile.
[0543] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (blue-based, yellow-based) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like. Specific embodiments are described below.
[0544] First, a user accesses the system's website or application and logs in to upload their own image. After logging in, the user selects their own image and clicks the upload button. At this time, the user's device sends the photo file to the server.
[0545] Next, the server receives the uploaded photo. It performs a quality check on the received photo to ensure it is suitable for analysis. Specifically, it checks whether the face is clearly visible, and whether the brightness and resolution are sufficient. This is done using OpenCV and other image analysis libraries.
[0546] Once the photos pass the quality check, they are fed into an AI model on a server that extracts facial features and analyzes skin tone, hair color, and bone shape to determine the person's personal color (blue-based or yellow-based) and bone type (straight, wavy, natural).
[0547] The server compares the analysis results obtained from the AI model with a database that suggests the best styling items for the user. Suggestions include clothing color and style, hair color, makeup products, colored contact lenses, etc. These suggestions are used to generate and store information to be provided to the user.
[0548] Furthermore, the server requests the generative AI model to generate an image that reflects the suggested styling items. The generative AI model creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, colored contact lenses). An example of a specific prompt is, "Based on this image, please generate an image of styling items (soft coral clothes, fresh peach-toned blush, warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type."
[0549] The generated image is received by the server and stored in association with the user data. Finally, the server sends the analysis results, proposals, and generated image to the terminal, which then displays this information on the user's screen.
[0550] Users can check the displayed information and purchase items they like from the online store. Purchase links are also provided for suggested styling items, making it easy for users to purchase the products.
[0551] As described above, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0552] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0553] Step 1: A user visits the system's website or application and logs in.
[0554] Specifically, the user enters their email address and password and clicks the login button, which causes the server to process the user's authentication data and output a login success or failure result.
[0555] Step 2: User selects their image and clicks the upload button.
[0556] Specifically, the user opens the photo library on their device and selects an appropriate image. When the user clicks the upload button, the device inputs the selected image file to be sent to the server. The server receives this request and outputs the image file to a temporary directory.
[0557] Step 3: The server checks the quality of the uploaded photos.
[0558] Specifically, the server uses an image analysis library such as OpenCV to check the image's resolution, brightness, and whether or not it contains a face. The input is the uploaded image file, and image analysis is performed on it. The output is the quality check result (appropriate or inappropriate). If it is determined to be inappropriate, the server notifies the user to upload the image again.
[0559] Step 4: The server inputs photos that pass the quality check into the AI model, which analyzes color attributes and bone structure type.
[0560] Specifically, the server sends the image to the AI model, which extracts facial features. The image must pass a quality check before it is input. The AI model analyzes skin tone, hair color, and bone structure, and outputs a personal color (blue-based or yellow-based) and bone structure type (straight, wavy, or natural).
[0561] Step 5: The server recommends styling items based on the analysis results obtained from the AI model.
[0562] Specifically, the server queries the database for styling items that fit the user's analysis results. The input is the analysis results, which are then compared with the database. The output is styling suggestions such as clothing color and style, hair color, makeup products, and colored contact lenses.
[0563] Step 6: The server generates a prompt sentence for the generated AI model and requests it to generate an image.
[0564] Specifically, the server generates a prompt and sends it to the generative AI model. The input is a styling suggestion, which is converted into text. An example of a prompt is, "Based on this image, please generate an image of styling items (soft coral clothing, fresh peach-toned blush, and warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type." The generative AI model generates an image based on this and sends it back to the server as output.
[0565] Step 7: The server receives the generated image and stores it in association with the user data.
[0566] Specifically, the server receives the image returned from the generative AI model, associates it with the user's profile, and stores it in a database. The input is the image from the generative AI model, and the output is a database update.
[0567] Step 8: The server sends the analysis results, proposals, and generated images to the terminal.
[0568] Specifically, the server compiles the generated information and generates an HTTP response to send to the user's device. The inputs include analysis results, styling suggestions, and images, and these data are sent together as a single response. The output is sent to the device.
[0569] Step 9: The terminal displays the received information on the user screen.
[0570] Specifically, the terminal analyzes the data received from the server and displays it in a format that is easy for the user to view. The input is the data from the server, and the output is displayed on the user interface.
[0571] Step 10: The user reviews the displayed information and purchases the items they like from the online store.
[0572] Specifically, the user clicks on a link to a suggested styling item and is taken to an online store. The input is the user's click, and the output is the browser displaying the online store page. The user then completes the purchase process and orders the product.
[0573] (Application example 1)
[0574] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0575] In conventional fashion try-on systems, users had to spend a lot of time and effort finding the perfect styling item for themselves. It was also difficult to receive styling advice without actually trying the items on. Furthermore, the lack of visual information to help users visualize the suggested items made the process of purchasing complicated, hindering the online shopping experience.
[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0577] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, generative AI model means for generating an image of the user's appearance when incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for providing the user with links to purchase items for realizing the suggested styling, means for generating an image using prompt text for the generative AI model, and means for checking the quality of the images and selecting images suitable for analysis. This enables users to efficiently find styling items that are best suited to them and instantly obtain specific images based on visual information, thereby improving the online shopping experience.
[0578] "User" refers to a person who uses this system to upload their own image and receive suggestions for the most suitable styling items.
[0579] "Means for uploading images" refers to the interface or process that allows users to send their own image data to the server.
[0580] "Means for analyzing images to determine a user's color attributes and bone structure type" refers to processes or algorithms that analyze a user's skin tone, hair color, and bone structure based on uploaded images to identify a user's color attributes and bone structure type.
[0581] "Means for suggesting styling items" refers to the process of algorithms and database searches that recommend the most suitable clothes, makeup products, accessories, etc. to users based on the analysis results.
[0582] "Generative AI Model" refers to the artificial intelligence model used to generate an image of what a user would look like when incorporating a suggested styling item.
[0583] The "means for displaying the proposed content and the generated image image" refers to a user interface or display for providing the user with the styling proposal and the generated image image in a visible form.
[0584] "Means for providing links to purchase items" refers to the process of providing a user with links to online stores or shopping lists for purchasing suggested styling items.
[0585] "Means for quality checks and selection of images suitable for analysis" refers to an algorithm that evaluates whether uploaded images are suitable for analysis, checking whether faces are clearly visible and whether the brightness and resolution are sufficient.
[0586] A "prompt" is a sentence or text that provides specific instructions or information to a generative AI model for generating an image.
[0587] A "virtual fitting room" refers to a system or application that allows users to upload their own images, receive suggestions for the most suitable styling items, and virtually try them on.
[0588] This detailed description of the present invention will explain in detail how the components of the system work together to generate and display styling item suggestions and images to the user.
[0589] System Program
[0590] The system allows users to upload their own images, analyzes them, and determines their color attributes and body type. Based on the results, it recommends optimal styling items and generates images incorporating the suggested items. This information is then displayed to the user, and if necessary, a link to purchase the suggested items is provided. It also uses a generative AI model to prompt users when generating specific images.
[0591] What the program does
[0592] 1. Upload a photo
[0593] Users access the system's web application and upload their own photos, which are then sent from devices such as smartphones and PCs to the server.
[0594] 2. Receiving and analyzing photos
[0595] The server checks the quality of the received photos, ensuring that the face is clearly visible and that the brightness and resolution are sufficient. To analyze photos that pass the quality check, an AI model specialized for image analysis is used. An AI model powered by TensorFlow is used to determine the user's color attributes and body type.
[0596] 3. Styling item suggestions
[0597] Based on the results of photo analysis, the system compares the results with a database to suggest the most suitable styling items. These styling items include clothing color and style, makeup, accessories, colored contact lenses, etc. The suggestions are generated on the server side and stored in association with user data.
[0598] 4. Image generation
[0599] The server requests the generative AI model to generate an image that reflects the suggested styling items. Specific prompts are used to provide instructions. For example, a prompt like this might be used: "Upload a photo of the user and determine their color attribute "cool undertones" and bone type "straight." Then, generate an image that incorporates a light blue blouse, natural makeup, and clear blue colored contact lenses."
[0600] 5. Displaying results and making purchasing suggestions
[0601] The server sends the analysis results, recommendations, and generated images to the device, which display them on the user's screen. The user can review the displayed information and purchase the items they like from the online store. A purchase link is provided for each suggested item.
[0602] Specific examples
[0603] When User A uploads a photo of himself / herself and the server receives and analyzes the photo, it determines that User A's color attribute is "cool-toned" and his / her bone structure type is "straight."
[0604] The server suggests the best styling items for user A: a light blue blouse, natural makeup, and clear blue colored contact lenses.
[0605] The server uses the generative AI model to generate an image of user A that reflects the suggested item and sends it to the device.
[0606] User A can check the images generated on their device and purchase items they like from the online store.
[0607] This invention allows users to efficiently find the styling items that are best suited to them and instantly get a concrete image of them, greatly improving the online shopping experience.
[0608] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0609] Step 1:
[0610] Users upload their own images.
[0611] A user accesses the system's web application, selects a photo file, and clicks the upload button. The terminal sends the selected photo file to the server. The input data is the user's photo file, and the output is the photo file sent to the server.
[0612] Step 2:
[0613] The server checks the quality of the received photos.
[0614] The server performs a quality check on the received photo before analyzing it. It checks whether the face is clearly visible and whether the brightness and resolution are sufficient. The input data is the user's photo file, and the output is a photo file that has passed the quality check. If the quality check is not passed, the server sends a notification to the user, encouraging them to upload again.
[0615] Step 3:
[0616] The server analyzes the photo and determines the user's color attributes and body type.
[0617] The server inputs photos that pass the quality check into the AI model for image analysis. The AI model extracts the user's facial features and analyzes their skin tone, hair color, and bone structure to determine their color attributes (cool or warm undertones) and bone structure type (straight, wavy, natural). The input data is a photo file that has passed the quality check, and the output is the analysis results of the user's color attributes and bone structure type.
[0618] Step 4:
[0619] The server will suggest styling items.
[0620] The server compares the acquired analysis results with the database and suggests styling items that are best suited to the user. Suggestions include clothing color and style, makeup products, colored contact lenses, etc. The input data is the analysis results of the user's color attributes and body type, and the output is a list of suggested styling items.
[0621] Step 5:
[0622] The server requests the generative AI model to generate an image.
[0623] The server sends a request to the generative AI model using a prompt to generate an image that reflects the suggested styling items. The input data is a list of suggested styling items and the prompt, and the output is the generated image. An example of a prompt is, "Please upload a photo of the user and determine the color attribute 'cool' and bone type 'straight'. Then, generate an image that includes a light blue blouse, natural makeup, and clear blue colored contact lenses."
[0624] Step 6:
[0625] The server displays the generated image to the user.
[0626] The server receives the generated image and stores it in association with the user data. It then transmits the analysis results, suggested styling items, and the generated image to the terminal. The input data is the user data associated with the generated image, and the output is the analysis results, suggested item list, and image displayed to the user.
[0627] Step 7:
[0628] The device will provide a link to purchase the suggested styling item.
[0629] The terminal displays a list of suggested styling items and a link to purchase them to the user. The user can click the link to purchase the suggested items from the online store. The input data is the list of suggested styling items, and the output is a user screen with the purchase link displayed.
[0630] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0631] The system of the present invention allows users to upload their own images, and the system analyzes the images to determine color attributes (cool or warm undertones) and bone structure (straight, wavy, natural), and then suggests the most suitable styling items. It also combines this with an emotion engine that recognizes the user's emotions to provide even more personalized styling suggestions. The system can also generate and display to the user an image of what the suggested styling items would look like when worn. It also provides a link to purchase the suggested styling items.
[0632] System Program Processing
[0633] 1. Upload a photo
[0634] A user accesses the system's website or application and logs in.
[0635] User clicks the image upload button, selects and uploads their own image.
[0636] The terminal transmits the selected image file to the server.
[0637] 2. Receiving and analyzing photos
[0638] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0639] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0640] 3. Emotion Recognition by Emotion Engine
[0641] The server simultaneously uses an emotion engine to analyze emotions from the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[0642] 4. Styling suggestion generation
[0643] The server compares the results of the image analysis and emotion engine with a database to suggest styling items suitable for the user. The suggestions include the following items:
[0644] Clothing color and style
[0645] hair color
[0646] Makeup products
[0647] colored contact lenses
[0648] The server uses these suggestions to generate and store information to provide to the user, which is tailored to take into account feedback on the user's emotional state.
[0649] 5. Image Generation
[0650] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0651] 6. Displaying results and making purchasing suggestions
[0652] The server sends the analysis results, proposals, and generated images to the terminal.
[0653] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0654] Analysis results of the user's personal color and bone structure type
[0655] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0656] Generated image
[0657] Links to purchase recommended items and access to our online store
[0658] Specific examples
[0659] 1. The user uploads their image to the system.
[0660] The device transfers the photo to the server.
[0661] 2. The server uses an AI model and emotion engine to analyze the photo, determine the user's color attributes as "spring warm skin tone" and body type as "wave," and further recognize the emotion of "joy" from the user's facial expression.
[0662] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[0663] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0664] The terminal displays this to the user.
[0665] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[0666] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0667] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[0668] The processing flow will be explained below.
[0669] Step 1:
[0670] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[0671] Step 2:
[0672] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[0673] Step 3:
[0674] The server receives the uploaded images. After receiving them, it checks the quality of the images to ensure that the face is clearly visible and that the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0675] Step 4:
[0676] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0677] Step 5:
[0678] At the same time, the server uses an emotion engine to analyze the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to determine emotions such as "happiness," "sadness," "surprise," and "anger."
[0679] Step 6:
[0680] The server compares the results of image analysis and emotion recognition with a database to suggest the most suitable styling items for the user. Specifically, it selects clothing colors and styles, hair colors, makeup products, colored contact lenses, etc. that correspond to the user's color attributes and body type. The suggestions are adjusted taking into account the user's emotional state.
[0681] Step 7:
[0682] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0683] Step 8:
[0684] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[0685] Step 9:
[0686] The server sends the analysis results, suggestions, and generated images to the device, along with a link to purchase the suggested styling items.
[0687] Step 10:
[0688] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0689] Analysis results of the user's personal color and bone structure type
[0690] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0691] Generated image
[0692] Links to purchase recommended items and access to our online store
[0693] Step 11:
[0694] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[0695] Specific examples
[0696] 1. The user uploads their image to the system.
[0697] The device transfers the photo to the server.
[0698] 2. The server analyzes the photo using an AI model to determine the attributes of "spring warm skin tone" and "wave skin tone." The emotion engine then recognizes "joy" from the facial expressions in the image.
[0699] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[0700] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0701] The terminal displays this to the user.
[0702] 4. The device displays a link to purchase items to help the user achieve the suggested styling.
[0703] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0704] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[0705] Example 2
[0706] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0707] Conventional styling suggestion systems can analyze a user's personal color and bone structure to suggest suitable styling items, but they cannot make suggestions that take into account the user's emotional state. As a result, the suggested styling items may not match the user's current emotions, making it difficult to increase user satisfaction. Another issue is that it is difficult for users to specifically imagine the suggested styling, which can reduce their motivation to purchase.
[0708] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0709] In this invention, the server includes a means for users to upload their own images, a means for analyzing the uploaded images and determining the user's color attributes and body type, a means for analyzing the user's emotions based on the uploaded images, a means for suggesting styling items suitable for the user based on the results of the determination and emotion analysis, a means for generating an image of the user incorporating the suggested styling items, and a means for displaying the suggested content and the generated image to the user. This enables personalized styling suggestions that take into account not only the user's personal color and body type but also their emotional state, thereby increasing user satisfaction. Furthermore, the generated image provides a concrete styling image, thereby increasing the user's desire to purchase.
[0710] "User" refers to a person who uses the system to upload their own image and receive styling suggestions.
[0711] "Image" refers to photographic data uploaded by a user to the system.
[0712] "Color attribute" refers to a personal color (e.g., cool-toned, warm-toned) that is classified based on the user's skin tone and hair color.
[0713] "Body type" refers to a styling type (e.g., straight, wavy, natural) that is classified based on the user's body shape and skeletal characteristics.
[0714] "Emotion" refers to the psychological state (e.g., joy, sadness, surprise, anger) that can be read from the user's facial expression.
[0715] "Styling items" refer to fashion-related items such as clothes, hair color, makeup products, and colored contact lenses that are recommended to users.
[0716] "Image" refers to image data used to generate a new visual for the user that reflects the suggested styling item.
[0717] "Server" refers to the central computing resource that performs the primary data processing and management of the system.
[0718] "Terminal" refers to a device such as a computer or smartphone that a user uses to access the system.
[0719] "Quality check" refers to the inspection process to ensure that uploaded images are suitable for analysis.
[0720] "Analysis" refers to the process of extracting features and recognizing emotions from uploaded images using AI models and emotion engines.
[0721] "Suggestion" refers to the act of providing the user with the most suitable styling items based on the analysis results.
[0722] "Generative AI" refers to artificial intelligence that generates image images that reflect suggested styling items in the user's image.
[0723] A "prompt sentence" refers to an instruction sentence that instructs the generation AI to generate a specific image.
[0724] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and then combines them with an emotion engine to provide personalized styling suggestions. The system can generate and display to the user an image of what the suggested styling items would look like. It also provides a link to purchase the suggested styling items.
[0725] The system uses the following major hardware and software:
[0726] 1. Hardware
[0727] Device: The computer or smartphone that a user uses to access the system.
[0728] Server: The central computing resource that handles the main data processing and management of the system.
[0729] 2. Software
[0730] Website or application: An interface where users can upload images and view analysis results and recommendations
[0731] AI model: a model for image analysis (e.g., a TensorFlow-based model)
[0732] Emotion engine: API for emotion recognition (e.g., Face API by Microsoft)
[0733] Generative AI: A model for generating images that reflect suggested styling items (e.g., DALL-E by OpenAI)
[0734] A specific embodiment of the system will be described below.
[0735] Uploading and receiving photos
[0736] Users access the system's website or application, log in, and then click the image upload button to select and upload their own image. At this time, the device sends the selected image file to the server. The image file is securely transmitted using HTTPS.
[0737] Image quality check and analysis
[0738] The server performs a quality check on the received images to determine whether they are suitable for analysis. For example, it uses OpenCV to check whether the face in the image is clearly visible and whether the resolution and brightness are appropriate. Images that pass the quality check are then analyzed using an AI model (e.g., TensorFlow). The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color attribute (e.g., cool-toned, warm-toned) and bone type (e.g., straight, wavy, natural).
[0739] emotion recognition
[0740] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze the user's emotions from the image. The emotion engine analyzes the user's facial expressions, eye movements, mouth features, etc., and recognizes emotions such as joy, sadness, surprise, and anger. The analysis results and emotion analysis results are then stored in a database.
[0741] Styling suggestions
[0742] The server compares the results of image analysis and emotion analysis with the system's styling database. For example, for a user with "spring warm skin tone," "waves," and the emotion "joy," it would suggest "soft coral" clothing and "fresh peach-toned blush." The server generates suggestions and saves a list of styling items suitable for the user in JSON format.
[0743] Image generation
[0744] The server requests the generation AI (e.g., DALL-E by OpenAI) to generate an image that reflects the suggested styling items. An example of a prompt to be input to the generation AI is "An image of a warm-toned, spring-season woman wearing soft coral clothing." The generation AI model creates a new image for the user that reflects the suggestions. The server retrieves the generated image and saves it by associating it with the user's profile.
[0745] Viewing and purchasing results
[0746] The server sends a dataset containing the analysis results, suggestions, and generated images to the device. The device displays the received information on the user's screen. The user sees:
[0747] Analysis results of the user's personal color and bone structure type
[0748] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0749] Generated image
[0750] Links to purchase recommended items and access to our online store
[0751] Specific examples
[0752] For example, a user uploads an image of themselves to the system. The device sends the uploaded image to the server, which analyzes the image using an AI model and emotion engine. If the analysis results indicate that the user's color attributes are "spring warm skin," their bone structure is "wave," and their emotion is "joy," the server will suggest "soft coral" clothing, "fresh peach-toned blush," and "warm brown" colored contact lenses as the optimal styling for the user. Using generative AI, a new image of the user incorporating these styling items is created and displayed to the user. The user can then purchase their favorite products from the online store.
[0753] In this way, the system of the present invention efficiently discovers the user's personal style, makes styling suggestions that take emotions into consideration, and allows the user to easily purchase specific styling items.
[0754] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0755] Step 1:
[0756] A user accesses a website or application of the system and logs in.
[0757] Input: User login information (email address, password)
[0758] How it works: A user enters their email address and password on the login screen and clicks the login button. The system authenticates the user and starts a login session.
[0759] Output: Successful login notification, user's home screen
[0760] Step 2:
[0761] User clicks the image upload button, selects and uploads their own image.
[0762] Input: User's image file
[0763] How it works: The user clicks the image upload button on the main screen. A file selection dialog opens and the user selects their image file. The device sends the selected image file to the server using an HTTP POST request.
[0764] Output: Notification of image file transfer to server
[0765] Step 3:
[0766] The server receives the uploaded images and performs a quality check.
[0767] Input: Uploaded image file
[0768] How it works: The server receives image files and performs a quality check on the images using a facial recognition algorithm (e.g., OpenCV). Check items include facial clarity, resolution, brightness, etc. Only image files that pass the quality check proceed to the next analysis step.
[0769] Output: Image files that pass the quality check or notification of poor quality
[0770] Step 4:
[0771] The server inputs images that pass the quality check into an AI model (e.g., TensorFlow) and performs image analysis.
[0772] Input: Image files that have passed the quality check
[0773] How it works: The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color profile (cool, warm) and bone type (straight, wavy, natural).
[0774] Output: Analysis results of color attributes and skeletal type
[0775] Step 5:
[0776] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze emotions from the uploaded image and the user's facial expressions.
[0777] Input: Uploaded image file
[0778] Behavior: The emotion engine analyzes the user's facial expressions, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[0779] Output: Emotion analysis results
[0780] Step 6:
[0781] The server suggests styling items suitable for the user based on the results of image analysis and the emotion engine.
[0782] Input: Color attribute and skeletal type analysis results, emotion analysis results
[0783] How it works: The system compares the styling database within the system and selects the styling items that are best suited to the user (e.g., "soft coral" clothing, "fresh peach-toned blush," etc.).
[0784] Output: A list of suggested styling items
[0785] Step 7:
[0786] The server uses a generation AI (e.g., DALL-E) to generate an image that reflects the suggested styling items.
[0787] Input: A list of suggested styling items
[0788] How it works: The user provides the AI with a prompt (e.g., "Image of a warm-toned spring woman wearing a soft coral dress") and requests that it generate an image. The AI then creates a new image for the user that reflects the suggestions.
[0789] Output: Generated image
[0790] Step 8:
[0791] The server sends the analysis results, proposals, and generated images to the terminal.
[0792] Input: Analysis results, proposals, generated images
[0793] How it works: The server sends these datasets in JSON format to the device.
[0794] Output: Notification of completion of dataset transmission to the terminal
[0795] Step 9:
[0796] The terminal displays the received information on the user screen.
[0797] Input: Analysis results, proposals, and generated images sent in JSON format
[0798] How it works: The device analyzes the received data and dynamically renders a user screen using a front-end framework such as Vue.js or React.js. The following information is displayed on the user screen: analysis results, recommendations, generated images, links to purchase recommended items, and links to access the online store.
[0799] Output: Styling suggestions displayed on the user's screen
[0800] Step 10:
[0801] The user clicks on the link for the item of interest and is taken to the specified online store page.
[0802] Input: User clicks
[0803] What it does: When the user clicks, the browser navigates to the product page in the specified online store.
[0804] Output: Online store product page
[0805] (Application example 2)
[0806] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0807] To quickly and accurately respond to today's diversifying customer needs, personalized styling suggestions are required. However, while conventional systems take into account a user's personal color and bone structure, they do not provide personalized styling that takes into account emotional fluctuations. Furthermore, there is a lack of methods for providing real-time suggestions in-store or instantly providing visually easy-to-understand images. The present invention aims to solve these problems and improve the customer experience.
[0808] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload their own image, means for analyzing the uploaded image and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, means for generating an image of the user when incorporating the suggested styling items, means for using an emotion engine that recognizes emotions from the user's facial expression, means for adjusting the suggestions taking into account the emotion recognition results, and means for displaying the suggestions and the generated image to the user. This enables the user to receive personalized styling suggestions in real time according to their emotional state.
[0809] A "user" is an individual or customer who uses the system.
[0810] An "image uploading means" is a method or device by which a user can send their image to the system.
[0811] "Means for analyzing images" refers to a method or device for identifying user information (color attributes and bone structure type) based on uploaded images.
[0812] "Color attribute" indicates the personal color (cool or warm) based on the user's skin tone.
[0813] "Body type" indicates a classification based on the shape of the user's body frame (straight, wavy, natural).
[0814] The "means for suggesting styling items" is a method or device for suggesting clothes, accessories, makeup products, etc. that are suitable for the user.
[0815] The "means for generating an image" is a method or device for creating a new image of the user incorporating the suggested styling item.
[0816] An "emotion engine" is a program or device that recognizes emotions (happiness, sadness, surprise, anger, etc.) from a user's facial expression.
[0817] The "means for adjusting the suggestions taking into account the emotion recognition results" is a method or apparatus for modifying or adjusting styling suggestions based on the emotions recognized by the emotion engine.
[0818] The "means for displaying the generated image" is a method or device for visually presenting new styling suggestions to the user.
[0819] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and uses an emotion engine to recognize emotions from the user's facial expressions. Based on the results of these analyses, the system suggests optimal styling items for the user and generates and displays images using the suggested items. The system aims to visually present the suggestions and the generated images to the user. Links to purchase the suggested styling items are also provided.
[0820] System hardware and software configuration
[0821] This system uses the following hardware and software:
[0822] Hardware:
[0823] Tablet device (e.g. iPad)
[0824] Smart mirrors (e.g., Novera Smart Mirror)
[0825] software:
[0826] AI image analysis engine (e.g. Google Cloud Vision API)
[0827] Emotion recognition engine (e.g. Microsoft Azure Face API)
[0828] Generative AI models (e.g., OpenAI DALL-E)
[0829] Data processing and calculation flow
[0830] Upload a photo
[0831] Users access the system from a tablet or smart mirror and upload their own images, which are then sent to the server.
[0832] Image reception and analysis
[0833] The server receives the uploaded images and performs a quality check. Images that pass the quality check are input into an AI image analysis engine (Google Cloud Vision API) to determine personal color and bone structure type.
[0834] Emotion recognition by emotion engine
[0835] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to analyze the user's facial expressions and recognize emotions such as joy, sadness, surprise, and anger.
[0836] Generate styling suggestions
[0837] The server then uses the results of image analysis and emotion recognition to suggest styling items suitable for the user, including clothing color and style, hair color, makeup products, and colored contact lenses, while also taking into account the user's emotional state.
[0838] Image generation
[0839] The server uses a generative AI model (OpenAI DALL-E) to generate new images for the user incorporating the suggested styling items.
[0840] Displaying results and making purchasing suggestions
[0841] The generated image and the proposed content are sent to the device, which then displays them on the user's screen. The displayed content includes the analysis results of personal color and bone structure type, recommended styling items, the generated image, and a link to purchase the items.
[0842] Specific examples
[0843] Below are some specific examples.
[0844] Example prompt sentence:
[0845] "Using Google Cloud Vision API and Microsoft Azure Face API, we analyze the customer's facial image to determine their personal color. We then analyze their facial expressions to recognize their emotions. Based on the results, we provide fashion styling and product suggestions, and generate new images."
[0846] Specifically, when a user uploads a photo of themselves using a tablet or smart mirror, the server analyzes the photo and determines the user's color attributes as "cool winter" and bone type as "straight," and then uses emotion recognition to identify the emotion of "surprise." Based on this, the server suggests a "sharp black jacket" and "cool-toned makeup products" to the user, and uses a generative AI model to create an image that reflects these suggestions and displays it on the device. The user can also receive a link to purchase the items from an online store based on the suggestions.
[0847] In this way, the system of the present invention allows users to easily receive personalized styling suggestions based on their individual characteristics and emotions, and by visually checking the generated image, the suggestions can be easily understood.
[0848] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0849] Step 1:
[0850] A user uploads their own image using a tablet device or smart mirror. The user opens a dedicated application, clicks the image upload button, and selects their own image file. The input is the user's image file, and the output is that the device sends this image to the server.
[0851] Step 2:
[0852] The server receives the uploaded image. The server performs a quality check on the image to ensure that the face is clearly visible, and that the brightness and resolution are sufficient. Only images that pass this quality check proceed to the next step. The input is the user's image file, and the output is an image that passes the quality check.
[0853] Step 3:
[0854] The server inputs images that pass the quality check into an AI image analysis engine (Google Cloud Vision API) to determine the user's color attributes and bone structure. During this process, the AI analyzes facial features and identifies skin tone, hair color, and bone structure. The input is an image that passes the quality check, and the output is the user's color attributes and bone structure data.
[0855] Step 4:
[0856] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to recognize emotions from the user's facial expressions. This process detects subtle facial features and identifies emotions such as joy, sadness, surprise, and anger. The input is a user image that has passed quality checks, and the output is the user's emotional data.
[0857] Step 5:
[0858] The server compares the results of image analysis and emotion recognition with a database of styling item suggestions. It then suggests the most suitable clothing, hair color, makeup products, and colored contact lenses for the user. The suggestions are adjusted based on the emotion data. The input is the user's color attributes, bone structure, and emotion data, and the output is a list of suggested styling items.
[0859] Step 6:
[0860] The server uses a generative AI model (OpenAI DALL-E) to generate an image that reflects the suggested styling items. In this process, a new image is created by combining styling items with the user's facial photo. The input is a list of suggested styling items and the user's image, and the output is the generated image.
[0861] Step 7:
[0862] The server sends the generated image and proposal details to the terminal, which then displays them on the user's screen. The displayed content includes the analysis results of the personal color and body type, recommended styling items, the generated image, and a link to purchase the items. The input is the generated image and proposal details, and the output is the specific styling proposal information displayed to the user.
[0863] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0864] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0865] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0866] [Third embodiment]
[0867] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0868] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0869] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0870] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0871] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0872] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0873] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0874] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0875] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0876] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0877] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0878] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0879] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (cool or warm skin tones) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like.
[0880] System Program Processing
[0881] 1. Upload a photo
[0882] A user accesses the system's website or application and logs in.
[0883] The user selects their image and clicks the upload button.
[0884] The device sends the photo file to the server.
[0885] 2. Receiving and analyzing photos
[0886] The server receives the uploaded photos.
[0887] The server checks the quality of the received photo to ensure it is suitable for analysis, specifically whether the face is clearly visible and whether the brightness and resolution are sufficient.
[0888] The server inputs photos that pass the quality check into the AI model, which analyzes the user's color attributes and bone structure. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure to determine the user's personal color (cool or warm undertones) and bone structure (straight, wavy, natural).
[0889] 3. Styling suggestion generation
[0890] The server compares the analysis results obtained from the AI model with a database to suggest the most suitable styling items for the user. The suggestions include:
[0891] Clothing color and style
[0892] hair color
[0893] Makeup products
[0894] colored contact lenses
[0895] The server uses these suggestions to generate and store information to provide to the user.
[0896] 4. Image generation
[0897] The server requests the generation AI to generate an image that reflects the proposed styling items.
[0898] The generative AI model creates a new image of the user incorporating each of the suggested elements (clothing, hair color, makeup, colored contact lenses).
[0899] The server receives the generated image and stores it in association with the user data.
[0900] 5. Displaying results and making purchasing suggestions
[0901] The server sends the analysis results, proposals, and generated images to the terminal.
[0902] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0903] Analysis results of the user's personal color and bone structure type
[0904] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0905] Generated image
[0906] Links to purchase recommended items and access to our online store
[0907] The user can check the displayed information and purchase the items they like from the online store.
[0908] Specific examples
[0909] 1. The user uploads their image to the system.
[0910] The device transfers the photo to the server.
[0911] 2. The server uses an AI model to analyze the photo and determine the user's color attribute as "spring warm skin tone" and their bone structure as "wave."
[0912] Based on this, the server will suggest "soft coral" as the best clothing color for the user, "fresh peach-toned blush" as a makeup product, and "warm brown" as colored contact lenses.
[0913] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[0914] The terminal displays this to the user.
[0915] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[0916] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[0917] In this way, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0918] The processing flow will be explained below.
[0919] Step 1:
[0920] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[0921] Step 2:
[0922] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[0923] Step 3:
[0924] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[0925] Step 4:
[0926] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[0927] Step 5:
[0928] The server then compares the image analysis results with a database to suggest styling items suitable for the user, such as clothing colors and styles, hair colors, makeup products, and colored contact lenses that match the user's color attributes and body type.
[0929] Step 6:
[0930] The server sends a request to the generative AI model to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[0931] Step 7:
[0932] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[0933] Step 8:
[0934] The server sends the analysis results, proposals, and generated images to the terminal.
[0935] Step 9:
[0936] The terminal displays the received information on the user's screen. The displayed content is as follows:
[0937] Analysis results of the user's personal color and bone structure type
[0938] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[0939] Generated image
[0940] Links to purchase recommended items and access to our online store
[0941] Step 10:
[0942] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[0943] Example 1
[0944] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0945] In conventional personal styling systems, even if users upload their own images, the system may not accurately determine color attributes or bone structure. It is also difficult for users to visualize how the suggested styling items will look on them, making it difficult to make a purchasing decision. Furthermore, depending on the quality of the image, the analysis results may be inaccurate, making it impossible to provide optimal styling suggestions.
[0946] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0947] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the results of the determination, means for generating an image of the user incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for checking the quality of the images and selecting images suitable for analysis, means for inputting a prompt to a generation AI model for generating an image reflecting the suggested styling items, and means for receiving the generated image and storing it in association with user data. This allows users to accurately grasp their own personal style, easily imagine the suggested styling items, and select and purchase the most suitable products.
[0948] A "user" is a person who accesses the system and receives suggestions for uploading images and styling items.
[0949] "Server" refers to a computer system that receives images uploaded by users, analyzes them, and generates and stores suggestions.
[0950] "Terminal" refers to the device a user uses to access the system, including smartphones and personal computers.
[0951] "Image" refers to a photo file uploaded by a user and used to identify facial features and color attributes.
[0952] "Analysis" refers to the process of determining color attributes and bone structure type based on uploaded images.
[0953] "Color attribute" refers to a personal color determined based on the user's skin tone and hair color, and includes "cool skin" and "warm skin."
[0954] "Body type" is a type classified based on the characteristics of the user's body frame and figure, and includes "straight," "wavy," "natural," and the like.
[0955] "Styling items" are items such as clothes, makeup products, and colored contact lenses that are suggested based on the user's color attributes and body type.
[0956] "Suggestion" refers to the act of providing styling items suitable for the user based on the results of the determination of color attributes and bone structure type.
[0957] An "image image" is a composite image that generates an image of the user wearing the suggested styling item.
[0958] A "generative AI model" is an artificial intelligence model that generates imagery incorporating suggested styling items based on an input prompt.
[0959] A "prompt sentence" is an input sentence to a generative AI model that defines the details of the image to be generated.
[0960] "Quality check" is the process of checking the resolution and brightness of uploaded images and selecting images suitable for analysis.
[0961] "Saving" refers to the act of recording the generated image and proposal content in a database and associating them with the user profile.
[0962] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (blue-based, yellow-based) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like. Specific embodiments are described below.
[0963] First, a user accesses the system's website or application and logs in to upload their own image. After logging in, the user selects their own image and clicks the upload button. At this time, the user's device sends the photo file to the server.
[0964] Next, the server receives the uploaded photo. It performs a quality check on the received photo to ensure it is suitable for analysis. Specifically, it checks whether the face is clearly visible, and whether the brightness and resolution are sufficient. This is done using OpenCV and other image analysis libraries.
[0965] Once the photos pass the quality check, they are fed into an AI model on a server that extracts facial features and analyzes skin tone, hair color, and bone shape to determine the person's personal color (blue-based or yellow-based) and bone type (straight, wavy, natural).
[0966] The server compares the analysis results obtained from the AI model with a database that suggests the best styling items for the user. Suggestions include clothing color and style, hair color, makeup products, colored contact lenses, etc. These suggestions are used to generate and store information to be provided to the user.
[0967] Furthermore, the server requests the generative AI model to generate an image that reflects the suggested styling items. The generative AI model creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, colored contact lenses). An example of a specific prompt is, "Based on this image, please generate an image of styling items (soft coral clothes, fresh peach-toned blush, warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type."
[0968] The generated image is received by the server and stored in association with the user data. Finally, the server sends the analysis results, proposals, and generated image to the terminal, which then displays this information on the user's screen.
[0969] Users can check the displayed information and purchase items they like from the online store. Purchase links are also provided for suggested styling items, making it easy for users to purchase the products.
[0970] As described above, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[0971] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0972] Step 1: A user visits the system's website or application and logs in.
[0973] Specifically, the user enters their email address and password and clicks the login button, which causes the server to process the user's authentication data and output a login success or failure result.
[0974] Step 2: User selects their image and clicks the upload button.
[0975] Specifically, the user opens the photo library on their device and selects an appropriate image. When the user clicks the upload button, the device inputs the selected image file to be sent to the server. The server receives this request and outputs the image file to a temporary directory.
[0976] Step 3: The server checks the quality of the uploaded photos.
[0977] Specifically, the server uses an image analysis library such as OpenCV to check the image's resolution, brightness, and whether or not it contains a face. The input is the uploaded image file, and image analysis is performed on it. The output is the quality check result (appropriate or inappropriate). If it is determined to be inappropriate, the server notifies the user to upload the image again.
[0978] Step 4: The server inputs photos that pass the quality check into the AI model, which analyzes color attributes and bone structure type.
[0979] Specifically, the server sends the image to the AI model, which extracts facial features. The image must pass a quality check before it is input. The AI model analyzes skin tone, hair color, and bone structure, and outputs a personal color (blue-based or yellow-based) and bone structure type (straight, wavy, or natural).
[0980] Step 5: The server recommends styling items based on the analysis results obtained from the AI model.
[0981] Specifically, the server queries the database for styling items that fit the user's analysis results. The input is the analysis results, which are then compared with the database. The output is styling suggestions such as clothing color and style, hair color, makeup products, and colored contact lenses.
[0982] Step 6: The server generates a prompt sentence for the generated AI model and requests it to generate an image.
[0983] Specifically, the server generates a prompt and sends it to the generative AI model. The input is a styling suggestion, which is converted into text. An example of a prompt is, "Based on this image, please generate an image of styling items (soft coral clothing, fresh peach-toned blush, and warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type." The generative AI model generates an image based on this and sends it back to the server as output.
[0984] Step 7: The server receives the generated image and stores it in association with the user data.
[0985] Specifically, the server receives the image returned from the generative AI model, associates it with the user's profile, and stores it in a database. The input is the image from the generative AI model, and the output is a database update.
[0986] Step 8: The server sends the analysis results, proposals, and generated images to the terminal.
[0987] Specifically, the server compiles the generated information and generates an HTTP response to send to the user's device. The inputs include analysis results, styling suggestions, and images, and these data are sent together as a single response. The output is sent to the device.
[0988] Step 9: The terminal displays the received information on the user screen.
[0989] Specifically, the terminal analyzes the data received from the server and displays it in a format that is easy for the user to view. The input is the data from the server, and the output is displayed on the user interface.
[0990] Step 10: The user reviews the displayed information and purchases the items they like from the online store.
[0991] Specifically, the user clicks on a link to a suggested styling item and is taken to an online store. The input is the user's click, and the output is the browser displaying the online store page. The user then completes the purchase process and orders the product.
[0992] (Application example 1)
[0993] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0994] In conventional fashion try-on systems, users had to spend a lot of time and effort finding the perfect styling item for themselves. It was also difficult to receive styling advice without actually trying the items on. Furthermore, the lack of visual information to help users visualize the suggested items made the process of purchasing complicated, hindering the online shopping experience.
[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0996] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, generative AI model means for generating an image of the user's appearance when incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for providing the user with links to purchase items for realizing the suggested styling, means for generating an image using prompt text for the generative AI model, and means for checking the quality of the images and selecting images suitable for analysis. This enables users to efficiently find styling items that are best suited to them and instantly obtain specific images based on visual information, thereby improving the online shopping experience.
[0997] "User" refers to a person who uses this system to upload their own image and receive suggestions for the most suitable styling items.
[0998] "Means for uploading images" refers to the interface or process that allows users to send their own image data to the server.
[0999] "Means for analyzing images to determine a user's color attributes and bone structure type" refers to processes or algorithms that analyze a user's skin tone, hair color, and bone structure based on uploaded images to identify a user's color attributes and bone structure type.
[1000] "Means for suggesting styling items" refers to the process of algorithms and database searches that recommend the most suitable clothes, makeup products, accessories, etc. to users based on the analysis results.
[1001] "Generative AI Model" refers to the artificial intelligence model used to generate an image of what a user would look like when incorporating a suggested styling item.
[1002] The "means for displaying the proposed content and the generated image image" refers to a user interface or display for providing the user with the styling proposal and the generated image image in a visible form.
[1003] "Means for providing links to purchase items" refers to the process of providing a user with links to online stores or shopping lists for purchasing suggested styling items.
[1004] "Means for quality checks and selection of images suitable for analysis" refers to an algorithm that evaluates whether uploaded images are suitable for analysis, checking whether faces are clearly visible and whether the brightness and resolution are sufficient.
[1005] A "prompt" is a sentence or text that provides specific instructions or information to a generative AI model for generating an image.
[1006] A "virtual fitting room" refers to a system or application that allows users to upload their own images, receive suggestions for the most suitable styling items, and virtually try them on.
[1007] This detailed description of the present invention will explain in detail how the components of the system work together to generate and display styling item suggestions and images to the user.
[1008] System Program
[1009] The system allows users to upload their own images, analyzes them, and determines their color attributes and body type. Based on the results, it recommends optimal styling items and generates images incorporating the suggested items. This information is then displayed to the user, and if necessary, a link to purchase the suggested items is provided. It also uses a generative AI model to prompt users when generating specific images.
[1010] What the program does
[1011] 1. Upload a photo
[1012] Users access the system's web application and upload their own photos, which are then sent from devices such as smartphones and PCs to the server.
[1013] 2. Receiving and analyzing photos
[1014] The server checks the quality of the received photos, ensuring that the face is clearly visible and that the brightness and resolution are sufficient. To analyze photos that pass the quality check, an AI model specialized for image analysis is used. An AI model powered by TensorFlow is used to determine the user's color attributes and body type.
[1015] 3. Styling item suggestions
[1016] Based on the results of photo analysis, the system compares the results with a database to suggest the most suitable styling items. These styling items include clothing color and style, makeup, accessories, colored contact lenses, etc. The suggestions are generated on the server side and stored in association with user data.
[1017] 4. Image generation
[1018] The server requests the generative AI model to generate an image that reflects the suggested styling items. Specific prompts are used to provide instructions. For example, a prompt like this might be used: "Upload a photo of the user and determine their color attribute "cool undertones" and bone type "straight." Then, generate an image that incorporates a light blue blouse, natural makeup, and clear blue colored contact lenses."
[1019] 5. Displaying results and making purchasing suggestions
[1020] The server sends the analysis results, recommendations, and generated images to the device, which display them on the user's screen. The user can review the displayed information and purchase the items they like from the online store. A purchase link is provided for each suggested item.
[1021] Specific examples
[1022] When User A uploads a photo of himself / herself and the server receives and analyzes the photo, it determines that User A's color attribute is "cool-toned" and his / her bone structure type is "straight."
[1023] The server suggests the best styling items for user A: a light blue blouse, natural makeup, and clear blue colored contact lenses.
[1024] The server uses the generative AI model to generate an image of user A that reflects the suggested item and sends it to the device.
[1025] User A can check the images generated on their device and purchase items they like from the online store.
[1026] This invention allows users to efficiently find the styling items that are best suited to them and instantly get a concrete image of them, greatly improving the online shopping experience.
[1027] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1028] Step 1:
[1029] Users upload their own images.
[1030] A user accesses the system's web application, selects a photo file, and clicks the upload button. The terminal sends the selected photo file to the server. The input data is the user's photo file, and the output is the photo file sent to the server.
[1031] Step 2:
[1032] The server checks the quality of the received photos.
[1033] The server performs a quality check on the received photo before analyzing it. It checks whether the face is clearly visible and whether the brightness and resolution are sufficient. The input data is the user's photo file, and the output is a photo file that has passed the quality check. If the quality check is not passed, the server sends a notification to the user, encouraging them to upload again.
[1034] Step 3:
[1035] The server analyzes the photo and determines the user's color attributes and body type.
[1036] The server inputs photos that pass the quality check into the AI model for image analysis. The AI model extracts the user's facial features and analyzes their skin tone, hair color, and bone structure to determine their color attributes (cool or warm undertones) and bone structure type (straight, wavy, natural). The input data is a photo file that has passed the quality check, and the output is the analysis results of the user's color attributes and bone structure type.
[1037] Step 4:
[1038] The server will suggest styling items.
[1039] The server compares the acquired analysis results with the database and suggests styling items that are best suited to the user. Suggestions include clothing color and style, makeup products, colored contact lenses, etc. The input data is the analysis results of the user's color attributes and body type, and the output is a list of suggested styling items.
[1040] Step 5:
[1041] The server requests the generative AI model to generate an image.
[1042] The server sends a request to the generative AI model using a prompt to generate an image that reflects the suggested styling items. The input data is a list of suggested styling items and the prompt, and the output is the generated image. An example of a prompt is, "Please upload a photo of the user and determine the color attribute 'cool' and bone type 'straight'. Then, generate an image that includes a light blue blouse, natural makeup, and clear blue colored contact lenses."
[1043] Step 6:
[1044] The server displays the generated image to the user.
[1045] The server receives the generated image and stores it in association with the user data. It then transmits the analysis results, suggested styling items, and the generated image to the terminal. The input data is the user data associated with the generated image, and the output is the analysis results, suggested item list, and image displayed to the user.
[1046] Step 7:
[1047] The device will provide a link to purchase the suggested styling item.
[1048] The terminal displays a list of suggested styling items and a link to purchase them to the user. The user can click the link to purchase the suggested items from the online store. The input data is the list of suggested styling items, and the output is a user screen with the purchase link displayed.
[1049] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1050] The system of the present invention allows users to upload their own images, and the system analyzes the images to determine color attributes (cool or warm undertones) and bone structure (straight, wavy, natural), and then suggests the most suitable styling items. It also combines this with an emotion engine that recognizes the user's emotions to provide even more personalized styling suggestions. The system can also generate and display to the user an image of what the suggested styling items would look like when worn. It also provides a link to purchase the suggested styling items.
[1051] System Program Processing
[1052] 1. Upload a photo
[1053] A user accesses the system's website or application and logs in.
[1054] User clicks the image upload button, selects and uploads their own image.
[1055] The terminal transmits the selected image file to the server.
[1056] 2. Receiving and analyzing photos
[1057] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[1058] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[1059] 3. Emotion Recognition by Emotion Engine
[1060] The server simultaneously uses an emotion engine to analyze emotions from the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[1061] 4. Styling suggestion generation
[1062] The server compares the results of the image analysis and emotion engine with a database to suggest styling items suitable for the user. The suggestions include the following items:
[1063] Clothing color and style
[1064] hair color
[1065] Makeup products
[1066] colored contact lenses
[1067] The server uses these suggestions to generate and store information to provide to the user, which is tailored to take into account feedback on the user's emotional state.
[1068] 5. Image Generation
[1069] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[1070] 6. Displaying results and making purchasing suggestions
[1071] The server sends the analysis results, proposals, and generated images to the terminal.
[1072] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1073] Analysis results of the user's personal color and bone structure type
[1074] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1075] Generated image
[1076] Links to purchase recommended items and access to our online store
[1077] Specific examples
[1078] 1. The user uploads their image to the system.
[1079] The device transfers the photo to the server.
[1080] 2. The server uses an AI model and emotion engine to analyze the photo, determine the user's color attributes as "spring warm skin tone" and body type as "wave," and further recognize the emotion of "joy" from the user's facial expression.
[1081] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[1082] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[1083] The terminal displays this to the user.
[1084] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[1085] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[1086] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[1087] The processing flow will be explained below.
[1088] Step 1:
[1089] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[1090] Step 2:
[1091] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[1092] Step 3:
[1093] The server receives the uploaded images. After receiving them, it checks the quality of the images to ensure that the face is clearly visible and that the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[1094] Step 4:
[1095] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[1096] Step 5:
[1097] At the same time, the server uses an emotion engine to analyze the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to determine emotions such as "happiness," "sadness," "surprise," and "anger."
[1098] Step 6:
[1099] The server compares the results of image analysis and emotion recognition with a database to suggest the most suitable styling items for the user. Specifically, it selects clothing colors and styles, hair colors, makeup products, colored contact lenses, etc. that correspond to the user's color attributes and body type. The suggestions are adjusted taking into account the user's emotional state.
[1100] Step 7:
[1101] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[1102] Step 8:
[1103] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[1104] Step 9:
[1105] The server sends the analysis results, suggestions, and generated images to the device, along with a link to purchase the suggested styling items.
[1106] Step 10:
[1107] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1108] Analysis results of the user's personal color and bone structure type
[1109] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1110] Generated image
[1111] Links to purchase recommended items and access to our online store
[1112] Step 11:
[1113] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[1114] Specific examples
[1115] 1. The user uploads their image to the system.
[1116] The device transfers the photo to the server.
[1117] 2. The server analyzes the photo using an AI model to determine the attributes of "spring warm skin tone" and "wave skin tone." The emotion engine then recognizes "joy" from the facial expressions in the image.
[1118] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[1119] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[1120] The terminal displays this to the user.
[1121] 4. The device displays a link to purchase items to help the user achieve the suggested styling.
[1122] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[1123] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[1124] Example 2
[1125] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1126] Conventional styling suggestion systems can analyze a user's personal color and bone structure to suggest suitable styling items, but they cannot make suggestions that take into account the user's emotional state. As a result, the suggested styling items may not match the user's current emotions, making it difficult to increase user satisfaction. Another issue is that it is difficult for users to specifically imagine the suggested styling, which can reduce their motivation to purchase.
[1127] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1128] In this invention, the server includes a means for users to upload their own images, a means for analyzing the uploaded images and determining the user's color attributes and body type, a means for analyzing the user's emotions based on the uploaded images, a means for suggesting styling items suitable for the user based on the results of the determination and emotion analysis, a means for generating an image of the user incorporating the suggested styling items, and a means for displaying the suggested content and the generated image to the user. This enables personalized styling suggestions that take into account not only the user's personal color and body type but also their emotional state, thereby increasing user satisfaction. Furthermore, the generated image provides a concrete styling image, thereby increasing the user's desire to purchase.
[1129] "User" refers to a person who uses the system to upload their own image and receive styling suggestions.
[1130] "Image" refers to photographic data uploaded by a user to the system.
[1131] "Color attribute" refers to a personal color (e.g., cool-toned, warm-toned) that is classified based on the user's skin tone and hair color.
[1132] "Body type" refers to a styling type (e.g., straight, wavy, natural) that is classified based on the user's body shape and skeletal characteristics.
[1133] "Emotion" refers to the psychological state (e.g., joy, sadness, surprise, anger) that can be read from the user's facial expression.
[1134] "Styling items" refer to fashion-related items such as clothes, hair color, makeup products, and colored contact lenses that are recommended to users.
[1135] "Image" refers to image data used to generate a new visual for the user that reflects the suggested styling item.
[1136] "Server" refers to the central computing resource that performs the primary data processing and management of the system.
[1137] "Terminal" refers to a device such as a computer or smartphone that a user uses to access the system.
[1138] "Quality check" refers to the inspection process to ensure that uploaded images are suitable for analysis.
[1139] "Analysis" refers to the process of extracting features and recognizing emotions from uploaded images using AI models and emotion engines.
[1140] "Suggestion" refers to the act of providing the user with the most suitable styling items based on the analysis results.
[1141] "Generative AI" refers to artificial intelligence that generates image images that reflect suggested styling items in the user's image.
[1142] A "prompt sentence" refers to an instruction sentence that instructs the generation AI to generate a specific image.
[1143] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and then combines them with an emotion engine to provide personalized styling suggestions. The system can generate and display to the user an image of what the suggested styling items would look like. It also provides a link to purchase the suggested styling items.
[1144] The system uses the following major hardware and software:
[1145] 1. Hardware
[1146] Device: The computer or smartphone that a user uses to access the system.
[1147] Server: The central computing resource that handles the main data processing and management of the system.
[1148] 2. Software
[1149] Website or application: An interface where users can upload images and view analysis results and recommendations
[1150] AI model: a model for image analysis (e.g., a TensorFlow-based model)
[1151] Emotion engine: API for emotion recognition (e.g., Face API by Microsoft)
[1152] Generative AI: A model for generating images that reflect suggested styling items (e.g., DALL-E by OpenAI)
[1153] A specific embodiment of the system will be described below.
[1154] Uploading and receiving photos
[1155] Users access the system's website or application, log in, and then click the image upload button to select and upload their own image. At this time, the device sends the selected image file to the server. The image file is securely transmitted using HTTPS.
[1156] Image quality check and analysis
[1157] The server performs a quality check on the received images to determine whether they are suitable for analysis. For example, it uses OpenCV to check whether the face in the image is clearly visible and whether the resolution and brightness are appropriate. Images that pass the quality check are then analyzed using an AI model (e.g., TensorFlow). The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color attribute (e.g., cool-toned, warm-toned) and bone type (e.g., straight, wavy, natural).
[1158] emotion recognition
[1159] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze the user's emotions from the image. The emotion engine analyzes the user's facial expressions, eye movements, mouth features, etc., and recognizes emotions such as joy, sadness, surprise, and anger. The analysis results and emotion analysis results are then stored in a database.
[1160] Styling suggestions
[1161] The server compares the results of image analysis and emotion analysis with the system's styling database. For example, for a user with "spring warm skin tone," "waves," and the emotion "joy," it would suggest "soft coral" clothing and "fresh peach-toned blush." The server generates suggestions and saves a list of styling items suitable for the user in JSON format.
[1162] Image generation
[1163] The server requests the generation AI (e.g., DALL-E by OpenAI) to generate an image that reflects the suggested styling items. An example of a prompt to be input to the generation AI is "An image of a warm-toned, spring-season woman wearing soft coral clothing." The generation AI model creates a new image for the user that reflects the suggestions. The server retrieves the generated image and saves it by associating it with the user's profile.
[1164] Viewing and purchasing results
[1165] The server sends a dataset containing the analysis results, suggestions, and generated images to the device. The device displays the received information on the user's screen. The user sees:
[1166] Analysis results of the user's personal color and bone structure type
[1167] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1168] Generated image
[1169] Links to purchase recommended items and access to our online store
[1170] Specific examples
[1171] For example, a user uploads an image of themselves to the system. The device sends the uploaded image to the server, which analyzes the image using an AI model and emotion engine. If the analysis results indicate that the user's color attributes are "spring warm skin," their bone structure is "wave," and their emotion is "joy," the server will suggest "soft coral" clothing, "fresh peach-toned blush," and "warm brown" colored contact lenses as the optimal styling for the user. Using generative AI, a new image of the user incorporating these styling items is created and displayed to the user. The user can then purchase their favorite products from the online store.
[1172] In this way, the system of the present invention efficiently discovers the user's personal style, makes styling suggestions that take emotions into consideration, and allows the user to easily purchase specific styling items.
[1173] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1174] Step 1:
[1175] A user accesses a website or application of the system and logs in.
[1176] Input: User login information (email address, password)
[1177] How it works: A user enters their email address and password on the login screen and clicks the login button. The system authenticates the user and starts a login session.
[1178] Output: Successful login notification, user's home screen
[1179] Step 2:
[1180] User clicks the image upload button, selects and uploads their own image.
[1181] Input: User's image file
[1182] How it works: The user clicks the image upload button on the main screen. A file selection dialog opens and the user selects their image file. The device sends the selected image file to the server using an HTTP POST request.
[1183] Output: Notification of image file transfer to server
[1184] Step 3:
[1185] The server receives the uploaded images and performs a quality check.
[1186] Input: Uploaded image file
[1187] How it works: The server receives image files and performs a quality check on the images using a facial recognition algorithm (e.g., OpenCV). Check items include facial clarity, resolution, brightness, etc. Only image files that pass the quality check proceed to the next analysis step.
[1188] Output: Image files that pass the quality check or notification of poor quality
[1189] Step 4:
[1190] The server inputs images that pass the quality check into an AI model (e.g., TensorFlow) and performs image analysis.
[1191] Input: Image files that have passed the quality check
[1192] How it works: The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color profile (cool, warm) and bone type (straight, wavy, natural).
[1193] Output: Analysis results of color attributes and skeletal type
[1194] Step 5:
[1195] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze emotions from the uploaded image and the user's facial expressions.
[1196] Input: Uploaded image file
[1197] Behavior: The emotion engine analyzes the user's facial expressions, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[1198] Output: Emotion analysis results
[1199] Step 6:
[1200] The server suggests styling items suitable for the user based on the results of image analysis and the emotion engine.
[1201] Input: Color attribute and skeletal type analysis results, emotion analysis results
[1202] How it works: The system compares the styling database within the system and selects the styling items that are best suited to the user (e.g., "soft coral" clothing, "fresh peach-toned blush," etc.).
[1203] Output: A list of suggested styling items
[1204] Step 7:
[1205] The server uses a generation AI (e.g., DALL-E) to generate an image that reflects the suggested styling items.
[1206] Input: A list of suggested styling items
[1207] How it works: The user provides the AI with a prompt (e.g., "Image of a warm-toned spring woman wearing a soft coral dress") and requests that it generate an image. The AI then creates a new image for the user that reflects the suggestions.
[1208] Output: Generated image
[1209] Step 8:
[1210] The server sends the analysis results, proposals, and generated images to the terminal.
[1211] Input: Analysis results, proposals, generated images
[1212] How it works: The server sends these datasets in JSON format to the device.
[1213] Output: Notification of completion of dataset transmission to the terminal
[1214] Step 9:
[1215] The terminal displays the received information on the user screen.
[1216] Input: Analysis results, proposals, and generated images sent in JSON format
[1217] How it works: The device analyzes the received data and dynamically renders a user screen using a front-end framework such as Vue.js or React.js. The following information is displayed on the user screen: analysis results, recommendations, generated images, links to purchase recommended items, and links to access the online store.
[1218] Output: Styling suggestions displayed on the user's screen
[1219] Step 10:
[1220] The user clicks on the link for the item of interest and is taken to the specified online store page.
[1221] Input: User clicks
[1222] What it does: When the user clicks, the browser navigates to the product page in the specified online store.
[1223] Output: Online store product page
[1224] (Application example 2)
[1225] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1226] To quickly and accurately respond to today's diversifying customer needs, personalized styling suggestions are required. However, while conventional systems take into account a user's personal color and bone structure, they do not provide personalized styling that takes into account emotional fluctuations. Furthermore, there is a lack of methods for providing real-time suggestions in-store or instantly providing visually easy-to-understand images. The present invention aims to solve these problems and improve the customer experience.
[1227] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload their own image, means for analyzing the uploaded image and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, means for generating an image of the user when incorporating the suggested styling items, means for using an emotion engine that recognizes emotions from the user's facial expression, means for adjusting the suggestions taking into account the emotion recognition results, and means for displaying the suggestions and the generated image to the user. This enables the user to receive personalized styling suggestions in real time according to their emotional state.
[1228] A "user" is an individual or customer who uses the system.
[1229] An "image uploading means" is a method or device by which a user can send their image to the system.
[1230] "Means for analyzing images" refers to a method or device for identifying user information (color attributes and bone structure type) based on uploaded images.
[1231] "Color attribute" indicates the personal color (cool or warm) based on the user's skin tone.
[1232] "Body type" indicates a classification based on the shape of the user's body frame (straight, wavy, natural).
[1233] The "means for suggesting styling items" is a method or device for suggesting clothes, accessories, makeup products, etc. that are suitable for the user.
[1234] The "means for generating an image" is a method or device for creating a new image of the user incorporating the suggested styling item.
[1235] An "emotion engine" is a program or device that recognizes emotions (happiness, sadness, surprise, anger, etc.) from a user's facial expression.
[1236] The "means for adjusting the suggestions taking into account the emotion recognition results" is a method or apparatus for modifying or adjusting styling suggestions based on the emotions recognized by the emotion engine.
[1237] The "means for displaying the generated image" is a method or device for visually presenting new styling suggestions to the user.
[1238] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and uses an emotion engine to recognize emotions from the user's facial expressions. Based on the results of these analyses, the system suggests optimal styling items for the user and generates and displays images using the suggested items. The system aims to visually present the suggestions and the generated images to the user. Links to purchase the suggested styling items are also provided.
[1239] System hardware and software configuration
[1240] This system uses the following hardware and software:
[1241] Hardware:
[1242] Tablet device (e.g. iPad)
[1243] Smart mirrors (e.g., Novera Smart Mirror)
[1244] software:
[1245] AI image analysis engine (e.g. Google Cloud Vision API)
[1246] Emotion recognition engine (e.g. Microsoft Azure Face API)
[1247] Generative AI models (e.g., OpenAI DALL-E)
[1248] Data processing and calculation flow
[1249] Upload a photo
[1250] Users access the system from a tablet or smart mirror and upload their own images, which are then sent to the server.
[1251] Image reception and analysis
[1252] The server receives the uploaded images and performs a quality check. Images that pass the quality check are input into an AI image analysis engine (Google Cloud Vision API) to determine personal color and bone structure type.
[1253] Emotion recognition by emotion engine
[1254] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to analyze the user's facial expressions and recognize emotions such as joy, sadness, surprise, and anger.
[1255] Generate styling suggestions
[1256] The server then uses the results of image analysis and emotion recognition to suggest styling items suitable for the user, including clothing color and style, hair color, makeup products, and colored contact lenses, while also taking into account the user's emotional state.
[1257] Image generation
[1258] The server uses a generative AI model (OpenAI DALL-E) to generate new images for the user incorporating the suggested styling items.
[1259] Displaying results and making purchasing suggestions
[1260] The generated image and the proposed content are sent to the device, which then displays them on the user's screen. The displayed content includes the analysis results of personal color and bone structure type, recommended styling items, the generated image, and a link to purchase the items.
[1261] Specific examples
[1262] Below are some specific examples.
[1263] Example prompt sentence:
[1264] "Using Google Cloud Vision API and Microsoft Azure Face API, we analyze the customer's facial image to determine their personal color. We then analyze their facial expressions to recognize their emotions. Based on the results, we provide fashion styling and product suggestions, and generate new images."
[1265] Specifically, when a user uploads a photo of themselves using a tablet or smart mirror, the server analyzes the photo and determines the user's color attributes as "cool winter" and bone type as "straight," and then uses emotion recognition to identify the emotion of "surprise." Based on this, the server suggests a "sharp black jacket" and "cool-toned makeup products" to the user, and uses a generative AI model to create an image that reflects these suggestions and displays it on the device. The user can also receive a link to purchase the items from an online store based on the suggestions.
[1266] In this way, the system of the present invention allows users to easily receive personalized styling suggestions based on their individual characteristics and emotions, and by visually checking the generated image, the suggestions can be easily understood.
[1267] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1268] Step 1:
[1269] A user uploads their own image using a tablet device or smart mirror. The user opens a dedicated application, clicks the image upload button, and selects their own image file. The input is the user's image file, and the output is that the device sends this image to the server.
[1270] Step 2:
[1271] The server receives the uploaded image. The server performs a quality check on the image to ensure that the face is clearly visible, and that the brightness and resolution are sufficient. Only images that pass this quality check proceed to the next step. The input is the user's image file, and the output is an image that passes the quality check.
[1272] Step 3:
[1273] The server inputs images that pass the quality check into an AI image analysis engine (Google Cloud Vision API) to determine the user's color attributes and bone structure. During this process, the AI analyzes facial features and identifies skin tone, hair color, and bone structure. The input is an image that passes the quality check, and the output is the user's color attributes and bone structure data.
[1274] Step 4:
[1275] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to recognize emotions from the user's facial expressions. This process detects subtle facial features and identifies emotions such as joy, sadness, surprise, and anger. The input is a user image that has passed quality checks, and the output is the user's emotional data.
[1276] Step 5:
[1277] The server compares the results of image analysis and emotion recognition with a database of styling item suggestions. It then suggests the most suitable clothing, hair color, makeup products, and colored contact lenses for the user. The suggestions are adjusted based on the emotion data. The input is the user's color attributes, bone structure, and emotion data, and the output is a list of suggested styling items.
[1278] Step 6:
[1279] The server uses a generative AI model (OpenAI DALL-E) to generate an image that reflects the suggested styling items. In this process, a new image is created by combining styling items with the user's facial photo. The input is a list of suggested styling items and the user's image, and the output is the generated image.
[1280] Step 7:
[1281] The server sends the generated image and proposal details to the terminal, which then displays them on the user's screen. The displayed content includes the analysis results of the personal color and body type, recommended styling items, the generated image, and a link to purchase the items. The input is the generated image and proposal details, and the output is the specific styling proposal information displayed to the user.
[1282] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1283] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1284] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1285] [Fourth embodiment]
[1286] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1287] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1288] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1289] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1290] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1291] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1292] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1293] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1294] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1295] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1296] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1297] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1298] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1299] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (cool or warm skin tones) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like.
[1300] System Program Processing
[1301] 1. Upload a photo
[1302] A user accesses the system's website or application and logs in.
[1303] The user selects their image and clicks the upload button.
[1304] The device sends the photo file to the server.
[1305] 2. Receiving and analyzing photos
[1306] The server receives the uploaded photos.
[1307] The server checks the quality of the received photo to ensure it is suitable for analysis, specifically whether the face is clearly visible and whether the brightness and resolution are sufficient.
[1308] The server inputs photos that pass the quality check into the AI model, which analyzes the user's color attributes and bone structure. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure to determine the user's personal color (cool or warm undertones) and bone structure (straight, wavy, natural).
[1309] 3. Styling suggestion generation
[1310] The server compares the analysis results obtained from the AI model with a database to suggest the most suitable styling items for the user. The suggestions include:
[1311] Clothing color and style
[1312] hair color
[1313] Makeup products
[1314] colored contact lenses
[1315] The server uses these suggestions to generate and store information to provide to the user.
[1316] 4. Image generation
[1317] The server requests the generation AI to generate an image that reflects the proposed styling items.
[1318] The generative AI model creates a new image of the user incorporating each of the suggested elements (clothing, hair color, makeup, colored contact lenses).
[1319] The server receives the generated image and stores it in association with the user data.
[1320] 5. Displaying results and making purchasing suggestions
[1321] The server sends the analysis results, proposals, and generated images to the terminal.
[1322] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1323] Analysis results of the user's personal color and bone structure type
[1324] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1325] Generated image
[1326] Links to purchase recommended items and access to our online store
[1327] The user can check the displayed information and purchase the items they like from the online store.
[1328] Specific examples
[1329] 1. The user uploads their image to the system.
[1330] The device transfers the photo to the server.
[1331] 2. The server uses an AI model to analyze the photo and determine the user's color attribute as "spring warm skin tone" and their bone structure as "wave."
[1332] Based on this, the server will suggest "soft coral" as the best clothing color for the user, "fresh peach-toned blush" as a makeup product, and "warm brown" as colored contact lenses.
[1333] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[1334] The terminal displays this to the user.
[1335] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[1336] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[1337] In this way, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[1338] The processing flow will be explained below.
[1339] Step 1:
[1340] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[1341] Step 2:
[1342] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[1343] Step 3:
[1344] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[1345] Step 4:
[1346] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[1347] Step 5:
[1348] The server then compares the image analysis results with a database to suggest styling items suitable for the user, such as clothing colors and styles, hair colors, makeup products, and colored contact lenses that match the user's color attributes and body type.
[1349] Step 6:
[1350] The server sends a request to the generative AI model to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[1351] Step 7:
[1352] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[1353] Step 8:
[1354] The server sends the analysis results, proposals, and generated images to the terminal.
[1355] Step 9:
[1356] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1357] Analysis results of the user's personal color and bone structure type
[1358] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1359] Generated image
[1360] Links to purchase recommended items and access to our online store
[1361] Step 10:
[1362] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[1363] Example 1
[1364] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1365] In conventional personal styling systems, even if users upload their own images, the system may not accurately determine color attributes or bone structure. It is also difficult for users to visualize how the suggested styling items will look on them, making it difficult to make a purchasing decision. Furthermore, depending on the quality of the image, the analysis results may be inaccurate, making it impossible to provide optimal styling suggestions.
[1366] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1367] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the results of the determination, means for generating an image of the user incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for checking the quality of the images and selecting images suitable for analysis, means for inputting a prompt to a generation AI model for generating an image reflecting the suggested styling items, and means for receiving the generated image and storing it in association with user data. This allows users to accurately grasp their own personal style, easily imagine the suggested styling items, and select and purchase the most suitable products.
[1368] A "user" is a person who accesses the system and receives suggestions for uploading images and styling items.
[1369] "Server" refers to a computer system that receives images uploaded by users, analyzes them, and generates and stores suggestions.
[1370] "Terminal" refers to the device a user uses to access the system, including smartphones and personal computers.
[1371] "Image" refers to a photo file uploaded by a user and used to identify facial features and color attributes.
[1372] "Analysis" refers to the process of determining color attributes and bone structure type based on uploaded images.
[1373] "Color attribute" refers to a personal color determined based on the user's skin tone and hair color, and includes "cool skin" and "warm skin."
[1374] "Body type" is a type classified based on the characteristics of the user's body frame and figure, and includes "straight," "wavy," "natural," and the like.
[1375] "Styling items" are items such as clothes, makeup products, and colored contact lenses that are suggested based on the user's color attributes and body type.
[1376] "Suggestion" refers to the act of providing styling items suitable for the user based on the results of the determination of color attributes and bone structure type.
[1377] An "image image" is a composite image that generates an image of the user wearing the suggested styling item.
[1378] A "generative AI model" is an artificial intelligence model that generates imagery incorporating suggested styling items based on an input prompt.
[1379] A "prompt sentence" is an input sentence to a generative AI model that defines the details of the image to be generated.
[1380] "Quality check" is the process of checking the resolution and brightness of uploaded images and selecting images suitable for analysis.
[1381] "Saving" refers to the act of recording the generated image and proposal content in a database and associating them with the user profile.
[1382] The system of the present invention allows users to upload their own images, analyzes the images, determines color attributes (blue-based, yellow-based) and bone structure (straight, wavy, natural), and suggests optimal styling items. It can also generate and display to the user an image of what the suggested styling items would look like. Specific embodiments are described below.
[1383] First, a user accesses the system's website or application and logs in to upload their own image. After logging in, the user selects their own image and clicks the upload button. At this time, the user's device sends the photo file to the server.
[1384] Next, the server receives the uploaded photo. It performs a quality check on the received photo to ensure it is suitable for analysis. Specifically, it checks whether the face is clearly visible, and whether the brightness and resolution are sufficient. This is done using OpenCV and other image analysis libraries.
[1385] Once the photos pass the quality check, they are fed into an AI model on a server that extracts facial features and analyzes skin tone, hair color, and bone shape to determine the person's personal color (blue-based or yellow-based) and bone type (straight, wavy, natural).
[1386] The server compares the analysis results obtained from the AI model with a database that suggests the best styling items for the user. Suggestions include clothing color and style, hair color, makeup products, colored contact lenses, etc. These suggestions are used to generate and store information to be provided to the user.
[1387] Furthermore, the server requests the generative AI model to generate an image that reflects the suggested styling items. The generative AI model creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, colored contact lenses). An example of a specific prompt is, "Based on this image, please generate an image of styling items (soft coral clothes, fresh peach-toned blush, warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type."
[1388] The generated image is received by the server and stored in association with the user data. Finally, the server sends the analysis results, proposals, and generated image to the terminal, which then displays this information on the user's screen.
[1389] Users can check the displayed information and purchase items they like from the online store. Purchase links are also provided for suggested styling items, making it easy for users to purchase the products.
[1390] As described above, the system of the present invention allows users to efficiently discover their own personal style and easily purchase specific styling items.
[1391] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1392] Step 1: A user visits the system's website or application and logs in.
[1393] Specifically, the user enters their email address and password and clicks the login button, which causes the server to process the user's authentication data and output a login success or failure result.
[1394] Step 2: User selects their image and clicks the upload button.
[1395] Specifically, the user opens the photo library on their device and selects an appropriate image. When the user clicks the upload button, the device inputs the selected image file to be sent to the server. The server receives this request and outputs the image file to a temporary directory.
[1396] Step 3: The server checks the quality of the uploaded photos.
[1397] Specifically, the server uses an image analysis library such as OpenCV to check the image's resolution, brightness, and whether or not it contains a face. The input is the uploaded image file, and image analysis is performed on it. The output is the quality check result (appropriate or inappropriate). If it is determined to be inappropriate, the server notifies the user to upload the image again.
[1398] Step 4: The server inputs photos that pass the quality check into the AI model, which analyzes color attributes and bone structure type.
[1399] Specifically, the server sends the image to the AI model, which extracts facial features. The image must pass a quality check before it is input. The AI model analyzes skin tone, hair color, and bone structure, and outputs a personal color (blue-based or yellow-based) and bone structure type (straight, wavy, or natural).
[1400] Step 5: The server recommends styling items based on the analysis results obtained from the AI model.
[1401] Specifically, the server queries the database for styling items that fit the user's analysis results. The input is the analysis results, which are then compared with the database. The output is styling suggestions such as clothing color and style, hair color, makeup products, and colored contact lenses.
[1402] Step 6: The server generates a prompt sentence for the generated AI model and requests it to generate an image.
[1403] Specifically, the server generates a prompt and sends it to the generative AI model. The input is a styling suggestion, which is converted into text. An example of a prompt is, "Based on this image, please generate an image of styling items (soft coral clothing, fresh peach-toned blush, and warm brown colored contact lenses) that match the warm-toned spring color type and wavy bone structure type." The generative AI model generates an image based on this and sends it back to the server as output.
[1404] Step 7: The server receives the generated image and stores it in association with the user data.
[1405] Specifically, the server receives the image returned from the generative AI model, associates it with the user's profile, and stores it in a database. The input is the image from the generative AI model, and the output is a database update.
[1406] Step 8: The server sends the analysis results, proposals, and generated images to the terminal.
[1407] Specifically, the server compiles the generated information and generates an HTTP response to send to the user's device. The inputs include analysis results, styling suggestions, and images, and these data are sent together as a single response. The output is sent to the device.
[1408] Step 9: The terminal displays the received information on the user screen.
[1409] Specifically, the terminal analyzes the data received from the server and displays it in a format that is easy for the user to view. The input is the data from the server, and the output is displayed on the user interface.
[1410] Step 10: The user reviews the displayed information and purchases the items they like from the online store.
[1411] Specifically, the user clicks on a link to a suggested styling item and is taken to an online store. The input is the user's click, and the output is the browser displaying the online store page. The user then completes the purchase process and orders the product.
[1412] (Application example 1)
[1413] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1414] In conventional fashion try-on systems, users had to spend a lot of time and effort finding the perfect styling item for themselves. It was also difficult to receive styling advice without actually trying the items on. Furthermore, the lack of visual information to help users visualize the suggested items made the process of purchasing complicated, hindering the online shopping experience.
[1415] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1416] In this invention, the server includes means for users to upload their own images, means for analyzing the uploaded images and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, generative AI model means for generating an image of the user's appearance when incorporating the suggested styling items, means for displaying the suggestions and the generated image to the user, means for providing the user with links to purchase items for realizing the suggested styling, means for generating an image using prompt text for the generative AI model, and means for checking the quality of the images and selecting images suitable for analysis. This enables users to efficiently find styling items that are best suited to them and instantly obtain specific images based on visual information, thereby improving the online shopping experience.
[1417] "User" refers to a person who uses this system to upload their own image and receive suggestions for the most suitable styling items.
[1418] "Means for uploading images" refers to the interface or process that allows users to send their own image data to the server.
[1419] "Means for analyzing images to determine a user's color attributes and bone structure type" refers to processes or algorithms that analyze a user's skin tone, hair color, and bone structure based on uploaded images to identify a user's color attributes and bone structure type.
[1420] "Means for suggesting styling items" refers to the process of algorithms and database searches that recommend the most suitable clothes, makeup products, accessories, etc. to users based on the analysis results.
[1421] "Generative AI Model" refers to the artificial intelligence model used to generate an image of what a user would look like when incorporating a suggested styling item.
[1422] The "means for displaying the proposed content and the generated image image" refers to a user interface or display for providing the user with the styling proposal and the generated image image in a visible form.
[1423] "Means for providing links to purchase items" refers to the process of providing a user with links to online stores or shopping lists for purchasing suggested styling items.
[1424] "Means for quality checks and selection of images suitable for analysis" refers to an algorithm that evaluates whether uploaded images are suitable for analysis, checking whether faces are clearly visible and whether the brightness and resolution are sufficient.
[1425] A "prompt" is a sentence or text that provides specific instructions or information to a generative AI model for generating an image.
[1426] A "virtual fitting room" refers to a system or application that allows users to upload their own images, receive suggestions for the most suitable styling items, and virtually try them on.
[1427] This detailed description of the present invention will explain in detail how the components of the system work together to generate and display styling item suggestions and images to the user.
[1428] System Program
[1429] The system allows users to upload their own images, analyzes them, and determines their color attributes and body type. Based on the results, it recommends optimal styling items and generates images incorporating the suggested items. This information is then displayed to the user, and if necessary, a link to purchase the suggested items is provided. It also uses a generative AI model to prompt users when generating specific images.
[1430] What the program does
[1431] 1. Upload a photo
[1432] Users access the system's web application and upload their own photos, which are then sent from devices such as smartphones and PCs to the server.
[1433] 2. Receiving and analyzing photos
[1434] The server checks the quality of the received photos, ensuring that the face is clearly visible and that the brightness and resolution are sufficient. To analyze photos that pass the quality check, an AI model specialized for image analysis is used. An AI model powered by TensorFlow is used to determine the user's color attributes and body type.
[1435] 3. Styling item suggestions
[1436] Based on the results of photo analysis, the system compares the results with a database to suggest the most suitable styling items. These styling items include clothing color and style, makeup, accessories, colored contact lenses, etc. The suggestions are generated on the server side and stored in association with user data.
[1437] 4. Image generation
[1438] The server requests the generative AI model to generate an image that reflects the suggested styling items. Specific prompts are used to provide instructions. For example, a prompt like this might be used: "Upload a photo of the user and determine their color attribute "cool undertones" and bone type "straight." Then, generate an image that incorporates a light blue blouse, natural makeup, and clear blue colored contact lenses."
[1439] 5. Displaying results and making purchasing suggestions
[1440] The server sends the analysis results, recommendations, and generated images to the device, which display them on the user's screen. The user can review the displayed information and purchase the items they like from the online store. A purchase link is provided for each suggested item.
[1441] Specific examples
[1442] When User A uploads a photo of himself / herself and the server receives and analyzes the photo, it determines that User A's color attribute is "cool-toned" and his / her bone structure type is "straight."
[1443] The server suggests the best styling items for user A: a light blue blouse, natural makeup, and clear blue colored contact lenses.
[1444] The server uses the generative AI model to generate an image of user A that reflects the suggested item and sends it to the device.
[1445] User A can check the images generated on their device and purchase items they like from the online store.
[1446] This invention allows users to efficiently find the styling items that are best suited to them and instantly get a concrete image of them, greatly improving the online shopping experience.
[1447] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1448] Step 1:
[1449] Users upload their own images.
[1450] A user accesses the system's web application, selects a photo file, and clicks the upload button. The terminal sends the selected photo file to the server. The input data is the user's photo file, and the output is the photo file sent to the server.
[1451] Step 2:
[1452] The server checks the quality of the received photos.
[1453] The server performs a quality check on the received photo before analyzing it. It checks whether the face is clearly visible and whether the brightness and resolution are sufficient. The input data is the user's photo file, and the output is a photo file that has passed the quality check. If the quality check is not passed, the server sends a notification to the user, encouraging them to upload again.
[1454] Step 3:
[1455] The server analyzes the photo and determines the user's color attributes and body type.
[1456] The server inputs photos that pass the quality check into the AI model for image analysis. The AI model extracts the user's facial features and analyzes their skin tone, hair color, and bone structure to determine their color attributes (cool or warm undertones) and bone structure type (straight, wavy, natural). The input data is a photo file that has passed the quality check, and the output is the analysis results of the user's color attributes and bone structure type.
[1457] Step 4:
[1458] The server will suggest styling items.
[1459] The server compares the acquired analysis results with the database and suggests styling items that are best suited to the user. Suggestions include clothing color and style, makeup products, colored contact lenses, etc. The input data is the analysis results of the user's color attributes and body type, and the output is a list of suggested styling items.
[1460] Step 5:
[1461] The server requests the generative AI model to generate an image.
[1462] The server sends a request to the generative AI model using a prompt to generate an image that reflects the suggested styling items. The input data is a list of suggested styling items and the prompt, and the output is the generated image. An example of a prompt is, "Please upload a photo of the user and determine the color attribute 'cool' and bone type 'straight'. Then, generate an image that includes a light blue blouse, natural makeup, and clear blue colored contact lenses."
[1463] Step 6:
[1464] The server displays the generated image to the user.
[1465] The server receives the generated image and stores it in association with the user data. It then transmits the analysis results, suggested styling items, and the generated image to the terminal. The input data is the user data associated with the generated image, and the output is the analysis results, suggested item list, and image displayed to the user.
[1466] Step 7:
[1467] The device will provide a link to purchase the suggested styling item.
[1468] The terminal displays a list of suggested styling items and a link to purchase them to the user. The user can click the link to purchase the suggested items from the online store. The input data is the list of suggested styling items, and the output is a user screen with the purchase link displayed.
[1469] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1470] The system of the present invention allows users to upload their own images, and the system analyzes the images to determine color attributes (cool or warm undertones) and bone structure (straight, wavy, natural), and then suggests the most suitable styling items. It also combines this with an emotion engine that recognizes the user's emotions to provide even more personalized styling suggestions. The system can also generate and display to the user an image of what the suggested styling items would look like when worn. It also provides a link to purchase the suggested styling items.
[1471] System Program Processing
[1472] 1. Upload a photo
[1473] A user accesses the system's website or application and logs in.
[1474] User clicks the image upload button, selects and uploads their own image.
[1475] The terminal transmits the selected image file to the server.
[1476] 2. Receiving and analyzing photos
[1477] The server receives the uploaded images. After receiving them, it checks the quality of the images. Specifically, it checks whether the face is clearly visible and whether the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[1478] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[1479] 3. Emotion Recognition by Emotion Engine
[1480] The server simultaneously uses an emotion engine to analyze emotions from the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[1481] 4. Styling suggestion generation
[1482] The server compares the results of the image analysis and emotion engine with a database to suggest styling items suitable for the user. The suggestions include the following items:
[1483] Clothing color and style
[1484] hair color
[1485] Makeup products
[1486] colored contact lenses
[1487] The server uses these suggestions to generate and store information to provide to the user, which is tailored to take into account feedback on the user's emotional state.
[1488] 5. Image Generation
[1489] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[1490] 6. Displaying results and making purchasing suggestions
[1491] The server sends the analysis results, proposals, and generated images to the terminal.
[1492] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1493] Analysis results of the user's personal color and bone structure type
[1494] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1495] Generated image
[1496] Links to purchase recommended items and access to our online store
[1497] Specific examples
[1498] 1. The user uploads their image to the system.
[1499] The device transfers the photo to the server.
[1500] 2. The server uses an AI model and emotion engine to analyze the photo, determine the user's color attributes as "spring warm skin tone" and body type as "wave," and further recognize the emotion of "joy" from the user's facial expression.
[1501] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[1502] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[1503] The terminal displays this to the user.
[1504] 4. The device displays links to where users can purchase items to achieve the suggested styling.
[1505] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[1506] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[1507] The processing flow will be explained below.
[1508] Step 1:
[1509] A user accesses the system's website or application and logs in. Logging in can be done using an email address or a social networking account.
[1510] Step 2:
[1511] The user clicks the image upload button, selects and uploads their own image, and the device sends the selected image file to the server.
[1512] Step 3:
[1513] The server receives the uploaded images. After receiving them, it checks the quality of the images to ensure that the face is clearly visible and that the brightness and resolution are sufficient. Only images that pass the quality check are allowed to proceed to the next analysis step.
[1514] Step 4:
[1515] The server inputs images that pass the quality check into the AI model for image analysis. The AI model extracts facial features and analyzes skin tone, hair color, and bone structure. This determines the user's color attributes (cool or warm skin tone) and bone structure type (straight, wavy, natural).
[1516] Step 5:
[1517] At the same time, the server uses an emotion engine to analyze the uploaded image and the user's facial expression. The emotion engine analyzes the user's facial expression, eye movements, mouth, and other characteristics to determine emotions such as "happiness," "sadness," "surprise," and "anger."
[1518] Step 6:
[1519] The server compares the results of image analysis and emotion recognition with a database to suggest the most suitable styling items for the user. Specifically, it selects clothing colors and styles, hair colors, makeup products, colored contact lenses, etc. that correspond to the user's color attributes and body type. The suggestions are adjusted taking into account the user's emotional state.
[1520] Step 7:
[1521] The server requests the generative AI to generate an image that reflects the suggested styling items. The generative AI model then creates a new image for the user that incorporates each of the suggested elements (clothes, hair color, makeup, and colored contact lenses).
[1522] Step 8:
[1523] The server receives the generated image, associates it with the user data, and saves it. It also compiles information to be provided to the user along with the proposal.
[1524] Step 9:
[1525] The server sends the analysis results, suggestions, and generated images to the device, along with a link to purchase the suggested styling items.
[1526] Step 10:
[1527] The terminal displays the received information on the user's screen. The displayed content is as follows:
[1528] Analysis results of the user's personal color and bone structure type
[1529] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1530] Generated image
[1531] Links to purchase recommended items and access to our online store
[1532] Step 11:
[1533] The user can check the displayed information and purchase the items they like from the online store by clicking the purchase link.
[1534] Specific examples
[1535] 1. The user uploads their image to the system.
[1536] The device transfers the photo to the server.
[1537] 2. The server analyzes the photo using an AI model to determine the attributes of "spring warm skin tone" and "wave skin tone." The emotion engine then recognizes "joy" from the facial expressions in the image.
[1538] Based on this, the server will suggest the user's best clothing color, "soft coral," makeup product, "fresh peach-toned blush," and colored contact lenses, "warm brown." The suggestions are adjusted to emphasize the user's emotion of "joy."
[1539] 3. The server uses generative AI to create an image of the user incorporating the suggested styling and sends it to the device.
[1540] The terminal displays this to the user.
[1541] 4. The device displays a link to purchase items to help the user achieve the suggested styling.
[1542] Users can click on the link to purchase the suggested clothing, makeup products, or colored contact lenses from the online store.
[1543] In this way, the system of the present invention allows users to efficiently discover their own personal style, provides styling suggestions that take into account their emotions at the time, and allows them to easily purchase specific styling items.
[1544] Example 2
[1545] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1546] Conventional styling suggestion systems can analyze a user's personal color and bone structure to suggest suitable styling items, but they cannot make suggestions that take into account the user's emotional state. As a result, the suggested styling items may not match the user's current emotions, making it difficult to increase user satisfaction. Another issue is that it is difficult for users to specifically imagine the suggested styling, which can reduce their motivation to purchase.
[1547] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1548] In this invention, the server includes a means for users to upload their own images, a means for analyzing the uploaded images and determining the user's color attributes and body type, a means for analyzing the user's emotions based on the uploaded images, a means for suggesting styling items suitable for the user based on the results of the determination and emotion analysis, a means for generating an image of the user incorporating the suggested styling items, and a means for displaying the suggested content and the generated image to the user. This enables personalized styling suggestions that take into account not only the user's personal color and body type but also their emotional state, thereby increasing user satisfaction. Furthermore, the generated image provides a concrete styling image, thereby increasing the user's desire to purchase.
[1549] "User" refers to a person who uses the system to upload their own image and receive styling suggestions.
[1550] "Image" refers to photographic data uploaded by a user to the system.
[1551] "Color attribute" refers to a personal color (e.g., cool-toned, warm-toned) that is classified based on the user's skin tone and hair color.
[1552] "Body type" refers to a styling type (e.g., straight, wavy, natural) that is classified based on the user's body shape and skeletal characteristics.
[1553] "Emotion" refers to the psychological state (e.g., joy, sadness, surprise, anger) that can be read from the user's facial expression.
[1554] "Styling items" refer to fashion-related items such as clothes, hair color, makeup products, and colored contact lenses that are recommended to users.
[1555] "Image" refers to image data used to generate a new visual for the user that reflects the suggested styling item.
[1556] "Server" refers to the central computing resource that performs the primary data processing and management of the system.
[1557] "Terminal" refers to a device such as a computer or smartphone that a user uses to access the system.
[1558] "Quality check" refers to the inspection process to ensure that uploaded images are suitable for analysis.
[1559] "Analysis" refers to the process of extracting features and recognizing emotions from uploaded images using AI models and emotion engines.
[1560] "Suggestion" refers to the act of providing the user with the most suitable styling items based on the analysis results.
[1561] "Generative AI" refers to artificial intelligence that generates image images that reflect suggested styling items in the user's image.
[1562] A "prompt sentence" refers to an instruction sentence that instructs the generation AI to generate a specific image.
[1563] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and then combines them with an emotion engine to provide personalized styling suggestions. The system can generate and display to the user an image of what the suggested styling items would look like. It also provides a link to purchase the suggested styling items.
[1564] The system uses the following major hardware and software:
[1565] 1. Hardware
[1566] Device: The computer or smartphone that a user uses to access the system.
[1567] Server: The central computing resource that handles the main data processing and management of the system.
[1568] 2. Software
[1569] Website or application: An interface where users can upload images and view analysis results and recommendations
[1570] AI model: a model for image analysis (e.g., a TensorFlow-based model)
[1571] Emotion engine: API for emotion recognition (e.g., Face API by Microsoft)
[1572] Generative AI: A model for generating images that reflect suggested styling items (e.g., DALL-E by OpenAI)
[1573] A specific embodiment of the system will be described below.
[1574] Uploading and receiving photos
[1575] Users access the system's website or application, log in, and then click the image upload button to select and upload their own image. At this time, the device sends the selected image file to the server. The image file is securely transmitted using HTTPS.
[1576] Image quality check and analysis
[1577] The server performs a quality check on the received images to determine whether they are suitable for analysis. For example, it uses OpenCV to check whether the face in the image is clearly visible and whether the resolution and brightness are appropriate. Images that pass the quality check are then analyzed using an AI model (e.g., TensorFlow). The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color attribute (e.g., cool-toned, warm-toned) and bone type (e.g., straight, wavy, natural).
[1578] emotion recognition
[1579] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze the user's emotions from the image. The emotion engine analyzes the user's facial expressions, eye movements, mouth features, etc., and recognizes emotions such as joy, sadness, surprise, and anger. The analysis results and emotion analysis results are then stored in a database.
[1580] Styling suggestions
[1581] The server compares the results of image analysis and emotion analysis with the system's styling database. For example, for a user with "spring warm skin tone," "waves," and the emotion "joy," it would suggest "soft coral" clothing and "fresh peach-toned blush." The server generates suggestions and saves a list of styling items suitable for the user in JSON format.
[1582] Image generation
[1583] The server requests the generation AI (e.g., DALL-E by OpenAI) to generate an image that reflects the suggested styling items. An example of a prompt to be input to the generation AI is "An image of a warm-toned, spring-season woman wearing soft coral clothing." The generation AI model creates a new image for the user that reflects the suggestions. The server retrieves the generated image and saves it by associating it with the user's profile.
[1584] Viewing and purchasing results
[1585] The server sends a dataset containing the analysis results, suggestions, and generated images to the device. The device displays the received information on the user's screen. The user sees:
[1586] Analysis results of the user's personal color and bone structure type
[1587] A list of recommended clothes, hair colors, makeup products, and colored contact lenses
[1588] Generated image
[1589] Links to purchase recommended items and access to our online store
[1590] Specific examples
[1591] For example, a user uploads an image of themselves to the system. The device sends the uploaded image to the server, which analyzes the image using an AI model and emotion engine. If the analysis results indicate that the user's color attributes are "spring warm skin," their bone structure is "wave," and their emotion is "joy," the server will suggest "soft coral" clothing, "fresh peach-toned blush," and "warm brown" colored contact lenses as the optimal styling for the user. Using generative AI, a new image of the user incorporating these styling items is created and displayed to the user. The user can then purchase their favorite products from the online store.
[1592] In this way, the system of the present invention efficiently discovers the user's personal style, makes styling suggestions that take emotions into consideration, and allows the user to easily purchase specific styling items.
[1593] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1594] Step 1:
[1595] A user accesses a website or application of the system and logs in.
[1596] Input: User login information (email address, password)
[1597] How it works: A user enters their email address and password on the login screen and clicks the login button. The system authenticates the user and starts a login session.
[1598] Output: Successful login notification, user's home screen
[1599] Step 2:
[1600] User clicks the image upload button, selects and uploads their own image.
[1601] Input: User's image file
[1602] How it works: The user clicks the image upload button on the main screen. A file selection dialog opens and the user selects their image file. The device sends the selected image file to the server using an HTTP POST request.
[1603] Output: Notification of image file transfer to server
[1604] Step 3:
[1605] The server receives the uploaded images and performs a quality check.
[1606] Input: Uploaded image file
[1607] How it works: The server receives image files and performs a quality check on the images using a facial recognition algorithm (e.g., OpenCV). Check items include facial clarity, resolution, brightness, etc. Only image files that pass the quality check proceed to the next analysis step.
[1608] Output: Image files that pass the quality check or notification of poor quality
[1609] Step 4:
[1610] The server inputs images that pass the quality check into an AI model (e.g., TensorFlow) and performs image analysis.
[1611] Input: Image files that have passed the quality check
[1612] How it works: The AI model extracts facial features and analyzes skin tone, hair color, and bone shape to determine the user's color profile (cool, warm) and bone type (straight, wavy, natural).
[1613] Output: Analysis results of color attributes and skeletal type
[1614] Step 5:
[1615] At the same time, the server uses an emotion engine (e.g., Face API by Microsoft) to analyze emotions from the uploaded image and the user's facial expressions.
[1616] Input: Uploaded image file
[1617] Behavior: The emotion engine analyzes the user's facial expressions, eye movements, mouth, and other characteristics to recognize emotions such as joy, sadness, surprise, and anger.
[1618] Output: Emotion analysis results
[1619] Step 6:
[1620] The server suggests styling items suitable for the user based on the results of image analysis and the emotion engine.
[1621] Input: Color attribute and skeletal type analysis results, emotion analysis results
[1622] How it works: The system compares the styling database within the system and selects the styling items that are best suited to the user (e.g., "soft coral" clothing, "fresh peach-toned blush," etc.).
[1623] Output: A list of suggested styling items
[1624] Step 7:
[1625] The server uses a generation AI (e.g., DALL-E) to generate an image that reflects the suggested styling items.
[1626] Input: A list of suggested styling items
[1627] How it works: The user provides the AI with a prompt (e.g., "Image of a warm-toned spring woman wearing a soft coral dress") and requests that it generate an image. The AI then creates a new image for the user that reflects the suggestions.
[1628] Output: Generated image
[1629] Step 8:
[1630] The server sends the analysis results, proposals, and generated images to the terminal.
[1631] Input: Analysis results, proposals, generated images
[1632] How it works: The server sends these datasets in JSON format to the device.
[1633] Output: Notification of completion of dataset transmission to the terminal
[1634] Step 9:
[1635] The terminal displays the received information on the user screen.
[1636] Input: Analysis results, proposals, and generated images sent in JSON format
[1637] How it works: The device analyzes the received data and dynamically renders a user screen using a front-end framework such as Vue.js or React.js. The following information is displayed on the user screen: analysis results, recommendations, generated images, links to purchase recommended items, and links to access the online store.
[1638] Output: Styling suggestions displayed on the user's screen
[1639] Step 10:
[1640] The user clicks on the link for the item of interest and is taken to the specified online store page.
[1641] Input: User clicks
[1642] What it does: When the user clicks, the browser navigates to the product page in the specified online store.
[1643] Output: Online store product page
[1644] (Application example 2)
[1645] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1646] To quickly and accurately respond to today's diversifying customer needs, personalized styling suggestions are required. However, while conventional systems take into account a user's personal color and bone structure, they do not provide personalized styling that takes into account emotional fluctuations. Furthermore, there is a lack of methods for providing real-time suggestions in-store or instantly providing visually easy-to-understand images. The present invention aims to solve these problems and improve the customer experience.
[1647] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to upload their own image, means for analyzing the uploaded image and determining the user's color attributes and body type, means for suggesting styling items suitable for the user based on the determination results, means for generating an image of the user when incorporating the suggested styling items, means for using an emotion engine that recognizes emotions from the user's facial expression, means for adjusting the suggestions taking into account the emotion recognition results, and means for displaying the suggestions and the generated image to the user. This enables the user to receive personalized styling suggestions in real time according to their emotional state.
[1648] A "user" is an individual or customer who uses the system.
[1649] An "image uploading means" is a method or device by which a user can send their image to the system.
[1650] "Means for analyzing images" refers to a method or device for identifying user information (color attributes and bone structure type) based on uploaded images.
[1651] "Color attribute" indicates the personal color (cool or warm) based on the user's skin tone.
[1652] "Body type" indicates a classification based on the shape of the user's body frame (straight, wavy, natural).
[1653] The "means for suggesting styling items" is a method or device for suggesting clothes, accessories, makeup products, etc. that are suitable for the user.
[1654] The "means for generating an image" is a method or device for creating a new image of the user incorporating the suggested styling item.
[1655] An "emotion engine" is a program or device that recognizes emotions (happiness, sadness, surprise, anger, etc.) from a user's facial expression.
[1656] The "means for adjusting the suggestions taking into account the emotion recognition results" is a method or apparatus for modifying or adjusting styling suggestions based on the emotions recognized by the emotion engine.
[1657] The "means for displaying the generated image" is a method or device for visually presenting new styling suggestions to the user.
[1658] The system of the present invention allows users to upload their own images, analyzes the images to determine color attributes and body type, and uses an emotion engine to recognize emotions from the user's facial expressions. Based on the results of these analyses, the system suggests optimal styling items for the user and generates and displays images using the suggested items. The system aims to visually present the suggestions and the generated images to the user. Links to purchase the suggested styling items are also provided.
[1659] System hardware and software configuration
[1660] This system uses the following hardware and software:
[1661] Hardware:
[1662] Tablet device (e.g. iPad)
[1663] Smart mirrors (e.g., Novera Smart Mirror)
[1664] software:
[1665] AI image analysis engine (e.g. Google Cloud Vision API)
[1666] Emotion recognition engine (e.g. Microsoft Azure Face API)
[1667] Generative AI models (e.g., OpenAI DALL-E)
[1668] Data processing and calculation flow
[1669] Upload a photo
[1670] Users access the system from a tablet or smart mirror and upload their own images, which are then sent to the server.
[1671] Image reception and analysis
[1672] The server receives the uploaded images and performs a quality check. Images that pass the quality check are input into an AI image analysis engine (Google Cloud Vision API) to determine personal color and bone structure type.
[1673] Emotion recognition by emotion engine
[1674] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to analyze the user's facial expressions and recognize emotions such as joy, sadness, surprise, and anger.
[1675] Generate styling suggestions
[1676] The server then uses the results of image analysis and emotion recognition to suggest styling items suitable for the user, including clothing color and style, hair color, makeup products, and colored contact lenses, while also taking into account the user's emotional state.
[1677] Image generation
[1678] The server uses a generative AI model (OpenAI DALL-E) to generate new images for the user incorporating the suggested styling items.
[1679] Displaying results and making purchasing suggestions
[1680] The generated image and the proposed content are sent to the device, which then displays them on the user's screen. The displayed content includes the analysis results of personal color and bone structure type, recommended styling items, the generated image, and a link to purchase the items.
[1681] Specific examples
[1682] Below are some specific examples.
[1683] Example prompt sentence:
[1684] "Using Google Cloud Vision API and Microsoft Azure Face API, we analyze the customer's facial image to determine their personal color. We then analyze their facial expressions to recognize their emotions. Based on the results, we provide fashion styling and product suggestions, and generate new images."
[1685] Specifically, when a user uploads a photo of themselves using a tablet or smart mirror, the server analyzes the photo and determines the user's color attributes as "cool winter" and bone type as "straight," and then uses emotion recognition to identify the emotion of "surprise." Based on this, the server suggests a "sharp black jacket" and "cool-toned makeup products" to the user, and uses a generative AI model to create an image that reflects these suggestions and displays it on the device. The user can also receive a link to purchase the items from an online store based on the suggestions.
[1686] In this way, the system of the present invention allows users to easily receive personalized styling suggestions based on their individual characteristics and emotions, and by visually checking the generated image, the suggestions can be easily understood.
[1687] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1688] Step 1:
[1689] A user uploads their own image using a tablet device or smart mirror. The user opens a dedicated application, clicks the image upload button, and selects their own image file. The input is the user's image file, and the output is that the device sends this image to the server.
[1690] Step 2:
[1691] The server receives the uploaded image. The server performs a quality check on the image to ensure that the face is clearly visible, and that the brightness and resolution are sufficient. Only images that pass this quality check proceed to the next step. The input is the user's image file, and the output is an image that passes the quality check.
[1692] Step 3:
[1693] The server inputs images that pass the quality check into an AI image analysis engine (Google Cloud Vision API) to determine the user's color attributes and bone structure. During this process, the AI analyzes facial features and identifies skin tone, hair color, and bone structure. The input is an image that passes the quality check, and the output is the user's color attributes and bone structure data.
[1694] Step 4:
[1695] At the same time, the server uses an emotion recognition engine (Microsoft Azure Face API) to recognize emotions from the user's facial expressions. This process detects subtle facial features and identifies emotions such as joy, sadness, surprise, and anger. The input is a user image that has passed quality checks, and the output is the user's emotional data.
[1696] Step 5:
[1697] The server compares the results of image analysis and emotion recognition with a database of styling item suggestions. It then suggests the most suitable clothing, hair color, makeup products, and colored contact lenses for the user. The suggestions are adjusted based on the emotion data. The input is the user's color attributes, bone structure, and emotion data, and the output is a list of suggested styling items.
[1698] Step 6:
[1699] The server uses a generative AI model (OpenAI DALL-E) to generate an image that reflects the suggested styling items. In this process, a new image is created by combining styling items with the user's facial photo. The input is a list of suggested styling items and the user's image, and the output is the generated image.
[1700] Step 7:
[1701] The server sends the generated image and proposal details to the terminal, which then displays them on the user's screen. The displayed content includes the analysis results of the personal color and body type, recommended styling items, the generated image, and a link to purchase the items. The input is the generated image and proposal details, and the output is the specific styling proposal information displayed to the user.
[1702] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1703] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1704] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1705] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1706] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1707] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1708] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1709] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1710] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1711] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1712] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1713] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1714] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1715] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1716] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1717] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1718] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1719] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1720] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1721] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1722] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1723] The following is further disclosed regarding the above embodiment.
[1724] (Claim 1)
[1725] a means for users to upload images of themselves;
[1726] means for analyzing the uploaded image and determining the user's color attribute and bone structure type;
[1727] A means for suggesting styling items suitable for the user based on the determination result;
[1728] A means for generating an image of the user when incorporating the suggested styling item;
[1729] The system includes means for displaying the suggestions and the generated images to the user.
[1730] (Claim 2)
[1731] 10. The system of claim 1, further comprising: means for providing a purchase link for the suggested styling item.
[1732] (Claim 3)
[1733] 10. The system of claim 1, further comprising means for performing a quality check on the images and selecting images suitable for analysis.
[1734] "Example 1"
[1735] (Claim 1)
[1736] a means for users to upload images of themselves;
[1737] means for analyzing the uploaded image and determining the user's color attribute and bone structure type;
[1738] A means for suggesting styling items suitable for the user based on the determination result;
[1739] A means for generating an image of the user when incorporating the suggested styling item;
[1740] means for displaying the proposal and the generated image to the user;
[1741] A means of checking the quality of images and selecting images suitable for analysis;
[1742] a means for inputting a prompt sentence into a generative AI model for generating an image image reflecting the suggested styling item;
[1743] The system includes means for receiving the generated image and storing it in association with user data.
[1744] (Claim 2)
[1745] 10. The system of claim 1, further comprising: means for providing a purchase link for the suggested styling item.
[1746] (Claim 3)
[1747] 10. The system of claim 1, further comprising quality check means for checking the resolution and brightness of the uploaded image.
[1748] "Application Example 1"
[1749] (Claim 1)
[1750] a means for users to upload images of themselves;
[1751] means for analyzing the uploaded image and determining the user's color attribute and bone structure type;
[1752] A means for suggesting styling items suitable for the user based on the determination result;
[1753] A generative AI model means for generating an image of the user when incorporating the suggested styling items;
[1754] means for displaying the proposal and the generated image to the user;
[1755] The system includes a means for providing a user with a link to purchase items for realizing the suggested styling.
[1756] (Claim 2)
[1757] 10. The system of claim 1, further comprising means for performing a quality check on the images and selecting images suitable for analysis.
[1758] (Claim 3)
[1759] The system according to claim 1, further comprising means for generating an image that reflects the suggested styling items using a prompt sentence for the generation AI model.
[1760] "Example 2: Combining Emotion Engines"
[1761] (Claim 1)
[1762] a means for users to upload images of themselves;
[1763] means for analyzing the uploaded image and determining the user's color attribute and bone structure type;
[1764] means for analyzing a user's emotions based on the uploaded images;
[1765] A means for suggesting styling items suitable for a user based on the judgment result and the emotion analysis result;
[1766] A means for generating an image of the user when incorporating the suggested styling item;
[1767] A means for displaying the proposal and generated image to the user
[1768] A system including:
[1769] (Claim 2)
[1770] 10. The system of claim 1, further comprising: means for providing a purchase link for the suggested styling item.
[1771] (Claim 3)
[1772] 10. The system of claim 1, further comprising means for performing a quality check on the images and selecting images suitable for analysis.
[1773] "Application example 2 when combining emotion engines"
[1774] (Claim 1)
[1775] a means for users to upload images of themselves;
[1776] means for analyzing the uploaded image and determining the user's color attribute and bone structure type;
[1777] A means for suggesting styling items suitable for the user based on the determination result;
[1778] A means for generating an image of the user when incorporating the suggested styling item;
[1779] means for using an emotion engine to recognize emotions from the user's facial expressions;
[1780] a means for adjusting the content of the suggestions in consideration of the emotion recognition results;
[1781] The system includes means for displaying the suggestions and the generated images to the user.
[1782] (Claim 2)
[1783] 10. The system of claim 1, further comprising: means for providing a purchase link for the suggested styling item.
[1784] (Claim 3)
[1785] 10. The system of claim 1, further comprising means for performing a quality check on the images and selecting images suitable for analysis. [Explanation of symbols]
[1786] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for users to upload images of themselves; means for analyzing the uploaded image and determining the user's color attribute and bone structure type; A means for suggesting styling items suitable for the user based on the determination result; A means for generating an image of the user when incorporating the suggested styling item; The system includes means for displaying the suggestions and the generated images to the user.
2. The system of claim 1 , further comprising means for providing a purchase link for the suggested styling item.
3. 10. The system of claim 1, further comprising means for performing a quality check of the images and selecting images suitable for analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A