system

The system addresses the challenge of selecting suitable clothes in online shopping by generating personalized clothing images based on user input and feedback, improving user satisfaction.

JP2026014901APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116375
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Traditional online shopping systems make it difficult for users to easily choose clothes that suit their style and body type, leading to lower user satisfaction and a decline in online shopping due to the lack of personalized feedback in clothing selections.

Method used

A system that generates clothing images using an image generation API based on user-entered style and body type information, allows for user feedback, and reflects this feedback in subsequent image generations to provide more personalized suggestions.

Benefits of technology

Improves the online shopping experience by ensuring users can easily find clothes that fit their style and body type, enhancing user satisfaction through personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014901000001_ABST
    Figure 2026014901000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving information on a style and a body type input by a user; means for generating a clothing image by calling an image generation API based on the information on the style and the body type; means for transmitting the generated clothing image to a terminal of the user; and means for receiving feedback from the user and reflecting the feedback in generation of a next clothing image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional online shopping systems make it difficult for users to easily choose clothes that suit their style and body type. Users cannot make clothing selections based on feedback, which increases the risk of purchasing clothes that do not fit them. These issues can lead to lower user satisfaction and a decline in online shopping. [Means for solving the problem]

[0005] The present invention is a system that generates clothing images using an image generation API based on style and body type information entered by the user and transmits the generated clothing images to the user's terminal. It also includes a means for receiving feedback from the user, and by reflecting this feedback in the next clothing image generation, it is possible to provide suggestions that are more suited to the user's individual needs. The present invention also includes a means for formatting the style and body type information and converting it into a format suitable for the image generation API, thereby achieving more accurate image generation. It also includes a means for providing a rating and feedback interface along with the generated clothing images, allowing users to easily provide feedback. This can improve user satisfaction and promote online shopping.

[0006] "User" refers to a person who uses the system to select clothing based on style and body type.

[0007] "Style" refers to the type of clothing a user prefers and fashion trends.

[0008] "Body type" refers to the user's body shape and characteristics, typically including height, weight, and body characteristics.

[0009] "Image Generation API" refers to an application programming interface that generates images based on input text or data.

[0010] "Clothing Image" refers to a visual representation of clothing based on the User's style and body type generated by the image generation API.

[0011] "Terminal" refers to the device used by a user to input information, display images, and provide feedback.

[0012] "Feedback" refers to the user's opinions and evaluations of the generated clothing images.

[0013] The term "system" refers to a collection of hardware and software for realizing a series of functions of the present invention. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information on their style and body type. The system includes a server, a terminal, and an image generation API.

[0036] Entering user information

[0037] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0038] Sending input information to the server

[0039] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0040] json

[0041] {

[0042] "style": "casual",

[0043] "height": 170,

[0044] "weight": 65,

[0045] "body_shape": "slim"

[0046] }

[0047] DALL-E API call and image generation

[0048] The server calls the image generation API based on the received user information, converting the information into an appropriate format and using it as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit."

[0049] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0050] Sending generated images from the server to the device

[0051] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0052] Displaying images and collecting user feedback

[0053] The terminal displays the received clothing image to the user, and further provides the user with an interface for inputting ratings and feedback, allowing the user to provide feedback on their opinions and ratings of the displayed outfit image.

[0054] Sending and using feedback to the server

[0055] Feedback from users is sent to the server via their devices. The server analyzes this feedback and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0056] Specific examples

[0057] For example, if a user prefers casual styles and has a slim build, the user enters their height of 170cm, weight of 65kg, and style of "casual" on the device. The device sends this information to the server, which then generates a prompt for "casual outfit for a 170cm, 65kg, slim build" and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0058] This allows users to easily find the clothes that suit them best, greatly improving their online shopping experience.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0062] Step 2:

[0063] The device compiles the user's input information into JSON format. For example:

[0064] json

[0065] {

[0066] "style": "casual",

[0067] "height": 170,

[0068] "weight": 65,

[0069] "body_shape": "slim"

[0070] }

[0071] Step 3:

[0072] The device sends the user information compiled in JSON format to the server.

[0073] Step 4:

[0074] The server analyzes the received user information and generates an appropriate prompt. Example: "170cm, 65kg, slim build, casual outfit."

[0075] Step 5:

[0076] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0077] Step 6:

[0078] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0079] Step 7:

[0080] The server sends the clothing image received from the image generation API to the user's device.

[0081] Step 8:

[0082] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0083] Step 9:

[0084] Users can input feedback about the outfit images displayed, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[0085] Step 10:

[0086] The device sends the user's feedback to the server.

[0087] Step 11:

[0088] The server receives user feedback and stores it in a database, which is then reflected in the next image generation.

[0089] This allows users to easily find the clothes that best suit their style and body type, enhancing their online shopping experience.

[0090] Example 1

[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0092] Traditional online shopping has the problem of making it difficult for users to find clothes that perfectly fit their style and body type. This results in users wasting time and effort, and they often fail to find the products they want. Another issue is that if the generated clothing image does not meet the user's expectations, feedback is not reflected in the next image generation, resulting in the same problem repeating itself.

[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0094] In this invention, the server includes means for accepting style and body type information entered by the user, means for generating a prompt based on the style and body type information, calling an image generation API, and generating a clothing image, and means for sending the generated clothing image to the user's device. This allows the user to efficiently obtain clothing images that suit their style and body type. Furthermore, by including means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, more appropriate suggestions that reflect the user's preferences and requests can be made.

[0095] A "user" is a person who uses the system to input style and body type information and view the generated clothing images.

[0096] "Style" refers to fashion categories based on the user's preferences, such as casual, formal, and sporty.

[0097] "Body type" refers to information about the user's physical appearance, such as their height, weight, and physical characteristics.

[0098] An "image generation API" is an application programming interface that generates an image based on a specified prompt.

[0099] A "prompt sentence" is a text sentence that issues a specific request to the image generation API.

[0100] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[0101] A "terminal" is an electronic device such as a computer or smartphone that a user uses to input information and receive and display generated clothing images.

[0102] "Feedback" refers to the evaluation or opinion provided by the user regarding the generated clothing image, which is reflected in the next image generation.

[0103] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. The system includes a server, a terminal, and an image generation API.

[0104] Users input information about their style and body type using their device. Specifically, style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device compiles this information in a predefined format and sends it to the server in JSON format.

[0105] The server generates a prompt based on the received user information and calls an image generation API (e.g., an API using an image generation model). In this case, a specific prompt such as "170 cm, 65 kg, slim build, casual outfit" is used. The server then sends a request including the generated prompt to the DALL-E API, requesting image generation.

[0106] The image generation API generates a clothing image based on the specified prompt and returns the result to the server. The server receives the clothing image returned from the image generation API and sends it to the user's device as an HTTP response.

[0107] The device displays the received clothing image on the screen and provides the user with an interface for inputting ratings and feedback. The user inputs their opinions and ratings about the displayed outfit image as feedback and presses the send button. Feedback can include specific comments such as "I wish the pants were a lighter color."

[0108] The device converts the user's feedback information into JSON format and sends it to the server as an HTTP POST request. The server then analyzes the feedback and stores it in a database. The saved feedback information is reflected in the next image generation, making suggestions that match the user's preferences.

[0109] For example, if a user likes casual style and has a slim build, the user can enter their height (170cm), weight (65kg), and style ("casual") into the terminal. When the user presses the "Submit" button, the terminal will generate the following prompt:

[0110] "Generate a casual outfit for a slim person who is 170cm tall and weighs 65kg."

[0111] This prompt is sent to the DALL-E API, and the generated image is sent to the device via the server. The user can review the generated image and provide feedback to reflect it in the next image generation.

[0112] In this way, users can easily find the clothes that suit them best, greatly improving their online shopping experience.

[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0114] Step 1:

[0115] The user uses the device to input information about their own style and body type. Specifically, they input their style (e.g., casual, formal), height, weight, and body type characteristics into a form displayed on the device screen. The user's input is then acquired by the device.

[0116] Step 2:

[0117] The device converts the style and body information into a predefined format (JSON). Specifically, it formats the input information as follows:

[0118] json

[0119] {

[0120] "style": "casual",

[0121] "height": 170,

[0122] "weight": 65,

[0123] "body_shape": "slim"

[0124] }

[0125] The input here is the user's style and body type information, and the output is JSON format data.

[0126] Step 3:

[0127] The terminal sends the generated JSON data to the server. Specifically, it sends data to the server as an HTTP POST request. The input is JSON data, and the output is a message to the server indicating that data was successfully sent.

[0128] Step 4:

[0129] The server generates a prompt based on the received user information. Specifically, it uses a template in the server script to create a prompt like this:

[0130] "Casual outfit for a slim figure, 170cm, 65kg"

[0131] The input is the user's JSON data and the output is the prompt text.

[0132] Step 5:

[0133] The server sends the generated prompt text to the image generation API. Specifically, it sends the prompt text to the DALL-E API as an HTTP POST request. The input is the prompt text, and the output is a request sending success message to the image generation API.

[0134] Step 6:

[0135] The image generation API generates a clothing image based on the prompt received from the server and sends the result back to the server. Specifically, it generates image data using a generative AI model. The input is the prompt, and the output is the generated clothing image.

[0136] Step 7:

[0137] The server receives the clothing image returned from the image generation API and transfers it to the user's device. Specifically, it sends the image data to the user's device as an HTTP response. The input is the generated clothing image, and the output is a message that the image data was successfully sent to the user's device.

[0138] Step 8:

[0139] The device displays the received clothing image on the screen. Specifically, it renders an image display UI to allow the user to check the outfit image. The input is the generated clothing image, and the output is the display of the image.

[0140] Step 9:

[0141] The user provides feedback on the displayed outfit image. Specifically, the user inputs their opinion or rating into the rating and feedback input interface and presses the send button. The input is the feedback content, and the output is the feedback sending action.

[0142] Step 10:

[0143] The device converts the feedback information provided by the user into JSON format and sends it to the server. Specifically, it sends the feedback data as an HTTP POST request. The input is the feedback content, and the output is a feedback transmission success message to the server.

[0144] Step 11:

[0145] The server analyzes the feedback and stores it in a database. Specifically, it analyzes the feedback data and stores it as user preference data. The input is the feedback data, and the output is a message indicating that the analyzed data has been saved in the database.

[0146] Step 12:

[0147] The server reflects the feedback information when generating the next image. Specifically, it improves the next prompt generation based on the feedback stored in the database. The input is the stored feedback data, and the output is the improved prompt.

[0148] This series of processes provides the user with clothing images that are optimally suited to their style and body type, and by incorporating user feedback into future recommendations, the system delivers a more accurate and personalized online shopping experience.

[0149] (Application example 1)

[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0151] It is difficult to find the perfect clothes when shopping online. In particular, the process of selecting clothes that fit the user's body type and style is complicated, resulting in an unsatisfactory shopping experience. Another issue is that personalization based on user feedback is not fully implemented, making it difficult to provide suggestions that match the user's preferences.

[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0153] In this invention, the server includes means for receiving style and body type information input by the user, means for generating a prompt for the generation AI model based on the style and body type information and for generating clothing images by calling an image generation API, means for sending the generated clothing images to the user's device, and means for receiving user feedback and reflecting the feedback in generating the next clothing image. This allows the user to easily find clothing that suits them and receive even more personalized suggestions through the feedback.

[0154] "User" refers to an individual who uses this system to input their own style and body type information and receive suggestions for suitable clothing images.

[0155] "Style" refers to categories that indicate the type of clothing and fashion trends that users prefer, such as casual, formal, or sporty.

[0156] "Body type" refers to the user's physical characteristics and includes information about height, weight, and body type such as slim or chubby.

[0157] A "generative AI model" refers to an artificial intelligence model that has the technology to generate clothing images based on prompt text provided by the user.

[0158] A "prompt sentence" refers to a text request sentence that is generated based on the user's body type and style information and sent to the image generation API.

[0159] "Image generation API" refers to a programming interface for generating an image that meets specified conditions based on the received prompt text.

[0160] "Clothing Images" refers to images of clothing generated based on the user's body type and style using a generative AI model.

[0161] "Feedback" refers to the user inputting their opinions and evaluations of the generated clothing images, and refers to information that will be used to improve the next proposal.

[0162] "Terminal" refers to a device used by a user to input information and receive generated clothing images.

[0163] "Server" refers to the central computer system that receives user input information, calls the image generation API to generate garment images, and manages the collected feedback.

[0164] This invention provides a mail-order system that helps users find clothes that suit their body type and style when shopping online. The system includes a server, a terminal, and an image generation API.

[0165] First, the user uses the device to input their body type and style information, including height, weight, body characteristics, preferred style (e.g., casual, formal, sporty), etc. The device then formats the input information into a predefined format and sends it to the server.

[0166] Based on the received information, the server generates an appropriate prompt for the generative AI model. For example, a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is sent to the image generation API, which generates a clothing image based on the prompt. The software used for this is an HTTP client such as axios and the DALL-E API.

[0167] The generated clothing image is sent from the server to the user's device, where the user can view the image. In addition, an interface for inputting evaluations and feedback is provided to the user. The user can input their opinions and evaluations of the generated clothing image. For example, they can provide feedback such as "I would like the color of the pants to be lighter." This feedback information is sent to the server and reflected the next time an image is generated.

[0168] A concrete example is as follows: If a user prefers a casual style and has a slim build, the user enters their height as 170cm, weight as 65kg, and style as "casual" on their device. The device sends this information to the server, which then generates a prompt message saying "Casual outfit for a 170cm, 65kg, slim build" and sends it to the image generation API. The generated image is then sent to the device via the server, allowing the user to check the suitable outfit. At this time, the user enters feedback such as "I'd like the pants to be a lighter color," and the server takes this information into consideration the next time it generates a clothing image.

[0169] In this way, users can efficiently find the clothes that best suit their tastes and body type, and further personalization can be achieved through feedback, greatly improving the online shopping experience.

[0170] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0171] Step 1:

[0172] The user uses the device to enter their own style and body type information. Specifically, the user enters their height, weight, body type characteristics, and preferred style (e.g., casual, formal, sporty) into the device's input form. The entered information is formatted into a JSON format defined by the system. The input here is the user's body type and style information, and the output is formatted JSON data.

[0173] Step 2:

[0174] The device sends the entered information to the server. The data sent is in JSON format, for example:

[0175] json

[0176] {

[0177] "style": "casual",

[0178] "height": 170,

[0179] "weight": 65,

[0180] "body_shape": "slim"

[0181] }

[0182] The server receives this data, where the input is formatted JSON data and the output is the received user information.

[0183] Step 3:

[0184] The server generates a prompt based on the information it receives. Specifically, it creates a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is the request sent to the generative AI model (image generation API). The input here is JSON data of the user information, and the output is the prompt.

[0185] Step 4:

[0186] The server uses the generated prompt text to call an image generation API (e.g., DALL-E API). Specifically, it sends an API request including the prompt text and receives the URL of the generated clothing image. The input here is the prompt text, and the output is the URL of the generated clothing image.

[0187] Step 5:

[0188] The server sends the URL of the clothing image received from the image generation API to the user's device. The user's device receives this URL and displays the image on the screen. The input here is the URL of the generated clothing image, and the output is the clothing image displayed on the user's device.

[0189] Step 6:

[0190] The user inputs their evaluation and feedback for the displayed clothing image. Specifically, the user inputs feedback such as "I would like the color of the pants to be lighter" into the evaluation form and sends it from the terminal to the server. The input here is the user's feedback, and the output is the feedback data sent to the server.

[0191] Step 7:

[0192] The server analyzes the received feedback and reflects it in the next prompt generation. Specifically, the feedback data is stored in a database and taken into consideration when generating the next prompt for personalization. The input here is the user's feedback data, and the output is an updated prompt generation algorithm.

[0193] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0194] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. In addition to the basic functions of inputting user information, calling an image generation API, displaying generated images, and collecting feedback, the system also incorporates an emotion engine that recognizes the user's emotions.

[0195] Entering user information

[0196] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0197] Sending input information to the server

[0198] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0199] json

[0200] {

[0201] "style": "casual",

[0202] "height": 170,

[0203] "weight": 65,

[0204] "body_shape": "slim"

[0205] }

[0206] Emotion Engine Data Collection and Analysis

[0207] The device is equipped with a camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine uses this data to analyze the user's emotions. For example, if the user is smiling, it is recognized as a positive emotion.

[0208] DALL-E API call and image generation

[0209] The server calls the image generation API based on the received user information and the analysis results from the emotion engine. At this time, the information is converted into an appropriate format and used as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[0210] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0211] Sending generated images from the server to the device

[0212] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0213] Displaying images and collecting user feedback

[0214] The device displays the received clothing images to the user. It also provides the user with an interface for inputting ratings and feedback. The emotion engine then analyzes the user's emotions and uses the results to improve the displayed content. The user can provide feedback on their opinions and ratings of the displayed outfit images.

[0215] Sending and using feedback to the server

[0216] User feedback is sent to the server via the device. The server analyzes this feedback along with emotional data and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0217] Specific examples

[0218] For example, if a user prefers casual styles and has a slim build, they can enter their height of 170cm, weight of 65kg, and style of "casual" on their device. At this time, the device captures the user's emotions, and the emotion engine determines that they are reacting positively. The device sends this information to the server, which then generates a prompt that reads, "Casual outfit for a 170cm, 65kg, slim build. User smiling," and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0219] This will allow users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of the emotion engine will enable more personalized suggestions based on the user's emotions, aiming to further increase user satisfaction.

[0220] The processing flow will be explained below.

[0221] Step 1:

[0222] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0223] Step 2:

[0224] The device formats the entered user information into JSON format. Example:

[0225] json

[0226] {

[0227] "style": "casual",

[0228] "height": 170,

[0229] "weight": 65,

[0230] "body_shape": "slim"

[0231] }

[0232] Step 3:

[0233] The terminal transmits the formatted user information to the server.

[0234] Step 4:

[0235] The device captures the user's facial expressions and voice, and collects emotional data using the device's built-in camera and microphone.

[0236] Step 5:

[0237] The emotion engine analyzes data collected by the device to determine the user's emotional state. For example, a smiling user is recognized as a positive emotion.

[0238] Step 6:

[0239] The server generates a prompt based on the received user information and emotion data. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[0240] Step 7:

[0241] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0242] Step 8:

[0243] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0244] Step 9:

[0245] The server sends the clothing image received from the image generation API to the user's device.

[0246] Step 10:

[0247] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0248] Step 11:

[0249] The device continues to use the emotion engine to monitor the user's emotional state, detecting the user's reaction to the displayed image (e.g., smiling, making a displeased face).

[0250] Step 12:

[0251] Users can input feedback on outfit images, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[0252] Step 13:

[0253] The device sends the user's feedback to the server.

[0254] Step 14:

[0255] The server stores the received feedback along with the emotion data and reflects it in the next image generation. The server updates the feedback database and uses it to generate the next prompt.

[0256] In this way, users can easily find the clothes that best suit their style and body type, improving their online shopping experience. In addition, the emotion engine can make suggestions based on the user's emotional state, further improving user satisfaction.

[0257] Example 2

[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0259] In today's online shopping environment, users often find it difficult to find clothing that suits their style and body type. Furthermore, the lack of technology that makes personalized recommendations based on user emotions makes it difficult to provide a satisfying shopping experience. To address this issue, a system is needed that not only generates clothing images based on a user's style and body type information, but also makes personalized recommendations that take the user's emotions into account.

[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0261] In this invention, the server includes means for accepting style and body type information input by the user, means for generating clothing images by invoking an image generation algorithm based on the style and body type information, and means for transmitting the generated clothing images to the user's information processing device. This makes it possible to provide accurate clothing images based on the user's style and body type, and to reflect feedback based on the results. Furthermore, by incorporating means for collecting emotional information using an emotion analysis engine that analyzes user emotions, and means for providing the emotional information along with the style and body type information as prompts to the generation AI model, personalized suggestions based on the user's emotions can be made, significantly improving the satisfaction of the online shopping experience.

[0262] "Style" is information that refers to the clothing category and design characteristics that a user prefers.

[0263] "Body type" refers to information that refers to the user's physical proportions, such as the user's height, weight, and body shape characteristics.

[0264] An "image generation algorithm" is a program or system that generates images of suitable clothing based on style and body type information provided by the user.

[0265] An "information processing device" is a terminal used by a user, specifically a device such as a smartphone or personal computer.

[0266] "Feedback" refers to information that indicates the evaluation or opinion that a user provides regarding the generated clothing image.

[0267] An "emotion analysis engine" is a program or system that collects and analyzes emotional information from a user's facial expressions and voice in real time.

[0268] "Emotional information" is data that indicates the emotional state of a user collected from their facial expressions and voice.

[0269] A "prompt" is input information given to a generative AI model, and refers to the form of data that includes style, body type, and emotional information.

[0270] A "generative AI model" is an artificial intelligence system that generates content, such as images, based on given prompts.

[0271] The virtual personal shopper system of the present invention aims to generate appropriate clothing images based on user input of personal style and body shape information. The system's main hardware consists of the user's device (smartphone or personal computer), a server, and a camera and microphone for running the emotion analysis engine. The software includes an image generation algorithm (e.g., DALL-E API), an emotion analysis engine for emotion analysis, and back-end services for data management and transmission.

[0272] Entering user information

[0273] The user uses the device to input their style (e.g., casual, formal, sporty) and body type (e.g., height, weight, body characteristics). The input information is compiled in JSON format and formatted as follows:

[0274] json

[0275] {

[0276] "style": "casual",

[0277] "height": 170,

[0278] "weight": 65,

[0279] "body_shape": "slim"

[0280] }

[0281] Sending input information to the server

[0282] Once the user has completed entering their information, the device sends it to the server, where the data is encrypted and transmitted over the internet, ensuring the security of the user's information.

[0283] Emotion Engine Data Collection and Analysis

[0284] The device's built-in camera and microphone are activated to capture the user's facial expressions and voice in real time. The collected data is sent to an emotion analysis engine, which generates emotional information from the user's facial expressions and voice. For example, if the user is smiling, it is recognized as a positive emotion.

[0285] DALL-E API call and image generation

[0286] The server generates a prompt based on the style and body type information entered by the user and the emotional information received from the emotion analysis engine. An example of a specific prompt is "170cm, 65kg, slim build, casual outfit. User is smiling." Using this prompt, the server sends a request to the DALL-E API. The DALL-E API generates a clothing image according to the specified prompt and sends the generated image back to the server.

[0287] Sending generated images from the server to the device

[0288] The server receives the image returned from the DALL-E API and sends it to the user's device, allowing the user to check the generated clothing image.

[0289] Displaying images and collecting user feedback

[0290] The device then displays the received clothing image to the user. During display, the emotion analysis engine continues to analyze the user's emotions and collects data on the user's reactions as appropriate. The user can then enter their ratings and feedback on the displayed image.

[0291] Sending and using feedback to the server

[0292] The user's feedback is sent to the server via the device. The server analyzes the received feedback and stores it in a database. This feedback information is used the next time an image is generated, and suggestions that better fit the user's preferences are made.

[0293] Specific examples

[0294] For example, if a user prefers casual style, is slim, 170cm tall, and weighs 65kg, the user enters this information on their device. The emotion analysis engine detects the user's smile and sends the data to the server as a positive emotion. The server then creates a prompt that reads, "170cm, 65kg, slim, casual outfit. User smiling," and sends it to the DALL-E API. The generated image is then sent to the device via the server, where the user can review it and provide specific feedback, such as "I'd like the pants to be a lighter color." This feedback is reflected in the next image generation.

[0295] The system aims to enable users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of a sentiment analysis engine will make personalized suggestions based on user emotions, further increasing user satisfaction.

[0296] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0297] Step 1:

[0298] The user uses the terminal to input information about his or her style (casual, formal, sporty, etc.) and body type (height, weight, body characteristics).

[0299] The input information is formatted into JSON format by the terminal, resulting in the following data:

[0300] json

[0301] {

[0302] "style": "casual",

[0303] "height": 170,

[0304] "weight": 65,

[0305] "body_shape": "slim"

[0306] }

[0307] The formatted data is sent in the next step.

[0308] Step 2:

[0309] The device sends the formatted user information to the server via an HTTP POST request, and the information is encrypted to ensure its security.

[0310] Input: User information in JSON format.

[0311] Output: User information sent to the server.

[0312] Step 3:

[0313] The device's camera and microphone are activated to capture the user's facial expressions and voice in real time.

[0314] The emotion analysis engine analyzes this data and generates the user's emotional information. For example, if the user is smiling, it will be analyzed as a "positive emotion."

[0315] Input: Facial and vocal data collected in real time.

[0316] Output: Emotion information from the emotion analysis engine.

[0317] Step 4:

[0318] The server generates a prompt based on the received user information and emotion information. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[0319] Input: User information, emotion information.

[0320] Output: The generated prompt.

[0321] Step 5:

[0322] The server sends the generated prompt to the image generation algorithm (DALL-E API) and requests it to generate a clothing image. The DALL-E API generates an image based on this prompt and returns the result to the server.

[0323] Input: The generated prompt.

[0324] Output: Generated image from DALL-E API.

[0325] Step 6:

[0326] The server receives the generated image returned from the DALL-E API and sends it to the user's device.

[0327] Input: Generated images from DALL-E API.

[0328] Output: The generated image sent to the device.

[0329] Step 7:

[0330] The terminal displays the received clothing image to the user.

[0331] The emotion analysis engine continues to analyze the user's facial expressions and voice, and the user enters their ratings and opinions through the feedback interface.

[0332] Input: Generated images from the server, user feedback.

[0333] Output: User input of ratings and feedback.

[0334] Step 8:

[0335] The device sends the feedback information provided by the user to the server, which receives the feedback information and stores it in a database. The feedback information is then reflected in future image generation and suggestions.

[0336] Input: User feedback information.

[0337] Output: Feedback information stored in a database.

[0338] This series of steps allows users to easily find the best clothing images based on their style and body type, significantly improving their online shopping experience. Furthermore, the sentiment analysis engine enables personalized recommendations, increasing user satisfaction.

[0339] (Application example 2)

[0340] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0341] In conventional online shopping, users cannot actually try on clothes, making it difficult to choose clothes that fit their body type and style. Furthermore, because the system does not take into account the user's emotions, personalized suggestions are not provided, resulting in an inconvenient shopping experience. The present invention aims to solve these problems by providing a system that incorporates virtual try-on and emotion recognition functions.

[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0343] In this invention, the server includes means for accepting style and body type information input by the user, means for calling an image generation API based on the style and body type information to generate clothing images, means for sending the generated clothing images to the user's terminal, means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, means for recognizing the user's emotions in real time and optimizing image generation prompts based on the results, and means for converting the generated clothing images into 3D models and providing a virtual try-on. This allows users to easily select clothing that best suits their body type and style, significantly improving their shopping experience.

[0344] "Style and body type information entered by the user" refers to information that the user specifies about their preferred clothing style, height, weight, and body type characteristics.

[0345] An "image generation API" is an application programming interface for generating new images based on input information.

[0346] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[0347] "User device" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[0348] "User feedback" refers to the evaluations and opinions that users give to the generated clothing images.

[0349] "Recognizing emotions in real time" means analyzing the user's facial expressions, voice, etc. to instantly determine their current emotional state.

[0350] An "image generation prompt" is a textual instruction that instructs the image generation API to generate a particular image.

[0351] "3D modeling" is the process of converting two-dimensional image data into three-dimensional, three-dimensional data.

[0352] "Virtual try-on" is an experience that uses virtual reality and augmented reality technology to allow users to try on clothes in a virtual space without actually having to physically try them on.

[0353] This system generates appropriate clothing images based on style and body type information entered by the user, and then provides a virtual try-on experience based on the images. The system uses emotion recognition technology to analyze the user's real-time reactions and provide personalized suggestions.

[0354] System configuration

[0355] This system consists of a user terminal, a server, an emotion engine, an image generation API, and a virtual try-on engine.

[0356] User Device

[0357] The user device can be a smartphone, tablet, or head-mounted display (HMD), which provides an interface for users to input information about their style and body shape, and collects data through a camera and microphone for emotion recognition.

[0358] server

[0359] The server receives the information sent by the user and performs the necessary data processing. The server is responsible for the following:

[0360] 1. Receiving and formatting user information: The server receives the user information and converts it into a format suitable for the image generation API.

[0361] 2. Prompt generation: Generate prompts based on data from the emotion engine. For example, generate a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[0362] 3. Calling the image generation API: Call the image generation API using the generated prompt to generate an appropriate clothing image.

[0363] 4. 3D modeling of images: The generated clothing images are converted into 3D models and sent to the virtual fitting engine.

[0364] 5. Feedback analysis: Receive feedback from users, analyze it, and reflect it in future suggestions.

[0365] Emotion Engine

[0366] The emotion engine captures the user's facial expressions and voice in real time and analyzes their emotional state, optimizing prompt generation based on the user's preferences and reactions.

[0367] Image Generation API

[0368] The image generation API uses, for example, the OpenAI API, which generates 2D clothing images based on prompts received from the server.

[0369] Virtual Try-On Engine

[0370] The virtual try-on engine converts the generated 2D images into 3D models, allowing users to try on clothes in a virtual space.

[0371] Specific examples

[0372] If the user enters "170cm, 65kg, slim build, casual style" and the emotion engine detects a smiley face response, the server will generate the following prompt:

[0373] 170cm, 65kg, slim body, casual outfit. User smiles.

[0374] Using this prompt, the image generation API generates clothing images, converts them into 3D models, and sends them to the user's device. The virtual fitting engine allows users to enjoy a virtual fitting experience. At the same time, user feedback can be collected and reflected in future proposals, providing a more personalized experience.

[0375] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0376] Step 1:

[0377] The user uses the device to input information about their style and body type, including style (e.g., casual, formal), height, weight, and body characteristics. The input data is saved in JSON format.

[0378] Step 2:

[0379] The device sends the entered user information to the server. The server receives this information, formats it, and converts it into a format suitable for the image generation API. For example, if a user is 170cm tall, weighs 65kg, and has a casual style, it generates a prompt like this: "170cm tall, 65kg, slim build, casual outfit."

[0380] Step 3:

[0381] The device uses a camera and microphone to capture the user's emotions in real time and transmits them to the server. The emotion engine analyzes this data and determines the user's emotional state. For example, if the user is smiling, it will recognize the emotion as positive.

[0382] Step 4:

[0383] The server optimizes the image generation prompt based on data from the emotion engine. It adds emotion data and regenerates the prompt. For example, it adds the element "user smiling" to the prompt, such as "170cm, 65kg, slim build, casual outfit. User smiling."

[0384] Step 5:

[0385] The server sends the generated prompt to the image generation API, which generates an appropriate clothing image. The image generation API generates a clothing image based on the specified prompt and returns the result to the server.

[0386] Step 6:

[0387] The server converts the received clothing image into a 3D model and sends it to the virtual fitting engine, which converts the 2D clothing image into three-dimensional data.

[0388] Step 7:

[0389] The server sends the 3D modeled clothing data to the user's device, allowing the user to virtually try on the clothing. The device then displays the generated clothing image to the user, providing a virtual fitting experience through the virtual fitting engine.

[0390] Step 8:

[0391] The user enters feedback on the clothing generated through the virtual try-on experience, including opinions and ratings on the color and design of the clothing.

[0392] Step 9:

[0393] The device sends user feedback to the server, which analyzes it and uses it to generate the next image. The feedback information is stored in a database and used to make suggestions that better suit the user's preferences.

[0394] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0395] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0396] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0397] [Second embodiment]

[0398] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0399] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0400] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0401] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0402] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0404] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0405] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0406] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0407] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0408] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0409] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0410] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information on their style and body type. The system includes a server, a terminal, and an image generation API.

[0411] Entering user information

[0412] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0413] Sending input information to the server

[0414] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0415] json

[0416] {

[0417] "style": "casual",

[0418] "height": 170,

[0419] "weight": 65,

[0420] "body_shape": "slim"

[0421] }

[0422] DALL-E API call and image generation

[0423] The server calls the image generation API based on the received user information, converting the information into an appropriate format and using it as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit."

[0424] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0425] Sending generated images from the server to the device

[0426] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0427] Displaying images and collecting user feedback

[0428] The terminal displays the received clothing image to the user, and further provides the user with an interface for inputting ratings and feedback, allowing the user to provide feedback on their opinions and ratings of the displayed outfit image.

[0429] Sending and using feedback to the server

[0430] Feedback from users is sent to the server via their devices. The server analyzes this feedback and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0431] Specific examples

[0432] For example, if a user prefers casual styles and has a slim build, the user enters their height of 170cm, weight of 65kg, and style of "casual" on the device. The device sends this information to the server, which then generates a prompt for "casual outfit for a 170cm, 65kg, slim build" and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0433] This allows users to easily find the clothes that suit them best, greatly improving their online shopping experience.

[0434] The processing flow will be explained below.

[0435] Step 1:

[0436] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0437] Step 2:

[0438] The device compiles the user's input information into JSON format. For example:

[0439] json

[0440] {

[0441] "style": "casual",

[0442] "height": 170,

[0443] "weight": 65,

[0444] "body_shape": "slim"

[0445] }

[0446] Step 3:

[0447] The device sends the user information compiled in JSON format to the server.

[0448] Step 4:

[0449] The server analyzes the received user information and generates an appropriate prompt. Example: "170cm, 65kg, slim build, casual outfit."

[0450] Step 5:

[0451] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0452] Step 6:

[0453] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0454] Step 7:

[0455] The server sends the clothing image received from the image generation API to the user's device.

[0456] Step 8:

[0457] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0458] Step 9:

[0459] Users can input feedback about the outfit images displayed, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[0460] Step 10:

[0461] The device sends the user's feedback to the server.

[0462] Step 11:

[0463] The server receives user feedback and stores it in a database, which is then reflected in the next image generation.

[0464] This allows users to easily find the clothes that best suit their style and body type, enhancing their online shopping experience.

[0465] Example 1

[0466] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0467] Traditional online shopping has the problem of making it difficult for users to find clothes that perfectly fit their style and body type. This results in users wasting time and effort, and they often fail to find the products they want. Another issue is that if the generated clothing image does not meet the user's expectations, feedback is not reflected in the next image generation, resulting in the same problem repeating itself.

[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0469] In this invention, the server includes means for accepting style and body type information entered by the user, means for generating a prompt based on the style and body type information, calling an image generation API, and generating a clothing image, and means for sending the generated clothing image to the user's device. This allows the user to efficiently obtain clothing images that suit their style and body type. Furthermore, by including means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, more appropriate suggestions that reflect the user's preferences and requests can be made.

[0470] A "user" is a person who uses the system to input style and body type information and view the generated clothing images.

[0471] "Style" refers to fashion categories based on the user's preferences, such as casual, formal, and sporty.

[0472] "Body type" refers to information about the user's physical appearance, such as their height, weight, and physical characteristics.

[0473] An "image generation API" is an application programming interface that generates an image based on a specified prompt.

[0474] A "prompt sentence" is a text sentence that issues a specific request to the image generation API.

[0475] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[0476] A "terminal" is an electronic device such as a computer or smartphone that a user uses to input information and receive and display generated clothing images.

[0477] "Feedback" refers to the evaluation or opinion provided by the user regarding the generated clothing image, which is reflected in the next image generation.

[0478] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. The system includes a server, a terminal, and an image generation API.

[0479] Users input information about their style and body type using their device. Specifically, style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device compiles this information in a predefined format and sends it to the server in JSON format.

[0480] The server generates a prompt based on the received user information and calls an image generation API (e.g., an API using an image generation model). In this case, a specific prompt such as "170 cm, 65 kg, slim build, casual outfit" is used. The server then sends a request including the generated prompt to the DALL-E API, requesting image generation.

[0481] The image generation API generates a clothing image based on the specified prompt and returns the result to the server. The server receives the clothing image returned from the image generation API and sends it to the user's device as an HTTP response.

[0482] The device displays the received clothing image on the screen and provides the user with an interface for inputting ratings and feedback. The user inputs their opinions and ratings about the displayed outfit image as feedback and presses the send button. Feedback can include specific comments such as "I wish the pants were a lighter color."

[0483] The device converts the user's feedback information into JSON format and sends it to the server as an HTTP POST request. The server then analyzes the feedback and stores it in a database. The saved feedback information is reflected in the next image generation, making suggestions that match the user's preferences.

[0484] For example, if a user likes casual style and has a slim build, the user can enter their height (170cm), weight (65kg), and style ("casual") into the terminal. When the user presses the "Submit" button, the terminal will generate the following prompt:

[0485] "Generate a casual outfit for a slim person who is 170cm tall and weighs 65kg."

[0486] This prompt is sent to the DALL-E API, and the generated image is sent to the device via the server. The user can review the generated image and provide feedback to reflect it in the next image generation.

[0487] In this way, users can easily find the clothes that suit them best, greatly improving their online shopping experience.

[0488] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0489] Step 1:

[0490] The user uses the device to input information about their own style and body type. Specifically, they input their style (e.g., casual, formal), height, weight, and body type characteristics into a form displayed on the device screen. The user's input is then acquired by the device.

[0491] Step 2:

[0492] The device converts the style and body information into a predefined format (JSON). Specifically, it formats the input information as follows:

[0493] json

[0494] {

[0495] "style": "casual",

[0496] "height": 170,

[0497] "weight": 65,

[0498] "body_shape": "slim"

[0499] }

[0500] The input here is the user's style and body type information, and the output is JSON format data.

[0501] Step 3:

[0502] The terminal sends the generated JSON data to the server. Specifically, it sends data to the server as an HTTP POST request. The input is JSON data, and the output is a message to the server indicating that data was successfully sent.

[0503] Step 4:

[0504] The server generates a prompt based on the received user information. Specifically, it uses a template in the server script to create a prompt like this:

[0505] "Casual outfit for a slim figure, 170cm, 65kg"

[0506] The input is the user's JSON data and the output is the prompt text.

[0507] Step 5:

[0508] The server sends the generated prompt text to the image generation API. Specifically, it sends the prompt text to the DALL-E API as an HTTP POST request. The input is the prompt text, and the output is a request sending success message to the image generation API.

[0509] Step 6:

[0510] The image generation API generates a clothing image based on the prompt received from the server and sends the result back to the server. Specifically, it generates image data using a generative AI model. The input is the prompt, and the output is the generated clothing image.

[0511] Step 7:

[0512] The server receives the clothing image returned from the image generation API and transfers it to the user's device. Specifically, it sends the image data to the user's device as an HTTP response. The input is the generated clothing image, and the output is a message that the image data was successfully sent to the user's device.

[0513] Step 8:

[0514] The device displays the received clothing image on the screen. Specifically, it renders an image display UI to allow the user to check the outfit image. The input is the generated clothing image, and the output is the display of the image.

[0515] Step 9:

[0516] The user provides feedback on the displayed outfit image. Specifically, the user inputs their opinion or rating into the rating and feedback input interface and presses the send button. The input is the feedback content, and the output is the feedback sending action.

[0517] Step 10:

[0518] The device converts the feedback information provided by the user into JSON format and sends it to the server. Specifically, it sends the feedback data as an HTTP POST request. The input is the feedback content, and the output is a feedback transmission success message to the server.

[0519] Step 11:

[0520] The server analyzes the feedback and stores it in a database. Specifically, it analyzes the feedback data and stores it as user preference data. The input is the feedback data, and the output is a message indicating that the analyzed data has been saved in the database.

[0521] Step 12:

[0522] The server reflects the feedback information when generating the next image. Specifically, it improves the next prompt generation based on the feedback stored in the database. The input is the stored feedback data, and the output is the improved prompt.

[0523] This series of processes provides the user with clothing images that are optimally suited to their style and body type, and by incorporating user feedback into future recommendations, the system delivers a more accurate and personalized online shopping experience.

[0524] (Application example 1)

[0525] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0526] It is difficult to find the perfect clothes when shopping online. In particular, the process of selecting clothes that fit the user's body type and style is complicated, resulting in an unsatisfactory shopping experience. Another issue is that personalization based on user feedback is not fully implemented, making it difficult to provide suggestions that match the user's preferences.

[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0528] In this invention, the server includes means for receiving style and body type information input by the user, means for generating a prompt for the generation AI model based on the style and body type information and for generating clothing images by calling an image generation API, means for sending the generated clothing images to the user's device, and means for receiving user feedback and reflecting the feedback in generating the next clothing image. This allows the user to easily find clothing that suits them and receive even more personalized suggestions through the feedback.

[0529] "User" refers to an individual who uses this system to input their own style and body type information and receive suggestions for suitable clothing images.

[0530] "Style" refers to categories that indicate the type of clothing and fashion trends that users prefer, such as casual, formal, or sporty.

[0531] "Body type" refers to the user's physical characteristics and includes information about height, weight, and body type such as slim or chubby.

[0532] A "generative AI model" refers to an artificial intelligence model that has the technology to generate clothing images based on prompt text provided by the user.

[0533] A "prompt sentence" refers to a text request sentence that is generated based on the user's body type and style information and sent to the image generation API.

[0534] "Image generation API" refers to a programming interface for generating an image that meets specified conditions based on the received prompt text.

[0535] "Clothing Images" refers to images of clothing generated based on the user's body type and style using a generative AI model.

[0536] "Feedback" refers to the user inputting their opinions and evaluations of the generated clothing images, and refers to information that will be used to improve the next proposal.

[0537] "Terminal" refers to a device used by a user to input information and receive generated clothing images.

[0538] "Server" refers to the central computer system that receives user input information, calls the image generation API to generate garment images, and manages the collected feedback.

[0539] This invention provides a mail-order system that helps users find clothes that suit their body type and style when shopping online. The system includes a server, a terminal, and an image generation API.

[0540] First, the user uses the device to input their body type and style information, including height, weight, body characteristics, preferred style (e.g., casual, formal, sporty), etc. The device then formats the input information into a predefined format and sends it to the server.

[0541] Based on the received information, the server generates an appropriate prompt for the generative AI model. For example, a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is sent to the image generation API, which generates a clothing image based on the prompt. The software used for this is an HTTP client such as axios and the DALL-E API.

[0542] The generated clothing image is sent from the server to the user's device, where the user can view the image. In addition, an interface for inputting evaluations and feedback is provided to the user. The user can input their opinions and evaluations of the generated clothing image. For example, they can provide feedback such as "I would like the color of the pants to be lighter." This feedback information is sent to the server and reflected the next time an image is generated.

[0543] A concrete example is as follows: If a user prefers a casual style and has a slim build, the user enters their height as 170cm, weight as 65kg, and style as "casual" on their device. The device sends this information to the server, which then generates a prompt message saying "Casual outfit for a 170cm, 65kg, slim build" and sends it to the image generation API. The generated image is then sent to the device via the server, allowing the user to check the suitable outfit. At this time, the user enters feedback such as "I'd like the pants to be a lighter color," and the server takes this information into consideration the next time it generates a clothing image.

[0544] In this way, users can efficiently find the clothes that best suit their tastes and body type, and further personalization can be achieved through feedback, greatly improving the online shopping experience.

[0545] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0546] Step 1:

[0547] The user uses the device to enter their own style and body type information. Specifically, the user enters their height, weight, body type characteristics, and preferred style (e.g., casual, formal, sporty) into the device's input form. The entered information is formatted into a JSON format defined by the system. The input here is the user's body type and style information, and the output is formatted JSON data.

[0548] Step 2:

[0549] The device sends the entered information to the server. The data sent is in JSON format, for example:

[0550] json

[0551] {

[0552] "style": "casual",

[0553] "height": 170,

[0554] "weight": 65,

[0555] "body_shape": "slim"

[0556] }

[0557] The server receives this data, where the input is formatted JSON data and the output is the received user information.

[0558] Step 3:

[0559] The server generates a prompt based on the information it receives. Specifically, it creates a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is the request sent to the generative AI model (image generation API). The input here is JSON data of the user information, and the output is the prompt.

[0560] Step 4:

[0561] The server uses the generated prompt text to call an image generation API (e.g., DALL-E API). Specifically, it sends an API request including the prompt text and receives the URL of the generated clothing image. The input here is the prompt text, and the output is the URL of the generated clothing image.

[0562] Step 5:

[0563] The server sends the URL of the clothing image received from the image generation API to the user's device. The user's device receives this URL and displays the image on the screen. The input here is the URL of the generated clothing image, and the output is the clothing image displayed on the user's device.

[0564] Step 6:

[0565] The user inputs their evaluation and feedback for the displayed clothing image. Specifically, the user inputs feedback such as "I would like the color of the pants to be lighter" into the evaluation form and sends it from the terminal to the server. The input here is the user's feedback, and the output is the feedback data sent to the server.

[0566] Step 7:

[0567] The server analyzes the received feedback and reflects it in the next prompt generation. Specifically, the feedback data is stored in a database and taken into consideration when generating the next prompt for personalization. The input here is the user's feedback data, and the output is an updated prompt generation algorithm.

[0568] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0569] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. In addition to the basic functions of inputting user information, calling an image generation API, displaying generated images, and collecting feedback, the system also incorporates an emotion engine that recognizes the user's emotions.

[0570] Entering user information

[0571] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0572] Sending input information to the server

[0573] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0574] json

[0575] {

[0576] "style": "casual",

[0577] "height": 170,

[0578] "weight": 65,

[0579] "body_shape": "slim"

[0580] }

[0581] Emotion Engine Data Collection and Analysis

[0582] The device is equipped with a camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine uses this data to analyze the user's emotions. For example, if the user is smiling, it is recognized as a positive emotion.

[0583] DALL-E API call and image generation

[0584] The server calls the image generation API based on the received user information and the analysis results from the emotion engine. At this time, the information is converted into an appropriate format and used as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[0585] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0586] Sending generated images from the server to the device

[0587] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0588] Displaying images and collecting user feedback

[0589] The device displays the received clothing images to the user. It also provides the user with an interface for inputting ratings and feedback. The emotion engine then analyzes the user's emotions and uses the results to improve the displayed content. The user can provide feedback on their opinions and ratings of the displayed outfit images.

[0590] Sending and using feedback to the server

[0591] User feedback is sent to the server via the device. The server analyzes this feedback along with emotional data and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0592] Specific examples

[0593] For example, if a user prefers casual styles and has a slim build, they can enter their height of 170cm, weight of 65kg, and style of "casual" on their device. At this time, the device captures the user's emotions, and the emotion engine determines that they are reacting positively. The device sends this information to the server, which then generates a prompt that reads, "Casual outfit for a 170cm, 65kg, slim build. User smiling," and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0594] This will allow users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of the emotion engine will enable more personalized suggestions based on the user's emotions, aiming to further increase user satisfaction.

[0595] The processing flow will be explained below.

[0596] Step 1:

[0597] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0598] Step 2:

[0599] The device formats the entered user information into JSON format. Example:

[0600] json

[0601] {

[0602] "style": "casual",

[0603] "height": 170,

[0604] "weight": 65,

[0605] "body_shape": "slim"

[0606] }

[0607] Step 3:

[0608] The terminal transmits the formatted user information to the server.

[0609] Step 4:

[0610] The device captures the user's facial expressions and voice, and collects emotional data using the device's built-in camera and microphone.

[0611] Step 5:

[0612] The emotion engine analyzes data collected by the device to determine the user's emotional state. For example, a smiling user is recognized as a positive emotion.

[0613] Step 6:

[0614] The server generates a prompt based on the received user information and emotion data. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[0615] Step 7:

[0616] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0617] Step 8:

[0618] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0619] Step 9:

[0620] The server sends the clothing image received from the image generation API to the user's device.

[0621] Step 10:

[0622] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0623] Step 11:

[0624] The device continues to use the emotion engine to monitor the user's emotional state, detecting the user's reaction to the displayed image (e.g., smiling, making a displeased face).

[0625] Step 12:

[0626] Users can input feedback on outfit images, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[0627] Step 13:

[0628] The device sends the user's feedback to the server.

[0629] Step 14:

[0630] The server stores the received feedback along with the emotion data and reflects it in the next image generation. The server updates the feedback database and uses it to generate the next prompt.

[0631] In this way, users can easily find the clothes that best suit their style and body type, improving their online shopping experience. In addition, the emotion engine can make suggestions based on the user's emotional state, further improving user satisfaction.

[0632] Example 2

[0633] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0634] In today's online shopping environment, users often find it difficult to find clothing that suits their style and body type. Furthermore, the lack of technology that makes personalized recommendations based on user emotions makes it difficult to provide a satisfying shopping experience. To address this issue, a system is needed that not only generates clothing images based on a user's style and body type information, but also makes personalized recommendations that take the user's emotions into account.

[0635] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0636] In this invention, the server includes means for accepting style and body type information input by the user, means for generating clothing images by invoking an image generation algorithm based on the style and body type information, and means for transmitting the generated clothing images to the user's information processing device. This makes it possible to provide accurate clothing images based on the user's style and body type, and to reflect feedback based on the results. Furthermore, by incorporating means for collecting emotional information using an emotion analysis engine that analyzes user emotions, and means for providing the emotional information along with the style and body type information as prompts to the generation AI model, personalized suggestions based on the user's emotions can be made, significantly improving the satisfaction of the online shopping experience.

[0637] "Style" is information that refers to the clothing category and design characteristics that a user prefers.

[0638] "Body type" refers to information that refers to the user's physical proportions, such as the user's height, weight, and body shape characteristics.

[0639] An "image generation algorithm" is a program or system that generates images of suitable clothing based on style and body type information provided by the user.

[0640] An "information processing device" is a terminal used by a user, specifically a device such as a smartphone or personal computer.

[0641] "Feedback" refers to information that indicates the evaluation or opinion that a user provides regarding the generated clothing image.

[0642] An "emotion analysis engine" is a program or system that collects and analyzes emotional information from a user's facial expressions and voice in real time.

[0643] "Emotional information" is data that indicates the emotional state of a user collected from their facial expressions and voice.

[0644] A "prompt" is input information given to a generative AI model, and refers to the form of data that includes style, body type, and emotional information.

[0645] A "generative AI model" is an artificial intelligence system that generates content, such as images, based on given prompts.

[0646] The virtual personal shopper system of the present invention aims to generate appropriate clothing images based on user input of personal style and body shape information. The system's main hardware consists of the user's device (smartphone or personal computer), a server, and a camera and microphone for running the emotion analysis engine. The software includes an image generation algorithm (e.g., DALL-E API), an emotion analysis engine for emotion analysis, and back-end services for data management and transmission.

[0647] Entering user information

[0648] The user uses the device to input their style (e.g., casual, formal, sporty) and body type (e.g., height, weight, body characteristics). The input information is compiled in JSON format and formatted as follows:

[0649] json

[0650] {

[0651] "style": "casual",

[0652] "height": 170,

[0653] "weight": 65,

[0654] "body_shape": "slim"

[0655] }

[0656] Sending input information to the server

[0657] Once the user has completed entering their information, the device sends it to the server, where the data is encrypted and transmitted over the internet, ensuring the security of the user's information.

[0658] Emotion Engine Data Collection and Analysis

[0659] The device's built-in camera and microphone are activated to capture the user's facial expressions and voice in real time. The collected data is sent to an emotion analysis engine, which generates emotional information from the user's facial expressions and voice. For example, if the user is smiling, it is recognized as a positive emotion.

[0660] DALL-E API call and image generation

[0661] The server generates a prompt based on the style and body type information entered by the user and the emotional information received from the emotion analysis engine. An example of a specific prompt is "170cm, 65kg, slim build, casual outfit. User is smiling." Using this prompt, the server sends a request to the DALL-E API. The DALL-E API generates a clothing image according to the specified prompt and sends the generated image back to the server.

[0662] Sending generated images from the server to the device

[0663] The server receives the image returned from the DALL-E API and sends it to the user's device, allowing the user to check the generated clothing image.

[0664] Displaying images and collecting user feedback

[0665] The device then displays the received clothing image to the user. During display, the emotion analysis engine continues to analyze the user's emotions and collects data on the user's reactions as appropriate. The user can then enter their ratings and feedback on the displayed image.

[0666] Sending and using feedback to the server

[0667] The user's feedback is sent to the server via the device. The server analyzes the received feedback and stores it in a database. This feedback information is used the next time an image is generated, and suggestions that better fit the user's preferences are made.

[0668] Specific examples

[0669] For example, if a user prefers casual style, is slim, 170cm tall, and weighs 65kg, the user enters this information on their device. The emotion analysis engine detects the user's smile and sends the data to the server as a positive emotion. The server then creates a prompt that reads, "170cm, 65kg, slim, casual outfit. User smiling," and sends it to the DALL-E API. The generated image is then sent to the device via the server, where the user can review it and provide specific feedback, such as "I'd like the pants to be a lighter color." This feedback is reflected in the next image generation.

[0670] The system aims to enable users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of a sentiment analysis engine will make personalized suggestions based on user emotions, further increasing user satisfaction.

[0671] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0672] Step 1:

[0673] The user uses the terminal to input information about his or her style (casual, formal, sporty, etc.) and body type (height, weight, body characteristics).

[0674] The input information is formatted into JSON format by the terminal, resulting in the following data:

[0675] json

[0676] {

[0677] "style": "casual",

[0678] "height": 170,

[0679] "weight": 65,

[0680] "body_shape": "slim"

[0681] }

[0682] The formatted data is sent in the next step.

[0683] Step 2:

[0684] The device sends the formatted user information to the server via an HTTP POST request, and the information is encrypted to ensure its security.

[0685] Input: User information in JSON format.

[0686] Output: User information sent to the server.

[0687] Step 3:

[0688] The device's camera and microphone are activated to capture the user's facial expressions and voice in real time.

[0689] The emotion analysis engine analyzes this data and generates the user's emotional information. For example, if the user is smiling, it will be analyzed as a "positive emotion."

[0690] Input: Facial and vocal data collected in real time.

[0691] Output: Emotion information from the emotion analysis engine.

[0692] Step 4:

[0693] The server generates a prompt based on the received user information and emotion information. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[0694] Input: User information, emotion information.

[0695] Output: The generated prompt.

[0696] Step 5:

[0697] The server sends the generated prompt to the image generation algorithm (DALL-E API) and requests it to generate a clothing image. The DALL-E API generates an image based on this prompt and returns the result to the server.

[0698] Input: The generated prompt.

[0699] Output: Generated image from DALL-E API.

[0700] Step 6:

[0701] The server receives the generated image returned from the DALL-E API and sends it to the user's device.

[0702] Input: Generated images from DALL-E API.

[0703] Output: The generated image sent to the device.

[0704] Step 7:

[0705] The terminal displays the received clothing image to the user.

[0706] The emotion analysis engine continues to analyze the user's facial expressions and voice, and the user enters their ratings and opinions through the feedback interface.

[0707] Input: Generated images from the server, user feedback.

[0708] Output: User input of ratings and feedback.

[0709] Step 8:

[0710] The device sends the feedback information provided by the user to the server, which receives the feedback information and stores it in a database. The feedback information is then reflected in future image generation and suggestions.

[0711] Input: User feedback information.

[0712] Output: Feedback information stored in a database.

[0713] This series of steps allows users to easily find the best clothing images based on their style and body type, significantly improving their online shopping experience. Furthermore, the sentiment analysis engine enables personalized recommendations, increasing user satisfaction.

[0714] (Application example 2)

[0715] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] In conventional online shopping, users cannot actually try on clothes, making it difficult to choose clothes that fit their body type and style. Furthermore, because the system does not take into account the user's emotions, personalized suggestions are not provided, resulting in an inconvenient shopping experience. The present invention aims to solve these problems by providing a system that incorporates virtual try-on and emotion recognition functions.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0718] In this invention, the server includes means for accepting style and body type information input by the user, means for calling an image generation API based on the style and body type information to generate clothing images, means for sending the generated clothing images to the user's terminal, means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, means for recognizing the user's emotions in real time and optimizing image generation prompts based on the results, and means for converting the generated clothing images into 3D models and providing a virtual try-on. This allows users to easily select clothing that best suits their body type and style, significantly improving their shopping experience.

[0719] "Style and body type information entered by the user" refers to information that the user specifies about their preferred clothing style, height, weight, and body type characteristics.

[0720] An "image generation API" is an application programming interface for generating new images based on input information.

[0721] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[0722] "User device" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[0723] "User feedback" refers to the evaluations and opinions that users give to the generated clothing images.

[0724] "Recognizing emotions in real time" means analyzing the user's facial expressions, voice, etc. to instantly determine their current emotional state.

[0725] An "image generation prompt" is a textual instruction that instructs the image generation API to generate a particular image.

[0726] "3D modeling" is the process of converting two-dimensional image data into three-dimensional, three-dimensional data.

[0727] "Virtual try-on" is an experience that uses virtual reality and augmented reality technology to allow users to try on clothes in a virtual space without actually having to physically try them on.

[0728] This system generates appropriate clothing images based on style and body type information entered by the user, and then provides a virtual try-on experience based on the images. The system uses emotion recognition technology to analyze the user's real-time reactions and provide personalized suggestions.

[0729] System configuration

[0730] This system consists of a user terminal, a server, an emotion engine, an image generation API, and a virtual try-on engine.

[0731] User Device

[0732] The user device can be a smartphone, tablet, or head-mounted display (HMD), which provides an interface for users to input information about their style and body shape, and collects data through a camera and microphone for emotion recognition.

[0733] server

[0734] The server receives the information sent by the user and performs the necessary data processing. The server is responsible for the following:

[0735] 1. Receiving and formatting user information: The server receives the user information and converts it into a format suitable for the image generation API.

[0736] 2. Prompt generation: Generate prompts based on data from the emotion engine. For example, generate a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[0737] 3. Calling the image generation API: Call the image generation API using the generated prompt to generate an appropriate clothing image.

[0738] 4. 3D modeling of images: The generated clothing images are converted into 3D models and sent to the virtual fitting engine.

[0739] 5. Feedback analysis: Receive feedback from users, analyze it, and reflect it in future suggestions.

[0740] Emotion Engine

[0741] The emotion engine captures the user's facial expressions and voice in real time and analyzes their emotional state, optimizing prompt generation based on the user's preferences and reactions.

[0742] Image Generation API

[0743] The image generation API uses, for example, the OpenAI API, which generates 2D clothing images based on prompts received from the server.

[0744] Virtual Try-On Engine

[0745] The virtual try-on engine converts the generated 2D images into 3D models, allowing users to try on clothes in a virtual space.

[0746] Specific examples

[0747] If the user enters "170cm, 65kg, slim build, casual style" and the emotion engine detects a smiley face response, the server will generate the following prompt:

[0748] 170cm, 65kg, slim body, casual outfit. User smiles.

[0749] Using this prompt, the image generation API generates clothing images, converts them into 3D models, and sends them to the user's device. The virtual fitting engine allows users to enjoy a virtual fitting experience. At the same time, user feedback can be collected and reflected in future proposals, providing a more personalized experience.

[0750] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0751] Step 1:

[0752] The user uses the device to input information about their style and body type, including style (e.g., casual, formal), height, weight, and body characteristics. The input data is saved in JSON format.

[0753] Step 2:

[0754] The device sends the entered user information to the server. The server receives this information, formats it, and converts it into a format suitable for the image generation API. For example, if a user is 170cm tall, weighs 65kg, and has a casual style, it generates a prompt like this: "170cm tall, 65kg, slim build, casual outfit."

[0755] Step 3:

[0756] The device uses a camera and microphone to capture the user's emotions in real time and transmits them to the server. The emotion engine analyzes this data and determines the user's emotional state. For example, if the user is smiling, it will recognize the emotion as positive.

[0757] Step 4:

[0758] The server optimizes the image generation prompt based on data from the emotion engine. It adds emotion data and regenerates the prompt. For example, it adds the element "user smiling" to the prompt, such as "170cm, 65kg, slim build, casual outfit. User smiling."

[0759] Step 5:

[0760] The server sends the generated prompt to the image generation API, which generates an appropriate clothing image. The image generation API generates a clothing image based on the specified prompt and returns the result to the server.

[0761] Step 6:

[0762] The server converts the received clothing image into a 3D model and sends it to the virtual fitting engine, which converts the 2D clothing image into three-dimensional data.

[0763] Step 7:

[0764] The server sends the 3D modeled clothing data to the user's device, allowing the user to virtually try on the clothing. The device then displays the generated clothing image to the user, providing a virtual fitting experience through the virtual fitting engine.

[0765] Step 8:

[0766] The user enters feedback on the clothing generated through the virtual try-on experience, including opinions and ratings on the color and design of the clothing.

[0767] Step 9:

[0768] The device sends user feedback to the server, which analyzes it and uses it to generate the next image. The feedback information is stored in a database and used to make suggestions that better suit the user's preferences.

[0769] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0770] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0771] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0772] [Third embodiment]

[0773] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0774] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0775] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0776] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0777] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0778] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0779] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0780] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0781] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0782] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0783] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0784] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0785] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information on their style and body type. The system includes a server, a terminal, and an image generation API.

[0786] Entering user information

[0787] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0788] Sending input information to the server

[0789] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0790] json

[0791] {

[0792] "style": "casual",

[0793] "height": 170,

[0794] "weight": 65,

[0795] "body_shape": "slim"

[0796] }

[0797] DALL-E API call and image generation

[0798] The server calls the image generation API based on the received user information, converting the information into an appropriate format and using it as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit."

[0799] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0800] Sending generated images from the server to the device

[0801] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0802] Displaying images and collecting user feedback

[0803] The terminal displays the received clothing image to the user, and further provides the user with an interface for inputting ratings and feedback, allowing the user to provide feedback on their opinions and ratings of the displayed outfit image.

[0804] Sending and using feedback to the server

[0805] Feedback from users is sent to the server via their devices. The server analyzes this feedback and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0806] Specific examples

[0807] For example, if a user prefers casual styles and has a slim build, the user enters their height of 170cm, weight of 65kg, and style of "casual" on the device. The device sends this information to the server, which then generates a prompt for "casual outfit for a 170cm, 65kg, slim build" and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0808] This allows users to easily find the clothes that suit them best, greatly improving their online shopping experience.

[0809] The processing flow will be explained below.

[0810] Step 1:

[0811] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0812] Step 2:

[0813] The device compiles the user's input information into JSON format. For example:

[0814] json

[0815] {

[0816] "style": "casual",

[0817] "height": 170,

[0818] "weight": 65,

[0819] "body_shape": "slim"

[0820] }

[0821] Step 3:

[0822] The device sends the user information compiled in JSON format to the server.

[0823] Step 4:

[0824] The server analyzes the received user information and generates an appropriate prompt. Example: "170cm, 65kg, slim build, casual outfit."

[0825] Step 5:

[0826] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0827] Step 6:

[0828] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0829] Step 7:

[0830] The server sends the clothing image received from the image generation API to the user's device.

[0831] Step 8:

[0832] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0833] Step 9:

[0834] Users can input feedback about the outfit images displayed, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[0835] Step 10:

[0836] The device sends the user's feedback to the server.

[0837] Step 11:

[0838] The server receives user feedback and stores it in a database, which is then reflected in the next image generation.

[0839] This allows users to easily find the clothes that best suit their style and body type, enhancing their online shopping experience.

[0840] Example 1

[0841] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0842] Traditional online shopping has the problem of making it difficult for users to find clothes that perfectly fit their style and body type. This results in users wasting time and effort, and they often fail to find the products they want. Another issue is that if the generated clothing image does not meet the user's expectations, feedback is not reflected in the next image generation, resulting in the same problem repeating itself.

[0843] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0844] In this invention, the server includes means for accepting style and body type information entered by the user, means for generating a prompt based on the style and body type information, calling an image generation API, and generating a clothing image, and means for sending the generated clothing image to the user's device. This allows the user to efficiently obtain clothing images that suit their style and body type. Furthermore, by including means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, more appropriate suggestions that reflect the user's preferences and requests can be made.

[0845] A "user" is a person who uses the system to input style and body type information and view the generated clothing images.

[0846] "Style" refers to fashion categories based on the user's preferences, such as casual, formal, and sporty.

[0847] "Body type" refers to information about the user's physical appearance, such as their height, weight, and physical characteristics.

[0848] An "image generation API" is an application programming interface that generates an image based on a specified prompt.

[0849] A "prompt sentence" is a text sentence that issues a specific request to the image generation API.

[0850] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[0851] A "terminal" is an electronic device such as a computer or smartphone that a user uses to input information and receive and display generated clothing images.

[0852] "Feedback" refers to the evaluation or opinion provided by the user regarding the generated clothing image, which is reflected in the next image generation.

[0853] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. The system includes a server, a terminal, and an image generation API.

[0854] Users input information about their style and body type using their device. Specifically, style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device compiles this information in a predefined format and sends it to the server in JSON format.

[0855] The server generates a prompt based on the received user information and calls an image generation API (e.g., an API using an image generation model). In this case, a specific prompt such as "170 cm, 65 kg, slim build, casual outfit" is used. The server then sends a request including the generated prompt to the DALL-E API, requesting image generation.

[0856] The image generation API generates a clothing image based on the specified prompt and returns the result to the server. The server receives the clothing image returned from the image generation API and sends it to the user's device as an HTTP response.

[0857] The device displays the received clothing image on the screen and provides the user with an interface for inputting ratings and feedback. The user inputs their opinions and ratings about the displayed outfit image as feedback and presses the send button. Feedback can include specific comments such as "I wish the pants were a lighter color."

[0858] The device converts the user's feedback information into JSON format and sends it to the server as an HTTP POST request. The server then analyzes the feedback and stores it in a database. The saved feedback information is reflected in the next image generation, making suggestions that match the user's preferences.

[0859] For example, if a user likes casual style and has a slim build, the user can enter their height (170cm), weight (65kg), and style ("casual") into the terminal. When the user presses the "Submit" button, the terminal will generate the following prompt:

[0860] "Generate a casual outfit for a slim person who is 170cm tall and weighs 65kg."

[0861] This prompt is sent to the DALL-E API, and the generated image is sent to the device via the server. The user can review the generated image and provide feedback to reflect it in the next image generation.

[0862] In this way, users can easily find the clothes that suit them best, greatly improving their online shopping experience.

[0863] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0864] Step 1:

[0865] The user uses the device to input information about their own style and body type. Specifically, they input their style (e.g., casual, formal), height, weight, and body type characteristics into a form displayed on the device screen. The user's input is then acquired by the device.

[0866] Step 2:

[0867] The device converts the style and body information into a predefined format (JSON). Specifically, it formats the input information as follows:

[0868] json

[0869] {

[0870] "style": "casual",

[0871] "height": 170,

[0872] "weight": 65,

[0873] "body_shape": "slim"

[0874] }

[0875] The input here is the user's style and body type information, and the output is JSON format data.

[0876] Step 3:

[0877] The terminal sends the generated JSON data to the server. Specifically, it sends data to the server as an HTTP POST request. The input is JSON data, and the output is a message to the server indicating that data was successfully sent.

[0878] Step 4:

[0879] The server generates a prompt based on the received user information. Specifically, it uses a template in the server script to create a prompt like this:

[0880] "Casual outfit for a slim figure, 170cm, 65kg"

[0881] The input is the user's JSON data and the output is the prompt text.

[0882] Step 5:

[0883] The server sends the generated prompt text to the image generation API. Specifically, it sends the prompt text to the DALL-E API as an HTTP POST request. The input is the prompt text, and the output is a request sending success message to the image generation API.

[0884] Step 6:

[0885] The image generation API generates a clothing image based on the prompt received from the server and sends the result back to the server. Specifically, it generates image data using a generative AI model. The input is the prompt, and the output is the generated clothing image.

[0886] Step 7:

[0887] The server receives the clothing image returned from the image generation API and transfers it to the user's device. Specifically, it sends the image data to the user's device as an HTTP response. The input is the generated clothing image, and the output is a message that the image data was successfully sent to the user's device.

[0888] Step 8:

[0889] The device displays the received clothing image on the screen. Specifically, it renders an image display UI to allow the user to check the outfit image. The input is the generated clothing image, and the output is the display of the image.

[0890] Step 9:

[0891] The user provides feedback on the displayed outfit image. Specifically, the user inputs their opinion or rating into the rating and feedback input interface and presses the send button. The input is the feedback content, and the output is the feedback sending action.

[0892] Step 10:

[0893] The device converts the feedback information provided by the user into JSON format and sends it to the server. Specifically, it sends the feedback data as an HTTP POST request. The input is the feedback content, and the output is a feedback transmission success message to the server.

[0894] Step 11:

[0895] The server analyzes the feedback and stores it in a database. Specifically, it analyzes the feedback data and stores it as user preference data. The input is the feedback data, and the output is a message indicating that the analyzed data has been saved in the database.

[0896] Step 12:

[0897] The server reflects the feedback information when generating the next image. Specifically, it improves the next prompt generation based on the feedback stored in the database. The input is the stored feedback data, and the output is the improved prompt.

[0898] This series of processes provides the user with clothing images that are optimally suited to their style and body type, and by incorporating user feedback into future recommendations, the system delivers a more accurate and personalized online shopping experience.

[0899] (Application example 1)

[0900] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0901] It is difficult to find the perfect clothes when shopping online. In particular, the process of selecting clothes that fit the user's body type and style is complicated, resulting in an unsatisfactory shopping experience. Another issue is that personalization based on user feedback is not fully implemented, making it difficult to provide suggestions that match the user's preferences.

[0902] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0903] In this invention, the server includes means for receiving style and body type information input by the user, means for generating a prompt for the generation AI model based on the style and body type information and for generating clothing images by calling an image generation API, means for sending the generated clothing images to the user's device, and means for receiving user feedback and reflecting the feedback in generating the next clothing image. This allows the user to easily find clothing that suits them and receive even more personalized suggestions through the feedback.

[0904] "User" refers to an individual who uses this system to input their own style and body type information and receive suggestions for suitable clothing images.

[0905] "Style" refers to categories that indicate the type of clothing and fashion trends that users prefer, such as casual, formal, or sporty.

[0906] "Body type" refers to the user's physical characteristics and includes information about height, weight, and body type such as slim or chubby.

[0907] A "generative AI model" refers to an artificial intelligence model that has the technology to generate clothing images based on prompt text provided by the user.

[0908] A "prompt sentence" refers to a text request sentence that is generated based on the user's body type and style information and sent to the image generation API.

[0909] "Image generation API" refers to a programming interface for generating an image that meets specified conditions based on the received prompt text.

[0910] "Clothing Images" refers to images of clothing generated based on the user's body type and style using a generative AI model.

[0911] "Feedback" refers to the user inputting their opinions and evaluations of the generated clothing images, and refers to information that will be used to improve the next proposal.

[0912] "Terminal" refers to a device used by a user to input information and receive generated clothing images.

[0913] "Server" refers to the central computer system that receives user input information, calls the image generation API to generate garment images, and manages the collected feedback.

[0914] This invention provides a mail-order system that helps users find clothes that suit their body type and style when shopping online. The system includes a server, a terminal, and an image generation API.

[0915] First, the user uses the device to input their body type and style information, including height, weight, body characteristics, preferred style (e.g., casual, formal, sporty), etc. The device then formats the input information into a predefined format and sends it to the server.

[0916] Based on the received information, the server generates an appropriate prompt for the generative AI model. For example, a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is sent to the image generation API, which generates a clothing image based on the prompt. The software used for this is an HTTP client such as axios and the DALL-E API.

[0917] The generated clothing image is sent from the server to the user's device, where the user can view the image. In addition, an interface for inputting evaluations and feedback is provided to the user. The user can input their opinions and evaluations of the generated clothing image. For example, they can provide feedback such as "I would like the color of the pants to be lighter." This feedback information is sent to the server and reflected the next time an image is generated.

[0918] A concrete example is as follows: If a user prefers a casual style and has a slim build, the user enters their height as 170cm, weight as 65kg, and style as "casual" on their device. The device sends this information to the server, which then generates a prompt message saying "Casual outfit for a 170cm, 65kg, slim build" and sends it to the image generation API. The generated image is then sent to the device via the server, allowing the user to check the suitable outfit. At this time, the user enters feedback such as "I'd like the pants to be a lighter color," and the server takes this information into consideration the next time it generates a clothing image.

[0919] In this way, users can efficiently find the clothes that best suit their tastes and body type, and further personalization can be achieved through feedback, greatly improving the online shopping experience.

[0920] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0921] Step 1:

[0922] The user uses the device to enter their own style and body type information. Specifically, the user enters their height, weight, body type characteristics, and preferred style (e.g., casual, formal, sporty) into the device's input form. The entered information is formatted into a JSON format defined by the system. The input here is the user's body type and style information, and the output is formatted JSON data.

[0923] Step 2:

[0924] The device sends the entered information to the server. The data sent is in JSON format, for example:

[0925] json

[0926] {

[0927] "style": "casual",

[0928] "height": 170,

[0929] "weight": 65,

[0930] "body_shape": "slim"

[0931] }

[0932] The server receives this data, where the input is formatted JSON data and the output is the received user information.

[0933] Step 3:

[0934] The server generates a prompt based on the information it receives. Specifically, it creates a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is the request sent to the generative AI model (image generation API). The input here is JSON data of the user information, and the output is the prompt.

[0935] Step 4:

[0936] The server uses the generated prompt text to call an image generation API (e.g., DALL-E API). Specifically, it sends an API request including the prompt text and receives the URL of the generated clothing image. The input here is the prompt text, and the output is the URL of the generated clothing image.

[0937] Step 5:

[0938] The server sends the URL of the clothing image received from the image generation API to the user's device. The user's device receives this URL and displays the image on the screen. The input here is the URL of the generated clothing image, and the output is the clothing image displayed on the user's device.

[0939] Step 6:

[0940] The user inputs their evaluation and feedback for the displayed clothing image. Specifically, the user inputs feedback such as "I would like the color of the pants to be lighter" into the evaluation form and sends it from the terminal to the server. The input here is the user's feedback, and the output is the feedback data sent to the server.

[0941] Step 7:

[0942] The server analyzes the received feedback and reflects it in the next prompt generation. Specifically, the feedback data is stored in a database and taken into consideration when generating the next prompt for personalization. The input here is the user's feedback data, and the output is an updated prompt generation algorithm.

[0943] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0944] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. In addition to the basic functions of inputting user information, calling an image generation API, displaying generated images, and collecting feedback, the system also incorporates an emotion engine that recognizes the user's emotions.

[0945] Entering user information

[0946] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[0947] Sending input information to the server

[0948] Once the input is complete, the device sends this information to the server in JSON format, for example:

[0949] json

[0950] {

[0951] "style": "casual",

[0952] "height": 170,

[0953] "weight": 65,

[0954] "body_shape": "slim"

[0955] }

[0956] Emotion Engine Data Collection and Analysis

[0957] The device is equipped with a camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine uses this data to analyze the user's emotions. For example, if the user is smiling, it is recognized as a positive emotion.

[0958] DALL-E API call and image generation

[0959] The server calls the image generation API based on the received user information and the analysis results from the emotion engine. At this time, the information is converted into an appropriate format and used as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[0960] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[0961] Sending generated images from the server to the device

[0962] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[0963] Displaying images and collecting user feedback

[0964] The device displays the received clothing images to the user. It also provides the user with an interface for inputting ratings and feedback. The emotion engine then analyzes the user's emotions and uses the results to improve the displayed content. The user can provide feedback on their opinions and ratings of the displayed outfit images.

[0965] Sending and using feedback to the server

[0966] User feedback is sent to the server via the device. The server analyzes this feedback along with emotional data and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[0967] Specific examples

[0968] For example, if a user prefers casual styles and has a slim build, they can enter their height of 170cm, weight of 65kg, and style of "casual" on their device. At this time, the device captures the user's emotions, and the emotion engine determines that they are reacting positively. The device sends this information to the server, which then generates a prompt that reads, "Casual outfit for a 170cm, 65kg, slim build. User smiling," and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[0969] This will allow users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of the emotion engine will enable more personalized suggestions based on the user's emotions, aiming to further increase user satisfaction.

[0970] The processing flow will be explained below.

[0971] Step 1:

[0972] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[0973] Step 2:

[0974] The device formats the entered user information into JSON format. Example:

[0975] json

[0976] {

[0977] "style": "casual",

[0978] "height": 170,

[0979] "weight": 65,

[0980] "body_shape": "slim"

[0981] }

[0982] Step 3:

[0983] The terminal transmits the formatted user information to the server.

[0984] Step 4:

[0985] The device captures the user's facial expressions and voice, and collects emotional data using the device's built-in camera and microphone.

[0986] Step 5:

[0987] The emotion engine analyzes data collected by the device to determine the user's emotional state. For example, a smiling user is recognized as a positive emotion.

[0988] Step 6:

[0989] The server generates a prompt based on the received user information and emotion data. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[0990] Step 7:

[0991] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[0992] Step 8:

[0993] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[0994] Step 9:

[0995] The server sends the clothing image received from the image generation API to the user's device.

[0996] Step 10:

[0997] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[0998] Step 11:

[0999] The device continues to use the emotion engine to monitor the user's emotional state, detecting the user's reaction to the displayed image (e.g., smiling, making a displeased face).

[1000] Step 12:

[1001] Users can input feedback on outfit images, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[1002] Step 13:

[1003] The device sends the user's feedback to the server.

[1004] Step 14:

[1005] The server stores the received feedback along with the emotion data and reflects it in the next image generation. The server updates the feedback database and uses it to generate the next prompt.

[1006] In this way, users can easily find the clothes that best suit their style and body type, improving their online shopping experience. In addition, the emotion engine can make suggestions based on the user's emotional state, further improving user satisfaction.

[1007] Example 2

[1008] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1009] In today's online shopping environment, users often find it difficult to find clothing that suits their style and body type. Furthermore, the lack of technology that makes personalized recommendations based on user emotions makes it difficult to provide a satisfying shopping experience. To address this issue, a system is needed that not only generates clothing images based on a user's style and body type information, but also makes personalized recommendations that take the user's emotions into account.

[1010] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1011] In this invention, the server includes means for accepting style and body type information input by the user, means for generating clothing images by invoking an image generation algorithm based on the style and body type information, and means for transmitting the generated clothing images to the user's information processing device. This makes it possible to provide accurate clothing images based on the user's style and body type, and to reflect feedback based on the results. Furthermore, by incorporating means for collecting emotional information using an emotion analysis engine that analyzes user emotions, and means for providing the emotional information along with the style and body type information as prompts to the generation AI model, personalized suggestions based on the user's emotions can be made, significantly improving the satisfaction of the online shopping experience.

[1012] "Style" is information that refers to the clothing category and design characteristics that a user prefers.

[1013] "Body type" refers to information that refers to the user's physical proportions, such as the user's height, weight, and body shape characteristics.

[1014] An "image generation algorithm" is a program or system that generates images of suitable clothing based on style and body type information provided by the user.

[1015] An "information processing device" is a terminal used by a user, specifically a device such as a smartphone or personal computer.

[1016] "Feedback" refers to information that indicates the evaluation or opinion that a user provides regarding the generated clothing image.

[1017] An "emotion analysis engine" is a program or system that collects and analyzes emotional information from a user's facial expressions and voice in real time.

[1018] "Emotional information" is data that indicates the emotional state of a user collected from their facial expressions and voice.

[1019] A "prompt" is input information given to a generative AI model, and refers to the form of data that includes style, body type, and emotional information.

[1020] A "generative AI model" is an artificial intelligence system that generates content, such as images, based on given prompts.

[1021] The virtual personal shopper system of the present invention aims to generate appropriate clothing images based on user input of personal style and body shape information. The system's main hardware consists of the user's device (smartphone or personal computer), a server, and a camera and microphone for running the emotion analysis engine. The software includes an image generation algorithm (e.g., DALL-E API), an emotion analysis engine for emotion analysis, and back-end services for data management and transmission.

[1022] Entering user information

[1023] The user uses the device to input their style (e.g., casual, formal, sporty) and body type (e.g., height, weight, body characteristics). The input information is compiled in JSON format and formatted as follows:

[1024] json

[1025] {

[1026] "style": "casual",

[1027] "height": 170,

[1028] "weight": 65,

[1029] "body_shape": "slim"

[1030] }

[1031] Sending input information to the server

[1032] Once the user has completed entering their information, the device sends it to the server, where the data is encrypted and transmitted over the internet, ensuring the security of the user's information.

[1033] Emotion Engine Data Collection and Analysis

[1034] The device's built-in camera and microphone are activated to capture the user's facial expressions and voice in real time. The collected data is sent to an emotion analysis engine, which generates emotional information from the user's facial expressions and voice. For example, if the user is smiling, it is recognized as a positive emotion.

[1035] DALL-E API call and image generation

[1036] The server generates a prompt based on the style and body type information entered by the user and the emotional information received from the emotion analysis engine. An example of a specific prompt is "170cm, 65kg, slim build, casual outfit. User is smiling." Using this prompt, the server sends a request to the DALL-E API. The DALL-E API generates a clothing image according to the specified prompt and sends the generated image back to the server.

[1037] Sending generated images from the server to the device

[1038] The server receives the image returned from the DALL-E API and sends it to the user's device, allowing the user to check the generated clothing image.

[1039] Displaying images and collecting user feedback

[1040] The device then displays the received clothing image to the user. During display, the emotion analysis engine continues to analyze the user's emotions and collects data on the user's reactions as appropriate. The user can then enter their ratings and feedback on the displayed image.

[1041] Sending and using feedback to the server

[1042] The user's feedback is sent to the server via the device. The server analyzes the received feedback and stores it in a database. This feedback information is used the next time an image is generated, and suggestions that better fit the user's preferences are made.

[1043] Specific examples

[1044] For example, if a user prefers casual style, is slim, 170cm tall, and weighs 65kg, the user enters this information on their device. The emotion analysis engine detects the user's smile and sends the data to the server as a positive emotion. The server then creates a prompt that reads, "170cm, 65kg, slim, casual outfit. User smiling," and sends it to the DALL-E API. The generated image is then sent to the device via the server, where the user can review it and provide specific feedback, such as "I'd like the pants to be a lighter color." This feedback is reflected in the next image generation.

[1045] The system aims to enable users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of a sentiment analysis engine will make personalized suggestions based on user emotions, further increasing user satisfaction.

[1046] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1047] Step 1:

[1048] The user uses the terminal to input information about his or her style (casual, formal, sporty, etc.) and body type (height, weight, body characteristics).

[1049] The input information is formatted into JSON format by the terminal, resulting in the following data:

[1050] json

[1051] {

[1052] "style": "casual",

[1053] "height": 170,

[1054] "weight": 65,

[1055] "body_shape": "slim"

[1056] }

[1057] The formatted data is sent in the next step.

[1058] Step 2:

[1059] The device sends the formatted user information to the server via an HTTP POST request, and the information is encrypted to ensure its security.

[1060] Input: User information in JSON format.

[1061] Output: User information sent to the server.

[1062] Step 3:

[1063] The device's camera and microphone are activated to capture the user's facial expressions and voice in real time.

[1064] The emotion analysis engine analyzes this data and generates the user's emotional information. For example, if the user is smiling, it will be analyzed as a "positive emotion."

[1065] Input: Facial and vocal data collected in real time.

[1066] Output: Emotion information from the emotion analysis engine.

[1067] Step 4:

[1068] The server generates a prompt based on the received user information and emotion information. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[1069] Input: User information, emotion information.

[1070] Output: The generated prompt.

[1071] Step 5:

[1072] The server sends the generated prompt to the image generation algorithm (DALL-E API) and requests it to generate a clothing image. The DALL-E API generates an image based on this prompt and returns the result to the server.

[1073] Input: The generated prompt.

[1074] Output: Generated image from DALL-E API.

[1075] Step 6:

[1076] The server receives the generated image returned from the DALL-E API and sends it to the user's device.

[1077] Input: Generated images from DALL-E API.

[1078] Output: The generated image sent to the device.

[1079] Step 7:

[1080] The terminal displays the received clothing image to the user.

[1081] The emotion analysis engine continues to analyze the user's facial expressions and voice, and the user enters their ratings and opinions through the feedback interface.

[1082] Input: Generated images from the server, user feedback.

[1083] Output: User input of ratings and feedback.

[1084] Step 8:

[1085] The device sends the feedback information provided by the user to the server, which receives the feedback information and stores it in a database. The feedback information is then reflected in future image generation and suggestions.

[1086] Input: User feedback information.

[1087] Output: Feedback information stored in a database.

[1088] This series of steps allows users to easily find the best clothing images based on their style and body type, significantly improving their online shopping experience. Furthermore, the sentiment analysis engine enables personalized recommendations, increasing user satisfaction.

[1089] (Application example 2)

[1090] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1091] In conventional online shopping, users cannot actually try on clothes, making it difficult to choose clothes that fit their body type and style. Furthermore, because the system does not take into account the user's emotions, personalized suggestions are not provided, resulting in an inconvenient shopping experience. The present invention aims to solve these problems by providing a system that incorporates virtual try-on and emotion recognition functions.

[1092] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1093] In this invention, the server includes means for accepting style and body type information input by the user, means for calling an image generation API based on the style and body type information to generate clothing images, means for sending the generated clothing images to the user's terminal, means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, means for recognizing the user's emotions in real time and optimizing image generation prompts based on the results, and means for converting the generated clothing images into 3D models and providing a virtual try-on. This allows users to easily select clothing that best suits their body type and style, significantly improving their shopping experience.

[1094] "Style and body type information entered by the user" refers to information that the user specifies about their preferred clothing style, height, weight, and body type characteristics.

[1095] An "image generation API" is an application programming interface for generating new images based on input information.

[1096] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[1097] "User device" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[1098] "User feedback" refers to the evaluations and opinions that users give to the generated clothing images.

[1099] "Recognizing emotions in real time" means analyzing the user's facial expressions, voice, etc. to instantly determine their current emotional state.

[1100] An "image generation prompt" is a textual instruction that instructs the image generation API to generate a particular image.

[1101] "3D modeling" is the process of converting two-dimensional image data into three-dimensional, three-dimensional data.

[1102] "Virtual try-on" is an experience that uses virtual reality and augmented reality technology to allow users to try on clothes in a virtual space without actually having to physically try them on.

[1103] This system generates appropriate clothing images based on style and body type information entered by the user, and then provides a virtual try-on experience based on the images. The system uses emotion recognition technology to analyze the user's real-time reactions and provide personalized suggestions.

[1104] System configuration

[1105] This system consists of a user terminal, a server, an emotion engine, an image generation API, and a virtual try-on engine.

[1106] User Device

[1107] The user device can be a smartphone, tablet, or head-mounted display (HMD), which provides an interface for users to input information about their style and body shape, and collects data through a camera and microphone for emotion recognition.

[1108] server

[1109] The server receives the information sent by the user and performs the necessary data processing. The server is responsible for the following:

[1110] 1. Receiving and formatting user information: The server receives the user information and converts it into a format suitable for the image generation API.

[1111] 2. Prompt generation: Generate prompts based on data from the emotion engine. For example, generate a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[1112] 3. Calling the image generation API: Call the image generation API using the generated prompt to generate an appropriate clothing image.

[1113] 4. 3D modeling of images: The generated clothing images are converted into 3D models and sent to the virtual fitting engine.

[1114] 5. Feedback analysis: Receive feedback from users, analyze it, and reflect it in future suggestions.

[1115] Emotion Engine

[1116] The emotion engine captures the user's facial expressions and voice in real time and analyzes their emotional state, optimizing prompt generation based on the user's preferences and reactions.

[1117] Image Generation API

[1118] The image generation API uses, for example, the OpenAI API, which generates 2D clothing images based on prompts received from the server.

[1119] Virtual Try-On Engine

[1120] The virtual try-on engine converts the generated 2D images into 3D models, allowing users to try on clothes in a virtual space.

[1121] Specific examples

[1122] If the user enters "170cm, 65kg, slim build, casual style" and the emotion engine detects a smiley face response, the server will generate the following prompt:

[1123] 170cm, 65kg, slim body, casual outfit. User smiles.

[1124] Using this prompt, the image generation API generates clothing images, converts them into 3D models, and sends them to the user's device. The virtual fitting engine allows users to enjoy a virtual fitting experience. At the same time, user feedback can be collected and reflected in future proposals, providing a more personalized experience.

[1125] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1126] Step 1:

[1127] The user uses the device to input information about their style and body type, including style (e.g., casual, formal), height, weight, and body characteristics. The input data is saved in JSON format.

[1128] Step 2:

[1129] The device sends the entered user information to the server. The server receives this information, formats it, and converts it into a format suitable for the image generation API. For example, if a user is 170cm tall, weighs 65kg, and has a casual style, it generates a prompt like this: "170cm tall, 65kg, slim build, casual outfit."

[1130] Step 3:

[1131] The device uses a camera and microphone to capture the user's emotions in real time and transmits them to the server. The emotion engine analyzes this data and determines the user's emotional state. For example, if the user is smiling, it will recognize the emotion as positive.

[1132] Step 4:

[1133] The server optimizes the image generation prompt based on data from the emotion engine. It adds emotion data and regenerates the prompt. For example, it adds the element "user smiling" to the prompt, such as "170cm, 65kg, slim build, casual outfit. User smiling."

[1134] Step 5:

[1135] The server sends the generated prompt to the image generation API, which generates an appropriate clothing image. The image generation API generates a clothing image based on the specified prompt and returns the result to the server.

[1136] Step 6:

[1137] The server converts the received clothing image into a 3D model and sends it to the virtual fitting engine, which converts the 2D clothing image into three-dimensional data.

[1138] Step 7:

[1139] The server sends the 3D modeled clothing data to the user's device, allowing the user to virtually try on the clothing. The device then displays the generated clothing image to the user, providing a virtual fitting experience through the virtual fitting engine.

[1140] Step 8:

[1141] The user enters feedback on the clothing generated through the virtual try-on experience, including opinions and ratings on the color and design of the clothing.

[1142] Step 9:

[1143] The device sends user feedback to the server, which analyzes it and uses it to generate the next image. The feedback information is stored in a database and used to make suggestions that better suit the user's preferences.

[1144] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1145] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1146] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1147] [Fourth embodiment]

[1148] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1149] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1150] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1151] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1152] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1153] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1154] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1155] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1156] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1157] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1158] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1159] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1160] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1161] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information on their style and body type. The system includes a server, a terminal, and an image generation API.

[1162] Entering user information

[1163] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[1164] Sending input information to the server

[1165] Once the input is complete, the device sends this information to the server in JSON format, for example:

[1166] json

[1167] {

[1168] "style": "casual",

[1169] "height": 170,

[1170] "weight": 65,

[1171] "body_shape": "slim"

[1172] }

[1173] DALL-E API call and image generation

[1174] The server calls the image generation API based on the received user information, converting the information into an appropriate format and using it as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit."

[1175] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[1176] Sending generated images from the server to the device

[1177] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[1178] Displaying images and collecting user feedback

[1179] The terminal displays the received clothing image to the user, and further provides the user with an interface for inputting ratings and feedback, allowing the user to provide feedback on their opinions and ratings of the displayed outfit image.

[1180] Sending and using feedback to the server

[1181] Feedback from users is sent to the server via their devices. The server analyzes this feedback and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[1182] Specific examples

[1183] For example, if a user prefers casual styles and has a slim build, the user enters their height of 170cm, weight of 65kg, and style of "casual" on the device. The device sends this information to the server, which then generates a prompt for "casual outfit for a 170cm, 65kg, slim build" and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[1184] This allows users to easily find the clothes that suit them best, greatly improving their online shopping experience.

[1185] The processing flow will be explained below.

[1186] Step 1:

[1187] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[1188] Step 2:

[1189] The device compiles the user's input information into JSON format. For example:

[1190] json

[1191] {

[1192] "style": "casual",

[1193] "height": 170,

[1194] "weight": 65,

[1195] "body_shape": "slim"

[1196] }

[1197] Step 3:

[1198] The device sends the user information compiled in JSON format to the server.

[1199] Step 4:

[1200] The server analyzes the received user information and generates an appropriate prompt. Example: "170cm, 65kg, slim build, casual outfit."

[1201] Step 5:

[1202] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[1203] Step 6:

[1204] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[1205] Step 7:

[1206] The server sends the clothing image received from the image generation API to the user's device.

[1207] Step 8:

[1208] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[1209] Step 9:

[1210] Users can input feedback about the outfit images displayed, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[1211] Step 10:

[1212] The device sends the user's feedback to the server.

[1213] Step 11:

[1214] The server receives user feedback and stores it in a database, which is then reflected in the next image generation.

[1215] This allows users to easily find the clothes that best suit their style and body type, enhancing their online shopping experience.

[1216] Example 1

[1217] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1218] Traditional online shopping has the problem of making it difficult for users to find clothes that perfectly fit their style and body type. This results in users wasting time and effort, and they often fail to find the products they want. Another issue is that if the generated clothing image does not meet the user's expectations, feedback is not reflected in the next image generation, resulting in the same problem repeating itself.

[1219] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1220] In this invention, the server includes means for accepting style and body type information entered by the user, means for generating a prompt based on the style and body type information, calling an image generation API, and generating a clothing image, and means for sending the generated clothing image to the user's device. This allows the user to efficiently obtain clothing images that suit their style and body type. Furthermore, by including means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, more appropriate suggestions that reflect the user's preferences and requests can be made.

[1221] A "user" is a person who uses the system to input style and body type information and view the generated clothing images.

[1222] "Style" refers to fashion categories based on the user's preferences, such as casual, formal, and sporty.

[1223] "Body type" refers to information about the user's physical appearance, such as their height, weight, and physical characteristics.

[1224] An "image generation API" is an application programming interface that generates an image based on a specified prompt.

[1225] A "prompt sentence" is a text sentence that issues a specific request to the image generation API.

[1226] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[1227] A "terminal" is an electronic device such as a computer or smartphone that a user uses to input information and receive and display generated clothing images.

[1228] "Feedback" refers to the evaluation or opinion provided by the user regarding the generated clothing image, which is reflected in the next image generation.

[1229] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. The system includes a server, a terminal, and an image generation API.

[1230] Users input information about their style and body type using their device. Specifically, style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device compiles this information in a predefined format and sends it to the server in JSON format.

[1231] The server generates a prompt based on the received user information and calls an image generation API (e.g., an API using an image generation model). In this case, a specific prompt such as "170 cm, 65 kg, slim build, casual outfit" is used. The server then sends a request including the generated prompt to the DALL-E API, requesting image generation.

[1232] The image generation API generates a clothing image based on the specified prompt and returns the result to the server. The server receives the clothing image returned from the image generation API and sends it to the user's device as an HTTP response.

[1233] The device displays the received clothing image on the screen and provides the user with an interface for inputting ratings and feedback. The user inputs their opinions and ratings about the displayed outfit image as feedback and presses the send button. Feedback can include specific comments such as "I wish the pants were a lighter color."

[1234] The device converts the user's feedback information into JSON format and sends it to the server as an HTTP POST request. The server then analyzes the feedback and stores it in a database. The saved feedback information is reflected in the next image generation, making suggestions that match the user's preferences.

[1235] For example, if a user likes casual style and has a slim build, the user can enter their height (170cm), weight (65kg), and style ("casual") into the terminal. When the user presses the "Submit" button, the terminal will generate the following prompt:

[1236] "Generate a casual outfit for a slim person who is 170cm tall and weighs 65kg."

[1237] This prompt is sent to the DALL-E API, and the generated image is sent to the device via the server. The user can review the generated image and provide feedback to reflect it in the next image generation.

[1238] In this way, users can easily find the clothes that suit them best, greatly improving their online shopping experience.

[1239] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1240] Step 1:

[1241] The user uses the device to input information about their own style and body type. Specifically, they input their style (e.g., casual, formal), height, weight, and body type characteristics into a form displayed on the device screen. The user's input is then acquired by the device.

[1242] Step 2:

[1243] The device converts the style and body information into a predefined format (JSON). Specifically, it formats the input information as follows:

[1244] json

[1245] {

[1246] "style": "casual",

[1247] "height": 170,

[1248] "weight": 65,

[1249] "body_shape": "slim"

[1250] }

[1251] The input here is the user's style and body type information, and the output is JSON format data.

[1252] Step 3:

[1253] The terminal sends the generated JSON data to the server. Specifically, it sends data to the server as an HTTP POST request. The input is JSON data, and the output is a message to the server indicating that data was successfully sent.

[1254] Step 4:

[1255] The server generates a prompt based on the received user information. Specifically, it uses a template in the server script to create a prompt like this:

[1256] "Casual outfit for a slim figure, 170cm, 65kg"

[1257] The input is the user's JSON data and the output is the prompt text.

[1258] Step 5:

[1259] The server sends the generated prompt text to the image generation API. Specifically, it sends the prompt text to the DALL-E API as an HTTP POST request. The input is the prompt text, and the output is a request sending success message to the image generation API.

[1260] Step 6:

[1261] The image generation API generates a clothing image based on the prompt received from the server and sends the result back to the server. Specifically, it generates image data using a generative AI model. The input is the prompt, and the output is the generated clothing image.

[1262] Step 7:

[1263] The server receives the clothing image returned from the image generation API and transfers it to the user's device. Specifically, it sends the image data to the user's device as an HTTP response. The input is the generated clothing image, and the output is a message that the image data was successfully sent to the user's device.

[1264] Step 8:

[1265] The device displays the received clothing image on the screen. Specifically, it renders an image display UI to allow the user to check the outfit image. The input is the generated clothing image, and the output is the display of the image.

[1266] Step 9:

[1267] The user provides feedback on the displayed outfit image. Specifically, the user inputs their opinion or rating into the rating and feedback input interface and presses the send button. The input is the feedback content, and the output is the feedback sending action.

[1268] Step 10:

[1269] The device converts the feedback information provided by the user into JSON format and sends it to the server. Specifically, it sends the feedback data as an HTTP POST request. The input is the feedback content, and the output is a feedback transmission success message to the server.

[1270] Step 11:

[1271] The server analyzes the feedback and stores it in a database. Specifically, it analyzes the feedback data and stores it as user preference data. The input is the feedback data, and the output is a message indicating that the analyzed data has been saved in the database.

[1272] Step 12:

[1273] The server reflects the feedback information when generating the next image. Specifically, it improves the next prompt generation based on the feedback stored in the database. The input is the stored feedback data, and the output is the improved prompt.

[1274] This series of processes provides the user with clothing images that are optimally suited to their style and body type, and by incorporating user feedback into future recommendations, the system delivers a more accurate and personalized online shopping experience.

[1275] (Application example 1)

[1276] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1277] It is difficult to find the perfect clothes when shopping online. In particular, the process of selecting clothes that fit the user's body type and style is complicated, resulting in an unsatisfactory shopping experience. Another issue is that personalization based on user feedback is not fully implemented, making it difficult to provide suggestions that match the user's preferences.

[1278] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1279] In this invention, the server includes means for receiving style and body type information input by the user, means for generating a prompt for the generation AI model based on the style and body type information and for generating clothing images by calling an image generation API, means for sending the generated clothing images to the user's device, and means for receiving user feedback and reflecting the feedback in generating the next clothing image. This allows the user to easily find clothing that suits them and receive even more personalized suggestions through the feedback.

[1280] "User" refers to an individual who uses this system to input their own style and body type information and receive suggestions for suitable clothing images.

[1281] "Style" refers to categories that indicate the type of clothing and fashion trends that users prefer, such as casual, formal, or sporty.

[1282] "Body type" refers to the user's physical characteristics and includes information about height, weight, and body type such as slim or chubby.

[1283] A "generative AI model" refers to an artificial intelligence model that has the technology to generate clothing images based on prompt text provided by the user.

[1284] A "prompt sentence" refers to a text request sentence that is generated based on the user's body type and style information and sent to the image generation API.

[1285] "Image generation API" refers to a programming interface for generating an image that meets specified conditions based on the received prompt text.

[1286] "Clothing Images" refers to images of clothing generated based on the user's body type and style using a generative AI model.

[1287] "Feedback" refers to the user inputting their opinions and evaluations of the generated clothing images, and refers to information that will be used to improve the next proposal.

[1288] "Terminal" refers to a device used by a user to input information and receive generated clothing images.

[1289] "Server" refers to the central computer system that receives user input information, calls the image generation API to generate garment images, and manages the collected feedback.

[1290] This invention provides a mail-order system that helps users find clothes that suit their body type and style when shopping online. The system includes a server, a terminal, and an image generation API.

[1291] First, the user uses the device to input their body type and style information, including height, weight, body characteristics, preferred style (e.g., casual, formal, sporty), etc. The device then formats the input information into a predefined format and sends it to the server.

[1292] Based on the received information, the server generates an appropriate prompt for the generative AI model. For example, a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is sent to the image generation API, which generates a clothing image based on the prompt. The software used for this is an HTTP client such as axios and the DALL-E API.

[1293] The generated clothing image is sent from the server to the user's device, where the user can view the image. In addition, an interface for inputting evaluations and feedback is provided to the user. The user can input their opinions and evaluations of the generated clothing image. For example, they can provide feedback such as "I would like the color of the pants to be lighter." This feedback information is sent to the server and reflected the next time an image is generated.

[1294] A concrete example is as follows: If a user prefers a casual style and has a slim build, the user enters their height as 170cm, weight as 65kg, and style as "casual" on their device. The device sends this information to the server, which then generates a prompt message saying "Casual outfit for a 170cm, 65kg, slim build" and sends it to the image generation API. The generated image is then sent to the device via the server, allowing the user to check the suitable outfit. At this time, the user enters feedback such as "I'd like the pants to be a lighter color," and the server takes this information into consideration the next time it generates a clothing image.

[1295] In this way, users can efficiently find the clothes that best suit their tastes and body type, and further personalization can be achieved through feedback, greatly improving the online shopping experience.

[1296] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1297] Step 1:

[1298] The user uses the device to enter their own style and body type information. Specifically, the user enters their height, weight, body type characteristics, and preferred style (e.g., casual, formal, sporty) into the device's input form. The entered information is formatted into a JSON format defined by the system. The input here is the user's body type and style information, and the output is formatted JSON data.

[1299] Step 2:

[1300] The device sends the entered information to the server. The data sent is in JSON format, for example:

[1301] json

[1302] {

[1303] "style": "casual",

[1304] "height": 170,

[1305] "weight": 65,

[1306] "body_shape": "slim"

[1307] }

[1308] The server receives this data, where the input is formatted JSON data and the output is the received user information.

[1309] Step 3:

[1310] The server generates a prompt based on the information it receives. Specifically, it creates a prompt such as "170cm, 65kg, slim build, casual outfit." This prompt is the request sent to the generative AI model (image generation API). The input here is JSON data of the user information, and the output is the prompt.

[1311] Step 4:

[1312] The server uses the generated prompt text to call an image generation API (e.g., DALL-E API). Specifically, it sends an API request including the prompt text and receives the URL of the generated clothing image. The input here is the prompt text, and the output is the URL of the generated clothing image.

[1313] Step 5:

[1314] The server sends the URL of the clothing image received from the image generation API to the user's device. The user's device receives this URL and displays the image on the screen. The input here is the URL of the generated clothing image, and the output is the clothing image displayed on the user's device.

[1315] Step 6:

[1316] The user inputs their evaluation and feedback for the displayed clothing image. Specifically, the user inputs feedback such as "I would like the color of the pants to be lighter" into the evaluation form and sends it from the terminal to the server. The input here is the user's feedback, and the output is the feedback data sent to the server.

[1317] Step 7:

[1318] The server analyzes the received feedback and reflects it in the next prompt generation. Specifically, the feedback data is stored in a database and taken into consideration when generating the next prompt for personalization. The input here is the user's feedback data, and the output is an updated prompt generation algorithm.

[1319] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1320] The virtual personal shopper system of the present invention aims to improve the online shopping experience by generating appropriate clothing images based on the user's input of information about their style and body type. In addition to the basic functions of inputting user information, calling an image generation API, displaying generated images, and collecting feedback, the system also incorporates an emotion engine that recognizes the user's emotions.

[1321] Entering user information

[1322] Users use the device to input information about their style and body type. Style includes categories such as casual, formal, and sporty, and body type information includes height, weight, and body characteristics. The device then compiles this information into a predefined format.

[1323] Sending input information to the server

[1324] Once the input is complete, the device sends this information to the server in JSON format, for example:

[1325] json

[1326] {

[1327] "style": "casual",

[1328] "height": 170,

[1329] "weight": 65,

[1330] "body_shape": "slim"

[1331] }

[1332] Emotion Engine Data Collection and Analysis

[1333] The device is equipped with a camera and microphone to capture the user's facial expressions and voice in real time. The emotion engine uses this data to analyze the user's emotions. For example, if the user is smiling, it is recognized as a positive emotion.

[1334] DALL-E API call and image generation

[1335] The server calls the image generation API based on the received user information and the analysis results from the emotion engine. At this time, the information is converted into an appropriate format and used as a prompt. For example, it generates a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[1336] The server then sends a request containing this prompt to the image generation API, which generates a clothing image based on the specified prompt and returns the result to the server.

[1337] Sending generated images from the server to the device

[1338] The server receives the clothing image returned from the image generation API and sends it to the user's device, allowing the user to check the generated clothing image.

[1339] Displaying images and collecting user feedback

[1340] The device displays the received clothing images to the user. It also provides the user with an interface for inputting ratings and feedback. The emotion engine then analyzes the user's emotions and uses the results to improve the displayed content. The user can provide feedback on their opinions and ratings of the displayed outfit images.

[1341] Sending and using feedback to the server

[1342] User feedback is sent to the server via the device. The server analyzes this feedback along with emotional data and stores it in a database. This feedback information is reflected in the next image generation, making suggestions that better suit the user's preferences.

[1343] Specific examples

[1344] For example, if a user prefers casual styles and has a slim build, they can enter their height of 170cm, weight of 65kg, and style of "casual" on their device. At this time, the device captures the user's emotions, and the emotion engine determines that they are reacting positively. The device sends this information to the server, which then generates a prompt that reads, "Casual outfit for a 170cm, 65kg, slim build. User smiling," and sends a request to the image generation API. The generated image is sent to the device via the server, allowing the user to check the suitable outfit. At the same time, if the user provides feedback such as "I wish the pants were a lighter color," the server can use this opinion as a reference for the next generation.

[1345] This will allow users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of the emotion engine will enable more personalized suggestions based on the user's emotions, aiming to further increase user satisfaction.

[1346] The processing flow will be explained below.

[1347] Step 1:

[1348] The user inputs information about their style and body type into the device. The user inputs information about their preferred style (e.g., casual, formal, sporty, etc.) and body type (height, weight, body characteristics, etc.) into the device's input form.

[1349] Step 2:

[1350] The device formats the entered user information into JSON format. Example:

[1351] json

[1352] {

[1353] "style": "casual",

[1354] "height": 170,

[1355] "weight": 65,

[1356] "body_shape": "slim"

[1357] }

[1358] Step 3:

[1359] The terminal transmits the formatted user information to the server.

[1360] Step 4:

[1361] The device captures the user's facial expressions and voice, and collects emotional data using the device's built-in camera and microphone.

[1362] Step 5:

[1363] The emotion engine analyzes data collected by the device to determine the user's emotional state. For example, a smiling user is recognized as a positive emotion.

[1364] Step 6:

[1365] The server generates a prompt based on the received user information and emotion data. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[1366] Step 7:

[1367] A request containing the server-generated prompt is sent to the image generation API in the form of an HTTP POST request.

[1368] Step 8:

[1369] The image generation API generates a clothing image based on the prompt and sends it back to the server.

[1370] Step 9:

[1371] The server sends the clothing image received from the image generation API to the user's device.

[1372] Step 10:

[1373] The device displays the received clothing image to the user, who can then view the generated outfit image on the device.

[1374] Step 11:

[1375] The device continues to use the emotion engine to monitor the user's emotional state, detecting the user's reaction to the displayed image (e.g., smiling, making a displeased face).

[1376] Step 12:

[1377] Users can input feedback on outfit images, such as "I'd like to change the color of the shirt" or "I'd like to change the style of the pants to a slim fit."

[1378] Step 13:

[1379] The device sends the user's feedback to the server.

[1380] Step 14:

[1381] The server stores the received feedback along with the emotion data and reflects it in the next image generation. The server updates the feedback database and uses it to generate the next prompt.

[1382] In this way, users can easily find the clothes that best suit their style and body type, improving their online shopping experience. In addition, the emotion engine can make suggestions based on the user's emotional state, further improving user satisfaction.

[1383] Example 2

[1384] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1385] In today's online shopping environment, users often find it difficult to find clothing that suits their style and body type. Furthermore, the lack of technology that makes personalized recommendations based on user emotions makes it difficult to provide a satisfying shopping experience. To address this issue, a system is needed that not only generates clothing images based on a user's style and body type information, but also makes personalized recommendations that take the user's emotions into account.

[1386] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1387] In this invention, the server includes means for accepting style and body type information input by the user, means for generating clothing images by invoking an image generation algorithm based on the style and body type information, and means for transmitting the generated clothing images to the user's information processing device. This makes it possible to provide accurate clothing images based on the user's style and body type, and to reflect feedback based on the results. Furthermore, by incorporating means for collecting emotional information using an emotion analysis engine that analyzes user emotions, and means for providing the emotional information along with the style and body type information as prompts to the generation AI model, personalized suggestions based on the user's emotions can be made, significantly improving the satisfaction of the online shopping experience.

[1388] "Style" is information that refers to the clothing category and design characteristics that a user prefers.

[1389] "Body type" refers to information that refers to the user's physical proportions, such as the user's height, weight, and body shape characteristics.

[1390] An "image generation algorithm" is a program or system that generates images of suitable clothing based on style and body type information provided by the user.

[1391] An "information processing device" is a terminal used by a user, specifically a device such as a smartphone or personal computer.

[1392] "Feedback" refers to information that indicates the evaluation or opinion that a user provides regarding the generated clothing image.

[1393] An "emotion analysis engine" is a program or system that collects and analyzes emotional information from a user's facial expressions and voice in real time.

[1394] "Emotional information" is data that indicates the emotional state of a user collected from their facial expressions and voice.

[1395] A "prompt" is input information given to a generative AI model, and refers to the form of data that includes style, body type, and emotional information.

[1396] A "generative AI model" is an artificial intelligence system that generates content, such as images, based on given prompts.

[1397] The virtual personal shopper system of the present invention aims to generate appropriate clothing images based on user input of personal style and body shape information. The system's main hardware consists of the user's device (smartphone or personal computer), a server, and a camera and microphone for running the emotion analysis engine. The software includes an image generation algorithm (e.g., DALL-E API), an emotion analysis engine for emotion analysis, and back-end services for data management and transmission.

[1398] Entering user information

[1399] The user uses the device to input their style (e.g., casual, formal, sporty) and body type (e.g., height, weight, body characteristics). The input information is compiled in JSON format and formatted as follows:

[1400] json

[1401] {

[1402] "style": "casual",

[1403] "height": 170,

[1404] "weight": 65,

[1405] "body_shape": "slim"

[1406] }

[1407] Sending input information to the server

[1408] Once the user has completed entering their information, the device sends it to the server, where the data is encrypted and transmitted over the internet, ensuring the security of the user's information.

[1409] Emotion Engine Data Collection and Analysis

[1410] The device's built-in camera and microphone are activated to capture the user's facial expressions and voice in real time. The collected data is sent to an emotion analysis engine, which generates emotional information from the user's facial expressions and voice. For example, if the user is smiling, it is recognized as a positive emotion.

[1411] DALL-E API call and image generation

[1412] The server generates a prompt based on the style and body type information entered by the user and the emotional information received from the emotion analysis engine. An example of a specific prompt is "170cm, 65kg, slim build, casual outfit. User is smiling." Using this prompt, the server sends a request to the DALL-E API. The DALL-E API generates a clothing image according to the specified prompt and sends the generated image back to the server.

[1413] Sending generated images from the server to the device

[1414] The server receives the image returned from the DALL-E API and sends it to the user's device, allowing the user to check the generated clothing image.

[1415] Displaying images and collecting user feedback

[1416] The device then displays the received clothing image to the user. During display, the emotion analysis engine continues to analyze the user's emotions and collects data on the user's reactions as appropriate. The user can then enter their ratings and feedback on the displayed image.

[1417] Sending and using feedback to the server

[1418] The user's feedback is sent to the server via the device. The server analyzes the received feedback and stores it in a database. This feedback information is used the next time an image is generated, and suggestions that better fit the user's preferences are made.

[1419] Specific examples

[1420] For example, if a user prefers casual style, is slim, 170cm tall, and weighs 65kg, the user enters this information on their device. The emotion analysis engine detects the user's smile and sends the data to the server as a positive emotion. The server then creates a prompt that reads, "170cm, 65kg, slim, casual outfit. User smiling," and sends it to the DALL-E API. The generated image is then sent to the device via the server, where the user can review it and provide specific feedback, such as "I'd like the pants to be a lighter color." This feedback is reflected in the next image generation.

[1421] The system aims to enable users to easily find the clothes that best suit them, significantly improving the online shopping experience. The introduction of a sentiment analysis engine will make personalized suggestions based on user emotions, further increasing user satisfaction.

[1422] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1423] Step 1:

[1424] The user uses the terminal to input information about his or her style (casual, formal, sporty, etc.) and body type (height, weight, body characteristics).

[1425] The input information is formatted into JSON format by the terminal, resulting in the following data:

[1426] json

[1427] {

[1428] "style": "casual",

[1429] "height": 170,

[1430] "weight": 65,

[1431] "body_shape": "slim"

[1432] }

[1433] The formatted data is sent in the next step.

[1434] Step 2:

[1435] The device sends the formatted user information to the server via an HTTP POST request, and the information is encrypted to ensure its security.

[1436] Input: User information in JSON format.

[1437] Output: User information sent to the server.

[1438] Step 3:

[1439] The device's camera and microphone are activated to capture the user's facial expressions and voice in real time.

[1440] The emotion analysis engine analyzes this data and generates the user's emotional information. For example, if the user is smiling, it will be analyzed as a "positive emotion."

[1441] Input: Facial and vocal data collected in real time.

[1442] Output: Emotion information from the emotion analysis engine.

[1443] Step 4:

[1444] The server generates a prompt based on the received user information and emotion information. For example, it creates a prompt such as "170cm, 65kg, slim build, casual outfit, user smiling."

[1445] Input: User information, emotion information.

[1446] Output: The generated prompt.

[1447] Step 5:

[1448] The server sends the generated prompt to the image generation algorithm (DALL-E API) and requests it to generate a clothing image. The DALL-E API generates an image based on this prompt and returns the result to the server.

[1449] Input: The generated prompt.

[1450] Output: Generated image from DALL-E API.

[1451] Step 6:

[1452] The server receives the generated image returned from the DALL-E API and sends it to the user's device.

[1453] Input: Generated images from DALL-E API.

[1454] Output: The generated image sent to the device.

[1455] Step 7:

[1456] The terminal displays the received clothing image to the user.

[1457] The emotion analysis engine continues to analyze the user's facial expressions and voice, and the user enters their ratings and opinions through the feedback interface.

[1458] Input: Generated images from the server, user feedback.

[1459] Output: User input of ratings and feedback.

[1460] Step 8:

[1461] The device sends the feedback information provided by the user to the server, which receives the feedback information and stores it in a database. The feedback information is then reflected in future image generation and suggestions.

[1462] Input: User feedback information.

[1463] Output: Feedback information stored in a database.

[1464] This series of steps allows users to easily find the best clothing images based on their style and body type, significantly improving their online shopping experience. Furthermore, the sentiment analysis engine enables personalized recommendations, increasing user satisfaction.

[1465] (Application example 2)

[1466] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1467] In conventional online shopping, users cannot actually try on clothes, making it difficult to choose clothes that fit their body type and style. Furthermore, because the system does not take into account the user's emotions, personalized suggestions are not provided, resulting in an inconvenient shopping experience. The present invention aims to solve these problems by providing a system that incorporates virtual try-on and emotion recognition functions.

[1468] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1469] In this invention, the server includes means for accepting style and body type information input by the user, means for calling an image generation API based on the style and body type information to generate clothing images, means for sending the generated clothing images to the user's terminal, means for accepting user feedback and reflecting the feedback in the generation of the next clothing image, means for recognizing the user's emotions in real time and optimizing image generation prompts based on the results, and means for converting the generated clothing images into 3D models and providing a virtual try-on. This allows users to easily select clothing that best suits their body type and style, significantly improving their shopping experience.

[1470] "Style and body type information entered by the user" refers to information that the user specifies about their preferred clothing style, height, weight, and body type characteristics.

[1471] An "image generation API" is an application programming interface for generating new images based on input information.

[1472] A "clothing image" is an image of virtual clothing generated based on the user's style and body type information.

[1473] "User device" refers to an electronic device used by a user, such as a smartphone, tablet, or computer.

[1474] "User feedback" refers to the evaluations and opinions that users give to the generated clothing images.

[1475] "Recognizing emotions in real time" means analyzing the user's facial expressions, voice, etc. to instantly determine their current emotional state.

[1476] An "image generation prompt" is a textual instruction that instructs the image generation API to generate a particular image.

[1477] "3D modeling" is the process of converting two-dimensional image data into three-dimensional, three-dimensional data.

[1478] "Virtual try-on" is an experience that uses virtual reality and augmented reality technology to allow users to try on clothes in a virtual space without actually having to physically try them on.

[1479] This system generates appropriate clothing images based on style and body type information entered by the user, and then provides a virtual try-on experience based on the images. The system uses emotion recognition technology to analyze the user's real-time reactions and provide personalized suggestions.

[1480] System configuration

[1481] This system consists of a user terminal, a server, an emotion engine, an image generation API, and a virtual try-on engine.

[1482] User Device

[1483] The user device can be a smartphone, tablet, or head-mounted display (HMD), which provides an interface for users to input information about their style and body shape, and collects data through a camera and microphone for emotion recognition.

[1484] server

[1485] The server receives the information sent by the user and performs the necessary data processing. The server is responsible for the following:

[1486] 1. Receiving and formatting user information: The server receives the user information and converts it into a format suitable for the image generation API.

[1487] 2. Prompt generation: Generate prompts based on data from the emotion engine. For example, generate a prompt such as "170cm, 65kg, slim build, casual outfit. User is smiling."

[1488] 3. Calling the image generation API: Call the image generation API using the generated prompt to generate an appropriate clothing image.

[1489] 4. 3D modeling of images: The generated clothing images are converted into 3D models and sent to the virtual fitting engine.

[1490] 5. Feedback analysis: Receive feedback from users, analyze it, and reflect it in future suggestions.

[1491] Emotion Engine

[1492] The emotion engine captures the user's facial expressions and voice in real time and analyzes their emotional state, optimizing prompt generation based on the user's preferences and reactions.

[1493] Image Generation API

[1494] The image generation API uses, for example, the OpenAI API, which generates 2D clothing images based on prompts received from the server.

[1495] Virtual Try-On Engine

[1496] The virtual try-on engine converts the generated 2D images into 3D models, allowing users to try on clothes in a virtual space.

[1497] Specific examples

[1498] If the user enters "170cm, 65kg, slim build, casual style" and the emotion engine detects a smiley face response, the server will generate the following prompt:

[1499] 170cm, 65kg, slim body, casual outfit. User smiles.

[1500] Using this prompt, the image generation API generates clothing images, converts them into 3D models, and sends them to the user's device. The virtual fitting engine allows users to enjoy a virtual fitting experience. At the same time, user feedback can be collected and reflected in future proposals, providing a more personalized experience.

[1501] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1502] Step 1:

[1503] The user uses the device to input information about their style and body type, including style (e.g., casual, formal), height, weight, and body characteristics. The input data is saved in JSON format.

[1504] Step 2:

[1505] The device sends the entered user information to the server. The server receives this information, formats it, and converts it into a format suitable for the image generation API. For example, if a user is 170cm tall, weighs 65kg, and has a casual style, it generates a prompt like this: "170cm tall, 65kg, slim build, casual outfit."

[1506] Step 3:

[1507] The device uses a camera and microphone to capture the user's emotions in real time and transmits them to the server. The emotion engine analyzes this data and determines the user's emotional state. For example, if the user is smiling, it will recognize the emotion as positive.

[1508] Step 4:

[1509] The server optimizes the image generation prompt based on data from the emotion engine. It adds emotion data and regenerates the prompt. For example, it adds the element "user smiling" to the prompt, such as "170cm, 65kg, slim build, casual outfit. User smiling."

[1510] Step 5:

[1511] The server sends the generated prompt to the image generation API, which generates an appropriate clothing image. The image generation API generates a clothing image based on the specified prompt and returns the result to the server.

[1512] Step 6:

[1513] The server converts the received clothing image into a 3D model and sends it to the virtual fitting engine, which converts the 2D clothing image into three-dimensional data.

[1514] Step 7:

[1515] The server sends the 3D modeled clothing data to the user's device, allowing the user to virtually try on the clothing. The device then displays the generated clothing image to the user, providing a virtual fitting experience through the virtual fitting engine.

[1516] Step 8:

[1517] The user enters feedback on the clothing generated through the virtual try-on experience, including opinions and ratings on the color and design of the clothing.

[1518] Step 9:

[1519] The device sends user feedback to the server, which analyzes it and uses it to generate the next image. The feedback information is stored in a database and used to make suggestions that better suit the user's preferences.

[1520] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1521] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1522] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1523] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1524] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1525] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1526] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1527] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1528] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1529] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1530] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1531] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1532] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1533] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1534] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1535] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1536] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1537] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1538] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1539] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1540] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1541] The following is further disclosed regarding the above embodiment.

[1542] (Claim 1)

[1543] means for accepting user-entered style and body type information;

[1544] a means for generating a clothing image by calling an image generation API based on the style and body type information;

[1545] means for transmitting the generated clothing image to a user's terminal;

[1546] a means for receiving user feedback and reflecting the feedback in generating the next clothing image;

[1547] A system including:

[1548] (Claim 2)

[1549] The system of claim 1 , further comprising means for formatting and converting the style and shape information into a format suitable for an image generation API.

[1550] (Claim 3)

[1551] The system according to claim 1 , further comprising: means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal.

[1552] "Example 1"

[1553] (Claim 1)

[1554] means for accepting user-entered style and body type information;

[1555] a means for generating a prompt sentence based on the style and body type information, calling an image generation API, and generating a clothing image;

[1556] means for transmitting the generated clothing image to a user's terminal;

[1557] a means for receiving user feedback and reflecting the feedback in generating the next clothing image;

[1558] A system including:

[1559] (Claim 2)

[1560] The system of claim 1 , further comprising means for formatting and converting the style and shape information into a format suitable for an image generation API.

[1561] (Claim 3)

[1562] The system according to claim 1 , further comprising: means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal.

[1563] "Application Example 1"

[1564] (Claim 1)

[1565] means for accepting user-entered style and body type information;

[1566] A means for generating a prompt sentence for a generation AI model based on the style and body type information, and calling an image generation API to generate a clothing image;

[1567] means for transmitting the generated clothing image to a user's terminal;

[1568] a means for receiving user feedback and reflecting the feedback in generating the next clothing image;

[1569] A system including:

[1570] (Claim 2)

[1571] The system of claim 1 , further comprising means for formatting the style and shape information and converting it into a prompt sentence suitable for an image generation API.

[1572] (Claim 3)

[1573] The system according to claim 1 , further comprising: means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal.

[1574] "Example 2: Combining Emotion Engines"

[1575] (Claim 1)

[1576] means for accepting user-entered style and body type information;

[1577] means for generating a clothing image by calling an image generation algorithm based on the style and body type information;

[1578] means for transmitting the generated clothing image to a user's information processing device;

[1579] a means for receiving user feedback and reflecting the feedback in generating the next clothing image;

[1580] A means for collecting emotion information by using an emotion analysis engine that analyzes user emotions;

[1581] means for providing said emotion information together with said style and body type information as prompts to a generative AI model;

[1582] A system including:

[1583] (Claim 2)

[1584] 10. The system of claim 1, further comprising means for formatting and converting said style and shape information and emotional information into a format suitable for an image generation algorithm.

[1585] (Claim 3)

[1586] The system according to claim 1 , further comprising: means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal.

[1587] "Application example 2 when combining emotion engines"

[1588] (Claim 1)

[1589] means for accepting user-entered style and body type information;

[1590] a means for generating a clothing image by calling an image generation API based on the style and body type information;

[1591] means for transmitting the generated clothing image to a user's terminal;

[1592] a means for receiving user feedback and reflecting the feedback in generating the next clothing image;

[1593] A means of recognizing user emotions in real time and optimizing image generation prompts based on the results;

[1594] A means for converting the generated clothing image into a 3D model and providing a virtual try-on;

[1595] A system including:

[1596] (Claim 2)

[1597] The system of claim 1 , further comprising means for formatting and converting the style and shape information into a format suitable for an image generation API.

[1598] (Claim 3)

[1599] The system according to claim 1 , further comprising: means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal. [Explanation of symbols]

[1600] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for accepting user-entered style and body type information; a means for generating a clothing image by calling an image generation API based on the style and body type information; means for transmitting the generated clothing image to a user's terminal; a means for receiving user feedback and reflecting the feedback in generating the next clothing image; A system including:

2. The system of claim 1 , further comprising means for formatting and converting the style and shape information into a format suitable for an image generation API.

3. The system according to claim 1 , further comprising means for providing an interface for evaluation and feedback input together with displaying the image when transmitting the generated clothing image to a user terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A