system
The system addresses the challenge of obtaining personalized beauty and fashion advice by analyzing facial features and preferences, integrating service booking, and facilitating payments, enhancing user experience through efficient and tailored suggestions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Consumers face difficulties in obtaining personalized beauty and fashion advice tailored to their individual characteristics, and the reservation and payment processes for salons and stores are complex and inefficient.
A system that analyzes user facial photographs and style preferences to suggest personalized styling, integrates service booking, and facilitates payment through a unified process using a generative AI model and third-party payment services.
Enables users to easily find suitable beauty and fashion styles, streamline service reservations, and complete payments efficiently, improving the user experience by personalizing suggestions based on individual characteristics.
Smart Images

Figure 2026068308000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern consumers spend a lot of time and effort finding the most suitable style for themselves, but there is a problem that it is difficult to obtain appropriate beauty and fashion advice based on individual characteristics during this process. In addition, since reservation and settlement procedures for beauty salons and fashion stores have to be carried out separately, there is also a problem of increased complexity. Such a situation has become a factor increasing the burden on users and deteriorating the usage experience.
Means for Solving the Problems
[0005] This invention provides a means for analyzing a user's facial photograph and style preferences to propose personalized styling, thereby enabling the suggestion of the optimal style tailored to individual characteristics. Furthermore, by incorporating a means for booking services based on the proposed style and a payment method related to the booking, it aims to unify complex procedures. With this system, users can easily find beauty and fashion styles that suit them and consistently book and pay for services.
[0006] "User" refers to an individual who uses this system to receive beauty and fashion services.
[0007] "Image data" refers to visual information, such as facial photographs, provided by users and used for analysis.
[0008] "Input means" refers to the interface through which users provide image data and preference information to this system.
[0009] "Analysis means" refers to the part that has the function of analyzing the user's facial features based on the input image data.
[0010] "Suggestion method" refers to a function that presents styling suitable for the user based on analyzed facial features and preference information.
[0011] "Reservation method" refers to a function for making reservations for services such as hair salons and fashion stores based on the proposed style.
[0012] "Payment method" refers to a function for making payments for reserved services, and includes those that can be linked with third-party payment services. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] Embodiments of the present invention include a system involving three main entities: a user, a terminal, and a server. This system enables users to efficiently find their own fashion style, book related beauty and fashion services, and complete the payment process in a single, integrated manner.
[0035] Specifically, the user first opens an application on their device, takes a photo of their face, and enters their personal fashion and makeup preferences. This information is securely sent to the server by the device.
[0036] The server performs analysis based on the received image data and preferences. A generative AI model analyzes facial features in detail and generates styling suggestions tailored to the user's individuality. For example, it can suggest optimal makeup and hairstyles based on the user's face shape and skin tone.
[0037] Next, the server returns the suggested style to the user, displaying it on the terminal. The user can review the suggested style and, if necessary, make a reservation at the suggested hair salon or fashion store. This reservation process involves the user entering their desired date and time through the terminal, and the server registering the reservation information in the relevant store's system.
[0038] Furthermore, users can complete payment simultaneously with their reservation using the payment methods provided on their device. This payment process is handled securely and quickly through a server in conjunction with a third-party payment service.
[0039] For example, if a user prefers blue-toned clothing and has a round face, the server analyzes these features and suggests an outfit such as "blue top and slim jeans." The user can then book a specific haircut service at a salon that matches that style. Through this entire process, users can quickly select a style that suits them and streamline the booking and payment process.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user launches the application on their device and opens the screen for taking a photo of their face. Following the instructions, the user takes the photo and registers it. The user also selects or enters their fashion and makeup preferences.
[0043] Step 2:
[0044] The device encrypts the user's provided facial photo and entered preference data, and sends it to the server via secure communication.
[0045] Step 3:
[0046] The server inputs the received facial photo data into an AI model, which analyzes features such as facial shape and skin tone. Furthermore, it analyzes the user's preferences and combines them with the analysis results to generate optimal makeup, hairstyles, and fashion coordinates.
[0047] Step 4:
[0048] The server compiles the styling suggestions into a data format and sends it to the terminal. This result includes specific style images suggested to the user and related service options.
[0049] Step 5:
[0050] The terminal displays the suggestions sent from the server within the application, allowing the user to review the suggestions in detail.
[0051] Step 6:
[0052] If the user likes the suggested style and decides to book the corresponding service, they proceed to the application's booking page. The user selects a service provider, enters their desired date and time, and submits a booking request.
[0053] Step 7:
[0054] The terminal sends the user's reservation information to the server. The server registers the reservation information in the relevant facility's system and obtains a reservation confirmation.
[0055] Step 8:
[0056] The terminal or server executes the payment using the selected payment method. The payment information is shared with a third-party payment provider, and the process is completed. The user receives a notification on the application that the reservation and payment have been completed.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] In modern society, individual fashion styles and beauty choices are diversifying, and there is a need to quickly and reliably find the optimal style that suits each person's individuality and preferences. However, traditional methods have the problem of being time-consuming and troublesome each time individual services or stores are used. Furthermore, ensuring safety and efficiency in the reservation and payment processes has also been a challenge.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for using a generative AI model to analyze the user's image information in detail, suggestion means for making suggestions based on recognized features and preference data, and means for reflecting the user's preference information using prompt sentences. This allows the user to quickly receive style suggestions tailored to them and to complete the entire process from reservation to payment securely and consistently.
[0062] "User image information" refers to digital data, including facial photographs, taken by the user using their device, and is used for personal identification and styling suggestions.
[0063] A "generative AI model" is an artificial intelligence model that analyzes facial features in detail and suggests the most suitable fashion style for the user.
[0064] A "server" is a central computer system that processes user image information and preference data, uses a generative AI model to suggest styles, and works in conjunction with other systems to handle reservations and payments.
[0065] A "prompt message" is an instruction used to have a generative AI model perform a specific analysis or make a suggestion, and it is an input in a sentence format that reflects the user's preferences and characteristics.
[0066] The "proposal method" is a function that uses a generative AI model to generate the optimal style based on the user's facial features and preference data, and then provides it to the user.
[0067] "Reservation method" refers to a function that allows users to make reservations by linking with the systems of stores and facilities that provide services related to the proposed style.
[0068] "Payment method" refers to a function that allows for secure and efficient payment related to reservations through a third-party payment service.
[0069] To implement this invention, a system is configured in which a user, a terminal, and a server work together. First, the user launches a dedicated application on their terminal. The terminal accepts a facial photograph taken by the user using the camera as input. The terminal can also input preference data, such as preferred fashion styles and makeup preferences.
[0070] The device encrypts and securely transmits collected facial photos and preference data to the server. The server utilizes a generative AI model to analyze the received data, performing a detailed analysis of features such as facial shape and skin tone. Based on this analysis, it generates suggestions for the user's optimal fashion style and beauty routine.
[0071] The generative AI model performs more effective analysis using prompts. A concrete example of a prompt would be, "Please suggest a suitable fashion style based on the user's facial features and preferences." This allows the AI model to suggest styles that maximize the user's unique characteristics.
[0072] The server sends the suggested style back to the terminal, allowing the user to review it. Based on the suggested style, the user can book a reservation at a relevant beauty salon or fashion store. This reservation is made through the terminal in conjunction with the system of the relevant facility.
[0073] Furthermore, payment processing can also be done via the terminal. The server completes payments securely and smoothly through external payment services. This allows users to efficiently select the style that best suits them and make reservations and payments easily through a series of processes.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The user launches a dedicated application on their device and takes a photo of their face. The photo is saved as image data on the device. The user also enters information about their fashion and makeup preferences through an input form on the device. This creates a single dataset containing both the photo and the preference data.
[0077] Step 2:
[0078] The device encrypts the collected facial images and preference data and sends them to the server using a secure communication protocol. This input data includes image files and text information. The dataset received by the server is used as foundational information for facial feature analysis and preference matching.
[0079] Step 3:
[0080] The server passes the received image data to a generating AI model, which analyzes facial features such as shape and skin tone. The AI model uses prompts to generate style suggestions tailored to the user. For example, it might use a prompt such as, "Please suggest the best fashion style based on the user's facial features."
[0081] Step 4:
[0082] The server compiles the generated style suggestions and sends them back to the user. This output data is formatted as fashion style and beauty suggestions and displayed on the terminal. The user reviews the displayed style suggestions and makes selections that suit them.
[0083] Step 5:
[0084] The user reviews the suggested style and makes a reservation at the suggested beauty salon or fashion store via their device. This reservation information is sent from the device to the server, which registers the reservation information in the system of the affiliated facility.
[0085] Step 6:
[0086] Users make reservations and payments simultaneously using their devices. The server receives the payment information transmitted from the device and securely processes the payment via an external payment service. This ensures that the entire process is completed smoothly.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Traditional technologies lack sufficient support for users to discover their own style, select appropriate products in physical stores, and make efficient purchases. This problem not only detracts from the user experience but also reduces purchasing intent. Therefore, there is a need for a system that allows users to easily find products that suit them in physical stores and learn about optimal coordination.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes an input means for acquiring user image data, an analysis means for analyzing the image data and recognizing facial features, a suggestion means for proposing an appropriate style based on the recognized features, a product identification means for acquiring product information and identifying products that match the user's features, and a presentation means for presenting coordination information related to the products identified by the product identification means. This makes it possible for users to improve their in-store shopping experience and efficiently select products that suit them.
[0092] A "user" refers to a person who operates the system and provides image data.
[0093] "Image data" refers to information that visually records a user's face or a product, and is the data that is subject to analysis.
[0094] "Input means" refers to devices or technologies used by a user to supply image data to a system.
[0095] "Analysis means" refers to devices or technologies that process acquired image data and analyze the user's facial features and preferences.
[0096] "Suggested means" refers to devices or technologies that present appropriate fashion styles and products to the user based on the analysis results.
[0097] A "reservation method" refers to a device or technology for reserving related services based on the proposed style.
[0098] A "payment method" refers to the equipment or technology used to process and pay for reserved services or goods.
[0099] "Product identification means" refers to devices or technologies used to identify products that match the characteristics of a user.
[0100] "Presentation means" refers to a device or technology that displays or notifies a user of information related to a specified product.
[0101] To realize this invention, a system is needed in which three entities—the user, the terminal, and the server—work together. The system begins with the user launching an application installed on the terminal.
[0102] A smartphone is recommended as the device, and it is equipped with a camera that can acquire image data in real time. Using this camera, users can take pictures of their faces or products they are interested in. The device securely transmits the captured image data to the server.
[0103] The server is equipped with advanced image analysis algorithms, using the open-source software OpenCV and TENSORFLOW® to perform highly accurate facial feature and preference analysis. The server also uses a generative AI model to suggest fashion styles based on the user's face and preferences. The generative AI model has the ability to output a variety of styling options based on the input of prompts. For example, with the prompt "Considering the user's face photo and fashion preferences, suggest outfits that go well with a denim jacket," the model generates the most suitable options for the user.
[0104] The analysis results and generated styling suggestions are fed back to the terminal and visually displayed to the user. This display includes images of suggested products and recommended styles. Based on this, the user can find the relevant products in physical stores and obtain additional information. The terminal can also be used to make reservations for suggested services and, if necessary, to complete payments. Payments are completed securely using a third-party payment service.
[0105] This entire process allows users to smoothly select the fashion items that best suit them and is supported throughout the purchase process.
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The user launches the application on their smartphone and uses the camera to capture images of their face or the product they are considering purchasing. The input is the image data obtained through the camera. The output is the image data itself. This process is completed by the user operating the camera according to the instructions on their device.
[0109] Step 2:
[0110] The terminal transmits the acquired image data to the server using a secure communication protocol. The input is the image data obtained by the terminal, and the output is the image data sent to the server. Encrypted communication is used for secure transmission.
[0111] Step 3:
[0112] The server analyzes the received image data using OpenCV and TensorFlow, performing data processing to extract features of the user's face and attributes of the product. The input is the image data sent to the server, and the output is the analyzed feature information. This operation uses a feature extraction algorithm.
[0113] Step 4:
[0114] The server uses the analyzed feature information to provide prompts to a generating AI model, which then generates styling suggestions best suited to the user. The input consists of the analyzed feature information and prompts, while the output is the generated styling suggestions. The prompt used is: "Considering the user's facial photo and fashion preferences, please suggest an outfit that goes well with a denim jacket."
[0115] Step 5:
[0116] The server sends the generated styling suggestions back to the terminal, which then visually displays them to the user. The input is the styling suggestions received from the server, and the output is the displayed suggestions. The terminal uses a UI-based display to present the suggestions clearly to the user.
[0117] Step 6:
[0118] The user reviews the presented styling suggestions and selects the products and services they like. Based on this selection, they decide whether or not to book the suggested services or products. The input is the displayed styling suggestions, and the output is the user's selections. This process is completed when the user makes a decision using the terminal's interface.
[0119] Step 7:
[0120] When a user makes a reservation, the terminal sends the reservation information to the server, which then registers the reservation information in conjunction with the system of the relevant service provider. The input is reservation information based on the user's selection, and the output is a notification to the service provider that registration is complete. The server performs the registration process with the facility's system via an API.
[0121] Step 8:
[0122] After completing the reservation, the user uses the same device to quickly and securely complete the associated payment. The device interacts with a third-party payment service via a server to complete the payment process. The input is the payment information related to the reservation, and the output is a notification that the payment has been completed. This operation is performed using the payment method selected by the user.
[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0124] The embodiment of this invention is a system centered on three components: a user, a terminal, and a server, with an emotion recognition engine incorporated as part of it. This system proposes appropriate beauty and fashion styles to the user and also enables the user to handle everything from service reservations to payment in a consistent manner.
[0125] The user first uses an application on their device to take a photo of their face and input their fashion and makeup preferences. In addition, there is a function to capture the user's facial expressions in real time, and emotional data obtained from this is collected. The device encrypts this information and sends it to the server.
[0126] The server analyzes the received facial photograph and extracts the user's characteristics. It also combines and analyzes the user's preference information and emotional data to suggest a style that best suits the user's current emotional state. For example, if the user has a relaxed expression, it can suggest a relaxed style.
[0127] The suggested styling is sent back to the terminal and presented to the user on the terminal. Here, the user can review the suggested style and make reservations at hair salons and related fashion stores. The reservation process is carried out by the server coordinating with the store's reservation system and registering the necessary information.
[0128] Furthermore, users can pay for their bookings instantly using their chosen payment method. This is done through a third-party payment service via the server, ensuring smooth processing.
[0129] For example, suppose a user wants to try something a little different from their usual style. Based on their facial photo and emotional data, the server suggests slightly brighter coloring and develops a fashion style that is consistent with the underlying intentions of their facial expressions. Based on this suggestion, the user selects a beauty service, and the booking and payment are completed instantly. This ensures that the user receives the service in a way that is best suited to them.
[0130] The following describes the processing flow.
[0131] Step 1:
[0132] The user launches an application on their device, takes a photo of their face, and records their current facial expression through the camera. At the same time, the user inputs their preferences regarding fashion and makeup.
[0133] Step 2:
[0134] The device encrypts the user's facial image, preference information, and real-time captured emotional data and sends it to the server.
[0135] Step 3:
[0136] The server first analyzes facial image data to extract features such as the user's face shape and skin tone. It also uses an emotion recognition engine to identify the user's emotions from their facial expression data. This information is then combined with preference data for a comprehensive analysis.
[0137] Step 4:
[0138] Based on the analysis results, the server generates styling that best suits not only the user's personality but also their emotional state. For example, if a depressed expression is detected, a colorful style that will cheer the user up will be suggested.
[0139] Step 5:
[0140] The server sends the suggested content back to the terminal, which then displays it to the user. The user reviews the suggested style, and if they like it, selects from the related service providers and proceeds with the reservation.
[0141] Step 6:
[0142] The terminal sends the reservation information selected by the user to the server, and the server works with the reservation system to complete the reservation for the desired beauty service.
[0143] Step 7:
[0144] The terminal or server processes the payment using the payment method selected by the user. The payment process is carried out through integration with a third-party payment service.
[0145] Step 8:
[0146] Users can confirm on their device that their reservation and payment have been successfully completed. User feedback is also recorded and used to improve future suggestions.
[0147] (Example 2)
[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0149] In response to the demands of modern, personalized beauty and fashion, there is a lack of systems that can accurately recognize users' emotions and preferences in real time and integrate everything from quick and appropriate style suggestions to booking and payment based on that information. Conventional systems lack the ability to suggest styles that take into account the user's real-time emotions, making it difficult to optimize the user experience.
[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0151] In this invention, the server includes an input means for acquiring user video information, an analysis means for analyzing the video information and extracting facial features, and a suggestion means for proposing an appropriate style based on the analyzed facial features and emotional information. This makes it possible to propose a personalized style that reflects the user's emotional state.
[0152] "Input means" refers to a device or method for acquiring video information from a user.
[0153] "Analysis means" refers to a device or method for processing acquired video information and extracting facial features and emotional information.
[0154] "Suggestion means" refers to a device or method for suggesting personalized styles to a user based on analyzed facial features and emotional information.
[0155] "Reservation means" refers to a device or method for reserving necessary operations based on the proposed format.
[0156] "Payment processing means" refers to a device or method for properly settling the costs associated with a reservation.
[0157] This system aims to automate beauty and fashion style suggestions by incorporating an emotion recognition engine, centered around three parties: the user, the device, and the server. Users take a photo of their face and input their fashion preferences through a dedicated app installed on a device such as a smartphone or tablet. The device encrypts this image information and preference data, along with real-time emotion data, using secure protocols such as SSL / TLS, and transmits it to the server.
[0158] The server decodes the received data and uses image analysis software to extract facial features. It also utilizes an emotion recognition engine to evaluate the user's emotional state and uses a generative AI model to suggest the optimal fashion style based on the collected data. For example, a relaxed mood might be suggested with a casual style, while a more adventurous mood might be suggested with a style featuring vibrant colors.
[0159] The suggested styles are sent back to the device and visually presented to the user. The user can review the suggestions and make reservations at hair salons or fashion establishments via the app. The server links the reservation information with the establishment's information system and registers it. It also links with payment service providers to process payments, allowing users to pay immediately after making a reservation. This enables users to smoothly enjoy the experience based on their chosen style.
[0160] For example, if a user enters "I want a bright and adventurous style," the server analyzes their facial photo and the frequency of their smiles, and suggests clothing in bright colors. Based on this suggestion, the user can select a suitable facility, make a reservation immediately, and complete the payment.
[0161] Examples of prompts for a generative AI model include the following:
[0162] "Please suggest beauty and fashion styles based on user emotional data and fashion preferences."
[0163] This embodiment enables personalized style selection based on emotions, improving the user experience.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] The user launches a dedicated application on their device and takes a photo of their face. This photo becomes the input data. Furthermore, the user inputs their preferences regarding fashion and makeup within the app. A real-time emotion capture function is also activated, collecting emotional data. All of this data is compiled into a single dataset.
[0167] Step 2:
[0168] The device encrypts the captured facial image, preferences, and emotional data. This process uses encryption protocols such as SSL / TLS to ensure data security. The encrypted data is then converted into data packets and sent to the server.
[0169] Step 3:
[0170] The server decodes the received data packets and analyzes the facial image to extract the user's facial features. An image analysis algorithm is used for this process. The input data is a facial image, and the output is data on the user's facial features. In addition, an emotion recognition engine is used to process data to determine the user's emotional state. The output in this process is data indicating the specific emotion and its intensity.
[0171] Step 4:
[0172] The server uses a generative AI model to suggest the optimal fashion style based on analyzed facial feature data and emotion data. This process takes into account the user's preferences and emotional state. For example, a calm mood will result in casual suggestions, while an adventurous mood will generate vibrant suggestions. The AI model generates specific suggestions based on prompt messages. The system then compiles these suggestions into a dataset and sends it to the terminal.
[0173] Step 5:
[0174] The terminal displays fashion style suggestions to the user. The user can review the suggested styles and decide to book beauty or fashion services based on them. Based on the user's selection, booking details are generated, and the process proceeds to the next step.
[0175] Step 6:
[0176] The server links information with the service provider's reservation system based on the user's selection. This information processing ensures the reservation is correctly registered. Furthermore, by linking with a third-party payment processing service via the server, a system is activated that allows the user to pay immediately upon booking. In this step, the reservation and payment information are the final outputs. As a result, the user is ready to receive the service.
[0177] (Application Example 2)
[0178] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0179] In modern retail, there is a demand for personalized services tailored to each customer's preferences and mood. However, achieving this requires significant human resources and is difficult to do instantly. Furthermore, there is a challenge in the lack of adequately developed systems that enable customers to make optimal choices immediately and increase their satisfaction.
[0180] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0181] In this invention, the server includes an input / output device for acquiring user image data, an information processing device for analyzing the image data and recognizing facial features, and a style suggestion device for suggesting appropriate styles based on the recognized features and user emotion data. This makes it possible to suggest optimal fashion styles based on the customer's emotions and preferences.
[0182] An "input / output device for acquiring user image data" is a device that takes image information from a user as input and provides output to the user as needed.
[0183] An "information processing device that analyzes image data to recognize facial features" is a device that receives image data, analyzes that data, and identifies specific facial features.
[0184] A "style suggestion device" is a device that has the function of suggesting fashion styles and other styles suitable for the user based on recognized facial features and user emotion data.
[0185] A "service support device" is a device that, based on a proposed style, assists users with selecting and trying on products at the service location, as well as providing other customer support.
[0186] A "reservation and payment device" is a device that accepts reservations for services selected by users and performs the associated payment processing.
[0187] An "information processing device" is a device that has the basic functions to receive and analyze various types of data and perform processing based on the results.
[0188] An "electronic commerce device" is a device that has the necessary functions to conduct transactions online and is used to carry out processes related to electronic commerce, such as payment and purchase procedures.
[0189] This system works in conjunction with the user, terminal, and server to provide the user with the optimal fashion style. The user acquires image data using input / output devices installed in the terminal. For example, a smartphone or smart glasses function as this terminal. The terminal transmits the acquired image data to an information processing unit. This information processing unit is equipped with an analysis system for recognizing facial features, and uses software such as OpenCV and AWS® Rekognition to analyze facial expressions and emotion data.
[0190] Based on analyzed feature and emotion data, the server suggests a suitable style to the user via a style suggestion device. This style suggestion uses a generative AI model to determine the fashion style best suited to the user's current emotions and preferences. For example, a user judged to be in a relaxed state might be suggested a casual and comfortable style.
[0191] Based on the proposed style, users can select and try on products through physical stores or online service support devices. A reservation and payment device operates through the terminal, allowing users to reserve and pay for their selected products and services. Smooth online transactions are ensured by integration with third-party payment systems such as Stripe.
[0192] An example of a prompt message is as follows: "Here is the customer's facial expression data. Please use the emotion recognition engine to suggest a fashion style that best suits the customer's current mood."
[0193] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0194] Step 1:
[0195] The user acquires their own image data using a device. Specifically, they take a photo of their face using the camera function of a smartphone or smart glasses. In this case, the input is the user's facial image data, and the output is the same image file.
[0196] Step 2:
[0197] The terminal transmits the acquired image data to the information processing device. Here, the terminal encrypts the image data and securely transmits it to the server. In this process, the input is facial image data, and the output is encrypted image data.
[0198] Step 3:
[0199] The server uses an information processing device to analyze the received image data. Specifically, it uses OpenCV and AWS Rekognition to extract facial features from the images and analyze emotions. The input to this process is encrypted image data, and the output is the analyzed facial features and emotion data.
[0200] Step 4:
[0201] The server uses a generative AI model based on the analyzed data to suggest a style suitable for the user. Here, the input is facial features and emotion data, and the output is a recommended fashion style. As a concrete example, a user with a relaxed expression might be suggested a casual style.
[0202] Step 5:
[0203] The terminal presents the user with style suggestions received from the server. After reviewing the suggested styles, the user decides whether to try them on in a store or make a purchase. The input to this process is the suggested fashion styles, and the output is the style information selected by the user.
[0204] Step 6:
[0205] The terminal operates the reservation and payment device based on the user's selection to make reservations and payments for services and products. Reservation information is transmitted from the terminal to the store's information processing device. The input to this process is the selected style information, and the output is reservation confirmation and payment completion information.
[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0209] [Second Embodiment]
[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0222] Embodiments of the present invention include a system involving three main entities: a user, a terminal, and a server. This system enables users to efficiently find their own fashion style, book related beauty and fashion services, and complete the payment process in a single, integrated manner.
[0223] Specifically, the user first opens an application on their device, takes a photo of their face, and enters their personal fashion and makeup preferences. This information is securely sent to the server by the device.
[0224] The server performs analysis based on the received image data and preferences. A generative AI model analyzes facial features in detail and generates styling suggestions tailored to the user's individuality. For example, it can suggest optimal makeup and hairstyles based on the user's face shape and skin tone.
[0225] Next, the server returns the suggested style to the user, displaying it on the terminal. The user can review the suggested style and, if necessary, make a reservation at the suggested hair salon or fashion store. This reservation process involves the user entering their desired date and time through the terminal, and the server registering the reservation information in the relevant store's system.
[0226] Furthermore, users can complete payment simultaneously with their reservation using the payment methods provided on their device. This payment process is handled securely and quickly through a server in conjunction with a third-party payment service.
[0227] For example, if a user prefers blue-toned clothing and has a round face, the server analyzes these features and suggests an outfit such as "blue top and slim jeans." The user can then book a specific haircut service at a salon that matches that style. Through this entire process, users can quickly select a style that suits them and streamline the booking and payment process.
[0228] The following describes the processing flow.
[0229] Step 1:
[0230] The user launches the application on their device and opens the screen for taking a photo of their face. Following the instructions, the user takes the photo and registers it. The user also selects or enters their fashion and makeup preferences.
[0231] Step 2:
[0232] The device encrypts the user's provided facial photo and entered preference data, and sends it to the server via secure communication.
[0233] Step 3:
[0234] The server inputs the received facial photo data into an AI model, which analyzes features such as facial shape and skin tone. Furthermore, it analyzes the user's preferences and combines them with the analysis results to generate optimal makeup, hairstyles, and fashion coordinates.
[0235] Step 4:
[0236] The server compiles the styling suggestions into a data format and sends it to the terminal. This result includes specific style images suggested to the user and related service options.
[0237] Step 5:
[0238] The terminal displays the suggestions sent from the server within the application, allowing the user to review the suggestions in detail.
[0239] Step 6:
[0240] If the user likes the suggested style and decides to book the corresponding service, they proceed to the application's booking page. The user selects a service provider, enters their desired date and time, and submits a booking request.
[0241] Step 7:
[0242] The terminal sends the user's reservation information to the server. The server registers the reservation information in the relevant facility's system and obtains a reservation confirmation.
[0243] Step 8:
[0244] The terminal or server executes the payment using the selected payment method. The payment information is shared with a third-party payment provider, and the process is completed. The user receives a notification on the application that the reservation and payment have been completed.
[0245] (Example 1)
[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0247] In modern society, individual fashion styles and beauty choices are diversifying, and there is a need to quickly and reliably find the optimal style that suits each person's individuality and preferences. However, traditional methods have the problem of being time-consuming and troublesome each time individual services or stores are used. Furthermore, ensuring safety and efficiency in the reservation and payment processes has also been a challenge.
[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0249] In this invention, the server includes means for using a generative AI model to analyze the user's image information in detail, suggestion means for making suggestions based on recognized features and preference data, and means for reflecting the user's preference information using prompt sentences. This allows the user to quickly receive style suggestions tailored to them and to complete the entire process from reservation to payment securely and consistently.
[0250] "User image information" refers to digital data, including facial photographs, taken by the user using their device, and is used for personal identification and styling suggestions.
[0251] A "generative AI model" is an artificial intelligence model that analyzes facial features in detail and suggests the most suitable fashion style for the user.
[0252] A "server" is a central computer system that processes user image information and preference data, uses a generative AI model to suggest styles, and works in conjunction with other systems to handle reservations and payments.
[0253] A "prompt message" is an instruction used to have a generative AI model perform a specific analysis or make a suggestion, and it is an input in a sentence format that reflects the user's preferences and characteristics.
[0254] The "proposal method" is a function that uses a generative AI model to generate the optimal style based on the user's facial features and preference data, and then provides it to the user.
[0255] "Reservation method" refers to a function that allows users to make reservations by linking with the systems of stores and facilities that provide services related to the proposed style.
[0256] "Payment method" refers to a function that allows for secure and efficient payment related to reservations through a third-party payment service.
[0257] To implement this invention, a system is configured in which a user, a terminal, and a server work together. First, the user launches a dedicated application on their terminal. The terminal accepts a facial photograph taken by the user using the camera as input. The terminal can also input preference data, such as preferred fashion styles and makeup preferences.
[0258] The device encrypts and securely transmits collected facial photos and preference data to the server. The server utilizes a generative AI model to analyze the received data, performing a detailed analysis of features such as facial shape and skin tone. Based on this analysis, it generates suggestions for the user's optimal fashion style and beauty routine.
[0259] The generative AI model performs more effective analysis using prompts. A concrete example of a prompt would be, "Please suggest a suitable fashion style based on the user's facial features and preferences." This allows the AI model to suggest styles that maximize the user's unique characteristics.
[0260] The server sends the suggested style back to the terminal, allowing the user to review it. Based on the suggested style, the user can book a reservation at a relevant beauty salon or fashion store. This reservation is made through the terminal in conjunction with the system of the relevant facility.
[0261] Furthermore, payment processing can also be done via the terminal. The server completes payments securely and smoothly through external payment services. This allows users to efficiently select the style that best suits them and make reservations and payments easily through a series of processes.
[0262] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0263] Step 1:
[0264] The user launches a dedicated application on their device and takes a photo of their face. The photo is saved as image data on the device. The user also enters information about their fashion and makeup preferences through an input form on the device. This creates a single dataset containing both the photo and the preference data.
[0265] Step 2:
[0266] The device encrypts the collected facial images and preference data and sends them to the server using a secure communication protocol. This input data includes image files and text information. The dataset received by the server is used as foundational information for facial feature analysis and preference matching.
[0267] Step 3:
[0268] The server passes the received image data to a generating AI model, which analyzes facial features such as shape and skin tone. The AI model uses prompts to generate style suggestions tailored to the user. For example, it might use a prompt such as, "Please suggest the best fashion style based on the user's facial features."
[0269] Step 4:
[0270] The server compiles the generated style suggestions and sends them back to the user. This output data is formatted as fashion style and beauty suggestions and displayed on the terminal. The user reviews the displayed style suggestions and makes selections that suit them.
[0271] Step 5:
[0272] The user reviews the suggested style and makes a reservation at the suggested beauty salon or fashion store via their device. This reservation information is sent from the device to the server, which registers the reservation information in the system of the affiliated facility.
[0273] Step 6:
[0274] Users make reservations and payments simultaneously using their devices. The server receives the payment information transmitted from the device and securely processes the payment via an external payment service. This ensures that the entire process is completed smoothly.
[0275] (Application Example 1)
[0276] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0277] In the prior art, there is a lack of support for users to find their own styles, select appropriate products in physical stores, and purchase them efficiently. This problem not only impairs the user experience but also becomes a factor in reducing the willingness to purchase. Therefore, there is a need for a system that enables users to easily find products suitable for themselves in physical stores and know the optimal coordination.
[0278] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0279] In this invention, the server includes an input means for acquiring the user's image data, an analysis means for analyzing the image data to recognize facial features, a proposal means for proposing an appropriate style based on the recognized features, a product identification means for acquiring product information and identifying products that match the user's features, and a presentation means for presenting coordination information related to the products identified by the product identification means. As a result, it becomes possible for users to improve their purchasing experience in physical stores and efficiently select products suitable for themselves.
[0280] The "user" refers to a person who operates the system and provides image data.
[0281] The "image data" is information that visually records the user's face and products, and is the data to be analyzed.
[0282] The "input means" is a device or technology used by the user to supply image data to the system.
[0283] The "analysis means" is a device or technology that processes the acquired image data and analyzes the user's facial features and preferences.
[0284] The "proposal means" is a device or technology that presents an appropriate fashion style and products to the user based on the analysis results.
[0285] The "reservation means" is a device or technology for reserving related services based on the proposed style.
[0286] The "payment means" is a device or technology for processing and paying the price of reserved services or goods.
[0287] The "product identification means" is a device or technology for identifying products that match the characteristics of the user.
[0288] The "presentation means" is a device or technology for displaying or notifying the user of information related to the identified product.
[0289] To realize the present invention, a system in which three entities, namely a user, a terminal, and a server, cooperate to function is required. The system starts when the user first launches an application installed on the terminal.
[0290] A smartphone is recommended for the terminal, and here it is equipped with a camera capable of acquiring image data in real time. Using this camera, the user can take pictures of their face or products of interest. The terminal securely transmits the captured image data to the server.
[0291] The server is equipped with advanced image analysis algorithms. Here, using the open-source OpenCV and TensorFlow, it performs high-precision face feature and preference analysis. The server further uses a generative AI model to propose a fashion style based on the user's face and preferences. The generative AI model has a function of outputting various styling options by inputting a prompt sentence. Taking the prompt "Please propose a coordination that suits a denim jacket considering the user's face photo and fashion preference." as an example, the model generates an optimal option for the user.
[0292] The analysis results and generated styling suggestions are fed back to the terminal and visually displayed to the user. This display includes images of suggested products and recommended styles. Based on this, the user can find the relevant products in physical stores and obtain additional information. The terminal can also be used to make reservations for suggested services and, if necessary, to complete payments. Payments are completed securely using a third-party payment service.
[0293] This entire process allows users to smoothly select the fashion items that best suit them and is supported throughout the purchase process.
[0294] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0295] Step 1:
[0296] The user launches the application on their smartphone and uses the camera to capture images of their face or the product they are considering purchasing. The input is the image data obtained through the camera. The output is the image data itself. This process is completed by the user operating the camera according to the instructions on their device.
[0297] Step 2:
[0298] The terminal transmits the acquired image data to the server using a secure communication protocol. The input is the image data obtained by the terminal, and the output is the image data sent to the server. Encrypted communication is used for secure transmission.
[0299] Step 3:
[0300] The server analyzes the received image data using OpenCV and TensorFlow, performing data processing to extract features of the user's face and attributes of the product. The input is the image data sent to the server, and the output is the analyzed feature information. This operation uses a feature extraction algorithm.
[0301] Step 4:
[0302] Based on the analyzed feature information, the server gives a prompt sentence to the generative AI model to generate an optimal styling proposal for the user. The input is the analyzed feature information and the prompt sentence, and the output is the generated styling proposal. A prompt sentence like "Please propose a coordination suitable for a denim jacket considering the user's face photo and fashion preference." is used.
[0303] Step 5:
[0304] The server returns the generated styling proposal to the terminal, and the terminal visually displays it to the user. The input is the styling proposal received from the server, and the output is the displayed proposal content. The terminal performs a display based on the UI to present the proposal to the user in an understandable manner.
[0305] Step 6:
[0306] The user checks the presented styling proposal and selects the products or services they like. Based on this selection, it is determined whether to make a reservation for the proposed services or products. The input is the displayed styling proposal, and the output is the user's selection content. This operation is completed when the user makes a decision using the terminal interface.
[0307] Step 7:
[0308] If the user makes a reservation, the terminal sends the reservation information to the server, and the server cooperates with the system of the relevant service facility to register the reservation information. The input is the reservation information based on the user's selection, and the output is the registration completion notification to the service facility. The server performs the registration process for the facility system via the API.
[0309] Step 8:
[0310] After completing the reservation, the user uses the same device to quickly and securely complete the associated payment. The device interacts with a third-party payment service via a server to complete the payment process. The input is the payment information related to the reservation, and the output is a notification that the payment has been completed. This operation is performed using the payment method selected by the user.
[0311] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0312] The embodiment of this invention is a system centered on three components: a user, a terminal, and a server, with an emotion recognition engine incorporated as part of it. This system proposes appropriate beauty and fashion styles to the user and also enables the user to handle everything from service reservations to payment in a consistent manner.
[0313] The user first uses an application on their device to take a photo of their face and input their fashion and makeup preferences. In addition, there is a function to capture the user's facial expressions in real time, and emotional data obtained from this is collected. The device encrypts this information and sends it to the server.
[0314] The server analyzes the received facial photograph and extracts the user's characteristics. It also combines and analyzes the user's preference information and emotional data to suggest a style that best suits the user's current emotional state. For example, if the user has a relaxed expression, it can suggest a relaxed style.
[0315] The suggested styling is sent back to the terminal and presented to the user on the terminal. Here, the user can review the suggested style and make reservations at hair salons and related fashion stores. The reservation process is carried out by the server coordinating with the store's reservation system and registering the necessary information.
[0316] Furthermore, users can pay for their bookings instantly using their chosen payment method. This is done through a third-party payment service via the server, ensuring smooth processing.
[0317] For example, suppose a user wants to try something a little different from their usual style. Based on their facial photo and emotional data, the server suggests slightly brighter coloring and develops a fashion style that is consistent with the underlying intentions of their facial expressions. Based on this suggestion, the user selects a beauty service, and the booking and payment are completed instantly. This ensures that the user receives the service in a way that is best suited to them.
[0318] The following describes the processing flow.
[0319] Step 1:
[0320] The user launches an application on their device, takes a photo of their face, and records their current facial expression through the camera. At the same time, the user inputs their preferences regarding fashion and makeup.
[0321] Step 2:
[0322] The device encrypts the user's facial image, preference information, and real-time captured emotional data and sends it to the server.
[0323] Step 3:
[0324] The server first analyzes facial image data to extract features such as the user's face shape and skin tone. It also uses an emotion recognition engine to identify the user's emotions from their facial expression data. This information is then combined with preference data for a comprehensive analysis.
[0325] Step 4:
[0326] Based on the analysis results, the server generates styling that best suits not only the user's personality but also their emotional state. For example, if a depressed expression is detected, a colorful style that will cheer the user up will be suggested.
[0327] Step 5:
[0328] The server sends the suggested content back to the terminal, which then displays it to the user. The user reviews the suggested style, and if they like it, selects from the related service providers and proceeds with the reservation.
[0329] Step 6:
[0330] The terminal sends the reservation information selected by the user to the server, and the server works with the reservation system to complete the reservation for the desired beauty service.
[0331] Step 7:
[0332] The terminal or server processes the payment using the payment method selected by the user. The payment process is carried out through integration with a third-party payment service.
[0333] Step 8:
[0334] Users can confirm on their device that their reservation and payment have been successfully completed. User feedback is also recorded and used to improve future suggestions.
[0335] (Example 2)
[0336] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0337] In response to the demands of modern, personalized beauty and fashion, there is a lack of systems that can accurately recognize users' emotions and preferences in real time and integrate everything from quick and appropriate style suggestions to booking and payment based on that information. Conventional systems lack the ability to suggest styles that take into account the user's real-time emotions, making it difficult to optimize the user experience.
[0338] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0339] In this invention, the server includes an input means for acquiring user video information, an analysis means for analyzing the video information and extracting facial features, and a suggestion means for proposing an appropriate style based on the analyzed facial features and emotional information. This makes it possible to propose a personalized style that reflects the user's emotional state.
[0340] "Input means" refers to a device or method for acquiring video information from a user.
[0341] "Analysis means" refers to a device or method for processing acquired video information and extracting facial features and emotional information.
[0342] "Suggestion means" refers to a device or method for suggesting personalized styles to a user based on analyzed facial features and emotional information.
[0343] "Reservation means" refers to a device or method for reserving necessary operations based on the proposed format.
[0344] "Payment processing means" refers to a device or method for properly settling the costs associated with a reservation.
[0345] This system aims to automate beauty and fashion style suggestions by incorporating an emotion recognition engine, centered around three parties: the user, the device, and the server. Users take a photo of their face and input their fashion preferences through a dedicated app installed on a device such as a smartphone or tablet. The device encrypts this image information and preference data, along with real-time emotion data, using secure protocols such as SSL / TLS, and transmits it to the server.
[0346] The server decodes the received data and uses image analysis software to extract facial features. It also utilizes an emotion recognition engine to evaluate the user's emotional state and uses a generative AI model to suggest the optimal fashion style based on the collected data. For example, a relaxed mood might be suggested with a casual style, while a more adventurous mood might be suggested with a style featuring vibrant colors.
[0347] The suggested styles are sent back to the device and visually presented to the user. The user can review the suggestions and make reservations at hair salons or fashion establishments via the app. The server links the reservation information with the establishment's information system and registers it. It also links with payment service providers to process payments, allowing users to pay immediately after making a reservation. This enables users to smoothly enjoy the experience based on their chosen style.
[0348] For example, if a user enters "I want a bright and adventurous style," the server analyzes their facial photo and the frequency of their smiles, and suggests clothing in bright colors. Based on this suggestion, the user can select a suitable facility, make a reservation immediately, and complete the payment.
[0349] Examples of prompts for a generative AI model include the following:
[0350] "Please suggest beauty and fashion styles based on user emotional data and fashion preferences."
[0351] This embodiment enables personalized style selection based on emotions, improving the user experience.
[0352] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0353] Step 1:
[0354] The user launches a dedicated application on their device and takes a photo of their face. This photo becomes the input data. Furthermore, the user inputs their preferences regarding fashion and makeup within the app. A real-time emotion capture function is also activated, collecting emotional data. All of this data is compiled into a single dataset.
[0355] Step 2:
[0356] The device encrypts the captured facial image, preferences, and emotional data. This process uses encryption protocols such as SSL / TLS to ensure data security. The encrypted data is then converted into data packets and sent to the server.
[0357] Step 3:
[0358] The server decodes the received data packets and analyzes the facial image to extract the user's facial features. An image analysis algorithm is used for this process. The input data is a facial image, and the output is data on the user's facial features. In addition, an emotion recognition engine is used to process data to determine the user's emotional state. The output in this process is data indicating the specific emotion and its intensity.
[0359] Step 4:
[0360] The server uses a generative AI model to suggest the optimal fashion style based on analyzed facial feature data and emotion data. This process takes into account the user's preferences and emotional state. For example, a calm mood will result in casual suggestions, while an adventurous mood will generate vibrant suggestions. The AI model generates specific suggestions based on prompt messages. The system then compiles these suggestions into a dataset and sends it to the terminal.
[0361] Step 5:
[0362] The terminal displays fashion style suggestions to the user. The user can review the suggested styles and decide to book beauty or fashion services based on them. Based on the user's selection, booking details are generated, and the process proceeds to the next step.
[0363] Step 6:
[0364] The server links information with the service provider's reservation system based on the user's selection. This information processing ensures the reservation is correctly registered. Furthermore, by linking with a third-party payment processing service via the server, a system is activated that allows the user to pay immediately upon booking. In this step, the reservation and payment information are the final outputs. As a result, the user is ready to receive the service.
[0365] (Application Example 2)
[0366] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0367] In modern retail, there is a demand for personalized services tailored to each customer's preferences and mood. However, achieving this requires significant human resources and is difficult to do instantly. Furthermore, there is a challenge in the lack of adequately developed systems that enable customers to make optimal choices immediately and increase their satisfaction.
[0368] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0369] In this invention, the server includes an input / output device for acquiring user image data, an information processing device for analyzing the image data and recognizing facial features, and a style suggestion device for suggesting appropriate styles based on the recognized features and user emotion data. This makes it possible to suggest optimal fashion styles based on the customer's emotions and preferences.
[0370] An "input / output device for acquiring user image data" is a device that takes image information from a user as input and provides output to the user as needed.
[0371] An "information processing device that analyzes image data to recognize facial features" is a device that receives image data, analyzes that data, and identifies specific facial features.
[0372] A "style suggestion device" is a device that has the function of suggesting fashion styles and other styles suitable for the user based on recognized facial features and user emotion data.
[0373] A "service support device" is a device that, based on a proposed style, assists users with selecting and trying on products at the service location, as well as providing other customer support.
[0374] A "reservation and payment device" is a device that accepts reservations for services selected by users and performs the associated payment processing.
[0375] An "information processing device" is a device that has the basic functions to receive and analyze various types of data and perform processing based on the results.
[0376] An "electronic commerce device" is a device that has the necessary functions to conduct transactions online and is used to carry out processes related to electronic commerce, such as payment and purchase procedures.
[0377] This system works in conjunction with the user, terminal, and server to provide the user with the optimal fashion style. The user acquires image data using input / output devices installed in the terminal. For example, a smartphone or smart glasses function as this terminal. The terminal transmits the acquired image data to an information processing unit. This information processing unit is equipped with an analysis system for recognizing facial features, and uses software such as OpenCV or AWS Rekognition to analyze facial expressions and emotion data.
[0378] Based on analyzed feature and emotion data, the server suggests a suitable style to the user via a style suggestion device. This style suggestion uses a generative AI model to determine the fashion style best suited to the user's current emotions and preferences. For example, a user judged to be in a relaxed state might be suggested a casual and comfortable style.
[0379] Based on the proposed style, users can select and try on products through physical stores or online service support devices. A reservation and payment device operates through the terminal, allowing users to reserve and pay for their selected products and services. Smooth online transactions are ensured by integration with third-party payment systems such as Stripe.
[0380] An example of a prompt message is as follows: "Here is the customer's facial expression data. Please use the emotion recognition engine to suggest a fashion style that best suits the customer's current mood."
[0381] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0382] Step 1:
[0383] The user acquires their own image data using a device. Specifically, they take a photo of their face using the camera function of a smartphone or smart glasses. In this case, the input is the user's facial image data, and the output is the same image file.
[0384] Step 2:
[0385] The terminal transmits the acquired image data to the information processing device. Here, the terminal encrypts the image data and securely transmits it to the server. In this process, the input is facial image data, and the output is encrypted image data.
[0386] Step 3:
[0387] The server uses an information processing device to analyze the received image data. Specifically, it uses OpenCV and AWS Rekognition to extract facial features from the images and analyze emotions. The input to this process is encrypted image data, and the output is the analyzed facial features and emotion data.
[0388] Step 4:
[0389] The server uses a generative AI model based on the analyzed data to suggest a style suitable for the user. Here, the input is facial features and emotion data, and the output is a recommended fashion style. As a concrete example, a user with a relaxed expression might be suggested a casual style.
[0390] Step 5:
[0391] The terminal presents the user with style suggestions received from the server. After reviewing the suggested styles, the user decides whether to try them on in a store or make a purchase. The input to this process is the suggested fashion styles, and the output is the style information selected by the user.
[0392] Step 6:
[0393] The terminal operates the reservation and payment device based on the user's selection to make reservations and payments for services and products. Reservation information is transmitted from the terminal to the store's information processing device. The input to this process is the selected style information, and the output is reservation confirmation and payment completion information.
[0394] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0395] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0396] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0397] [Third Embodiment]
[0398] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0399] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0400] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0401] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0402] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0404] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0405] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0406] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0407] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0408] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0409] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0410] Embodiments of the present invention include a system involving three main entities: a user, a terminal, and a server. This system enables users to efficiently find their own fashion style, book related beauty and fashion services, and complete the payment process in a single, integrated manner.
[0411] Specifically, the user first opens an application on their device, takes a photo of their face, and enters their personal fashion and makeup preferences. This information is securely sent to the server by the device.
[0412] The server performs analysis based on the received image data and preferences. A generative AI model analyzes facial features in detail and generates styling suggestions tailored to the user's individuality. For example, it can suggest optimal makeup and hairstyles based on the user's face shape and skin tone.
[0413] Next, the server returns the suggested style to the user, displaying it on the terminal. The user can review the suggested style and, if necessary, make a reservation at the suggested hair salon or fashion store. This reservation process involves the user entering their desired date and time through the terminal, and the server registering the reservation information in the relevant store's system.
[0414] Furthermore, users can complete payment simultaneously with their reservation using the payment methods provided on their device. This payment process is handled securely and quickly through a server in conjunction with a third-party payment service.
[0415] For example, if a user prefers blue-toned clothing and has a round face, the server analyzes these features and suggests an outfit such as "blue top and slim jeans." The user can then book a specific haircut service at a salon that matches that style. Through this entire process, users can quickly select a style that suits them and streamline the booking and payment process.
[0416] The following describes the processing flow.
[0417] Step 1:
[0418] The user launches the application on their device and opens the screen for taking a photo of their face. Following the instructions, the user takes the photo and registers it. The user also selects or enters their fashion and makeup preferences.
[0419] Step 2:
[0420] The device encrypts the user's provided facial photo and entered preference data, and sends it to the server via secure communication.
[0421] Step 3:
[0422] The server inputs the received facial photo data into an AI model, which analyzes features such as facial shape and skin tone. Furthermore, it analyzes the user's preferences and combines them with the analysis results to generate optimal makeup, hairstyles, and fashion coordinates.
[0423] Step 4:
[0424] The server compiles the styling suggestions into a data format and sends it to the terminal. This result includes specific style images suggested to the user and related service options.
[0425] Step 5:
[0426] The terminal displays the suggestions sent from the server within the application, allowing the user to review the suggestions in detail.
[0427] Step 6:
[0428] If the user likes the suggested style and decides to book the corresponding service, they proceed to the application's booking page. The user selects a service provider, enters their desired date and time, and submits a booking request.
[0429] Step 7:
[0430] The terminal sends the user's reservation information to the server. The server registers the reservation information in the relevant facility's system and obtains a reservation confirmation.
[0431] Step 8:
[0432] The terminal or server executes the payment using the selected payment method. The payment information is shared with a third-party payment provider, and the process is completed. The user receives a notification on the application that the reservation and payment have been completed.
[0433] (Example 1)
[0434] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0435] In modern society, individual fashion styles and beauty choices are diversifying, and there is a need to quickly and reliably find the optimal style that suits each person's individuality and preferences. However, traditional methods have the problem of being time-consuming and troublesome each time individual services or stores are used. Furthermore, ensuring safety and efficiency in the reservation and payment processes has also been a challenge.
[0436] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0437] In this invention, the server includes means for using a generative AI model to analyze the user's image information in detail, suggestion means for making suggestions based on recognized features and preference data, and means for reflecting the user's preference information using prompt sentences. This allows the user to quickly receive style suggestions tailored to them and to complete the entire process from reservation to payment securely and consistently.
[0438] "User image information" refers to digital data, including facial photographs, taken by the user using their device, and is used for personal identification and styling suggestions.
[0439] A "generative AI model" is an artificial intelligence model that analyzes facial features in detail and suggests the most suitable fashion style for the user.
[0440] A "server" is a central computer system that processes user image information and preference data, uses a generative AI model to suggest styles, and works in conjunction with other systems to handle reservations and payments.
[0441] A "prompt message" is an instruction used to have a generative AI model perform a specific analysis or make a suggestion, and it is an input in a sentence format that reflects the user's preferences and characteristics.
[0442] The "proposal method" is a function that uses a generative AI model to generate the optimal style based on the user's facial features and preference data, and then provides it to the user.
[0443] "Reservation method" refers to a function that allows users to make reservations by linking with the systems of stores and facilities that provide services related to the proposed style.
[0444] "Payment method" refers to a function that allows for secure and efficient payment related to reservations through a third-party payment service.
[0445] To implement this invention, a system is configured in which a user, a terminal, and a server work together. First, the user launches a dedicated application on their terminal. The terminal accepts a facial photograph taken by the user using the camera as input. The terminal can also input preference data, such as preferred fashion styles and makeup preferences.
[0446] The device encrypts and securely transmits collected facial photos and preference data to the server. The server utilizes a generative AI model to analyze the received data, performing a detailed analysis of features such as facial shape and skin tone. Based on this analysis, it generates suggestions for the user's optimal fashion style and beauty routine.
[0447] The generative AI model performs more effective analysis using prompts. A concrete example of a prompt would be, "Please suggest a suitable fashion style based on the user's facial features and preferences." This allows the AI model to suggest styles that maximize the user's unique characteristics.
[0448] The server sends the suggested style back to the terminal, allowing the user to review it. Based on the suggested style, the user can book a reservation at a relevant beauty salon or fashion store. This reservation is made through the terminal in conjunction with the system of the relevant facility.
[0449] Furthermore, payment processing can also be done via the terminal. The server completes payments securely and smoothly through external payment services. This allows users to efficiently select the style that best suits them and make reservations and payments easily through a series of processes.
[0450] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0451] Step 1:
[0452] The user launches a dedicated application on their device and takes a photo of their face. The photo is saved as image data on the device. The user also enters information about their fashion and makeup preferences through an input form on the device. This creates a single dataset containing both the photo and the preference data.
[0453] Step 2:
[0454] The device encrypts the collected facial images and preference data and sends them to the server using a secure communication protocol. This input data includes image files and text information. The dataset received by the server is used as foundational information for facial feature analysis and preference matching.
[0455] Step 3:
[0456] The server passes the received image data to a generating AI model, which analyzes facial features such as shape and skin tone. The AI model uses prompts to generate style suggestions tailored to the user. For example, it might use a prompt such as, "Please suggest the best fashion style based on the user's facial features."
[0457] Step 4:
[0458] The server compiles the generated style suggestions and sends them back to the user. This output data is formatted as fashion style and beauty suggestions and displayed on the terminal. The user reviews the displayed style suggestions and makes selections that suit them.
[0459] Step 5:
[0460] The user reviews the suggested style and makes a reservation at the suggested beauty salon or fashion store via their device. This reservation information is sent from the device to the server, which registers the reservation information in the system of the affiliated facility.
[0461] Step 6:
[0462] Users make reservations and payments simultaneously using their devices. The server receives the payment information transmitted from the device and securely processes the payment via an external payment service. This ensures that the entire process is completed smoothly.
[0463] (Application Example 1)
[0464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] Traditional technologies lack sufficient support for users to discover their own style, select appropriate products in physical stores, and make efficient purchases. This problem not only detracts from the user experience but also reduces purchasing intent. Therefore, there is a need for a system that allows users to easily find products that suit them in physical stores and learn about optimal coordination.
[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0467] In this invention, the server includes an input means for acquiring user image data, an analysis means for analyzing the image data and recognizing facial features, a suggestion means for proposing an appropriate style based on the recognized features, a product identification means for acquiring product information and identifying products that match the user's features, and a presentation means for presenting coordination information related to the products identified by the product identification means. This makes it possible for users to improve their in-store shopping experience and efficiently select products that suit them.
[0468] A "user" refers to a person who operates the system and provides image data.
[0469] "Image data" refers to information that visually records a user's face or a product, and is the data that is subject to analysis.
[0470] "Input means" refers to devices or technologies used by a user to supply image data to a system.
[0471] "Analysis means" refers to devices or technologies that process acquired image data and analyze the user's facial features and preferences.
[0472] "Suggested means" refers to devices or technologies that present appropriate fashion styles and products to the user based on the analysis results.
[0473] A "reservation method" refers to a device or technology for reserving related services based on the proposed style.
[0474] A "payment method" refers to the equipment or technology used to process and pay for reserved services or goods.
[0475] "Product identification means" refers to devices or technologies used to identify products that match the characteristics of a user.
[0476] "Presentation means" refers to a device or technology that displays or notifies a user of information related to a specified product.
[0477] To realize this invention, a system is needed in which three entities—the user, the terminal, and the server—work together. The system begins with the user launching an application installed on the terminal.
[0478] A smartphone is recommended as the device, and it is equipped with a camera that can acquire image data in real time. Using this camera, users can take pictures of their faces or products they are interested in. The device securely transmits the captured image data to the server.
[0479] The server is equipped with advanced image analysis algorithms, using open-source OpenCV and TensorFlow to perform highly accurate facial feature and preference analysis. The server also uses a generative AI model to suggest fashion styles based on the user's face and preferences. The generative AI model has the ability to output a variety of styling options based on the input of a prompt. For example, with the prompt "Consider the user's face photo and fashion preferences, suggest an outfit that goes well with a denim jacket," the model generates the most suitable options for the user.
[0480] The analysis results and generated styling suggestions are fed back to the terminal and visually displayed to the user. This display includes images of suggested products and recommended styles. Based on this, the user can find the relevant products in physical stores and obtain additional information. The terminal can also be used to make reservations for suggested services and, if necessary, to complete payments. Payments are completed securely using a third-party payment service.
[0481] This entire process allows users to smoothly select the fashion items that best suit them and is supported throughout the purchase process.
[0482] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0483] Step 1:
[0484] The user launches the application on their smartphone and uses the camera to capture images of their face or the product they are considering purchasing. The input is the image data obtained through the camera. The output is the image data itself. This process is completed by the user operating the camera according to the instructions on their device.
[0485] Step 2:
[0486] The terminal transmits the acquired image data to the server using a secure communication protocol. The input is the image data obtained by the terminal, and the output is the image data sent to the server. Encrypted communication is used for secure transmission.
[0487] Step 3:
[0488] The server analyzes the received image data using OpenCV and TensorFlow, performing data processing to extract features of the user's face and attributes of the product. The input is the image data sent to the server, and the output is the analyzed feature information. This operation uses a feature extraction algorithm.
[0489] Step 4:
[0490] The server uses the analyzed feature information to provide prompts to a generating AI model, which then generates styling suggestions best suited to the user. The input consists of the analyzed feature information and prompts, while the output is the generated styling suggestions. The prompt used is: "Considering the user's facial photo and fashion preferences, please suggest an outfit that goes well with a denim jacket."
[0491] Step 5:
[0492] The server sends the generated styling suggestions back to the terminal, which then visually displays them to the user. The input is the styling suggestions received from the server, and the output is the displayed suggestions. The terminal uses a UI-based display to present the suggestions clearly to the user.
[0493] Step 6:
[0494] The user reviews the presented styling suggestions and selects the products and services they like. Based on this selection, they decide whether or not to book the suggested services or products. The input is the displayed styling suggestions, and the output is the user's selections. This process is completed when the user makes a decision using the terminal's interface.
[0495] Step 7:
[0496] When a user makes a reservation, the terminal sends the reservation information to the server, which then registers the reservation information in conjunction with the system of the relevant service provider. The input is reservation information based on the user's selection, and the output is a notification to the service provider that registration is complete. The server performs the registration process with the facility's system via an API.
[0497] Step 8:
[0498] After completing the reservation, the user uses the same device to quickly and securely complete the associated payment. The device interacts with a third-party payment service via a server to complete the payment process. The input is the payment information related to the reservation, and the output is a notification that the payment has been completed. This operation is performed using the payment method selected by the user.
[0499] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0500] The embodiment of this invention is a system centered on three components: a user, a terminal, and a server, with an emotion recognition engine incorporated as part of it. This system proposes appropriate beauty and fashion styles to the user and also enables the user to handle everything from service reservations to payment in a consistent manner.
[0501] The user first uses an application on their device to take a photo of their face and input their fashion and makeup preferences. In addition, there is a function to capture the user's facial expressions in real time, and emotional data obtained from this is collected. The device encrypts this information and sends it to the server.
[0502] The server analyzes the received facial photograph and extracts the user's characteristics. It also combines and analyzes the user's preference information and emotional data to suggest a style that best suits the user's current emotional state. For example, if the user has a relaxed expression, it can suggest a relaxed style.
[0503] The suggested styling is sent back to the terminal and presented to the user on the terminal. Here, the user can review the suggested style and make reservations at hair salons and related fashion stores. The reservation process is carried out by the server coordinating with the store's reservation system and registering the necessary information.
[0504] Furthermore, users can pay for their bookings instantly using their chosen payment method. This is done through a third-party payment service via the server, ensuring smooth processing.
[0505] For example, suppose a user wants to try something a little different from their usual style. Based on their facial photo and emotional data, the server suggests slightly brighter coloring and develops a fashion style that is consistent with the underlying intentions of their facial expressions. Based on this suggestion, the user selects a beauty service, and the booking and payment are completed instantly. This ensures that the user receives the service in a way that is best suited to them.
[0506] The following describes the processing flow.
[0507] Step 1:
[0508] The user launches an application on their device, takes a photo of their face, and records their current facial expression through the camera. At the same time, the user inputs their preferences regarding fashion and makeup.
[0509] Step 2:
[0510] The device encrypts the user's facial image, preference information, and real-time captured emotional data and sends it to the server.
[0511] Step 3:
[0512] The server first analyzes facial image data to extract features such as the user's face shape and skin tone. It also uses an emotion recognition engine to identify the user's emotions from their facial expression data. This information is then combined with preference data for a comprehensive analysis.
[0513] Step 4:
[0514] Based on the analysis results, the server generates styling that best suits not only the user's personality but also their emotional state. For example, if a depressed expression is detected, a colorful style that will cheer the user up will be suggested.
[0515] Step 5:
[0516] The server sends the suggested content back to the terminal, which then displays it to the user. The user reviews the suggested style, and if they like it, selects from the related service providers and proceeds with the reservation.
[0517] Step 6:
[0518] The terminal sends the reservation information selected by the user to the server, and the server works with the reservation system to complete the reservation for the desired beauty service.
[0519] Step 7:
[0520] The terminal or server processes the payment using the payment method selected by the user. The payment process is carried out through integration with a third-party payment service.
[0521] Step 8:
[0522] Users can confirm on their device that their reservation and payment have been successfully completed. User feedback is also recorded and used to improve future suggestions.
[0523] (Example 2)
[0524] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0525] In response to the demands of modern, personalized beauty and fashion, there is a lack of systems that can accurately recognize users' emotions and preferences in real time and integrate everything from quick and appropriate style suggestions to booking and payment based on that information. Conventional systems lack the ability to suggest styles that take into account the user's real-time emotions, making it difficult to optimize the user experience.
[0526] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0527] In this invention, the server includes an input means for acquiring user video information, an analysis means for analyzing the video information and extracting facial features, and a suggestion means for proposing an appropriate style based on the analyzed facial features and emotional information. This makes it possible to propose a personalized style that reflects the user's emotional state.
[0528] "Input means" refers to a device or method for acquiring video information from a user.
[0529] "Analysis means" refers to a device or method for processing acquired video information and extracting facial features and emotional information.
[0530] "Suggestion means" refers to a device or method for suggesting personalized styles to a user based on analyzed facial features and emotional information.
[0531] "Reservation means" refers to a device or method for reserving necessary operations based on the proposed format.
[0532] "Payment processing means" refers to a device or method for properly settling the costs associated with a reservation.
[0533] This system aims to automate beauty and fashion style suggestions by incorporating an emotion recognition engine, centered around three parties: the user, the device, and the server. Users take a photo of their face and input their fashion preferences through a dedicated app installed on a device such as a smartphone or tablet. The device encrypts this image information and preference data, along with real-time emotion data, using secure protocols such as SSL / TLS, and transmits it to the server.
[0534] The server decodes the received data and uses image analysis software to extract facial features. It also utilizes an emotion recognition engine to evaluate the user's emotional state and uses a generative AI model to suggest the optimal fashion style based on the collected data. For example, a relaxed mood might be suggested with a casual style, while a more adventurous mood might be suggested with a style featuring vibrant colors.
[0535] The suggested styles are sent back to the device and visually presented to the user. The user can review the suggestions and make reservations at hair salons or fashion establishments via the app. The server links the reservation information with the establishment's information system and registers it. It also links with payment service providers to process payments, allowing users to pay immediately after making a reservation. This enables users to smoothly enjoy the experience based on their chosen style.
[0536] For example, if a user enters "I want a bright and adventurous style," the server analyzes their facial photo and the frequency of their smiles, and suggests clothing in bright colors. Based on this suggestion, the user can select a suitable facility, make a reservation immediately, and complete the payment.
[0537] Examples of prompts for a generative AI model include the following:
[0538] "Please suggest beauty and fashion styles based on user emotional data and fashion preferences."
[0539] This embodiment enables personalized style selection based on emotions, improving the user experience.
[0540] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0541] Step 1:
[0542] The user launches a dedicated application on their device and takes a photo of their face. This photo becomes the input data. Furthermore, the user inputs their preferences regarding fashion and makeup within the app. A real-time emotion capture function is also activated, collecting emotional data. All of this data is compiled into a single dataset.
[0543] Step 2:
[0544] The device encrypts the captured facial image, preferences, and emotional data. This process uses encryption protocols such as SSL / TLS to ensure data security. The encrypted data is then converted into data packets and sent to the server.
[0545] Step 3:
[0546] The server decodes the received data packets and analyzes the facial image to extract the user's facial features. An image analysis algorithm is used for this process. The input data is a facial image, and the output is data on the user's facial features. In addition, an emotion recognition engine is used to process data to determine the user's emotional state. The output in this process is data indicating the specific emotion and its intensity.
[0547] Step 4:
[0548] The server uses a generative AI model to suggest the optimal fashion style based on analyzed facial feature data and emotion data. This process takes into account the user's preferences and emotional state. For example, a calm mood will result in casual suggestions, while an adventurous mood will generate vibrant suggestions. The AI model generates specific suggestions based on prompt messages. The system then compiles these suggestions into a dataset and sends it to the terminal.
[0549] Step 5:
[0550] The terminal displays fashion style suggestions to the user. The user can review the suggested styles and decide to book beauty or fashion services based on them. Based on the user's selection, booking details are generated, and the process proceeds to the next step.
[0551] Step 6:
[0552] The server links information with the service provider's reservation system based on the user's selection. This information processing ensures the reservation is correctly registered. Furthermore, by linking with a third-party payment processing service via the server, a system is activated that allows the user to pay immediately upon booking. In this step, the reservation and payment information are the final outputs. As a result, the user is ready to receive the service.
[0553] (Application Example 2)
[0554] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0555] In modern retail, there is a demand for personalized services tailored to each customer's preferences and mood. However, achieving this requires significant human resources and is difficult to do instantly. Furthermore, there is a challenge in the lack of adequately developed systems that enable customers to make optimal choices immediately and increase their satisfaction.
[0556] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0557] In this invention, the server includes an input / output device for acquiring user image data, an information processing device for analyzing the image data and recognizing facial features, and a style suggestion device for suggesting appropriate styles based on the recognized features and user emotion data. This makes it possible to suggest optimal fashion styles based on the customer's emotions and preferences.
[0558] An "input / output device for acquiring user image data" is a device that takes image information from a user as input and provides output to the user as needed.
[0559] An "information processing device that analyzes image data to recognize facial features" is a device that receives image data, analyzes that data, and identifies specific facial features.
[0560] A "style suggestion device" is a device that has the function of suggesting fashion styles and other styles suitable for the user based on recognized facial features and user emotion data.
[0561] A "service support device" is a device that, based on a proposed style, assists users with selecting and trying on products at the service location, as well as providing other customer support.
[0562] A "reservation and payment device" is a device that accepts reservations for services selected by users and performs the associated payment processing.
[0563] An "information processing device" is a device that has the basic functions to receive and analyze various types of data and perform processing based on the results.
[0564] An "electronic commerce device" is a device that has the necessary functions to conduct transactions online and is used to carry out processes related to electronic commerce, such as payment and purchase procedures.
[0565] This system works in conjunction with the user, terminal, and server to provide the user with the optimal fashion style. The user acquires image data using input / output devices installed in the terminal. For example, a smartphone or smart glasses function as this terminal. The terminal transmits the acquired image data to an information processing unit. This information processing unit is equipped with an analysis system for recognizing facial features, and uses software such as OpenCV or AWS Rekognition to analyze facial expressions and emotion data.
[0566] Based on analyzed feature and emotion data, the server suggests a suitable style to the user via a style suggestion device. This style suggestion uses a generative AI model to determine the fashion style best suited to the user's current emotions and preferences. For example, a user judged to be in a relaxed state might be suggested a casual and comfortable style.
[0567] Based on the proposed style, users can select and try on products through physical stores or online service support devices. A reservation and payment device operates through the terminal, allowing users to reserve and pay for their selected products and services. Smooth online transactions are ensured by integration with third-party payment systems such as Stripe.
[0568] An example of a prompt message is as follows: "Here is the customer's facial expression data. Please use the emotion recognition engine to suggest a fashion style that best suits the customer's current mood."
[0569] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0570] Step 1:
[0571] The user acquires their own image data using a device. Specifically, they take a photo of their face using the camera function of a smartphone or smart glasses. In this case, the input is the user's facial image data, and the output is the same image file.
[0572] Step 2:
[0573] The terminal transmits the acquired image data to the information processing device. Here, the terminal encrypts the image data and securely transmits it to the server. In this process, the input is facial image data, and the output is encrypted image data.
[0574] Step 3:
[0575] The server uses an information processing device to analyze the received image data. Specifically, it uses OpenCV and AWS Rekognition to extract facial features from the images and analyze emotions. The input to this process is encrypted image data, and the output is the analyzed facial features and emotion data.
[0576] Step 4:
[0577] The server uses a generative AI model based on the analyzed data to suggest a style suitable for the user. Here, the input is facial features and emotion data, and the output is a recommended fashion style. As a concrete example, a user with a relaxed expression might be suggested a casual style.
[0578] Step 5:
[0579] The terminal presents the user with style suggestions received from the server. After reviewing the suggested styles, the user decides whether to try them on in a store or make a purchase. The input to this process is the suggested fashion styles, and the output is the style information selected by the user.
[0580] Step 6:
[0581] The terminal operates the reservation and payment device based on the user's selection to make reservations and payments for services and products. Reservation information is transmitted from the terminal to the store's information processing device. The input to this process is the selected style information, and the output is reservation confirmation and payment completion information.
[0582] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0583] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0584] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0585] [Fourth Embodiment]
[0586] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0587] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0588] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0589] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0590] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0591] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0592] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0593] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0594] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0595] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0596] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0597] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0598] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0599] Embodiments of the present invention include a system involving three main entities: a user, a terminal, and a server. This system enables users to efficiently find their own fashion style, book related beauty and fashion services, and complete the payment process in a single, integrated manner.
[0600] Specifically, the user first opens an application on their device, takes a photo of their face, and enters their personal fashion and makeup preferences. This information is securely sent to the server by the device.
[0601] The server performs analysis based on the received image data and preferences. A generative AI model analyzes facial features in detail and generates styling suggestions tailored to the user's individuality. For example, it can suggest optimal makeup and hairstyles based on the user's face shape and skin tone.
[0602] Next, the server returns the suggested style to the user, displaying it on the terminal. The user can review the suggested style and, if necessary, make a reservation at the suggested hair salon or fashion store. This reservation process involves the user entering their desired date and time through the terminal, and the server registering the reservation information in the relevant store's system.
[0603] Furthermore, users can complete payment simultaneously with their reservation using the payment methods provided on their device. This payment process is handled securely and quickly through a server in conjunction with a third-party payment service.
[0604] For example, if a user prefers blue-toned clothing and has a round face, the server analyzes these features and suggests an outfit such as "blue top and slim jeans." The user can then book a specific haircut service at a salon that matches that style. Through this entire process, users can quickly select a style that suits them and streamline the booking and payment process.
[0605] The following describes the processing flow.
[0606] Step 1:
[0607] The user launches the application on their device and opens the screen for taking a photo of their face. Following the instructions, the user takes the photo and registers it. The user also selects or enters their fashion and makeup preferences.
[0608] Step 2:
[0609] The device encrypts the user's provided facial photo and entered preference data, and sends it to the server via secure communication.
[0610] Step 3:
[0611] The server inputs the received facial photo data into an AI model, which analyzes features such as facial shape and skin tone. Furthermore, it analyzes the user's preferences and combines them with the analysis results to generate optimal makeup, hairstyles, and fashion coordinates.
[0612] Step 4:
[0613] The server compiles the styling suggestions into a data format and sends it to the terminal. This result includes specific style images suggested to the user and related service options.
[0614] Step 5:
[0615] The terminal displays the suggestions sent from the server within the application, allowing the user to review the suggestions in detail.
[0616] Step 6:
[0617] If the user likes the suggested style and decides to book the corresponding service, they proceed to the application's booking page. The user selects a service provider, enters their desired date and time, and submits a booking request.
[0618] Step 7:
[0619] The terminal sends the user's reservation information to the server. The server registers the reservation information in the relevant facility's system and obtains a reservation confirmation.
[0620] Step 8:
[0621] The terminal or server executes the payment using the selected payment method. The payment information is shared with a third-party payment provider, and the process is completed. The user receives a notification on the application that the reservation and payment have been completed.
[0622] (Example 1)
[0623] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0624] In modern society, individual fashion styles and beauty choices are diversifying, and there is a need to quickly and reliably find the optimal style that suits each person's individuality and preferences. However, traditional methods have the problem of being time-consuming and troublesome each time individual services or stores are used. Furthermore, ensuring safety and efficiency in the reservation and payment processes has also been a challenge.
[0625] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0626] In this invention, the server includes means for using a generative AI model to analyze the user's image information in detail, suggestion means for making suggestions based on recognized features and preference data, and means for reflecting the user's preference information using prompt sentences. This allows the user to quickly receive style suggestions tailored to them and to complete the entire process from reservation to payment securely and consistently.
[0627] "User image information" refers to digital data, including facial photographs, taken by the user using their device, and is used for personal identification and styling suggestions.
[0628] A "generative AI model" is an artificial intelligence model that analyzes facial features in detail and suggests the most suitable fashion style for the user.
[0629] A "server" is a central computer system that processes user image information and preference data, uses a generative AI model to suggest styles, and works in conjunction with other systems to handle reservations and payments.
[0630] A "prompt message" is an instruction used to have a generative AI model perform a specific analysis or make a suggestion, and it is an input in a sentence format that reflects the user's preferences and characteristics.
[0631] The "proposal method" is a function that uses a generative AI model to generate the optimal style based on the user's facial features and preference data, and then provides it to the user.
[0632] "Reservation method" refers to a function that allows users to make reservations by linking with the systems of stores and facilities that provide services related to the proposed style.
[0633] "Payment method" refers to a function that allows for secure and efficient payment related to reservations through a third-party payment service.
[0634] To implement this invention, a system is configured in which a user, a terminal, and a server work together. First, the user launches a dedicated application on their terminal. The terminal accepts a facial photograph taken by the user using the camera as input. The terminal can also input preference data, such as preferred fashion styles and makeup preferences.
[0635] The device encrypts and securely transmits collected facial photos and preference data to the server. The server utilizes a generative AI model to analyze the received data, performing a detailed analysis of features such as facial shape and skin tone. Based on this analysis, it generates suggestions for the user's optimal fashion style and beauty routine.
[0636] The generative AI model performs more effective analysis using prompts. A concrete example of a prompt would be, "Please suggest a suitable fashion style based on the user's facial features and preferences." This allows the AI model to suggest styles that maximize the user's unique characteristics.
[0637] The server sends the suggested style back to the terminal, allowing the user to review it. Based on the suggested style, the user can book a reservation at a relevant beauty salon or fashion store. This reservation is made through the terminal in conjunction with the system of the relevant facility.
[0638] Furthermore, payment processing can also be done via the terminal. The server completes payments securely and smoothly through external payment services. This allows users to efficiently select the style that best suits them and make reservations and payments easily through a series of processes.
[0639] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0640] Step 1:
[0641] The user launches a dedicated application on their device and takes a photo of their face. The photo is saved as image data on the device. The user also enters information about their fashion and makeup preferences through an input form on the device. This creates a single dataset containing both the photo and the preference data.
[0642] Step 2:
[0643] The device encrypts the collected facial images and preference data and sends them to the server using a secure communication protocol. This input data includes image files and text information. The dataset received by the server is used as foundational information for facial feature analysis and preference matching.
[0644] Step 3:
[0645] The server passes the received image data to a generating AI model, which analyzes facial features such as shape and skin tone. The AI model uses prompts to generate style suggestions tailored to the user. For example, it might use a prompt such as, "Please suggest the best fashion style based on the user's facial features."
[0646] Step 4:
[0647] The server compiles the generated style suggestions and sends them back to the user. This output data is formatted as fashion style and beauty suggestions and displayed on the terminal. The user reviews the displayed style suggestions and makes selections that suit them.
[0648] Step 5:
[0649] The user reviews the suggested style and makes a reservation at the suggested beauty salon or fashion store via their device. This reservation information is sent from the device to the server, which registers the reservation information in the system of the affiliated facility.
[0650] Step 6:
[0651] Users make reservations and payments simultaneously using their devices. The server receives the payment information transmitted from the device and securely processes the payment via an external payment service. This ensures that the entire process is completed smoothly.
[0652] (Application Example 1)
[0653] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0654] Traditional technologies lack sufficient support for users to discover their own style, select appropriate products in physical stores, and make efficient purchases. This problem not only detracts from the user experience but also reduces purchasing intent. Therefore, there is a need for a system that allows users to easily find products that suit them in physical stores and learn about optimal coordination.
[0655] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0656] In this invention, the server includes an input means for acquiring user image data, an analysis means for analyzing the image data and recognizing facial features, a suggestion means for proposing an appropriate style based on the recognized features, a product identification means for acquiring product information and identifying products that match the user's features, and a presentation means for presenting coordination information related to the products identified by the product identification means. This makes it possible for users to improve their in-store shopping experience and efficiently select products that suit them.
[0657] A "user" refers to a person who operates the system and provides image data.
[0658] "Image data" refers to information that visually records a user's face or a product, and is the data that is subject to analysis.
[0659] "Input means" refers to devices or technologies used by a user to supply image data to a system.
[0660] "Analysis means" refers to devices or technologies that process acquired image data and analyze the user's facial features and preferences.
[0661] "Suggested means" refers to devices or technologies that present appropriate fashion styles and products to the user based on the analysis results.
[0662] A "reservation method" refers to a device or technology for reserving related services based on the proposed style.
[0663] A "payment method" refers to the equipment or technology used to process and pay for reserved services or goods.
[0664] "Product identification means" refers to devices or technologies used to identify products that match the characteristics of a user.
[0665] "Presentation means" refers to a device or technology that displays or notifies a user of information related to a specified product.
[0666] To realize this invention, a system is needed in which three entities—the user, the terminal, and the server—work together. The system begins with the user launching an application installed on the terminal.
[0667] A smartphone is recommended as the device, and it is equipped with a camera that can acquire image data in real time. Using this camera, users can take pictures of their faces or products they are interested in. The device securely transmits the captured image data to the server.
[0668] The server is equipped with advanced image analysis algorithms, using open-source OpenCV and TensorFlow to perform highly accurate facial feature and preference analysis. The server also uses a generative AI model to suggest fashion styles based on the user's face and preferences. The generative AI model has the ability to output a variety of styling options based on the input of a prompt. For example, with the prompt "Consider the user's face photo and fashion preferences, suggest an outfit that goes well with a denim jacket," the model generates the most suitable options for the user.
[0669] The analysis results and generated styling suggestions are fed back to the terminal and visually displayed to the user. This display includes images of suggested products and recommended styles. Based on this, the user can find the relevant products in physical stores and obtain additional information. The terminal can also be used to make reservations for suggested services and, if necessary, to complete payments. Payments are completed securely using a third-party payment service.
[0670] This entire process allows users to smoothly select the fashion items that best suit them and is supported throughout the purchase process.
[0671] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0672] Step 1:
[0673] The user launches the application on their smartphone and uses the camera to capture images of their face or the product they are considering purchasing. The input is the image data obtained through the camera. The output is the image data itself. This process is completed by the user operating the camera according to the instructions on their device.
[0674] Step 2:
[0675] The terminal transmits the acquired image data to the server using a secure communication protocol. The input is the image data obtained by the terminal, and the output is the image data sent to the server. Encrypted communication is used for secure transmission.
[0676] Step 3:
[0677] The server analyzes the received image data using OpenCV and TensorFlow, performing data processing to extract features of the user's face and attributes of the product. The input is the image data sent to the server, and the output is the analyzed feature information. This operation uses a feature extraction algorithm.
[0678] Step 4:
[0679] The server uses the analyzed feature information to provide prompts to a generating AI model, which then generates styling suggestions best suited to the user. The input consists of the analyzed feature information and prompts, while the output is the generated styling suggestions. The prompt used is: "Considering the user's facial photo and fashion preferences, please suggest an outfit that goes well with a denim jacket."
[0680] Step 5:
[0681] The server sends the generated styling suggestions back to the terminal, which then visually displays them to the user. The input is the styling suggestions received from the server, and the output is the displayed suggestions. The terminal uses a UI-based display to present the suggestions clearly to the user.
[0682] Step 6:
[0683] The user reviews the presented styling suggestions and selects the products and services they like. Based on this selection, they decide whether or not to book the suggested services or products. The input is the displayed styling suggestions, and the output is the user's selections. This process is completed when the user makes a decision using the terminal's interface.
[0684] Step 7:
[0685] When a user makes a reservation, the terminal sends the reservation information to the server, which then registers the reservation information in conjunction with the system of the relevant service provider. The input is reservation information based on the user's selection, and the output is a notification to the service provider that registration is complete. The server performs the registration process with the facility's system via an API.
[0686] Step 8:
[0687] After completing the reservation, the user uses the same device to quickly and securely complete the associated payment. The device interacts with a third-party payment service via a server to complete the payment process. The input is the payment information related to the reservation, and the output is a notification that the payment has been completed. This operation is performed using the payment method selected by the user.
[0688] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0689] The embodiment of this invention is a system centered on three components: a user, a terminal, and a server, with an emotion recognition engine incorporated as part of it. This system proposes appropriate beauty and fashion styles to the user and also enables the user to handle everything from service reservations to payment in a consistent manner.
[0690] The user first uses an application on their device to take a photo of their face and input their fashion and makeup preferences. In addition, there is a function to capture the user's facial expressions in real time, and emotional data obtained from this is collected. The device encrypts this information and sends it to the server.
[0691] The server analyzes the received facial photograph and extracts the user's characteristics. It also combines and analyzes the user's preference information and emotional data to suggest a style that best suits the user's current emotional state. For example, if the user has a relaxed expression, it can suggest a relaxed style.
[0692] The suggested styling is sent back to the terminal and presented to the user on the terminal. Here, the user can review the suggested style and make reservations at hair salons and related fashion stores. The reservation process is carried out by the server coordinating with the store's reservation system and registering the necessary information.
[0693] Furthermore, users can pay for their bookings instantly using their chosen payment method. This is done through a third-party payment service via the server, ensuring smooth processing.
[0694] For example, suppose a user wants to try something a little different from their usual style. Based on their facial photo and emotional data, the server suggests slightly brighter coloring and develops a fashion style that is consistent with the underlying intentions of their facial expressions. Based on this suggestion, the user selects a beauty service, and the booking and payment are completed instantly. This ensures that the user receives the service in a way that is best suited to them.
[0695] The following describes the processing flow.
[0696] Step 1:
[0697] The user launches an application on their device, takes a photo of their face, and records their current facial expression through the camera. At the same time, the user inputs their preferences regarding fashion and makeup.
[0698] Step 2:
[0699] The device encrypts the user's facial image, preference information, and real-time captured emotional data and sends it to the server.
[0700] Step 3:
[0701] The server first analyzes facial image data to extract features such as the user's face shape and skin tone. It also uses an emotion recognition engine to identify the user's emotions from their facial expression data. This information is then combined with preference data for a comprehensive analysis.
[0702] Step 4:
[0703] Based on the analysis results, the server generates styling that best suits not only the user's personality but also their emotional state. For example, if a depressed expression is detected, a colorful style that will cheer the user up will be suggested.
[0704] Step 5:
[0705] The server sends the suggested content back to the terminal, which then displays it to the user. The user reviews the suggested style, and if they like it, selects from the related service providers and proceeds with the reservation.
[0706] Step 6:
[0707] The terminal sends the reservation information selected by the user to the server, and the server works with the reservation system to complete the reservation for the desired beauty service.
[0708] Step 7:
[0709] The terminal or server processes the payment using the payment method selected by the user. The payment process is carried out through integration with a third-party payment service.
[0710] Step 8:
[0711] Users can confirm on their device that their reservation and payment have been successfully completed. User feedback is also recorded and used to improve future suggestions.
[0712] (Example 2)
[0713] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0714] In response to the demands of modern, personalized beauty and fashion, there is a lack of systems that can accurately recognize users' emotions and preferences in real time and integrate everything from quick and appropriate style suggestions to booking and payment based on that information. Conventional systems lack the ability to suggest styles that take into account the user's real-time emotions, making it difficult to optimize the user experience.
[0715] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0716] In this invention, the server includes an input means for acquiring user video information, an analysis means for analyzing the video information and extracting facial features, and a suggestion means for proposing an appropriate style based on the analyzed facial features and emotional information. This makes it possible to propose a personalized style that reflects the user's emotional state.
[0717] "Input means" refers to a device or method for acquiring video information from a user.
[0718] "Analysis means" refers to a device or method for processing acquired video information and extracting facial features and emotional information.
[0719] "Suggestion means" refers to a device or method for suggesting personalized styles to a user based on analyzed facial features and emotional information.
[0720] "Reservation means" refers to a device or method for reserving necessary operations based on the proposed format.
[0721] "Payment processing means" refers to a device or method for properly settling the costs associated with a reservation.
[0722] This system aims to automate beauty and fashion style suggestions by incorporating an emotion recognition engine, centered around three parties: the user, the device, and the server. Users take a photo of their face and input their fashion preferences through a dedicated app installed on a device such as a smartphone or tablet. The device encrypts this image information and preference data, along with real-time emotion data, using secure protocols such as SSL / TLS, and transmits it to the server.
[0723] The server decodes the received data and uses image analysis software to extract facial features. It also utilizes an emotion recognition engine to evaluate the user's emotional state and uses a generative AI model to suggest the optimal fashion style based on the collected data. For example, a relaxed mood might be suggested with a casual style, while a more adventurous mood might be suggested with a style featuring vibrant colors.
[0724] The suggested styles are sent back to the device and visually presented to the user. The user can review the suggestions and make reservations at hair salons or fashion establishments via the app. The server links the reservation information with the establishment's information system and registers it. It also links with payment service providers to process payments, allowing users to pay immediately after making a reservation. This enables users to smoothly enjoy the experience based on their chosen style.
[0725] For example, if a user enters "I want a bright and adventurous style," the server analyzes their facial photo and the frequency of their smiles, and suggests clothing in bright colors. Based on this suggestion, the user can select a suitable facility, make a reservation immediately, and complete the payment.
[0726] Examples of prompts for a generative AI model include the following:
[0727] "Please suggest beauty and fashion styles based on user emotional data and fashion preferences."
[0728] This embodiment enables personalized style selection based on emotions, improving the user experience.
[0729] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0730] Step 1:
[0731] The user launches a dedicated application on their device and takes a photo of their face. This photo becomes the input data. Furthermore, the user inputs their preferences regarding fashion and makeup within the app. A real-time emotion capture function is also activated, collecting emotional data. All of this data is compiled into a single dataset.
[0732] Step 2:
[0733] The device encrypts the captured facial image, preferences, and emotional data. This process uses encryption protocols such as SSL / TLS to ensure data security. The encrypted data is then converted into data packets and sent to the server.
[0734] Step 3:
[0735] The server decodes the received data packets and analyzes the facial image to extract the user's facial features. An image analysis algorithm is used for this process. The input data is a facial image, and the output is data on the user's facial features. In addition, an emotion recognition engine is used to process data to determine the user's emotional state. The output in this process is data indicating the specific emotion and its intensity.
[0736] Step 4:
[0737] The server uses a generative AI model to suggest the optimal fashion style based on analyzed facial feature data and emotion data. This process takes into account the user's preferences and emotional state. For example, a calm mood will result in casual suggestions, while an adventurous mood will generate vibrant suggestions. The AI model generates specific suggestions based on prompt messages. The system then compiles these suggestions into a dataset and sends it to the terminal.
[0738] Step 5:
[0739] The terminal displays fashion style suggestions to the user. The user can review the suggested styles and decide to book beauty or fashion services based on them. Based on the user's selection, booking details are generated, and the process proceeds to the next step.
[0740] Step 6:
[0741] The server links information with the service provider's reservation system based on the user's selection. This information processing ensures the reservation is correctly registered. Furthermore, by linking with a third-party payment processing service via the server, a system is activated that allows the user to pay immediately upon booking. In this step, the reservation and payment information are the final outputs. As a result, the user is ready to receive the service.
[0742] (Application Example 2)
[0743] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] In modern retail, there is a demand for personalized services tailored to each customer's preferences and mood. However, achieving this requires significant human resources and is difficult to do instantly. Furthermore, there is a challenge in the lack of adequately developed systems that enable customers to make optimal choices immediately and increase their satisfaction.
[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0746] In this invention, the server includes an input / output device for acquiring user image data, an information processing device for analyzing the image data and recognizing facial features, and a style suggestion device for suggesting appropriate styles based on the recognized features and user emotion data. This makes it possible to suggest optimal fashion styles based on the customer's emotions and preferences.
[0747] An "input / output device for acquiring user image data" is a device that takes image information from a user as input and provides output to the user as needed.
[0748] An "information processing device that analyzes image data to recognize facial features" is a device that receives image data, analyzes that data, and identifies specific facial features.
[0749] A "style suggestion device" is a device that has the function of suggesting fashion styles and other styles suitable for the user based on recognized facial features and user emotion data.
[0750] A "service support device" is a device that, based on a proposed style, assists users with selecting and trying on products at the service location, as well as providing other customer support.
[0751] A "reservation and payment device" is a device that accepts reservations for services selected by users and performs the associated payment processing.
[0752] An "information processing device" is a device that has the basic functions to receive and analyze various types of data and perform processing based on the results.
[0753] An "electronic commerce device" is a device that has the necessary functions to conduct transactions online and is used to carry out processes related to electronic commerce, such as payment and purchase procedures.
[0754] This system works in conjunction with the user, terminal, and server to provide the user with the optimal fashion style. The user acquires image data using input / output devices installed in the terminal. For example, a smartphone or smart glasses function as this terminal. The terminal transmits the acquired image data to an information processing unit. This information processing unit is equipped with an analysis system for recognizing facial features, and uses software such as OpenCV or AWS Rekognition to analyze facial expressions and emotion data.
[0755] Based on analyzed feature and emotion data, the server suggests a suitable style to the user via a style suggestion device. This style suggestion uses a generative AI model to determine the fashion style best suited to the user's current emotions and preferences. For example, a user judged to be in a relaxed state might be suggested a casual and comfortable style.
[0756] Based on the proposed style, users can select and try on products through physical stores or online service support devices. A reservation and payment device operates through the terminal, allowing users to reserve and pay for their selected products and services. Smooth online transactions are ensured by integration with third-party payment systems such as Stripe.
[0757] An example of a prompt message is as follows: "Here is the customer's facial expression data. Please use the emotion recognition engine to suggest a fashion style that best suits the customer's current mood."
[0758] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0759] Step 1:
[0760] The user acquires their own image data using a device. Specifically, they take a photo of their face using the camera function of a smartphone or smart glasses. In this case, the input is the user's facial image data, and the output is the same image file.
[0761] Step 2:
[0762] The terminal transmits the acquired image data to the information processing device. Here, the terminal encrypts the image data and securely transmits it to the server. In this process, the input is facial image data, and the output is encrypted image data.
[0763] Step 3:
[0764] The server uses an information processing device to analyze the received image data. Specifically, it uses OpenCV and AWS Rekognition to extract facial features from the images and analyze emotions. The input to this process is encrypted image data, and the output is the analyzed facial features and emotion data.
[0765] Step 4:
[0766] The server uses a generative AI model based on the analyzed data to suggest a style suitable for the user. Here, the input is facial features and emotion data, and the output is a recommended fashion style. As a concrete example, a user with a relaxed expression might be suggested a casual style.
[0767] Step 5:
[0768] The terminal presents the user with style suggestions received from the server. After reviewing the suggested styles, the user decides whether to try them on in a store or make a purchase. The input to this process is the suggested fashion styles, and the output is the style information selected by the user.
[0769] Step 6:
[0770] The terminal operates the reservation and payment device based on the user's selection to make reservations and payments for services and products. Reservation information is transmitted from the terminal to the store's information processing device. The input to this process is the selected style information, and the output is reservation confirmation and payment completion information.
[0771] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0772] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0773] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0774] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0775] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0776] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0777] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0778] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0779] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0780] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0781] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0782] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0783] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0784] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0785] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0786] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0787] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0788] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0789] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0790] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0791] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0792] The following is further disclosed regarding the embodiments described above.
[0793] (Claim 1)
[0794] An input means for acquiring user image data,
[0795] An analysis method that analyzes image data to recognize facial features,
[0796] A suggestion method that proposes an appropriate style based on recognized characteristics,
[0797] A booking method for reserving the necessary services based on the proposed style,
[0798] Payment methods for making payments related to reservations,
[0799] A system that includes this.
[0800] (Claim 2)
[0801] The system according to claim 1, characterized in that the analysis means further analyzes user preference information, and the suggestion means proposes a style that takes the preference information into consideration.
[0802] (Claim 3)
[0803] The system according to claim 1, characterized in that the reservation method registers reservation information in conjunction with the system of the service provider, and the payment method performs payment in conjunction with a third-party payment service.
[0804] "Example 1"
[0805] (Claim 1)
[0806] A terminal for acquiring user image information,
[0807] A server using a generative AI model that analyzes image information in detail to recognize facial features,
[0808] A suggestion method that proposes an appropriate style based on recognized features and user preference data,
[0809] A reservation method for booking services such as beauty salons and fashion stores based on a suggested style,
[0810] Payment methods that involve making reservations through third-party payment services,
[0811] A system that includes this.
[0812] (Claim 2)
[0813] The system according to claim 1, characterized in that the generating AI model further analyzes user preference information and provides a suggestion means for generating a style that takes preference information into account using prompt sentences.
[0814] (Claim 3)
[0815] The system according to claim 1, characterized in that the reservation method registers reservation information in cooperation with the service provider's system, and the payment method performs secure payments in cooperation with an external payment institution.
[0816] "Application Example 1"
[0817] (Claim 1)
[0818] An input means for acquiring user image data,
[0819] An analysis method that analyzes image data to recognize facial features,
[0820] A suggestion method that proposes an appropriate style based on recognized characteristics,
[0821] A booking method for reserving the necessary services based on the proposed style,
[0822] Payment methods for making payments related to reservations,
[0823] A product identification method that obtains product information and identifies products that match the user's characteristics,
[0824] A presentation means that presents coordination information related to a product identified by a product identification means,
[0825] A system that includes this.
[0826] (Claim 2)
[0827] The system according to claim 1, characterized in that the analysis means further analyzes user preference information, and the suggestion means proposes a style that takes the preference information into consideration.
[0828] (Claim 3)
[0829] The system according to claim 1, characterized in that the reservation method registers reservation information in conjunction with the system of the service provider, and the payment method performs payment in conjunction with a third-party payment service.
[0830] "Example 2 of combining an emotion engine"
[0831] (Claim 1)
[0832] An input means for acquiring user video information,
[0833] An analysis method that analyzes video information to extract facial features,
[0834] A proposal means for suggesting an appropriate style based on analyzed facial features and emotional information,
[0835] A reservation method for booking the necessary operations based on the proposed format,
[0836] A payment processing method for processing payments related to reservations,
[0837] A system that includes this.
[0838] (Claim 2)
[0839] The system according to claim 1, characterized in that the analysis means further analyzes the user's preference information, and the proposal means proposes a style that takes the preference information into consideration.
[0840] (Claim 3)
[0841] The system according to claim 1, characterized in that the reservation means registers reservation information in cooperation with the system of the facility providing the service, and the payment processing means processes payments in cooperation with a third-party payment processing service.
[0842] "Application example 2 when combining with an emotional engine"
[0843] (Claim 1)
[0844] An input / output device for acquiring user image data,
[0845] An information processing device that analyzes image data to recognize facial features,
[0846] A style suggestion device that proposes an appropriate style based on recognized features and user emotion data,
[0847] A service support device that assists in selecting and trying on items at the service location based on the proposed style,
[0848] A reservation and payment device for making reservations and payments for services,
[0849] A system that includes this.
[0850] (Claim 2)
[0851] The system according to claim 1, characterized in that it analyzes user preference information and emotional data, and the style suggestion device suggests a style that takes that data into consideration.
[0852] (Claim 3)
[0853] The system according to claim 1, characterized in that the reservation and payment device registers information in cooperation with an information processing device at the service provision location and performs payment in cooperation with a third-party e-commerce device. [Explanation of Symbols]
[0854] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input means for acquiring user image data, An analysis method that analyzes image data to recognize facial features, A suggestion method that proposes an appropriate style based on recognized characteristics, A booking method for reserving the necessary services based on the proposed style, Payment methods for making payments related to reservations, A system that includes this.
2. The system according to claim 1, characterized in that the analysis means further analyzes user preference information, and the suggestion means proposes a style that takes the preference information into consideration.
3. The system according to claim 1, characterized in that the reservation method registers reservation information in conjunction with the system of the service provider, and the payment method performs payment in conjunction with a third-party payment service.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A