system
A smartphone-based AI system automates menu creation, product description generation, and translation, addressing inefficiencies and language barriers to enhance restaurant operations and customer satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-09-19
- Publication Date
- 2026-04-21
Smart Images

Figure 0007849429000001 
Figure 0007849429000002 
Figure 0007849429000003
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the operation of restaurants, a great deal of time and labor are required for menu creation, product description generation, translation work, etc. Also, on the customer side, there are stressful experiences such as waiting in line in front of the host or cashier, and doubts when considering the menu. In particular, services for foreign tourists are not sufficiently provided due to language barriers.
Means for Solving the Problems
[0005] This invention provides a means for creating menus using a smartphone, an AI means for generating product descriptions from smartphone photos, and a means for AI to generate translations. This reduces the operational burden on restaurants and improves the customer experience. Furthermore, the AI explains questions that arise when considering menus and provides recommendations based on preferences, thereby offering a more personalized service. In addition, the multilingual translation AI enhances services for foreign tourists. [Brief explanation of the drawing]
[0006] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 1 of Example 1. [Figure 12]This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2. [Figure 15] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 3 of Example 3. [Figure 16] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3. [Figure 17] This is a sequence diagram showing the processing flow of the data processing system in Example 1 of the Form 1 when an emotion engine is combined. [Figure 18] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of the Form 2 when an emotion engine is combined. [Figure 20] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] This is a sequence diagram showing the processing flow of the data processing system in Example 3 of the Form 3 when an emotion engine is combined. [Figure 22] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. [Modes for carrying out the invention]
[0007] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0008] First, the terms used in the following description will be explained.
[0009] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)), etc.
[0010] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0011] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0012] In the following embodiments, a communication I / F (Interface) with a reference numeral is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0014] [First Embodiment]
[0015] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0016] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0017] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0018] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0019] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0021] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0022] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0023] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0024] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0025] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0026] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0027] "Example of form 1"
[0028] In one embodiment of the present invention, a restaurant operator creates a menu using a smartphone. Specifically, the operator uploads photos of products taken with the smartphone's camera, and the AI then processes those photos.
[0029] The system automatically recognizes the characteristics and features of a product. Based on this recognition, it generates a product description. This product description is then displayed to customers on the restaurant's website or application.
[0030] "Example of form 2"
[0031] Furthermore, in this embodiment of the present invention, the AI generates the translation work. Specifically, the generated product description is translated into multiple languages. This makes the product description easier for foreign tourists to understand. For example, it is possible to translate into major tourist languages such as English, Chinese, and Korean.
[0032] "Example of form 3"
[0033] Furthermore, in this embodiment of the present invention, the customer experience is also improved. Specifically, when a customer is considering a menu, the AI explains their questions and recommends products according to their preferences. For example, if a customer inputs information such as "I like spicy food," the AI will recommend spicy dishes. Also, if a customer asks a question such as "What are the ingredients in this dish?", the AI will provide an answer to that question.
[0034] The following describes the processing flow for each example of the form.
[0035] "Example of form 1"
[0036] Step 1: The restaurant operator takes photos of the products using their smartphone camera.
[0037] Step 2: Upload the photos you've taken to the system.
[0038] Step 3: The AI within the system automatically recognizes the product's characteristics and features from the photograph.
[0039] Step 4: The AI generates a product description based on the recognition results.
[0040] Step 5: The generated product description is displayed to customers on the restaurant's website or application.
[0041] "Example of form 2"
[0042] Step 1: Obtain the product description generated by the AI.
[0043] Step 2: The AI translates the product description into multiple languages.
[0044] Step 3: The translated product description is displayed to foreign tourists on the restaurant's website or application.
[0045] "Example of form 3"
[0046] Step 1: When customers are considering the menu, they input their questions into the AI.
[0047] Step 2: The AI generates the answer to that question.
[0048] Step 3: The AI recommends products based on the customer's preferences.
[0049] Step 4: The AI-generated answers and recommendations are displayed to the customer.
[0050] (Example 1)
[0051] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0052] When restaurant operators create menus, the process of taking photos of products and then recognizing their characteristics and features to generate product descriptions is time-consuming. Furthermore, there is a lack of multilingual support for foreign tourists and other means to enhance the customer experience. Therefore, there is a need for increased efficiency in menu creation and improved customer satisfaction.
[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0054] In this invention, the server includes means for creating menus using a smartphone, means for artificial intelligence to generate product descriptions from photos taken with a smartphone, means for artificial intelligence to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for uploading product photos taken with a smartphone camera and for artificial intelligence to automatically recognize the characteristics and features of the products from those photos, and means for generating product descriptions based on the recognition results. This makes it possible to improve the efficiency of menu creation and enhance customer satisfaction.
[0055] A "smartphone" is a multi-functional mobile device that, in addition to the functions of a mobile phone, is capable of internet connectivity and the use of applications.
[0056] "Menu creation methods" refer to the methods and tools that restaurant operators use to create menus, and especially those that utilize smartphones.
[0057] "Artificial intelligence methods" refer to technologies that use machine learning and data analysis to automatically perform specific tasks.
[0058] "Means of generating translation work using artificial intelligence" refers to methods and technologies that use artificial intelligence to translate text into multiple languages.
[0059] "Methods for improving the customer experience" refer to methods and tools for improving the convenience and satisfaction customers experience when using a service.
[0060] "Methods for AI to explain questions when considering menus" refers to methods and technologies in which artificial intelligence automatically provides answers to questions that customers may have when considering menus.
[0061] "Methods of recommending based on preferences" refer to methods and technologies that recommend appropriate products and services based on a customer's preferences and past choices.
[0062] "Means for automatically recognizing the characteristics and features of a product" refers to methods and technologies that use artificial intelligence to automatically extract the characteristics and features of a product from photographs or data.
[0063] "Means for generating product descriptions" refers to methods and technologies for creating product descriptions in natural language based on recognized product characteristics and features.
[0064] This invention is a system that allows restaurant operators to create menus using their smartphones. Specifically, users upload photos of products taken with their smartphone cameras, and artificial intelligence (AI) automatically recognizes the characteristics and features of the products from the photos. Based on the recognition results, the system generates product descriptions. These product descriptions are then displayed to customers on the restaurant's website or application.
[0065] Hardware and software to be used
[0066] Smartphone: A device used for taking and uploading photos.
[0067] Server: Performs data processing and runs AI models.
[0068] Artificial intelligence models: Image recognition models and natural language generation models based on TENSORFLOW® and PyTorch (e.g., GPT-3®, BERT).
[0069] Data processing and data calculation
[0070] 1. The user takes a photo of the product with their smartphone.
[0071] The user launches their smartphone's camera app and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[0072] 2. The user uploads photos to the system.
[0073] The user opens a dedicated application, selects the photos they have taken, and presses the upload button. The photos are then sent to the server via the internet.
[0074] 3. The server receives the photos and inputs them into the AI model.
[0075] The server receives the uploaded photos and inputs them into an AI model for image processing. The AI model used here is based on TensorFlow or PyTorch.
[0076] 4. The server uses an AI model to recognize the characteristics and features of the product.
[0077] The server uses an AI model to recognize the characteristics and features of a product from a photograph. For example, it extracts the type of dessert, main ingredients, and visual features.
[0078] 5. The server generates a product description based on the recognition results.
[0079] The server generates product descriptions based on the AI's recognition results. These product descriptions are created using natural language generation technology. The software used includes GPT-3 and BERT.
[0080] 6. The server displays the product description on the website or application.
[0081] The server displays the generated product descriptions on the restaurant's website or application. Customers can then view them.
[0082] Specific example
[0083] A user takes a photo of a new dessert, "Chocolate Cake," and uploads it to the system. The server receives the photo and uses an AI model to recognize features such as "Chocolate Cake," "Cream Topping," and "Berry Decoration." The server uses GPT-3 to generate a product description such as, "This chocolate cake features rich chocolate and creamy toppings. The berry decoration makes it visually appealing," and displays it on the website.
[0084] Example of a prompt
[0085] Please upload a photo of your new dessert. AI will recognize the product's characteristics and features from the photo and generate a product description.
[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0087] Step 1:
[0088] The user takes a photo of the product with their smartphone.
[0089] Input: Product (e.g., Chocolate cake)
[0090] Output: Photos of the photographed product
[0091] Specific action: The user launches the camera app on their smartphone and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[0092] Step 2:
[0093] The user uploads a photo to the system.
[0094] Input: Photos of the product
[0095] Output: Photo data sent to the server
[0096] Specific operation: The user opens a dedicated application, selects the photo they have taken, and presses the upload button. The photo is sent to the server via the internet.
[0097] Step 3:
[0098] The server receives the photos and inputs them into the AI model.
[0099] Input: Photo data sent to the server
[0100] Output: Photo data input to the AI model
[0101] Specific operation: The server receives an HTTP request, temporarily stores the photo data, and inputs it into a TensorFlow or PyTorch model.
[0102] Step 4:
[0103] The server uses an AI model to recognize the characteristics and features of the product.
[0104] Input: Photo data entered into the AI model
[0105] Output: Characteristics and features of the recognized product (e.g., chocolate cake, cream topping, berry decoration)
[0106] Specific operation: The server runs an AI model to extract product characteristics and features from a photograph. For example, it recognizes the type of dessert, main ingredients, and visual features.
[0107] Step 5:
[0108] The server generates a product description based on the recognition results.
[0109] Input: Characteristics and features of the recognized product
[0110] Output: Generated product description (Example: "This chocolate cake features rich chocolate and creamy toppings. The berry decorations make it visually appealing.")
[0111] Specific operation: The server uses GPT-3 or BERT to generate product descriptions based on the recognition results as input.
[0112] Step 6:
[0113] The server displays product descriptions on websites and applications.
[0114] Input: Generated product description
[0115] Output: Product description displayed on the website or application
[0116] Specific operation: The server saves the product description in the website's database and sends the data to the front-end for display. Customers can then view this data.
[0117] (Application Example 1)
[0118] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0119] Traditional restaurant menu creation was often done manually, which was time-consuming and labor-intensive, and made it difficult to adequately convey the appeal of the products. Furthermore, the lack of multilingual support for foreign tourists and insufficient recommendation features tailored to customer preferences were also problems. In addition, there were limited means to improve the in-store customer experience, such as having to wait for waitstaff or in line at the register, highlighting the need for increased customer satisfaction.
[0120] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0121] In this invention, the server includes means for creating menus using a smartphone, means for generating product descriptions from photos taken with the smartphone, means for the AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for the AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on the recognized features using a natural language generation model. This enables efficient menu creation and automatic generation of attractive product descriptions, as well as multilingual support, improved customer experience, and personalized recommendations.
[0122] "A method for creating menus using a smartphone" refers to a method for efficiently creating restaurant menus using the camera and applications of a smartphone.
[0123] "AI methods for generating product descriptions from smartphone photos" refers to artificial intelligence methods that analyze product photos taken with a smartphone, recognize their characteristics and features, and automatically generate product descriptions.
[0124] "Methods for AI to generate translation work" refers to methods of using artificial intelligence to translate product descriptions and menu contents into multiple languages.
[0125] "Methods to improve the customer experience by eliminating the need to wait for waitstaff or in line at the register" refers to methods that allow customers to receive service smoothly without having to wait for waitstaff or in line at the register.
[0126] "An AI-powered solution for questions during menu consideration" refers to a method in which artificial intelligence automatically provides answers to questions that customers may have when considering menu options.
[0127] "Methods of recommending based on preferences" refer to methods of recommending appropriate products or menus based on a customer's past choices and preferences.
[0128] "Methods for recognizing the characteristics and features of a product using an image recognition model" refers to methods for automatically recognizing the characteristics and features of a product from a photograph using image recognition technology.
[0129] "Means for generating product descriptions based on features recognized using a natural language generation model" refers to means for automatically generating attractive product descriptions using natural language generation technology based on recognized product characteristics and features.
[0130] A system for carrying out this invention includes means for creating menus using a smartphone, means for generating product descriptions from photos on a smartphone, means for AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the cash register, means for AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on recognized features using a natural language generation model.
[0131] System program
[0132] The server implements the system using the following hardware and software.
[0133] Hardware:
[0134] Smartphone (with camera)
[0135] Server (for hosting AI models)
[0136] software:
[0137] TensorFlow: A library for image recognition
[0138] OpenAI(registered trademark) GPT-3: API for natural language generation
[0139] Explanation of the process
[0140] Image recognition:
[0141] Users take photos of their food with their smartphone cameras and upload them to a server via an application. The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of the food from the image. For example, features such as "chicken curry, spicy, tomato-based" might be recognized.
[0142] Natural language generation:
[0143] The server converts the recognized features into a string and sends it to OpenAI GPT-3 as a prompt. An example of a prompt is, "Generate an appealing product description for a dish with the following features: Chicken curry, spicy, tomato-based." GPT-3 generates an appealing product description based on this prompt. For example, it might generate a product description such as, "This spicy chicken curry features juicy chicken simmered in a tomato-based sauce. The aromatic spices will whet your appetite."
[0144] Translation work:
[0145] The generated product descriptions are translated into multiple languages as needed. The server uses AI to perform translations in multiple languages, such as English and Chinese.
[0146] Improving the customer experience:
[0147] Customers can use their smartphones to select menu items and place orders in-store. This reduces waiting times at the counter and cashier, enabling smoother service.
[0148] Explanation of the question and recommendations:
[0149] The AI automatically provides answers to questions that customers may have when considering menu options. It also recommends appropriate products and menu items based on the customer's past choices and preferences.
[0150] In this way, menu creation becomes more efficient, attractive product descriptions are automatically generated, and multilingual support, improved customer experience, and personalized recommendations become possible.
[0151] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0152] Step 1:
[0153] The user takes a photo of the food with their smartphone camera and uploads it to the server through the application. The input is the photo of the food, and the output is the image data sent to the server.
[0154] Step 2:
[0155] The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of dishes from uploaded image data. The input is image data, and the output is a list of the recognized characteristics and features. Specifically, the image data is preprocessed and then input into the model to obtain prediction results.
[0156] Step 3:
[0157] The server converts the recognized traits and features into a string and sends it to OpenAI GPT-3 as a prompt. The input is a list of traits and features, and the output is a prompt statement. Specifically, it formats the traits and features and generates a prompt statement.
[0158] Step 4:
[0159] The server receives a response from GPT-3 and generates an attractive product description. The input is a prompt, and the output is the generated product description. Specifically, it calls the GPT-3 API, parses the response, and retrieves the product description.
[0160] Step 5:
[0161] The server translates the generated product description into multiple languages as needed. The input is the product description, and the output is the translated product description. Specifically, it calls a translation API to translate into multiple languages.
[0162] Step 6:
[0163] Users use their smartphones to select menu items and place orders in the store. The input is a generated product description, and the output is the user's order information. Specifically, the application displays product descriptions, and the user selects their order.
[0164] Step 7:
[0165] The server automatically provides answers to questions that arise when users consider the menu, using AI. The input is the user's question, and the output is the AI's answer. Specifically, it analyzes the question and generates an appropriate answer.
[0166] Step 8:
[0167] The server recommends appropriate products and menu items based on the user's past choices and preferences. The input is the user's past selection data, and the output is the recommended products and menu items. Specifically, it analyzes past selection data and uses a recommendation algorithm to suggest products.
[0168] (Example 2)
[0169] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0170] Traditional systems often required manual translation of product descriptions into multiple languages, resulting in time-consuming and labor-intensive processes. Furthermore, services for foreign tourists were inadequate, making it difficult for them to understand product descriptions. Additionally, insufficient efforts were made to improve the customer experience and address questions during menu selection, potentially leading to decreased customer satisfaction.
[0171] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0172] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for the artificial intelligence to generate translation tasks, means for improving the customer experience, means for the artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into the generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results. As a result, multilingual translation of product descriptions is automated, improving services for foreign tourists, enhancing the customer experience, and resolving questions when considering menus.
[0173] A "smart device" is a portable electronic device, such as a smartphone or tablet, that possesses advanced computing power and communication capabilities.
[0174] "Menu creation methods" refer to functions and applications that use smart devices to create menus for restaurants and retail stores.
[0175] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[0176] "A means of generating translation work using artificial intelligence" refers to a function that uses artificial intelligence to translate text data into other languages.
[0177] "Means of improving the customer experience" refer to functions and methods that enhance the convenience and satisfaction customers experience when using products or services.
[0178] "A means for artificial intelligence to explain questions when considering menus" refers to a function in which artificial intelligence automatically provides answers to questions and concerns that arise when customers are considering menus.
[0179] "Means of recommending based on preferences" refers to a function that recommends appropriate products and services based on the customer's past choices and preferences.
[0180] "Means for translating product descriptions into multiple languages" refers to a function for translating product descriptions into multiple languages.
[0181] A "generative AI model" is an artificial intelligence model trained to perform tasks such as text generation and translation.
[0182] A "prompt statement" is an instruction given to a generative AI model to perform a specific task.
[0183] "Means for saving translation results to a database" refers to the function of saving translation results generated by a generative AI model to a database.
[0184] "Means for sending translation results to the terminal" refers to a function that sends translation results stored in the database to the user's terminal.
[0185] "Means by which the terminal displays the translation results" refers to a function in which the user's terminal displays the received translation results on the screen.
[0186] This invention is a system that includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into a generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results.
[0187] Hardware and software to be used
[0188] Hardware:
[0189] Smart devices (smartphones, tablets, etc.)
[0190] Server (a computer with high-performance computing capabilities)
[0191] software:
[0192] Generative AI models (e.g., OpenAI's GPT-4®)
[0193] Database Management System
[0194] Communication protocol (e.g., HTTP / HTTPS)
[0195] Data processing and data calculation
[0196] Product description generation:
[0197] The user takes a picture of a product using a smart device and sends the image to a server. The server uses artificial intelligence to generate a product description from the image. Specifically, it uses an image analysis algorithm to recognize the characteristics and features of the product and generates a product description based on that.
[0198] Prompt message generation:
[0199] The server generates prompt messages using a generative AI model. The prompt message specifies the target language for translation (e.g., English, Chinese, Korean).
[0200] Perform the translation task:
[0201] The server inputs the generated prompt text and product description into the AI model and retrieves the translation results. The AI model then performs multilingual translation based on the input text.
[0202] Saving and sending translation results:
[0203] The server saves the acquired translation results to a database. The saved translation results are sent to the user's terminal, which then displays the translation results.
[0204] Specific example
[0205] Specific example:
[0206] The user enters the Japanese product description: "This product uses high-quality materials."
[0207] The server generates the prompt message: "Translate the following product description into English: This product uses high-quality materials."
[0208] The server inputs the prompt text and product description into the generated AI model and retrieves the English translation result, "This product uses high-quality materials."
[0209] The server saves the translation results to a database and sends them to the terminal.
[0210] The device displays the translation result, and the user confirms it.
[0211] In this way, the server, terminal, and user work together to achieve multilingual translation of product descriptions.
[0212] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0213] Step 1:
[0214] The user enters the product description.
[0215] The user enters the product description in text format using a smart device. The entered product description is then sent from the smart device to the server.
[0216] Input: Product description text
[0217] Output: Product description text sent to the server
[0218] Step 2:
[0219] The server receives the product description.
[0220] The server receives the product description text sent by the user and temporarily stores it in memory.
[0221] Input: Product description text submitted by the user
[0222] Output: Product description text stored in memory
[0223] Step 3:
[0224] The server generates the prompt message.
[0225] The server generates prompt messages based on the languages to be translated. For example, if translation is needed for English, Chinese, and Korean, it will generate prompt messages corresponding to each language.
[0226] Input: Product description text, target language for translation
[0227] Output: Generated prompt message
[0228] Step 4:
[0229] The server inputs prompt text and product description into the generated AI model.
[0230] The server inputs the generated prompt text and product description text into the AI model. The AI model then performs translation based on the input text.
[0231] Input: Prompt text, product description text
[0232] Output: Data input to the generative AI model
[0233] Step 5:
[0234] The server retrieves the translation results.
[0235] The server retrieves translation results from the generative AI model. The retrieved translation results are separated by language.
[0236] Input: Data entered into the generating AI model
[0237] Output: Translated text
[0238] Step 6:
[0239] The server saves the translation results to the database.
[0240] The server saves the retrieved translation results to a database. When saving, it associates the original Japanese product description with the translated result.
[0241] Input: Translated text, original product description text
[0242] Output: Translation results stored in the database
[0243] Step 7:
[0244] The server sends the translation result to the terminal.
[0245] The server sends the saved translation results to the user's device. The transmitted data is provided in a format that the user can access.
[0246] Input: Translation results stored in the database
[0247] Output: Translation results sent to the terminal
[0248] Step 8:
[0249] The device displays the translation result.
[0250] The terminal receives the translation results sent from the server and displays them to the user. The user can review the displayed translation results and make corrections or re-translates as needed.
[0251] Input: Translation result sent from the server
[0252] Output: Translation results displayed on the terminal
[0253] (Application Example 2)
[0254] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0255] In modern brick-and-mortar stores, foreign tourists face language barriers when purchasing goods, making it difficult for them to understand product descriptions. Furthermore, visually impaired individuals and those with reading and writing difficulties also have limited means of understanding product descriptions. This can lead to a diminished customer experience and potentially impact sales.
[0256] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for creating a menu using a smartphone, an artificial intelligence means for generating product descriptions from photos on the smartphone, a means for the artificial intelligence to generate translation work, a means for improving the customer experience, a means for the artificial intelligence to explain questions when considering menus, a means for making recommendations according to preferences, and a means for translating product descriptions into multiple languages and displaying and playing them aloud in the user's set language. This makes it possible to make product descriptions easier to understand for foreign tourists, visually impaired people, and people who have difficulty reading and writing.
[0257] "A method for creating menus using a smartphone" refers to a method for creating menus using a smartphone.
[0258] "An artificial intelligence method for generating product descriptions from smartphone photos" refers to a method that uses artificial intelligence technology to automatically generate product descriptions from photos taken with a smartphone.
[0259] "Methods for generating translation work using artificial intelligence" refers to methods of translating text into multiple languages using artificial intelligence.
[0260] "Methods for improving the customer experience" refer to measures taken to improve the customer's experience when purchasing a product.
[0261] "A means for artificial intelligence to explain questions that arise when customers are considering menu options" refers to a method in which artificial intelligence answers questions that customers may have when considering menu options.
[0262] "Methods of recommending based on preferences" refer to methods of recommending products and services based on customer preferences.
[0263] "Means for translating product descriptions into multiple languages and displaying and playing them in the user's preferred language" refers to means for translating product descriptions into multiple languages and displaying and playing them in the language specified by the user.
[0264] To implement this invention, it is necessary to build a system that combines a smartphone, artificial intelligence technology, a translation API, and a voice playback function.
[0265] First, smartphones have a camera function that allows users to take pictures of products. When a user takes a picture of a product, the image data is sent to a server. The server generates a product description using image recognition technology. This image recognition technology could include, for example, Google® Cloud Vision API.
[0266] Next, the generated product description is translated into multiple languages using artificial intelligence on the server. Translation services such as the Google Cloud Translation API are used for this translation. The translated text is sent to the smartphone and displayed based on the user's language settings.
[0267] Furthermore, the translated product descriptions are played back using the smartphone's audio playback function. This makes the product accessible to visually impaired individuals and those who have difficulty reading or writing.
[0268] As a concrete example, consider a tourist trying to purchase a leather wallet at a physical store in Japan. When the tourist scans the leather wallet's tag with their smartphone camera, the app automatically recognizes the product description and translates it into the user's chosen language (for example, English). The translated product description is displayed as "This is a high-quality leather wallet," and is also played aloud.
[0269] Examples of prompt messages include the following:
[0270] I am developing an application to translate product descriptions into multiple languages. Please translate the following text into the specified languages.
[0271] Text: "This is a high-quality leather wallet."
[0272] Target Language: Japanese
[0273] In this way, when foreign tourists, visually impaired people, or people who have difficulty reading and writing purchase products at physical stores, they can understand product descriptions without feeling the language barrier.
[0274] The flow of the specific process in Application Example 2 will be described using FIG. 14.
[0275] Step 1:
[0276] The user takes a photo of the product using the camera of the smartphone. The input is the product image taken with the camera of the smartphone. The output is the captured product image data.
[0277] Step 2:
[0278] The terminal sends the captured product image data to the server. The input is the product image data. The output is the product image data sent to the server.
[0279] Step 3:
[0280] The server generates a product description using image recognition technology. The input is the product image data. The server extracts text information from the image using, for example, the Google Cloud Vision API and generates a product description. The output is the generated product description text.
[0281] Step 4:
[0282] The server translates the generated product description text into multiple languages. The input is the product description text. The server translates it into the specified language using the Google Cloud Translation API. The output is the translated product description text.
[0283] Step 5:
[0284] The server sends the translated product description text to the terminal. The input is the translated product description text. The output is the translated product description text sent to the terminal.
[0285] Step 6:
[0286] The device displays the translated product description text. The input is the translated product description text. The output is the translated product description displayed on the smartphone screen.
[0287] Step 7:
[0288] The device plays the translated product description text as audio. The input is the translated product description text. The device uses the smartphone's audio playback function to convert the text to audio and play it. The output is the translated product description played as audio.
[0289] (Example 3)
[0290] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0291] Traditional menu creation and customer experience improvement systems fail to adequately answer users' questions when considering menu items and to provide product recommendations tailored to individual preferences. Furthermore, they are insufficient in reducing waiting times and providing multilingual services. This can lead to decreased customer satisfaction and negatively impact store sales.
[0292] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0293] In this invention, the server includes means for the user to input information, means for transmitting the input information to the server, means for the server to analyze the input information using a generated AI model, means for generating recommendations and answers based on the analysis results, means for transmitting the generated recommendations and answers to a terminal, and means for the terminal to display the recommendations and answers to the user. This enables appropriate answers to questions the user has when considering a menu and product recommendations tailored to individual preferences. It also enables reduced waiting times and the provision of multilingual services, leading to improved customer satisfaction and increased store sales.
[0294] A "smart device" is an electronic device used by users to input information and communicate with a server, and includes smartphones and tablets.
[0295] "Artificial intelligence tools" refer to algorithms and models that analyze input data and generate appropriate responses or recommendations.
[0296] A "generative AI model" refers to a machine learning model that analyzes user input information and generates appropriate responses and recommendations.
[0297] "Methods for improving the customer experience" refer to measures that enable customers to reduce waiting times and use services comfortably.
[0298] "A means for artificial intelligence to explain questions when considering menu options" refers to a method by which artificial intelligence provides appropriate answers to questions that users may have when considering menu options.
[0299] "Methods for recommending products according to preferences" refers to methods for recommending appropriate products based on the user's preferences.
[0300] "Means for users to input information" refers to an interface that allows users to input information in formats such as text or voice.
[0301] The means for "transmitting the input information to the server" refers to the communication means for transmitting the information input by the user to the server.
[0302] The means for "the server to analyze the input information using the generated AI model" refers to the means for the server to analyze the information received from the user using the generated AI model.
[0303] The means for "generating recommendations or answers based on the analysis results" refers to the means for generating recommendations or answers for the user based on the analysis results of the generated AI model.
[0304] The means for "transmitting the generated recommendations or answers to the terminal" refers to the communication means for the server to transmit the generated recommendations or answers to the user's terminal.
[0305] The means for "the terminal to display recommendations or answers to the user" refers to the means for the user's terminal to display the recommendations or answers received from the server to the user.
[0306] This invention is a system where a user inputs information using a smart device, the server analyzes the information using a generated AI model, generates appropriate recommendations or answers, and provides them to the user. The following describes specific embodiments of this system.
[0307] First, the user uses a smart device such as a smartphone or tablet to input information to the system. For example, the user inputs "likes spicy food". This information is input through the interface of the smart device.
[0308] ]> Next, the smart device transmits the input information to the server. At this time, the smart device converts the input data into an appropriate format (e.g., JSON format) and transmits it to the server through the network.
[0309] The server analyzes the received user input information using a generation AI model (for example, OpenAI's GPT-4). Specifically, it analyzes information such as "I like spicy food" and searches a database related to spicy food.
[0310] The server recommends products that match the user's preferences based on the analysis results. For example, it generates a list of spicy dishes and presents it to the user. Also, if the user asks, "What are the ingredients in this dish?", the server generates an answer to that question.
[0311] The generated recommendations and responses are sent from the server to the smart device. The server converts the data into the appropriate format and sends it to the smart device over the network.
[0312] Smart devices display recommendations and answers received from a server to the user. For example, they might display, "Recommended spicy dishes are Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot."
[0313] As a concrete example, the following prompt statement is shown.
[0314] "I like spicy food. Can you recommend some dishes?"
[0315] "What are the ingredients in this dish?"
[0316] In this way, a system is realized in which the server, smart device, and user work together to improve the customer experience. Through this system, users can receive appropriate answers to questions when considering menus and product recommendations tailored to their individual preferences. It can also reduce waiting times and provide multilingual services, leading to improved customer satisfaction and increased store sales. The flow of specific processing in Example 3 will be explained using Figure 15.
[0317] Step 1:
[0318] The user enters information using a terminal.
[0319] The user opens the application on their smartphone or tablet, types "I like spicy food" into the text box, and presses the submit button. The input data is in text format and represents information about the user's preferences.
[0320] Step 2:
[0321] The terminal sends the entered information to the server.
[0322] The terminal converts the user's input, "I like spicy food," into JSON format and sends it to the server using the HTTPS protocol. The input data is user preference information in text format, and the output data is request data in JSON format.
[0323] Step 3:
[0324] The server analyzes the input information using a generated AI model.
[0325] The server parses the received JSON data and inputs the information "likes spicy food" into a generative AI model. The generative AI model (e.g., GPT-4) searches a database of spicy dishes and generates a list of related dishes. The input data is request data in JSON format, and the output data is a list of dishes as a result of the analysis.
[0326] Step 4:
[0327] The server generates recommendations and responses based on the analysis results.
[0328] The server generates a list of spicy dishes based on the analysis results obtained from the generated AI model. For example, it might list dish names such as "Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot." The input data is the list of dishes resulting from the analysis, and the output data is a recommendation list presented to the user.
[0329] Step 5:
[0330] The server sends recommendations and responses it generates to the device.
[0331] The server converts the generated list of spicy dishes into JSON format and sends it to the terminal using the HTTPS protocol. The input data is a recommendation list, and the output data is a response in JSON format.
[0332] Step 6:
[0333] The device displays recommendations and answers to the user.
[0334] The terminal parses the received JSON data and displays to the user, "Recommended spicy dishes are Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot." The input data is response data in JSON format, and the output data is text information displayed to the user.
[0335] (Application Example 3)
[0336] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0337] Traditional food delivery systems have problems such as difficulty for customers to obtain detailed information when choosing a menu and a lack of personalized recommendations. Furthermore, the lack of a means to provide real-time information on ingredients and nutritional content has been a factor in lowering customer satisfaction.
[0338] In Application Example 3, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a menu creation means using a smart device, an artificial intelligence means for generating product descriptions from images on the smart device, a means for the artificial intelligence to generate translation work, a means for improving the customer experience by shortening waiting times, a means for the artificial intelligence to explain questions when considering menus, a means for making recommendations according to preferences, a means for recommending dishes based on customer preferences, and a means for providing information on the ingredients of dishes. As a result, customers can obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[0339] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone or tablet.
[0340] A "menu creation method" is a function that generates a list of dishes and drinks that customers can choose from.
[0341] "An artificial intelligence method for generating product descriptions from images" refers to a function that analyzes images taken with a smart device and automatically generates product descriptions based on their content.
[0342] "Means of generating translation work using artificial intelligence" refers to a function that automatically translates text between different languages.
[0343] "Methods to improve the customer experience by reducing waiting times" refer to functions that reduce the waiting time customers have when using a service.
[0344] "A means for artificial intelligence to explain questions when considering menus" refers to a function in which artificial intelligence provides appropriate answers to questions that customers may have when choosing from a menu.
[0345] "Methods for recommending based on preferences" refer to functions that recommend appropriate products and services based on the customer's preferences and past choices.
[0346] "A method for recommending dishes based on customer preferences" refers to a function that recommends appropriate dishes based on the preference information entered by the customer.
[0347] "Means of providing information on the ingredients of a dish" refers to a function that provides information about the ingredients and components contained in a particular dish.
[0348] A system for carrying out this invention includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images on a smart device, means for artificial intelligence to generate translation work, means for improving the customer experience by reducing waiting time, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recommending dishes based on customer preferences, and means for providing information on the ingredients of dishes.
[0349] System program
[0350] The system's program is implemented using Python and the OpenAI API. The server receives customer input and uses a generative AI model to provide appropriate recommendations and information.
[0351] Explanation of the process
[0352] The server receives customer preferences and questions sent from smart devices and generates prompts based on them. These generated prompts are input into a generation AI model via the OpenAI API, which then generates appropriate answers and recommendations. The generated information is then sent back to the smart device and provided to the customer.
[0353] The hardware used will be smart devices such as smartphones and tablets. The software used will be Python and the OpenAI API.
[0354] Specific example
[0355] For example, if a customer enters "I like spicy food," the server will generate the following prompt:
[0356] "Please recommend some dishes for people who like spicy food."
[0357] When this prompt is entered into the OpenAI API, the generating AI model will recommend spicy dishes such as "Mapo Tofu" and "Kimchi Hot Pot."
[0358] Furthermore, if a customer asks, "What are the ingredients in Mapo Tofu?", the server will generate the following prompt:
[0359] "What are the ingredients in Mapo Tofu?"
[0360] When this prompt is entered into the OpenAI API, the generating AI model provides ingredient information such as "tofu, ground pork, green onions, garlic, ginger, chili bean paste, sweet bean paste, soy sauce, sake, sugar, and chicken broth."
[0361] This allows customers to obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[0362] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[0363] Step 1:
[0364] The user launches the application using a smart device and enters their preferences and answers questions. The entered information is sent to the server in text format. Examples of input include "I like spicy food" and "What are the ingredients in mapo tofu?".
[0365] Step 2:
[0366] The server analyzes the input information received from the user and generates an appropriate prompt. For example, in response to the input "I like spicy food," it generates the prompt "Please recommend some dishes for someone who likes spicy food." This prompt is then input into the generation AI model.
[0367] Step 3:
[0368] The server sends the generated prompt to the OpenAI API and retrieves appropriate answers and recommendations based on the generating AI model. For example, in response to the prompt "Please recommend some dishes for someone who likes spicy food," answers such as "Mapo Tofu" and "Kimchi Hot Pot" are generated.
[0369] Step 4:
[0370] The server analyzes the responses and recommendations obtained from the generating AI model and converts them into a format for the user. For example, it formats responses such as "Mapo Tofu" and "Kimchi Hot Pot" into a list.
[0371] Step 5:
[0372] The server sends formatted responses and recommendations to the smart device. The user can then view this information on the smart device's screen.
[0373] Step 6:
[0374] Users can select menu items based on the information provided or ask further detailed questions. For example, they might re-enter the question, "What are the ingredients in Mapo Tofu?"
[0375] Step 7:
[0376] The server receives input from the user again and repeats the same process. Specifically, it generates a prompt again, inputs it into the generation AI model, retrieves the answer, formats it, and provides it to the user.
[0377] This allows users to obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[0378] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0379] "Example of form 1"
[0380] One embodiment of the present invention provides a system that incorporates an emotion engine. This system recognizes the user's emotions and recommends products accordingly. Specifically, when a user is choosing a product, the system recognizes their emotions from their facial expressions and tone of voice and recommends products that match those emotions. For example, if a user shows a joyful expression, the system recommends products that are likely to make that user happy.
[0381] "Example of form 2"
[0382] Furthermore, the emotion engine generates product descriptions based on the user's emotions. Specifically, it recognizes the user's emotions and selects words that match those emotions to generate a product description. For example, if the user is feeling down, it will select words that will cheer them up to generate a product description. (Example 3)
[0383] Furthermore, the emotion engine recognizes the user's emotions and provides services accordingly. Specifically, it recognizes the user's emotions and provides services that match those emotions. For example, if a user is showing anger, it will provide services such as offering an apology to that user.
[0384] The following describes the processing flow for each example of the form.
[0385] "Example of form 1"
[0386] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[0387] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[0388] Step 3: The system recommends products that match the recognized emotions.
[0389] "Example of form 2"
[0390] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[0391] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[0392] Step 3: The system selects words that match the recognized emotion and generates a product description.
[0393] "Example of form 3"
[0394] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[0395] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[0396] Step 3: The system provides services that match the recognized emotions.
[0397] (Example 1)
[0398] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0399] Traditional restaurant menu creation and customer experience improvement have been time-consuming and labor-intensive. In particular, generating product descriptions from photos and providing personalized recommendations consume significant human resources. Furthermore, multilingual services for foreign tourists are inadequate, highlighting the need for improved customer satisfaction. Additionally, the inability to provide product recommendations based on customer emotions can lead to a decline in the quality of the customer experience. An efficient system is needed to address these challenges.
[0400] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0401] In this invention, the server includes means for creating menus using a smartphone, means for artificial intelligence to generate product descriptions from photos taken on the smartphone, means for artificial intelligence to generate translation tasks, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the cash register, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, and means for recognizing the user's emotions and recommending products based on those emotions. This enables more efficient menu creation, product recommendations tailored to customer preferences, enhanced service for foreign tourists through multilingual support, and product recommendations based on customer emotions.
[0402] "A method for creating menus using a smartphone" refers to a method for creating restaurant menus using the camera and applications of a smartphone.
[0403] "An artificial intelligence method for generating product descriptions from smartphone photos" refers to a method that uses artificial intelligence technology to analyze product photos taken with a smartphone, recognize their characteristics and features, and automatically generate product descriptions.
[0404] "Methods for generating translation work using artificial intelligence" refers to methods of translating product descriptions and menu contents into multiple languages using artificial intelligence technology.
[0405] "Methods to improve the customer experience by eliminating the need to wait for waitstaff or in line at the register" refers to methods that allow customers to receive service smoothly without having to wait for waitstaff or in line at the register.
[0406] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a method in which artificial intelligence automatically provides answers to questions that customers may have when considering menu options.
[0407] "Methods of recommending based on preferences" refer to methods of recommending appropriate products based on a customer's past choices and preferences.
[0408] "A method for recognizing user emotions and recommending products based on those emotions" refers to a method of recognizing emotions by analyzing the user's facial expressions and tone of voice, and then recommending products that match those emotions.
[0409] This invention is a system that allows restaurant operators to create menus using smartphones and improve the customer experience. A specific embodiment of this system is described below.
[0410] First, the user takes a picture of the product using their smartphone camera. For example, to take a picture of a new dessert, they launch the smartphone's camera app, frame the dessert, and press the shutter button.
[0411] Next, the device uploads the captured photos to the server. Specifically, the smartphone application sends the photo data to the server via an HTTP request. At this time, the photo data is sent in an appropriate format (e.g., JPEG, PNG).
[0412] The server analyzes the received photos using image recognition software (e.g., Google Cloud Vision API). The server recognizes the characteristics of the product from the photo (e.g., chocolate cake, fruit parfait). For example, the Google Cloud Vision API might return the label "chocolate cake".
[0413] The server then generates a product description based on the recognized characteristics. Using a generation AI model (e.g., OpenAI GPT-3), it generates a product description such as, "A rich chocolate cake. It is not too sweet and has a moist texture."
[0414] The generated product descriptions are displayed on the restaurant's website or application by the server. Specifically, the product descriptions are sent to the front-end in HTML or JSON format and displayed in the user interface.
[0415] Furthermore, when a user selects a product, the device uses its camera and microphone to capture the user's facial expressions and voice tone. For example, the smartphone's camera captures the user's smile, and the microphone records the user's voice tone.
[0416] The server analyzes the captured data using an emotion engine (e.g., Microsoft® Azure® Emotion API) to recognize the user's emotions. Based on the recognized emotions, the server recommends products that are suitable for the user. For example, if the user shows a joyful expression, the server will recommend a "fruit parfait."
[0417] As a concrete example, consider a scenario where a user takes a photo of a new dessert with their smartphone and uploads it to the system. The server uses the Google Cloud Vision API to analyze the photo and recognizes its characteristic as "chocolate cake." The AI then generates a product description such as, "A rich chocolate cake. It's not too sweet and has a moist texture."
[0418] Furthermore, if a user displays an expression of joy while viewing the menu using their smartphone camera, the Microsoft Azure Emotion API recognizes that expression. The server then recommends a "fruit parfait" that the user is likely to enjoy.
[0419] Example of a prompt:
[0420] "Please upload photos of the product taken with your smartphone. Our AI will automatically recognize the product's characteristics and generate a product description. It will also use your camera and microphone to recommend products based on your emotions."
[0421] In this way, a system is realized in which servers, terminals, and users work together to efficiently create restaurant menus and recommend products to customers.
[0422] The flow of the specific processing in Example 1 will be explained using Figure 17.
[0423] Step 1:
[0424] The user takes a photo of the product.
[0425] The user takes a picture of the product using their smartphone camera. For example, to take a picture of a new dessert, the user launches the smartphone's camera app, frames the dessert, and presses the shutter button. The input is the photograph taken, and the output is the image data stored on the smartphone.
[0426] Step 2:
[0427] The device uploads the photo to the server.
[0428] The device uploads the captured photos to the server. Specifically, the smartphone application sends the photo data to the server via an HTTP request. The photo data is sent in an appropriate format (e.g., JPEG, PNG). The input is the image data on the smartphone, and the output is the image data sent to the server.
[0429] Step 3:
[0430] The server analyzes the photos and recognizes the characteristics of the product.
[0431] The server analyzes the received photograph using image recognition software (e.g., an image recognition API). The server recognizes the characteristics of the product from the photograph (e.g., chocolate cake, fruit parfait). For example, the image recognition API might return the label "chocolate cake". The input is the image data sent to the server, and the output is the recognized product characteristic information.
[0432] Step 4:
[0433] The server generates the product description.
[0434] The server generates a product description based on the recognized characteristics. Using a generative AI model (e.g., Generative AI Model), it generates a product description such as, "A rich chocolate cake. It is not too sweet and has a moist texture." The input is product characteristic information, and the output is the generated product description.
[0435] Step 5:
[0436] The server displays product descriptions on websites and applications.
[0437] The server displays the generated product descriptions on the restaurant's website or application. Specifically, it sends the product descriptions to the front-end in HTML or JSON format and displays them in the user interface. The input is the generated product description, and the output is the product description displayed on the website or application.
[0438] Step 6:
[0439] When a user chooses a product, the device captures the user's emotions.
[0440] When a user selects a product, the device uses its camera and microphone to capture the user's facial expressions and voice tone. For example, a smartphone's camera captures the user's smile, and the microphone records the user's voice tone. The input is the user's facial expressions and voice tone, and the output is the captured emotional data.
[0441] Step 7:
[0442] The server analyzes emotions and recommends appropriate products.
[0443] The server analyzes the captured data using an emotion engine (e.g., an emotion analysis API) to recognize the user's emotions. Based on the recognized emotions, the server recommends products that are suitable for the user. For example, if the user shows a joyful expression, the server will recommend a "fruit parfait." The input is the captured emotion data, and the output is the recommended product information.
[0444] (Application Example 1)
[0445] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0446] Traditional restaurants faced significant challenges in menu creation and improving the customer experience, requiring considerable time and effort. Specifically, there were multiple obstacles, including menu creation, product description generation, personalized recommendations, and multilingual support. Furthermore, providing service that resonated with customers' emotions was difficult, making improving customer satisfaction a challenge.
[0447] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a means for creating menus using a smart device, an AI means for generating product descriptions from photos on the smart device, a means for the AI to generate translation work, a means for improving the customer experience, a means for the AI to explain questions when considering menus, a means for making recommendations according to preferences, and a means for recognizing customer emotions and recommending products based on those emotions. This enables more efficient menu creation, automatic generation of product descriptions, multilingual support, improved customer experience, and emotion-based recommendations.
[0448] A "smart device" is a portable electronic device with advanced functions, such as a smartphone or smart glasses.
[0449] "Menu creation method" refers to the functions and methods for creating restaurant menus using smart devices.
[0450] "AI methods for generating product descriptions" refers to artificial intelligence technology that automatically recognizes the characteristics and features of a product from a photo taken with a smart device and generates a product description based on that.
[0451] "A means of generating translation work using AI" refers to a function that automatically performs multilingual translation using artificial intelligence.
[0452] "Methods for improving the customer experience" refer to methods and functions that enhance the customer experience, such as eliminating the need to wait in line with waitstaff or at the cash register.
[0453] "AI-powered explanations for questions during menu consideration" refers to a function where artificial intelligence automatically explains questions that arise when customers are considering menu options.
[0454] "A means of recommending based on preferences" refers to a function that recommends appropriate products based on the customer's preferences.
[0455] "A means of recognizing customer emotions and recommending products based on those emotions" refers to a function that recognizes customer emotions from facial expressions, tone of voice, etc., and recommends products that match those emotions.
[0456] A system for carrying out this invention includes means for creating menus using a smart device, means for generating product descriptions from photos on a smart device, means for AI to generate translations, means for improving the customer experience, means for AI to explain questions when considering menus, means for making recommendations according to preferences, and means for recognizing customer emotions and recommending products based on those emotions.
[0457] Program Processing Description
[0458] hardware
[0459] The server will be equipped with a high-performance GPU. Smart devices such as smartphones and smart glasses will be used as terminals.
[0460] software
[0461] The server uses the following software:
[0462] OpenCV: Image Processing Library
[0463] Keras: A deep learning library
[0464] Transformers: Generative AI models (such as GPT-3)
[0465] Data processing and calculations
[0466] 1. Menu creation method:
[0467] Take a picture of the food with the device's camera and upload it to the server.
[0468] The server uses OpenCV to preprocess images and Keras to recognize product characteristics and features from the images.
[0469] Based on the recognition results, product descriptions are generated using Transformers.
[0470] 2. Methods for AI to generate translation work:
[0471] The server uses a generative AI model to translate the generated product descriptions into multiple languages.
[0472] In particular, we will implement multilingual support to enhance services for foreign tourists.
[0473] 3. Means of improving the customer experience:
[0474] The terminals allow customers to complete ordering and payment on their smart devices, eliminating the need to wait in line with waitstaff or at the cash register.
[0475] 4. How AI can explain questions that arise when considering menu options:
[0476] The terminal sends questions that customers have when considering the menu to the server, and the server automatically generates explanations using AI.
[0477] 5. Methods for recommending based on preferences:
[0478] The server recommends appropriate products based on the customer's past order history and preferences.
[0479] 6. Means for recognizing customer emotions and recommending products based on those emotions:
[0480] The server analyzes the customer's facial expressions and tone of voice through the device's camera, and uses Keras to recognize their emotions.
[0481] Based on recognized emotions, Transformers are used to recommend appropriate products.
[0482] Specific examples and prompt statements
[0483] Specific example:
[0484] When a customer is wearing smart glasses and looking at a menu, the smart glasses analyze the customer's facial expression and display a message saying, "You seem happy, so we recommend dessert!"
[0485] Examples of prompts for a generative AI model:
[0486] "This dish looks like delicious pasta. Please generate a detailed product description."
[0487] Thus, the embodiment of the invention involves combining smart devices and AI technology to enable menu creation and improve the customer experience in restaurants.
[0488] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[0489] Step 1:
[0490] The user takes a photo of the food using a smart device (smartphone or smart glasses). The photo is saved on the device.
[0491] Step 2:
[0492] The device uploads the captured photos to the server. The server uses OpenCV to preprocess the images, specifically performing tasks such as resizing and noise reduction.
[0493] Step 3:
[0494] The server analyzes pre-processed images using Keras to recognize the characteristics and features of the product. The input is a pre-processed image, and the output is data indicating the characteristics and features of the product.
[0495] Step 4:
[0496] The server generates product descriptions using Transformers based on recognized characteristics and features. The input is data describing the characteristics and features of the product, and the output is the generated product description.
[0497] Step 5:
[0498] The server uses a generative AI model to translate the generated product descriptions into multiple languages. The input is the generated product description, and the output is the product description translated into multiple languages.
[0499] Step 6:
[0500] The terminal allows customers to complete their orders and payments on their smart devices, eliminating the need to wait in line with waiters or at the cash register. The input is the customer's order information, and the output is a notification that the order has been completed.
[0501] Step 7:
[0502] The terminal sends questions that arise when customers are considering the menu to the server. The server automatically generates explanations using AI. The input is the customer's question, and the output is the generated explanation.
[0503] Step 8:
[0504] The server recommends appropriate products based on the customer's past order history and preferences. The input is the customer's order history and preference data, and the output is the recommended products.
[0505] Step 9:
[0506] The server analyzes the customer's facial expressions and voice tone through the terminal's camera, and recognizes their emotions using Keras. The input is data on the customer's facial expressions and voice tone, and the output is the recognized emotion.
[0507] Step 10:
[0508] The server uses Transformers to recommend appropriate products based on the recognized emotions. The input is the recognized emotions, and the output is the recommended products.
[0509] (Example 2)
[0510] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0511] Modern consumers demand quick and accurate product descriptions, as well as multilingual information. They also require personalized product descriptions that cater to their emotions. However, traditional systems struggle to meet these demands, making the improvement of services, particularly for foreign tourists, a significant challenge.
[0512] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0513] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages. As a result, users can quickly and accurately understand product descriptions, and information can be provided in multiple languages. Furthermore, by providing personalized product descriptions that correspond to the user's emotions, customer satisfaction can be improved.
[0514] A "smart device" is a portable electronic device that has internet connectivity and can run applications.
[0515] "Menu creation method" refers to a function that allows restaurants and service businesses to create menus using smart devices.
[0516] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[0517] "Means of generating translation work using artificial intelligence" refers to artificial intelligence technology for translating generated product descriptions into multiple languages.
[0518] "Means of improving the customer experience" refer to features that enhance the convenience and satisfaction customers experience when using a service.
[0519] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a function that allows artificial intelligence to provide appropriate answers to questions that arise when customers are considering menu options.
[0520] "Methods for recommending based on preferences" refer to functions that recommend appropriate products and services based on a customer's past choices and preferences.
[0521] "A means of recognizing user emotions and generating product descriptions that correspond to those emotions" refers to artificial intelligence technology that analyzes user emotions from their input and actions and generates product descriptions that are appropriate for those emotions.
[0522] "Means for translating generated product descriptions into multiple languages" refers to a function for translating generated product descriptions into multiple languages.
[0523] Modes for carrying out the invention
[0524] This invention is a system that includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages.
[0525] Hardware and software to be used
[0526] The server generates product descriptions using a generative AI model (e.g., OpenAI's GPT-4). The generated product descriptions are translated into multiple languages using a translation engine (e.g., Google Translate API). To recognize user emotions, an emotion engine (e.g., IBM Watson®'s Sentiment Analysis API) is used. Smart devices are portable electronic devices with internet connectivity capable of running applications.
[0527] Explanation of the program's processing
[0528] The server receives text entered by the user through the terminal and generates a product description using a generative AI model. For example, if the user enters "This product is high quality and long-lasting," the server will generate a product description based on this text.
[0529] The generated product descriptions are translated into multiple languages using a translation engine. For example, they are translated into major tourist languages such as English, Chinese, and Korean. Specifically, a description like "This product is high quality and long-lasting" is translated.
[0530] When a user inputs emotional text or audio data through their device, the server uses an emotion engine to recognize the user's emotions. For example, if a user inputs "I'm feeling down today," the server analyzes this text and recognizes the user's emotion as "down."
[0531] The server uses a generative AI model to generate emotionally appropriate product descriptions based on the user's recognized emotions. For example, if the user is feeling down, it will select uplifting words to generate a product description.
[0532] Original description: "This product is high quality and long-lasting."
[0533] Description tailored to the emotion: "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[0534] Examples of specific cases and prompt statements
[0535] As a concrete example, consider the following scenario.
[0536] Scenario 1: Multilingual Translation
[0537] If a user enters "This product is high quality and long-lasting" in Japanese, the server will translate this product description into English, Chinese, and Korean.
[0538] Scenario 2: Product Description Based on Emotions
[0539] If a user enters "I'm feeling down today," the server uses an emotion engine to recognize the user's emotions and generates a product description that will cheer them up.
[0540] Original description: "This product is high quality and long-lasting."
[0541] Description tailored to the emotion: "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[0542] Example of a prompt
[0543] Examples of prompt statements for a generative AI model are as follows:
[0544] Please translate this Japanese product description into English, Chinese, and Korean: 'This product is high quality and long-lasting.'
[0545] "Please generate a product description that will cheer up a user who is feeling down. The original description is 'This product is high quality and long-lasting.'"
[0546] The above describes the embodiments for carrying out this invention.
[0547] The flow of the specific processing in Example 2 will be explained using Figure 19.
[0548] Step 1:
[0549] The user uses a terminal to input the text that will form the basis of the product description. For example, the user might input, "This product is high quality and long-lasting." The entered text is then sent to the server.
[0550] Step 2:
[0551] The server uses a generative AI model to generate a product description based on the text entered by the user. Specifically, the server sends the following prompt to the generative AI model: "Generate a product description based on the user's input, 'This product is high quality and long-lasting.'" The generative AI model generates a product description based on this prompt and returns it to the server. The output is the generated product description.
[0552] Step 3:
[0553] The server sends the generated product description to the translation engine for translation into multiple languages. Specifically, the server sends the following prompt to the translation engine: "Translate this product description into English, Chinese, and Korean: 'This product is high quality and long-lasting.'" The translation engine translates based on this prompt and returns it to the server. The output is the product description translated into multiple languages.
[0554] Step 4:
[0555] The user inputs text or audio data related to their emotions through their device. For example, the user might input, "I'm feeling down today." The input text or audio data is then sent to the server.
[0556] Step 5:
[0557] The server uses an emotion engine to recognize the user's emotions. Specifically, the server sends the following prompt to the emotion engine: "Analyze the user's input 'I'm feeling down today' and recognize the emotion." Based on this prompt, the emotion engine analyzes the emotion and returns it to the server. The output is the recognized emotion.
[0558] Step 6:
[0559] Based on the recognized user's emotions, the server uses a generative AI model to generate an emotion-responsive product description. Specifically, the server sends the following prompt to the generative AI model: "Generate an uplifting product description for a user who is feeling down. The original description is 'This product is high quality and long-lasting.'" Based on this prompt, the generative AI model generates an emotion-responsive product description and returns it to the server. The output is an emotion-responsive product description.
[0560] Step 7:
[0561] The server sends a product description to the user's device that corresponds to the generated emotion. The user can then view the uplifting product description through their device. Specifically, if the user enters "I'm feeling down today," the server will provide the user with a product description such as, "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[0562] The above is a detailed explanation of the program's processing flow.
[0563] (Application Example 2)
[0564] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0565] Traditional e-commerce sites often provide product descriptions in a single language, which posed a problem for foreign tourists and multilingual users, making them difficult to understand. Furthermore, the lack of emotionally resonant product descriptions made it difficult to increase user purchase intent. Additionally, users sometimes spent a significant amount of time trying to understand product descriptions, resulting in a less-than-ideal shopping experience.
[0566] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0567] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages. This makes it possible to provide product descriptions that correspond to the user's emotions in multiple languages, and to provide product descriptions that are easy to understand for foreign tourists and multilingual users. Furthermore, it can increase the user's desire to purchase and improve the purchasing experience.
[0568] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.
[0569] "Menu creation methods" refer to functions or software that allow users to create menus for products and services using smart devices.
[0570] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[0571] "Means of generating translation work using artificial intelligence" refers to artificial intelligence technology for translating generated product descriptions into multiple languages.
[0572] "Methods for improving the customer experience" refer to functions and software that enhance the customer's experience when using a product or service.
[0573] "A means for artificial intelligence to explain questions when considering menu options" refers to a function in which artificial intelligence provides appropriate explanations for questions that arise when users are considering menu options.
[0574] "Preference-based recommendation methods" refer to features that recommend appropriate products and services based on the user's preferences and past behavior.
[0575] "A means of recognizing user emotions and generating product descriptions that correspond to those emotions" refers to artificial intelligence technology that analyzes user emotions and generates product descriptions that are appropriate for those emotions.
[0576] "Means for translating generated product descriptions into multiple languages" refers to functions or software for translating generated product descriptions into multiple languages.
[0577] This invention will now describe embodiments for carrying out this invention. The system includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages.
[0578] System Configuration
[0579] hardware
[0580] Smart devices: Portable electronic devices with internet connectivity, such as smartphones, tablets, and smart glasses.
[0581] Server: A high-performance computer responsible for data processing and storage.
[0582] software
[0583] Artificial intelligence model: Generates product descriptions using the OpenAI API.
[0584] Translation software: Multilingual translation is performed using the Google Translate API.
[0585] Emotion recognition software: Use EmotionRecognizer to recognize the user's emotions.
[0586] Data processing and data calculation
[0587] 1. Image Analysis: Images of products taken with a smart device are sent to the server. The server uses an image analysis algorithm to extract product characteristics and features and generate a product description.
[0588] 2. Emotion Recognition: The smart device's camera and microphone are used to analyze the user's facial expressions and voice tone. EmotionRecognizer software recognizes the user's emotions and sends the results to the server.
[0589] 3. Product Description Generation: The server generates prompt text based on the user's emotions and uses the OpenAI API to generate a product description appropriate to those emotions.
[0590] 4. Multilingual Translation: The generated product description is translated into multiple languages using the Google Translate API. The translation results support major languages such as English, Chinese, and Korean.
[0591] Specific example
[0592] For example, if a user is feeling down, the smart device's camera captures the user's facial expression, and EmotionRecognizer recognizes it as "feeling down." The server then generates the following prompt:
[0593] "Generate a product description that will cheer up a user when they're feeling down: This product is made from high-quality materials and will last a long time."
[0594] This prompt is sent to the OpenAI API to generate a sentiment-appropriate product description. The generated product description will look like this:
[0595] "This product was created to brighten your life, even just a little. It's made from high-quality materials and is built to last."
[0596] Next, we will translate this product description into multiple languages using the Google Translate API.
[0597] In this way, it becomes possible to provide product descriptions in multiple languages that are tailored to the user's emotions.
[0598] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[0599] Step 1:
[0600] The user takes a picture of the product with a smart device.
[0601] Input: Product image
[0602] Output: Captured image data
[0603] Specific operation: The user uses a smart device such as a smartphone or tablet to take a picture of the product they are considering purchasing. This image data is stored on the smart device.
[0604] Step 2:
[0605] The device sends the captured image to the server.
[0606] Input: Captured image data
[0607] Output: Image data sent to the server
[0608] Specific operation: The smart device sends the captured image data to a server via the internet. The HTTP protocol and other protocols are used for transmission.
[0609] Step 3:
[0610] The server performs image analysis to extract the product's characteristics and features.
[0611] Input: Sent image data
[0612] Output: Characteristics and features of the extracted products
[0613] Specific operation: The server uses image analysis algorithms to automatically extract product characteristics and features from image data. For example, it recognizes the product's color, shape, brand logo, etc.
[0614] Step 4:
[0615] The server generates a product description based on the extracted characteristics and features.
[0616] Input: Characteristics and features of the extracted products
[0617] Output: Generated product description
[0618] Specific operation: The server uses a generative AI model to generate product descriptions based on extracted characteristics and features. For example, a description such as "This product is made from high-quality materials and will last a long time" might be generated.
[0619] Step 5:
[0620] The device recognizes the user's emotions.
[0621] Input: User's facial expressions and tone of voice
[0622] Output: Recognized user emotions
[0623] Specific operation: Using the camera and microphone of a smart device, the EmotionRecognizer software captures the user's facial expressions and voice tone, and recognizes the user's emotions. For example, emotions such as "depressed" or "happy" can be recognized.
[0624] Step 6:
[0625] The server generates prompt messages that respond to the user's emotions.
[0626] Input: Recognized user sentiment, generated product description
[0627] Output: Prompts tailored to your emotions
[0628] Specific operation: Based on the recognized user's emotions, the server generates prompt messages to send to the generative AI model. For example, a prompt message such as "Generate a product description to cheer up a user who is feeling down: This product is made from high-quality materials and will last a long time." might be generated.
[0629] Step 7:
[0630] The server uses an AI model to generate product descriptions that are appropriate for the user's emotions.
[0631] Input: Sentiment-based prompt text
[0632] Output: Product description suitable for emotions
[0633] Specific operation: The server uses a generative AI model (e.g., the OpenAI API) to generate emotionally appropriate product descriptions based on the prompt text. For example, a description such as "This product was created to brighten your life a little. It is made from high-quality materials and is durable." might be generated.
[0634] Step 8:
[0635] The server translates the generated product description into multiple languages.
[0636] Input: Generated product description
[0637] Output: Product description translated into multiple languages
[0638] Specific operation: The server uses the Google Translate API to translate the generated product description into multiple languages, such as English, Chinese, and Korean. For example, a translation result such as "This product is made to brighten your life a little. It is made of high-quality materials and lasts long." might be obtained.
[0639] Step 9:
[0640] The server sends product descriptions translated into multiple languages to the device.
[0641] Input: Product description translated into multiple languages
[0642] Output: Translation results sent to the terminal
[0643] Specific operation: The server sends product descriptions translated into multiple languages to smart devices via the internet. HTTP protocol and similar protocols are used for transmission.
[0644] Step 10:
[0645] The device displays product descriptions translated into multiple languages to the user.
[0646] Input: Translation result sent to the device
[0647] Output: Multilingual product description displayed to the user
[0648] Specific operation: The smart device displays the product description, translated into multiple languages, to the user. This allows the user to understand the emotionally relevant product description in multiple languages.
[0649] (Example 3)
[0650] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0651] Traditional customer experience improvement systems often rely on manual processes for menu creation, product descriptions, and translation, resulting in inefficiency. Furthermore, they may fail to adequately address customer questions and provide personalized recommendations, potentially leading to decreased customer satisfaction. Additionally, the lack of features to recognize and respond to customer emotions hinders the improvement of the overall customer experience.
[0652] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 3 is realized by the following means. In this invention, the server includes a menu creation means using a smart device, an artificial intelligence means for generating product descriptions from images on the smart device, a means for the artificial intelligence to generate translation work, a means for improving the customer experience by reducing waiting time, a means for the artificial intelligence to explain questions when considering a menu, a means for recommending products according to preferences, and a means for recognizing the user's emotions and providing services that correspond to those emotions. This makes it possible to quickly answer questions from customers when considering a menu, recommend products according to their preferences, and further improve customer satisfaction by providing services that correspond to the customer's emotions.
[0653] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone or tablet.
[0654] "Menu creation method" refers to a function that allows restaurants and service businesses to create menus using smart devices.
[0655] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[0656] "Means of generating translation work using artificial intelligence" refers to a function that automatically performs text and audio translation using artificial intelligence.
[0657] "Methods for improving customer experience by reducing waiting times" refer to functions that shorten the waiting time customers have to wait when using a service, thereby providing a smoother experience.
[0658] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a function in which artificial intelligence provides appropriate answers to questions that customers may have when considering menu options.
[0659] "A method for recommending products based on preferences" refers to a function in which artificial intelligence recommends appropriate products based on the customer's preferences and past choices.
[0660] "Means of recognizing user emotions and providing services that respond to those emotions" refers to a function that recognizes user emotions from their facial expressions and voice and provides appropriate services that respond to those emotions.
[0661] This invention is a system for improving the customer experience and includes means for creating menus using smart devices, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translations, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and means for recognizing user emotions and providing services that respond to those emotions.
[0662] Hardware and software to be used
[0663] hardware
[0664] Smart devices: Portable electronic devices with internet connectivity, such as smartphones and tablets.
[0665] Server: A high-performance server (e.g., a server with an NVIDIA GPU).
[0666] software
[0667] Generative AI model: GPT-4 is used as an example.
[0668] Emotion recognition engine: Affectiva SDK is used as an example.
[0669] Program Processing Description
[0670] The server receives information sent from the smart device and performs analysis using a generative AI model. For example, if the user inputs "I like spicy food," the generative AI model generates a list of spicy dishes, and the server sends that list to the smart device. The smart device then displays the received list to the user.
[0671] Furthermore, when a user asks, "What are the ingredients in this dish?", the server uses a generative AI model to generate ingredient information and sends it to the smart device. The smart device then displays the received ingredient information to the user.
[0672] Furthermore, smart devices use their built-in cameras and microphones to capture the user's facial expressions and voice, and use an emotion recognition engine to recognize the user's emotions. For example, if the user shows an angry expression, the smart device sends that information to a server. The server uses a generative AI model to generate an apology and sends it to the smart device. The smart device then displays the received apology to the user.
[0673] Examples of specific cases and prompt statements
[0674] Example 1: User enters "I like spicy food"
[0675] The user opens the app on their smartphone and enters "I like spicy food."
[0676] The server generates a list of spicy dishes using an AI model.
[0677] The server generates a list and sends it to the smart device.
[0678] A smart device displays a list of spicy dishes to the user.
[0679] Example prompt: "Please list dishes you would recommend to a customer who likes spicy food."
[0680] Example 2: The user asks, "What are the ingredients in this dish?"
[0681] The user opens the app on their tablet and asks, "What are the ingredients in this dish?"
[0682] The server generates ingredient information using an AI model.
[0683] The server generates ingredient information and sends it to the smart device.
[0684] A smart device displays ingredient information to the user.
[0685] Example of a prompt: "What are the ingredients in this dish?"
[0686] Example 3: The user shows an angry expression.
[0687] The user displays an angry expression while using their smartphone.
[0688] Smart devices use cameras to capture the user's facial expressions, which are then analyzed by an emotion recognition engine.
[0689] The smart device sends information about the anger it recognizes to the server.
[0690] The server generates an apology using an AI model.
[0691] The server generates an apology message and sends it to the smart device.
[0692] Smart devices display an apology message to the user.
[0693] Example prompt: "Generate an apology for when the user is angry."
[0694] In this way, the server, smart device, and user work together to operate the system and improve the customer experience. The flow of a specific process in Example 3 will be explained using Figure 21.
[0695] Step 1:
[0696] The user inputs information through their device. The user uses a smartphone or tablet to input questions and preferences into the system. For example, they might input "I like spicy food." The input data is saved on the device in text format.
[0697] Step 2:
[0698] The terminal sends the input information to the server. The terminal sends the information entered by the user to the server via the internet. The input data is sent to the server in text format.
[0699] Step 3:
[0700] The server analyzes the input information using a generative AI model. The server passes the received input information to the generative AI model (e.g., GPT-4) for analysis. The generative AI model generates appropriate answers or recommendations based on the input information. For example, given the information "I like spicy food," it generates a list of spicy dishes. The input data is in text format, and the output data is in list format.
[0701] Step 4:
[0702] The server generates appropriate answers and recommendations based on the analysis results. The server generates answers and recommendations for the user based on the analysis results obtained from the generating AI model. For example, it might generate a list of spicy dishes. The output data is in list format.
[0703] Step 5:
[0704] The server sends the generated responses and recommendations to the device. The server sends the generated responses and recommendations to the device via the internet. The output data is sent to the device in list format.
[0705] Step 6:
[0706] The device displays answers and recommendations to the user. The device displays answers and recommendations received from the server to the user. The user can review the displayed information. The output data is displayed in list format.
[0707] Step 7:
[0708] The device recognizes the user's emotions using an emotion engine. The device captures the user's facial expressions and voice using its built-in camera and microphone, and recognizes the user's emotions using an emotion engine (e.g., Affectiva SDK). Input data is in image or audio format, and output data is in emotion information format.
[0709] Step 8:
[0710] The device sends recognized emotion information to the server. The device sends recognized emotion information to the server via the internet. Input data is sent to the server in emotion information format, and output data is sent to the server in emotion information format.
[0711] Step 9:
[0712] The server generates appropriate services based on emotional information. The server uses a generative AI model to generate appropriate services based on the received emotional information. For example, if the user shows an angry expression, it will generate an apology. Input data is in emotional information format, and output data is in text format.
[0713] Step 10:
[0714] The server sends the generated service to the terminal. The server sends the generated service to the terminal via the internet. The output data is sent to the terminal in text format.
[0715] Step 11:
[0716] The terminal provides services to the user. The terminal provides services received from the server to the user. For example, it displays an apology. The output data is displayed in text format.
[0717] (Application Example 3)
[0718] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0719] Traditional food delivery services faced challenges such as users spending a lot of time choosing from menus and difficulty in providing services tailored to user preferences and emotions. Furthermore, multilingual support for foreign tourists was insufficient, highlighting the need for improved user experience. There is a demand to solve these problems and provide a more comfortable and personalized service.
[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[0721] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and means for recognizing the user's emotions and providing services corresponding to those emotions. As a result, users can efficiently select menus and receive personalized services tailored to their individual preferences and emotions. Furthermore, multilingual support enhances services for foreign tourists.
[0722] A "smart device" is a portable electronic device with advanced computing capabilities, such as a smartphone or tablet.
[0723] A "menu creation method" refers to a method or device for generating a list of products or services that a user can select.
[0724] "An artificial intelligence method for generating product descriptions from images" refers to an artificial intelligence system that uses image recognition technology to automatically analyze the characteristics and features of a product and generates a product description based on that analysis.
[0725] "Means of generating translation work using artificial intelligence" refers to artificial intelligence systems that automatically translate text and audio between different languages.
[0726] "Customer experience improvement measures" refer to methods and devices that enhance the convenience and satisfaction users experience when using a service.
[0727] "A means for artificial intelligence to explain questions when considering menu options" refers to a system in which artificial intelligence provides appropriate answers to questions that arise when users are choosing from a menu.
[0728] A "method for recommending products based on preferences" is a system that recommends appropriate products and services based on the user's past choices and input information.
[0729] An "emotion recognition system" is a system that analyzes a user's emotions from their facial expressions, tone of voice, etc., and responds accordingly.
[0730] A system for carrying out this invention includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and emotion recognition means for recognizing the user's emotions and providing services in accordance with those emotions.
[0731] System program
[0732] The program in this system performs the following operations:
[0733] Hardware and software
[0734] Hardware:
[0735] Smart devices (smartphones, tablets, etc.)
[0736] Camera (to capture the user's facial expressions)
[0737] Computer (a processing unit for executing programs)
[0738] software:
[0739] OpenCV (image processing library)
[0740] Keras (deep learning library)
[0741] Transformers (a library of generative AI models)
[0742] Data processing and data calculation
[0743] 1. Menu creation method:
[0744] The server generates a list of products and services that users can select using their smart devices.
[0745] 2. Artificial intelligence means for generating product descriptions from images:
[0746] The server analyzes images captured by the smart device's camera and automatically recognizes the product's characteristics and features.
[0747] Based on the recognized information, a product description is generated.
[0748] 3. Means for artificial intelligence to generate translation work:
[0749] The server automatically translates text and audio between different languages.
[0750] In particular, we will provide multilingual support to enhance services for foreign tourists.
[0751] 4. Means of improving customer experience:
[0752] The server reduces waiting times for users when accessing the service, thereby improving convenience.
[0753] 5. How artificial intelligence can explain questions that arise when considering menu options:
[0754] The server uses a generative AI model to provide appropriate answers to questions that arise when users select menu items.
[0755] 6. Methods for recommending products based on preferences:
[0756] The server recommends appropriate products and services based on the user's past choices and input information.
[0757] 7. Emotion recognition means:
[0758] The server analyzes the user's facial expressions and voice tone captured by the camera to recognize their emotions.
[0759] Respond appropriately to the recognized emotions.
[0760] Specific example
[0761] When a user enters "I like spicy food," the server uses a generative AI model to recommend spicy dishes.
[0762] When a user asks, "What are the ingredients in this dish?", the server uses a generative AI model to provide information about the ingredients.
[0763] If the user shows an angry expression, the server uses emotion recognition to apologize with a message like, "I'm sorry. Is there a problem?"
[0764] Example of a prompt
[0765] "I like spicy food. What do you recommend?"
[0766] "What are the ingredients in this dish?"
[0767] "How do you respond if a user is showing signs of anger?"
[0768] In this way, users can efficiently select from the menu and receive personalized service tailored to their individual preferences and feelings. Furthermore, multilingual support enhances services for foreign tourists.
[0769] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[0770] Step 1:
[0771] The user opens the menu using a smart device.
[0772] Input: A request to display a menu initiated by the user.
[0773] Data processing: The server retrieves menu information from the database and sends it to the smart device.
[0774] Output: A menu is displayed on the smart device.
[0775] Step 2:
[0776] The user takes a picture of the food with the camera on their smart device.
[0777] Input: An image of a dish taken by the user.
[0778] Data processing: The server receives images and analyzes them using an image recognition algorithm (OpenCV).
[0779] Output: The analysis results extract the characteristics and features of the dishes.
[0780] Step 3:
[0781] The server generates a product description from the image.
[0782] Input: Data on the characteristics and features of the dish.
[0783] Data processing: The server uses a generative AI model to generate product descriptions based on characteristics and features.
[0784] Output: Product description text is generated and sent to the smart device.
[0785] Step 4:
[0786] Users enter questions when considering menu options.
[0787] Input: A question entered by the user (e.g., "What are the ingredients in this dish?").
[0788] Data processing: The server generates answers to questions using a generative AI model.
[0789] Output: The answer text is generated and displayed on the smart device.
[0790] Step 5:
[0791] The server recommends products based on the user's preferences.
[0792] Input: User's past selections and input information (e.g., "I like spicy food").
[0793] Data processing: The server uses a generative AI model to recommend products based on user preferences.
[0794] Output: A list of recommended products is displayed on the smart device.
[0795] Step 6:
[0796] The user captures their facial expressions using the camera on their smart device.
[0797] Input: User's facial expression image.
[0798] Data processing: The server analyzes emotions using a facial recognition algorithm (Keras).
[0799] Output: As an analysis result, user sentiment data is generated.
[0800] Step 7:
[0801] The server provides services that respond to the user's emotions.
[0802] Input: User emotion data (e.g., anger).
[0803] Data processing: The server generates appropriate responses based on sentiment data (e.g., apology messages).
[0804] Output: The corresponding message is displayed on the smart device.
[0805] Step 8:
[0806] The server performs the translation work.
[0807] Input: Text or audio data entered by the user.
[0808] Data processing: The server uses a translation algorithm to translate the input data into multiple languages.
[0809] Output: Translated text and audio data are displayed on the smart device.
[0810] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0811] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0812] Other examples of generative AI include Gemini® (registered trademark) (Internet search). <url: https: gemini.google.com ?hl="ja">) are some examples.
[0813] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0814] [Second Embodiment]
[0815] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0816] As shown in Figure 3, the data processing system 210 includes the data processing device 12 and the smart eye
[0817] It is equipped with a mirror 214. An example of a data processing device 12 is a server.
[0818] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0819] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0820] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0821] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0822] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0823] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0824] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0825] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0826] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is performed by the processor 46 executing the reception output on the RAM 48.
[0827] This is achieved by operating as a control unit 46A according to the power program 60.
[0828] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0829] "Example of form 1"
[0830] In one embodiment of the present invention, a restaurant operator creates a menu using a smartphone. Specifically, the operator uploads photos of products taken with the smartphone's camera, and AI automatically recognizes the characteristics and features of the products from the photos. Based on the recognition results, a product description is generated. This product description is then displayed to customers on the restaurant's website or application.
[0831] "Example of form 2"
[0832] Furthermore, in this embodiment of the present invention, the AI generates the translation work. Specifically, the generated product description is translated into multiple languages. This makes the product description easier for foreign tourists to understand. For example, it is possible to translate into major tourist languages such as English, Chinese, and Korean.
[0833] "Example of form 3"
[0834] Furthermore, in this embodiment of the present invention, the customer experience is also improved. Specifically, when a customer is considering a menu, the AI explains their questions and recommends products according to their preferences. For example, if a customer inputs information such as "I like spicy food," the AI will recommend spicy dishes. Also, if a customer asks a question such as "What are the ingredients in this dish?", the AI will provide an answer to that question.
[0835] The following describes the processing flow for each example of the form.
[0836] "Example of form 1"
[0837] Step 1: The restaurant operator takes photos of the products using their smartphone camera.
[0838] Step 2: Upload the photos you've taken to the system.
[0839] Step 3: The AI within the system automatically recognizes the product's characteristics and features from the photograph.
[0840] Step 4: The AI generates a product description based on the recognition results.
[0841] Step 5: The generated product description is displayed to customers on the restaurant's website or application.
[0842] "Example of form 2"
[0843] Step 1: Obtain the product description generated by the AI.
[0844] Step 2: The AI translates the product description into multiple languages.
[0845] Step 3: The translated product description is displayed to foreign tourists on the restaurant's website or application.
[0846] "Example of form 3"
[0847] Step 1: When customers are considering the menu, they input their questions into the AI.
[0848] Step 2: The AI generates the answer to that question.
[0849] Step 3: The AI recommends products based on the customer's preferences.
[0850] Step 4: The AI-generated answers and recommendations are displayed to the customer.
[0851] (Example 1)
[0852] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0853] When restaurant operators create menus, the process of taking photos of products and then recognizing their characteristics and features to generate product descriptions is time-consuming. Furthermore, there is a lack of multilingual support for foreign tourists and other means to enhance the customer experience. Therefore, there is a need for increased efficiency in menu creation and improved customer satisfaction.
[0854] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0855] In this invention, the server includes means for creating menus using a smartphone, means for artificial intelligence to generate product descriptions from photos taken with a smartphone, means for artificial intelligence to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for uploading product photos taken with a smartphone camera and for artificial intelligence to automatically recognize the characteristics and features of the products from those photos, and means for generating product descriptions based on the recognition results. This makes it possible to improve the efficiency of menu creation and enhance customer satisfaction.
[0856] A "smartphone" is a multi-functional mobile device that, in addition to the functions of a mobile phone, is capable of internet connectivity and the use of applications.
[0857] "Menu creation methods" refer to the methods and tools that restaurant operators use to create menus, and especially those that utilize smartphones.
[0858] "Artificial intelligence methods" refer to technologies that use machine learning and data analysis to automatically perform specific tasks.
[0859] "Means of generating translation work using artificial intelligence" refers to methods and technologies that use artificial intelligence to translate text into multiple languages.
[0860] "Methods for improving the customer experience" refer to methods and tools for improving the convenience and satisfaction customers experience when using a service.
[0861] "Methods for AI to explain questions when considering menus" refers to methods and technologies in which artificial intelligence automatically provides answers to questions that customers may have when considering menus.
[0862] "Methods of recommending based on preferences" refer to methods and technologies that recommend appropriate products and services based on a customer's preferences and past choices.
[0863] "Means for automatically recognizing the characteristics and features of a product" refers to methods and technologies that use artificial intelligence to automatically extract the characteristics and features of a product from photographs or data.
[0864] "Means for generating product descriptions" refers to methods and technologies for creating product descriptions in natural language based on recognized product characteristics and features.
[0865] This invention is a system that allows restaurant operators to create menus using their smartphones. Specifically, users upload photos of products taken with their smartphone cameras, and artificial intelligence (AI) automatically recognizes the characteristics and features of the products from the photos. Based on the recognition results, the system generates product descriptions. These product descriptions are then displayed to customers on the restaurant's website or application.
[0866] Hardware and software to be used
[0867] Smartphone: A device used for taking and uploading photos.
[0868] Server: Performs data processing and runs AI models.
[0869] Artificial intelligence models: Image recognition models and natural language generation models based on TensorFlow and PyTorch (e.g., GPT-3, BERT).
[0870] Data processing and data calculation
[0871] 1. The user takes a photo of the product with their smartphone.
[0872] The user launches their smartphone's camera app and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[0873] 2. The user uploads photos to the system.
[0874] The user opens a dedicated application, selects the photos they have taken, and presses the upload button. The photos are then sent to the server via the internet.
[0875] 3. The server receives the photos and inputs them into the AI model.
[0876] The server receives the uploaded photos and inputs them into an AI model for image processing. The AI model used here is based on TensorFlow or PyTorch.
[0877] 4. The server uses an AI model to recognize the characteristics and features of the product.
[0878] The server uses an AI model to recognize the characteristics and features of a product from a photograph. For example, it extracts the type of dessert, main ingredients, and visual features.
[0879] 5. The server generates a product description based on the recognition results.
[0880] The server generates product descriptions based on the AI's recognition results. These product descriptions are created using natural language generation technology. The software used includes GPT-3 and BERT.
[0881] 6. The server displays the product description on the website or application.
[0882] The server displays the generated product descriptions on the restaurant's website or application. Customers can then view them.
[0883] Specific example
[0884] A user takes a photo of a new dessert, "Chocolate Cake," and uploads it to the system. The server receives the photo and uses an AI model to recognize features such as "Chocolate Cake," "Cream Topping," and "Berry Decoration." The server uses GPT-3 to generate a product description such as, "This chocolate cake features rich chocolate and creamy toppings. The berry decoration makes it visually appealing," and displays it on the website.
[0885] Example of a prompt
[0886] Please upload a photo of your new dessert. AI will recognize the product's characteristics and features from the photo and generate a product description.
[0887] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0888] Step 1:
[0889] The user takes a photo of the product with their smartphone.
[0890] Input: Product (e.g., Chocolate cake)
[0891] Output: Photos of the photographed product
[0892] Specific action: The user launches the camera app on their smartphone and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[0893] Step 2:
[0894] The user uploads a photo to the system.
[0895] Input: Photos of the product
[0896] Output: Photo data sent to the server
[0897] Specific operation: The user opens a dedicated application, selects the photo they have taken, and presses the upload button. The photo is sent to the server via the internet.
[0898] Step 3:
[0899] The server receives the photos and inputs them into the AI model.
[0900] Input: Photo data sent to the server
[0901] Output: Photo data input to the AI model
[0902] Specific operation: The server receives an HTTP request, temporarily stores the photo data, and inputs it into a TensorFlow or PyTorch model.
[0903] Step 4:
[0904] The server uses an AI model to recognize the characteristics and features of the product.
[0905] Input: Photo data entered into the AI model
[0906] Output: Characteristics and features of the recognized product (e.g., chocolate cake, cream topping, berry decoration)
[0907] Specific operation: The server runs an AI model to extract product characteristics and features from a photograph. For example, it recognizes the type of dessert, main ingredients, and visual features.
[0908] Step 5:
[0909] The server generates a product description based on the recognition results.
[0910] Input: Characteristics and features of the recognized product
[0911] Output: Generated product description (Example: "This chocolate cake features rich chocolate and creamy toppings. The berry decorations make it visually appealing.")
[0912] Specific operation: The server uses GPT-3 or BERT to generate product descriptions based on the recognition results as input.
[0913] Step 6:
[0914] The server displays product descriptions on websites and applications.
[0915] Input: Generated product description
[0916] Output: Product description displayed on the website or application
[0917] Specific operation: The server saves the product description in the website's database and sends the data to the front-end for display. Customers can then view this data.
[0918] (Application Example 1)
[0919] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0920] Traditional restaurant menu creation was often done manually, which was time-consuming and labor-intensive, and made it difficult to adequately convey the appeal of the products. Furthermore, the lack of multilingual support for foreign tourists and insufficient recommendation features tailored to customer preferences were also problems. In addition, there were limited means to improve the in-store customer experience, such as having to wait for waitstaff or in line at the register, highlighting the need for increased customer satisfaction.
[0921] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0922] In this invention, the server includes means for creating menus using a smartphone, means for generating product descriptions from photos taken with the smartphone, means for the AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for the AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on the recognized features using a natural language generation model. This enables efficient menu creation and automatic generation of attractive product descriptions, as well as multilingual support, improved customer experience, and personalized recommendations.
[0923] "A method for creating menus using a smartphone" refers to a method for efficiently creating restaurant menus using the camera and applications of a smartphone.
[0924] "AI methods for generating product descriptions from smartphone photos" refers to artificial intelligence methods that analyze product photos taken with a smartphone, recognize their characteristics and features, and automatically generate product descriptions.
[0925] "Methods for AI to generate translation work" refers to methods of using artificial intelligence to translate product descriptions and menu contents into multiple languages.
[0926] "Methods to improve the customer experience by eliminating the need to wait for waitstaff or in line at the register" refers to methods that allow customers to receive service smoothly without having to wait for waitstaff or in line at the register.
[0927] "An AI-powered solution for questions during menu consideration" refers to a method in which artificial intelligence automatically provides answers to questions that customers may have when considering menu options.
[0928] "Methods of recommending based on preferences" refer to methods of recommending appropriate products or menus based on a customer's past choices and preferences.
[0929] "Methods for recognizing the characteristics and features of a product using an image recognition model" refers to methods for automatically recognizing the characteristics and features of a product from a photograph using image recognition technology.
[0930] "Means for generating product descriptions based on features recognized using a natural language generation model" refers to means for automatically generating attractive product descriptions using natural language generation technology based on recognized product characteristics and features.
[0931] A system for carrying out this invention includes means for creating menus using a smartphone, means for generating product descriptions from photos on a smartphone, means for AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the cash register, means for AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on recognized features using a natural language generation model.
[0932] System program
[0933] The server implements the system using the following hardware and software.
[0934] Hardware:
[0935] Smartphone (with camera)
[0936] Server (for hosting AI models)
[0937] software:
[0938] TensorFlow: A library for image recognition
[0939] OpenAI GPT-3: API for Natural Language Generation
[0940] Explanation of the process
[0941] Image recognition:
[0942] Users take photos of their food with their smartphone cameras and upload them to a server via an application. The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of the food from the image. For example, features such as "chicken curry, spicy, tomato-based" might be recognized.
[0943] Natural language generation:
[0944] The server converts the recognized features into a string and sends it to OpenAI GPT-3 as a prompt. An example of a prompt is, "Generate an appealing product description for a dish with the following features: Chicken curry, spicy, tomato-based." GPT-3 generates an appealing product description based on this prompt. For example, it might generate a product description such as, "This spicy chicken curry features juicy chicken simmered in a tomato-based sauce. The aromatic spices will whet your appetite."
[0945] Translation work:
[0946] The generated product descriptions are translated into multiple languages as needed. The server uses AI to perform translations in multiple languages, such as English and Chinese.
[0947] Improving the customer experience:
[0948] Customers can use their smartphones to select menu items and place orders in-store. This reduces waiting times at the counter and cashier, enabling smoother service.
[0949] Explanation of the question and recommendations:
[0950] The AI automatically provides answers to questions that customers may have when considering menu options. It also recommends appropriate products and menu items based on the customer's past choices and preferences.
[0951] In this way, menu creation becomes more efficient, attractive product descriptions are automatically generated, and multilingual support, improved customer experience, and personalized recommendations become possible.
[0952] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0953] Step 1:
[0954] The user takes a photo of the food with their smartphone camera and uploads it to the server through the application. The input is the photo of the food, and the output is the image data sent to the server.
[0955] Step 2:
[0956] The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of dishes from uploaded image data. The input is image data, and the output is a list of the recognized characteristics and features. Specifically, the image data is preprocessed and then input into the model to obtain prediction results.
[0957] Step 3:
[0958] The server converts the recognized traits and features into a string and sends it to OpenAI GPT-3 as a prompt. The input is a list of traits and features, and the output is a prompt statement. Specifically, it formats the traits and features and generates a prompt statement.
[0959] Step 4:
[0960] The server receives a response from GPT-3 and generates an attractive product description. The input is a prompt, and the output is the generated product description. Specifically, it calls the GPT-3 API, parses the response, and retrieves the product description.
[0961] Step 5:
[0962] The server translates the generated product description into multiple languages as needed. The input is the product description, and the output is the translated product description. Specifically, it calls a translation API to translate into multiple languages.
[0963] Step 6:
[0964] Users use their smartphones to select menu items and place orders in the store. The input is a generated product description, and the output is the user's order information. Specifically, the application displays product descriptions, and the user selects their order.
[0965] Step 7:
[0966] The server automatically provides answers to questions that arise when users consider the menu, using AI. The input is the user's question, and the output is the AI's answer. Specifically, it analyzes the question and generates an appropriate answer.
[0967] Step 8:
[0968] The server recommends appropriate products and menu items based on the user's past choices and preferences. The input is the user's past selection data, and the output is the recommended products and menu items. Specifically, it analyzes past selection data and uses a recommendation algorithm to suggest products.
[0969] (Example 2)
[0970] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0971] Traditional systems often required manual translation of product descriptions into multiple languages, resulting in time-consuming and labor-intensive processes. Furthermore, services for foreign tourists were inadequate, making it difficult for them to understand product descriptions. Additionally, insufficient efforts were made to improve the customer experience and address questions during menu selection, potentially leading to decreased customer satisfaction.
[0972] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0973] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for the artificial intelligence to generate translation tasks, means for improving the customer experience, means for the artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into the generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results. As a result, multilingual translation of product descriptions is automated, improving services for foreign tourists, enhancing the customer experience, and resolving questions when considering menus.
[0974] A "smart device" is a portable electronic device, such as a smartphone or tablet, that possesses advanced computing power and communication capabilities.
[0975] "Menu creation methods" refer to functions and applications that use smart devices to create menus for restaurants and retail stores.
[0976] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[0977] "A means of generating translation work using artificial intelligence" refers to a function that uses artificial intelligence to translate text data into other languages.
[0978] "Means of improving the customer experience" refer to functions and methods that enhance the convenience and satisfaction customers experience when using products or services.
[0979] "A means for artificial intelligence to explain questions when considering menus" refers to a function in which artificial intelligence automatically provides answers to questions and concerns that arise when customers are considering menus.
[0980] "Means of recommending based on preferences" refers to a function that recommends appropriate products and services based on the customer's past choices and preferences.
[0981] "Means for translating product descriptions into multiple languages" refers to a function for translating product descriptions into multiple languages.
[0982] A "generative AI model" is an artificial intelligence model trained to perform tasks such as text generation and translation.
[0983] A "prompt statement" is an instruction given to a generative AI model to perform a specific task.
[0984] "Means for saving translation results to a database" refers to the function of saving translation results generated by a generative AI model to a database.
[0985] "Means for sending translation results to the terminal" refers to a function that sends translation results stored in the database to the user's terminal.
[0986] "Means by which the terminal displays the translation results" refers to a function in which the user's terminal displays the received translation results on the screen.
[0987] This invention is a system that includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into a generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results.
[0988] Hardware and software to be used
[0989] Hardware:
[0990] Smart devices (smartphones, tablets, etc.)
[0991] Server (a computer with high-performance computing capabilities)
[0992] software:
[0993] Generative AI models (e.g., OpenAI's GPT-4)
[0994] Database Management System
[0995] Communication protocol (e.g., HTTP / HTTPS)
[0996] Data processing and data calculation
[0997] Product description generation:
[0998] The user takes a picture of a product using a smart device and sends the image to a server. The server uses artificial intelligence to generate a product description from the image. Specifically, it uses an image analysis algorithm to recognize the characteristics and features of the product and generates a product description based on that.
[0999] Prompt message generation:
[1000] The server generates prompt messages using a generative AI model. The prompt message specifies the target language for translation (e.g., English, Chinese, Korean).
[1001] Perform the translation task:
[1002] The server inputs the generated prompt text and product description into the AI model and retrieves the translation results. The AI model then performs multilingual translation based on the input text.
[1003] Saving and sending translation results:
[1004] The server saves the acquired translation results to a database. The saved translation results are sent to the user's terminal, which then displays the translation results.
[1005] Specific example
[1006] Specific example:
[1007] The user enters the Japanese product description: "This product uses high-quality materials."
[1008] The server generates the prompt message: "Translate the following product description into English: This product uses high-quality materials."
[1009] The server inputs the prompt text and product description into the generated AI model and retrieves the English translation result, "This product uses high-quality materials."
[1010] The server saves the translation results to a database and sends them to the terminal.
[1011] The device displays the translation result, and the user confirms it.
[1012] In this way, the server, terminal, and user work together to achieve multilingual translation of product descriptions.
[1013] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1014] Step 1:
[1015] The user enters the product description.
[1016] The user enters the product description in text format using a smart device. The entered product description is then sent from the smart device to the server.
[1017] Input: Product description text
[1018] Output: Product description text sent to the server
[1019] Step 2:
[1020] The server receives the product description.
[1021] The server receives the product description text sent by the user and temporarily stores it in memory.
[1022] Input: Product description text submitted by the user
[1023] Output: Product description text stored in memory
[1024] Step 3:
[1025] The server generates the prompt message.
[1026] The server generates prompt messages based on the languages to be translated. For example, if translation is needed for English, Chinese, and Korean, it will generate prompt messages corresponding to each language.
[1027] Input: Product description text, target language for translation
[1028] Output: Generated prompt message
[1029] Step 4:
[1030] The server inputs prompt text and product description into the generated AI model.
[1031] The server inputs the generated prompt text and product description text into the AI model. The AI model then performs translation based on the input text.
[1032] Input: Prompt text, product description text
[1033] Output: Data input to the generative AI model
[1034] Step 5:
[1035] The server retrieves the translation results.
[1036] The server retrieves translation results from the generative AI model. The retrieved translation results are separated by language.
[1037] Input: Data entered into the generating AI model
[1038] Output: Translated text
[1039] Step 6:
[1040] The server saves the translation results to the database.
[1041] The server saves the retrieved translation results to a database. When saving, it associates the original Japanese product description with the translated result.
[1042] Input: Translated text, original product description text
[1043] Output: Translation results stored in the database
[1044] Step 7:
[1045] The server sends the translation result to the terminal.
[1046] The server sends the saved translation results to the user's device. The transmitted data is provided in a format that the user can access.
[1047] Input: Translation results stored in the database
[1048] Output: Translation results sent to the terminal
[1049] Step 8:
[1050] The device displays the translation result.
[1051] The terminal receives the translation results sent from the server and displays them to the user. The user can review the displayed translation results and make corrections or re-translates as needed.
[1052] Input: Translation result sent from the server
[1053] Output: Translation results displayed on the terminal
[1054] (Application Example 2)
[1055] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1056] In modern brick-and-mortar stores, foreign tourists face language barriers when purchasing goods, making it difficult for them to understand product descriptions. Furthermore, visually impaired individuals and those with reading and writing difficulties also have limited means of understanding product descriptions. This can lead to a diminished customer experience and potentially impact sales.
[1057] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for creating a menu using a smartphone, an artificial intelligence means for generating product descriptions from photos on the smartphone, a means for the artificial intelligence to generate translation work, a means for improving the customer experience, a means for the artificial intelligence to explain questions when considering menus, a means for making recommendations according to preferences, and a means for translating product descriptions into multiple languages and displaying and playing them aloud in the user's set language. This makes it possible to make product descriptions easier to understand for foreign tourists, visually impaired people, and people who have difficulty reading and writing.
[1058] "A method for creating menus using a smartphone" refers to a method for creating menus using a smartphone.
[1059] "An artificial intelligence method for generating product descriptions from smartphone photos" refers to a method that uses artificial intelligence technology to automatically generate product descriptions from photos taken with a smartphone.
[1060] "Methods for generating translation work using artificial intelligence" refers to methods of translating text into multiple languages using artificial intelligence.
[1061] "Methods for improving the customer experience" refer to measures taken to improve the customer's experience when purchasing a product.
[1062] "A means for artificial intelligence to explain questions that arise when customers are considering menu options" refers to a method in which artificial intelligence answers questions that customers may have when considering menu options.
[1063] "Methods of recommending based on preferences" refer to methods of recommending products and services based on customer preferences.
[1064] "Means for translating product descriptions into multiple languages and displaying and playing them in the user's preferred language" refers to means for translating product descriptions into multiple languages and displaying and playing them in the language specified by the user.
[1065] To implement this invention, it is necessary to build a system that combines a smartphone, artificial intelligence technology, a translation API, and a voice playback function.
[1066] First, smartphones have a camera function that allows users to take pictures of products. When a user takes a picture of a product, the image data is sent to a server. The server generates a product description using image recognition technology. This image recognition technology could include, for example, the Google Cloud Vision API.
[1067] Next, the generated product description is translated into multiple languages using artificial intelligence on the server. Translation services such as the Google Cloud Translation API are used for this translation. The translated text is sent to the smartphone and displayed based on the user's language settings.
[1068] Furthermore, the translated product descriptions are played back using the smartphone's audio playback function. This makes the product accessible to visually impaired individuals and those who have difficulty reading or writing.
[1069] As a concrete example, consider a tourist trying to purchase a leather wallet at a physical store in Japan. When the tourist scans the leather wallet's tag with their smartphone camera, the app automatically recognizes the product description and translates it into the user's chosen language (for example, English). The translated product description is displayed as "This is a high-quality leather wallet," and is also played aloud.
[1070] Examples of prompt messages include the following:
[1071] I am developing an application to translate product descriptions into multiple languages. Please translate the following text into the specified languages.
[1072] Text: "This is a high-quality leather wallet."
[1073] Target language: Japanese
[1074] In this way, foreign tourists, visually impaired individuals, and people with reading and writing difficulties will be able to understand product descriptions without encountering language barriers when purchasing goods in physical stores.
[1075] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1076] Step 1:
[1077] The user takes a picture of the product using their smartphone camera. The input is the product image taken with the smartphone camera. The output is the captured product image data.
[1078] Step 2:
[1079] The device sends the captured product image data to the server. The input is the product image data. The output is the product image data sent to the server.
[1080] Step 3:
[1081] The server generates product descriptions using image recognition technology. The input is product image data. The server extracts text information from the images using APIs such as Google Cloud Vision API and generates the product description. The output is the generated product description text.
[1082] Step 4:
[1083] The server translates the generated product description text into multiple languages. The input is the product description text. The server uses the Google Cloud Translation API to translate it into the specified languages. The output is the translated product description text.
[1084] Step 5:
[1085] The server sends the translated product description text to the terminal. The input is the translated product description text. The output is the translated product description text sent to the terminal.
[1086] Step 6:
[1087] The device displays the translated product description text. The input is the translated product description text. The output is the translated product description displayed on the smartphone screen.
[1088] Step 7:
[1089] The device plays the translated product description text as audio. The input is the translated product description text. The device uses the smartphone's audio playback function to convert the text to audio and play it. The output is the translated product description played as audio.
[1090] (Example 3)
[1091] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1092] Traditional menu creation and customer experience improvement systems fail to adequately answer users' questions when considering menu items and to provide product recommendations tailored to individual preferences. Furthermore, they are insufficient in reducing waiting times and providing multilingual services. This can lead to decreased customer satisfaction and negatively impact store sales.
[1093] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[1094] In this invention, the server includes means for the user to input information, means for transmitting the input information to the server, means for the server to analyze the input information using a generated AI model, means for generating recommendations and answers based on the analysis results, means for transmitting the generated recommendations and answers to a terminal, and means for the terminal to display the recommendations and answers to the user. This enables appropriate answers to questions the user has when considering a menu and product recommendations tailored to individual preferences. It also enables reduced waiting times and the provision of multilingual services, leading to improved customer satisfaction and increased store sales.
[1095] A "smart device" is an electronic device used by users to input information and communicate with a server, and includes smartphones and tablets.
[1096] "Artificial intelligence tools" refer to algorithms and models that analyze input data and generate appropriate responses or recommendations.
[1097] A "generative AI model" refers to a machine learning model that analyzes user input information and generates appropriate responses and recommendations.
[1098] "Methods for improving the customer experience" refer to measures that enable customers to reduce waiting times and use services comfortably.
[1099] "A means for artificial intelligence to explain questions when considering menu options" refers to a method by which artificial intelligence provides appropriate answers to questions that users may have when considering menu options.
[1100] "Methods for recommending products according to preferences" refers to methods for recommending appropriate products based on the user's preferences.
[1101] "Means for users to input information" refers to an interface that allows users to input information in formats such as text or voice.
[1102] "Means for sending entered information to the server" refers to the means of communication used to send information entered by the user to the server.
[1103] "Means by which a server analyzes input information using a generated AI model" refers to means by which a server uses a generated AI model to analyze information received from a user.
[1104] "Means for generating recommendations and responses based on analysis results" refers to means of generating recommendations and responses for users based on the analysis results of a generation AI model.
[1105] "Means for sending generated recommendations and responses to the device" refers to the communication means used by the server to send generated recommendations and responses to the user's device.
[1106] "Means by which a device displays recommendations and answers to a user" refers to the means by which a user's device displays recommendations and answers received from a server to the user.
[1107] This invention is a system in which a user inputs information using a smart device, a server analyzes that information using a generated AI model, and then generates and provides appropriate recommendations and answers to the user. Specific embodiments of this system are described below.
[1108] First, the user inputs information into the system using a smart device such as a smartphone or tablet. For example, the user might input "I like spicy food." This information is entered through the smart device's interface.
[1109] Next, the smart device sends the entered information to the server. In this process, the smart device converts the input data into an appropriate format (e.g., JSON format) and sends it to the server via the network.
[1110] The server analyzes the received user input information using a generation AI model (for example, OpenAI's GPT-4). Specifically, it analyzes information such as "I like spicy food" and searches a database related to spicy food.
[1111] The server recommends products that match the user's preferences based on the analysis results. For example, it generates a list of spicy dishes and presents it to the user. Also, if the user asks, "What are the ingredients in this dish?", the server generates an answer to that question.
[1112] The generated recommendations and responses are sent from the server to the smart device. The server converts the data into the appropriate format and sends it to the smart device over the network.
[1113] Smart devices display recommendations and answers received from a server to the user. For example, they might display, "Recommended spicy dishes are Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot."
[1114] As a concrete example, the following prompt statement is shown.
[1115] "I like spicy food. Can you recommend some dishes?"
[1116] "What are the ingredients in this dish?"
[1117] In this way, a system is realized in which the server, smart device, and user work together to improve the customer experience. Through this system, users can receive appropriate answers to questions when considering menus and product recommendations tailored to their individual preferences. It can also reduce waiting times and provide multilingual services, leading to improved customer satisfaction and increased store sales. The flow of specific processing in Example 3 will be explained using Figure 15.
[1118] Step 1:
[1119] The user enters information using a terminal.
[1120] The user opens the application on their smartphone or tablet, types "I like spicy food" into the text box, and presses the submit button. The input data is in text format and represents information about the user's preferences.
[1121] Step 2:
[1122] The terminal sends the entered information to the server.
[1123] The terminal converts the user's input, "I like spicy food," into JSON format and sends it to the server using the HTTPS protocol. The input data is user preference information in text format, and the output data is request data in JSON format.
[1124] Step 3:
[1125] The server analyzes the input information using a generated AI model.
[1126] The server parses the received JSON data and inputs the information "likes spicy food" into a generative AI model. The generative AI model (e.g., GPT-4) searches a database of spicy dishes and generates a list of related dishes. The input data is request data in JSON format, and the output data is a list of dishes as a result of the analysis.
[1127] Step 4:
[1128] The server generates recommendations and responses based on the analysis results.
[1129] The server generates a list of spicy dishes based on the analysis results obtained from the generated AI model. For example, it might list dish names such as "Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot." The input data is the list of dishes resulting from the analysis, and the output data is a recommendation list presented to the user.
[1130] Step 5:
[1131] The server sends recommendations and responses it generates to the device.
[1132] The server converts the generated list of spicy dishes into JSON format and sends it to the terminal using the HTTPS protocol. The input data is a recommendation list, and the output data is a response in JSON format.
[1133] Step 6:
[1134] The device displays recommendations and answers to the user.
[1135] The terminal parses the received JSON data and displays to the user, "Recommended spicy dishes are Mapo Tofu, Kimchi Stew, and Sichuan-style Mala Hot Pot." The input data is response data in JSON format, and the output data is text information displayed to the user.
[1136] (Application Example 3)
[1137] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1138] Traditional food delivery systems have problems such as difficulty for customers to obtain detailed information when choosing a menu and a lack of personalized recommendations. Furthermore, the lack of a means to provide real-time information on ingredients and nutritional content has been a factor in lowering customer satisfaction.
[1139] In Application Example 3, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a menu creation means using a smart device, an artificial intelligence means for generating product descriptions from images on the smart device, a means for the artificial intelligence to generate translation work, a means for improving the customer experience by shortening waiting times, a means for the artificial intelligence to explain questions when considering menus, a means for making recommendations according to preferences, a means for recommending dishes based on customer preferences, and a means for providing information on the ingredients of dishes. As a result, customers can obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[1140] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone or tablet.
[1141] A "menu creation method" is a function that generates a list of dishes and drinks that customers can choose from.
[1142] "An artificial intelligence method for generating product descriptions from images" refers to a function that analyzes images taken with a smart device and automatically generates product descriptions based on their content.
[1143] "Means of generating translation work using artificial intelligence" refers to a function that automatically translates text between different languages.
[1144] "Methods to improve the customer experience by reducing waiting times" refer to functions that reduce the waiting time customers have when using a service.
[1145] "A means for artificial intelligence to explain questions when considering menus" refers to a function in which artificial intelligence provides appropriate answers to questions that customers may have when choosing from a menu.
[1146] "Methods for recommending based on preferences" refer to functions that recommend appropriate products and services based on the customer's preferences and past choices.
[1147] "A method for recommending dishes based on customer preferences" refers to a function that recommends appropriate dishes based on the preference information entered by the customer.
[1148] "Means of providing information on the ingredients of a dish" refers to a function that provides information about the ingredients and components contained in a particular dish.
[1149] A system for carrying out this invention includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images on a smart device, means for artificial intelligence to generate translation work, means for improving the customer experience by reducing waiting time, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recommending dishes based on customer preferences, and means for providing information on the ingredients of dishes.
[1150] System program
[1151] The system's program is implemented using Python and the OpenAI API. The server receives customer input and uses a generative AI model to provide appropriate recommendations and information.
[1152] Explanation of the process
[1153] The server receives customer preferences and questions sent from smart devices and generates prompts based on them. These generated prompts are input into a generation AI model via the OpenAI API, which then generates appropriate answers and recommendations. The generated information is then sent back to the smart device and provided to the customer.
[1154] The hardware used will be smart devices such as smartphones and tablets. The software used will be Python and the OpenAI API.
[1155] Specific example
[1156] For example, if a customer enters "I like spicy food," the server will generate the following prompt:
[1157] "Please recommend some dishes for people who like spicy food."
[1158] When this prompt is entered into the OpenAI API, the generating AI model will recommend spicy dishes such as "Mapo Tofu" and "Kimchi Hot Pot."
[1159] Furthermore, if a customer asks, "What are the ingredients in Mapo Tofu?", the server will generate the following prompt:
[1160] "What are the ingredients in Mapo Tofu?"
[1161] When this prompt is entered into the OpenAI API, the generating AI model provides ingredient information such as "tofu, ground pork, green onions, garlic, ginger, chili bean paste, sweet bean paste, soy sauce, sake, sugar, and chicken broth."
[1162] This allows customers to obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[1163] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[1164] Step 1:
[1165] The user launches the application using a smart device and enters their preferences and answers questions. The entered information is sent to the server in text format. Examples of input include "I like spicy food" and "What are the ingredients in mapo tofu?".
[1166] Step 2:
[1167] The server analyzes the input information received from the user and generates an appropriate prompt. For example, in response to the input "I like spicy food," it generates the prompt "Please recommend some dishes for someone who likes spicy food." This prompt is then input into the generation AI model.
[1168] Step 3:
[1169] The server sends the generated prompt to the OpenAI API and retrieves appropriate answers and recommendations based on the generating AI model. For example, in response to the prompt "Please recommend some dishes for someone who likes spicy food," answers such as "Mapo Tofu" and "Kimchi Hot Pot" are generated.
[1170] Step 4:
[1171] The server analyzes the responses and recommendations obtained from the generating AI model and converts them into a format for the user. For example, it formats responses such as "Mapo Tofu" and "Kimchi Hot Pot" into a list.
[1172] Step 5:
[1173] The server sends formatted responses and recommendations to the smart device. The user can then view this information on the smart device's screen.
[1174] Step 6:
[1175] Users can select menu items based on the information provided or ask further detailed questions. For example, they might re-enter the question, "What are the ingredients in Mapo Tofu?"
[1176] Step 7:
[1177] The server receives input from the user again and repeats the same process. Specifically, it generates a prompt again, inputs it into the generation AI model, retrieves the answer, formats it, and provides it to the user.
[1178] This allows users to obtain detailed menu information in real time and receive recommendations tailored to their individual preferences.
[1179] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1180] "Example of form 1"
[1181] One embodiment of the present invention provides a system that combines an emotion engine. This system recognizes the user's emotions and recommends products according to those emotions. Specifically, when a user chooses a product, the system recognizes the user's emotions from their facial expressions and tone of voice, and then...
[1182] The system recommends products that match the user's preferences. For example, if a user is showing a joyful expression, it will recommend products that are likely to bring joy to that user.
[1183] "Example of form 2"
[1184] Furthermore, the emotion engine generates product descriptions based on the user's emotions. Specifically, it recognizes the user's emotions and selects words that match those emotions to generate a product description. For example, if the user is feeling down, it will select words that will cheer them up to generate a product description. (Example 3)
[1185] Furthermore, the emotion engine recognizes the user's emotions and provides services accordingly. Specifically, it recognizes the user's emotions and provides services that match those emotions. For example, if a user is showing anger, it will provide services such as offering an apology to that user.
[1186] The following describes the processing flow for each example of the form.
[1187] "Example of form 1"
[1188] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[1189] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[1190] Step 3: The system recommends products that match the recognized emotions.
[1191] "Example of form 2"
[1192] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[1193] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[1194] Step 3: The system selects words that match the recognized emotion and generates a product description.
[1195] "Example of form 3"
[1196] Step 1: When a user selects a product, the system captures the user's facial expressions and voice tone.
[1197] Step 2: The emotion engine recognizes the user's emotions from the captured information.
[1198] Step 3: The system provides services that match the recognized emotions.
[1199] (Example 1)
[1200] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1201] Traditional restaurant menu creation and customer experience improvement have been time-consuming and labor-intensive. In particular, generating product descriptions from photos and providing personalized recommendations consume significant human resources. Furthermore, multilingual services for foreign tourists are inadequate, highlighting the need for improved customer satisfaction. Additionally, the inability to provide product recommendations based on customer emotions can lead to a decline in the quality of the customer experience. An efficient system is needed to address these challenges.
[1202] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1203] In this invention, the server includes means for creating menus using a smartphone, means for artificial intelligence to generate product descriptions from photos taken on the smartphone, means for artificial intelligence to generate translation tasks, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the cash register, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, and means for recognizing the user's emotions and recommending products based on those emotions. This enables more efficient menu creation, product recommendations tailored to customer preferences, enhanced service for foreign tourists through multilingual support, and product recommendations based on customer emotions.
[1204] "A method for creating menus using a smartphone" refers to a method for creating restaurant menus using the camera and applications of a smartphone.
[1205] "An artificial intelligence method for generating product descriptions from smartphone photos" refers to a method that uses artificial intelligence technology to analyze product photos taken with a smartphone, recognize their characteristics and features, and automatically generate product descriptions.
[1206] "Methods for generating translation work using artificial intelligence" refers to methods of translating product descriptions and menu contents into multiple languages using artificial intelligence technology.
[1207] "Methods to improve the customer experience by eliminating the need to wait for waitstaff or in line at the register" refers to methods that allow customers to receive service smoothly without having to wait for waitstaff or in line at the register.
[1208] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a method in which artificial intelligence automatically provides answers to questions that customers may have when considering menu options.
[1209] "Methods of recommending based on preferences" refer to methods of recommending appropriate products based on a customer's past choices and preferences.
[1210] "A method for recognizing user emotions and recommending products based on those emotions" refers to a method of recognizing emotions by analyzing the user's facial expressions and tone of voice, and then recommending products that match those emotions.
[1211] This invention is a system that allows restaurant operators to create menus using smartphones and improve the customer experience. A specific embodiment of this system is described below.
[1212] First, the user takes a picture of the product using their smartphone camera. For example, to take a picture of a new dessert, they launch the smartphone's camera app, frame the dessert, and press the shutter button.
[1213] Next, the device uploads the captured photos to the server. Specifically, the smartphone application sends the photo data to the server via an HTTP request. At this time, the photo data is sent in an appropriate format (e.g., JPEG, PNG).
[1214] The server analyzes the received photos using image recognition software (e.g., Google Cloud Vision API). The server recognizes the characteristics of the product from the photo (e.g., chocolate cake, fruit parfait). For example, the Google Cloud Vision API might return the label "chocolate cake".
[1215] The server then generates a product description based on the recognized characteristics. Using a generation AI model (e.g., OpenAI GPT-3), it generates a product description such as, "A rich chocolate cake. It is not too sweet and has a moist texture."
[1216] The generated product descriptions are displayed on the restaurant's website or application by the server. Specifically, the product descriptions are sent to the front-end in HTML or JSON format and displayed in the user interface.
[1217] Furthermore, when a user selects a product, the device uses its camera and microphone to capture the user's facial expressions and voice tone. For example, the smartphone's camera captures the user's smile, and the microphone records the user's voice tone.
[1218] The server analyzes the captured data using an emotion engine (e.g., Microsoft Azure Emotion API) to recognize the user's emotions. Based on the recognized emotions, the server recommends products that are suitable for the user. For example, if the user shows a joyful expression, the server will recommend a "fruit parfait."
[1219] As a concrete example, consider a scenario where a user takes a photo of a new dessert with their smartphone and uploads it to the system. The server uses the Google Cloud Vision API to analyze the photo and recognizes its characteristic as "chocolate cake." The AI then generates a product description such as, "A rich chocolate cake. It's not too sweet and has a moist texture."
[1220] Furthermore, if a user displays an expression of joy while viewing the menu using their smartphone camera, the Microsoft Azure Emotion API recognizes that expression. The server then recommends a "fruit parfait" that the user is likely to enjoy.
[1221] Example of a prompt:
[1222] "Please upload photos of the product taken with your smartphone. Our AI will automatically recognize the product's characteristics and generate a product description. It will also use your camera and microphone to recommend products based on your emotions."
[1223] In this way, a system is realized in which servers, terminals, and users work together to efficiently create restaurant menus and recommend products to customers.
[1224] The flow of the specific processing in Example 1 will be explained using Figure 17.
[1225] Step 1:
[1226] The user takes a photo of the product.
[1227] The user takes a picture of the product using their smartphone camera. For example, to take a picture of a new dessert, the user launches the smartphone's camera app, frames the dessert, and presses the shutter button. The input is the photograph taken, and the output is the image data stored on the smartphone.
[1228] Step 2:
[1229] The device uploads the photo to the server.
[1230] The device uploads the captured photos to the server. Specifically, the smartphone application sends the photo data to the server via an HTTP request. The photo data is sent in an appropriate format (e.g., JPEG, PNG). The input is the image data on the smartphone, and the output is the image data sent to the server.
[1231] Step 3:
[1232] The server analyzes the photos and recognizes the characteristics of the product.
[1233] The server analyzes the received photograph using image recognition software (e.g., an image recognition API). The server recognizes the characteristics of the product from the photograph (e.g., chocolate cake, fruit parfait). For example, the image recognition API might return the label "chocolate cake". The input is the image data sent to the server, and the output is the recognized product characteristic information.
[1234] Step 4:
[1235] The server generates the product description.
[1236] The server generates a product description based on the recognized characteristics. Using a generative AI model (e.g., Generative AI Model), it generates a product description such as, "A rich chocolate cake. It is not too sweet and has a moist texture." The input is product characteristic information, and the output is the generated product description.
[1237] Step 5:
[1238] The server displays product descriptions on websites and applications.
[1239] The server displays the generated product descriptions on the restaurant's website or application. Specifically, it sends the product descriptions to the front-end in HTML or JSON format and displays them in the user interface. The input is the generated product description, and the output is the product description displayed on the website or application.
[1240] Step 6:
[1241] When a user chooses a product, the device captures the user's emotions.
[1242] When a user selects a product, the device uses its camera and microphone to capture the user's facial expressions and voice tone. For example, a smartphone's camera captures the user's smile, and the microphone records the user's voice tone. The input is the user's facial expressions and voice tone, and the output is the captured emotional data.
[1243] Step 7:
[1244] The server analyzes emotions and recommends appropriate products.
[1245] The server analyzes the captured data using an emotion engine (e.g., an emotion analysis API) to recognize the user's emotions. Based on the recognized emotions, the server recommends products that are suitable for the user. For example, if the user shows a joyful expression, the server will recommend a "fruit parfait." The input is the captured emotion data, and the output is the recommended product information.
[1246] (Application Example 1)
[1247] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1248] Traditional restaurants faced significant challenges in menu creation and improving the customer experience, requiring considerable time and effort. Specifically, there were multiple obstacles, including menu creation, product description generation, personalized recommendations, and multilingual support. Furthermore, providing service that resonated with customers' emotions was difficult, making improving customer satisfaction a challenge.
[1249] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a means for creating menus using a smart device, an AI means for generating product descriptions from photos on the smart device, a means for the AI to generate translation work, a means for improving the customer experience, a means for the AI to explain questions when considering menus, a means for making recommendations according to preferences, and a means for recognizing customer emotions and recommending products based on those emotions. This enables more efficient menu creation, automatic generation of product descriptions, multilingual support, improved customer experience, and emotion-based recommendations.
[1250] A "smart device" is a portable electronic device with advanced functions, such as a smartphone or smart glasses.
[1251] "Menu creation method" refers to the functions and methods for creating restaurant menus using smart devices.
[1252] "AI methods for generating product descriptions" refers to artificial intelligence technology that automatically recognizes the characteristics and features of a product from a photo taken with a smart device and generates a product description based on that.
[1253] "A means of generating translation work using AI" refers to a function that automatically performs multilingual translation using artificial intelligence.
[1254] "Methods for improving the customer experience" refer to methods and functions that enhance the customer experience, such as eliminating the need to wait in line with waitstaff or at the cash register.
[1255] "AI-powered explanations for questions during menu consideration" refers to a function where artificial intelligence automatically explains questions that arise when customers are considering menu options.
[1256] "A means of recommending based on preferences" refers to a function that recommends appropriate products based on the customer's preferences.
[1257] "A means of recognizing customer emotions and recommending products based on those emotions" refers to a function that recognizes customer emotions from facial expressions, tone of voice, etc., and recommends products that match those emotions.
[1258] A system for carrying out this invention includes means for creating menus using a smart device, means for generating product descriptions from photos on a smart device, means for AI to generate translations, means for improving the customer experience, means for AI to explain questions when considering menus, means for making recommendations according to preferences, and means for recognizing customer emotions and recommending products based on those emotions.
[1259] Program Processing Description
[1260] hardware
[1261] The server will be equipped with a high-performance GPU. Smart devices such as smartphones and smart glasses will be used as terminals.
[1262] software
[1263] The server uses the following software:
[1264] OpenCV: Image Processing Library
[1265] Keras: A deep learning library
[1266] Transformers: Generative AI models (such as GPT-3)
[1267] Data processing and calculations
[1268] 1. Menu creation method:
[1269] Take a picture of the food with the device's camera and upload it to the server.
[1270] The server uses OpenCV to preprocess images and Keras to recognize product characteristics and features from the images.
[1271] Based on the recognition results, product descriptions are generated using Transformers.
[1272] 2. Methods for AI to generate translation work:
[1273] The server uses a generative AI model to translate the generated product descriptions into multiple languages.
[1274] In particular, we will implement multilingual support to enhance services for foreign tourists.
[1275] 3. Means of improving the customer experience:
[1276] The terminals allow customers to complete ordering and payment on their smart devices, eliminating the need to wait in line with waitstaff or at the cash register.
[1277] 4. How AI can explain questions that arise when considering menu options:
[1278] The terminal sends questions that customers have when considering the menu to the server, and the server automatically generates explanations using AI.
[1279] 5. Methods for recommending based on preferences:
[1280] The server recommends appropriate products based on the customer's past order history and preferences.
[1281] 6. Means for recognizing customer emotions and recommending products based on those emotions:
[1282] The server analyzes the customer's facial expressions and tone of voice through the device's camera, and uses Keras to recognize their emotions.
[1283] Based on recognized emotions, Transformers are used to recommend appropriate products.
[1284] Specific examples and prompt statements
[1285] Specific example:
[1286] When a customer is wearing smart glasses and looking at a menu, the smart glasses analyze the customer's facial expression and display a message saying, "You seem happy, so we recommend dessert!"
[1287] Examples of prompts for a generative AI model:
[1288] "This dish looks like delicious pasta. Please generate a detailed product description."
[1289] Thus, the embodiment of the invention involves combining smart devices and AI technology to enable menu creation and improve the customer experience in restaurants.
[1290] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[1291] Step 1:
[1292] The user takes a photo of the food using a smart device (smartphone or smart glasses). The photo is saved on the device.
[1293] Step 2:
[1294] The device uploads the captured photos to the server. The server uses OpenCV to preprocess the images, specifically performing tasks such as resizing and noise reduction.
[1295] Step 3:
[1296] The server analyzes pre-processed images using Keras to recognize the characteristics and features of the product. The input is a pre-processed image, and the output is data indicating the characteristics and features of the product.
[1297] Step 4:
[1298] The server generates product descriptions using Transformers based on recognized characteristics and features. The input is data describing the characteristics and features of the product, and the output is the generated product description.
[1299] Step 5:
[1300] The server uses a generative AI model to translate the generated product descriptions into multiple languages. The input is the generated product description, and the output is the product description translated into multiple languages.
[1301] Step 6:
[1302] The terminal allows customers to complete their orders and payments on their smart devices, eliminating the need to wait in line with waiters or at the cash register. The input is the customer's order information, and the output is a notification that the order has been completed.
[1303] Step 7:
[1304] The terminal sends questions that arise when customers are considering the menu to the server. The server automatically generates explanations using AI. The input is the customer's question, and the output is the generated explanation.
[1305] Step 8:
[1306] The server recommends appropriate products based on the customer's past order history and preferences. The input is the customer's order history and preference data, and the output is the recommended products.
[1307] Step 9:
[1308] The server analyzes the customer's facial expressions and voice tone through the terminal's camera, and recognizes their emotions using Keras. The input is data on the customer's facial expressions and voice tone, and the output is the recognized emotion.
[1309] Step 10:
[1310] The server uses Transformers to recommend appropriate products based on the recognized emotions. The input is the recognized emotions, and the output is the recommended products.
[1311] (Example 2)
[1312] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1313] Modern consumers demand quick and accurate product descriptions, as well as multilingual information. They also require personalized product descriptions that cater to their emotions. However, traditional systems struggle to meet these demands, making the improvement of services, particularly for foreign tourists, a significant challenge.
[1314] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1315] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages. As a result, users can quickly and accurately understand product descriptions, and information can be provided in multiple languages. Furthermore, by providing personalized product descriptions that correspond to the user's emotions, customer satisfaction can be improved.
[1316] A "smart device" is a portable electronic device that has internet connectivity and can run applications.
[1317] "Menu creation method" refers to a function that allows restaurants and service businesses to create menus using smart devices.
[1318] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[1319] "Means of generating translation work using artificial intelligence" refers to artificial intelligence technology for translating generated product descriptions into multiple languages.
[1320] "Means of improving the customer experience" refer to features that enhance the convenience and satisfaction customers experience when using a service.
[1321] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a function that allows artificial intelligence to provide appropriate answers to questions that arise when customers are considering menu options.
[1322] "Methods for recommending based on preferences" refer to functions that recommend appropriate products and services based on a customer's past choices and preferences.
[1323] "A means of recognizing user emotions and generating product descriptions that correspond to those emotions" refers to artificial intelligence technology that analyzes user emotions from their input and actions and generates product descriptions that are appropriate for those emotions.
[1324] "Means for translating generated product descriptions into multiple languages" refers to a function for translating generated product descriptions into multiple languages.
[1325] Modes for carrying out the invention
[1326] This invention is a system that includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages.
[1327] Hardware and software to be used
[1328] The server generates product descriptions using a generative AI model (e.g., OpenAI's GPT-4). The generated product descriptions are translated into multiple languages using a translation engine (e.g., Google Translate API). To recognize user emotions, an emotion engine (e.g., IBM Watson's Sentiment Analysis API) is used. Smart devices are portable electronic devices with internet connectivity capable of running applications.
[1329] Explanation of the program's processing
[1330] The server receives text entered by the user through the terminal and generates a product description using a generative AI model. For example, if the user enters "This product is high quality and long-lasting," the server will generate a product description based on this text.
[1331] The generated product descriptions are translated into multiple languages using a translation engine. For example, they are translated into major tourist languages such as English, Chinese, and Korean. Specifically, a description like "This product is high quality and long-lasting" is translated.
[1332] When a user inputs emotional text or audio data through their device, the server uses an emotion engine to recognize the user's emotions. For example, if a user inputs "I'm feeling down today," the server analyzes this text and recognizes the user's emotion as "down."
[1333] The server uses a generative AI model to generate emotionally appropriate product descriptions based on the user's recognized emotions. For example, if the user is feeling down, it will select uplifting words to generate a product description.
[1334] Original description: "This product is high quality and long-lasting."
[1335] Description tailored to the emotion: "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[1336] Examples of specific cases and prompt statements
[1337] As a concrete example, consider the following scenario.
[1338] Scenario 1: Multilingual Translation
[1339] If a user enters "This product is high quality and long-lasting" in Japanese, the server will translate this product description into English, Chinese, and Korean.
[1340] Scenario 2: Product Description Based on Emotions
[1341] If a user enters "I'm feeling down today," the server uses an emotion engine to recognize the user's emotions and generates a product description that will cheer them up.
[1342] Original description: "This product is high quality and long-lasting."
[1343] Description tailored to the emotion: "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[1344] Example of a prompt
[1345] Examples of prompt statements for a generative AI model are as follows:
[1346] Please translate this Japanese product description into English, Chinese, and Korean: 'This product is high quality and long-lasting.'
[1347] "Please generate a product description that will cheer up a user who is feeling down. The original description is 'This product is high quality and long-lasting.'"
[1348] The above describes the embodiments for carrying out this invention.
[1349] The flow of the specific processing in Example 2 will be explained using Figure 19.
[1350] Step 1:
[1351] The user uses a terminal to input the text that will form the basis of the product description. For example, the user might input, "This product is high quality and long-lasting." The entered text is then sent to the server.
[1352] Step 2:
[1353] The server uses a generative AI model to generate a product description based on the text entered by the user. Specifically, the server sends the following prompt to the generative AI model: "Generate a product description based on the user's input, 'This product is high quality and long-lasting.'" The generative AI model generates a product description based on this prompt and returns it to the server. The output is the generated product description.
[1354] Step 3:
[1355] The server sends the generated product description to the translation engine for translation into multiple languages. Specifically, the server sends the following prompt to the translation engine: "Translate this product description into English, Chinese, and Korean: 'This product is high quality and long-lasting.'" The translation engine translates based on this prompt and returns it to the server. The output is the product description translated into multiple languages.
[1356] Step 4:
[1357] The user inputs text or audio data related to their emotions through their device. For example, the user might input, "I'm feeling down today." The input text or audio data is then sent to the server.
[1358] Step 5:
[1359] The server uses an emotion engine to recognize the user's emotions. Specifically, the server sends the following prompt to the emotion engine: "Analyze the user's input 'I'm feeling down today' and recognize the emotion." Based on this prompt, the emotion engine analyzes the emotion and returns it to the server. The output is the recognized emotion.
[1360] Step 6:
[1361] Based on the recognized user's emotions, the server uses a generative AI model to generate an emotion-responsive product description. Specifically, the server sends the following prompt to the generative AI model: "Generate an uplifting product description for a user who is feeling down. The original description is 'This product is high quality and long-lasting.'" Based on this prompt, the generative AI model generates an emotion-responsive product description and returns it to the server. The output is an emotion-responsive product description.
[1362] Step 7:
[1363] The server sends a product description to the user's device that corresponds to the generated emotion. The user can then view the uplifting product description through their device. Specifically, if the user enters "I'm feeling down today," the server will provide the user with a product description such as, "This product is high quality and will brighten your everyday life. It's durable, so you can use it for a long time!"
[1364] The above is a detailed explanation of the program's processing flow.
[1365] (Application Example 2)
[1366] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1367] Traditional e-commerce sites often provide product descriptions in a single language, which posed a problem for foreign tourists and multilingual users, making them difficult to understand. Furthermore, the lack of emotionally resonant product descriptions made it difficult to increase user purchase intent. Additionally, users sometimes spent a significant amount of time trying to understand product descriptions, resulting in a less-than-ideal shopping experience.
[1368] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1369] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages. This makes it possible to provide product descriptions that correspond to the user's emotions in multiple languages, and to provide product descriptions that are easy to understand for foreign tourists and multilingual users. Furthermore, it can increase the user's desire to purchase and improve the purchasing experience.
[1370] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.
[1371] "Menu creation methods" refer to functions or software that allow users to create menus for products and services using smart devices.
[1372] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[1373] "Means of generating translation work using artificial intelligence" refers to artificial intelligence technology for translating generated product descriptions into multiple languages.
[1374] "Methods for improving the customer experience" refer to functions and software that enhance the customer's experience when using a product or service.
[1375] "A means for artificial intelligence to explain questions when considering menu options" refers to a function in which artificial intelligence provides appropriate explanations for questions that arise when users are considering menu options.
[1376] "Preference-based recommendation methods" refer to features that recommend appropriate products and services based on the user's preferences and past behavior.
[1377] "A means of recognizing user emotions and generating product descriptions that correspond to those emotions" refers to artificial intelligence technology that analyzes user emotions and generates product descriptions that are appropriate for those emotions.
[1378] "Means for translating generated product descriptions into multiple languages" refers to functions or software for translating generated product descriptions into multiple languages.
[1379] This invention will now describe embodiments for carrying out this invention. The system includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing the user's emotions and generating product descriptions corresponding to those emotions, and means for translating the generated product descriptions into multiple languages.
[1380] System Configuration
[1381] hardware
[1382] Smart devices: Portable electronic devices with internet connectivity, such as smartphones, tablets, and smart glasses.
[1383] Server: A high-performance computer responsible for data processing and storage.
[1384] software
[1385] Artificial intelligence model: Generates product descriptions using the OpenAI API.
[1386] Translation software: Multilingual translation is performed using the Google Translate API.
[1387] Emotion recognition software: Use EmotionRecognizer to recognize the user's emotions.
[1388] Data processing and data calculation
[1389] 1. Image Analysis: Images of products taken with a smart device are sent to the server. The server uses an image analysis algorithm to extract product characteristics and features and generate a product description.
[1390] 2. Emotion Recognition: The smart device's camera and microphone are used to analyze the user's facial expressions and voice tone. EmotionRecognizer software recognizes the user's emotions and sends the results to the server.
[1391] 3. Product Description Generation: The server generates prompt text based on the user's emotions and uses the OpenAI API to generate a product description appropriate to those emotions.
[1392] 4. Multilingual Translation: The generated product description is translated into multiple languages using the Google Translate API. The translation results support major languages such as English, Chinese, and Korean.
[1393] Specific example
[1394] For example, if a user is feeling down, the smart device's camera captures the user's facial expression, and EmotionRecognizer recognizes it as "feeling down." The server then generates the following prompt:
[1395] "Generate a product description that will cheer up a user when they're feeling down: This product is made from high-quality materials and will last a long time."
[1396] This prompt is sent to the OpenAI API to generate a sentiment-appropriate product description. The generated product description will look like this:
[1397] "This product was created to brighten your life, even just a little. It's made from high-quality materials and is built to last."
[1398] Next, we will translate this product description into multiple languages using the Google Translate API.
[1399] In this way, it becomes possible to provide product descriptions in multiple languages that are tailored to the user's emotions.
[1400] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[1401] Step 1:
[1402] The user takes a picture of the product with a smart device.
[1403] Input: Product image
[1404] Output: Captured image data
[1405] Specific operation: The user uses a smart device such as a smartphone or tablet to take a picture of the product they are considering purchasing. This image data is stored on the smart device.
[1406] Step 2:
[1407] The device sends the captured image to the server.
[1408] Input: Captured image data
[1409] Output: Image data sent to the server
[1410] Specific operation: The smart device sends the captured image data to a server via the internet. The HTTP protocol and other protocols are used for transmission.
[1411] Step 3:
[1412] The server performs image analysis to extract the product's characteristics and features.
[1413] Input: Sent image data
[1414] Output: Characteristics and features of the extracted products
[1415] Specific operation: The server uses image analysis algorithms to automatically extract product characteristics and features from image data. For example, it recognizes the product's color, shape, brand logo, etc.
[1416] Step 4:
[1417] The server generates a product description based on the extracted characteristics and features.
[1418] Input: Characteristics and features of the extracted products
[1419] Output: Generated product description
[1420] Specific operation: The server uses a generative AI model to generate product descriptions based on extracted characteristics and features. For example, a description such as "This product is made from high-quality materials and will last a long time" might be generated.
[1421] Step 5:
[1422] The device recognizes the user's emotions.
[1423] Input: User's facial expressions and tone of voice
[1424] Output: Recognized user emotions
[1425] Specific operation: Using the camera and microphone of a smart device, the EmotionRecognizer software captures the user's facial expressions and voice tone, and recognizes the user's emotions. For example, emotions such as "depressed" or "happy" can be recognized.
[1426] Step 6:
[1427] The server generates prompt messages that respond to the user's emotions.
[1428] Input: Recognized user sentiment, generated product description
[1429] Output: Prompts tailored to your emotions
[1430] Specific operation: Based on the recognized user's emotions, the server generates prompt messages to send to the generative AI model. For example, a prompt message such as "Generate a product description to cheer up a user who is feeling down: This product is made from high-quality materials and will last a long time." might be generated.
[1431] Step 7:
[1432] The server uses an AI model to generate product descriptions that are appropriate for the user's emotions.
[1433] Input: Sentiment-based prompt text
[1434] Output: Product description suitable for emotions
[1435] Specific operation: The server uses a generative AI model (e.g., the OpenAI API) to generate emotionally appropriate product descriptions based on the prompt text. For example, a description such as "This product was created to brighten your life a little. It is made from high-quality materials and is durable." might be generated.
[1436] Step 8:
[1437] The server translates the generated product description into multiple languages.
[1438] Input: Generated product description
[1439] Output: Product description translated into multiple languages
[1440] Specific operation: The server uses the Google Translate API to translate the generated product description into multiple languages, such as English, Chinese, and Korean. For example, a translation result such as "This product is made to brighten your life a little. It is made of high-quality materials and lasts long." might be obtained.
[1441] Step 9:
[1442] The server sends product descriptions translated into multiple languages to the device.
[1443] Input: Product description translated into multiple languages
[1444] Output: Translation results sent to the terminal
[1445] Specific operation: The server sends product descriptions translated into multiple languages to smart devices via the internet. HTTP protocol and similar protocols are used for transmission.
[1446] Step 10:
[1447] The device displays product descriptions translated into multiple languages to the user.
[1448] Input: Translation result sent to the device
[1449] Output: Multilingual product description displayed to the user
[1450] Specific operation: The smart device displays the product description, translated into multiple languages, to the user. This allows the user to understand the emotionally relevant product description in multiple languages.
[1451] (Example 3)
[1452] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1453] Traditional customer experience improvement systems often rely on manual processes for menu creation, product descriptions, and translation, resulting in inefficiency. Furthermore, they may fail to adequately address customer questions and provide personalized recommendations, potentially leading to decreased customer satisfaction. Additionally, the lack of features to recognize and respond to customer emotions hinders the improvement of the overall customer experience.
[1454] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 3 is realized by the following means. In this invention, the server includes a menu creation means using a smart device, an artificial intelligence means for generating product descriptions from images on the smart device, a means for the artificial intelligence to generate translation work, a means for improving the customer experience by reducing waiting time, a means for the artificial intelligence to explain questions when considering a menu, a means for recommending products according to preferences, and a means for recognizing the user's emotions and providing services that correspond to those emotions. This makes it possible to quickly answer questions from customers when considering a menu, recommend products according to their preferences, and further improve customer satisfaction by providing services that correspond to the customer's emotions.
[1455] A "smart device" is a portable electronic device with internet connectivity, such as a smartphone or tablet.
[1456] "Menu creation method" refers to a function that allows restaurants and service businesses to create menus using smart devices.
[1457] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[1458] "Means of generating translation work using artificial intelligence" refers to a function that automatically performs text and audio translation using artificial intelligence.
[1459] "Methods for improving customer experience by reducing waiting times" refer to functions that shorten the waiting time customers have to wait when using a service, thereby providing a smoother experience.
[1460] "A means for artificial intelligence to explain questions when customers are considering menu options" refers to a function in which artificial intelligence provides appropriate answers to questions that customers may have when considering menu options.
[1461] "A method for recommending products based on preferences" refers to a function in which artificial intelligence recommends appropriate products based on the customer's preferences and past choices.
[1462] "Means of recognizing user emotions and providing services that respond to those emotions" refers to a function that recognizes user emotions from their facial expressions and voice and provides appropriate services that respond to those emotions.
[1463] This invention is a system for improving the customer experience and includes means for creating menus using smart devices, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translations, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and means for recognizing user emotions and providing services that respond to those emotions.
[1464] Hardware and software to be used
[1465] hardware
[1466] Smart devices: Portable electronic devices with internet connectivity, such as smartphones and tablets.
[1467] Server: A high-performance server (e.g., a server with an NVIDIA GPU).
[1468] software
[1469] Generative AI model: GPT-4 is used as an example.
[1470] Emotion recognition engine: Affectiva SDK is used as an example.
[1471] Program Processing Description
[1472] The server receives information sent from the smart device and performs analysis using a generative AI model. For example, if the user inputs "I like spicy food," the generative AI model generates a list of spicy dishes, and the server sends that list to the smart device. The smart device then displays the received list to the user.
[1473] Furthermore, when a user asks, "What are the ingredients in this dish?", the server uses a generative AI model to generate ingredient information and sends it to the smart device. The smart device then displays the received ingredient information to the user.
[1474] Furthermore, smart devices use their built-in cameras and microphones to capture the user's facial expressions and voice, and use an emotion recognition engine to recognize the user's emotions. For example, if the user shows an angry expression, the smart device sends that information to a server. The server uses a generative AI model to generate an apology and sends it to the smart device. The smart device then displays the received apology to the user.
[1475] Examples of specific cases and prompt statements
[1476] Example 1: User enters "I like spicy food"
[1477] The user opens the app on their smartphone and enters "I like spicy food."
[1478] The server generates a list of spicy dishes using an AI model.
[1479] The server generates a list and sends it to the smart device.
[1480] A smart device displays a list of spicy dishes to the user.
[1481] Example prompt: "Please list dishes you would recommend to a customer who likes spicy food."
[1482] Example 2: The user asks, "What are the ingredients in this dish?"
[1483] The user opens the app on their tablet and asks, "What are the ingredients in this dish?"
[1484] The server generates ingredient information using an AI model.
[1485] The server generates ingredient information and sends it to the smart device.
[1486] A smart device displays ingredient information to the user.
[1487] Example of a prompt: "What are the ingredients in this dish?"
[1488] Example 3: The user shows an angry expression.
[1489] The user displays an angry expression while using their smartphone.
[1490] Smart devices use cameras to capture the user's facial expressions, which are then analyzed by an emotion recognition engine.
[1491] The smart device sends information about the anger it recognizes to the server.
[1492] The server generates an apology using an AI model.
[1493] The server generates an apology message and sends it to the smart device.
[1494] Smart devices display an apology message to the user.
[1495] Example prompt: "Generate an apology for when the user is angry."
[1496] In this way, the server, smart device, and user work together to operate the system and improve the customer experience. The flow of a specific process in Example 3 will be explained using Figure 21.
[1497] Step 1:
[1498] The user inputs information through their device. The user uses a smartphone or tablet to input questions and preferences into the system. For example, they might input "I like spicy food." The input data is saved on the device in text format.
[1499] Step 2:
[1500] The terminal sends the input information to the server. The terminal sends the information entered by the user to the server via the internet. The input data is sent to the server in text format.
[1501] Step 3:
[1502] The server analyzes the input information using a generative AI model. The server passes the received input information to the generative AI model (e.g., GPT-4) for analysis. The generative AI model generates appropriate answers or recommendations based on the input information. For example, given the information "I like spicy food," it generates a list of spicy dishes. The input data is in text format, and the output data is in list format.
[1503] Step 4:
[1504] The server generates appropriate answers and recommendations based on the analysis results. The server generates answers and recommendations for the user based on the analysis results obtained from the generating AI model. For example, it might generate a list of spicy dishes. The output data is in list format.
[1505] Step 5:
[1506] The server sends the generated responses and recommendations to the device. The server sends the generated responses and recommendations to the device via the internet. The output data is sent to the device in list format.
[1507] Step 6:
[1508] The device displays answers and recommendations to the user. The device displays answers and recommendations received from the server to the user. The user can review the displayed information. The output data is displayed in list format.
[1509] Step 7:
[1510] The device recognizes the user's emotions using an emotion engine. The device captures the user's facial expressions and voice using its built-in camera and microphone, and recognizes the user's emotions using an emotion engine (e.g., Affectiva SDK). Input data is in image or audio format, and output data is in emotion information format.
[1511] Step 8:
[1512] The device sends recognized emotion information to the server. The device sends recognized emotion information to the server via the internet. Input data is sent to the server in emotion information format, and output data is sent to the server in emotion information format.
[1513] Step 9:
[1514] The server generates appropriate services based on emotional information. The server uses a generative AI model to generate appropriate services based on the received emotional information. For example, if the user shows an angry expression, it will generate an apology. Input data is in emotional information format, and output data is in text format.
[1515] Step 10:
[1516] The server sends the generated service to the terminal. The server sends the generated service to the terminal via the internet. The output data is sent to the terminal in text format.
[1517] Step 11:
[1518] The terminal provides services to the user. The terminal provides services received from the server to the user. For example, it displays an apology. The output data is displayed in text format.
[1519] (Application Example 3)
[1520] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1521] Traditional food delivery services faced challenges such as users spending a lot of time choosing from menus and difficulty in providing services tailored to user preferences and emotions. Furthermore, multilingual support for foreign tourists was insufficient, highlighting the need for improved user experience. There is a demand to solve these problems and provide a more comfortable and personalized service.
[1522] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[1523] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for artificial intelligence to generate translation tasks, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and means for recognizing the user's emotions and providing services corresponding to those emotions. As a result, users can efficiently select menus and receive personalized services tailored to their individual preferences and emotions. Furthermore, multilingual support enhances services for foreign tourists.
[1524] A "smart device" is a portable electronic device with advanced computing capabilities, such as a smartphone or tablet.
[1525] A "menu creation method" refers to a method or device for generating a list of products or services that a user can select.
[1526] "An artificial intelligence method for generating product descriptions from images" refers to an artificial intelligence system that uses image recognition technology to automatically analyze the characteristics and features of a product and generates a product description based on that analysis.
[1527] "Means of generating translation work using artificial intelligence" refers to artificial intelligence systems that automatically translate text and audio between different languages.
[1528] "Customer experience improvement measures" refer to methods and devices that enhance the convenience and satisfaction users experience when using a service.
[1529] "A means for artificial intelligence to explain questions when considering menu options" refers to a system in which artificial intelligence provides appropriate answers to questions that arise when users are choosing from a menu.
[1530] A "method for recommending products based on preferences" is a system that recommends appropriate products and services based on the user's past choices and input information.
[1531] An "emotion recognition system" is a system that analyzes a user's emotions from their facial expressions, tone of voice, etc., and responds accordingly.
[1532] A system for carrying out this invention includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images, means for artificial intelligence to generate translation work, means for improving the customer experience by reducing waiting times, means for artificial intelligence to explain questions when considering menus, means for recommending products according to preferences, and emotion recognition means for recognizing the user's emotions and providing services in accordance with those emotions.
[1533] System program
[1534] The program in this system performs the following operations:
[1535] Hardware and software
[1536] Hardware:
[1537] Smart devices (smartphones, tablets, etc.)
[1538] Camera (to capture the user's facial expressions)
[1539] Computer (a processing unit for executing programs)
[1540] software:
[1541] OpenCV (image processing library)
[1542] Keras (deep learning library)
[1543] Transformers (a library of generative AI models)
[1544] Data processing and data calculation
[1545] 1. Menu creation method:
[1546] The server generates a list of products and services that users can select using their smart devices.
[1547] 2. Artificial intelligence means for generating product descriptions from images:
[1548] The server analyzes images captured by the smart device's camera and automatically recognizes the product's characteristics and features.
[1549] Based on the recognized information, a product description is generated.
[1550] 3. Means for artificial intelligence to generate translation work:
[1551] The server automatically translates text and audio between different languages.
[1552] In particular, we will provide multilingual support to enhance services for foreign tourists.
[1553] 4. Means of improving customer experience:
[1554] The server reduces waiting times for users when accessing the service, thereby improving convenience.
[1555] 5. How artificial intelligence can explain questions that arise when considering menu options:
[1556] The server uses a generative AI model to provide appropriate answers to questions that arise when users select menu items.
[1557] 6. Methods for recommending products based on preferences:
[1558] The server recommends appropriate products and services based on the user's past choices and input information.
[1559] 7. Emotion recognition means:
[1560] The server analyzes the user's facial expressions and voice tone captured by the camera to recognize their emotions.
[1561] Respond appropriately to the recognized emotions.
[1562] Specific example
[1563] When a user enters "I like spicy food," the server uses a generative AI model to recommend spicy dishes.
[1564] When a user asks, "What are the ingredients in this dish?", the server uses a generative AI model to provide information about the ingredients.
[1565] If the user shows an angry expression, the server uses emotion recognition to apologize with a message like, "I'm sorry. Is there a problem?"
[1566] Example of a prompt
[1567] "I like spicy food. What do you recommend?"
[1568] "What are the ingredients in this dish?"
[1569] "How do you respond if a user is showing signs of anger?"
[1570] In this way, users can efficiently select from the menu and receive personalized service tailored to their individual preferences and feelings. Furthermore, multilingual support enhances services for foreign tourists.
[1571] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[1572] Step 1:
[1573] The user opens the menu using a smart device.
[1574] Input: A request to display a menu initiated by the user.
[1575] Data processing: The server retrieves menu information from the database and sends it to the smart device.
[1576] Output: A menu is displayed on the smart device.
[1577] Step 2:
[1578] The user takes a picture of the food with the camera on their smart device.
[1579] Input: An image of a dish taken by the user.
[1580] Data processing: The server receives images and analyzes them using an image recognition algorithm (OpenCV).
[1581] Output: The analysis results extract the characteristics and features of the dishes.
[1582] Step 3:
[1583] The server generates a product description from the image.
[1584] Input: Data on the characteristics and features of the dish.
[1585] Data processing: The server uses a generative AI model to generate product descriptions based on characteristics and features.
[1586] Output: Product description text is generated and sent to the smart device.
[1587] Step 4:
[1588] Users enter questions when considering menu options.
[1589] Input: A question entered by the user (e.g., "What are the ingredients in this dish?").
[1590] Data processing: The server generates answers to questions using a generative AI model.
[1591] Output: The answer text is generated and displayed on the smart device.
[1592] Step 5:
[1593] The server recommends products based on the user's preferences.
[1594] Input: User's past selections and input information (e.g., "I like spicy food").
[1595] Data processing: The server uses a generative AI model to recommend products based on user preferences.
[1596] Output: A list of recommended products is displayed on the smart device.
[1597] Step 6:
[1598] The user captures their facial expressions using the camera on their smart device.
[1599] Input: User's facial expression image.
[1600] Data processing: The server analyzes emotions using a facial recognition algorithm (Keras).
[1601] Output: As an analysis result, user sentiment data is generated.
[1602] Step 7:
[1603] The server provides services that respond to the user's emotions.
[1604] Input: User emotion data (e.g., anger).
[1605] Data processing: The server generates appropriate responses based on sentiment data (e.g., apology messages).
[1606] Output: The corresponding message is displayed on the smart device.
[1607] Step 8:
[1608] The server performs the translation work.
[1609] Input: Text or audio data entered by the user.
[1610] Data processing: The server uses a translation algorithm to translate the input data into multiple languages.
[1611] Output: Translated text and audio data are displayed on the smart device.
[1612] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1613] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1614] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.
[1615] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1616] [Third Embodiment]
[1617] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1618] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1619] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1620] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1621] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1622] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1623] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1624] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1625] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1626] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1627] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1628] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[1629] "Example of form 1"
[1630] In one embodiment of the present invention, a restaurant operator creates a menu using a smartphone. Specifically, the operator uploads photos of products taken with the smartphone's camera, and AI automatically recognizes the characteristics and features of the products from the photos. Based on the recognition results, a product description is generated. This product description is then displayed to customers on the restaurant's website or application.
[1631] "Example of form 2"
[1632] Furthermore, in this embodiment of the present invention, the AI generates the translation work. Specifically, the generated product description is translated into multiple languages. This makes the product description easier for foreign tourists to understand. For example, it is possible to translate into major tourist languages such as English, Chinese, and Korean.
[1633] "Example of form 3"
[1634] Furthermore, in this embodiment of the present invention, the customer experience is also improved. Specifically, when a customer is considering a menu, the AI explains their questions and recommends products according to their preferences. For example, if a customer inputs information such as "I like spicy food," the AI will recommend spicy dishes. Also, if a customer asks a question such as "What are the ingredients in this dish?", the AI will provide an answer to that question.
[1635] The following describes the processing flow for each example of the form.
[1636] "Example of form 1"
[1637] Step 1: The restaurant operator takes photos of the products using their smartphone camera.
[1638] Step 2: Upload the photos you've taken to the system.
[1639] Step 3: The AI within the system automatically recognizes the product's characteristics and features from the photograph.
[1640] Step 4: The AI generates a product description based on the recognition results.
[1641] Step 5: The generated product description is displayed to customers on the restaurant's website or application.
[1642] "Example of form 2"
[1643] Step 1: Obtain the product description generated by the AI.
[1644] Step 2: The AI translates the product description into multiple languages.
[1645] Step 3: The translated product description is displayed to foreign tourists on the restaurant's website or application.
[1646] "Example of form 3"
[1647] Step 1: When customers are considering the menu, they input their questions into the AI.
[1648] Step 2: The AI generates the answer to that question.
[1649] Step 3: The AI recommends products based on the customer's preferences.
[1650] Step 4: The AI-generated answers and recommendations are displayed to the customer.
[1651] (Example 1)
[1652] Next, we will describe Embodiment 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1653] When restaurant operators create menus, the process of taking photos of products and then recognizing their characteristics and features to generate product descriptions is time-consuming. Furthermore, there is a lack of multilingual support for foreign tourists and other means to enhance the customer experience. Therefore, there is a need for increased efficiency in menu creation and improved customer satisfaction.
[1654] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1655] In this invention, the server includes means for creating menus using a smartphone, means for artificial intelligence to generate product descriptions from photos taken with a smartphone, means for artificial intelligence to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for uploading product photos taken with a smartphone camera and for artificial intelligence to automatically recognize the characteristics and features of the products from those photos, and means for generating product descriptions based on the recognition results. This makes it possible to improve the efficiency of menu creation and enhance customer satisfaction.
[1656] A "smartphone" is a multi-functional mobile device that, in addition to the functions of a mobile phone, is capable of internet connectivity and the use of applications.
[1657] "Menu creation methods" refer to the methods and tools that restaurant operators use to create menus, and especially those that utilize smartphones.
[1658] "Artificial intelligence methods" refer to technologies that use machine learning and data analysis to automatically perform specific tasks.
[1659] "Means of generating translation work using artificial intelligence" refers to methods and technologies that use artificial intelligence to translate text into multiple languages.
[1660] "Methods for improving the customer experience" refer to methods and tools for improving the convenience and satisfaction customers experience when using a service.
[1661] "Methods for AI to explain questions when considering menus" refers to methods and technologies in which artificial intelligence automatically provides answers to questions that customers may have when considering menus.
[1662] "Methods of recommending based on preferences" refer to methods and technologies that recommend appropriate products and services based on a customer's preferences and past choices.
[1663] "Means for automatically recognizing the characteristics and features of a product" refers to methods and technologies that use artificial intelligence to automatically extract the characteristics and features of a product from photographs or data.
[1664] "Means for generating product descriptions" refers to methods and technologies for creating product descriptions in natural language based on recognized product characteristics and features.
[1665] This invention is a system that allows restaurant operators to create menus using their smartphones. Specifically, users upload photos of products taken with their smartphone cameras, and artificial intelligence (AI) automatically recognizes the characteristics and features of the products from the photos. Based on the recognition results, the system generates product descriptions. These product descriptions are then displayed to customers on the restaurant's website or application.
[1666] Hardware and software to be used
[1667] Smartphone: A device used for taking and uploading photos.
[1668] Server: Performs data processing and runs AI models.
[1669] Artificial intelligence models: Image recognition models and natural language generation models based on TensorFlow and PyTorch (e.g., GPT-3, BERT).
[1670] Data processing and data calculation
[1671] 1. The user takes a photo of the product with their smartphone.
[1672] The user launches their smartphone's camera app and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[1673] 2. The user uploads photos to the system.
[1674] The user opens a dedicated application, selects the photos they have taken, and presses the upload button. The photos are then sent to the server via the internet.
[1675] 3. The server receives the photos and inputs them into the AI model.
[1676] The server receives the uploaded photos and inputs them into an AI model for image processing. The AI model used here is based on TensorFlow or PyTorch.
[1677] 4. The server uses an AI model to recognize the characteristics and features of the product.
[1678] The server uses an AI model to recognize the characteristics and features of a product from a photograph. For example, it extracts the type of dessert, main ingredients, and visual features.
[1679] 5. The server generates a product description based on the recognition results.
[1680] The server generates product descriptions based on the AI's recognition results. These product descriptions are created using natural language generation technology. The software used includes GPT-3 and BERT.
[1681] 6. The server displays the product description on the website or application.
[1682] The server displays the generated product descriptions on the restaurant's website or application. Customers can then view them.
[1683] Specific example
[1684] A user takes a photo of a new dessert, "Chocolate Cake," and uploads it to the system. The server receives the photo and uses an AI model to recognize features such as "Chocolate Cake," "Cream Topping," and "Berry Decoration." The server uses GPT-3 to generate a product description such as, "This chocolate cake features rich chocolate and creamy toppings. The berry decoration makes it visually appealing," and displays it on the website.
[1685] Example of a prompt
[1686] Please upload a photo of your new dessert. AI will recognize the product's characteristics and features from the photo and generate a product description.
[1687] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1688] Step 1:
[1689] The user takes a photo of the product with their smartphone.
[1690] Input: Product (e.g., Chocolate cake)
[1691] Output: Photos of the photographed product
[1692] Specific action: The user launches the camera app on their smartphone and takes a picture of the product. For example, they might take a picture of a new dessert called "Chocolate Cake."
[1693] Step 2:
[1694] The user uploads a photo to the system.
[1695] Input: Photos of the product
[1696] Output: Photo data sent to the server
[1697] Specific operation: The user opens a dedicated application, selects the photo they have taken, and presses the upload button. The photo is sent to the server via the internet.
[1698] Step 3:
[1699] The server receives the photos and inputs them into the AI model.
[1700] Input: Photo data sent to the server
[1701] Output: Photo data input to the AI model
[1702] Specific operation: The server receives an HTTP request, temporarily stores the photo data, and inputs it into a TensorFlow or PyTorch model.
[1703] Step 4:
[1704] The server uses an AI model to recognize the characteristics and features of the product.
[1705] Input: Photo data entered into the AI model
[1706] Output: Characteristics and features of the recognized product (e.g., chocolate cake, cream topping, berry decoration)
[1707] Specific operation: The server runs an AI model to extract product characteristics and features from a photograph. For example, it recognizes the type of dessert, main ingredients, and visual features.
[1708] Step 5:
[1709] The server generates a product description based on the recognition results.
[1710] Input: Characteristics and features of the recognized product
[1711] Output: Generated product description (Example: "This chocolate cake features rich chocolate and creamy toppings. The berry decorations make it visually appealing.")
[1712] Specific operation: The server uses GPT-3 or BERT to generate product descriptions based on the recognition results as input.
[1713] Step 6:
[1714] The server displays product descriptions on websites and applications.
[1715] Input: Generated product description
[1716] Output: Product description displayed on the website or application
[1717] Specific operation: The server saves the product description in the website's database and sends the data to the front-end for display. Customers can then view this data.
[1718] (Application Example 1)
[1719] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1720] Traditional restaurant menu creation was often done manually, which was time-consuming and labor-intensive, and made it difficult to adequately convey the appeal of the products. Furthermore, the lack of multilingual support for foreign tourists and insufficient recommendation features tailored to customer preferences were also problems. In addition, there were limited means to improve the in-store customer experience, such as having to wait for waitstaff or in line at the register, highlighting the need for increased customer satisfaction.
[1721] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1722] In this invention, the server includes means for creating menus using a smartphone, means for generating product descriptions from photos taken with the smartphone, means for the AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the register, means for the AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on the recognized features using a natural language generation model. This enables efficient menu creation and automatic generation of attractive product descriptions, as well as multilingual support, improved customer experience, and personalized recommendations.
[1723] "A method for creating menus using a smartphone" refers to a method for efficiently creating restaurant menus using the camera and applications of a smartphone.
[1724] "AI methods for generating product descriptions from smartphone photos" refers to artificial intelligence methods that analyze product photos taken with a smartphone, recognize their characteristics and features, and automatically generate product descriptions.
[1725] "Methods for AI to generate translation work" refers to methods of using artificial intelligence to translate product descriptions and menu contents into multiple languages.
[1726] "Methods to improve the customer experience by eliminating the need to wait for waitstaff or in line at the register" refers to methods that allow customers to receive service smoothly without having to wait for waitstaff or in line at the register.
[1727] "An AI-powered solution for questions during menu consideration" refers to a method in which artificial intelligence automatically provides answers to questions that customers may have when considering menu options.
[1728] "Methods of recommending based on preferences" refer to methods of recommending appropriate products or menus based on a customer's past choices and preferences.
[1729] "Methods for recognizing the characteristics and features of a product using an image recognition model" refers to methods for automatically recognizing the characteristics and features of a product from a photograph using image recognition technology.
[1730] "Means for generating product descriptions based on features recognized using a natural language generation model" refers to means for automatically generating attractive product descriptions using natural language generation technology based on recognized product characteristics and features.
[1731] A system for carrying out this invention includes means for creating menus using a smartphone, means for generating product descriptions from photos on a smartphone, means for AI to generate translations, means for improving the customer experience by eliminating the need to wait for waitstaff or queues at the cash register, means for AI to explain questions when considering menus, means for making recommendations according to preferences, means for recognizing product characteristics and features using an image recognition model, and means for generating product descriptions based on recognized features using a natural language generation model.
[1732] System program
[1733] The server implements the system using the following hardware and software.
[1734] Hardware:
[1735] Smartphone (with camera)
[1736] Server (for hosting AI models)
[1737] software:
[1738] TensorFlow: A library for image recognition
[1739] OpenAI GPT-3: API for Natural Language Generation
[1740] Explanation of the process
[1741] Image recognition:
[1742] Users take photos of their food with their smartphone cameras and upload them to a server via an application. The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of the food from the image. For example, features such as "chicken curry, spicy, tomato-based" might be recognized.
[1743] Natural language generation:
[1744] The server converts the recognized features into a string and sends it to OpenAI GPT-3 as a prompt. An example of a prompt is, "Generate an appealing product description for a dish with the following features: Chicken curry, spicy, tomato-based." GPT-3 generates an appealing product description based on this prompt. For example, it might generate a product description such as, "This spicy chicken curry features juicy chicken simmered in a tomato-based sauce. The aromatic spices will whet your appetite."
[1745] Translation work:
[1746] The generated product descriptions are translated into multiple languages as needed. The server uses AI to perform translations in multiple languages, such as English and Chinese.
[1747] Improving the customer experience:
[1748] Customers can use their smartphones to select menu items and place orders in-store. This reduces waiting times at the counter and cashier, enabling smoother service.
[1749] Explanation of the question and recommendations:
[1750] The AI automatically provides answers to questions that customers may have when considering menu options. It also recommends appropriate products and menu items based on the customer's past choices and preferences.
[1751] In this way, menu creation becomes more efficient, attractive product descriptions are automatically generated, and multilingual support, improved customer experience, and personalized recommendations become possible.
[1752] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1753] Step 1:
[1754] The user takes a photo of the food with their smartphone camera and uploads it to the server through the application. The input is the photo of the food, and the output is the image data sent to the server.
[1755] Step 2:
[1756] The server uses a TensorFlow ResNet50 model to recognize the characteristics and features of dishes from uploaded image data. The input is image data, and the output is a list of the recognized characteristics and features. Specifically, the image data is preprocessed and then input into the model to obtain prediction results.
[1757] Step 3:
[1758] The server converts the recognized traits and features into a string and sends it to OpenAI GPT-3 as a prompt. The input is a list of traits and features, and the output is a prompt statement. Specifically, it formats the traits and features and generates a prompt statement.
[1759] Step 4:
[1760] The server receives a response from GPT-3 and generates an attractive product description. The input is a prompt, and the output is the generated product description. Specifically, it calls the GPT-3 API, parses the response, and retrieves the product description.
[1761] Step 5:
[1762] The server translates the generated product description into multiple languages as needed. The input is the product description, and the output is the translated product description. Specifically, it calls a translation API to translate into multiple languages.
[1763] Step 6:
[1764] Users use their smartphones to select menu items and place orders in the store. The input is a generated product description, and the output is the user's order information. Specifically, the application displays product descriptions, and the user selects their order.
[1765] Step 7:
[1766] The server automatically provides answers to questions that arise when users consider the menu, using AI. The input is the user's question, and the output is the AI's answer. Specifically, it analyzes the question and generates an appropriate answer.
[1767] Step 8:
[1768] The server recommends appropriate products and menu items based on the user's past choices and preferences. The input is the user's past selection data, and the output is the recommended products and menu items. Specifically, it analyzes past selection data and uses a recommendation algorithm to suggest products.
[1769] (Example 2)
[1770] Next, we will describe Example 2 of the Form Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1771] Traditional systems often required manual translation of product descriptions into multiple languages, resulting in time-consuming and labor-intensive processes. Furthermore, services for foreign tourists were inadequate, making it difficult for them to understand product descriptions. Additionally, insufficient efforts were made to improve the customer experience and address questions during menu selection, potentially leading to decreased customer satisfaction.
[1772] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1773] In this invention, the server includes means for creating menus using a smart device, means for artificial intelligence to generate product descriptions from images on the smart device, means for the artificial intelligence to generate translation tasks, means for improving the customer experience, means for the artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into the generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results. As a result, multilingual translation of product descriptions is automated, improving services for foreign tourists, enhancing the customer experience, and resolving questions when considering menus.
[1774] A "smart device" is a portable electronic device, such as a smartphone or tablet, that possesses advanced computing power and communication capabilities.
[1775] "Menu creation methods" refer to functions and applications that use smart devices to create menus for restaurants and retail stores.
[1776] "An artificial intelligence method for generating product descriptions from images" refers to artificial intelligence technology that analyzes images taken with a smart device and automatically generates product descriptions based on those images.
[1777] "A means of generating translation work using artificial intelligence" refers to a function that uses artificial intelligence to translate text data into other languages.
[1778] "Means of improving the customer experience" refer to functions and methods that enhance the convenience and satisfaction customers experience when using products or services.
[1779] "A means for artificial intelligence to explain questions when considering menus" refers to a function in which artificial intelligence automatically provides answers to questions and concerns that arise when customers are considering menus.
[1780] "Means of recommending based on preferences" refers to a function that recommends appropriate products and services based on the customer's past choices and preferences.
[1781] "Means for translating product descriptions into multiple languages" refers to a function for translating product descriptions into multiple languages.
[1782] A "generative AI model" is an artificial intelligence model trained to perform tasks such as text generation and translation.
[1783] A "prompt statement" is an instruction given to a generative AI model to perform a specific task.
[1784] "Means for saving translation results to a database" refers to the function of saving translation results generated by a generative AI model to a database.
[1785] "Means for sending translation results to the terminal" refers to a function that sends translation results stored in the database to the user's terminal.
[1786] "Means by which the terminal displays the translation results" refers to a function in which the user's terminal displays the received translation results on the screen.
[1787] This invention is a system that includes means for creating menus using a smart device, artificial intelligence means for generating product descriptions from images, means for artificial intelligence to generate translation tasks, means for improving the customer experience, means for artificial intelligence to explain questions when considering menus, means for making recommendations according to preferences, means for translating product descriptions into multiple languages, means for generating prompt sentences using a generation AI model, means for inputting prompt sentences and product descriptions into a generation AI model and obtaining translation results, means for saving translation results in a database, means for sending translation results to a terminal, and means for the terminal to display translation results.
[1788] Hardware and software to be used
[1789] Hardware:
[1790] Smart devices (smartphones, tablets, etc.)
[1791] Server (a computer with high-performance computing capabilities)
[1792] software:
[1793] Generative AI models (e.g., OpenAI's GPT-4)
[1794] Database Management System
[1795] Communication protocol (e.g., HTTP / HTTPS)
[1796] Data processing and data calculation
[1797] Product description generation:
[1798] The user takes a picture of a product using a smart device and sends the image to a server. The server uses artificial intelligence to generate a product description from the image. Specifically, it uses an image analysis algorithm to recognize the characteristics and features of the product and generates a product description based on that.
[1799] Prompt message generation:
[1800] The server generates prompt messages using a generative AI model. The prompt message specifies the target language for translation (e.g., English, Chinese, Korean).
[1801] Perform the translation task:
[1802] The server inputs the generated prompt text and product description into the AI model and retrieves the translation results. The AI model then performs multilingual translation based on the input text.
[1803] Saving and sending translation results:
[1804] The server saves the acquired translation results to a database. The saved translation results are sent to the user's terminal, which then displays the translation results.
[1805] Specific example
[1806] Specific example:
[1807] The user enters the Japanese product description: "This product uses high-quality materials."
[1808] The server generates the prompt message: "Translate the following product description into English: This product uses high-quality materials."
[1809] The server inputs the prompt text and product description into the generated AI model and retrieves the English translation result, "This product uses high-quality materials."
[1810] The server saves the translation results to a database and sends them to the terminal.
[1811] The device displays the translation result, and the user confirms it.
[1812] In this way, the server, terminal, and user work together to achieve multilingual translation of product descriptions.
[1813] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1814] Step 1:
[1815] The user enters the product description.
[1816] The user enters the product description in text format using a smart device. The entered product description is then sent from the smart device to the server.
[1817] Input: Product description text
[1818] Output: Product description text sent to the server
[1819] Step 2:
[1820] The server receives the product description.
[1821] The server receives the product description text sent by the user and temporarily stores it in memory.
[1822] Input: Product description text submitted by the user
[1823] Output: Product description text stored in memory
[1824] Step 3:
[1825] The server generates the prompt message.
[1826] The server generates prompt messages based on the languages to be translated. For example, if translation is needed for English, Chinese, and Korean, it will generate prompt messages corresponding to each language.
[1827] Input: Product description text, target language for translation
[1828] Output: Generated prompt message
[1829] Step 4:
[1830] The server inputs prompt text and product description into the generated AI model.
[1831] The server inputs the generated prompt text and product description text into the AI model. The AI model then performs translation based on the input text.
[1832] Input: Prompt text, product description text
[1833] Output: Data input to the generative AI model
[1834] Step 5:
[1835] The server retrieves the translation results.
[1836] The server retrieves translation results from the generative AI model. The retrieved translation results are separated by language.
[1837] Input: Data entered into the generating AI model
[1838] Output: Translated text
[1839] Step 6:
[1840] The server saves the translation results to the database.
[1841] The server saves the retrieved translation results to a database. When saving, it associates the original Japanese product description with the translated result.
[1842] Input: Translated text, original product description text
[1843] Output: Translation results stored in the database
[1844] Step 7:
[1845] The server sends the translation result to the terminal.
[1846] The server sends the saved translation results to the user's device. The transmitted data is provided in a format that the user can access.
[1847] Input: Translation results stored in the database
[1848] Output: Translation results sent to the terminal
[1849] Step 8:
[1850] The device displays the translation result.
[1851] The terminal receives the translation results sent from the server and displays them to the user. The user can review the displayed translation results and make corrections or re-translates as needed.
[1852] Input: Translation result sent from the server
[1853] Output: Translation results displayed on the terminal
[1854] (Application Example 2)
[1855] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1856] In modern brick-and-mortar stores, foreign tourists face language barriers when purchasing goods, making it difficult for them to understand product descriptions. Furthermore, visually impaired individuals and those with reading and writing difficulties also have limited means of understanding product descriptions. This can lead to a diminished customer experience and potentially impact sales.
[1857] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for creating a menu using a smartphone, an artificial intelligence means for generating product descriptions from photos on the smartphone, a means for the artificial intelligence to generate translation work, a means for improving the customer experience, a means for the artificial intelligence to explain questions when considering menus, a means for making recommendations according to preferences, and a means for translating product descriptions into multiple languages and displaying and playing them aloud in the user's set language. This makes it possible to make produ...
Claims
[Claim 1] A means for automatically recognizing the characteristics or features of a product from image data of the product taken with a smartphone by using an image analysis algorithm, A means for converting the aforementioned characteristics or features into a string and inputting it as a prompt to the generating AI model, thereby causing the generating AI model to generate a product description; A means for recognizing the user's emotions by analyzing the user's facial expressions or voice data using an emotion engine, A means for causing the generating AI model to generate a product description corresponding to the emotion by inputting a prompt sentence containing the emotion and the product description text into the generating AI model, A means for inputting instructions to the generating AI model to translate the product description corresponding to the aforementioned emotion into multiple languages, and for generating the translated result of the product description corresponding to the aforementioned emotion, A means for transmitting the translated results of the emotion-appropriate product description, translated into multiple languages, to the smartphone, A means for inputting the user's questions during menu consideration into the generating AI model, thereby causing the generating AI model to generate answers to those questions, A means for analyzing the user's past selection data using a recommendation algorithm and generating recommendation information, A system that includes this.
Citation Information
Patent Citations
Information processor, information processing method, and program
JP2015194858A
Presenting translations of text depicted in images
JP2017033585A
Homepage management device, homepage management system, and program
JP2021002184A
Device, system, and method for outputting user assistance information
JP2022124203A
Persona chatbot control method and system
JP2022180282A