System

The system addresses language barriers in menu understanding by using image and voice recognition, natural language processing, and recommendation algorithms to facilitate smooth ordering and communication in foreign restaurants.

JP2026022475APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123992
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Travelers face difficulties in understanding foreign language menus and ordering food due to language barriers, with existing translation tools being inefficient and inadequate for recommending drinks and desserts or facilitating smooth communication with waiters.

Method used

A system utilizing image recognition to extract menu text, natural language processing to describe dishes, recommendation algorithms for drinks and desserts, voice recognition for additional information, and a user interface to facilitate ordering, supported by network communication.

Benefits of technology

Enables users to easily understand and order appropriate dishes from foreign language menus, providing seamless communication and personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022475000001_ABST
    Figure 2026022475000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: image recognition means; natural language processing means; recommendation means; speech recognition means; image generation means; network communication means; and user interface means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When traveling abroad, it can be difficult to understand the menu and order at restaurants. This issue is particularly pronounced when travelers do not understand the language of the restaurant. Current technology, such as translation apps and dictionaries, is inefficient and does not provide a sufficient solution. Furthermore, language barriers can make it difficult to receive recommendations for drinks and desserts or to communicate smoothly with waiters. A system that solves these problems is needed. [Means for solving the problem]

[0005] In order to solve the above-mentioned problems, the present invention provides a system including the following means: An image recognition means allows a user to take an image of a menu and extract text from the image using OCR technology; a natural language processing means generates a description of the dishes from the extracted text; a recommendation means suggests suitable drinks and desserts based on the menu selected by the user; a voice recognition means analyzes the conversation between the user and the waiter and obtains additional information such as out-of-stock information and additional orders; an image generation means generates or obtains images of dishes, and a user interface means displays and allows menu items to be selected; and data is sent and received using a network communication means. In this way, a system is provided that supports users in understanding foreign language menus and ordering appropriate dishes.

[0006] "Image recognition means" is a technology that identifies characters and figures from images captured using a camera and converts them into digital data.

[0007] "Natural language processing means" is a technology that analyzes text data written in natural language, understands the meaning and intent, and generates appropriate information.

[0008] "Recommendation means" is a technology that suggests related products and services based on a user's preferences and choices.

[0009] "Speech recognition means" is a technology that analyzes speech as text data, understands the content, and generates appropriate responses and instructions.

[0010] "Image generation means" refers to a technique for generating or obtaining relevant images from text or other data.

[0011] "Network communication means" refers to the Internet or other communication means for transmitting and receiving data between devices.

[0012] "User interface means" refers to a screen or input device that allows a user to interact with and operate the system. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] As an embodiment of the present invention, the system program is implemented according to the following procedure: The system is mainly composed of a server, a terminal, and a user.

[0035] 1. Photographing the menu

[0036] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0037] 2. Sending images

[0038] Device: Creates a request to send the captured menu image to the server.

[0039] Terminal: Uploads menu images to the server via the network.

[0040] 3. Image analysis (OCR processing)

[0041] Server: Passes the received menu image to the image processing module.

[0042] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0043] 4. Natural Language Processing (NLP)

[0044] Server: Passes the extracted string to the natural language processing module.

[0045] Server: Generates a dish description from a string, for example "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[0046] 5. Image Generation

[0047] Server: Based on the string, retrieve related images using a web search API or database.

[0048] Server: Sends the acquired image and description to the device.

[0049] 6. Displaying the menu

[0050] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0051] User: Check the menu that appears and tap the desired dish.

[0052] 7. Sending Selected Information

[0053] Terminal: Sends information about the menu item selected by the user to the server.

[0054] 8. Generating Recommendations

[0055] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0056] Server: Receives the recommendation results from the modules and selects the drinks and desserts that are best suited to the user.

[0057] 9. Display of Recommendations

[0058] Terminal: Displays the drink and dessert recommendations received from the server to the user.

[0059] User: Review the recommendations and indicate their intent to place an order.

[0060] 10. Confirming your order and obtaining additional information

[0061] User: Tells the waiter their order.

[0062] Terminal: Passes the conversation between the user and the store clerk to a speech recognition module, which converts the speech into text.

[0063] Server: Parse the converted text and extract additional information, such as order confirmation and out-of-stock information.

[0064] 11. Amendments to Order Details

[0065] Server: If the store clerk's response includes information such as "out of stock," generate an alternative.

[0066] Terminal: Inform the user of the alternatives.

[0067] User: Review the alternatives, select a new menu option, and confirm.

[0068] Specific examples

[0069] 1. Photographing the menu

[0070] User: Take a photo of the menu with their smartphone.

[0071] 2. Sending images

[0072] Device: Creates a request to send the captured menu image to the server.

[0073] On your device: Upload the menu image to the server.

[0074] 3. Image analysis (OCR processing)

[0075] Server: Passes the menu image to the OCR processing module.

[0076] Server: Extract the string "Pasta alla Carbonara".

[0077] 4. Natural Language Processing (NLP)

[0078] Server: Passes the extracted string to the NLP module.

[0079] Server: Generates the description "Italian pasta featuring a creamy sauce and bacon" from "Pasta alla Carbonara."

[0080] 5. Image Generation

[0081] Server: Uses a web search API based on the string to retrieve related images.

[0082] Server: Sends the description and image to the device.

[0083] 6. Displaying the menu

[0084] Terminal: Display "Pasta alla Carbonara," "Italian pasta featuring a creamy sauce and bacon," and related images.

[0085] 7. Sending Selected Information

[0086] User: Select "Pasta alla Carbonara."

[0087] Terminal: Sends the selection information to the server.

[0088] 8. Generating Recommendations

[0089] Server: Recommend suitable drinks and desserts based on your selections.

[0090] Server: Recommend red wine or tiramisu.

[0091] 9. Display of Recommendations

[0092] Terminal: Display recommendations for red wine and tiramisu to the user.

[0093] 10. Confirming your order and obtaining additional information

[0094] User: Tells the waiter their order.

[0095] Terminal: The speech recognition module converts the conversation into text.

[0096] Server: Extracts additional information from the parsed text.

[0097] 11. Amendments to Order Details

[0098] Server: Based on the information "Pasta alla Carbonara is out of stock", generate "Pasta Bolognese" as an alternative.

[0099] Terminal: Inform the user of the alternatives.

[0100] User: Review the alternatives, select "Pasta Bolognese" and confirm.

[0101] In this way, a system is provided that allows users to easily understand menus written in foreign languages ​​and order appropriate dishes.

[0102] The processing flow will be explained below.

[0103] Step 1:

[0104] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0105] Step 2:

[0106] Device: Creates a request to send the captured menu image to the server.

[0107] Step 3:

[0108] Terminal: Uploads menu images to the server via the network.

[0109] Step 4:

[0110] Server: Passes the received menu image to the image processing module.

[0111] Step 5:

[0112] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0113] Step 6:

[0114] Server: Passes the extracted text data to the natural language processing module.

[0115] Step 7:

[0116] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[0117] Step 8:

[0118] Server: Obtains images related to the dish along with the generated description.

[0119] Step 9:

[0120] Server: Sends the description and image to the device.

[0121] Step 10:

[0122] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0123] Step 11:

[0124] User: Review the menu items displayed and tap to select the desired dish.

[0125] Step 12:

[0126] Terminal: Sends information about the menu item selected by the user to the server.

[0127] Step 13:

[0128] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0129] Step 14:

[0130] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[0131] Step 15:

[0132] Server: Sends the recommendation results to the device.

[0133] Step 16:

[0134] Terminal: Display recommended drinks and desserts to the user.

[0135] Step 17:

[0136] User: Review the recommendations and indicate their intent to place an order.

[0137] Step 18:

[0138] User: Inform the store clerk of the confirmed order.

[0139] Step 19:

[0140] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[0141] Step 20:

[0142] Device: Sends captured audio data to the server.

[0143] Step 21:

[0144] Server: Passes the voice data to the voice recognition module and converts it into text data.

[0145] Step 22:

[0146] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[0147] Step 23:

[0148] Server: Based on the analysis results, generate alternatives as needed.

[0149] Step 24:

[0150] Server: Sends alternatives to the device.

[0151] Step 25:

[0152] Terminal: Present alternatives to the user.

[0153] Step 26:

[0154] User: Review alternatives, select new menu, confirm.

[0155] Example 1

[0156] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0157] In modern society, many people face significant challenges in understanding menus written in a foreign language and ordering the appropriate food. This problem is particularly pronounced for tourists and people unfamiliar with the foreign language, leading to ordering errors and confusion. Furthermore, selecting the right drink or dessert presents similar challenges. The present invention aims to solve these problems and provide a smooth ordering process.

[0158] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0159] In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, and user interface means, which enable the server to extract text from menus written in a foreign language to generate dish descriptions, recommend related drinks and desserts, convert voice dialogue into text, and display this information on the user interface.

[0160] "Image recognition means" refers to a device or program that performs processing to extract text from an image captured using a camera.

[0161] "Natural language processing means" refers to a device or program that analyzes the extracted text and generates descriptions of dishes and related information.

[0162] A "recommendation means" is a device or program for recommending related food and beverage items based on a menu item selected by a user.

[0163] "Speech recognition means" refers to a device or program for converting the dialogue between the user and the store clerk from voice to text.

[0164] The "image generating means" refers to a device or program for obtaining an image related to a menu item and displaying it to the user.

[0165] "Network communication means" refers to a communication device or program for transmitting and receiving data between a terminal and a server.

[0166] "User interface means" refers to a device or program that provides a display screen and input means for a user to operate the system.

[0167] The system of the present invention is composed of a user, a terminal, and a server. By using this system, the user can understand menus written in a foreign language and order the appropriate food.

[0168] Specific examples of hardware and software used

[0169] Hardware

[0170] 1. Camera-equipped devices (smartphones, tablets)

[0171] 2. Server (Cloud server, local server)

[0172] software

[0173] 1. OCR module (Google Cloud Vision API, Tesseract OCR)

[0174] 2. Natural language processing module (IBM Watson NLP, Google Cloud Natural Language API)

[0175] 3. Recommendation module (Collaborative Filtering, Content-Based Filtering algorithms)

[0176] 4. Speech Recognition Module (Google Speech-to-Text API, IBM Watson Speech to Text)

[0177] 5. Network communication module (HTTP / HTTPS communication library)

[0178] 6. User Interface Module (React Native, Flutter)

[0179] Specific operation of the system

[0180] 1. Photographing the menu

[0181] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[0182] 2. Sending images

[0183] The device encodes the captured menu image and creates a request to send to the server, specifically uploading the image to the server's endpoint using an HTTP POST request.

[0184] 3. Image analysis (OCR processing)

[0185] The server passes the received image to an OCR module, which extracts text from the image, for example, the string "Pasta alla Carbonara."

[0186] 4. Natural Language Processing (NLP)

[0187] The server passes the extracted text to a natural language processing module to generate a description of the dish, for example, "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[0188] 5. Image Generation

[0189] The server uses a web search API to retrieve related images based on the results of natural language processing, and sends the images and descriptions together to the device.

[0190] 6. Displaying the menu

[0191] The terminal displays the menu items, explanations, and related images received from the server on the user interface, providing information in a format that is easy for the user to select.

[0192] 7. Sending Selected Information

[0193] The user selects the dish of interest and confirms the information.

[0194] The terminal transmits the user's selection information to the server.

[0195] 8. Generating Recommendations

[0196] Based on the user's selection, the server uses a recommendation module to recommend suitable drinks and desserts, for example, red wine or tiramisu that go well with "Pasta alla Carbonara."

[0197] 9. Display of Recommendations

[0198] The terminal displays the information recommended by the server to the user, who can then confirm the recommendation and place an order.

[0199] 10. Confirming your order and obtaining additional information

[0200] The user tells the store clerk their order.

[0201] The terminal uses a voice recognition module to convert the conversation between the user and the store clerk into text and send it to the server.

[0202] The server analyzes the dialogue and extracts any additional information needed.

[0203] 11. Amendments to Order Details

[0204] The server generates alternative suggestions based on information such as "Pasta alla Carbonara is out of stock."

[0205] The terminal notifies the user of the alternatives, and the user selects and confirms the new menu.

[0206] Specific examples

[0207] 1. Menu photography:

[0208] A user takes a photo of a restaurant menu using a smartphone.

[0209] 2. Sending images:

[0210] The device encodes the captured menu image and sends it to the server's endpoint.

[0211] 3. Image analysis (OCR processing):

[0212] The server uses an OCR module to extract the string "Pasta alla Carbonara".

[0213] 4. Natural Language Processing (NLP):

[0214] The server generates a description for "Pasta alla Carbonara": "Italian pasta featuring a creamy sauce and bacon."

[0215] 5. Image generation:

[0216] The server uses a web search API to retrieve an image of "Pasta alla Carbonara."

[0217] 6. Display Menu:

[0218] The device displays "Pasta alla Carbonara," an Italian pasta dish featuring a creamy sauce and bacon, along with related images.

[0219] 7. Sending Selected Information:

[0220] The user selects "Pasta alla Carbonara" and the terminal transmits the selection information to the server.

[0221] 8. Generating recommendations:

[0222] The server will recommend suitable drinks and desserts, such as red wine and tiramisu.

[0223] 9. Display of Recommendations:

[0224] The device displays recommendations for red wine and tiramisu.

[0225] 10. Confirm your order and obtain additional information:

[0226] The user tells the store clerk their order.

[0227] The device uses a voice recognition module to convert the conversation into text.

[0228] 11. Order Modifications:

[0229] Based on the information that "Pasta alla Carbonara is out of stock," the server suggests "Pasta Bolognese."

[0230] The device presents the alternatives and the user selects "Pasta Bolognese."

[0231] This system makes it easier for users to understand foreign language menus and order the right food.

[0232] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0233] Step 1:

[0234] A user opens the camera on their smartphone or tablet and takes a picture of a restaurant menu. The input is the camera's photo image, and the output is an image file of the menu. Specifically, the user opens the camera app and presses the shutter button to capture the image.

[0235] Step 2:

[0236] The device encodes the captured menu image to send to the server and creates an HTTP POST request. The input is the menu image file, and the output is an HTTP request. Specifically, the device encodes the image into JPEG format and sends it to the server's endpoint (e.g., https: / / example.com / upload).

[0237] Step 3:

[0238] The server passes the received menu image to the OCR module, which extracts text from the image. The input is the image data of the menu, and the output is the extracted string. Specifically, the server sends the image data to the Google Cloud Vision API and obtains the text "Pasta alla Carbonara."

[0239] Step 4:

[0240] The server passes the extracted string to a natural language processing module to generate a description of the dish. The input is the extracted string, and the output is a description of the dish. Specifically, the server uses the NLP module to generate a description such as "Italian pasta with a creamy sauce and bacon" from "Pasta alla Carbonara."

[0241] Step 5:

[0242] The server uses a web search API to retrieve related images based on the results of natural language processing. The input is a description of the dish, and the output is an image of the dish. Specifically, the server retrieves images related to "Pasta alla Carbonara" using the Google Image Search API and selects the appropriate image.

[0243] Step 6:

[0244] The server compiles the obtained description and image in JSON format and sends it to the device. The input is the description and image of the dish, and the output is a JSON response. Specifically, the server generates JSON data such as { "name": "Pasta alla Carbonara", "description": "Italian pasta characterized by a creamy sauce and bacon", "image_url": "https: / / example.com / image.jpg"} and sends it to the device.

[0245] Step 7:

[0246] The device parses the JSON data received from the server and displays menu items, descriptions, and images on the user interface. The input is the JSON response received from the server, and the output is the user interface display. Specifically, the device uses an application to parse the received data and display it on the screen.

[0247] Step 8:

[0248] The user selects the dish of interest from the displayed menu, and the device sends the selection to the server. The input is the user's selection, and the output is an HTTP request. Specifically, the user taps "Pasta alla Carbonara" on the touchscreen, and the device sends the selection to the server.

[0249] Step 9:

[0250] The server uses a recommendation module to recommend related drinks and desserts based on the user's selection. The input is the user's selection, and the output is a list of recommended items. Specifically, the server uses a Collaborative Filtering algorithm to recommend red wine and tiramisu to go with "Pasta alla Carbonara."

[0251] Step 10:

[0252] The server compiles the generated recommendation results in JSON format and sends them to the device. The input is a list of recommended items, and the output is a JSON response. Specifically, the server generates JSON data such as { "recommendations": ["red wine", "tiramisu"]} and sends it to the device.

[0253] Step 11:

[0254] The device displays the recommendation information received from the server on the user interface, and the user confirms the displayed recommendation selection. The input is the JSON response received from the server, and the output is the display of the user interface and the user's selection action. Specifically, the device displays the received recommendation information, and the user taps the confirm button.

[0255] Step 12:

[0256] The user tells the store clerk their order, and the device converts the conversation into text using a voice recognition module and sends it to the server. The input is the voice conversation between the user and the store clerk, and the output is text data. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text.

[0257] Step 13:

[0258] The server analyzes the textual content of the conversation and obtains additional information. The input is text data, and the output is the analysis result. Specifically, the server analyzes the text and extracts information such as confirmation of order details and out-of-stock information.

[0259] Step 14:

[0260] The server generates alternatives as needed and notifies the terminal. The input is the analysis result information, and the output is the alternative information. In concrete terms, the server might suggest "Pasta Bolognese" based on the information "Pasta alla Carbonara is out of stock."

[0261] Step 15:

[0262] The device notifies the user of the alternatives, and the user selects and confirms the new menu. The input is the information about the alternatives, and the output is the user's selection action. Specifically, the device displays the alternatives, and the user selects a new menu and taps the confirm button.

[0263] (Application example 1)

[0264] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0265] In modern factories, many pieces of equipment and products use advanced technology, making it difficult for workers to grasp all the information and work efficiently. In particular, not being able to instantly obtain operating procedures and maintenance information related to equipment and products can lead to reduced production efficiency and the risk of operating errors. In addition, when working in a multilingual environment, language barriers become an additional obstacle. Therefore, there is a growing need for systems that can provide detailed information about equipment and products in the factory in real time.

[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0267] In this invention, the server includes an image recognition unit, a natural language processing unit, and a character analysis unit. This allows real-time access to information about equipment and products in a factory using smart glasses or a head-mounted display. This allows workers to instantly access detailed operation guides and maintenance information about the equipment and products, improving work efficiency. Furthermore, by utilizing a generative AI model with prompt sentences, multilingual support is possible, enabling information provision across language barriers.

[0268] "Image recognition means" refers to a technology that uses a camera or other photographic device to acquire image data and extract specific features from the image.

[0269] "Natural language processing means" refers to technology that analyzes character strings and sentences and converts or interprets them into natural language that humans can understand.

[0270] A "recommendation tool" is a system that automatically recommends appropriate products and services based on a user's past behavior and choices.

[0271] "Speech recognition means" is a technology that analyzes speech and converts it into text data.

[0272] "Image generation means" refers to technology that creates new images based on data and information.

[0273] "Network communication means" refers to technology for sending and receiving data over the Internet or other communication networks.

[0274] "User interface means" refers to an interface such as a screen or input device that allows a user to interact with the system.

[0275] "Information acquisition means" refers to technology for acquiring information about equipment and products within a factory using smart glasses or a head-mounted display.

[0276] "Character analysis means" is a technology that uses OCR technology to analyze photographed character strings and convert them into digital text.

[0277] A "prompt" is a short command that a generative AI model receives as input and is an instruction to generate specific information.

[0278] The following system is required to implement this invention: The system is mainly composed of a server, a terminal, and a user.

[0279] The server includes an image recognition means, a natural language processing means, a character analysis means, a recommendation means, a voice recognition means, an image generation means, a network communication means, a generative AI model using prompt sentences, and the like.

[0280] The terminal is equipped with smart glasses or a head-mounted display, a camera, a microphone, and a user interface means, and is operated by the user.

[0281] The server first analyzes the label image of the equipment or product sent from the terminal using image recognition, which includes character analysis using OCR technology. Specifically, it uses Tesseract OCR to extract character strings from the image.

[0282] The extracted text is analyzed using natural language processing to generate an appropriate description, which can be translated into multiple languages ​​as needed using translation services such as Google Translate API.

[0283] The generated description and related images are sent to the terminal via the network communication means and displayed to the user by the user interface means. For example, if the character string "product 12345" is extracted, the details displayed will read, "Product 12345: This part has been machined using a high-precision grinder. Handle with care. Wear protective equipment and follow the operating instructions when using."

[0284] In addition, the recommendation means recommends appropriate operation guides and maintenance procedures based on the user's past selections and behavioral data, thereby further improving work efficiency.

[0285] An example of using a generative AI model with prompts is the following prompt:

[0286] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[0287] "Product 12345: This part has been processed using a high-precision grinder."

[0288] The system configuration and processing procedures described above enable detailed information on equipment and products within the factory to be provided in real time, improving work efficiency. In addition, the system's multilingual support makes it suitable for international work environments.

[0289] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0290] Step 1:

[0291] A user takes a picture of the labels of equipment or products in a factory using a camera in smart glasses or a head-mounted display. Here, the input is the label image, and the output is the image data captured by the camera.

[0292] Step 2:

[0293] The terminal sends the captured label image to the server. The input is label image data, and the output is image data sent to the server via the network.

[0294] Step 3:

[0295] The server uses image recognition and OCR technology to extract text from the received label image. Specifically, it uses Tesseract OCR. The input is the label image data, and the output is the extracted text.

[0296] Step 4:

[0297] The server passes the extracted string to a natural language processing means to generate an appropriate description. It translates it into multiple languages ​​as needed using NLP technology or the Google Translate API. The input is the extracted string, and the output is the generated description.

[0298] Step 5:

[0299] The server uses an image generation means to retrieve associated images along with the generated description, for example, the associated images are retrieved from a database or the Internet, where the input is the description and the output is the associated images.

[0300] Step 6:

[0301] The server transmits the generated explanatory text and related images to the terminal using a network communication means. The input is the generated explanatory text and related images, and the output is the data transmitted to the terminal.

[0302] Step 7:

[0303] The terminal displays the received explanatory text and image to the user through a user interface means. The input is the explanatory text and image data, and the output is the information displayed on the user interface.

[0304] Step 8:

[0305] The user checks the displayed information and performs operations to obtain more detailed information as necessary. For example, if a detailed operation guide is required, a prompt sentence is generated and the request is sent from the terminal to the server. The input is the user's operation, and the output is the generated prompt sentence and its request.

[0306] Step 9:

[0307] The server uses a generative AI model to generate detailed operation guides based on the prompt sentences and provides them to the user. The input is the prompt sentence, and the output is the detailed operation guides.

[0308] For specific actions, the following example prompt sentence is used:

[0309] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[0310] "Product 12345: This part has been processed using a high-precision grinder."

[0311] Through the above steps, users can obtain detailed information about the equipment and products in the factory in real time, improving work efficiency.

[0312] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0313] As an embodiment of the present invention, the specific implementation procedure of a system using a user, a terminal, a server, and an emotion engine is described below. This system assists users in the process of understanding foreign language menus and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[0314] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to the server. The server then passes the image to an image processing module, which uses OCR technology to extract the text from the image. For example, the string "Pasta alla Carbonara" is extracted.

[0315] The extracted text data is then passed to a natural language processing module. The server generates a description of the dish from this text data. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via a web search API or database. The server then sends the retrieved description and images to the device.

[0316] The device displays the received menu items, descriptions, and images to the user. The user checks the displayed menu and taps on the desired dish to select it. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts. The server sends the recommendation results to the device, which displays them to the user. The user checks the recommendations and confirms the order.

[0317] Furthermore, the system incorporates an emotion engine, which recognizes emotions from the user's facial expressions and voice. For example, if the user expresses a confused expression, the emotion engine analyzes the information. This emotion information is sent to the server, and menu suggestions and recommendations are adjusted based on the user's emotions. For example, if the user is confused, the system adjusts to provide more detailed explanations and additional information.

[0318] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion engine will recognize the emotion and the server will provide additional detailed explanations of the recommendation and alternatives based on the emotion.

[0319] After the order is confirmed, the user conveys the order to the store clerk. The voice recognition module is activated again, capturing the conversation between the user and the store clerk. The voice data is sent to the server, where it is converted into text data by the voice recognition module. The converted text data is then analyzed to extract additional information, such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user. The user then confirms the alternative options, selects a new menu item, and confirms the selection.

[0320] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[0321] The processing flow will be explained below.

[0322] Step 1:

[0323] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0324] Step 2:

[0325] Device: Creates a request to send the captured menu image to the server.

[0326] Step 3:

[0327] Terminal: Uploads menu images to the server via the network.

[0328] Step 4:

[0329] Server: Passes the received menu image to the image processing module.

[0330] Step 5:

[0331] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0332] Step 6:

[0333] Server: Passes the extracted text data to the natural language processing module.

[0334] Step 7:

[0335] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[0336] Step 8:

[0337] Server: Obtains images related to the dish along with the generated description.

[0338] Step 9:

[0339] Server: Sends the description and image to the device.

[0340] Step 10:

[0341] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0342] Step 11:

[0343] User: Review the menu items displayed and tap to select the desired dish.

[0344] Step 12:

[0345] Terminal: Sends information about the menu item selected by the user to the server.

[0346] Step 13:

[0347] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0348] Step 14:

[0349] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[0350] Step 15:

[0351] Server: Sends the recommendation results to the device.

[0352] Step 16:

[0353] Terminal: Display recommended drinks and desserts to the user.

[0354] Step 17:

[0355] User: Review the recommendations and indicate their intent to place an order.

[0356] Step 18:

[0357] Device: Activates the emotion engine to recognize emotions from the user's facial expressions and voice.

[0358] Step 19:

[0359] Terminal: The emotion engine analyzes the user's emotions and sends the results to the server.

[0360] Step 20:

[0361] Server: Receives user emotion data and adjusts recommendations and menu suggestions.

[0362] Step 21:

[0363] Server: Generates additional information and alternative suggestions based on emotion data and sends them to the device.

[0364] Step 22:

[0365] On the device: Display additional information or alternative suggestions to the user based on their emotions.

[0366] Step 23:

[0367] User: Review the proposed information and confirm the order.

[0368] Step 24:

[0369] User: Tells the waiter their order.

[0370] Step 25:

[0371] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[0372] Step 26:

[0373] Device: Sends captured audio data to the server.

[0374] Step 27:

[0375] Server: Passes the voice data to the voice recognition module and converts it into text data.

[0376] Step 28:

[0377] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[0378] Step 29:

[0379] Server: Based on the analysis results, generate alternatives as needed.

[0380] Step 30:

[0381] Server: Sends alternatives to the device.

[0382] Step 31:

[0383] Terminal: Present alternatives to the user.

[0384] Step 32:

[0385] User: Review alternatives, select new menu, confirm.

[0386] Example 2

[0387] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0388] For many users, understanding menus in a foreign language and ordering the appropriate dishes is a difficult task. Furthermore, if recommendations are not appropriate based on the user's emotions and preferences, the user experience cannot be improved. To solve these challenges, a system is needed that not only extracts menu text and provides easy-to-understand information, but also makes recommendations that adapt to the user's emotional state.

[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0390] In this invention, the server includes image recognition means, natural language processing means, recommendation means, and emotion recognition means. This allows users to easily understand menus in foreign languages ​​and order appropriate dishes. Recommendations are also made that are adapted to the user's emotional state, allowing the user to enjoy personalized services.

[0391] "Image recognition means" refers to technology for analyzing an image and extracting specific information from its content.

[0392] "Natural language processing" refers to technology for understanding, generating, and manipulating human language.

[0393] "Recommendation means" refers to technology for recommending appropriate items based on a user's preferences and history.

[0394] "Speech recognition means" refers to technology for analyzing voice data and converting it into text data.

[0395] "Image generation means" refers to technology for generating new images based on input data.

[0396] "Emotion recognition means" refers to technology for analyzing and recognizing emotions from a user's facial expressions and voice.

[0397] "Network communication means" refers to technology for communicating over a network to send and receive data.

[0398] "User interface means" refers to technology that provides an interface for exchanging information between a user and a system.

[0399] The present invention relates to a system that assists users in the process of understanding menus in a foreign language and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[0400] Hardware and Software Configuration

[0401] Required Hardware

[0402] User's device: A smartphone or tablet equipped with a camera is used.

[0403] Server: A high-performance server is used for menu image analysis and recommendation processing.

[0404] Required software

[0405] Image Recognition Method: OCR technology is used to extract text from images, specifically Google Cloud Vision API.

[0406] Natural language processing tools: A generative AI model (e.g., GPT-4) is used to generate dish descriptions from the extracted text data.

[0407] Recommendation method: Use a recommendation algorithm to recommend drinks and desserts related to the dish.

[0408] Emotion Recognition: Emotion recognition software is used to analyze and recognize emotions from the user's facial expressions and voice.

[0409] Speech recognition means: Use speech recognition technology (e.g., Google Speech-to-Text API) to capture the conversation between the user and the store clerk as voice data and convert it into text data.

[0410] Image generation method: Use a web search API (e.g., Bing Image Search API) to obtain related images.

[0411] Network communication means: Uses communication technology to send and receive data between the user terminal and the server.

[0412] User interface means: Provides an interface for exchanging information between the user and the system.

[0413] Example of a system

[0414] 1. Menu Parsing Prompt Example

[0415] "Analyze a restaurant menu image and generate dish names and descriptions."

[0416] 2. Recommendation prompt examples

[0417] "Recommend drinks and desserts to go with the selected dish."

[0418] 3. Example of emotion-responsive prompt

[0419] "If the user looks confused, provide a detailed explanation."

[0420] System Operation

[0421] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to a server. The server then analyzes the received image using the Google Cloud Vision API and extracts the text from the image using OCR technology. For example, the string "Pasta alla Carbonara" is extracted.

[0422] The extracted text data is then passed to a generative AI model such as GPT-4 to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via the Bing Image Search API. The server then sends the retrieved description and image to the device, which then displays them to the user.

[0423] The user checks the displayed menu items, descriptions, and images, and taps to select the desired dish. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts, and sends the recommendation results to the device. The device then displays them to the user.

[0424] The app also has a built-in emotion engine that recognizes emotions from the user's facial expressions and voice. For example, if the user shows a confused expression, the emotion engine analyzes the information and sends it to the server. Based on this emotion information, the server adjusts the service to provide detailed explanations and additional information to the user.

[0425] Finally, the user confirms the order and conveys it to the store clerk. The voice recognition module is activated, capturing the conversation between the user and the store clerk, and the voice data is sent to the server. The server's voice recognition module converts the voice data into text data, which is then analyzed to extract additional information such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user.

[0426] In this way, the present invention allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's feelings.

[0427] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0428] Step 1:

[0429] Users take a photo of a restaurant menu using the camera on their smartphone or tablet.

[0430] Input: Menu Image

[0431] Specific actions: The user opens the camera app, points the lens at the menu, and presses the shutter button to take a photo.

[0432] Output: Menu image captured

[0433] Step 2:

[0434] The terminal transmits the captured menu image to the server.

[0435] Input: Menu image taken

[0436] Specific operation: After taking a photo, the image will be automatically uploaded and sent.

[0437] Output: Menu image sent to the server

[0438] Step 3:

[0439] The server analyzes the received menu image using OCR technology and extracts the text within the image.

[0440] Input: Menu image sent to the server

[0441] Data processing: Call the Google Cloud Vision API to extract character codes from images

[0442] What it does: The API analyzes the text in the menu image and extracts text such as "Pasta alla Carbonara."

[0443] Output: Extracted text data

[0444] Step 4:

[0445] The server passes the extracted text data to a natural language processing module to generate a description of the dish.

[0446] Input: Extracted text data

[0447] Data processing: Calling GPT-4 to generate explanatory text from text data

[0448] What it does: GPT-4 takes the text "Pasta alla Carbonara" and generates a description: "Italian pasta with a creamy sauce and bacon."

[0449] Output: Generated dish description

[0450] Step 5:

[0451] The server retrieves related images via a web search API.

[0452] Input: Description of the generated dish

[0453] Data processing: Call the Bing Image Search API to get related images

[0454] What it does: Uses the Bing Image Search API to search and retrieve relevant images based on the description of "Pasta alla Carbonara."

[0455] Output: Captured image

[0456] Step 6:

[0457] The server transmits the explanatory text and the image to the terminal, which displays them to the user.

[0458] Input: Generated description and retrieved image

[0459] Specific operation: Data is sent from the server to the device, and the device displays a pop-up explanation and image on the screen.

[0460] Output: Description and image displayed to the user

[0461] Step 7:

[0462] The user taps on the desired dish to select it, and the device sends the selection information to the server.

[0463] Input: Displayed description and image

[0464] What happens: The user taps "Pasta alla Carbonara" and the selection is sent to the server.

[0465] Output: Selections sent to the server

[0466] Step 8:

[0467] The server uses a recommendation module to recommend related drinks and desserts and transmits them to the terminal.

[0468] Input: Selections sent to the server

[0469] Data processing: Use recommendation algorithms to select relevant drinks and desserts

[0470] Specific operation: The server selects red wine and tiramisu, and the recommended results are displayed on the device screen.

[0471] Output: Recommendation results displayed to the user

[0472] Step 9:

[0473] The emotion engine recognizes emotions from the user's facial expressions and voice, and transmits the emotion information to the server.

[0474] Input: User's facial expressions and voice

[0475] Data processing: Using emotion recognition software, analyze and recognize the user's emotions.

[0476] Specific operation: When the user makes a confused expression, the camera analyzes the expression and sends the emotional information to the server, which then displays a detailed explanation.

[0477] Output: Additional explanation based on emotion information

[0478] Step 10:

[0479] The user confirms the final order and communicates it to the store attendant. The voice recognition module captures the conversation and sends it to the server.

[0480] Input: User and store clerk conversation

[0481] Data processing: Using voice recognition technology, converting voice data into text data

[0482] What it does: The user verbally places an order, and the speech is sent to the server, which converts it into text. The server then provides out-of-stock information and alternative suggestions.

[0483] Output: Textualized dialogue, out-of-stock information, and alternatives

[0484] Through this series of processes, users can easily understand menus in foreign languages ​​and order the most suitable dishes based on their emotions.

[0485] (Application example 2)

[0486] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0487] In existing systems, users may have difficulty understanding foreign language menus and selecting appropriate dishes, and they are unable to provide personalized recommendations based on the user's emotions. The present invention aims to solve these problems, enabling users to easily understand foreign language menus and providing personalized services based on the user's emotions.

[0488] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means. This allows the user to easily understand foreign language menus and further receive personalized recommendations based on their emotions.

[0489] "Image recognition means" is a device that extracts information from a captured image and analyzes its content.

[0490] "Natural language processing means" is a device that analyzes extracted text data, converts it into natural language format, and understands and processes it.

[0491] A "recommendation means" is a device that makes relevant suggestions or recommendations to the user based on the analyzed information.

[0492] A "voice recognition means" is a device that analyzes voice data and converts it into a linguistic text format.

[0493] The "image generating means" is a device that generates an image based on the analyzed information and generated data.

[0494] A "network communication means" is a communication device for transmitting and receiving data between systems.

[0495] "User interface means" refers to a device that allows a user to operate and interact with the system.

[0496] The "emotion analysis means" is a device that recognizes and analyzes the user's emotions from their facial expressions and tone of voice.

[0497] The present invention provides a system that assists users in understanding menus written in a foreign language and ordering appropriate dishes. The system includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means.

[0498] First, a user takes a photo of a restaurant menu using a smartphone or tablet device. The image taken with the device's camera is sent from the device to a server via network communication means. The server then uses image recognition means to extract text from the received image. For example, the text "Pasta alla Carbonara" is extracted.

[0499] Next, the extracted text data is analyzed using natural language processing to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta characterized by a creamy sauce and bacon." At the same time, related images are retrieved via a web search API or database. The server then sends the generated description and the retrieved images to the device via the network.

[0500] The terminal uses a user interface means to display the received menu items, descriptions, and images to the user. The user checks the displayed menu and selects the desired dish. Information about the selected menu item is again sent from the terminal to the server. The server uses a recommendation means to recommend additional related drinks and desserts. The recommendation results are sent to the terminal and displayed to the user.

[0501] The system also incorporates an emotion analysis mechanism that recognizes emotions from the user's facial expressions and voice and adjusts menu suggestions and recommendations based on that information. For example, if the user expresses confusion, the emotion analysis mechanism sends that information to the server, which then adjusts the system to provide detailed explanations and additional information.

[0502] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion analyzer recognizes the emotion and the server provides additional detailed explanations of the recommendation and alternatives.

[0503] After the order is confirmed, the user tells the store clerk what they want to order. At this time, the voice recognition means is activated and the conversation between the user and the store clerk is captured. The voice data is sent back to the server and converted into text data by the voice recognition means. The converted text data is analyzed to extract information such as out-of-stock items and additional orders. If necessary, alternative options are generated and notified to the user. The user then selects a new menu item and confirms the order.

[0504] An example prompt is, "Use the EmotionRecognition library to analyze emotions from the image menu.jpg and generate personalized dish descriptions based on that."

[0505] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[0506] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0507] Step 1:

[0508] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[0509] Input: Menu Image

[0510] Output: Raw image data in the device

[0511] Specific operation: A user opens the device's camera app and takes a picture of a restaurant menu. After taking the picture, the image data is temporarily stored in the device's memory.

[0512] Step 2:

[0513] The terminal transmits the photographed menu image to the server using the network communication means.

[0514] Input: Raw image data in the device

[0515] Output: Image data sent to the server

[0516] How it works: The device sends image data to the server via Wi-Fi or mobile data using the HTTP or HTTPS protocol.

[0517] Step 3:

[0518] The server uses image recognition means to extract text from the received image.

[0519] Input: Image data sent to the server

[0520] Output: Extracted text data

[0521] How it works: Optical Character Recognition (OCR) software on the server analyzes the image and extracts text information, which is then stored in a database on the server, such as "Pasta alla Carbonara."

[0522] Step 4:

[0523] The server uses natural language processing means to generate a description of the dish from the extracted text data.

[0524] Input: Extracted text data

[0525] Output: The generated dish description

[0526] What it does: A server-based natural language processing library analyzes the text data and generates a description of the corresponding dish, such as "Italian pasta with a creamy sauce and bacon."

[0527] Step 5:

[0528] The server retrieves related images via a web search API or database.

[0529] Input: Extracted text data

[0530] Output: Related images

[0531] Specific operation: The server calls the web search API, searches for relevant images based on the text data, selects the most appropriate image from the search results, and downloads it.

[0532] Step 6:

[0533] The server transmits the generated description and image to the terminal using a network communication means.

[0534] Input: Generated description and associated image

[0535] Output: Description and image sent to the device

[0536] Specific operation: The server uses the HTTP or HTTPS protocol to send the description and image to the terminal.

[0537] Step 7:

[0538] The terminal uses the user interface means to display the received menu items, explanations, and images to the user.

[0539] Input: Received description and image

[0540] Output: Menu items, descriptions, and images displayed in the user interface

[0541] Specific behavior: The device uses the layout template to properly arrange the displayed information on the screen, allowing the user to visually confirm the displayed information.

[0542] Step 8:

[0543] The user checks the displayed menu items and taps to select the desired dish.

[0544] Input: User tap input

[0545] Output: Information about the selected menu item

[0546] Specific operation: The information of the menu item selected by the user is stored in the terminal's memory in preparation for the next communication step.

[0547] Step 9:

[0548] The terminal transmits information about the selected menu item to the server using the network communication means.

[0549] Input: Information of the selected menu item

[0550] Output: Selections sent to the server

[0551] Specific operation: The device sends the selection information to the server via Wi-Fi or mobile data communication.

[0552] Step 10:

[0553] The server uses a recommendation means to recommend related drinks and desserts.

[0554] Input: Information of the selected menu item

[0555] Output: Recommended drinks and desserts

[0556] Specific operation: The server's recommendation engine searches for items related to the selected menu and generates a recommendation list.

[0557] Step 11:

[0558] The server transmits the recommendation results to the terminal using a network communication means.

[0559] Input: Recommended drink and dessert information

[0560] Output: Recommendation results sent to the device

[0561] Specific operation: The server uses HTTP or HTTPS protocol to send the recommendation results to the terminal.

[0562] Step 12:

[0563] The terminal uses the user interface means to display the recommendation results to the user.

[0564] Input: Received recommendation results

[0565] Output: Recommendation results displayed in the user interface

[0566] Specific behavior: The recommended drinks and desserts will be displayed on the device screen for the user to review.

[0567] Step 13:

[0568] The server uses emotion analysis means to recognize emotions from the user's facial expressions and voice.

[0569] Input: User's facial expressions or voice data

[0570] Output: Recognized emotion information

[0571] Specific operation: The server uses a facial expression recognition library and a voice analysis library to analyze the user's emotions and saves the recognition results in a database.

[0572] Step 14:

[0573] The server adjusts menu suggestions and recommendations based on the recognized emotional information.

[0574] Input: Recognized emotion information

[0575] Output: Tailored menu suggestions and recommendations

[0576] Specific operation: Based on the results of the sentiment analysis, the server generates detailed information and alternatives and provides them to the user.

[0577] Step 15:

[0578] The user finally confirms and confirms the menu item.

[0579] Input: User confirmed input

[0580] Output: Confirmed order information

[0581] Specific operation: When the user selects a menu item and taps the confirmation button, the order information is finally captured and saved in the device's memory.

[0582] Step 16:

[0583] When the terminal communicates the order information to the store clerk, the voice recognition means captures the content of the conversation.

[0584] Input: Voice data of the conversation between the user and the store clerk

[0585] Output: Captured audio data

[0586] Specific operation: The device's microphone records the conversation between the user and the store clerk and saves it as audio data.

[0587] Step 17:

[0588] The server receives the voice data and converts it into text data using a voice recognition means.

[0589] Input: Captured audio data

[0590] Output: Converted text data

[0591] Specific operation: A speech recognition library on the server analyzes the voice data and generates corresponding text.

[0592] Step 18:

[0593] The server analyzes the converted text data and extracts out-of-stock information and additional orders.

[0594] Input: Converted text data

[0595] Output: Out-of-stock information and reorder notifications

[0596] Specific operation: The server analyzes the text data, extracts the necessary information, and generates alternative menus and additional order information as necessary.

[0597] Step 19:

[0598] The server informs the user of the alternatives and the user selects a new menu.

[0599] Input: Out-of-stock information and reorder notifications

[0600] Output: New menu selection information

[0601] Specific operation: The user confirms the alternatives and sends the newly selected menu information from the terminal to the server.

[0602] In this way, the present invention helps users understand foreign language menus and provides personalized services based on emotions.

[0603] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0604] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0605] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0606] [Second embodiment]

[0607] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0608] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0609] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0610] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0611] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0612] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0613] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0614] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0615] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0616] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0617] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0618] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0619] As an embodiment of the present invention, the system program is implemented according to the following procedure: The system is mainly composed of a server, a terminal, and a user.

[0620] 1. Photographing the menu

[0621] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0622] 2. Sending images

[0623] Device: Creates a request to send the captured menu image to the server.

[0624] Terminal: Uploads menu images to the server via the network.

[0625] 3. Image analysis (OCR processing)

[0626] Server: Passes the received menu image to the image processing module.

[0627] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0628] 4. Natural Language Processing (NLP)

[0629] Server: Passes the extracted string to the natural language processing module.

[0630] Server: Generates a dish description from a string, for example "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[0631] 5. Image Generation

[0632] Server: Based on the string, retrieve related images using a web search API or database.

[0633] Server: Sends the acquired image and description to the device.

[0634] 6. Displaying the menu

[0635] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0636] User: Check the menu that appears and tap the desired dish.

[0637] 7. Sending Selected Information

[0638] Terminal: Sends information about the menu item selected by the user to the server.

[0639] 8. Generating Recommendations

[0640] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0641] Server: Receives the recommendation results from the modules and selects the drinks and desserts that are best suited to the user.

[0642] 9. Display of Recommendations

[0643] Terminal: Displays the drink and dessert recommendations received from the server to the user.

[0644] User: Review the recommendations and indicate their intent to place an order.

[0645] 10. Confirming your order and obtaining additional information

[0646] User: Tells the waiter their order.

[0647] Terminal: Passes the conversation between the user and the store clerk to a speech recognition module, which converts the speech into text.

[0648] Server: Parse the converted text and extract additional information, such as order confirmation and out-of-stock information.

[0649] 11. Amendments to Order Details

[0650] Server: If the store clerk's response includes information such as "out of stock," generate an alternative.

[0651] Terminal: Inform the user of the alternatives.

[0652] User: Review the alternatives, select a new menu option, and confirm.

[0653] Specific examples

[0654] 1. Photographing the menu

[0655] User: Take a photo of the menu with their smartphone.

[0656] 2. Sending images

[0657] Device: Creates a request to send the captured menu image to the server.

[0658] On your device: Upload the menu image to the server.

[0659] 3. Image analysis (OCR processing)

[0660] Server: Passes the menu image to the OCR processing module.

[0661] Server: Extract the string "Pasta alla Carbonara".

[0662] 4. Natural Language Processing (NLP)

[0663] Server: Passes the extracted string to the NLP module.

[0664] Server: Generates the description "Italian pasta featuring a creamy sauce and bacon" from "Pasta alla Carbonara."

[0665] 5. Image Generation

[0666] Server: Uses a web search API based on the string to retrieve related images.

[0667] Server: Sends the description and image to the device.

[0668] 6. Displaying the menu

[0669] Terminal: Display "Pasta alla Carbonara," "Italian pasta featuring a creamy sauce and bacon," and related images.

[0670] 7. Sending Selected Information

[0671] User: Select "Pasta alla Carbonara."

[0672] Terminal: Sends the selection information to the server.

[0673] 8. Generating Recommendations

[0674] Server: Recommend suitable drinks and desserts based on your selections.

[0675] Server: Recommend red wine or tiramisu.

[0676] 9. Display of Recommendations

[0677] Terminal: Display recommendations for red wine and tiramisu to the user.

[0678] 10. Confirming your order and obtaining additional information

[0679] User: Tells the waiter their order.

[0680] Terminal: The speech recognition module converts the conversation into text.

[0681] Server: Extracts additional information from the parsed text.

[0682] 11. Amendments to Order Details

[0683] Server: Based on the information "Pasta alla Carbonara is out of stock", generate "Pasta Bolognese" as an alternative.

[0684] Terminal: Inform the user of the alternatives.

[0685] User: Review the alternatives, select "Pasta Bolognese" and confirm.

[0686] In this way, a system is provided that allows users to easily understand menus written in foreign languages ​​and order appropriate dishes.

[0687] The processing flow will be explained below.

[0688] Step 1:

[0689] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0690] Step 2:

[0691] Device: Creates a request to send the captured menu image to the server.

[0692] Step 3:

[0693] Terminal: Uploads menu images to the server via the network.

[0694] Step 4:

[0695] Server: Passes the received menu image to the image processing module.

[0696] Step 5:

[0697] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0698] Step 6:

[0699] Server: Passes the extracted text data to the natural language processing module.

[0700] Step 7:

[0701] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[0702] Step 8:

[0703] Server: Obtains images related to the dish along with the generated description.

[0704] Step 9:

[0705] Server: Sends the description and image to the device.

[0706] Step 10:

[0707] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0708] Step 11:

[0709] User: Review the menu items displayed and tap to select the desired dish.

[0710] Step 12:

[0711] Terminal: Sends information about the menu item selected by the user to the server.

[0712] Step 13:

[0713] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0714] Step 14:

[0715] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[0716] Step 15:

[0717] Server: Sends the recommendation results to the device.

[0718] Step 16:

[0719] Terminal: Display recommended drinks and desserts to the user.

[0720] Step 17:

[0721] User: Review the recommendations and indicate their intent to place an order.

[0722] Step 18:

[0723] User: Inform the store clerk of the confirmed order.

[0724] Step 19:

[0725] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[0726] Step 20:

[0727] Device: Sends captured audio data to the server.

[0728] Step 21:

[0729] Server: Passes the voice data to the voice recognition module and converts it into text data.

[0730] Step 22:

[0731] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[0732] Step 23:

[0733] Server: Based on the analysis results, generate alternatives as needed.

[0734] Step 24:

[0735] Server: Sends alternatives to the device.

[0736] Step 25:

[0737] Terminal: Present alternatives to the user.

[0738] Step 26:

[0739] User: Review alternatives, select new menu, confirm.

[0740] Example 1

[0741] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0742] In modern society, many people face significant challenges in understanding menus written in a foreign language and ordering the appropriate food. This problem is particularly pronounced for tourists and people unfamiliar with the foreign language, leading to ordering errors and confusion. Furthermore, selecting the right drink or dessert presents similar challenges. The present invention aims to solve these problems and provide a smooth ordering process.

[0743] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0744] In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, and user interface means, which enable the server to extract text from menus written in a foreign language to generate dish descriptions, recommend related drinks and desserts, convert voice dialogue into text, and display this information on the user interface.

[0745] "Image recognition means" refers to a device or program that performs processing to extract text from an image captured using a camera.

[0746] "Natural language processing means" refers to a device or program that analyzes the extracted text and generates descriptions of dishes and related information.

[0747] A "recommendation means" is a device or program for recommending related food and beverage items based on a menu item selected by a user.

[0748] "Speech recognition means" refers to a device or program for converting the dialogue between the user and the store clerk from voice to text.

[0749] The "image generating means" refers to a device or program for obtaining an image related to a menu item and displaying it to the user.

[0750] "Network communication means" refers to a communication device or program for transmitting and receiving data between a terminal and a server.

[0751] "User interface means" refers to a device or program that provides a display screen and input means for a user to operate the system.

[0752] The system of the present invention is composed of a user, a terminal, and a server. By using this system, the user can understand menus written in a foreign language and order the appropriate food.

[0753] Specific examples of hardware and software used

[0754] Hardware

[0755] 1. Camera-equipped devices (smartphones, tablets)

[0756] 2. Server (Cloud server, local server)

[0757] software

[0758] 1. OCR module (Google Cloud Vision API, Tesseract OCR)

[0759] 2. Natural language processing module (IBM Watson NLP, Google Cloud Natural Language API)

[0760] 3. Recommendation module (Collaborative Filtering, Content-Based Filtering algorithms)

[0761] 4. Speech Recognition Module (Google Speech-to-Text API, IBM Watson Speech to Text)

[0762] 5. Network communication module (HTTP / HTTPS communication library)

[0763] 6. User Interface Module (React Native, Flutter)

[0764] Specific operation of the system

[0765] 1. Photographing the menu

[0766] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[0767] 2. Sending images

[0768] The device encodes the captured menu image and creates a request to send to the server, specifically uploading the image to the server's endpoint using an HTTP POST request.

[0769] 3. Image analysis (OCR processing)

[0770] The server passes the received image to an OCR module, which extracts text from the image, for example, the string "Pasta alla Carbonara."

[0771] 4. Natural Language Processing (NLP)

[0772] The server passes the extracted text to a natural language processing module to generate a description of the dish, for example, "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[0773] 5. Image Generation

[0774] The server uses a web search API to retrieve related images based on the results of natural language processing, and sends the images and descriptions together to the device.

[0775] 6. Displaying the menu

[0776] The terminal displays the menu items, explanations, and related images received from the server on the user interface, providing information in a format that is easy for the user to select.

[0777] 7. Sending Selected Information

[0778] The user selects the dish of interest and confirms the information.

[0779] The terminal transmits the user's selection information to the server.

[0780] 8. Generating Recommendations

[0781] Based on the user's selection, the server uses a recommendation module to recommend suitable drinks and desserts, for example, red wine or tiramisu that go well with "Pasta alla Carbonara."

[0782] 9. Display of Recommendations

[0783] The terminal displays the information recommended by the server to the user, who can then confirm the recommendation and place an order.

[0784] 10. Confirming your order and obtaining additional information

[0785] The user tells the store clerk their order.

[0786] The terminal uses a voice recognition module to convert the conversation between the user and the store clerk into text and send it to the server.

[0787] The server analyzes the dialogue and extracts any additional information needed.

[0788] 11. Amendments to Order Details

[0789] The server generates alternative suggestions based on information such as "Pasta alla Carbonara is out of stock."

[0790] The terminal notifies the user of the alternatives, and the user selects and confirms the new menu.

[0791] Specific examples

[0792] 1. Menu photography:

[0793] A user takes a photo of a restaurant menu using a smartphone.

[0794] 2. Sending images:

[0795] The device encodes the captured menu image and sends it to the server's endpoint.

[0796] 3. Image analysis (OCR processing):

[0797] The server uses an OCR module to extract the string "Pasta alla Carbonara".

[0798] 4. Natural Language Processing (NLP):

[0799] The server generates a description for "Pasta alla Carbonara": "Italian pasta featuring a creamy sauce and bacon."

[0800] 5. Image generation:

[0801] The server uses a web search API to retrieve an image of "Pasta alla Carbonara."

[0802] 6. Display Menu:

[0803] The device displays "Pasta alla Carbonara," an Italian pasta dish featuring a creamy sauce and bacon, along with related images.

[0804] 7. Sending Selected Information:

[0805] The user selects "Pasta alla Carbonara" and the terminal transmits the selection information to the server.

[0806] 8. Generating recommendations:

[0807] The server will recommend suitable drinks and desserts, such as red wine and tiramisu.

[0808] 9. Display of Recommendations:

[0809] The device displays recommendations for red wine and tiramisu.

[0810] 10. Confirm your order and obtain additional information:

[0811] The user tells the store clerk their order.

[0812] The device uses a voice recognition module to convert the conversation into text.

[0813] 11. Order Modifications:

[0814] Based on the information that "Pasta alla Carbonara is out of stock," the server suggests "Pasta Bolognese."

[0815] The device presents the alternatives and the user selects "Pasta Bolognese."

[0816] This system makes it easier for users to understand foreign language menus and order the right food.

[0817] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0818] Step 1:

[0819] A user opens the camera on their smartphone or tablet and takes a picture of a restaurant menu. The input is the camera's photo image, and the output is an image file of the menu. Specifically, the user opens the camera app and presses the shutter button to capture the image.

[0820] Step 2:

[0821] The device encodes the captured menu image to send to the server and creates an HTTP POST request. The input is the menu image file, and the output is an HTTP request. Specifically, the device encodes the image into JPEG format and sends it to the server's endpoint (e.g., https: / / example.com / upload).

[0822] Step 3:

[0823] The server passes the received menu image to the OCR module, which extracts text from the image. The input is the image data of the menu, and the output is the extracted string. Specifically, the server sends the image data to the Google Cloud Vision API and obtains the text "Pasta alla Carbonara."

[0824] Step 4:

[0825] The server passes the extracted string to a natural language processing module to generate a description of the dish. The input is the extracted string, and the output is a description of the dish. Specifically, the server uses the NLP module to generate a description such as "Italian pasta with a creamy sauce and bacon" from "Pasta alla Carbonara."

[0826] Step 5:

[0827] The server uses a web search API to retrieve related images based on the results of natural language processing. The input is a description of the dish, and the output is an image of the dish. Specifically, the server retrieves images related to "Pasta alla Carbonara" using the Google Image Search API and selects the appropriate image.

[0828] Step 6:

[0829] The server compiles the obtained description and image in JSON format and sends it to the device. The input is the description and image of the dish, and the output is a JSON response. Specifically, the server generates JSON data such as { "name": "Pasta alla Carbonara", "description": "Italian pasta characterized by a creamy sauce and bacon", "image_url": "https: / / example.com / image.jpg"} and sends it to the device.

[0830] Step 7:

[0831] The device parses the JSON data received from the server and displays menu items, descriptions, and images on the user interface. The input is the JSON response received from the server, and the output is the user interface display. Specifically, the device uses an application to parse the received data and display it on the screen.

[0832] Step 8:

[0833] The user selects the dish of interest from the displayed menu, and the device sends the selection to the server. The input is the user's selection, and the output is an HTTP request. Specifically, the user taps "Pasta alla Carbonara" on the touchscreen, and the device sends the selection to the server.

[0834] Step 9:

[0835] The server uses a recommendation module to recommend related drinks and desserts based on the user's selection. The input is the user's selection, and the output is a list of recommended items. Specifically, the server uses a Collaborative Filtering algorithm to recommend red wine and tiramisu to go with "Pasta alla Carbonara."

[0836] Step 10:

[0837] The server compiles the generated recommendation results in JSON format and sends them to the device. The input is a list of recommended items, and the output is a JSON response. Specifically, the server generates JSON data such as { "recommendations": ["red wine", "tiramisu"]} and sends it to the device.

[0838] Step 11:

[0839] The device displays the recommendation information received from the server on the user interface, and the user confirms the displayed recommendation selection. The input is the JSON response received from the server, and the output is the display of the user interface and the user's selection action. Specifically, the device displays the received recommendation information, and the user taps the confirm button.

[0840] Step 12:

[0841] The user tells the store clerk their order, and the device converts the conversation into text using a voice recognition module and sends it to the server. The input is the voice conversation between the user and the store clerk, and the output is text data. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text.

[0842] Step 13:

[0843] The server analyzes the textual content of the conversation and obtains additional information. The input is text data, and the output is the analysis result. Specifically, the server analyzes the text and extracts information such as confirmation of order details and out-of-stock information.

[0844] Step 14:

[0845] The server generates alternatives as needed and notifies the terminal. The input is the analysis result information, and the output is the alternative information. In concrete terms, the server might suggest "Pasta Bolognese" based on the information "Pasta alla Carbonara is out of stock."

[0846] Step 15:

[0847] The device notifies the user of the alternatives, and the user selects and confirms the new menu. The input is the information about the alternatives, and the output is the user's selection action. Specifically, the device displays the alternatives, and the user selects a new menu and taps the confirm button.

[0848] (Application example 1)

[0849] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0850] In modern factories, many pieces of equipment and products use advanced technology, making it difficult for workers to grasp all the information and work efficiently. In particular, not being able to instantly obtain operating procedures and maintenance information related to equipment and products can lead to reduced production efficiency and the risk of operating errors. In addition, when working in a multilingual environment, language barriers become an additional obstacle. Therefore, there is a growing need for systems that can provide detailed information about equipment and products in the factory in real time.

[0851] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0852] In this invention, the server includes an image recognition unit, a natural language processing unit, and a character analysis unit. This allows real-time access to information about equipment and products in a factory using smart glasses or a head-mounted display. This allows workers to instantly access detailed operation guides and maintenance information about the equipment and products, improving work efficiency. Furthermore, by utilizing a generative AI model with prompt sentences, multilingual support is possible, enabling information provision across language barriers.

[0853] "Image recognition means" refers to a technology that uses a camera or other photographic device to acquire image data and extract specific features from the image.

[0854] "Natural language processing means" refers to technology that analyzes character strings and sentences and converts or interprets them into natural language that humans can understand.

[0855] A "recommendation tool" is a system that automatically recommends appropriate products and services based on a user's past behavior and choices.

[0856] "Speech recognition means" is a technology that analyzes speech and converts it into text data.

[0857] "Image generation means" refers to technology that creates new images based on data and information.

[0858] "Network communication means" refers to technology for sending and receiving data over the Internet or other communication networks.

[0859] "User interface means" refers to an interface such as a screen or input device that allows a user to interact with the system.

[0860] "Information acquisition means" refers to technology for acquiring information about equipment and products within a factory using smart glasses or a head-mounted display.

[0861] "Character analysis means" is a technology that uses OCR technology to analyze photographed character strings and convert them into digital text.

[0862] A "prompt" is a short command that a generative AI model receives as input and is an instruction to generate specific information.

[0863] The following system is required to implement this invention: The system is mainly composed of a server, a terminal, and a user.

[0864] The server includes an image recognition means, a natural language processing means, a character analysis means, a recommendation means, a voice recognition means, an image generation means, a network communication means, a generative AI model using prompt sentences, and the like.

[0865] The terminal is equipped with smart glasses or a head-mounted display, a camera, a microphone, and a user interface means, and is operated by the user.

[0866] The server first analyzes the label image of the equipment or product sent from the terminal using image recognition, which includes character analysis using OCR technology. Specifically, it uses Tesseract OCR to extract character strings from the image.

[0867] The extracted text is analyzed using natural language processing to generate an appropriate description, which can be translated into multiple languages ​​as needed using translation services such as Google Translate API.

[0868] The generated description and related images are sent to the terminal via the network communication means and displayed to the user by the user interface means. For example, if the character string "product 12345" is extracted, the details displayed will read, "Product 12345: This part has been machined using a high-precision grinder. Handle with care. Wear protective equipment and follow the operating instructions when using."

[0869] In addition, the recommendation means recommends appropriate operation guides and maintenance procedures based on the user's past selections and behavioral data, thereby further improving work efficiency.

[0870] An example of using a generative AI model with prompts is the following prompt:

[0871] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[0872] "Product 12345: This part has been processed using a high-precision grinder."

[0873] The system configuration and processing procedures described above enable detailed information on equipment and products within the factory to be provided in real time, improving work efficiency. In addition, the system's multilingual support makes it suitable for international work environments.

[0874] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0875] Step 1:

[0876] A user takes a picture of the labels of equipment or products in a factory using a camera in smart glasses or a head-mounted display. Here, the input is the label image, and the output is the image data captured by the camera.

[0877] Step 2:

[0878] The terminal sends the captured label image to the server. The input is label image data, and the output is image data sent to the server via the network.

[0879] Step 3:

[0880] The server uses image recognition and OCR technology to extract text from the received label image. Specifically, it uses Tesseract OCR. The input is the label image data, and the output is the extracted text.

[0881] Step 4:

[0882] The server passes the extracted string to a natural language processing means to generate an appropriate description. It translates it into multiple languages ​​as needed using NLP technology or the Google Translate API. The input is the extracted string, and the output is the generated description.

[0883] Step 5:

[0884] The server uses an image generation means to retrieve associated images along with the generated description, for example, the associated images are retrieved from a database or the Internet, where the input is the description and the output is the associated images.

[0885] Step 6:

[0886] The server transmits the generated explanatory text and related images to the terminal using a network communication means. The input is the generated explanatory text and related images, and the output is the data transmitted to the terminal.

[0887] Step 7:

[0888] The terminal displays the received explanatory text and image to the user through a user interface means. The input is the explanatory text and image data, and the output is the information displayed on the user interface.

[0889] Step 8:

[0890] The user checks the displayed information and performs operations to obtain more detailed information as necessary. For example, if a detailed operation guide is required, a prompt sentence is generated and the request is sent from the terminal to the server. The input is the user's operation, and the output is the generated prompt sentence and its request.

[0891] Step 9:

[0892] The server uses a generative AI model to generate detailed operation guides based on the prompt sentences and provides them to the user. The input is the prompt sentence, and the output is the detailed operation guides.

[0893] For specific actions, the following example prompt sentence is used:

[0894] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[0895] "Product 12345: This part has been processed using a high-precision grinder."

[0896] Through the above steps, users can obtain detailed information about the equipment and products in the factory in real time, improving work efficiency.

[0897] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0898] As an embodiment of the present invention, the specific implementation procedure of a system using a user, a terminal, a server, and an emotion engine is described below. This system assists users in the process of understanding foreign language menus and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[0899] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to the server. The server then passes the image to an image processing module, which uses OCR technology to extract the text from the image. For example, the string "Pasta alla Carbonara" is extracted.

[0900] The extracted text data is then passed to a natural language processing module. The server generates a description of the dish from this text data. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via a web search API or database. The server then sends the retrieved description and images to the device.

[0901] The device displays the received menu items, descriptions, and images to the user. The user checks the displayed menu and taps on the desired dish to select it. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts. The server sends the recommendation results to the device, which displays them to the user. The user checks the recommendations and confirms the order.

[0902] Furthermore, the system incorporates an emotion engine, which recognizes emotions from the user's facial expressions and voice. For example, if the user expresses a confused expression, the emotion engine analyzes the information. This emotion information is sent to the server, and menu suggestions and recommendations are adjusted based on the user's emotions. For example, if the user is confused, the system adjusts to provide more detailed explanations and additional information.

[0903] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion engine will recognize the emotion and the server will provide additional detailed explanations of the recommendation and alternatives based on the emotion.

[0904] After the order is confirmed, the user conveys the order to the store clerk. The voice recognition module is activated again, capturing the conversation between the user and the store clerk. The voice data is sent to the server, where it is converted into text data by the voice recognition module. The converted text data is then analyzed to extract additional information, such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user. The user then confirms the alternative options, selects a new menu item, and confirms the selection.

[0905] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[0906] The processing flow will be explained below.

[0907] Step 1:

[0908] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[0909] Step 2:

[0910] Device: Creates a request to send the captured menu image to the server.

[0911] Step 3:

[0912] Terminal: Uploads menu images to the server via the network.

[0913] Step 4:

[0914] Server: Passes the received menu image to the image processing module.

[0915] Step 5:

[0916] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[0917] Step 6:

[0918] Server: Passes the extracted text data to the natural language processing module.

[0919] Step 7:

[0920] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[0921] Step 8:

[0922] Server: Obtains images related to the dish along with the generated description.

[0923] Step 9:

[0924] Server: Sends the description and image to the device.

[0925] Step 10:

[0926] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[0927] Step 11:

[0928] User: Review the menu items displayed and tap to select the desired dish.

[0929] Step 12:

[0930] Terminal: Sends information about the menu item selected by the user to the server.

[0931] Step 13:

[0932] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[0933] Step 14:

[0934] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[0935] Step 15:

[0936] Server: Sends the recommendation results to the device.

[0937] Step 16:

[0938] Terminal: Display recommended drinks and desserts to the user.

[0939] Step 17:

[0940] User: Review the recommendations and indicate their intent to place an order.

[0941] Step 18:

[0942] Device: Activates the emotion engine to recognize emotions from the user's facial expressions and voice.

[0943] Step 19:

[0944] Terminal: The emotion engine analyzes the user's emotions and sends the results to the server.

[0945] Step 20:

[0946] Server: Receives user emotion data and adjusts recommendations and menu suggestions.

[0947] Step 21:

[0948] Server: Generates additional information and alternative suggestions based on emotion data and sends them to the device.

[0949] Step 22:

[0950] On the device: Display additional information or alternative suggestions to the user based on their emotions.

[0951] Step 23:

[0952] User: Review the proposed information and confirm the order.

[0953] Step 24:

[0954] User: Tells the waiter their order.

[0955] Step 25:

[0956] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[0957] Step 26:

[0958] Device: Sends captured audio data to the server.

[0959] Step 27:

[0960] Server: Passes the voice data to the voice recognition module and converts it into text data.

[0961] Step 28:

[0962] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[0963] Step 29:

[0964] Server: Based on the analysis results, generate alternatives as needed.

[0965] Step 30:

[0966] Server: Sends alternatives to the device.

[0967] Step 31:

[0968] Terminal: Present alternatives to the user.

[0969] Step 32:

[0970] User: Review alternatives, select new menu, confirm.

[0971] Example 2

[0972] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0973] For many users, understanding menus in a foreign language and ordering the appropriate dishes is a difficult task. Furthermore, if recommendations are not appropriate based on the user's emotions and preferences, the user experience cannot be improved. To solve these challenges, a system is needed that not only extracts menu text and provides easy-to-understand information, but also makes recommendations that adapt to the user's emotional state.

[0974] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0975] In this invention, the server includes image recognition means, natural language processing means, recommendation means, and emotion recognition means. This allows users to easily understand menus in foreign languages ​​and order appropriate dishes. Recommendations are also made that are adapted to the user's emotional state, allowing the user to enjoy personalized services.

[0976] "Image recognition means" refers to technology for analyzing an image and extracting specific information from its content.

[0977] "Natural language processing" refers to technology for understanding, generating, and manipulating human language.

[0978] "Recommendation means" refers to technology for recommending appropriate items based on a user's preferences and history.

[0979] "Speech recognition means" refers to technology for analyzing voice data and converting it into text data.

[0980] "Image generation means" refers to technology for generating new images based on input data.

[0981] "Emotion recognition means" refers to technology for analyzing and recognizing emotions from a user's facial expressions and voice.

[0982] "Network communication means" refers to technology for communicating over a network to send and receive data.

[0983] "User interface means" refers to technology that provides an interface for exchanging information between a user and a system.

[0984] The present invention relates to a system that assists users in the process of understanding menus in a foreign language and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[0985] Hardware and Software Configuration

[0986] Required Hardware

[0987] User's device: A smartphone or tablet equipped with a camera is used.

[0988] Server: A high-performance server is used for menu image analysis and recommendation processing.

[0989] Required software

[0990] Image Recognition Method: OCR technology is used to extract text from images, specifically Google Cloud Vision API.

[0991] Natural language processing tools: A generative AI model (e.g., GPT-4) is used to generate dish descriptions from the extracted text data.

[0992] Recommendation method: Use a recommendation algorithm to recommend drinks and desserts related to the dish.

[0993] Emotion Recognition: Emotion recognition software is used to analyze and recognize emotions from the user's facial expressions and voice.

[0994] Speech recognition means: Use speech recognition technology (e.g., Google Speech-to-Text API) to capture the conversation between the user and the store clerk as voice data and convert it into text data.

[0995] Image generation method: Use a web search API (e.g., Bing Image Search API) to obtain related images.

[0996] Network communication means: Uses communication technology to send and receive data between the user terminal and the server.

[0997] User interface means: Provides an interface for exchanging information between the user and the system.

[0998] Example of a system

[0999] 1. Menu Parsing Prompt Example

[1000] "Analyze a restaurant menu image and generate dish names and descriptions."

[1001] 2. Recommendation prompt examples

[1002] "Recommend drinks and desserts to go with the selected dish."

[1003] 3. Example of emotion-responsive prompt

[1004] "If the user looks confused, provide a detailed explanation."

[1005] System Operation

[1006] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to a server. The server then analyzes the received image using the Google Cloud Vision API and extracts the text from the image using OCR technology. For example, the string "Pasta alla Carbonara" is extracted.

[1007] The extracted text data is then passed to a generative AI model such as GPT-4 to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via the Bing Image Search API. The server then sends the retrieved description and image to the device, which then displays them to the user.

[1008] The user checks the displayed menu items, descriptions, and images, and taps to select the desired dish. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts, and sends the recommendation results to the device. The device then displays them to the user.

[1009] The app also has a built-in emotion engine that recognizes emotions from the user's facial expressions and voice. For example, if the user shows a confused expression, the emotion engine analyzes the information and sends it to the server. Based on this emotion information, the server adjusts the service to provide detailed explanations and additional information to the user.

[1010] Finally, the user confirms the order and conveys it to the store clerk. The voice recognition module is activated, capturing the conversation between the user and the store clerk, and the voice data is sent to the server. The server's voice recognition module converts the voice data into text data, which is then analyzed to extract additional information such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user.

[1011] In this way, the present invention allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's feelings.

[1012] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1013] Step 1:

[1014] Users take a photo of a restaurant menu using the camera on their smartphone or tablet.

[1015] Input: Menu Image

[1016] Specific actions: The user opens the camera app, points the lens at the menu, and presses the shutter button to take a photo.

[1017] Output: Menu image captured

[1018] Step 2:

[1019] The terminal transmits the captured menu image to the server.

[1020] Input: Menu image taken

[1021] Specific operation: After taking a photo, the image will be automatically uploaded and sent.

[1022] Output: Menu image sent to the server

[1023] Step 3:

[1024] The server analyzes the received menu image using OCR technology and extracts the text within the image.

[1025] Input: Menu image sent to the server

[1026] Data processing: Call the Google Cloud Vision API to extract character codes from images

[1027] What it does: The API analyzes the text in the menu image and extracts text such as "Pasta alla Carbonara."

[1028] Output: Extracted text data

[1029] Step 4:

[1030] The server passes the extracted text data to a natural language processing module to generate a description of the dish.

[1031] Input: Extracted text data

[1032] Data processing: Calling GPT-4 to generate explanatory text from text data

[1033] What it does: GPT-4 takes the text "Pasta alla Carbonara" and generates a description: "Italian pasta with a creamy sauce and bacon."

[1034] Output: Generated dish description

[1035] Step 5:

[1036] The server retrieves related images via a web search API.

[1037] Input: Description of the generated dish

[1038] Data processing: Call the Bing Image Search API to get related images

[1039] What it does: Uses the Bing Image Search API to search and retrieve relevant images based on the description of "Pasta alla Carbonara."

[1040] Output: Captured image

[1041] Step 6:

[1042] The server transmits the explanatory text and the image to the terminal, which displays them to the user.

[1043] Input: Generated description and retrieved image

[1044] Specific operation: Data is sent from the server to the device, and the device displays a pop-up explanation and image on the screen.

[1045] Output: Description and image displayed to the user

[1046] Step 7:

[1047] The user taps on the desired dish to select it, and the device sends the selection information to the server.

[1048] Input: Displayed description and image

[1049] What happens: The user taps "Pasta alla Carbonara" and the selection is sent to the server.

[1050] Output: Selections sent to the server

[1051] Step 8:

[1052] The server uses a recommendation module to recommend related drinks and desserts and transmits them to the terminal.

[1053] Input: Selections sent to the server

[1054] Data processing: Use recommendation algorithms to select relevant drinks and desserts

[1055] Specific operation: The server selects red wine and tiramisu, and the recommended results are displayed on the device screen.

[1056] Output: Recommendation results displayed to the user

[1057] Step 9:

[1058] The emotion engine recognizes emotions from the user's facial expressions and voice, and transmits the emotion information to the server.

[1059] Input: User's facial expressions and voice

[1060] Data processing: Using emotion recognition software, analyze and recognize the user's emotions.

[1061] Specific operation: When the user makes a confused expression, the camera analyzes the expression and sends the emotional information to the server, which then displays a detailed explanation.

[1062] Output: Additional explanation based on emotion information

[1063] Step 10:

[1064] The user confirms the final order and communicates it to the store attendant. The voice recognition module captures the conversation and sends it to the server.

[1065] Input: User and store clerk conversation

[1066] Data processing: Using voice recognition technology, converting voice data into text data

[1067] What it does: The user verbally places an order, and the speech is sent to the server, which converts it into text. The server then provides out-of-stock information and alternative suggestions.

[1068] Output: Textualized dialogue, out-of-stock information, and alternatives

[1069] Through this series of processes, users can easily understand menus in foreign languages ​​and order the most suitable dishes based on their emotions.

[1070] (Application example 2)

[1071] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1072] In existing systems, users may have difficulty understanding foreign language menus and selecting appropriate dishes, and they are unable to provide personalized recommendations based on the user's emotions. The present invention aims to solve these problems, enabling users to easily understand foreign language menus and providing personalized services based on the user's emotions.

[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means. This allows the user to easily understand foreign language menus and further receive personalized recommendations based on their emotions.

[1074] "Image recognition means" is a device that extracts information from a captured image and analyzes its content.

[1075] "Natural language processing means" is a device that analyzes extracted text data, converts it into natural language format, and understands and processes it.

[1076] A "recommendation means" is a device that makes relevant suggestions or recommendations to the user based on the analyzed information.

[1077] A "voice recognition means" is a device that analyzes voice data and converts it into a linguistic text format.

[1078] The "image generating means" is a device that generates an image based on the analyzed information and generated data.

[1079] A "network communication means" is a communication device for transmitting and receiving data between systems.

[1080] "User interface means" refers to a device that allows a user to operate and interact with the system.

[1081] The "emotion analysis means" is a device that recognizes and analyzes the user's emotions from their facial expressions and tone of voice.

[1082] The present invention provides a system that assists users in understanding menus written in a foreign language and ordering appropriate dishes. The system includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means.

[1083] First, a user takes a photo of a restaurant menu using a smartphone or tablet device. The image taken with the device's camera is sent from the device to a server via network communication means. The server then uses image recognition means to extract text from the received image. For example, the text "Pasta alla Carbonara" is extracted.

[1084] Next, the extracted text data is analyzed using natural language processing to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta characterized by a creamy sauce and bacon." At the same time, related images are retrieved via a web search API or database. The server then sends the generated description and the retrieved images to the device via the network.

[1085] The terminal uses a user interface means to display the received menu items, descriptions, and images to the user. The user checks the displayed menu and selects the desired dish. Information about the selected menu item is again sent from the terminal to the server. The server uses a recommendation means to recommend additional related drinks and desserts. The recommendation results are sent to the terminal and displayed to the user.

[1086] The system also incorporates an emotion analysis mechanism that recognizes emotions from the user's facial expressions and voice and adjusts menu suggestions and recommendations based on that information. For example, if the user expresses confusion, the emotion analysis mechanism sends that information to the server, which then adjusts the system to provide detailed explanations and additional information.

[1087] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion analyzer recognizes the emotion and the server provides additional detailed explanations of the recommendation and alternatives.

[1088] After the order is confirmed, the user tells the store clerk what they want to order. At this time, the voice recognition means is activated and the conversation between the user and the store clerk is captured. The voice data is sent back to the server and converted into text data by the voice recognition means. The converted text data is analyzed to extract information such as out-of-stock items and additional orders. If necessary, alternative options are generated and notified to the user. The user then selects a new menu item and confirms the order.

[1089] An example prompt is, "Use the EmotionRecognition library to analyze emotions from the image menu.jpg and generate personalized dish descriptions based on that."

[1090] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[1091] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1092] Step 1:

[1093] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[1094] Input: Menu Image

[1095] Output: Raw image data in the device

[1096] Specific operation: A user opens the device's camera app and takes a picture of a restaurant menu. After taking the picture, the image data is temporarily stored in the device's memory.

[1097] Step 2:

[1098] The terminal transmits the photographed menu image to the server using the network communication means.

[1099] Input: Raw image data in the device

[1100] Output: Image data sent to the server

[1101] How it works: The device sends image data to the server via Wi-Fi or mobile data using the HTTP or HTTPS protocol.

[1102] Step 3:

[1103] The server uses image recognition means to extract text from the received image.

[1104] Input: Image data sent to the server

[1105] Output: Extracted text data

[1106] How it works: Optical Character Recognition (OCR) software on the server analyzes the image and extracts text information, which is then stored in a database on the server, such as "Pasta alla Carbonara."

[1107] Step 4:

[1108] The server uses natural language processing means to generate a description of the dish from the extracted text data.

[1109] Input: Extracted text data

[1110] Output: The generated dish description

[1111] What it does: A server-based natural language processing library analyzes the text data and generates a description of the corresponding dish, such as "Italian pasta with a creamy sauce and bacon."

[1112] Step 5:

[1113] The server retrieves related images via a web search API or database.

[1114] Input: Extracted text data

[1115] Output: Related images

[1116] Specific operation: The server calls the web search API, searches for relevant images based on the text data, selects the most appropriate image from the search results, and downloads it.

[1117] Step 6:

[1118] The server transmits the generated description and image to the terminal using a network communication means.

[1119] Input: Generated description and associated image

[1120] Output: Description and image sent to the device

[1121] Specific operation: The server uses the HTTP or HTTPS protocol to send the description and image to the terminal.

[1122] Step 7:

[1123] The terminal uses the user interface means to display the received menu items, explanations, and images to the user.

[1124] Input: Received description and image

[1125] Output: Menu items, descriptions, and images displayed in the user interface

[1126] Specific behavior: The device uses the layout template to properly arrange the displayed information on the screen, allowing the user to visually confirm the displayed information.

[1127] Step 8:

[1128] The user checks the displayed menu items and taps to select the desired dish.

[1129] Input: User tap input

[1130] Output: Information about the selected menu item

[1131] Specific operation: The information of the menu item selected by the user is stored in the terminal's memory in preparation for the next communication step.

[1132] Step 9:

[1133] The terminal transmits information about the selected menu item to the server using the network communication means.

[1134] Input: Information of the selected menu item

[1135] Output: Selections sent to the server

[1136] Specific operation: The device sends the selection information to the server via Wi-Fi or mobile data communication.

[1137] Step 10:

[1138] The server uses a recommendation means to recommend related drinks and desserts.

[1139] Input: Information of the selected menu item

[1140] Output: Recommended drinks and desserts

[1141] Specific operation: The server's recommendation engine searches for items related to the selected menu and generates a recommendation list.

[1142] Step 11:

[1143] The server transmits the recommendation results to the terminal using a network communication means.

[1144] Input: Recommended drink and dessert information

[1145] Output: Recommendation results sent to the device

[1146] Specific operation: The server uses HTTP or HTTPS protocol to send the recommendation results to the terminal.

[1147] Step 12:

[1148] The terminal uses the user interface means to display the recommendation results to the user.

[1149] Input: Received recommendation results

[1150] Output: Recommendation results displayed in the user interface

[1151] Specific behavior: The recommended drinks and desserts will be displayed on the device screen for the user to review.

[1152] Step 13:

[1153] The server uses emotion analysis means to recognize emotions from the user's facial expressions and voice.

[1154] Input: User's facial expressions or voice data

[1155] Output: Recognized emotion information

[1156] Specific operation: The server uses a facial expression recognition library and a voice analysis library to analyze the user's emotions and saves the recognition results in a database.

[1157] Step 14:

[1158] The server adjusts menu suggestions and recommendations based on the recognized emotional information.

[1159] Input: Recognized emotion information

[1160] Output: Tailored menu suggestions and recommendations

[1161] Specific operation: Based on the results of the sentiment analysis, the server generates detailed information and alternatives and provides them to the user.

[1162] Step 15:

[1163] The user finally confirms and confirms the menu item.

[1164] Input: User confirmed input

[1165] Output: Confirmed order information

[1166] Specific operation: When the user selects a menu item and taps the confirmation button, the order information is finally captured and saved in the device's memory.

[1167] Step 16:

[1168] When the terminal communicates the order information to the store clerk, the voice recognition means captures the content of the conversation.

[1169] Input: Voice data of the conversation between the user and the store clerk

[1170] Output: Captured audio data

[1171] Specific operation: The device's microphone records the conversation between the user and the store clerk and saves it as audio data.

[1172] Step 17:

[1173] The server receives the voice data and converts it into text data using a voice recognition means.

[1174] Input: Captured audio data

[1175] Output: Converted text data

[1176] Specific operation: A speech recognition library on the server analyzes the voice data and generates corresponding text.

[1177] Step 18:

[1178] The server analyzes the converted text data and extracts out-of-stock information and additional orders.

[1179] Input: Converted text data

[1180] Output: Out-of-stock information and reorder notifications

[1181] Specific operation: The server analyzes the text data, extracts the necessary information, and generates alternative menus and additional order information as necessary.

[1182] Step 19:

[1183] The server informs the user of the alternatives and the user selects a new menu.

[1184] Input: Out-of-stock information and reorder notifications

[1185] Output: New menu selection information

[1186] Specific operation: The user confirms the alternatives and sends the newly selected menu information from the terminal to the server.

[1187] In this way, the present invention helps users understand foreign language menus and provides personalized services based on emotions.

[1188] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1191] [Third embodiment]

[1192] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1193] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1194] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1195] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1196] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1198] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1199] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1200] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1201] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1202] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1203] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1204] As an embodiment of the present invention, the system program is implemented according to the following procedure: The system is mainly composed of a server, a terminal, and a user.

[1205] 1. Photographing the menu

[1206] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[1207] 2. Sending images

[1208] Device: Creates a request to send the captured menu image to the server.

[1209] Terminal: Uploads menu images to the server via the network.

[1210] 3. Image analysis (OCR processing)

[1211] Server: Passes the received menu image to the image processing module.

[1212] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[1213] 4. Natural Language Processing (NLP)

[1214] Server: Passes the extracted string to the natural language processing module.

[1215] Server: Generates a dish description from a string, for example "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[1216] 5. Image Generation

[1217] Server: Based on the string, retrieve related images using a web search API or database.

[1218] Server: Sends the acquired image and description to the device.

[1219] 6. Displaying the menu

[1220] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[1221] User: Check the menu that appears and tap the desired dish.

[1222] 7. Sending Selected Information

[1223] Terminal: Sends information about the menu item selected by the user to the server.

[1224] 8. Generating Recommendations

[1225] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[1226] Server: Receives the recommendation results from the modules and selects the drinks and desserts that are best suited to the user.

[1227] 9. Display of Recommendations

[1228] Terminal: Displays the drink and dessert recommendations received from the server to the user.

[1229] User: Review the recommendations and indicate their intent to place an order.

[1230] 10. Confirming your order and obtaining additional information

[1231] User: Tells the waiter their order.

[1232] Terminal: Passes the conversation between the user and the store clerk to a speech recognition module, which converts the speech into text.

[1233] Server: Parse the converted text and extract additional information, such as order confirmation and out-of-stock information.

[1234] 11. Amendments to Order Details

[1235] Server: If the store clerk's response includes information such as "out of stock," generate an alternative.

[1236] Terminal: Inform the user of the alternatives.

[1237] User: Review the alternatives, select a new menu option, and confirm.

[1238] Specific examples

[1239] 1. Photographing the menu

[1240] User: Take a photo of the menu with their smartphone.

[1241] 2. Sending images

[1242] Device: Creates a request to send the captured menu image to the server.

[1243] On your device: Upload the menu image to the server.

[1244] 3. Image analysis (OCR processing)

[1245] Server: Passes the menu image to the OCR processing module.

[1246] Server: Extract the string "Pasta alla Carbonara".

[1247] 4. Natural Language Processing (NLP)

[1248] Server: Passes the extracted string to the NLP module.

[1249] Server: Generates the description "Italian pasta featuring a creamy sauce and bacon" from "Pasta alla Carbonara."

[1250] 5. Image Generation

[1251] Server: Uses a web search API based on the string to retrieve related images.

[1252] Server: Sends the description and image to the device.

[1253] 6. Displaying the menu

[1254] Terminal: Display "Pasta alla Carbonara," "Italian pasta featuring a creamy sauce and bacon," and related images.

[1255] 7. Sending Selected Information

[1256] User: Select "Pasta alla Carbonara."

[1257] Terminal: Sends the selection information to the server.

[1258] 8. Generating Recommendations

[1259] Server: Recommend suitable drinks and desserts based on your selections.

[1260] Server: Recommend red wine or tiramisu.

[1261] 9. Display of Recommendations

[1262] Terminal: Display recommendations for red wine and tiramisu to the user.

[1263] 10. Confirming your order and obtaining additional information

[1264] User: Tells the waiter their order.

[1265] Terminal: The speech recognition module converts the conversation into text.

[1266] Server: Extracts additional information from the parsed text.

[1267] 11. Amendments to Order Details

[1268] Server: Based on the information "Pasta alla Carbonara is out of stock", generate "Pasta Bolognese" as an alternative.

[1269] Terminal: Inform the user of the alternatives.

[1270] User: Review the alternatives, select "Pasta Bolognese" and confirm.

[1271] In this way, a system is provided that allows users to easily understand menus written in foreign languages ​​and order appropriate dishes.

[1272] The processing flow will be explained below.

[1273] Step 1:

[1274] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[1275] Step 2:

[1276] Device: Creates a request to send the captured menu image to the server.

[1277] Step 3:

[1278] Terminal: Uploads menu images to the server via the network.

[1279] Step 4:

[1280] Server: Passes the received menu image to the image processing module.

[1281] Step 5:

[1282] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[1283] Step 6:

[1284] Server: Passes the extracted text data to the natural language processing module.

[1285] Step 7:

[1286] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[1287] Step 8:

[1288] Server: Obtains images related to the dish along with the generated description.

[1289] Step 9:

[1290] Server: Sends the description and image to the device.

[1291] Step 10:

[1292] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[1293] Step 11:

[1294] User: Review the menu items displayed and tap to select the desired dish.

[1295] Step 12:

[1296] Terminal: Sends information about the menu item selected by the user to the server.

[1297] Step 13:

[1298] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[1299] Step 14:

[1300] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[1301] Step 15:

[1302] Server: Sends the recommendation results to the device.

[1303] Step 16:

[1304] Terminal: Display recommended drinks and desserts to the user.

[1305] Step 17:

[1306] User: Review the recommendations and indicate their intent to place an order.

[1307] Step 18:

[1308] User: Inform the store clerk of the confirmed order.

[1309] Step 19:

[1310] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[1311] Step 20:

[1312] Device: Sends captured audio data to the server.

[1313] Step 21:

[1314] Server: Passes the voice data to the voice recognition module and converts it into text data.

[1315] Step 22:

[1316] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[1317] Step 23:

[1318] Server: Based on the analysis results, generate alternatives as needed.

[1319] Step 24:

[1320] Server: Sends alternatives to the device.

[1321] Step 25:

[1322] Terminal: Present alternatives to the user.

[1323] Step 26:

[1324] User: Review alternatives, select new menu, confirm.

[1325] Example 1

[1326] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1327] In modern society, many people face significant challenges in understanding menus written in a foreign language and ordering the appropriate food. This problem is particularly pronounced for tourists and people unfamiliar with the foreign language, leading to ordering errors and confusion. Furthermore, selecting the right drink or dessert presents similar challenges. The present invention aims to solve these problems and provide a smooth ordering process.

[1328] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1329] In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, and user interface means, which enable the server to extract text from menus written in a foreign language to generate dish descriptions, recommend related drinks and desserts, convert voice dialogue into text, and display this information on the user interface.

[1330] "Image recognition means" refers to a device or program that performs processing to extract text from an image captured using a camera.

[1331] "Natural language processing means" refers to a device or program that analyzes the extracted text and generates descriptions of dishes and related information.

[1332] A "recommendation means" is a device or program for recommending related food and beverage items based on a menu item selected by a user.

[1333] "Speech recognition means" refers to a device or program for converting the dialogue between the user and the store clerk from voice to text.

[1334] The "image generating means" refers to a device or program for obtaining an image related to a menu item and displaying it to the user.

[1335] "Network communication means" refers to a communication device or program for transmitting and receiving data between a terminal and a server.

[1336] "User interface means" refers to a device or program that provides a display screen and input means for a user to operate the system.

[1337] The system of the present invention is composed of a user, a terminal, and a server. By using this system, the user can understand menus written in a foreign language and order the appropriate food.

[1338] Specific examples of hardware and software used

[1339] Hardware

[1340] 1. Camera-equipped devices (smartphones, tablets)

[1341] 2. Server (Cloud server, local server)

[1342] software

[1343] 1. OCR module (Google Cloud Vision API, Tesseract OCR)

[1344] 2. Natural language processing module (IBM Watson NLP, Google Cloud Natural Language API)

[1345] 3. Recommendation module (Collaborative Filtering, Content-Based Filtering algorithms)

[1346] 4. Speech Recognition Module (Google Speech-to-Text API, IBM Watson Speech to Text)

[1347] 5. Network communication module (HTTP / HTTPS communication library)

[1348] 6. User Interface Module (React Native, Flutter)

[1349] Specific operation of the system

[1350] 1. Photographing the menu

[1351] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[1352] 2. Sending images

[1353] The device encodes the captured menu image and creates a request to send to the server, specifically uploading the image to the server's endpoint using an HTTP POST request.

[1354] 3. Image analysis (OCR processing)

[1355] The server passes the received image to an OCR module, which extracts text from the image, for example, the string "Pasta alla Carbonara."

[1356] 4. Natural Language Processing (NLP)

[1357] The server passes the extracted text to a natural language processing module to generate a description of the dish, for example, "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[1358] 5. Image Generation

[1359] The server uses a web search API to retrieve related images based on the results of natural language processing, and sends the images and descriptions together to the device.

[1360] 6. Displaying the menu

[1361] The terminal displays the menu items, explanations, and related images received from the server on the user interface, providing information in a format that is easy for the user to select.

[1362] 7. Sending Selected Information

[1363] The user selects the dish of interest and confirms the information.

[1364] The terminal transmits the user's selection information to the server.

[1365] 8. Generating Recommendations

[1366] Based on the user's selection, the server uses a recommendation module to recommend suitable drinks and desserts, for example, red wine or tiramisu that go well with "Pasta alla Carbonara."

[1367] 9. Display of Recommendations

[1368] The terminal displays the information recommended by the server to the user, who can then confirm the recommendation and place an order.

[1369] 10. Confirming your order and obtaining additional information

[1370] The user tells the store clerk their order.

[1371] The terminal uses a voice recognition module to convert the conversation between the user and the store clerk into text and send it to the server.

[1372] The server analyzes the dialogue and extracts any additional information needed.

[1373] 11. Amendments to Order Details

[1374] The server generates alternative suggestions based on information such as "Pasta alla Carbonara is out of stock."

[1375] The terminal notifies the user of the alternatives, and the user selects and confirms the new menu.

[1376] Specific examples

[1377] 1. Menu photography:

[1378] A user takes a photo of a restaurant menu using a smartphone.

[1379] 2. Sending images:

[1380] The device encodes the captured menu image and sends it to the server's endpoint.

[1381] 3. Image analysis (OCR processing):

[1382] The server uses an OCR module to extract the string "Pasta alla Carbonara".

[1383] 4. Natural Language Processing (NLP):

[1384] The server generates a description for "Pasta alla Carbonara": "Italian pasta featuring a creamy sauce and bacon."

[1385] 5. Image generation:

[1386] The server uses a web search API to retrieve an image of "Pasta alla Carbonara."

[1387] 6. Display Menu:

[1388] The device displays "Pasta alla Carbonara," an Italian pasta dish featuring a creamy sauce and bacon, along with related images.

[1389] 7. Sending Selected Information:

[1390] The user selects "Pasta alla Carbonara" and the terminal transmits the selection information to the server.

[1391] 8. Generating recommendations:

[1392] The server will recommend suitable drinks and desserts, such as red wine and tiramisu.

[1393] 9. Display of Recommendations:

[1394] The device displays recommendations for red wine and tiramisu.

[1395] 10. Confirm your order and obtain additional information:

[1396] The user tells the store clerk their order.

[1397] The device uses a voice recognition module to convert the conversation into text.

[1398] 11. Order Modifications:

[1399] Based on the information that "Pasta alla Carbonara is out of stock," the server suggests "Pasta Bolognese."

[1400] The device presents the alternatives and the user selects "Pasta Bolognese."

[1401] This system makes it easier for users to understand foreign language menus and order the right food.

[1402] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1403] Step 1:

[1404] A user opens the camera on their smartphone or tablet and takes a picture of a restaurant menu. The input is the camera's photo image, and the output is an image file of the menu. Specifically, the user opens the camera app and presses the shutter button to capture the image.

[1405] Step 2:

[1406] The device encodes the captured menu image to send to the server and creates an HTTP POST request. The input is the menu image file, and the output is an HTTP request. Specifically, the device encodes the image into JPEG format and sends it to the server's endpoint (e.g., https: / / example.com / upload).

[1407] Step 3:

[1408] The server passes the received menu image to the OCR module, which extracts text from the image. The input is the image data of the menu, and the output is the extracted string. Specifically, the server sends the image data to the Google Cloud Vision API and obtains the text "Pasta alla Carbonara."

[1409] Step 4:

[1410] The server passes the extracted string to a natural language processing module to generate a description of the dish. The input is the extracted string, and the output is a description of the dish. Specifically, the server uses the NLP module to generate a description such as "Italian pasta with a creamy sauce and bacon" from "Pasta alla Carbonara."

[1411] Step 5:

[1412] The server uses a web search API to retrieve related images based on the results of natural language processing. The input is a description of the dish, and the output is an image of the dish. Specifically, the server retrieves images related to "Pasta alla Carbonara" using the Google Image Search API and selects the appropriate image.

[1413] Step 6:

[1414] The server compiles the obtained description and image in JSON format and sends it to the device. The input is the description and image of the dish, and the output is a JSON response. Specifically, the server generates JSON data such as { "name": "Pasta alla Carbonara", "description": "Italian pasta characterized by a creamy sauce and bacon", "image_url": "https: / / example.com / image.jpg"} and sends it to the device.

[1415] Step 7:

[1416] The device parses the JSON data received from the server and displays menu items, descriptions, and images on the user interface. The input is the JSON response received from the server, and the output is the user interface display. Specifically, the device uses an application to parse the received data and display it on the screen.

[1417] Step 8:

[1418] The user selects the dish of interest from the displayed menu, and the device sends the selection to the server. The input is the user's selection, and the output is an HTTP request. Specifically, the user taps "Pasta alla Carbonara" on the touchscreen, and the device sends the selection to the server.

[1419] Step 9:

[1420] The server uses a recommendation module to recommend related drinks and desserts based on the user's selection. The input is the user's selection, and the output is a list of recommended items. Specifically, the server uses a Collaborative Filtering algorithm to recommend red wine and tiramisu to go with "Pasta alla Carbonara."

[1421] Step 10:

[1422] The server compiles the generated recommendation results in JSON format and sends them to the device. The input is a list of recommended items, and the output is a JSON response. Specifically, the server generates JSON data such as { "recommendations": ["red wine", "tiramisu"]} and sends it to the device.

[1423] Step 11:

[1424] The device displays the recommendation information received from the server on the user interface, and the user confirms the displayed recommendation selection. The input is the JSON response received from the server, and the output is the display of the user interface and the user's selection action. Specifically, the device displays the received recommendation information, and the user taps the confirm button.

[1425] Step 12:

[1426] The user tells the store clerk their order, and the device converts the conversation into text using a voice recognition module and sends it to the server. The input is the voice conversation between the user and the store clerk, and the output is text data. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text.

[1427] Step 13:

[1428] The server analyzes the textual content of the conversation and obtains additional information. The input is text data, and the output is the analysis result. Specifically, the server analyzes the text and extracts information such as confirmation of order details and out-of-stock information.

[1429] Step 14:

[1430] The server generates alternatives as needed and notifies the terminal. The input is the analysis result information, and the output is the alternative information. In concrete terms, the server might suggest "Pasta Bolognese" based on the information "Pasta alla Carbonara is out of stock."

[1431] Step 15:

[1432] The device notifies the user of the alternatives, and the user selects and confirms the new menu. The input is the information about the alternatives, and the output is the user's selection action. Specifically, the device displays the alternatives, and the user selects a new menu and taps the confirm button.

[1433] (Application example 1)

[1434] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1435] In modern factories, many pieces of equipment and products use advanced technology, making it difficult for workers to grasp all the information and work efficiently. In particular, not being able to instantly obtain operating procedures and maintenance information related to equipment and products can lead to reduced production efficiency and the risk of operating errors. In addition, when working in a multilingual environment, language barriers become an additional obstacle. Therefore, there is a growing need for systems that can provide detailed information about equipment and products in the factory in real time.

[1436] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1437] In this invention, the server includes an image recognition unit, a natural language processing unit, and a character analysis unit. This allows real-time access to information about equipment and products in a factory using smart glasses or a head-mounted display. This allows workers to instantly access detailed operation guides and maintenance information about the equipment and products, improving work efficiency. Furthermore, by utilizing a generative AI model with prompt sentences, multilingual support is possible, enabling information provision across language barriers.

[1438] "Image recognition means" refers to a technology that uses a camera or other photographic device to acquire image data and extract specific features from the image.

[1439] "Natural language processing means" refers to technology that analyzes character strings and sentences and converts or interprets them into natural language that humans can understand.

[1440] A "recommendation tool" is a system that automatically recommends appropriate products and services based on a user's past behavior and choices.

[1441] "Speech recognition means" is a technology that analyzes speech and converts it into text data.

[1442] "Image generation means" refers to technology that creates new images based on data and information.

[1443] "Network communication means" refers to technology for sending and receiving data over the Internet or other communication networks.

[1444] "User interface means" refers to an interface such as a screen or input device that allows a user to interact with the system.

[1445] "Information acquisition means" refers to technology for acquiring information about equipment and products within a factory using smart glasses or a head-mounted display.

[1446] "Character analysis means" is a technology that uses OCR technology to analyze photographed character strings and convert them into digital text.

[1447] A "prompt" is a short command that a generative AI model receives as input and is an instruction to generate specific information.

[1448] The following system is required to implement this invention: The system is mainly composed of a server, a terminal, and a user.

[1449] The server includes an image recognition means, a natural language processing means, a character analysis means, a recommendation means, a voice recognition means, an image generation means, a network communication means, a generative AI model using prompt sentences, and the like.

[1450] The terminal is equipped with smart glasses or a head-mounted display, a camera, a microphone, and a user interface means, and is operated by the user.

[1451] The server first analyzes the label image of the equipment or product sent from the terminal using image recognition, which includes character analysis using OCR technology. Specifically, it uses Tesseract OCR to extract character strings from the image.

[1452] The extracted text is analyzed using natural language processing to generate an appropriate description, which can be translated into multiple languages ​​as needed using translation services such as Google Translate API.

[1453] The generated description and related images are sent to the terminal via the network communication means and displayed to the user by the user interface means. For example, if the character string "product 12345" is extracted, the details displayed will read, "Product 12345: This part has been machined using a high-precision grinder. Handle with care. Wear protective equipment and follow the operating instructions when using."

[1454] In addition, the recommendation means recommends appropriate operation guides and maintenance procedures based on the user's past selections and behavioral data, thereby further improving work efficiency.

[1455] An example of using a generative AI model with prompts is the following prompt:

[1456] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[1457] "Product 12345: This part has been processed using a high-precision grinder."

[1458] The system configuration and processing procedures described above enable detailed information on equipment and products within the factory to be provided in real time, improving work efficiency. In addition, the system's multilingual support makes it suitable for international work environments.

[1459] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1460] Step 1:

[1461] A user takes a picture of the labels of equipment or products in a factory using a camera in smart glasses or a head-mounted display. Here, the input is the label image, and the output is the image data captured by the camera.

[1462] Step 2:

[1463] The terminal sends the captured label image to the server. The input is label image data, and the output is image data sent to the server via the network.

[1464] Step 3:

[1465] The server uses image recognition and OCR technology to extract text from the received label image. Specifically, it uses Tesseract OCR. The input is the label image data, and the output is the extracted text.

[1466] Step 4:

[1467] The server passes the extracted string to a natural language processing means to generate an appropriate description. It translates it into multiple languages ​​as needed using NLP technology or the Google Translate API. The input is the extracted string, and the output is the generated description.

[1468] Step 5:

[1469] The server uses an image generation means to retrieve associated images along with the generated description, for example, the associated images are retrieved from a database or the Internet, where the input is the description and the output is the associated images.

[1470] Step 6:

[1471] The server transmits the generated explanatory text and related images to the terminal using a network communication means. The input is the generated explanatory text and related images, and the output is the data transmitted to the terminal.

[1472] Step 7:

[1473] The terminal displays the received explanatory text and image to the user through a user interface means. The input is the explanatory text and image data, and the output is the information displayed on the user interface.

[1474] Step 8:

[1475] The user checks the displayed information and performs operations to obtain more detailed information as necessary. For example, if a detailed operation guide is required, a prompt sentence is generated and the request is sent from the terminal to the server. The input is the user's operation, and the output is the generated prompt sentence and its request.

[1476] Step 9:

[1477] The server uses a generative AI model to generate detailed operation guides based on the prompt sentences and provides them to the user. The input is the prompt sentence, and the output is the detailed operation guides.

[1478] For specific actions, the following example prompt sentence is used:

[1479] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[1480] "Product 12345: This part has been processed using a high-precision grinder."

[1481] Through the above steps, users can obtain detailed information about the equipment and products in the factory in real time, improving work efficiency.

[1482] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1483] As an embodiment of the present invention, the specific implementation procedure of a system using a user, a terminal, a server, and an emotion engine is described below. This system assists users in the process of understanding foreign language menus and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[1484] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to the server. The server then passes the image to an image processing module, which uses OCR technology to extract the text from the image. For example, the string "Pasta alla Carbonara" is extracted.

[1485] The extracted text data is then passed to a natural language processing module. The server generates a description of the dish from this text data. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via a web search API or database. The server then sends the retrieved description and images to the device.

[1486] The device displays the received menu items, descriptions, and images to the user. The user checks the displayed menu and taps on the desired dish to select it. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts. The server sends the recommendation results to the device, which displays them to the user. The user checks the recommendations and confirms the order.

[1487] Furthermore, the system incorporates an emotion engine, which recognizes emotions from the user's facial expressions and voice. For example, if the user expresses a confused expression, the emotion engine analyzes the information. This emotion information is sent to the server, and menu suggestions and recommendations are adjusted based on the user's emotions. For example, if the user is confused, the system adjusts to provide more detailed explanations and additional information.

[1488] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion engine will recognize the emotion and the server will provide additional detailed explanations of the recommendation and alternatives based on the emotion.

[1489] After the order is confirmed, the user conveys the order to the store clerk. The voice recognition module is activated again, capturing the conversation between the user and the store clerk. The voice data is sent to the server, where it is converted into text data by the voice recognition module. The converted text data is then analyzed to extract additional information, such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user. The user then confirms the alternative options, selects a new menu item, and confirms the selection.

[1490] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[1491] The processing flow will be explained below.

[1492] Step 1:

[1493] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[1494] Step 2:

[1495] Device: Creates a request to send the captured menu image to the server.

[1496] Step 3:

[1497] Terminal: Uploads menu images to the server via the network.

[1498] Step 4:

[1499] Server: Passes the received menu image to the image processing module.

[1500] Step 5:

[1501] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[1502] Step 6:

[1503] Server: Passes the extracted text data to the natural language processing module.

[1504] Step 7:

[1505] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[1506] Step 8:

[1507] Server: Obtains images related to the dish along with the generated description.

[1508] Step 9:

[1509] Server: Sends the description and image to the device.

[1510] Step 10:

[1511] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[1512] Step 11:

[1513] User: Review the menu items displayed and tap to select the desired dish.

[1514] Step 12:

[1515] Terminal: Sends information about the menu item selected by the user to the server.

[1516] Step 13:

[1517] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[1518] Step 14:

[1519] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[1520] Step 15:

[1521] Server: Sends the recommendation results to the device.

[1522] Step 16:

[1523] Terminal: Display recommended drinks and desserts to the user.

[1524] Step 17:

[1525] User: Review the recommendations and indicate their intent to place an order.

[1526] Step 18:

[1527] Device: Activates the emotion engine to recognize emotions from the user's facial expressions and voice.

[1528] Step 19:

[1529] Terminal: The emotion engine analyzes the user's emotions and sends the results to the server.

[1530] Step 20:

[1531] Server: Receives user emotion data and adjusts recommendations and menu suggestions.

[1532] Step 21:

[1533] Server: Generates additional information and alternative suggestions based on emotion data and sends them to the device.

[1534] Step 22:

[1535] On the device: Display additional information or alternative suggestions to the user based on their emotions.

[1536] Step 23:

[1537] User: Review the proposed information and confirm the order.

[1538] Step 24:

[1539] User: Tells the waiter their order.

[1540] Step 25:

[1541] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[1542] Step 26:

[1543] Device: Sends captured audio data to the server.

[1544] Step 27:

[1545] Server: Passes the voice data to the voice recognition module and converts it into text data.

[1546] Step 28:

[1547] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[1548] Step 29:

[1549] Server: Based on the analysis results, generate alternatives as needed.

[1550] Step 30:

[1551] Server: Sends alternatives to the device.

[1552] Step 31:

[1553] Terminal: Present alternatives to the user.

[1554] Step 32:

[1555] User: Review alternatives, select new menu, confirm.

[1556] Example 2

[1557] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1558] For many users, understanding menus in a foreign language and ordering the appropriate dishes is a difficult task. Furthermore, if recommendations are not appropriate based on the user's emotions and preferences, the user experience cannot be improved. To solve these challenges, a system is needed that not only extracts menu text and provides easy-to-understand information, but also makes recommendations that adapt to the user's emotional state.

[1559] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1560] In this invention, the server includes image recognition means, natural language processing means, recommendation means, and emotion recognition means. This allows users to easily understand menus in foreign languages ​​and order appropriate dishes. Recommendations are also made that are adapted to the user's emotional state, allowing the user to enjoy personalized services.

[1561] "Image recognition means" refers to technology for analyzing an image and extracting specific information from its content.

[1562] "Natural language processing" refers to technology for understanding, generating, and manipulating human language.

[1563] "Recommendation means" refers to technology for recommending appropriate items based on a user's preferences and history.

[1564] "Speech recognition means" refers to technology for analyzing voice data and converting it into text data.

[1565] "Image generation means" refers to technology for generating new images based on input data.

[1566] "Emotion recognition means" refers to technology for analyzing and recognizing emotions from a user's facial expressions and voice.

[1567] "Network communication means" refers to technology for communicating over a network to send and receive data.

[1568] "User interface means" refers to technology that provides an interface for exchanging information between a user and a system.

[1569] The present invention relates to a system that assists users in the process of understanding menus in a foreign language and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[1570] Hardware and Software Configuration

[1571] Required Hardware

[1572] User's device: A smartphone or tablet equipped with a camera is used.

[1573] Server: A high-performance server is used for menu image analysis and recommendation processing.

[1574] Required software

[1575] Image Recognition Method: OCR technology is used to extract text from images, specifically Google Cloud Vision API.

[1576] Natural language processing tools: A generative AI model (e.g., GPT-4) is used to generate dish descriptions from the extracted text data.

[1577] Recommendation method: Use a recommendation algorithm to recommend drinks and desserts related to the dish.

[1578] Emotion Recognition: Emotion recognition software is used to analyze and recognize emotions from the user's facial expressions and voice.

[1579] Speech recognition means: Use speech recognition technology (e.g., Google Speech-to-Text API) to capture the conversation between the user and the store clerk as voice data and convert it into text data.

[1580] Image generation method: Use a web search API (e.g., Bing Image Search API) to obtain related images.

[1581] Network communication means: Uses communication technology to send and receive data between the user terminal and the server.

[1582] User interface means: Provides an interface for exchanging information between the user and the system.

[1583] Example of a system

[1584] 1. Menu Parsing Prompt Example

[1585] "Analyze a restaurant menu image and generate dish names and descriptions."

[1586] 2. Recommendation prompt examples

[1587] "Recommend drinks and desserts to go with the selected dish."

[1588] 3. Example of emotion-responsive prompt

[1589] "If the user looks confused, provide a detailed explanation."

[1590] System Operation

[1591] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to a server. The server then analyzes the received image using the Google Cloud Vision API and extracts the text from the image using OCR technology. For example, the string "Pasta alla Carbonara" is extracted.

[1592] The extracted text data is then passed to a generative AI model such as GPT-4 to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via the Bing Image Search API. The server then sends the retrieved description and image to the device, which then displays them to the user.

[1593] The user checks the displayed menu items, descriptions, and images, and taps to select the desired dish. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts, and sends the recommendation results to the device. The device then displays them to the user.

[1594] The app also has a built-in emotion engine that recognizes emotions from the user's facial expressions and voice. For example, if the user shows a confused expression, the emotion engine analyzes the information and sends it to the server. Based on this emotion information, the server adjusts the service to provide detailed explanations and additional information to the user.

[1595] Finally, the user confirms the order and conveys it to the store clerk. The voice recognition module is activated, capturing the conversation between the user and the store clerk, and the voice data is sent to the server. The server's voice recognition module converts the voice data into text data, which is then analyzed to extract additional information such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user.

[1596] In this way, the present invention allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's feelings.

[1597] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1598] Step 1:

[1599] Users take a photo of a restaurant menu using the camera on their smartphone or tablet.

[1600] Input: Menu Image

[1601] Specific actions: The user opens the camera app, points the lens at the menu, and presses the shutter button to take a photo.

[1602] Output: Menu image captured

[1603] Step 2:

[1604] The terminal transmits the captured menu image to the server.

[1605] Input: Menu image taken

[1606] Specific operation: After taking a photo, the image will be automatically uploaded and sent.

[1607] Output: Menu image sent to the server

[1608] Step 3:

[1609] The server analyzes the received menu image using OCR technology and extracts the text within the image.

[1610] Input: Menu image sent to the server

[1611] Data processing: Call the Google Cloud Vision API to extract character codes from images

[1612] What it does: The API analyzes the text in the menu image and extracts text such as "Pasta alla Carbonara."

[1613] Output: Extracted text data

[1614] Step 4:

[1615] The server passes the extracted text data to a natural language processing module to generate a description of the dish.

[1616] Input: Extracted text data

[1617] Data processing: Calling GPT-4 to generate explanatory text from text data

[1618] What it does: GPT-4 takes the text "Pasta alla Carbonara" and generates a description: "Italian pasta with a creamy sauce and bacon."

[1619] Output: Generated dish description

[1620] Step 5:

[1621] The server retrieves related images via a web search API.

[1622] Input: Description of the generated dish

[1623] Data processing: Call the Bing Image Search API to get related images

[1624] What it does: Uses the Bing Image Search API to search and retrieve relevant images based on the description of "Pasta alla Carbonara."

[1625] Output: Captured image

[1626] Step 6:

[1627] The server transmits the explanatory text and the image to the terminal, which displays them to the user.

[1628] Input: Generated description and retrieved image

[1629] Specific operation: Data is sent from the server to the device, and the device displays a pop-up explanation and image on the screen.

[1630] Output: Description and image displayed to the user

[1631] Step 7:

[1632] The user taps on the desired dish to select it, and the device sends the selection information to the server.

[1633] Input: Displayed description and image

[1634] What happens: The user taps "Pasta alla Carbonara" and the selection is sent to the server.

[1635] Output: Selections sent to the server

[1636] Step 8:

[1637] The server uses a recommendation module to recommend related drinks and desserts and transmits them to the terminal.

[1638] Input: Selections sent to the server

[1639] Data processing: Use recommendation algorithms to select relevant drinks and desserts

[1640] Specific operation: The server selects red wine and tiramisu, and the recommended results are displayed on the device screen.

[1641] Output: Recommendation results displayed to the user

[1642] Step 9:

[1643] The emotion engine recognizes emotions from the user's facial expressions and voice, and transmits the emotion information to the server.

[1644] Input: User's facial expressions and voice

[1645] Data processing: Using emotion recognition software, analyze and recognize the user's emotions.

[1646] Specific operation: When the user makes a confused expression, the camera analyzes the expression and sends the emotional information to the server, which then displays a detailed explanation.

[1647] Output: Additional explanation based on emotion information

[1648] Step 10:

[1649] The user confirms the final order and communicates it to the store attendant. The voice recognition module captures the conversation and sends it to the server.

[1650] Input: User and store clerk conversation

[1651] Data processing: Using voice recognition technology, converting voice data into text data

[1652] What it does: The user verbally places an order, and the speech is sent to the server, which converts it into text. The server then provides out-of-stock information and alternative suggestions.

[1653] Output: Textualized dialogue, out-of-stock information, and alternatives

[1654] Through this series of processes, users can easily understand menus in foreign languages ​​and order the most suitable dishes based on their emotions.

[1655] (Application example 2)

[1656] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1657] In existing systems, users may have difficulty understanding foreign language menus and selecting appropriate dishes, and they are unable to provide personalized recommendations based on the user's emotions. The present invention aims to solve these problems, enabling users to easily understand foreign language menus and providing personalized services based on the user's emotions.

[1658] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means. This allows the user to easily understand foreign language menus and further receive personalized recommendations based on their emotions.

[1659] "Image recognition means" is a device that extracts information from a captured image and analyzes its content.

[1660] "Natural language processing means" is a device that analyzes extracted text data, converts it into natural language format, and understands and processes it.

[1661] A "recommendation means" is a device that makes relevant suggestions or recommendations to the user based on the analyzed information.

[1662] A "voice recognition means" is a device that analyzes voice data and converts it into a linguistic text format.

[1663] The "image generating means" is a device that generates an image based on the analyzed information and generated data.

[1664] A "network communication means" is a communication device for transmitting and receiving data between systems.

[1665] "User interface means" refers to a device that allows a user to operate and interact with the system.

[1666] The "emotion analysis means" is a device that recognizes and analyzes the user's emotions from their facial expressions and tone of voice.

[1667] The present invention provides a system that assists users in understanding menus written in a foreign language and ordering appropriate dishes. The system includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means.

[1668] First, a user takes a photo of a restaurant menu using a smartphone or tablet device. The image taken with the device's camera is sent from the device to a server via network communication means. The server then uses image recognition means to extract text from the received image. For example, the text "Pasta alla Carbonara" is extracted.

[1669] Next, the extracted text data is analyzed using natural language processing to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta characterized by a creamy sauce and bacon." At the same time, related images are retrieved via a web search API or database. The server then sends the generated description and the retrieved images to the device via the network.

[1670] The terminal uses a user interface means to display the received menu items, descriptions, and images to the user. The user checks the displayed menu and selects the desired dish. Information about the selected menu item is again sent from the terminal to the server. The server uses a recommendation means to recommend additional related drinks and desserts. The recommendation results are sent to the terminal and displayed to the user.

[1671] The system also incorporates an emotion analysis mechanism that recognizes emotions from the user's facial expressions and voice and adjusts menu suggestions and recommendations based on that information. For example, if the user expresses confusion, the emotion analysis mechanism sends that information to the server, which then adjusts the system to provide detailed explanations and additional information.

[1672] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion analyzer recognizes the emotion and the server provides additional detailed explanations of the recommendation and alternatives.

[1673] After the order is confirmed, the user tells the store clerk what they want to order. At this time, the voice recognition means is activated and the conversation between the user and the store clerk is captured. The voice data is sent back to the server and converted into text data by the voice recognition means. The converted text data is analyzed to extract information such as out-of-stock items and additional orders. If necessary, alternative options are generated and notified to the user. The user then selects a new menu item and confirms the order.

[1674] An example prompt is, "Use the EmotionRecognition library to analyze emotions from the image menu.jpg and generate personalized dish descriptions based on that."

[1675] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[1676] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1677] Step 1:

[1678] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[1679] Input: Menu Image

[1680] Output: Raw image data in the device

[1681] Specific operation: A user opens the device's camera app and takes a picture of a restaurant menu. After taking the picture, the image data is temporarily stored in the device's memory.

[1682] Step 2:

[1683] The terminal transmits the photographed menu image to the server using the network communication means.

[1684] Input: Raw image data in the device

[1685] Output: Image data sent to the server

[1686] How it works: The device sends image data to the server via Wi-Fi or mobile data using the HTTP or HTTPS protocol.

[1687] Step 3:

[1688] The server uses image recognition means to extract text from the received image.

[1689] Input: Image data sent to the server

[1690] Output: Extracted text data

[1691] How it works: Optical Character Recognition (OCR) software on the server analyzes the image and extracts text information, which is then stored in a database on the server, such as "Pasta alla Carbonara."

[1692] Step 4:

[1693] The server uses natural language processing means to generate a description of the dish from the extracted text data.

[1694] Input: Extracted text data

[1695] Output: The generated dish description

[1696] What it does: A server-based natural language processing library analyzes the text data and generates a description of the corresponding dish, such as "Italian pasta with a creamy sauce and bacon."

[1697] Step 5:

[1698] The server retrieves related images via a web search API or database.

[1699] Input: Extracted text data

[1700] Output: Related images

[1701] Specific operation: The server calls the web search API, searches for relevant images based on the text data, selects the most appropriate image from the search results, and downloads it.

[1702] Step 6:

[1703] The server transmits the generated description and image to the terminal using a network communication means.

[1704] Input: Generated description and associated image

[1705] Output: Description and image sent to the device

[1706] Specific operation: The server uses the HTTP or HTTPS protocol to send the description and image to the terminal.

[1707] Step 7:

[1708] The terminal uses the user interface means to display the received menu items, explanations, and images to the user.

[1709] Input: Received description and image

[1710] Output: Menu items, descriptions, and images displayed in the user interface

[1711] Specific behavior: The device uses the layout template to properly arrange the displayed information on the screen, allowing the user to visually confirm the displayed information.

[1712] Step 8:

[1713] The user checks the displayed menu items and taps to select the desired dish.

[1714] Input: User tap input

[1715] Output: Information about the selected menu item

[1716] Specific operation: The information of the menu item selected by the user is stored in the terminal's memory in preparation for the next communication step.

[1717] Step 9:

[1718] The terminal transmits information about the selected menu item to the server using the network communication means.

[1719] Input: Information of the selected menu item

[1720] Output: Selections sent to the server

[1721] Specific operation: The device sends the selection information to the server via Wi-Fi or mobile data communication.

[1722] Step 10:

[1723] The server uses a recommendation means to recommend related drinks and desserts.

[1724] Input: Information of the selected menu item

[1725] Output: Recommended drinks and desserts

[1726] Specific operation: The server's recommendation engine searches for items related to the selected menu and generates a recommendation list.

[1727] Step 11:

[1728] The server transmits the recommendation results to the terminal using a network communication means.

[1729] Input: Recommended drink and dessert information

[1730] Output: Recommendation results sent to the device

[1731] Specific operation: The server uses HTTP or HTTPS protocol to send the recommendation results to the terminal.

[1732] Step 12:

[1733] The terminal uses the user interface means to display the recommendation results to the user.

[1734] Input: Received recommendation results

[1735] Output: Recommendation results displayed in the user interface

[1736] Specific behavior: The recommended drinks and desserts will be displayed on the device screen for the user to review.

[1737] Step 13:

[1738] The server uses emotion analysis means to recognize emotions from the user's facial expressions and voice.

[1739] Input: User's facial expressions or voice data

[1740] Output: Recognized emotion information

[1741] Specific operation: The server uses a facial expression recognition library and a voice analysis library to analyze the user's emotions and saves the recognition results in a database.

[1742] Step 14:

[1743] The server adjusts menu suggestions and recommendations based on the recognized emotional information.

[1744] Input: Recognized emotion information

[1745] Output: Tailored menu suggestions and recommendations

[1746] Specific operation: Based on the results of the sentiment analysis, the server generates detailed information and alternatives and provides them to the user.

[1747] Step 15:

[1748] The user finally confirms and confirms the menu item.

[1749] Input: User confirmed input

[1750] Output: Confirmed order information

[1751] Specific operation: When the user selects a menu item and taps the confirmation button, the order information is finally captured and saved in the device's memory.

[1752] Step 16:

[1753] When the terminal communicates the order information to the store clerk, the voice recognition means captures the content of the conversation.

[1754] Input: Voice data of the conversation between the user and the store clerk

[1755] Output: Captured audio data

[1756] Specific operation: The device's microphone records the conversation between the user and the store clerk and saves it as audio data.

[1757] Step 17:

[1758] The server receives the voice data and converts it into text data using a voice recognition means.

[1759] Input: Captured audio data

[1760] Output: Converted text data

[1761] Specific operation: A speech recognition library on the server analyzes the voice data and generates corresponding text.

[1762] Step 18:

[1763] The server analyzes the converted text data and extracts out-of-stock information and additional orders.

[1764] Input: Converted text data

[1765] Output: Out-of-stock information and reorder notifications

[1766] Specific operation: The server analyzes the text data, extracts the necessary information, and generates alternative menus and additional order information as necessary.

[1767] Step 19:

[1768] The server informs the user of the alternatives and the user selects a new menu.

[1769] Input: Out-of-stock information and reorder notifications

[1770] Output: New menu selection information

[1771] Specific operation: The user confirms the alternatives and sends the newly selected menu information from the terminal to the server.

[1772] In this way, the present invention helps users understand foreign language menus and provides personalized services based on emotions.

[1773] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1774] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1775] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1776] [Fourth embodiment]

[1777] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1778] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1779] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1780] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1781] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1782] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1783] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1784] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1785] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1786] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1787] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1788] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1789] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1790] As an embodiment of the present invention, the system program is implemented according to the following procedure: The system is mainly composed of a server, a terminal, and a user.

[1791] 1. Photographing the menu

[1792] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[1793] 2. Sending images

[1794] Device: Creates a request to send the captured menu image to the server.

[1795] Terminal: Uploads menu images to the server via the network.

[1796] 3. Image analysis (OCR processing)

[1797] Server: Passes the received menu image to the image processing module.

[1798] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[1799] 4. Natural Language Processing (NLP)

[1800] Server: Passes the extracted string to the natural language processing module.

[1801] Server: Generates a dish description from a string, for example "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[1802] 5. Image Generation

[1803] Server: Based on the string, retrieve related images using a web search API or database.

[1804] Server: Sends the acquired image and description to the device.

[1805] 6. Displaying the menu

[1806] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[1807] User: Check the menu that appears and tap the desired dish.

[1808] 7. Sending Selected Information

[1809] Terminal: Sends information about the menu item selected by the user to the server.

[1810] 8. Generating Recommendations

[1811] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[1812] Server: Receives the recommendation results from the modules and selects the drinks and desserts that are best suited to the user.

[1813] 9. Display of Recommendations

[1814] Terminal: Displays the drink and dessert recommendations received from the server to the user.

[1815] User: Review the recommendations and indicate their intent to place an order.

[1816] 10. Confirming your order and obtaining additional information

[1817] User: Tells the waiter their order.

[1818] Terminal: Passes the conversation between the user and the store clerk to a speech recognition module, which converts the speech into text.

[1819] Server: Parse the converted text and extract additional information, such as order confirmation and out-of-stock information.

[1820] 11. Amendments to Order Details

[1821] Server: If the store clerk's response includes information such as "out of stock," generate an alternative.

[1822] Terminal: Inform the user of the alternatives.

[1823] User: Review the alternatives, select a new menu option, and confirm.

[1824] Specific examples

[1825] 1. Photographing the menu

[1826] User: Take a photo of the menu with their smartphone.

[1827] 2. Sending images

[1828] Device: Creates a request to send the captured menu image to the server.

[1829] On your device: Upload the menu image to the server.

[1830] 3. Image analysis (OCR processing)

[1831] Server: Passes the menu image to the OCR processing module.

[1832] Server: Extract the string "Pasta alla Carbonara".

[1833] 4. Natural Language Processing (NLP)

[1834] Server: Passes the extracted string to the NLP module.

[1835] Server: Generates the description "Italian pasta featuring a creamy sauce and bacon" from "Pasta alla Carbonara."

[1836] 5. Image Generation

[1837] Server: Uses a web search API based on the string to retrieve related images.

[1838] Server: Sends the description and image to the device.

[1839] 6. Displaying the menu

[1840] Terminal: Display "Pasta alla Carbonara," "Italian pasta featuring a creamy sauce and bacon," and related images.

[1841] 7. Sending Selected Information

[1842] User: Select "Pasta alla Carbonara."

[1843] Terminal: Sends the selection information to the server.

[1844] 8. Generating Recommendations

[1845] Server: Recommend suitable drinks and desserts based on your selections.

[1846] Server: Recommend red wine or tiramisu.

[1847] 9. Display of Recommendations

[1848] Terminal: Display recommendations for red wine and tiramisu to the user.

[1849] 10. Confirming your order and obtaining additional information

[1850] User: Tells the waiter their order.

[1851] Terminal: The speech recognition module converts the conversation into text.

[1852] Server: Extracts additional information from the parsed text.

[1853] 11. Amendments to Order Details

[1854] Server: Based on the information "Pasta alla Carbonara is out of stock", generate "Pasta Bolognese" as an alternative.

[1855] Terminal: Inform the user of the alternatives.

[1856] User: Review the alternatives, select "Pasta Bolognese" and confirm.

[1857] In this way, a system is provided that allows users to easily understand menus written in foreign languages ​​and order appropriate dishes.

[1858] The processing flow will be explained below.

[1859] Step 1:

[1860] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[1861] Step 2:

[1862] Device: Creates a request to send the captured menu image to the server.

[1863] Step 3:

[1864] Terminal: Uploads menu images to the server via the network.

[1865] Step 4:

[1866] Server: Passes the received menu image to the image processing module.

[1867] Step 5:

[1868] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[1869] Step 6:

[1870] Server: Passes the extracted text data to the natural language processing module.

[1871] Step 7:

[1872] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[1873] Step 8:

[1874] Server: Obtains images related to the dish along with the generated description.

[1875] Step 9:

[1876] Server: Sends the description and image to the device.

[1877] Step 10:

[1878] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[1879] Step 11:

[1880] User: Review the menu items displayed and tap to select the desired dish.

[1881] Step 12:

[1882] Terminal: Sends information about the menu item selected by the user to the server.

[1883] Step 13:

[1884] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[1885] Step 14:

[1886] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[1887] Step 15:

[1888] Server: Sends the recommendation results to the device.

[1889] Step 16:

[1890] Terminal: Display recommended drinks and desserts to the user.

[1891] Step 17:

[1892] User: Review the recommendations and indicate their intent to place an order.

[1893] Step 18:

[1894] User: Inform the store clerk of the confirmed order.

[1895] Step 19:

[1896] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[1897] Step 20:

[1898] Device: Sends captured audio data to the server.

[1899] Step 21:

[1900] Server: Passes the voice data to the voice recognition module and converts it into text data.

[1901] Step 22:

[1902] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[1903] Step 23:

[1904] Server: Based on the analysis results, generate alternatives as needed.

[1905] Step 24:

[1906] Server: Sends alternatives to the device.

[1907] Step 25:

[1908] Terminal: Present alternatives to the user.

[1909] Step 26:

[1910] User: Review alternatives, select new menu, confirm.

[1911] Example 1

[1912] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1913] In modern society, many people face significant challenges in understanding menus written in a foreign language and ordering the appropriate food. This problem is particularly pronounced for tourists and people unfamiliar with the foreign language, leading to ordering errors and confusion. Furthermore, selecting the right drink or dessert presents similar challenges. The present invention aims to solve these problems and provide a smooth ordering process.

[1914] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1915] In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, and user interface means, which enable the server to extract text from menus written in a foreign language to generate dish descriptions, recommend related drinks and desserts, convert voice dialogue into text, and display this information on the user interface.

[1916] "Image recognition means" refers to a device or program that performs processing to extract text from an image captured using a camera.

[1917] "Natural language processing means" refers to a device or program that analyzes the extracted text and generates descriptions of dishes and related information.

[1918] A "recommendation means" is a device or program for recommending related food and beverage items based on a menu item selected by a user.

[1919] "Speech recognition means" refers to a device or program for converting the dialogue between the user and the store clerk from voice to text.

[1920] The "image generating means" refers to a device or program for obtaining an image related to a menu item and displaying it to the user.

[1921] "Network communication means" refers to a communication device or program for transmitting and receiving data between a terminal and a server.

[1922] "User interface means" refers to a device or program that provides a display screen and input means for a user to operate the system.

[1923] The system of the present invention is composed of a user, a terminal, and a server. By using this system, the user can understand menus written in a foreign language and order the appropriate food.

[1924] Specific examples of hardware and software used

[1925] Hardware

[1926] 1. Camera-equipped devices (smartphones, tablets)

[1927] 2. Server (Cloud server, local server)

[1928] software

[1929] 1. OCR module (Google Cloud Vision API, Tesseract OCR)

[1930] 2. Natural language processing module (IBM Watson NLP, Google Cloud Natural Language API)

[1931] 3. Recommendation module (Collaborative Filtering, Content-Based Filtering algorithms)

[1932] 4. Speech Recognition Module (Google Speech-to-Text API, IBM Watson Speech to Text)

[1933] 5. Network communication module (HTTP / HTTPS communication library)

[1934] 6. User Interface Module (React Native, Flutter)

[1935] Specific operation of the system

[1936] 1. Photographing the menu

[1937] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[1938] 2. Sending images

[1939] The device encodes the captured menu image and creates a request to send to the server, specifically uploading the image to the server's endpoint using an HTTP POST request.

[1940] 3. Image analysis (OCR processing)

[1941] The server passes the received image to an OCR module, which extracts text from the image, for example, the string "Pasta alla Carbonara."

[1942] 4. Natural Language Processing (NLP)

[1943] The server passes the extracted text to a natural language processing module to generate a description of the dish, for example, "Pasta alla Carbonara" to "Italian pasta with a creamy sauce and bacon."

[1944] 5. Image Generation

[1945] The server uses a web search API to retrieve related images based on the results of natural language processing, and sends the images and descriptions together to the device.

[1946] 6. Displaying the menu

[1947] The terminal displays the menu items, explanations, and related images received from the server on the user interface, providing information in a format that is easy for the user to select.

[1948] 7. Sending Selected Information

[1949] The user selects the dish of interest and confirms the information.

[1950] The terminal transmits the user's selection information to the server.

[1951] 8. Generating Recommendations

[1952] Based on the user's selection, the server uses a recommendation module to recommend suitable drinks and desserts, for example, red wine or tiramisu that go well with "Pasta alla Carbonara."

[1953] 9. Display of Recommendations

[1954] The terminal displays the information recommended by the server to the user, who can then confirm the recommendation and place an order.

[1955] 10. Confirming your order and obtaining additional information

[1956] The user tells the store clerk their order.

[1957] The terminal uses a voice recognition module to convert the conversation between the user and the store clerk into text and send it to the server.

[1958] The server analyzes the dialogue and extracts any additional information needed.

[1959] 11. Amendments to Order Details

[1960] The server generates alternative suggestions based on information such as "Pasta alla Carbonara is out of stock."

[1961] The terminal notifies the user of the alternatives, and the user selects and confirms the new menu.

[1962] Specific examples

[1963] 1. Menu photography:

[1964] A user takes a photo of a restaurant menu using a smartphone.

[1965] 2. Sending images:

[1966] The device encodes the captured menu image and sends it to the server's endpoint.

[1967] 3. Image analysis (OCR processing):

[1968] The server uses an OCR module to extract the string "Pasta alla Carbonara".

[1969] 4. Natural Language Processing (NLP):

[1970] The server generates a description for "Pasta alla Carbonara": "Italian pasta featuring a creamy sauce and bacon."

[1971] 5. Image generation:

[1972] The server uses a web search API to retrieve an image of "Pasta alla Carbonara."

[1973] 6. Display Menu:

[1974] The device displays "Pasta alla Carbonara," an Italian pasta dish featuring a creamy sauce and bacon, along with related images.

[1975] 7. Sending Selected Information:

[1976] The user selects "Pasta alla Carbonara" and the terminal transmits the selection information to the server.

[1977] 8. Generating recommendations:

[1978] The server will recommend suitable drinks and desserts, such as red wine and tiramisu.

[1979] 9. Display of Recommendations:

[1980] The device displays recommendations for red wine and tiramisu.

[1981] 10. Confirm your order and obtain additional information:

[1982] The user tells the store clerk their order.

[1983] The device uses a voice recognition module to convert the conversation into text.

[1984] 11. Order Modifications:

[1985] Based on the information that "Pasta alla Carbonara is out of stock," the server suggests "Pasta Bolognese."

[1986] The device presents the alternatives and the user selects "Pasta Bolognese."

[1987] This system makes it easier for users to understand foreign language menus and order the right food.

[1988] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1989] Step 1:

[1990] A user opens the camera on their smartphone or tablet and takes a picture of a restaurant menu. The input is the camera's photo image, and the output is an image file of the menu. Specifically, the user opens the camera app and presses the shutter button to capture the image.

[1991] Step 2:

[1992] The device encodes the captured menu image to send to the server and creates an HTTP POST request. The input is the menu image file, and the output is an HTTP request. Specifically, the device encodes the image into JPEG format and sends it to the server's endpoint (e.g., https: / / example.com / upload).

[1993] Step 3:

[1994] The server passes the received menu image to the OCR module, which extracts text from the image. The input is the image data of the menu, and the output is the extracted string. Specifically, the server sends the image data to the Google Cloud Vision API and obtains the text "Pasta alla Carbonara."

[1995] Step 4:

[1996] The server passes the extracted string to a natural language processing module to generate a description of the dish. The input is the extracted string, and the output is a description of the dish. Specifically, the server uses the NLP module to generate a description such as "Italian pasta with a creamy sauce and bacon" from "Pasta alla Carbonara."

[1997] Step 5:

[1998] The server uses a web search API to retrieve related images based on the results of natural language processing. The input is a description of the dish, and the output is an image of the dish. Specifically, the server retrieves images related to "Pasta alla Carbonara" using the Google Image Search API and selects the appropriate image.

[1999] Step 6:

[2000] The server compiles the obtained description and image in JSON format and sends it to the device. The input is the description and image of the dish, and the output is a JSON response. Specifically, the server generates JSON data such as { "name": "Pasta alla Carbonara", "description": "Italian pasta characterized by a creamy sauce and bacon", "image_url": "https: / / example.com / image.jpg"} and sends it to the device.

[2001] Step 7:

[2002] The device parses the JSON data received from the server and displays menu items, descriptions, and images on the user interface. The input is the JSON response received from the server, and the output is the user interface display. Specifically, the device uses an application to parse the received data and display it on the screen.

[2003] Step 8:

[2004] The user selects the dish of interest from the displayed menu, and the device sends the selection to the server. The input is the user's selection, and the output is an HTTP request. Specifically, the user taps "Pasta alla Carbonara" on the touchscreen, and the device sends the selection to the server.

[2005] Step 9:

[2006] The server uses a recommendation module to recommend related drinks and desserts based on the user's selection. The input is the user's selection, and the output is a list of recommended items. Specifically, the server uses a Collaborative Filtering algorithm to recommend red wine and tiramisu to go with "Pasta alla Carbonara."

[2007] Step 10:

[2008] The server compiles the generated recommendation results in JSON format and sends them to the device. The input is a list of recommended items, and the output is a JSON response. Specifically, the server generates JSON data such as { "recommendations": ["red wine", "tiramisu"]} and sends it to the device.

[2009] Step 11:

[2010] The device displays the recommendation information received from the server on the user interface, and the user confirms the displayed recommendation selection. The input is the JSON response received from the server, and the output is the display of the user interface and the user's selection action. Specifically, the device displays the received recommendation information, and the user taps the confirm button.

[2011] Step 12:

[2012] The user tells the store clerk their order, and the device converts the conversation into text using a voice recognition module and sends it to the server. The input is the voice conversation between the user and the store clerk, and the output is text data. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text.

[2013] Step 13:

[2014] The server analyzes the textual content of the conversation and obtains additional information. The input is text data, and the output is the analysis result. Specifically, the server analyzes the text and extracts information such as confirmation of order details and out-of-stock information.

[2015] Step 14:

[2016] The server generates alternatives as needed and notifies the terminal. The input is the analysis result information, and the output is the alternative information. In concrete terms, the server might suggest "Pasta Bolognese" based on the information "Pasta alla Carbonara is out of stock."

[2017] Step 15:

[2018] The device notifies the user of the alternatives, and the user selects and confirms the new menu. The input is the information about the alternatives, and the output is the user's selection action. Specifically, the device displays the alternatives, and the user selects a new menu and taps the confirm button.

[2019] (Application example 1)

[2020] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2021] In modern factories, many pieces of equipment and products use advanced technology, making it difficult for workers to grasp all the information and work efficiently. In particular, not being able to instantly obtain operating procedures and maintenance information related to equipment and products can lead to reduced production efficiency and the risk of operating errors. In addition, when working in a multilingual environment, language barriers become an additional obstacle. Therefore, there is a growing need for systems that can provide detailed information about equipment and products in the factory in real time.

[2022] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2023] In this invention, the server includes an image recognition unit, a natural language processing unit, and a character analysis unit. This allows real-time access to information about equipment and products in a factory using smart glasses or a head-mounted display. This allows workers to instantly access detailed operation guides and maintenance information about the equipment and products, improving work efficiency. Furthermore, by utilizing a generative AI model with prompt sentences, multilingual support is possible, enabling information provision across language barriers.

[2024] "Image recognition means" refers to a technology that uses a camera or other photographic device to acquire image data and extract specific features from the image.

[2025] "Natural language processing means" refers to technology that analyzes character strings and sentences and converts or interprets them into natural language that humans can understand.

[2026] A "recommendation tool" is a system that automatically recommends appropriate products and services based on a user's past behavior and choices.

[2027] "Speech recognition means" is a technology that analyzes speech and converts it into text data.

[2028] "Image generation means" refers to technology that creates new images based on data and information.

[2029] "Network communication means" refers to technology for sending and receiving data over the Internet or other communication networks.

[2030] "User interface means" refers to an interface such as a screen or input device that allows a user to interact with the system.

[2031] "Information acquisition means" refers to technology for acquiring information about equipment and products within a factory using smart glasses or a head-mounted display.

[2032] "Character analysis means" is a technology that uses OCR technology to analyze photographed character strings and convert them into digital text.

[2033] A "prompt" is a short command that a generative AI model receives as input and is an instruction to generate specific information.

[2034] The following system is required to implement this invention: The system is mainly composed of a server, a terminal, and a user.

[2035] The server includes an image recognition means, a natural language processing means, a character analysis means, a recommendation means, a voice recognition means, an image generation means, a network communication means, a generative AI model using prompt sentences, and the like.

[2036] The terminal is equipped with smart glasses or a head-mounted display, a camera, a microphone, and a user interface means, and is operated by the user.

[2037] The server first analyzes the label image of the equipment or product sent from the terminal using image recognition, which includes character analysis using OCR technology. Specifically, it uses Tesseract OCR to extract character strings from the image.

[2038] The extracted text is analyzed using natural language processing to generate an appropriate description, which can be translated into multiple languages ​​as needed using translation services such as Google Translate API.

[2039] The generated description and related images are sent to the terminal via the network communication means and displayed to the user by the user interface means. For example, if the character string "product 12345" is extracted, the details displayed will read, "Product 12345: This part has been machined using a high-precision grinder. Handle with care. Wear protective equipment and follow the operating instructions when using."

[2040] In addition, the recommendation means recommends appropriate operation guides and maintenance procedures based on the user's past selections and behavioral data, thereby further improving work efficiency.

[2041] An example of using a generative AI model with prompts is the following prompt:

[2042] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[2043] "Product 12345: This part has been processed using a high-precision grinder."

[2044] The system configuration and processing procedures described above enable detailed information on equipment and products within the factory to be provided in real time, improving work efficiency. In addition, the system's multilingual support makes it suitable for international work environments.

[2045] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2046] Step 1:

[2047] A user takes a picture of the labels of equipment or products in a factory using a camera in smart glasses or a head-mounted display. Here, the input is the label image, and the output is the image data captured by the camera.

[2048] Step 2:

[2049] The terminal sends the captured label image to the server. The input is label image data, and the output is image data sent to the server via the network.

[2050] Step 3:

[2051] The server uses image recognition and OCR technology to extract text from the received label image. Specifically, it uses Tesseract OCR. The input is the label image data, and the output is the extracted text.

[2052] Step 4:

[2053] The server passes the extracted string to a natural language processing means to generate an appropriate description. It translates it into multiple languages ​​as needed using NLP technology or the Google Translate API. The input is the extracted string, and the output is the generated description.

[2054] Step 5:

[2055] The server uses an image generation means to retrieve associated images along with the generated description, for example, the associated images are retrieved from a database or the Internet, where the input is the description and the output is the associated images.

[2056] Step 6:

[2057] The server transmits the generated explanatory text and related images to the terminal using a network communication means. The input is the generated explanatory text and related images, and the output is the data transmitted to the terminal.

[2058] Step 7:

[2059] The terminal displays the received explanatory text and image to the user through a user interface means, with the input being the explanatory text and image data, and the output being the information displayed on the user interface.

[2060] Step 8:

[2061] The user checks the displayed information and performs operations to obtain more detailed information as necessary. For example, if a detailed operation guide is required, a prompt sentence is generated and the request is sent from the terminal to the server. The input is the user's operation, and the output is the generated prompt sentence and its request.

[2062] Step 9:

[2063] The server uses a generative AI model based on the prompt sentence to generate a detailed operation guide and provide it to the user. The input is the prompt sentence, and the output is the detailed operation guide.

[2064] For specific actions, the following example prompt sentence is used:

[2065] Translate the following English text into Japanese and provide a detailed operational guide based on the content:

[2066] "Product 12345: This part has been processed using a high-precision grinder."

[2067] Through the above steps, users can obtain detailed information about the equipment and products in the factory in real time, improving work efficiency.

[2068] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2069] As an embodiment of the present invention, the specific implementation procedure of a system using a user, a terminal, a server, and an emotion engine is described below. This system assists users in the process of understanding foreign language menus and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[2070] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to the server. The server then passes the image to an image processing module, which uses OCR technology to extract the text from the image. For example, the string "Pasta alla Carbonara" is extracted.

[2071] The extracted text data is then passed to a natural language processing module. The server generates a description of the dish from this text data. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via a web search API or database. The server then sends the retrieved description and images to the device.

[2072] The device displays the received menu items, descriptions, and images to the user. The user checks the displayed menu and taps on the desired dish to select it. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts. The server sends the recommendation results to the device, which displays them to the user. The user checks the recommendations and confirms the order.

[2073] Furthermore, the system incorporates an emotion engine, which recognizes emotions from the user's facial expressions and voice. For example, if the user expresses a confused expression, the emotion engine analyzes the information. This emotion information is sent to the server, and menu suggestions and recommendations are adjusted based on the user's emotions. For example, if the user is confused, the system adjusts to provide more detailed explanations and additional information.

[2074] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion engine will recognize the emotion and the server will provide additional detailed explanations of the recommendation and alternatives based on the emotion.

[2075] After the order is confirmed, the user conveys the order to the store clerk. The voice recognition module is activated again, capturing the conversation between the user and the store clerk. The voice data is sent to the server, where it is converted into text data by the voice recognition module. The converted text data is then analyzed to extract additional information, such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user. The user then confirms the alternative options, selects a new menu item, and confirms the selection.

[2076] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[2077] The processing flow will be explained below.

[2078] Step 1:

[2079] User: Activates the camera on their smartphone or tablet and takes a picture of a restaurant menu.

[2080] Step 2:

[2081] Device: Creates a request to send the captured menu image to the server.

[2082] Step 3:

[2083] Terminal: Uploads menu images to the server via the network.

[2084] Step 4:

[2085] Server: Passes the received menu image to the image processing module.

[2086] Step 5:

[2087] Server: Uses OCR technology to extract text from an image, for example, the string "Pasta alla Carbonara."

[2088] Step 6:

[2089] Server: Passes the extracted text data to the natural language processing module.

[2090] Step 7:

[2091] Server: Generates a description of a dish from the extracted strings. For example, "Pasta alla Carbonara" generates the description "Italian pasta with a creamy sauce and bacon."

[2092] Step 8:

[2093] Server: Obtains images related to the dish along with the generated description.

[2094] Step 9:

[2095] Server: Sends the description and image to the device.

[2096] Step 10:

[2097] Terminal: Displays the menu items, descriptions, and images received from the server to the user.

[2098] Step 11:

[2099] User: Review the menu items displayed and tap to select the desired dish.

[2100] Step 12:

[2101] Terminal: Sends information about the menu item selected by the user to the server.

[2102] Step 13:

[2103] Server: Sends a request to the recommendation module to recommend related drinks and desserts based on the menu items selected by the user.

[2104] Step 14:

[2105] Server: Receives the recommendation results from the recommendation module and selects the best drinks and desserts for the user.

[2106] Step 15:

[2107] Server: Sends the recommendation results to the device.

[2108] Step 16:

[2109] Terminal: Display recommended drinks and desserts to the user.

[2110] Step 17:

[2111] User: Review the recommendations and indicate their intent to place an order.

[2112] Step 18:

[2113] Device: Activates the emotion engine to recognize emotions from the user's facial expressions and voice.

[2114] Step 19:

[2115] Terminal: The emotion engine analyzes the user's emotions and sends the results to the server.

[2116] Step 20:

[2117] Server: Receives user emotion data and adjusts recommendations and menu suggestions.

[2118] Step 21:

[2119] Server: Generates additional information and alternative suggestions based on emotion data and sends them to the device.

[2120] Step 22:

[2121] On the device: Display additional information or alternative suggestions to the user based on their emotions.

[2122] Step 23:

[2123] User: Review the proposed information and confirm the order.

[2124] Step 24:

[2125] User: Tells the waiter their order.

[2126] Step 25:

[2127] Terminal: Activates the voice recognition module and captures the conversation between the user and the store clerk as voice data.

[2128] Step 26:

[2129] Device: Sends captured audio data to the server.

[2130] Step 27:

[2131] Server: Passes the voice data to the voice recognition module and converts it into text data.

[2132] Step 28:

[2133] Server: Analyzes the converted text data and extracts additional information such as order confirmation and out-of-stock information.

[2134] Step 29:

[2135] Server: Based on the analysis results, generate alternatives as needed.

[2136] Step 30:

[2137] Server: Sends alternatives to the device.

[2138] Step 31:

[2139] Terminal: Present alternatives to the user.

[2140] Step 32:

[2141] User: Review alternatives, select new menu, confirm.

[2142] Example 2

[2143] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2144] For many users, understanding menus in a foreign language and ordering the appropriate dishes is a difficult task. Furthermore, if recommendations are not appropriate based on the user's emotions and preferences, the user experience cannot be improved. To solve these challenges, a system is needed that not only extracts menu text and provides easy-to-understand information, but also makes recommendations that adapt to the user's emotional state.

[2145] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2146] In this invention, the server includes image recognition means, natural language processing means, recommendation means, and emotion recognition means. This allows users to easily understand menus in foreign languages ​​and order appropriate dishes. Recommendations are also made that are adapted to the user's emotional state, allowing the user to enjoy personalized services.

[2147] "Image recognition means" refers to technology for analyzing an image and extracting specific information from its content.

[2148] "Natural language processing" refers to technology for understanding, generating, and manipulating human language.

[2149] "Recommendation means" refers to technology for recommending appropriate items based on a user's preferences and history.

[2150] "Speech recognition means" refers to technology for analyzing voice data and converting it into text data.

[2151] "Image generation means" refers to technology for generating new images based on input data.

[2152] "Emotion recognition means" refers to technology for analyzing and recognizing emotions from a user's facial expressions and voice.

[2153] "Network communication means" refers to technology for communicating over a network to send and receive data.

[2154] "User interface means" refers to technology that provides an interface for exchanging information between a user and a system.

[2155] The present invention relates to a system that assists users in the process of understanding menus in a foreign language and ordering appropriate dishes, and provides personalized services based on the user's emotions.

[2156] Hardware and Software Configuration

[2157] Required Hardware

[2158] User's device: A smartphone or tablet equipped with a camera is used.

[2159] Server: A high-performance server is used for menu image analysis and recommendation processing.

[2160] Required software

[2161] Image Recognition Method: OCR technology is used to extract text from images, specifically Google Cloud Vision API.

[2162] Natural language processing tools: A generative AI model (e.g., GPT-4) is used to generate dish descriptions from the extracted text data.

[2163] Recommendation method: Use a recommendation algorithm to recommend drinks and desserts related to the dish.

[2164] Emotion Recognition: Emotion recognition software is used to analyze and recognize emotions from the user's facial expressions and voice.

[2165] Speech recognition means: Use speech recognition technology (e.g., Google Speech-to-Text API) to capture the conversation between the user and the store clerk as voice data and convert it into text data.

[2166] Image generation method: Use a web search API (e.g., Bing Image Search API) to obtain related images.

[2167] Network communication means: Uses communication technology to send and receive data between the user terminal and the server.

[2168] User interface means: Provides an interface for exchanging information between the user and the system.

[2169] Example of a system

[2170] 1. Menu Parsing Prompt Example

[2171] "Analyze a restaurant menu image and generate dish names and descriptions."

[2172] 2. Recommendation prompt examples

[2173] "Recommend drinks and desserts to go with the selected dish."

[2174] 3. Example of emotion-responsive prompt

[2175] "If the user looks confused, provide a detailed explanation."

[2176] System Operation

[2177] First, a user takes a photo of a restaurant menu using the camera on their smartphone or tablet. The device then sends the image to a server. The server then analyzes the received image using the Google Cloud Vision API and extracts the text from the image using OCR technology. For example, the string "Pasta alla Carbonara" is extracted.

[2178] The extracted text data is then passed to a generative AI model such as GPT-4 to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta featuring a creamy sauce and bacon." Along with the generated description, related images are retrieved via the Bing Image Search API. The server then sends the retrieved description and image to the device, which then displays them to the user.

[2179] The user checks the displayed menu items, descriptions, and images, and taps to select the desired dish. Information about the menu item selected by the user is sent from the device to the server. Based on this information, the server uses a recommendation module to recommend related drinks and desserts, and sends the recommendation results to the device. The device then displays them to the user.

[2180] The app also has a built-in emotion engine that recognizes emotions from the user's facial expressions and voice. For example, if the user shows a confused expression, the emotion engine analyzes the information and sends it to the server. Based on this emotion information, the server adjusts the service to provide detailed explanations and additional information to the user.

[2181] Finally, the user confirms the order and conveys it to the store clerk. The voice recognition module is activated, capturing the conversation between the user and the store clerk, and the voice data is sent to the server. The server's voice recognition module converts the voice data into text data, which is then analyzed to extract additional information such as out-of-stock information and additional orders. If necessary, alternative options are generated and notified to the user.

[2182] In this way, the present invention allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's feelings.

[2183] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2184] Step 1:

[2185] Users take a photo of a restaurant menu using the camera on their smartphone or tablet.

[2186] Input: Menu Image

[2187] Specific actions: The user opens the camera app, points the lens at the menu, and presses the shutter button to take a photo.

[2188] Output: Menu image captured

[2189] Step 2:

[2190] The terminal transmits the captured menu image to the server.

[2191] Input: Menu image taken

[2192] Specific operation: After taking a photo, the image will be automatically uploaded and sent.

[2193] Output: Menu image sent to the server

[2194] Step 3:

[2195] The server analyzes the received menu image using OCR technology and extracts the text within the image.

[2196] Input: Menu image sent to the server

[2197] Data processing: Call the Google Cloud Vision API to extract character codes from images

[2198] What it does: The API analyzes the text in the menu image and extracts text such as "Pasta alla Carbonara."

[2199] Output: Extracted text data

[2200] Step 4:

[2201] The server passes the extracted text data to a natural language processing module to generate a description of the dish.

[2202] Input: Extracted text data

[2203] Data processing: Calling GPT-4 to generate explanatory text from text data

[2204] What it does: GPT-4 takes the text "Pasta alla Carbonara" and generates a description: "Italian pasta with a creamy sauce and bacon."

[2205] Output: Generated dish description

[2206] Step 5:

[2207] The server retrieves related images via a web search API.

[2208] Input: Description of the generated dish

[2209] Data processing: Call the Bing Image Search API to get related images

[2210] What it does: Uses the Bing Image Search API to search and retrieve relevant images based on the description of "Pasta alla Carbonara."

[2211] Output: Captured image

[2212] Step 6:

[2213] The server transmits the explanatory text and the image to the terminal, which displays them to the user.

[2214] Input: Generated description and retrieved image

[2215] Specific operation: Data is sent from the server to the device, and the device displays a pop-up explanation and image on the screen.

[2216] Output: Description and image displayed to the user

[2217] Step 7:

[2218] The user taps on the desired dish to select it, and the device sends the selection information to the server.

[2219] Input: Displayed description and image

[2220] What happens: The user taps "Pasta alla Carbonara" and the selection is sent to the server.

[2221] Output: Selections sent to the server

[2222] Step 8:

[2223] The server uses a recommendation module to recommend related drinks and desserts and transmits them to the terminal.

[2224] Input: Selections sent to the server

[2225] Data processing: Use recommendation algorithms to select relevant drinks and desserts

[2226] Specific operation: The server selects red wine and tiramisu, and the recommended results are displayed on the device screen.

[2227] Output: Recommendation results displayed to the user

[2228] Step 9:

[2229] The emotion engine recognizes emotions from the user's facial expressions and voice, and transmits the emotion information to the server.

[2230] Input: User's facial expressions and voice

[2231] Data processing: Using emotion recognition software, analyze and recognize the user's emotions.

[2232] Specific operation: When the user makes a confused expression, the camera analyzes the expression and sends the emotional information to the server, which then displays a detailed explanation.

[2233] Output: Additional explanation based on emotion information

[2234] Step 10:

[2235] The user confirms the final order and communicates it to the store attendant. The voice recognition module captures the conversation and sends it to the server.

[2236] Input: User and store clerk conversation

[2237] Data processing: Using voice recognition technology, converting voice data into text data

[2238] What it does: The user verbally places an order, and the speech is sent to the server, which converts it into text. The server then provides out-of-stock information and alternative suggestions.

[2239] Output: Textualized dialogue, out-of-stock information, and alternatives

[2240] Through this series of processes, users can easily understand menus in foreign languages ​​and order the most suitable dishes based on their emotions.

[2241] (Application example 2)

[2242] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2243] In existing systems, users may have difficulty understanding foreign language menus and selecting appropriate dishes, and they are unable to provide personalized recommendations based on the user's emotions. The present invention aims to solve these problems, enabling users to easily understand foreign language menus and providing personalized services based on the user's emotions.

[2244] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means. This allows the user to easily understand foreign language menus and further receive personalized recommendations based on their emotions.

[2245] "Image recognition means" is a device that extracts information from a captured image and analyzes its content.

[2246] "Natural language processing means" is a device that analyzes extracted text data, converts it into natural language format, and understands and processes it.

[2247] A "recommendation means" is a device that makes relevant suggestions or recommendations to the user based on the analyzed information.

[2248] A "voice recognition means" is a device that analyzes voice data and converts it into a linguistic text format.

[2249] The "image generating means" is a device that generates an image based on the analyzed information and generated data.

[2250] A "network communication means" is a communication device for transmitting and receiving data between systems.

[2251] "User interface means" refers to a device that allows a user to operate and interact with the system.

[2252] The "emotion analysis means" is a device that recognizes and analyzes the user's emotions from their facial expressions and tone of voice.

[2253] The present invention provides a system that assists users in understanding menus written in a foreign language and ordering appropriate dishes. The system includes image recognition means, natural language processing means, recommendation means, voice recognition means, image generation means, network communication means, user interface means, and emotion analysis means.

[2254] First, a user takes a photo of a restaurant menu using a smartphone or tablet device. The image taken with the device's camera is sent from the device to a server via network communication means. The server then uses image recognition means to extract text from the received image. For example, the text "Pasta alla Carbonara" is extracted.

[2255] Next, the extracted text data is analyzed using natural language processing to generate a description of the dish. For example, "Pasta alla Carbonara" generates the description "Italian pasta characterized by a creamy sauce and bacon." At the same time, related images are retrieved via a web search API or database. The server then sends the generated description and the retrieved images to the device via the network.

[2256] The terminal uses a user interface means to display the received menu items, descriptions, and images to the user. The user checks the displayed menu and selects the desired dish. Information about the selected menu item is again sent from the terminal to the server. The server uses a recommendation means to recommend additional related drinks and desserts. The recommendation results are sent to the terminal and displayed to the user.

[2257] The system also incorporates an emotion analysis mechanism that recognizes emotions from the user's facial expressions and voice and adjusts menu suggestions and recommendations based on that information. For example, if the user expresses confusion, the emotion analysis mechanism sends that information to the server, which then adjusts the system to provide detailed explanations and additional information.

[2258] As a concrete example, consider a case where a user selects "Pasta alla Carbonara" and the server recommends red wine and tiramisu. If the user shows confusion while making the selection, the emotion analyzer recognizes the emotion and the server provides additional detailed explanations of the recommendation and alternatives.

[2259] After the order is confirmed, the user tells the store clerk what they want to order. At this time, the voice recognition means is activated and the conversation between the user and the store clerk is captured. The voice data is sent back to the server and converted into text data by the voice recognition means. The converted text data is analyzed to extract information such as out-of-stock items and additional orders. If necessary, alternative options are generated and notified to the user. The user then selects a new menu item and confirms the order.

[2260] An example prompt is, "Use the EmotionRecognition library to analyze emotions from the image menu.jpg and generate personalized dish descriptions based on that."

[2261] In this way, the present invention provides a system that allows users to easily understand menus written in foreign languages, order appropriate dishes, and enjoy personalized service based on the user's emotions.

[2262] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2263] Step 1:

[2264] A user takes a photo of a restaurant menu using the camera on their smartphone or tablet.

[2265] Input: Menu Image

[2266] Output: Raw image data in the device

[2267] Specific operation: A user opens the device's camera app and takes a picture of a restaurant menu. After taking the picture, the image data is temporarily stored in the device's memory.

[2268] Step 2:

[2269] The terminal transmits the photographed menu image to the server using the network communication means.

[2270] Input: Raw image data in the device

[2271] Output: Image data sent to the server

[2272] How it works: The device sends image data to the server via Wi-Fi or mobile data using the HTTP or HTTPS protocol.

[2273] Step 3:

[2274] The server uses image recognition means to extract text from the received image.

[2275] Input: Image data sent to the server

[2276] Output: Extracted text data

[2277] How it works: Optical Character Recognition (OCR) software on the server analyzes the image and extracts text information, which is then stored in a database on the server, such as "Pasta alla Carbonara."

[2278] Step 4:

[2279] The server uses natural language processing means to generate a description of the dish from the extracted text data.

[2280] Input: Extracted text data

[2281] Output: The generated dish description

[2282] What it does: A server-based natural language processing library analyzes the text data and generates a description of the corresponding dish, such as "Italian pasta with a creamy sauce and bacon."

[2283] Step 5:

[2284] The server retrieves related images via a web search API or database.

[2285] Input: Extracted text data

[2286] Output: Related images

[2287] Specific operation: The server calls the web search API, searches for relevant images based on the text data, selects the most appropriate image from the search results, and downloads it.

[2288] Step 6:

[2289] The server transmits the generated description and image to the terminal using a network communication means.

[2290] Input: Generated description and associated image

[2291] Output: Description and image sent to the device

[2292] Specific operation: The server uses the HTTP or HTTPS protocol to send the description and image to the terminal.

[2293] Step 7:

[2294] The terminal uses the user interface means to display the received menu items, explanations, and images to the user.

[2295] Input: Received description and image

[2296] Output: Menu items, descriptions, and images displayed in the user interface

[2297] Specific behavior: The device uses the layout template to properly arrange the displayed information on the screen, allowing the user to visually confirm the displayed information.

[2298] Step 8:

[2299] The user checks the displayed menu items and taps to select the desired dish.

[2300] Input: User tap input

[2301] Output: Information about the selected menu item

[2302] Specific operation: The information of the menu item selected by the user is stored in the terminal's memory in preparation for the next communication step.

[2303] Step 9:

[2304] The terminal transmits information about the selected menu item to the server using the network communication means.

[2305] Input: Information of the selected menu item

[2306] Output: Selections sent to the server

[2307] Specific operation: The device sends the selection information to the server via Wi-Fi or mobile data communication.

[2308] Step 10:

[2309] The server uses a recommendation means to recommend related drinks and desserts.

[2310] Input: Information of the selected menu item

[2311] Output: Recommended drinks and desserts

[2312] Specific operation: The server's recommendation engine searches for items related to the selected menu and generates a recommendation list.

[2313] Step 11:

[2314] The server transmits the recommendation results to the terminal using a network communication means.

[2315] Input: Recommended drink and dessert information

[2316] Output: Recommendation results sent to the device

[2317] Specific operation: The server uses HTTP or HTTPS protocol to send the recommendation results to the terminal.

[2318] Step 12:

[2319] The terminal uses the user interface means to display the recommendation results to the user.

[2320] Input: Received recommendation results

[2321] Output: Recommendation results displayed in the user interface

[2322] Specific behavior: The recommended drinks and desserts will be displayed on the device screen for the user to review.

[2323] Step 13:

[2324] The server uses emotion analysis means to recognize emotions from the user's facial expressions and voice.

[2325] Input: User's facial expressions or voice data

[2326] Output: Recognized emotion information

[2327] Specific operation: The server uses a facial expression recognition library and a voice analysis library to analyze the user's emotions and saves the recognition results in a database.

[2328] Step 14:

[2329] The server adjusts menu suggestions and recommendations based on the recognized emotional information.

[2330] Input: Recognized emotion information

[2331] Output: Tailored menu suggestions and recommendations

[2332] Specific operation: Based on the results of the sentiment analysis, the server generates detailed information and alternatives and provides them to the user.

[2333] Step 15:

[2334] The user finally confirms and confirms the menu item.

[2335] Input: User confirmed input

[2336] Output: Confirmed order information

[2337] Specific operation: When the user selects a menu item and taps the confirmation button, the order information is finally captured and saved in the device's memory.

[2338] Step 16:

[2339] When the terminal communicates the order information to the store clerk, the voice recognition means captures the content of the conversation.

[2340] Input: Voice data of the conversation between the user and the store clerk

[2341] Output: Captured audio data

[2342] Specific operation: The device's microphone records the conversation between the user and the store clerk and saves it as audio data.

[2343] Step 17:

[2344] The server receives the voice data and converts it into text data using a voice recognition means.

[2345] Input: Captured audio data

[2346] Output: Converted text data

[2347] Specific operation: A speech recognition library on the server analyzes the voice data and generates corresponding text.

[2348] Step 18:

[2349] The server analyzes the converted text data and extracts out-of-stock information and additional orders.

[2350] Input: Converted text data

[2351] Output: Out-of-stock information and reorder notifications

[2352] Specific operation: The server analyzes the text data, extracts the necessary information, and generates alternative menus and additional order information as necessary.

[2353] Step 19:

[2354] The server informs the user of the alternatives and the user selects a new menu.

[2355] Input: Out-of-stock information and reorder notifications

[2356] Output: New menu selection information

[2357] Specific operation: The user confirms the alternatives and sends the newly selected menu information from the terminal to the server.

[2358] In this way, the present invention helps users understand foreign language menus and provides personalized services based on emotions.

[2359] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2360] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2361] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2362] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2363] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2364] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2365] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2366] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2367] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2368] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2369] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2370] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2371] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2372] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2373] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2374] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2375] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2376] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2377] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2378] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2379] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2380] The following is further disclosed regarding the above embodiment.

[2381] (Claim 1)

[2382] Image recognition means;

[2383] natural language processing means;

[2384] Recommendation means;

[2385] a voice recognition means;

[2386] image generating means;

[2387] a network communication means;

[2388] user interface means;

[2389] A system including:

[2390] (Claim 2)

[2391] 2. The system of claim 1, wherein the image recognition means is for taking an image of a...

Claims

1. Image recognition means; natural language processing means; Recommendation means; a voice recognition means; image generating means; a network communication means; user interface means; A system including:

2. 2. The system of claim 1, wherein said image recognition means is for taking an image of a menu and extracting text from the image.

3. 2. The system of claim 1, wherein the natural language processing means is for generating a description of a dish from the extracted text.

4. The system according to claim 1 , wherein the recommendation means is for suggesting drinks and desserts suitable for a dish.

5. 2. The system according to claim 1, wherein said speech recognition means analyzes a conversation between a user and a store clerk to obtain additional information.

6. 2. The system according to claim 1, wherein the image generating means generates or acquires an image of a dish.

7. 2. The system according to claim 1, wherein said network communication means is for transmitting and receiving data.

8. 2. The system of claim 1, wherein said user interface means is for displaying and selecting menu items.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A