System

A system using image capture, cloud processing, and secure transmission of menu information addresses the challenge of foreign language menus by providing clear dish descriptions, enhancing tourist dining experiences.

JP2026019172APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120581
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Foreign tourists often struggle to understand menus written in local languages, and simple text translations fail to provide clear dish descriptions, making it difficult to select appropriate dishes, especially when menus are not available in English or other languages, detracting from the tourist experience.

Method used

A system that captures menu images with a user's device, processes them through a cloud-based server using optical character recognition, translation, internet search for food images and descriptions, and generates concise summaries, which are then displayed on the user's device, ensuring secure transmission.

Benefits of technology

Enables users to intuitively understand and select dishes despite language barriers, improving convenience and enhancing the dining experience for tourists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019172000001_ABST
    Figure 2026019172000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for capturing an image of a menu with a camera; means for transmitting the image to a server on a cloud; optical character recognition means for extracting character information from the received image; translation means for translating the extracted character information; means for searching the Internet for an image and an explanation of a food based on a translation result; means for transmitting the image and the explanation of the food to a user terminal; and means for displaying the transmitted information on the user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] At tourist destinations, foreign tourists often have difficulty understanding menus written in the local language and deciding which dishes to choose. Furthermore, simple text translations do not provide a clear understanding of the specific contents and characteristics of the dishes, making it difficult to make an appropriate selection. Furthermore, in many cases, menus written in English or other languages ​​are not available, making this inconvenient and detracting from the tourist experience. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for capturing menu images with a camera on a user's device and a means for transmitting the images to a cloud-based server. The server includes an optical character recognition means for extracting text information from the transmitted image and a translation means for translating the extracted text information. The system also includes a means for searching the Internet for food images and descriptions based on the translation results and a means for generating food images and descriptions from the retrieved information. This information is transmitted to the user's device and displayed on the user's device. The system further includes a means for summarizing the food descriptions and a means for using encryption in transmitting the information, thereby providing a system that allows users to select appropriate dishes based on visual and text information.

[0006] A "camera" is a photographic device that a user uses to take an image of a menu.

[0007] An "image" is a digital representation of visual information captured via a photographic device such as a camera.

[0008] A "cloud server" is a remote computing resource accessible via the Internet, and is a device for processing images and analyzing data.

[0009] The "transmitting means" is a communication means for transferring the captured image data to a server on the cloud.

[0010] "Optical Character Recognition (OCR) means" means techniques and devices that extract textual information from an image and convert it into a digital form.

[0011] "Translation means" refers to technology and devices that convert extracted text information into another specified language.

[0012] "Means for searching from the Internet" refers to technology and devices that use databases and APIs on the Internet to obtain images and information related to specific dishes.

[0013] "Generation means" refers to the technology and devices that create images and descriptions of dishes based on the acquired information.

[0014] A "user terminal" is a computing device that is directly operated by a user, such as a smartphone or tablet.

[0015] The "display means" refers to the technology and device for visually presenting received information to the user at the user terminal.

[0016] "Summarizing means" refers to techniques and devices that extract important parts from the acquired food information and summarize them concisely.

[0017] "Means using encryption" refers to techniques and devices that encrypt data so that third parties cannot decipher the information being transmitted. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The following system and program are provided as embodiments of the present invention.

[0040] The system is broadly composed of the user's terminal, a server on the cloud, and a database connected to the Internet.

[0041] User terminal

[0042] The user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user device is to capture images of the menu and send them to a server on the cloud.

[0043] server

[0044] The cloud server processes the received image data using advanced AI technology. The specific processing performed by the server is shown below.

[0045] 1. Image and Character Recognition (OCR):

[0046] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, obtaining the names and descriptions of the dishes on the menu as digital text.

[0047] 2. Text Translation:

[0048] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, the Japanese character for "tempura" is translated into English.

[0049] 3. Internet Search:

[0050] The server uses the translated dish name to search image and information databases on the Internet to retrieve images and descriptions of the corresponding dish, such as images related to "Tempura," along with information about the dish's characteristics and summary.

[0051] 4. Summary sentence generation:

[0052] AI technology is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, flavor characteristics, etc. For example, a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0053] 5. Encryption and Transmission of Information:

[0054] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[0055] Display on user device

[0056] The user device displays the information received from the server through a dedicated application. This allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0057] Specific examples

[0058] Specific examples are shown below.

[0059] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0060] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[0064] Step 2:

[0065] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[0066] Step 3:

[0067] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[0068] Step 4:

[0069] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[0070] Step 5:

[0071] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[0072] Step 6:

[0073] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0074] Step 7:

[0075] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[0076] Step 8:

[0077] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[0078] Step 9:

[0079] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[0080] Step 10:

[0081] The user decides which dish to order based on the displayed information. For example, after seeing an image of tempura and a summary, the user decides whether to order that dish.

[0082] Through this series of steps, the user can easily and intuitively understand the menu contents and select dishes.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] When travelers or users unfamiliar with foreign languages ​​visit restaurants overseas, they often encounter problems due to language barriers and cultural differences when trying to understand the menu and select the appropriate dish. These issues reduce convenience for travelers and detract from their dining experience.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes optical character recognition means for extracting text information from images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, and means for generating summaries of the searched food images and descriptions using a generation AI model, thereby enabling users to quickly and accurately understand menus in foreign languages ​​and easily select dishes.

[0088] A "camera" is a device for taking pictures or videos.

[0089] A "cloud server" is a remotely located computer system that can be accessed via the Internet for storing data and performing computations.

[0090] "Optical character recognition" refers to technology or devices that extract character information from an image as digital text.

[0091] A "translation tool" is a technique or device for converting text expressed in a particular language into another language.

[0092] A "generative AI model" is an artificial intelligence algorithm or set of systems that learns from large amounts of data and performs tasks such as text generation, classification, and translation.

[0093] A "means for generating a summary" is a technology or device that extracts important points from the provided information and summarizes them in a concise sentence.

[0094] An "encryption method" is a technology or device that converts data based on a specific algorithm to prevent unauthorized decryption by third parties.

[0095] A "user terminal" is a computer system operated by a user, and includes mobile phones, smartphones, tablets, and the like.

[0096] "Internet search tools" are techniques or devices that retrieve information from online databases or sources based on specific keywords.

[0097] The following system and program are provided as embodiments of the present invention: The system is broadly composed of a user terminal, a server on the cloud, and a database connected to the Internet.

[0098] User terminal

[0099] A user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user's device is to capture images of the menu and send them to a server on the cloud. The user takes a picture of the menu using the camera function and sends it to the server via the dedicated application.

[0100] server

[0101] The cloud server processes the received image data using advanced AI technology. Below are some examples of specific hardware and software used on the server.

[0102] 1. Image and Character Recognition (OCR):

[0103] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, using libraries such as TensorFlow and OpenCV. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[0104] 2. Text Translation:

[0105] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, the Japanese character for "tempura" (Tempura) is translated into English.

[0106] 3. Internet Search:

[0107] The server then uses the translated dish name to search image and information databases on the Internet, using the Bing Search API and other services. For example, it retrieves images related to "Tempura" and information about the dish, including its general description and characteristics.

[0108] 4. Summary sentence generation:

[0109] A generative AI model (e.g., GPT-4) is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, and flavor characteristics. A prompt sentence is used to input the model, and a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0110] Example prompt sentence:

[0111] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0112] 5. Encryption and Transmission of Information:

[0113] The generated image and summary are encrypted using the SSL / TLS protocol and sent to the user's device. The server then encrypts the data using an encryption library and sends it via an HTTP POST request.

[0114] Display on user device

[0115] The user device decrypts the encrypted data received from the server and displays it through a dedicated application. This process allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0116] Specific examples

[0117] Specific examples are shown below.

[0118] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0119] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0121] Step 1:

[0122] Image capture and transmission

[0123] Users take a photo of a restaurant menu with their smartphone or tablet, and the device acquires the captured image and sends it to a cloud server via a dedicated application using an HTTP POST request.

[0124] Input: A user-taken image of the menu

[0125] Output: Image data sent to a cloud server

[0126] Step 2:

[0127] Image and character recognition (OCR)

[0128] The server processes the received image data and extracts the text information from the image using optical character recognition (OCR) technology. This process uses libraries such as TensorFlow and OpenCV. The server analyzes each pixel of the image and identifies areas that can be recognized as text.

[0129] Input: Image data sent to a cloud server

[0130] Output: Extracted text information (e.g. "tempura")

[0131] Step 3:

[0132] Text Translation

[0133] The server inputs the extracted text information into the Google Translate API and translates it into the specified language. It creates an API request, sends the extracted text information, and receives the translation result. For example, the Japanese text information "tempura" is translated into English "Tempura."

[0134] Input: Extracted text information (e.g., "tempura")

[0135] Output: Translated text (e.g. "Tempura")

[0136] Step 4:

[0137] Internet search

[0138] The server uses the translated text as a key to search for related information on the Internet using the Bing Search API. Specifically, it sends an API request and receives search results, which retrieves information such as images and descriptions related to "Tempura."

[0139] Input: Translated text (e.g. "Tempura")

[0140] Output: Images and descriptions of the dishes retrieved as search results

[0141] Step 5:

[0142] Summary sentence generation

[0143] Based on the acquired information, the server uses a generative AI model (e.g., GPT-4) to generate a concise summary. The server creates a prompt sentence and inputs it into the model to generate a summary. For example, the summary generated is, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0144] Input: Images and descriptions of dishes obtained as search results

[0145] Output: Generated summary

[0146] Example prompt sentence:

[0147] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0148] Step 6:

[0149] Encryption and transmission of information

[0150] The server encrypts the generated image and summary using the SSL / TLS protocol and sends it to the user's terminal. The server also encrypts the data using an encryption library and then sends it via an HTTP POST request.

[0151] Input: Generated summary and food image

[0152] Output: Sent to the user's device as encrypted data

[0153] Step 7:

[0154] Display on user device

[0155] The user device decrypts the encrypted data received from the server and displays it using a dedicated application. Specifically, the device performs decryption processing based on the SSL / TLS protocol and displays an image of tempura and a summary to the user.

[0156] Input: Encrypted data sent by the server

[0157] Output: A picture of the food and a summary of it displayed on the user's device

[0158] By following the above steps, the user can easily understand and select from the menu at a restaurant overseas.

[0159] (Application example 1)

[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0161] Modern restaurants face the challenge of making it difficult for visitors to intuitively understand the dishes and their details on the menu in multiple languages. This is particularly true for tourists who cannot understand foreign languages, as they are unable to immediately understand the dish descriptions, allergen information, prices, and ratings. This can lead to inconvenience when ordering and a decrease in customer satisfaction. Therefore, there is a need for a system that provides multilingual and visually easy-to-understand information about dishes.

[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0163] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a cloud server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for summarizing the generated food images and descriptions as detailed information about dishes in the restaurant, means for transmitting the summarized information to a user terminal, means for displaying the transmitted information on the user terminal, and means for allowing the user to check related image data and rating information on the terminal. This allows visitors to instantly understand detailed information and ratings about dishes in multiple languages, eliminating inconvenience when ordering and improving customer satisfaction.

[0164] "Means for capturing images of menus with a camera" is a function that allows a user to take a photo of a restaurant menu using a mobile device such as a smartphone or tablet and obtain the image data.

[0165] The "means for transmitting the image to a server on the cloud" is a function for uploading image data taken by a mobile terminal to a remote cloud server via the Internet.

[0166] "Optical character recognition means" refers to a technology for extracting character information from received image data, and in particular, uses OCR (optical character recognition) technology.

[0167] The "translation means" is a function for converting extracted text information into a different language as needed.

[0168] The "means for searching on the Internet" is a function for searching an online database based on the translated text information and obtaining images and descriptions of related dishes.

[0169] The "means for generating" is a function for processing the searched food image and description into the format required to provide it to the user.

[0170] The "means for summarizing detailed dish information" is a function that extracts and displays a concise explanation and essential points from the generated image and description of the dish so that the user can easily understand it.

[0171] The "means for transmitting to the user terminal" is a function for transmitting the generated image of the dish and a summarized description to the user's mobile terminal.

[0172] The "means for displaying on the user terminal" is a function for visually displaying the transmitted image of the dish and the summarized description on the user's mobile terminal.

[0173] The "means for checking related image data and evaluation information" is a function that allows the user to refer to image data related to the displayed dish information and evaluation information from other users.

[0174] As an embodiment of the present invention, the following system and program are provided. This system is composed of a user's mobile terminal, a cloud server, and a database connected to the Internet. Specifically, each component functions as follows:

[0175] User terminal

[0176] A user's mobile device, such as a smartphone or tablet, is equipped with a camera, Internet connectivity, and a dedicated application. The main role of the user device is to capture images of menus and send them to a server on the cloud. The user takes a photo of a restaurant menu using the camera and uploads the image to the server via the application.

[0177] server

[0178] The server processes the received image data using advanced AI technology. The server performs the processing using the following specific hardware and software:

[0179] Hardware: High-performance processor (e.g., Intel Xeon), large memory capacity

[0180] Software: Tesseract (OCR engine), Google Translate API, Python

[0181] Image Recognition and Translation

[0182] The server first uses OCR technology to extract text information from the menu image sent from the device. The extracted text information is then translated into the specified language using the Google Translate API. For example, the Japanese character for "tempura" is translated into "Tempura."

[0183] Data Search and Generation

[0184] The translated text is then used to search an online database to retrieve images and descriptions of the corresponding dishes. For example, images related to "Tempura" are retrieved, along with information about the dish's characteristics. This information is then used to summarize the details of the restaurant's dishes. The summary includes information about the main ingredients, cooking methods, and flavor characteristics.

[0185] Sending and Displaying Information

[0186] The summarized images and descriptions are encrypted for security reasons and sent to the user's device. The user's device displays the information received from the server through a dedicated application. This allows users to view images of dishes and brief descriptions simply by pointing their camera at the menu. They can also view related image data and review information.

[0187] Specific examples

[0188] Below is a specific example. When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server uses OCR technology to extract the text information for "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and the image of tempura, a description, and user ratings are displayed on the user's application.

[0189] Example prompts for generative AI models

[0190] "Please translate the Japanese menu items in this image into English."

[0191] "Get detailed information about ramen, including ingredients, cooking method, and flavor characteristics."

[0192] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0193] Step 1:

[0194] A user takes a picture of a restaurant menu using the camera on their mobile device, and the image is sent to a cloud server via the application.

[0195] Input: A menu image taken by the user.

[0196] Output: Menu images sent to a server on the cloud.

[0197] Step 2:

[0198] The server processes the received image data using an OCR engine (Tesseract) to extract the text information from the image, and as a result, the menu items are converted into text data.

[0199] Input: Menu image sent to a server on the cloud.

[0200] Output: Character information (text data) extracted by the OCR engine.

[0201] Step 3:

[0202] The extracted text information is translated into the specified language using the Google Translate API.

[0203] Input: Extracted character information (text data).

[0204] Output: Translated text information (translated text data).

[0205] Step 4:

[0206] The server searches an online database based on the translated text information and retrieves images and descriptions of the corresponding dishes.

[0207] Input: Translated text.

[0208] Output: Images and descriptions of dishes retrieved through internet searches.

[0209] Step 5:

[0210] The server generates a summary from the acquired cooking information, including the main ingredients, cooking method, flavor characteristics, etc.

[0211] Input: Food images and descriptions obtained from an internet search.

[0212] Output: A condensed description of the dish.

[0213] Step 6:

[0214] The generated image and summarized description are encrypted for security reasons and sent to the user's terminal.

[0215] Input: Abridged dish description and image.

[0216] Output: Encrypted information (abridged description and image of the dish).

[0217] Step 7:

[0218] The user terminal decodes the received information using a dedicated application and displays an image of the dish, a summary of its description, related image data, and rating information.

[0219] Input: Encrypted information.

[0220] Output: Decoded and displayed dish image on the user's device, along with a summary description, associated image data, and rating information.

[0221] Specific operation example

[0222] When a user points their camera at the Japanese menu item "tempura" and takes a photo, the image is sent to a cloud server. The server uses OCR technology to extract the text information for "tempura" and translates it into "Tempura" using the Google Translate API. It then searches the internet for images and descriptions related to tempura and summarizes it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." This information is then encrypted and sent to the user's device, where it is displayed on the application.

[0223] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0224] This invention combines a system that captures menu images, extracts and translates text information from the images, and searches the internet for related food images and descriptions, with an emotion engine that recognizes the user's emotions. This system makes it possible to provide information according to the user's emotions, providing a more personalized experience.

[0225] User terminal

[0226] The user's device is a smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The main function of the user's device is to capture menu images and send them to a cloud server. In addition, the device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[0227] server

[0228] The cloud server has the main function of processing the received image data using advanced AI technology. The specific processing details are shown below.

[0229] 1. Image and Character Recognition (OCR):

[0230] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition technology. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[0231] 2. Text Translation:

[0232] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[0233] 3. Internet Search:

[0234] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database, such as images and descriptions related to "Tempura."

[0235] 4. Summary sentence generation:

[0236] From the acquired information, AI technology is used to generate a concise summary, such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0237] 5. Encryption and Transmission of Information:

[0238] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[0239] Emotion Engine

[0240] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns. This makes it possible to provide information tailored to the user's interests and preferences.

[0241] Display on user device

[0242] The user device displays the information received from the server through the application. The display content is adjusted according to the user's emotions by the emotion engine. This allows the user to obtain the most appropriate information and helps them make food selections. For example, if a user points the camera at a menu item called "tempura," a description and visual information will be displayed immediately, along with related recommendations based on the user's emotions.

[0243] Specific examples

[0244] Specific examples are shown below.

[0245] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve information related to the tempura image and generates a summary. The summary, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture," is created, encrypted, and sent to the user's device. Furthermore, the emotion engine recognizes interest from the user's facial expression and displays other recommended dishes. This allows the user to view information related to tempura visually and through text, making the best choice.

[0246] Through this system, users can intuitively understand the contents of the dishes and receive personalized information, making it easier to select dishes.

[0247] The processing flow will be explained below.

[0248] Step 1:

[0249] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[0250] Step 2:

[0251] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[0252] Step 3:

[0253] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[0254] Step 4:

[0255] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[0256] Step 5:

[0257] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[0258] Step 6:

[0259] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0260] Step 7:

[0261] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[0262] Step 8:

[0263] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[0264] Step 9:

[0265] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[0266] Step 10:

[0267] The emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, if the user's facial expression indicates joy or interest, that information is recorded.

[0268] Step 11:

[0269] The emotion engine uses the user's emotion information to provide additional information tailored to the user's interests and preferences. For example, if a user is interested in tempura, it will suggest other fried dishes and Japanese food options.

[0270] Step 12:

[0271] The device displays information provided by the emotion engine, and the user can review it and select a dish. For example, in addition to information about tempura, images and descriptions of katsudon and karaage are also displayed.

[0272] This series of steps allows users to intuitively understand the menu contents and select the best dish based on personalized information that responds to their emotions.

[0273] Example 2

[0274] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0275] Conventional menu translation systems have the problem of not providing information that takes into account the user's emotions and interests, and can only display uniform information. As a result, it is difficult for users to quickly obtain the information they really want, and satisfaction cannot be increased. In addition, there are limitations to the accuracy of translation and the reliability of the information, so there is a demand for more accurate information provision.

[0276] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0277] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a remote server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching a network for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for transmitting the food images and descriptions to a user terminal, means for displaying the transmitted information on the user terminal, means for recognizing emotions by analyzing a user's facial expressions and tone of voice, and means for adjusting information based on the emotion recognition, thereby enabling the provision of personalized information that takes into account the user's emotions and interests.

[0278] The "means for capturing an image of a menu using a camera" is a function for taking a digital image of a physical menu using a camera mounted on a user terminal.

[0279] The "means for transmitting the image to a remote server" is a function for transmitting captured image data from a user terminal to a server on the cloud using an internet connection.

[0280] "Optical character recognition means for extracting character information from received images" refers to technology for recognizing character information from transmitted image data and converting it into digital text.

[0281] "Translation means for translating extracted text information" refers to technology for converting text information obtained by OCR into another language.

[0282] "Means for searching the network for food images and descriptions based on the translation results" is a technology that uses the translated text information as input to obtain related images and descriptions using the Internet.

[0283] "Means for generating images and descriptions of searched dishes" is a function that creates images of dishes and related descriptions based on information obtained from the network.

[0284] The "means for transmitting the image and description of the dish to the user terminal" is a technique for transferring the generated image and description of the dish from a remote server to the user terminal.

[0285] The "means for displaying the transmitted information on the user terminal" is a function for visually displaying the image and description of the dish received on the user terminal.

[0286] "Means for recognizing emotions by analyzing a user's facial expressions and tone of voice" refers to technology that uses a camera and microphone to analyze the user's facial expressions and voice to determine the user's emotional state.

[0287] The "means for adjusting information based on emotion recognition" is a technology that dynamically changes the content of the information to be displayed according to the user's recognized emotion.

[0288] The system of the present invention allows users to capture images of menus, extract and translate textual information from the images, search the internet for related food images and descriptions, and even recognize the user's emotions to provide personalized information, providing users with more intuitive and individually tailored information.

[0289] Hardware and software used

[0290] User's device: A smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The user captures an image of the menu with the device's camera and sends it to a cloud server. The device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[0291] Server: The server on the cloud processes the received image data using advanced AI technology and has the following functions:

[0292] 1. OCR (Optical Character Recognition): Uses OCR technology, such as Google Cloud Vision API, to extract text information from received images.

[0293] 2. Translation: To translate the extracted text information into multiple languages, use, for example, the Google Translate API.

[0294] 3. Internet search: Using the translated text, you can search for food images and descriptions on the Internet, for example, using the Bing Search API.

[0295] 4. Summary generation: A generative AI model such as GPT-3 is used to generate a summary from the acquired information.

[0296] 5. Information encryption: The generated information is encrypted and sent to the user terminal using the Advanced Encryption Standard (AES).

[0297] Specific examples

[0298] Consider a scenario where a user points their camera at a Japanese menu item that says "tempura" at a Japanese restaurant. The user's device sends the image to a server in the cloud. The server analyzes the image and extracts the text "tempura" using OCR technology. The server then translates "tempura" to "Tempura" using the Google Translate API. The server then retrieves images and descriptions related to "Tempura" from the Internet using the Bing Search API. Based on the retrieved information, the server uses GPT-3 to generate a summary: "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture." This information is then encrypted using AES and sent to the user's device. The user's device decrypts the received information and displays it to the user through an application. The emotion engine recognizes the user's facial expression as an indication of interest and simultaneously displays other related recommended dishes.

[0299] Prompt Sentence Examples

[0300] "Let the user capture an image of a restaurant menu, extract and translate text from the image, and provide relevant dish images and descriptions. Also, recognize the user's emotions and adjust to provide a personalized information experience."

[0301] This system allows users to intuitively understand the contents of the dish and receive information that is tailored to their individual needs.

[0302] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0303] Step 1:

[0304] The user captures an image of the menu

[0305] The user takes a picture of a restaurant menu using the camera on their smartphone or tablet. This image becomes the initial input for the system. Specifically, the user launches the device's camera application and captures an image of the menu. The image data is then saved via a dedicated application, and the system proceeds to the next step.

[0306] Step 2:

[0307] The device sends the image to the server

[0308] After image data is captured by the camera application, a dedicated application sends the image to a server on the cloud. The input is the captured image data, and the output is the image data sent to the server. Specifically, the application on the device sends the image data to a specified endpoint on the server via an Internet connection.

[0309] Step 3:

[0310] The server performs image recognition and OCR processing

[0311] The server receives image data sent from the device. The input is the sent image data, and the output is the extracted text information. The image recognition module in the server analyzes the image and converts the text information in the image into digital text using OCR (optical character recognition) technology. Specifically, the server uses the "Google Cloud Vision API" to extract the text information in the image.

[0312] Step 4:

[0313] The server translates the extracted text.

[0314] The server takes the character information extracted by the OCR process as input and sends this information to a translation engine. The output is the translated text. The server uses the "Google Translate API" to translate the text into the specified language. Specifically, the server translates the extracted Japanese character information "Tempura" into "Tempura."

[0315] Step 5:

[0316] The server performs an internet search based on the translation results

[0317] Using the translated text as input, the server performs an internet search to gather images and descriptions of related dishes. The output is the retrieved images and descriptions. The server uses the Bing Search API to search for information about "Tempura." Specifically, the server gathers images and descriptions related to "Tempura" from multiple websites.

[0318] Step 6:

[0319] The server generates a summary

[0320] Using the acquired information as input, the server uses a generative AI model to generate a concise summary. The output is summarized text. The server uses GPT-3 to summarize the information, generating a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." Specifically, the server performs a process to summarize the collected descriptions using the generative AI model.

[0321] Step 7:

[0322] The server encrypts the information and sends it to the user's device.

[0323] The server takes the generated image and summary as input and encrypts this information. The output is encrypted data. The server encrypts the data using AES (Advanced Encryption Standard) and sends it to the user's device. Specifically, the server uses AES encryption to keep the data secure and sends it to the device via the Internet.

[0324] Step 8:

[0325] Display information received by the device

[0326] The device receives encrypted information from the server as input, decrypts it, and visually displays it to the user. The output is an image of the displayed dish and a summary. A dedicated application on the device decrypts the encrypted data and displays an image of the dish along with a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture." Specifically, the device uses the received key data to decrypt the data and displays it through the application interface.

[0327] Step 9:

[0328] The device recognizes the user's emotions and adjusts the information accordingly.

[0329] Once again, the emotion engine uses the user's facial expression and tone of voice as input to recognize the emotion and adjust the information. The output is additional information based on the user's emotion. The emotion engine uses the "Emotion API" to analyze the user's emotion, and if the user shows interest, it displays additional information about related dishes or recommendations. Specifically, the device uses the camera and microphone to analyze the user's facial expression and voice in real time, dynamically changing the information displayed.

[0330] (Application example 2)

[0331] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0332] In today's increasingly internationalized world, understanding the menu contents of restaurants abroad is an important factor for travelers. However, language barriers and lack of knowledge about cuisine can make it difficult for travelers to accurately understand the menu. Furthermore, conventional menu translation systems lack the ability to provide personalized information that takes into account the user's interests and preferences, making it difficult for them to choose a meal. To resolve this situation, it is necessary to recognize the user's emotions and provide information based on them.

[0333] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0334] In this invention, the server includes optical character recognition means for extracting text information from received images, translation means for translating the extracted text information, and means for searching the Internet for food images and descriptions based on the translation results, thereby making it easier to understand menu contents beyond language barriers and enabling personalized food suggestions.

[0335] A "camera" is a device that captures light and records images or videos.

[0336] A "cloud server" is a remote computing resource accessed via the Internet, a computer system that processes and stores data.

[0337] "Optical character recognition means" is a technology that recognizes character information from image data and extracts it as text data.

[0338] A "translation tool" is a technique for converting text in one language into text in another language.

[0339] "Means of searching the Internet" refers to the technology of obtaining information that is publicly available online through search engines or APIs.

[0340] The "means for generating food images and descriptions" is a technology that constructs visual and text descriptions based on information obtained from search results.

[0341] "Means for transmitting to the user terminal" refers to the technology for transferring data from a server on the cloud to the user's device.

[0342] "Emotion recognition means" is a technology that analyzes and identifies a user's emotional state from facial expressions, tone of voice, etc.

[0343] "Means for making personalized dish suggestions" refers to technology that makes individually optimized dish suggestions based on the user's emotional data and preferences.

[0344] "Means for displaying on the user terminal" refers to a technique for visually displaying the provided information on the user's device.

[0345] This invention combines a system that captures menu images with a camera, extracts and translates text information from the images, and searches the internet for related food images and descriptions to display them, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system are described below.

[0346] System Configuration

[0347] User device:

[0348] The user's device is typically a smartphone or tablet. The device is equipped with a camera and is connected to the Internet. A dedicated application is installed on the device, which has the function of capturing images of menus and sending them to a cloud server. Additionally, the device is equipped with an emotion engine, which can analyze the user's facial expressions and tone of voice.

[0349] Cloud Server:

[0350] The cloud server has high-performance computing resources and performs the following processes:

[0351] 1. Image and Character Recognition (OCR):

[0352] The server receives the menu image sent from the user device and uses optical character recognition software (e.g., pytesseract) to extract the text information within the image. This process results in the names and descriptions of the dishes on the menu being captured as digital text.

[0353] 2. Text Translation:

[0354] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[0355] 3. Internet Search:

[0356] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database (e.g., Google search), such as images and descriptions related to "Tempura."

[0357] 4. Information generation:

[0358] Using the acquired information, AI technology generates a concise description, such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture."

[0359] 5. Encryption and Transmission of Information:

[0360] The generated image and description are encrypted as a security measure and sent to the user's terminal.

[0361] Emotion Engine:

[0362] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns, enabling it to provide information tailored to the user's interests and preferences.

[0363] System Operation

[0364] When a user captures a menu image with their camera, the image is sent to a cloud server. The server extracts text information from the image and translates it into the specified language. Based on the translation results, the server searches the internet for related food images and descriptions, and generates a summary of the information.

[0365] The generated information is tailored based on the user's emotional data. For example, if the emotion engine recognizes that the user is interested, it will provide related information and recommended dishes. This allows users to get personalized information that will help them make better food choices.

[0366] Examples:

[0367] When a user points their camera at the Japanese menu item "Tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text information "Tempura" from the received image and translates it to "Tempura." The server then uses the Internet to obtain information related to the image of tempura and generates a description. The description it creates is "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture," and the information is encrypted and sent to the user's device. The emotion engine recognizes interest from the user's facial expression and displays other recommended dishes.

[0368] Example prompt sentence:

[0369] "Generate a Python program that captures an image of a dish called tempura on a menu at a Japanese restaurant, extracts and translates the text from the image, analyzes the user's emotions from their facial expressions, and provides the user with the most appropriate food information based on their emotions."

[0370] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0371] Step 1:

[0372] The device captures an image of the menu with its camera.

[0373] Input: Restaurant menu image

[0374] Output: Captured image data

[0375] Specific action: The user takes a photo of the menu using the camera on their smartphone or tablet.

[0376] Step 2:

[0377] The device sends the captured image to a server on the cloud.

[0378] Input: Captured image data

[0379] Output: Image data sent to the server

[0380] Specific operation: The device uses an internet connection to upload image data to a server in the cloud.

[0381] Step 3:

[0382] The server extracts text information from the received image using optical character recognition (OCR).

[0383] Input: Image data sent to the server

[0384] Output: Extracted text information (text data)

[0385] Specific operation: The server uses OCR software (e.g., pytesseract) to analyze the characters in the image and generate text data.

[0386] Step 4:

[0387] The server translates the extracted text information into the specified language using a translation engine.

[0388] Input: Extracted text information (text data)

[0389] Output: Translated text information (text data)

[0390] What happens: The server uses a translation engine, such as the Google Translate API, to translate the text into the specified language.

[0391] Step 5:

[0392] The server searches the internet for images and descriptions of related dishes based on the translation results.

[0393] Input: Translated text information (text data)

[0394] Output: Images and descriptions of the dishes obtained

[0395] Specific operation: The server uses a search engine API to retrieve information about dishes related to the translated text from the Internet.

[0396] Step 6:

[0397] The server generates an image and description of the dish from the information it obtains.

[0398] Input: Image and description of the food obtained

[0399] Output: Generated food images and abbreviated descriptions

[0400] Specific operation: The server uses AI technology to organize the acquired information and generate visually easy-to-understand images and concise explanations.

[0401] Step 7:

[0402] The server encrypts the generated information and transmits it to the user terminal.

[0403] Input: Generated food image and description

[0404] Output: Encrypted dish image and description

[0405] Specific operation: The server encrypts the information as a security measure and sends it to the user's terminal via the Internet.

[0406] Step 8:

[0407] The device analyzes the user's facial expressions and tone of voice to recognize emotions.

[0408] Input: User's facial expression and voice data

[0409] Output: Recognized emotion data

[0410] Specific operation: The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them through an emotion engine.

[0411] Step 9:

[0412] The server makes personalized dish suggestions based on emotion recognition results.

[0413] Input: Recognized emotion data, user behavior history

[0414] Output: Personalized food suggestions

[0415] Specific operation: The server generates individually optimized dish suggestions based on the emotion data and the user's previous history and sends them to the user's device.

[0416] Step 10:

[0417] The terminal displays the transmitted information to the user.

[0418] Input: Encrypted food images, abbreviated descriptions, and personalized food suggestions

[0419] Output: Food images, descriptions, and suggestions displayed on the user's device

[0420] Specific operation: The device decodes the received information and displays it visually to the user.

[0421] The above is a specific processing flow for carrying out the present invention.

[0422] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0423] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0424] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0425] [Second embodiment]

[0426] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0427] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0428] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0429] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0430] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0431] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0432] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0433] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0434] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0435] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0436] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0437] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0438] The following system and program are provided as embodiments of the present invention.

[0439] The system is broadly composed of the user's terminal, a server on the cloud, and a database connected to the Internet.

[0440] User terminal

[0441] The user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user device is to capture images of the menu and send them to a server on the cloud.

[0442] server

[0443] The cloud server processes the received image data using advanced AI technology. The specific processing performed by the server is shown below.

[0444] 1. Image and Character Recognition (OCR):

[0445] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, obtaining the names and descriptions of the dishes on the menu as digital text.

[0446] 2. Text Translation:

[0447] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, the Japanese character for "tempura" is translated into English.

[0448] 3. Internet Search:

[0449] The server uses the translated dish name to search image and information databases on the Internet to retrieve images and descriptions of the corresponding dish, such as images related to "Tempura," along with information about the dish's characteristics and summary.

[0450] 4. Summary sentence generation:

[0451] AI technology is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, flavor characteristics, etc. For example, a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0452] 5. Encryption and Transmission of Information:

[0453] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[0454] Display on user device

[0455] The user device displays the information received from the server through a dedicated application. This allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0456] Specific examples

[0457] Specific examples are shown below.

[0458] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0459] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0460] The processing flow will be explained below.

[0461] Step 1:

[0462] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[0463] Step 2:

[0464] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[0465] Step 3:

[0466] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[0467] Step 4:

[0468] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[0469] Step 5:

[0470] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[0471] Step 6:

[0472] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0473] Step 7:

[0474] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[0475] Step 8:

[0476] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[0477] Step 9:

[0478] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[0479] Step 10:

[0480] The user decides which dish to order based on the displayed information. For example, after seeing an image of tempura and a summary, the user decides whether to order that dish.

[0481] Through this series of steps, the user can easily and intuitively understand the menu contents and select dishes.

[0482] Example 1

[0483] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0484] When travelers or users unfamiliar with foreign languages ​​visit restaurants overseas, they often encounter problems due to language barriers and cultural differences when trying to understand the menu and select the appropriate dish. These issues reduce convenience for travelers and detract from their dining experience.

[0485] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0486] In this invention, the server includes optical character recognition means for extracting text information from images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, and means for generating summaries of the searched food images and descriptions using a generation AI model, thereby enabling users to quickly and accurately understand menus in foreign languages ​​and easily select dishes.

[0487] A "camera" is a device for taking pictures or videos.

[0488] A "cloud server" is a remotely located computer system that can be accessed via the Internet for storing data and performing computations.

[0489] "Optical character recognition" refers to technology or devices that extract character information from an image as digital text.

[0490] A "translation tool" is a technique or device for converting text expressed in a particular language into another language.

[0491] A "generative AI model" is an artificial intelligence algorithm or set of systems that learns from large amounts of data and performs tasks such as text generation, classification, and translation.

[0492] A "means for generating a summary" is a technology or device that extracts important points from the provided information and summarizes them in a concise sentence.

[0493] An "encryption method" is a technology or device that converts data based on a specific algorithm to prevent unauthorized decryption by third parties.

[0494] A "user terminal" is a computer system operated by a user, and includes mobile phones, smartphones, tablets, and the like.

[0495] "Internet search tools" are techniques or devices that retrieve information from online databases or sources based on specific keywords.

[0496] The following system and program are provided as embodiments of the present invention: The system is broadly composed of a user terminal, a server on the cloud, and a database connected to the Internet.

[0497] User terminal

[0498] A user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user's device is to capture images of the menu and send them to a server on the cloud. The user takes a picture of the menu using the camera function and sends it to the server via the dedicated application.

[0499] server

[0500] The cloud server processes the received image data using advanced AI technology. Below are some examples of specific hardware and software used on the server.

[0501] 1. Image and Character Recognition (OCR):

[0502] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, using libraries such as TensorFlow and OpenCV. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[0503] 2. Text Translation:

[0504] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, the Japanese character for "tempura" (Tempura) is translated into English.

[0505] 3. Internet Search:

[0506] The server then uses the translated dish name to search image and information databases on the Internet, using the Bing Search API and other services. For example, it retrieves images related to "Tempura" and information about the dish, including its general description and characteristics.

[0507] 4. Summary sentence generation:

[0508] A generative AI model (e.g., GPT-4) is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, and flavor characteristics. A prompt sentence is used to input the model, and a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0509] Example prompt sentence:

[0510] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0511] 5. Encryption and Transmission of Information:

[0512] The generated image and summary are encrypted using the SSL / TLS protocol and sent to the user's device. The server then encrypts the data using an encryption library and sends it via an HTTP POST request.

[0513] Display on user device

[0514] The user device decrypts the encrypted data received from the server and displays it through a dedicated application. This process allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0515] Specific examples

[0516] Specific examples are shown below.

[0517] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0518] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0519] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0520] Step 1:

[0521] Image capture and transmission

[0522] Users take a photo of a restaurant menu with their smartphone or tablet, and the device acquires the captured image and sends it to a cloud server via a dedicated application using an HTTP POST request.

[0523] Input: A user-taken image of the menu

[0524] Output: Image data sent to a cloud server

[0525] Step 2:

[0526] Image and character recognition (OCR)

[0527] The server processes the received image data and extracts the text information from the image using optical character recognition (OCR) technology. This process uses libraries such as TensorFlow and OpenCV. The server analyzes each pixel of the image and identifies areas that can be recognized as text.

[0528] Input: Image data sent to a cloud server

[0529] Output: Extracted text information (e.g. "tempura")

[0530] Step 3:

[0531] Text Translation

[0532] The server inputs the extracted text information into the Google Translate API and translates it into the specified language. It creates an API request, sends the extracted text information, and receives the translation result. For example, the Japanese text information "tempura" is translated into English "Tempura."

[0533] Input: Extracted text information (e.g., "tempura")

[0534] Output: Translated text (e.g. "Tempura")

[0535] Step 4:

[0536] Internet search

[0537] The server uses the translated text as a key to search for related information on the Internet using the Bing Search API. Specifically, it sends an API request and receives search results, which retrieves information such as images and descriptions related to "Tempura."

[0538] Input: Translated text (e.g. "Tempura")

[0539] Output: Images and descriptions of the dishes retrieved as search results

[0540] Step 5:

[0541] Summary sentence generation

[0542] Based on the acquired information, the server uses a generative AI model (e.g., GPT-4) to generate a concise summary. The server creates a prompt sentence and inputs it into the model to generate a summary. For example, the summary generated is, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0543] Input: Images and descriptions of dishes obtained as search results

[0544] Output: Generated summary

[0545] Example prompt sentence:

[0546] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0547] Step 6:

[0548] Encryption and transmission of information

[0549] The server encrypts the generated image and summary using the SSL / TLS protocol and sends it to the user's terminal. The server also encrypts the data using an encryption library and then sends it via an HTTP POST request.

[0550] Input: Generated summary and food image

[0551] Output: Sent to the user's device as encrypted data

[0552] Step 7:

[0553] Display on user device

[0554] The user device decrypts the encrypted data received from the server and displays it using a dedicated application. Specifically, the device performs decryption processing based on the SSL / TLS protocol and displays an image of tempura and a summary to the user.

[0555] Input: Encrypted data sent by the server

[0556] Output: A picture of the food and a summary of it displayed on the user's device

[0557] By following the above steps, the user can easily understand and select from the menu at a restaurant overseas.

[0558] (Application example 1)

[0559] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0560] Modern restaurants face the challenge of making it difficult for visitors to intuitively understand the dishes and their details on the menu in multiple languages. This is particularly true for tourists who cannot understand foreign languages, as they are unable to immediately understand the dish descriptions, allergen information, prices, and ratings. This can lead to inconvenience when ordering and a decrease in customer satisfaction. Therefore, there is a need for a system that provides multilingual and visually easy-to-understand information about dishes.

[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0562] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a cloud server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for summarizing the generated food images and descriptions as detailed information about dishes in the restaurant, means for transmitting the summarized information to a user terminal, means for displaying the transmitted information on the user terminal, and means for allowing the user to check related image data and rating information on the terminal. This allows visitors to instantly understand detailed information and ratings about dishes in multiple languages, eliminating inconvenience when ordering and improving customer satisfaction.

[0563] "Means for capturing images of menus with a camera" is a function that allows a user to take a photo of a restaurant menu using a mobile device such as a smartphone or tablet and obtain the image data.

[0564] The "means for transmitting the image to a server on the cloud" is a function for uploading image data taken by a mobile terminal to a remote cloud server via the Internet.

[0565] "Optical character recognition means" refers to a technology for extracting character information from received image data, and in particular, uses OCR (optical character recognition) technology.

[0566] The "translation means" is a function for converting extracted text information into a different language as needed.

[0567] The "means for searching on the Internet" is a function for searching an online database based on the translated text information and obtaining images and descriptions of related dishes.

[0568] The "means for generating" is a function for processing the searched food image and description into the format required to provide it to the user.

[0569] The "means for summarizing detailed dish information" is a function that extracts and displays a concise explanation and essential points from the generated image and description of the dish so that the user can easily understand it.

[0570] The "means for transmitting to the user terminal" is a function for transmitting the generated image of the dish and a summarized description to the user's mobile terminal.

[0571] The "means for displaying on the user terminal" is a function for visually displaying the transmitted image of the dish and the summarized description on the user's mobile terminal.

[0572] The "means for checking related image data and evaluation information" is a function that allows the user to refer to image data related to the displayed dish information and evaluation information from other users.

[0573] As an embodiment of the present invention, the following system and program are provided. This system is composed of a user's mobile terminal, a cloud server, and a database connected to the Internet. Specifically, each component functions as follows:

[0574] User terminal

[0575] A user's mobile device, such as a smartphone or tablet, is equipped with a camera, Internet connectivity, and a dedicated application. The main role of the user device is to capture images of menus and send them to a server on the cloud. The user takes a photo of a restaurant menu using the camera and uploads the image to the server via the application.

[0576] server

[0577] The server processes the received image data using advanced AI technology. The server performs the processing using the following specific hardware and software:

[0578] Hardware: High-performance processor (e.g., Intel Xeon), large memory capacity

[0579] Software: Tesseract (OCR engine), Google Translate API, Python

[0580] Image Recognition and Translation

[0581] The server first uses OCR technology to extract text information from the menu image sent from the device. The extracted text information is then translated into the specified language using the Google Translate API. For example, the Japanese character for "tempura" is translated into "Tempura."

[0582] Data Search and Generation

[0583] The translated text is then used to search an online database to retrieve images and descriptions of the corresponding dishes. For example, images related to "Tempura" are retrieved, along with information about the dish's characteristics. This information is then used to summarize the details of the restaurant's dishes. The summary includes information about the main ingredients, cooking methods, and flavor characteristics.

[0584] Sending and Displaying Information

[0585] The summarized images and descriptions are encrypted for security reasons and sent to the user's device. The user's device displays the information received from the server through a dedicated application. This allows users to view images of dishes and brief descriptions simply by pointing their camera at the menu. They can also view related image data and review information.

[0586] Specific examples

[0587] Below is a specific example. When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server uses OCR technology to extract the text information for "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and the image of tempura, a description, and user ratings are displayed on the user's application.

[0588] Example prompts for generative AI models

[0589] "Please translate the Japanese menu items in this image into English."

[0590] "Get detailed information about ramen, including ingredients, cooking method, and flavor characteristics."

[0591] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0592] Step 1:

[0593] A user takes a picture of a restaurant menu using the camera on their mobile device, and the image is sent to a cloud server via the application.

[0594] Input: A menu image taken by the user.

[0595] Output: Menu images sent to a server on the cloud.

[0596] Step 2:

[0597] The server processes the received image data using an OCR engine (Tesseract) to extract the text information from the image, and as a result, the menu items are converted into text data.

[0598] Input: Menu image sent to a server on the cloud.

[0599] Output: Character information (text data) extracted by the OCR engine.

[0600] Step 3:

[0601] The extracted text information is translated into the specified language using the Google Translate API.

[0602] Input: Extracted character information (text data).

[0603] Output: Translated text information (translated text data).

[0604] Step 4:

[0605] The server searches an online database based on the translated text information and retrieves images and descriptions of the corresponding dishes.

[0606] Input: Translated text.

[0607] Output: Images and descriptions of dishes retrieved through internet searches.

[0608] Step 5:

[0609] The server generates a summary from the acquired cooking information, including the main ingredients, cooking method, flavor characteristics, etc.

[0610] Input: Food images and descriptions obtained from an internet search.

[0611] Output: A condensed description of the dish.

[0612] Step 6:

[0613] The generated image and summarized description are encrypted for security reasons and sent to the user's terminal.

[0614] Input: Abridged dish description and image.

[0615] Output: Encrypted information (abridged description and image of the dish).

[0616] Step 7:

[0617] The user terminal decodes the received information using a dedicated application and displays an image of the dish, a summary of its description, related image data, and rating information.

[0618] Input: Encrypted information.

[0619] Output: Decoded and displayed dish image on the user's device, along with a summary description, associated image data, and rating information.

[0620] Specific operation example

[0621] When a user points their camera at the Japanese menu item "tempura" and takes a photo, the image is sent to a cloud server. The server uses OCR technology to extract the text information for "tempura" and translates it into "Tempura" using the Google Translate API. It then searches the internet for images and descriptions related to tempura and summarizes it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." This information is then encrypted and sent to the user's device, where it is displayed on the application.

[0622] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0623] This invention combines a system that captures menu images, extracts and translates text information from the images, and searches the internet for related food images and descriptions, with an emotion engine that recognizes the user's emotions. This system makes it possible to provide information according to the user's emotions, providing a more personalized experience.

[0624] User terminal

[0625] The user's device is a smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The main function of the user's device is to capture menu images and send them to a cloud server. In addition, the device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[0626] server

[0627] The cloud server has the main function of processing the received image data using advanced AI technology. The specific processing details are shown below.

[0628] 1. Image and Character Recognition (OCR):

[0629] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition technology. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[0630] 2. Text Translation:

[0631] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[0632] 3. Internet Search:

[0633] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database, such as images and descriptions related to "Tempura."

[0634] 4. Summary sentence generation:

[0635] From the acquired information, AI technology is used to generate a concise summary, such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0636] 5. Encryption and Transmission of Information:

[0637] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[0638] Emotion Engine

[0639] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns. This makes it possible to provide information tailored to the user's interests and preferences.

[0640] Display on user device

[0641] The user device displays the information received from the server through the application. The display content is adjusted according to the user's emotions by the emotion engine. This allows the user to obtain the most appropriate information and helps them make food selections. For example, if a user points the camera at a menu item called "tempura," a description and visual information will be displayed immediately, along with related recommendations based on the user's emotions.

[0642] Specific examples

[0643] Specific examples are shown below.

[0644] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve information related to the tempura image and generates a summary. The summary, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture," is created, encrypted, and sent to the user's device. Furthermore, the emotion engine recognizes interest from the user's facial expression and displays other recommended dishes. This allows the user to view information related to tempura visually and through text, making the best choice.

[0645] Through this system, users can intuitively understand the contents of the dishes and receive personalized information, making it easier to select dishes.

[0646] The processing flow will be explained below.

[0647] Step 1:

[0648] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[0649] Step 2:

[0650] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[0651] Step 3:

[0652] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[0653] Step 4:

[0654] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[0655] Step 5:

[0656] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[0657] Step 6:

[0658] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0659] Step 7:

[0660] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[0661] Step 8:

[0662] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[0663] Step 9:

[0664] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[0665] Step 10:

[0666] The emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, if the user's facial expression indicates joy or interest, that information is recorded.

[0667] Step 11:

[0668] The emotion engine uses the user's emotion information to provide additional information tailored to the user's interests and preferences. For example, if a user is interested in tempura, it will suggest other fried dishes and Japanese food options.

[0669] Step 12:

[0670] The device displays information provided by the emotion engine, and the user can review it and select a dish. For example, in addition to information about tempura, images and descriptions of katsudon and karaage are also displayed.

[0671] This series of steps allows users to intuitively understand the menu contents and select the best dish based on personalized information that responds to their emotions.

[0672] Example 2

[0673] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0674] Conventional menu translation systems have the problem of not providing information that takes into account the user's emotions and interests, and can only display uniform information. As a result, it is difficult for users to quickly obtain the information they really want, and satisfaction cannot be increased. In addition, there are limitations to the accuracy of translation and the reliability of the information, so there is a demand for more accurate information provision.

[0675] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0676] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a remote server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching a network for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for transmitting the food images and descriptions to a user terminal, means for displaying the transmitted information on the user terminal, means for recognizing emotions by analyzing a user's facial expressions and tone of voice, and means for adjusting information based on the emotion recognition, thereby enabling the provision of personalized information that takes into account the user's emotions and interests.

[0677] The "means for capturing an image of a menu using a camera" is a function for taking a digital image of a physical menu using a camera mounted on a user terminal.

[0678] The "means for transmitting the image to a remote server" is a function for transmitting captured image data from a user terminal to a server on the cloud using an internet connection.

[0679] "Optical character recognition means for extracting character information from received images" refers to technology for recognizing character information from transmitted image data and converting it into digital text.

[0680] "Translation means for translating extracted text information" refers to technology for converting text information obtained by OCR into another language.

[0681] "Means for searching the network for food images and descriptions based on the translation results" is a technology that uses the translated text information as input to obtain related images and descriptions using the Internet.

[0682] "Means for generating images and descriptions of searched dishes" is a function that creates images of dishes and related descriptions based on information obtained from the network.

[0683] The "means for transmitting the image and description of the dish to the user terminal" is a technique for transferring the generated image and description of the dish from a remote server to the user terminal.

[0684] The "means for displaying the transmitted information on the user terminal" is a function for visually displaying the image and description of the dish received on the user terminal.

[0685] "Means for recognizing emotions by analyzing a user's facial expressions and tone of voice" refers to technology that uses a camera and microphone to analyze the user's facial expressions and voice to determine the user's emotional state.

[0686] The "means for adjusting information based on emotion recognition" is a technology that dynamically changes the content of the information to be displayed according to the user's recognized emotion.

[0687] The system of the present invention allows users to capture images of menus, extract and translate textual information from the images, search the internet for related food images and descriptions, and even recognize the user's emotions to provide personalized information, providing users with more intuitive and individually tailored information.

[0688] Hardware and software used

[0689] User's device: A smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The user captures an image of the menu with the device's camera and sends it to a cloud server. The device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[0690] Server: The server on the cloud processes the received image data using advanced AI technology and has the following functions:

[0691] 1. OCR (Optical Character Recognition): Uses OCR technology, such as Google Cloud Vision API, to extract text information from received images.

[0692] 2. Translation: To translate the extracted text information into multiple languages, use, for example, the Google Translate API.

[0693] 3. Internet search: Using the translated text, you can search for food images and descriptions on the Internet, for example, using the Bing Search API.

[0694] 4. Summary generation: A generative AI model such as GPT-3 is used to generate a summary from the acquired information.

[0695] 5. Information encryption: The generated information is encrypted and sent to the user terminal using the Advanced Encryption Standard (AES).

[0696] Specific examples

[0697] Consider a scenario where a user points their camera at a Japanese menu item that says "tempura" at a Japanese restaurant. The user's device sends the image to a server in the cloud. The server analyzes the image and extracts the text "tempura" using OCR technology. The server then translates "tempura" to "Tempura" using the Google Translate API. The server then retrieves images and descriptions related to "Tempura" from the Internet using the Bing Search API. Based on the retrieved information, the server uses GPT-3 to generate a summary: "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture." This information is then encrypted using AES and sent to the user's device. The user's device decrypts the received information and displays it to the user through an application. The emotion engine recognizes the user's facial expression as an indication of interest and simultaneously displays other related recommended dishes.

[0698] Prompt Sentence Examples

[0699] "Let the user capture an image of a restaurant menu, extract and translate text from the image, and provide relevant dish images and descriptions. Also, recognize the user's emotions and adjust to provide a personalized information experience."

[0700] This system allows users to intuitively understand the contents of the dish and receive information that is tailored to their individual needs.

[0701] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0702] Step 1:

[0703] The user captures an image of the menu

[0704] The user takes a picture of a restaurant menu using the camera on their smartphone or tablet. This image becomes the initial input for the system. Specifically, the user launches the device's camera application and captures an image of the menu. The image data is then saved via a dedicated application, and the system proceeds to the next step.

[0705] Step 2:

[0706] The device sends the image to the server

[0707] After image data is captured by the camera application, a dedicated application sends the image to a server on the cloud. The input is the captured image data, and the output is the image data sent to the server. Specifically, the application on the device sends the image data to a specified endpoint on the server via an Internet connection.

[0708] Step 3:

[0709] The server performs image recognition and OCR processing

[0710] The server receives image data sent from the device. The input is the sent image data, and the output is the extracted text information. The image recognition module in the server analyzes the image and converts the text information in the image into digital text using OCR (optical character recognition) technology. Specifically, the server uses the "Google Cloud Vision API" to extract the text information in the image.

[0711] Step 4:

[0712] The server translates the extracted text.

[0713] The server takes the character information extracted by the OCR process as input and sends this information to a translation engine. The output is the translated text. The server uses the "Google Translate API" to translate the text into the specified language. Specifically, the server translates the extracted Japanese character information "Tempura" into "Tempura."

[0714] Step 5:

[0715] The server performs an internet search based on the translation results

[0716] Using the translated text as input, the server performs an internet search to gather images and descriptions of related dishes. The output is the retrieved images and descriptions. The server uses the Bing Search API to search for information about "Tempura." Specifically, the server gathers images and descriptions related to "Tempura" from multiple websites.

[0717] Step 6:

[0718] The server generates a summary

[0719] Using the acquired information as input, the server uses a generative AI model to generate a concise summary. The output is summarized text. The server uses GPT-3 to summarize the information, generating a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." Specifically, the server performs a process to summarize the collected descriptions using the generative AI model.

[0720] Step 7:

[0721] The server encrypts the information and sends it to the user's device.

[0722] The server takes the generated image and summary as input and encrypts this information. The output is encrypted data. The server encrypts the data using AES (Advanced Encryption Standard) and sends it to the user's device. Specifically, the server uses AES encryption to keep the data secure and sends it to the device via the Internet.

[0723] Step 8:

[0724] Display information received by the device

[0725] The device receives encrypted information from the server as input, decrypts it, and visually displays it to the user. The output is an image of the displayed dish and a summary. A dedicated application on the device decrypts the encrypted data and displays an image of the dish along with a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture." Specifically, the device uses the received key data to decrypt the data and displays it through the application interface.

[0726] Step 9:

[0727] The device recognizes the user's emotions and adjusts the information accordingly.

[0728] Once again, the emotion engine uses the user's facial expression and tone of voice as input to recognize the emotion and adjust the information. The output is additional information based on the user's emotion. The emotion engine uses the "Emotion API" to analyze the user's emotion, and if the user shows interest, it displays additional information about related dishes or recommendations. Specifically, the device uses the camera and microphone to analyze the user's facial expression and voice in real time, dynamically changing the information displayed.

[0729] (Application example 2)

[0730] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] In today's increasingly internationalized world, understanding the menu contents of restaurants abroad is an important factor for travelers. However, language barriers and lack of knowledge about cuisine can make it difficult for travelers to accurately understand the menu. Furthermore, conventional menu translation systems lack the ability to provide personalized information that takes into account the user's interests and preferences, making it difficult for them to choose a meal. To resolve this situation, it is necessary to recognize the user's emotions and provide information based on them.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0733] In this invention, the server includes optical character recognition means for extracting text information from received images, translation means for translating the extracted text information, and means for searching the Internet for food images and descriptions based on the translation results, thereby making it easier to understand menu contents beyond language barriers and enabling personalized food suggestions.

[0734] A "camera" is a device that captures light and records images or videos.

[0735] A "cloud server" is a remote computing resource accessed via the Internet, a computer system that processes and stores data.

[0736] "Optical character recognition means" is a technology that recognizes character information from image data and extracts it as text data.

[0737] A "translation tool" is a technique for converting text in one language into text in another language.

[0738] "Means of searching the Internet" refers to the technology of obtaining information that is publicly available online through search engines or APIs.

[0739] The "means for generating food images and descriptions" is a technology that constructs visual and text descriptions based on information obtained from search results.

[0740] "Means for transmitting to the user terminal" refers to the technology for transferring data from a server on the cloud to the user's device.

[0741] "Emotion recognition means" is a technology that analyzes and identifies a user's emotional state from facial expressions, tone of voice, etc.

[0742] "Means for making personalized dish suggestions" refers to technology that makes individually optimized dish suggestions based on the user's emotional data and preferences.

[0743] "Means for displaying on the user terminal" refers to a technique for visually displaying the provided information on the user's device.

[0744] This invention combines a system that captures menu images with a camera, extracts and translates text information from the images, and searches the internet for related food images and descriptions to display them, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system are described below.

[0745] System Configuration

[0746] User device:

[0747] The user's device is typically a smartphone or tablet. The device is equipped with a camera and is connected to the Internet. A dedicated application is installed on the device, which has the function of capturing images of menus and sending them to a cloud server. Additionally, the device is equipped with an emotion engine, which can analyze the user's facial expressions and tone of voice.

[0748] Cloud Server:

[0749] The cloud server has high-performance computing resources and performs the following processes:

[0750] 1. Image and Character Recognition (OCR):

[0751] The server receives the menu image sent from the user device and uses optical character recognition software (e.g., pytesseract) to extract the text information within the image. This process results in the names and descriptions of the dishes on the menu being captured as digital text.

[0752] 2. Text Translation:

[0753] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[0754] 3. Internet Search:

[0755] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database (e.g., Google search), such as images and descriptions related to "Tempura."

[0756] 4. Information generation:

[0757] Using the acquired information, AI technology generates a concise description, such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture."

[0758] 5. Encryption and Transmission of Information:

[0759] The generated image and description are encrypted as a security measure and sent to the user's terminal.

[0760] Emotion Engine:

[0761] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns, enabling it to provide information tailored to the user's interests and preferences.

[0762] System Operation

[0763] When a user captures a menu image with their camera, the image is sent to a cloud server. The server extracts text information from the image and translates it into the specified language. Based on the translation results, the server searches the internet for related food images and descriptions, and generates a summary of the information.

[0764] The generated information is tailored based on the user's emotional data. For example, if the emotion engine recognizes that the user is interested, it will provide related information and recommended dishes. This allows users to get personalized information that will help them make better food choices.

[0765] Examples:

[0766] When a user points their camera at the Japanese menu item "Tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text information "Tempura" from the received image and translates it to "Tempura." The server then uses the Internet to obtain information related to the image of tempura and generates a description. The description it creates is "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture," and the information is encrypted and sent to the user's device. The emotion engine recognizes interest from the user's facial expression and displays other recommended dishes.

[0767] Example prompt sentence:

[0768] "Generate a Python program that captures an image of a dish called tempura on a menu at a Japanese restaurant, extracts and translates the text from the image, analyzes the user's emotions from their facial expressions, and provides the user with the most appropriate food information based on their emotions."

[0769] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0770] Step 1:

[0771] The device captures an image of the menu with its camera.

[0772] Input: Restaurant menu image

[0773] Output: Captured image data

[0774] Specific action: The user takes a photo of the menu using the camera on their smartphone or tablet.

[0775] Step 2:

[0776] The device sends the captured image to a server on the cloud.

[0777] Input: Captured image data

[0778] Output: Image data sent to the server

[0779] Specific operation: The device uses an internet connection to upload image data to a server in the cloud.

[0780] Step 3:

[0781] The server extracts text information from the received image using optical character recognition (OCR).

[0782] Input: Image data sent to the server

[0783] Output: Extracted text information (text data)

[0784] Specific operation: The server uses OCR software (e.g., pytesseract) to analyze the characters in the image and generate text data.

[0785] Step 4:

[0786] The server translates the extracted text information into the specified language using a translation engine.

[0787] Input: Extracted text information (text data)

[0788] Output: Translated text information (text data)

[0789] What happens: The server uses a translation engine, such as the Google Translate API, to translate the text into the specified language.

[0790] Step 5:

[0791] The server searches the internet for images and descriptions of related dishes based on the translation results.

[0792] Input: Translated text information (text data)

[0793] Output: Images and descriptions of the dishes obtained

[0794] Specific operation: The server uses a search engine API to retrieve information about dishes related to the translated text from the Internet.

[0795] Step 6:

[0796] The server generates an image and description of the dish from the information it obtains.

[0797] Input: Image and description of the food obtained

[0798] Output: Generated food images and abbreviated descriptions

[0799] Specific operation: The server uses AI technology to organize the acquired information and generate visually easy-to-understand images and concise explanations.

[0800] Step 7:

[0801] The server encrypts the generated information and transmits it to the user terminal.

[0802] Input: Generated food image and description

[0803] Output: Encrypted dish image and description

[0804] Specific operation: The server encrypts the information as a security measure and sends it to the user's terminal via the Internet.

[0805] Step 8:

[0806] The device analyzes the user's facial expressions and tone of voice to recognize emotions.

[0807] Input: User's facial expression and voice data

[0808] Output: Recognized emotion data

[0809] Specific operation: The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them through an emotion engine.

[0810] Step 9:

[0811] The server makes personalized dish suggestions based on emotion recognition results.

[0812] Input: Recognized emotion data, user behavior history

[0813] Output: Personalized food suggestions

[0814] Specific operation: The server generates individually optimized dish suggestions based on the emotion data and the user's previous history and sends them to the user's device.

[0815] Step 10:

[0816] The terminal displays the transmitted information to the user.

[0817] Input: Encrypted food images, abbreviated descriptions, and personalized food suggestions

[0818] Output: Food images, descriptions, and suggestions displayed on the user's device

[0819] Specific operation: The device decodes the received information and displays it visually to the user.

[0820] The above is a specific processing flow for carrying out the present invention.

[0821] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0822] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0823] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0824] [Third embodiment]

[0825] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0826] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0827] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0828] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0829] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0830] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0831] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0832] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0833] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0834] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0835] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0836] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0837] The following system and program are provided as embodiments of the present invention.

[0838] The system is broadly composed of the user's terminal, a server on the cloud, and a database connected to the Internet.

[0839] User terminal

[0840] The user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user device is to capture images of the menu and send them to a server on the cloud.

[0841] server

[0842] The cloud server processes the received image data using advanced AI technology. The specific processing performed by the server is shown below.

[0843] 1. Image and Character Recognition (OCR):

[0844] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, obtaining the names and descriptions of the dishes on the menu as digital text.

[0845] 2. Text Translation:

[0846] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, the Japanese character for "tempura" is translated into English.

[0847] 3. Internet Search:

[0848] The server uses the translated dish name to search image and information databases on the Internet to retrieve images and descriptions of the corresponding dish, such as images related to "Tempura," along with information about the dish's characteristics and summary.

[0849] 4. Summary sentence generation:

[0850] AI technology is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, flavor characteristics, etc. For example, a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0851] 5. Encryption and Transmission of Information:

[0852] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[0853] Display on user device

[0854] The user device displays the information received from the server through a dedicated application. This allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0855] Specific examples

[0856] Specific examples are shown below.

[0857] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0858] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0859] The processing flow will be explained below.

[0860] Step 1:

[0861] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[0862] Step 2:

[0863] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[0864] Step 3:

[0865] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[0866] Step 4:

[0867] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[0868] Step 5:

[0869] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[0870] Step 6:

[0871] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0872] Step 7:

[0873] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[0874] Step 8:

[0875] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[0876] Step 9:

[0877] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[0878] Step 10:

[0879] The user decides which dish to order based on the displayed information. For example, after seeing an image of tempura and a summary, the user decides whether to order that dish.

[0880] Through this series of steps, the user can easily and intuitively understand the menu contents and select dishes.

[0881] Example 1

[0882] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0883] When travelers or users unfamiliar with foreign languages ​​visit restaurants overseas, they often encounter problems due to language barriers and cultural differences when trying to understand the menu and select the appropriate dish. These issues reduce convenience for travelers and detract from their dining experience.

[0884] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0885] In this invention, the server includes optical character recognition means for extracting text information from images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, and means for generating summaries of the searched food images and descriptions using a generation AI model, thereby enabling users to quickly and accurately understand menus in foreign languages ​​and easily select dishes.

[0886] A "camera" is a device for taking pictures or videos.

[0887] A "cloud server" is a remotely located computer system that can be accessed via the Internet for storing data and performing computations.

[0888] "Optical character recognition" refers to technology or devices that extract character information from an image as digital text.

[0889] A "translation tool" is a technique or device for converting text expressed in a particular language into another language.

[0890] A "generative AI model" is an artificial intelligence algorithm or set of systems that learns from large amounts of data and performs tasks such as text generation, classification, and translation.

[0891] A "means for generating a summary" is a technology or device that extracts important points from the provided information and summarizes them in a concise sentence.

[0892] An "encryption method" is a technology or device that converts data based on a specific algorithm to prevent unauthorized decryption by third parties.

[0893] A "user terminal" is a computer system operated by a user, and includes mobile phones, smartphones, tablets, and the like.

[0894] "Internet search tools" are techniques or devices that retrieve information from online databases or sources based on specific keywords.

[0895] The following system and program are provided as embodiments of the present invention: The system is broadly composed of a user terminal, a server on the cloud, and a database connected to the Internet.

[0896] User terminal

[0897] A user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user's device is to capture images of the menu and send them to a server on the cloud. The user takes a picture of the menu using the camera function and sends it to the server via the dedicated application.

[0898] server

[0899] The cloud server processes the received image data using advanced AI technology. Below are some examples of specific hardware and software used on the server.

[0900] 1. Image and Character Recognition (OCR):

[0901] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, using libraries such as TensorFlow and OpenCV. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[0902] 2. Text Translation:

[0903] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, the Japanese character for "tempura" (Tempura) is translated into English.

[0904] 3. Internet Search:

[0905] The server then uses the translated dish name to search image and information databases on the Internet, using the Bing Search API and other services. For example, it retrieves images related to "Tempura" and information about the dish, including its general description and characteristics.

[0906] 4. Summary sentence generation:

[0907] A generative AI model (e.g., GPT-4) is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, and flavor characteristics. A prompt sentence is used to input the model, and a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[0908] Example prompt sentence:

[0909] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0910] 5. Encryption and Transmission of Information:

[0911] The generated image and summary are encrypted using the SSL / TLS protocol and sent to the user's device. The server then encrypts the data using an encryption library and sends it via an HTTP POST request.

[0912] Display on user device

[0913] The user device decrypts the encrypted data received from the server and displays it through a dedicated application. This process allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[0914] Specific examples

[0915] Specific examples are shown below.

[0916] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[0917] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[0918] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0919] Step 1:

[0920] Image capture and transmission

[0921] Users take a photo of a restaurant menu with their smartphone or tablet, and the device acquires the captured image and sends it to a cloud server via a dedicated application using an HTTP POST request.

[0922] Input: A user-taken image of the menu

[0923] Output: Image data sent to a cloud server

[0924] Step 2:

[0925] Image and character recognition (OCR)

[0926] The server processes the received image data and extracts the text information from the image using optical character recognition (OCR) technology. This process uses libraries such as TensorFlow and OpenCV. The server analyzes each pixel of the image and identifies areas that can be recognized as text.

[0927] Input: Image data sent to a cloud server

[0928] Output: Extracted text information (e.g. "tempura")

[0929] Step 3:

[0930] Text Translation

[0931] The server inputs the extracted text information into the Google Translate API and translates it into the specified language. It creates an API request, sends the extracted text information, and receives the translation result. For example, the Japanese text information "tempura" is translated into English "Tempura."

[0932] Input: Extracted text information (e.g., "tempura")

[0933] Output: Translated text (e.g. "Tempura")

[0934] Step 4:

[0935] Internet search

[0936] The server uses the translated text as a key to search for related information on the Internet using the Bing Search API. Specifically, it sends an API request and receives search results, which retrieves information such as images and descriptions related to "Tempura."

[0937] Input: Translated text (e.g. "Tempura")

[0938] Output: Images and descriptions of the dishes retrieved as search results

[0939] Step 5:

[0940] Summary sentence generation

[0941] Based on the acquired information, the server uses a generative AI model (e.g., GPT-4) to generate a concise summary. The server creates a prompt sentence and inputs it into the model to generate a summary. For example, the summary generated is, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[0942] Input: Images and descriptions of dishes obtained as search results

[0943] Output: Generated summary

[0944] Example prompt sentence:

[0945] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[0946] Step 6:

[0947] Encryption and transmission of information

[0948] The server encrypts the generated image and summary using the SSL / TLS protocol and sends it to the user's terminal. The server also encrypts the data using an encryption library and then sends it via an HTTP POST request.

[0949] Input: Generated summary and food image

[0950] Output: Sent to the user's device as encrypted data

[0951] Step 7:

[0952] Display on user device

[0953] The user device decrypts the encrypted data received from the server and displays it using a dedicated application. Specifically, the device performs decryption processing based on the SSL / TLS protocol and displays an image of tempura and a summary to the user.

[0954] Input: Encrypted data sent by the server

[0955] Output: A picture of the food and a summary of it displayed on the user's device

[0956] By following the above steps, the user can easily understand and select from the menu at a restaurant overseas.

[0957] (Application example 1)

[0958] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0959] Modern restaurants face the challenge of making it difficult for visitors to intuitively understand the dishes and their details on the menu in multiple languages. This is particularly true for tourists who cannot understand foreign languages, as they are unable to immediately understand the dish descriptions, allergen information, prices, and ratings. This can lead to inconvenience when ordering and a decrease in customer satisfaction. Therefore, there is a need for a system that provides multilingual and visually easy-to-understand information about dishes.

[0960] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0961] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a cloud server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for summarizing the generated food images and descriptions as detailed information about dishes in the restaurant, means for transmitting the summarized information to a user terminal, means for displaying the transmitted information on the user terminal, and means for allowing the user to check related image data and rating information on the terminal. This allows visitors to instantly understand detailed information and ratings about dishes in multiple languages, eliminating inconvenience when ordering and improving customer satisfaction.

[0962] "Means for capturing images of menus with a camera" is a function that allows a user to take a photo of a restaurant menu using a mobile device such as a smartphone or tablet and obtain the image data.

[0963] The "means for transmitting the image to a server on the cloud" is a function for uploading image data taken by a mobile terminal to a remote cloud server via the Internet.

[0964] "Optical character recognition means" refers to a technology for extracting character information from received image data, and in particular, uses OCR (optical character recognition) technology.

[0965] The "translation means" is a function for converting extracted text information into a different language as needed.

[0966] The "means for searching on the Internet" is a function for searching an online database based on the translated text information and obtaining images and descriptions of related dishes.

[0967] The "means for generating" is a function for processing the searched food image and description into the format required to provide it to the user.

[0968] The "means for summarizing detailed dish information" is a function that extracts and displays a concise explanation and essential points from the generated image and description of the dish so that the user can easily understand it.

[0969] The "means for transmitting to the user terminal" is a function for transmitting the generated image of the dish and a summarized description to the user's mobile terminal.

[0970] The "means for displaying on the user terminal" is a function for visually displaying the transmitted image of the dish and the summarized description on the user's mobile terminal.

[0971] The "means for checking related image data and evaluation information" is a function that allows the user to refer to image data related to the displayed dish information and evaluation information from other users.

[0972] As an embodiment of the present invention, the following system and program are provided. This system is composed of a user's mobile terminal, a cloud server, and a database connected to the Internet. Specifically, each component functions as follows:

[0973] User terminal

[0974] A user's mobile device, such as a smartphone or tablet, is equipped with a camera, Internet connectivity, and a dedicated application. The main role of the user device is to capture images of menus and send them to a server on the cloud. The user takes a photo of a restaurant menu using the camera and uploads the image to the server via the application.

[0975] server

[0976] The server processes the received image data using advanced AI technology. The server performs the processing using the following specific hardware and software:

[0977] Hardware: High-performance processor (e.g., Intel Xeon), large memory capacity

[0978] Software: Tesseract (OCR engine), Google Translate API, Python

[0979] Image Recognition and Translation

[0980] The server first uses OCR technology to extract text information from the menu image sent from the device. The extracted text information is then translated into the specified language using the Google Translate API. For example, the Japanese character for "tempura" is translated into "Tempura."

[0981] Data Search and Generation

[0982] The translated text is then used to search an online database to retrieve images and descriptions of the corresponding dishes. For example, images related to "Tempura" are retrieved, along with information about the dish's characteristics. This information is then used to summarize the details of the restaurant's dishes. The summary includes information about the main ingredients, cooking methods, and flavor characteristics.

[0983] Sending and Displaying Information

[0984] The summarized images and descriptions are encrypted for security reasons and sent to the user's device. The user's device displays the information received from the server through a dedicated application. This allows users to view images of dishes and brief descriptions simply by pointing their camera at the menu. They can also view related image data and review information.

[0985] Specific examples

[0986] Below is a specific example. When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server uses OCR technology to extract the text information for "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and the image of tempura, a description, and user ratings are displayed on the user's application.

[0987] Example prompts for generative AI models

[0988] "Please translate the Japanese menu items in this image into English."

[0989] "Get detailed information about ramen, including ingredients, cooking method, and flavor characteristics."

[0990] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0991] Step 1:

[0992] A user takes a picture of a restaurant menu using the camera on their mobile device, and the image is sent to a cloud server via the application.

[0993] Input: A menu image taken by the user.

[0994] Output: Menu images sent to a server on the cloud.

[0995] Step 2:

[0996] The server processes the received image data using an OCR engine (Tesseract) to extract the text information from the image, and as a result, the menu items are converted into text data.

[0997] Input: Menu image sent to a server on the cloud.

[0998] Output: Character information (text data) extracted by the OCR engine.

[0999] Step 3:

[1000] The extracted text information is translated into the specified language using the Google Translate API.

[1001] Input: Extracted character information (text data).

[1002] Output: Translated text information (translated text data).

[1003] Step 4:

[1004] The server searches an online database based on the translated text information and retrieves images and descriptions of the corresponding dishes.

[1005] Input: Translated text.

[1006] Output: Images and descriptions of dishes retrieved through internet searches.

[1007] Step 5:

[1008] The server generates a summary from the acquired cooking information, including the main ingredients, cooking method, flavor characteristics, etc.

[1009] Input: Food images and descriptions obtained from an internet search.

[1010] Output: A condensed description of the dish.

[1011] Step 6:

[1012] The generated image and summarized description are encrypted for security reasons and sent to the user's terminal.

[1013] Input: Abridged dish description and image.

[1014] Output: Encrypted information (abridged description and image of the dish).

[1015] Step 7:

[1016] The user terminal decodes the received information using a dedicated application and displays an image of the dish, a summary of its description, related image data, and rating information.

[1017] Input: Encrypted information.

[1018] Output: Decoded and displayed dish image on the user's device, along with a summary description, associated image data, and rating information.

[1019] Specific operation example

[1020] When a user points their camera at the Japanese menu item "tempura" and takes a photo, the image is sent to a cloud server. The server uses OCR technology to extract the text information for "tempura" and translates it into "Tempura" using the Google Translate API. It then searches the internet for images and descriptions related to tempura and summarizes it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." This information is then encrypted and sent to the user's device, where it is displayed on the application.

[1021] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1022] This invention combines a system that captures menu images, extracts and translates text information from the images, and searches the internet for related food images and descriptions, with an emotion engine that recognizes the user's emotions. This system makes it possible to provide information according to the user's emotions, providing a more personalized experience.

[1023] User terminal

[1024] The user's device is a smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The main function of the user's device is to capture menu images and send them to a cloud server. In addition, the device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[1025] server

[1026] The cloud server has the main function of processing the received image data using advanced AI technology. The specific processing details are shown below.

[1027] 1. Image and Character Recognition (OCR):

[1028] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition technology. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[1029] 2. Text Translation:

[1030] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[1031] 3. Internet Search:

[1032] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database, such as images and descriptions related to "Tempura."

[1033] 4. Summary sentence generation:

[1034] From the acquired information, AI technology is used to generate a concise summary, such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1035] 5. Encryption and Transmission of Information:

[1036] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[1037] Emotion Engine

[1038] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns. This makes it possible to provide information tailored to the user's interests and preferences.

[1039] Display on user device

[1040] The user device displays the information received from the server through the application. The display content is adjusted according to the user's emotions by the emotion engine. This allows the user to obtain the most appropriate information and helps them make food selections. For example, if a user points the camera at a menu item called "tempura," a description and visual information will be displayed immediately, along with related recommendations based on the user's emotions.

[1041] Specific examples

[1042] Specific examples are shown below.

[1043] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve information related to the tempura image and generates a summary. The summary, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture," is created, encrypted, and sent to the user's device. Furthermore, the emotion engine recognizes interest from the user's facial expression and displays other recommended dishes. This allows the user to view information related to tempura visually and through text, making the best choice.

[1044] Through this system, users can intuitively understand the contents of the dishes and receive personalized information, making it easier to select dishes.

[1045] The processing flow will be explained below.

[1046] Step 1:

[1047] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[1048] Step 2:

[1049] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[1050] Step 3:

[1051] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[1052] Step 4:

[1053] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[1054] Step 5:

[1055] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[1056] Step 6:

[1057] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1058] Step 7:

[1059] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[1060] Step 8:

[1061] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[1062] Step 9:

[1063] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[1064] Step 10:

[1065] The emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, if the user's facial expression indicates joy or interest, that information is recorded.

[1066] Step 11:

[1067] The emotion engine uses the user's emotion information to provide additional information tailored to the user's interests and preferences. For example, if a user is interested in tempura, it will suggest other fried dishes and Japanese food options.

[1068] Step 12:

[1069] The device displays information provided by the emotion engine, and the user can review it and select a dish. For example, in addition to information about tempura, images and descriptions of katsudon and karaage are also displayed.

[1070] This series of steps allows users to intuitively understand the menu contents and select the best dish based on personalized information that responds to their emotions.

[1071] Example 2

[1072] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1073] Conventional menu translation systems have the problem of not providing information that takes into account the user's emotions and interests, and can only display uniform information. As a result, it is difficult for users to quickly obtain the information they really want, and satisfaction cannot be increased. In addition, there are limitations to the accuracy of translation and the reliability of the information, so there is a demand for more accurate information provision.

[1074] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1075] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a remote server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching a network for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for transmitting the food images and descriptions to a user terminal, means for displaying the transmitted information on the user terminal, means for recognizing emotions by analyzing a user's facial expressions and tone of voice, and means for adjusting information based on the emotion recognition, thereby enabling the provision of personalized information that takes into account the user's emotions and interests.

[1076] The "means for capturing an image of a menu using a camera" is a function for taking a digital image of a physical menu using a camera mounted on a user terminal.

[1077] The "means for transmitting the image to a remote server" is a function for transmitting captured image data from a user terminal to a server on the cloud using an internet connection.

[1078] "Optical character recognition means for extracting character information from received images" refers to technology for recognizing character information from transmitted image data and converting it into digital text.

[1079] "Translation means for translating extracted text information" refers to technology for converting text information obtained by OCR into another language.

[1080] "Means for searching the network for food images and descriptions based on the translation results" is a technology that uses the translated text information as input to obtain related images and descriptions using the Internet.

[1081] "Means for generating images and descriptions of searched dishes" is a function that creates images of dishes and related descriptions based on information obtained from the network.

[1082] The "means for transmitting the image and description of the dish to the user terminal" is a technique for transferring the generated image and description of the dish from a remote server to the user terminal.

[1083] The "means for displaying the transmitted information on the user terminal" is a function for visually displaying the image and description of the dish received on the user terminal.

[1084] "Means for recognizing emotions by analyzing a user's facial expressions and tone of voice" refers to technology that uses a camera and microphone to analyze the user's facial expressions and voice to determine the user's emotional state.

[1085] The "means for adjusting information based on emotion recognition" is a technology that dynamically changes the content of the information to be displayed according to the user's recognized emotion.

[1086] The system of the present invention allows users to capture images of menus, extract and translate textual information from the images, search the internet for related food images and descriptions, and even recognize the user's emotions to provide personalized information, providing users with more intuitive and individually tailored information.

[1087] Hardware and software used

[1088] User's device: A smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The user captures an image of the menu with the device's camera and sends it to a cloud server. The device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[1089] Server: The server on the cloud processes the received image data using advanced AI technology and has the following functions:

[1090] 1. OCR (Optical Character Recognition): Uses OCR technology, such as Google Cloud Vision API, to extract text information from received images.

[1091] 2. Translation: To translate the extracted text information into multiple languages, use, for example, the Google Translate API.

[1092] 3. Internet search: Using the translated text, you can search for food images and descriptions on the Internet, for example, using the Bing Search API.

[1093] 4. Summary generation: A generative AI model such as GPT-3 is used to generate a summary from the acquired information.

[1094] 5. Information encryption: The generated information is encrypted and sent to the user terminal using the Advanced Encryption Standard (AES).

[1095] Specific examples

[1096] Consider a scenario where a user points their camera at a Japanese menu item that says "tempura" at a Japanese restaurant. The user's device sends the image to a server in the cloud. The server analyzes the image and extracts the text "tempura" using OCR technology. The server then translates "tempura" to "Tempura" using the Google Translate API. The server then retrieves images and descriptions related to "Tempura" from the Internet using the Bing Search API. Based on the retrieved information, the server uses GPT-3 to generate a summary: "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture." This information is then encrypted using AES and sent to the user's device. The user's device decrypts the received information and displays it to the user through an application. The emotion engine recognizes the user's facial expression as an indication of interest and simultaneously displays other related recommended dishes.

[1097] Prompt Sentence Examples

[1098] "Let the user capture an image of a restaurant menu, extract and translate text from the image, and provide relevant dish images and descriptions. Also, recognize the user's emotions and adjust to provide a personalized information experience."

[1099] This system allows users to intuitively understand the contents of the dish and receive information that is tailored to their individual needs.

[1100] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1101] Step 1:

[1102] The user captures an image of the menu

[1103] The user takes a picture of a restaurant menu using the camera on their smartphone or tablet. This image becomes the initial input for the system. Specifically, the user launches the device's camera application and captures an image of the menu. The image data is then saved via a dedicated application, and the system proceeds to the next step.

[1104] Step 2:

[1105] The device sends the image to the server

[1106] After image data is captured by the camera application, a dedicated application sends the image to a server on the cloud. The input is the captured image data, and the output is the image data sent to the server. Specifically, the application on the device sends the image data to a specified endpoint on the server via an Internet connection.

[1107] Step 3:

[1108] The server performs image recognition and OCR processing

[1109] The server receives image data sent from the device. The input is the sent image data, and the output is the extracted text information. The image recognition module in the server analyzes the image and converts the text information in the image into digital text using OCR (optical character recognition) technology. Specifically, the server uses the "Google Cloud Vision API" to extract the text information in the image.

[1110] Step 4:

[1111] The server translates the extracted text.

[1112] The server takes the character information extracted by the OCR process as input and sends this information to a translation engine. The output is the translated text. The server uses the "Google Translate API" to translate the text into the specified language. Specifically, the server translates the extracted Japanese character information "Tempura" into "Tempura."

[1113] Step 5:

[1114] The server performs an internet search based on the translation results

[1115] Using the translated text as input, the server performs an internet search to gather images and descriptions of related dishes. The output is the retrieved images and descriptions. The server uses the Bing Search API to search for information about "Tempura." Specifically, the server gathers images and descriptions related to "Tempura" from multiple websites.

[1116] Step 6:

[1117] The server generates a summary

[1118] Using the acquired information as input, the server uses a generative AI model to generate a concise summary. The output is summarized text. The server uses GPT-3 to summarize the information, generating a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." Specifically, the server performs a process to summarize the collected descriptions using the generative AI model.

[1119] Step 7:

[1120] The server encrypts the information and sends it to the user's device.

[1121] The server takes the generated image and summary as input and encrypts this information. The output is encrypted data. The server encrypts the data using AES (Advanced Encryption Standard) and sends it to the user's device. Specifically, the server uses AES encryption to keep the data secure and sends it to the device via the Internet.

[1122] Step 8:

[1123] Display information received by the device

[1124] The device receives encrypted information from the server as input, decrypts it, and visually displays it to the user. The output is an image of the displayed dish and a summary. A dedicated application on the device decrypts the encrypted data and displays an image of the dish along with a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture." Specifically, the device uses the received key data to decrypt the data and displays it through the application interface.

[1125] Step 9:

[1126] The device recognizes the user's emotions and adjusts the information accordingly.

[1127] Once again, the emotion engine uses the user's facial expression and tone of voice as input to recognize the emotion and adjust the information. The output is additional information based on the user's emotion. The emotion engine uses the "Emotion API" to analyze the user's emotion, and if the user shows interest, it displays additional information about related dishes or recommendations. Specifically, the device uses the camera and microphone to analyze the user's facial expression and voice in real time, dynamically changing the information displayed.

[1128] (Application example 2)

[1129] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1130] In today's increasingly internationalized world, understanding the menu contents of restaurants abroad is an important factor for travelers. However, language barriers and lack of knowledge about cuisine can make it difficult for travelers to accurately understand the menu. Furthermore, conventional menu translation systems lack the ability to provide personalized information that takes into account the user's interests and preferences, making it difficult for them to choose a meal. To resolve this situation, it is necessary to recognize the user's emotions and provide information based on them.

[1131] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1132] In this invention, the server includes optical character recognition means for extracting text information from received images, translation means for translating the extracted text information, and means for searching the Internet for food images and descriptions based on the translation results, thereby making it easier to understand menu contents beyond language barriers and enabling personalized food suggestions.

[1133] A "camera" is a device that captures light and records images or videos.

[1134] A "cloud server" is a remote computing resource accessed via the Internet, a computer system that processes and stores data.

[1135] "Optical character recognition means" is a technology that recognizes character information from image data and extracts it as text data.

[1136] A "translation tool" is a technique for converting text in one language into text in another language.

[1137] "Means of searching the Internet" refers to the technology of obtaining information that is publicly available online through search engines or APIs.

[1138] The "means for generating food images and descriptions" is a technology that constructs visual and text descriptions based on information obtained from search results.

[1139] "Means for transmitting to the user terminal" refers to the technology for transferring data from a server on the cloud to the user's device.

[1140] "Emotion recognition means" is a technology that analyzes and identifies a user's emotional state from facial expressions, tone of voice, etc.

[1141] "Means for making personalized dish suggestions" refers to technology that makes individually optimized dish suggestions based on the user's emotional data and preferences.

[1142] "Means for displaying on the user terminal" refers to a technique for visually displaying the provided information on the user's device.

[1143] This invention combines a system that captures menu images with a camera, extracts and translates text information from the images, and searches the internet for related food images and descriptions to display them, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system are described below.

[1144] System Configuration

[1145] User device:

[1146] The user's device is typically a smartphone or tablet. The device is equipped with a camera and is connected to the Internet. A dedicated application is installed on the device, which has the function of capturing images of menus and sending them to a cloud server. Additionally, the device is equipped with an emotion engine, which can analyze the user's facial expressions and tone of voice.

[1147] Cloud Server:

[1148] The cloud server has high-performance computing resources and performs the following processes:

[1149] 1. Image and Character Recognition (OCR):

[1150] The server receives the menu image sent from the user device and uses optical character recognition software (e.g., pytesseract) to extract the text information within the image. This process results in the names and descriptions of the dishes on the menu being captured as digital text.

[1151] 2. Text Translation:

[1152] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[1153] 3. Internet Search:

[1154] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database (e.g., Google search), such as images and descriptions related to "Tempura."

[1155] 4. Information generation:

[1156] Using the acquired information, AI technology generates a concise description, such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture."

[1157] 5. Encryption and Transmission of Information:

[1158] The generated image and description are encrypted as a security measure and sent to the user's terminal.

[1159] Emotion Engine:

[1160] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns, enabling it to provide information tailored to the user's interests and preferences.

[1161] System Operation

[1162] When a user captures a menu image with their camera, the image is sent to a cloud server. The server extracts text information from the image and translates it into the specified language. Based on the translation results, the server searches the internet for related food images and descriptions, and generates a summary of the information.

[1163] The generated information is tailored based on the user's emotional data. For example, if the emotion engine recognizes that the user is interested, it will provide related information and recommended dishes. This allows users to get personalized information that will help them make better food choices.

[1164] Examples:

[1165] When a user points their camera at the Japanese menu item "Tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text information "Tempura" from the received image and translates it to "Tempura." The server then uses the Internet to obtain information related to the image of tempura and generates a description. The description it creates is "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture," and the information is encrypted and sent to the user's device. The emotion engine recognizes interest from the user's facial expression and displays other recommended dishes.

[1166] Example prompt sentence:

[1167] "Generate a Python program that captures an image of a dish called tempura on a menu at a Japanese restaurant, extracts and translates the text from the image, analyzes the user's emotions from their facial expressions, and provides the user with the most appropriate food information based on their emotions."

[1168] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1169] Step 1:

[1170] The device captures an image of the menu with its camera.

[1171] Input: Restaurant menu image

[1172] Output: Captured image data

[1173] Specific action: The user takes a photo of the menu using the camera on their smartphone or tablet.

[1174] Step 2:

[1175] The device sends the captured image to a server on the cloud.

[1176] Input: Captured image data

[1177] Output: Image data sent to the server

[1178] Specific operation: The device uses an internet connection to upload image data to a server in the cloud.

[1179] Step 3:

[1180] The server extracts text information from the received image using optical character recognition (OCR).

[1181] Input: Image data sent to the server

[1182] Output: Extracted text information (text data)

[1183] Specific operation: The server uses OCR software (e.g., pytesseract) to analyze the characters in the image and generate text data.

[1184] Step 4:

[1185] The server translates the extracted text information into the specified language using a translation engine.

[1186] Input: Extracted text information (text data)

[1187] Output: Translated text information (text data)

[1188] What happens: The server uses a translation engine, such as the Google Translate API, to translate the text into the specified language.

[1189] Step 5:

[1190] The server searches the internet for images and descriptions of related dishes based on the translation results.

[1191] Input: Translated text information (text data)

[1192] Output: Images and descriptions of the dishes obtained

[1193] Specific operation: The server uses a search engine API to retrieve information about dishes related to the translated text from the Internet.

[1194] Step 6:

[1195] The server generates an image and description of the dish from the information it obtains.

[1196] Input: Image and description of the food obtained

[1197] Output: Generated food images and abbreviated descriptions

[1198] Specific operation: The server uses AI technology to organize the acquired information and generate visually easy-to-understand images and concise explanations.

[1199] Step 7:

[1200] The server encrypts the generated information and transmits it to the user terminal.

[1201] Input: Generated food image and description

[1202] Output: Encrypted dish image and description

[1203] Specific operation: The server encrypts the information as a security measure and sends it to the user's terminal via the Internet.

[1204] Step 8:

[1205] The device analyzes the user's facial expressions and tone of voice to recognize emotions.

[1206] Input: User's facial expression and voice data

[1207] Output: Recognized emotion data

[1208] Specific operation: The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them through an emotion engine.

[1209] Step 9:

[1210] The server makes personalized dish suggestions based on emotion recognition results.

[1211] Input: Recognized emotion data, user behavior history

[1212] Output: Personalized food suggestions

[1213] Specific operation: The server generates individually optimized dish suggestions based on the emotion data and the user's previous history and sends them to the user's device.

[1214] Step 10:

[1215] The terminal displays the transmitted information to the user.

[1216] Input: Encrypted food images, abbreviated descriptions, and personalized food suggestions

[1217] Output: Food images, descriptions, and suggestions displayed on the user's device

[1218] Specific operation: The device decodes the received information and displays it visually to the user.

[1219] The above is a specific processing flow for carrying out the present invention.

[1220] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1221] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1222] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1223] [Fourth embodiment]

[1224] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1225] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1226] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1227] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1228] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1230] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1231] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1232] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1233] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1234] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1235] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1236] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1237] The following system and program are provided as embodiments of the present invention.

[1238] The system is broadly composed of the user's terminal, a server on the cloud, and a database connected to the Internet.

[1239] User terminal

[1240] The user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user device is to capture images of the menu and send them to a server on the cloud.

[1241] server

[1242] The cloud server processes the received image data using advanced AI technology. The specific processing performed by the server is shown below.

[1243] 1. Image and Character Recognition (OCR):

[1244] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, obtaining the names and descriptions of the dishes on the menu as digital text.

[1245] 2. Text Translation:

[1246] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, the Japanese character for "tempura" is translated into English.

[1247] 3. Internet Search:

[1248] The server uses the translated dish name to search image and information databases on the Internet to retrieve images and descriptions of the corresponding dish, such as images related to "Tempura," along with information about the dish's characteristics and summary.

[1249] 4. Summary sentence generation:

[1250] AI technology is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, flavor characteristics, etc. For example, a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[1251] 5. Encryption and Transmission of Information:

[1252] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[1253] Display on user device

[1254] The user device displays the information received from the server through a dedicated application. This allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[1255] Specific examples

[1256] Specific examples are shown below.

[1257] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[1258] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[1259] The processing flow will be explained below.

[1260] Step 1:

[1261] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[1262] Step 2:

[1263] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[1264] Step 3:

[1265] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[1266] Step 4:

[1267] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[1268] Step 5:

[1269] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[1270] Step 6:

[1271] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1272] Step 7:

[1273] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[1274] Step 8:

[1275] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[1276] Step 9:

[1277] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[1278] Step 10:

[1279] The user decides which dish to order based on the displayed information. For example, after seeing an image of tempura and a summary, the user decides whether to order that dish.

[1280] Through this series of steps, the user can easily and intuitively understand the menu contents and select dishes.

[1281] Example 1

[1282] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1283] When travelers or users unfamiliar with foreign languages ​​visit restaurants overseas, they often encounter problems due to language barriers and cultural differences when trying to understand the menu and select the appropriate dish. These issues reduce convenience for travelers and detract from their dining experience.

[1284] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1285] In this invention, the server includes optical character recognition means for extracting text information from images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, and means for generating summaries of the searched food images and descriptions using a generation AI model, thereby enabling users to quickly and accurately understand menus in foreign languages ​​and easily select dishes.

[1286] A "camera" is a device for taking pictures or videos.

[1287] A "cloud server" is a remotely located computer system that can be accessed via the Internet for storing data and performing computations.

[1288] "Optical character recognition" refers to technology or devices that extract character information from an image as digital text.

[1289] A "translation tool" is a technique or device for converting text expressed in a particular language into another language.

[1290] A "generative AI model" is an artificial intelligence algorithm or set of systems that learns from large amounts of data and performs tasks such as text generation, classification, and translation.

[1291] A "means for generating a summary" is a technology or device that extracts important points from the provided information and summarizes them in a concise sentence.

[1292] An "encryption method" is a technology or device that converts data based on a specific algorithm to prevent unauthorized decryption by third parties.

[1293] A "user terminal" is a computer system operated by a user, and includes mobile phones, smartphones, tablets, and the like.

[1294] "Internet search tools" are techniques or devices that retrieve information from online databases or sources based on specific keywords.

[1295] The following system and program are provided as embodiments of the present invention: The system is broadly composed of a user terminal, a server on the cloud, and a database connected to the Internet.

[1296] User terminal

[1297] A user's device, such as a smartphone or tablet, is equipped with a camera, an internet connection, and a dedicated application. The main role of the user's device is to capture images of the menu and send them to a server on the cloud. The user takes a picture of the menu using the camera function and sends it to the server via the dedicated application.

[1298] server

[1299] The cloud server processes the received image data using advanced AI technology. Below are some examples of specific hardware and software used on the server.

[1300] 1. Image and Character Recognition (OCR):

[1301] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition (OCR) technology, using libraries such as TensorFlow and OpenCV. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[1302] 2. Text Translation:

[1303] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, the Japanese character for "tempura" (Tempura) is translated into English.

[1304] 3. Internet Search:

[1305] The server then uses the translated dish name to search image and information databases on the Internet, using the Bing Search API and other services. For example, it retrieves images related to "Tempura" and information about the dish, including its general description and characteristics.

[1306] 4. Summary sentence generation:

[1307] A generative AI model (e.g., GPT-4) is used to generate a concise summary from the acquired dish information. This summary includes the main ingredients, cooking method, and flavor characteristics. A prompt sentence is used to input the model, and a summary such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture" is generated.

[1308] Example prompt sentence:

[1309] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[1310] 5. Encryption and Transmission of Information:

[1311] The generated image and summary are encrypted using the SSL / TLS protocol and sent to the user's device. The server then encrypts the data using an encryption library and sends it via an HTTP POST request.

[1312] Display on user device

[1313] The user device decrypts the encrypted data received from the server and displays it through a dedicated application. This process allows users to simply point their camera at the menu to see images of dishes and brief descriptions. For example, if a user points their camera at a menu item called "tempura," the description and visual information will be instantly displayed, allowing them to decide which dish to order.

[1314] Specific examples

[1315] Specific examples are shown below.

[1316] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to a server. The server extracts the text information "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and an image of tempura and a description are displayed on the user's application.

[1317] This system allows users to easily select dishes, regardless of language barriers. This series of processes and system configuration allows users to intuitively understand local menus, greatly improving convenience for tourists.

[1318] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1319] Step 1:

[1320] Image capture and transmission

[1321] Users take a photo of a restaurant menu with their smartphone or tablet, and the device acquires the captured image and sends it to a cloud server via a dedicated application using an HTTP POST request.

[1322] Input: A user-taken image of the menu

[1323] Output: Image data sent to a cloud server

[1324] Step 2:

[1325] Image and character recognition (OCR)

[1326] The server processes the received image data and extracts the text information from the image using optical character recognition (OCR) technology. This process uses libraries such as TensorFlow and OpenCV. The server analyzes each pixel of the image and identifies areas that can be recognized as text.

[1327] Input: Image data sent to a cloud server

[1328] Output: Extracted text information (e.g. "tempura")

[1329] Step 3:

[1330] Text Translation

[1331] The server inputs the extracted text information into the Google Translate API and translates it into the specified language. It creates an API request, sends the extracted text information, and receives the translation result. For example, the Japanese text information "tempura" is translated into English "Tempura."

[1332] Input: Extracted text information (e.g., "tempura")

[1333] Output: Translated text (e.g. "Tempura")

[1334] Step 4:

[1335] Internet search

[1336] The server uses the translated text as a key to search for related information on the Internet using the Bing Search API. Specifically, it sends an API request and receives search results, which retrieves information such as images and descriptions related to "Tempura."

[1337] Input: Translated text (e.g. "Tempura")

[1338] Output: Images and descriptions of the dishes retrieved as search results

[1339] Step 5:

[1340] Summary sentence generation

[1341] Based on the acquired information, the server uses a generative AI model (e.g., GPT-4) to generate a concise summary. The server creates a prompt sentence and inputs it into the model to generate a summary. For example, the summary generated is, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1342] Input: Images and descriptions of dishes obtained as search results

[1343] Output: Generated summary

[1344] Example prompt sentence:

[1345] "Generate a concise summary of the given dish details. Examples include key ingredients, cooking method, and flavor characteristics. The dish details are: [insert the details you obtained here]"

[1346] Step 6:

[1347] Encryption and transmission of information

[1348] The server encrypts the generated image and summary using the SSL / TLS protocol and sends it to the user's terminal. The server also encrypts the data using an encryption library and then sends it via an HTTP POST request.

[1349] Input: Generated summary and food image

[1350] Output: Sent to the user's device as encrypted data

[1351] Step 7:

[1352] Display on user device

[1353] The user device decrypts the encrypted data received from the server and displays it using a dedicated application. Specifically, the device performs decryption processing based on the SSL / TLS protocol and displays an image of tempura and a summary to the user.

[1354] Input: Encrypted data sent by the server

[1355] Output: A picture of the food and a summary of it displayed on the user's device

[1356] By following the above steps, the user can easily understand and select from the menu at a restaurant overseas.

[1357] (Application example 1)

[1358] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1359] Modern restaurants face the challenge of making it difficult for visitors to intuitively understand the dishes and their details on the menu in multiple languages. This is particularly true for tourists who cannot understand foreign languages, as they are unable to immediately understand the dish descriptions, allergen information, prices, and ratings. This can lead to inconvenience when ordering and a decrease in customer satisfaction. Therefore, there is a need for a system that provides multilingual and visually easy-to-understand information about dishes.

[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1361] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a cloud server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching the Internet for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for summarizing the generated food images and descriptions as detailed information about dishes in the restaurant, means for transmitting the summarized information to a user terminal, means for displaying the transmitted information on the user terminal, and means for allowing the user to check related image data and rating information on the terminal. This allows visitors to instantly understand detailed information and ratings about dishes in multiple languages, eliminating inconvenience when ordering and improving customer satisfaction.

[1362] "Means for capturing images of menus with a camera" is a function that allows a user to take a photo of a restaurant menu using a mobile device such as a smartphone or tablet and obtain the image data.

[1363] The "means for transmitting the image to a server on the cloud" is a function for uploading image data taken by a mobile terminal to a remote cloud server via the Internet.

[1364] "Optical character recognition means" refers to a technology for extracting character information from received image data, and in particular, uses OCR (optical character recognition) technology.

[1365] The "translation means" is a function for converting extracted text information into a different language as needed.

[1366] The "means for searching on the Internet" is a function for searching an online database based on the translated text information and obtaining images and descriptions of related dishes.

[1367] The "means for generating" is a function for processing the searched food image and description into the format required to provide it to the user.

[1368] The "means for summarizing detailed dish information" is a function that extracts and displays a concise explanation and essential points from the generated image and description of the dish so that the user can easily understand it.

[1369] The "means for transmitting to the user terminal" is a function for transmitting the generated image of the dish and a summarized description to the user's mobile terminal.

[1370] The "means for displaying on the user terminal" is a function for visually displaying the transmitted image of the dish and the summarized description on the user's mobile terminal.

[1371] The "means for checking related image data and evaluation information" is a function that allows the user to refer to image data related to the displayed dish information and evaluation information from other users.

[1372] As an embodiment of the present invention, the following system and program are provided. This system is composed of a user's mobile terminal, a cloud server, and a database connected to the Internet. Specifically, each component functions as follows:

[1373] User terminal

[1374] A user's mobile device, such as a smartphone or tablet, is equipped with a camera, Internet connectivity, and a dedicated application. The main role of the user device is to capture images of menus and send them to a server on the cloud. The user takes a photo of a restaurant menu using the camera and uploads the image to the server via the application.

[1375] server

[1376] The server processes the received image data using advanced AI technology. The server performs the processing using the following specific hardware and software:

[1377] Hardware: High-performance processor (e.g., Intel Xeon), large memory capacity

[1378] Software: Tesseract (OCR engine), Google Translate API, Python

[1379] Image Recognition and Translation

[1380] The server first uses OCR technology to extract text information from the menu image sent from the device. The extracted text information is then translated into the specified language using the Google Translate API. For example, the Japanese character for "tempura" is translated into "Tempura."

[1381] Data Search and Generation

[1382] The translated text is then used to search an online database to retrieve images and descriptions of the corresponding dishes. For example, images related to "Tempura" are retrieved, along with information about the dish's characteristics. This information is then used to summarize the details of the restaurant's dishes. The summary includes information about the main ingredients, cooking methods, and flavor characteristics.

[1383] Sending and Displaying Information

[1384] The summarized images and descriptions are encrypted for security reasons and sent to the user's device. The user's device displays the information received from the server through a dedicated application. This allows users to view images of dishes and brief descriptions simply by pointing their camera at the menu. They can also view related image data and review information.

[1385] Specific examples

[1386] Below is a specific example. When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server uses OCR technology to extract the text information for "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve the image of tempura and related information, summarizing it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." The encrypted information is sent to the user's device, and the image of tempura, a description, and user ratings are displayed on the user's application.

[1387] Example prompts for generative AI models

[1388] "Please translate the Japanese menu items in this image into English."

[1389] "Get detailed information about ramen, including ingredients, cooking method, and flavor characteristics."

[1390] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1391] Step 1:

[1392] A user takes a picture of a restaurant menu using the camera on their mobile device, and the image is sent to a cloud server via the application.

[1393] Input: A menu image taken by the user.

[1394] Output: Menu images sent to a server on the cloud.

[1395] Step 2:

[1396] The server processes the received image data using an OCR engine (Tesseract) to extract the text information from the image, and as a result, the menu items are converted into text data.

[1397] Input: Menu image sent to a server on the cloud.

[1398] Output: Character information (text data) extracted by the OCR engine.

[1399] Step 3:

[1400] The extracted text information is translated into the specified language using the Google Translate API.

[1401] Input: Extracted character information (text data).

[1402] Output: Translated text information (translated text data).

[1403] Step 4:

[1404] The server searches an online database based on the translated text information and retrieves images and descriptions of the corresponding dishes.

[1405] Input: Translated text.

[1406] Output: Images and descriptions of dishes retrieved through internet searches.

[1407] Step 5:

[1408] The server generates a summary from the acquired cooking information, including the main ingredients, cooking method, flavor characteristics, etc.

[1409] Input: Food images and descriptions obtained from an internet search.

[1410] Output: A condensed description of the dish.

[1411] Step 6:

[1412] The generated image and summarized description are encrypted for security reasons and sent to the user's terminal.

[1413] Input: Abridged dish description and image.

[1414] Output: Encrypted information (abridged description and image of the dish).

[1415] Step 7:

[1416] The user terminal decodes the received information using a dedicated application and displays an image of the dish, a summary of its description, related image data, and rating information.

[1417] Input: Encrypted information.

[1418] Output: Decoded and displayed dish image on the user's device, along with a summary description, associated image data, and rating information.

[1419] Specific operation example

[1420] When a user points their camera at the Japanese menu item "tempura" and takes a photo, the image is sent to a cloud server. The server uses OCR technology to extract the text information for "tempura" and translates it into "Tempura" using the Google Translate API. It then searches the internet for images and descriptions related to tempura and summarizes it as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." This information is then encrypted and sent to the user's device, where it is displayed on the application.

[1421] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1422] This invention combines a system that captures menu images, extracts and translates text information from the images, and searches the internet for related food images and descriptions, with an emotion engine that recognizes the user's emotions. This system makes it possible to provide information according to the user's emotions, providing a more personalized experience.

[1423] User terminal

[1424] The user's device is a smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The main function of the user's device is to capture menu images and send them to a cloud server. In addition, the device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[1425] server

[1426] The cloud server has the main function of processing the received image data using advanced AI technology. The specific processing details are shown below.

[1427] 1. Image and Character Recognition (OCR):

[1428] The server receives the menu image sent from the device and first extracts the text information from the image using optical character recognition technology. Through this process, the names and descriptions of the dishes on the menu are obtained as digital text.

[1429] 2. Text Translation:

[1430] The extracted text information is input into a translation engine built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[1431] 3. Internet Search:

[1432] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database, such as images and descriptions related to "Tempura."

[1433] 4. Summary sentence generation:

[1434] From the acquired information, AI technology is used to generate a concise summary, such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1435] 5. Encryption and Transmission of Information:

[1436] The generated image and summary are encrypted for security reasons and sent to the user's terminal.

[1437] Emotion Engine

[1438] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns. This makes it possible to provide information tailored to the user's interests and preferences.

[1439] Display on user device

[1440] The user device displays the information received from the server through the application. The display content is adjusted according to the user's emotions by the emotion engine. This allows the user to obtain the most appropriate information and helps them make food selections. For example, if a user points the camera at a menu item called "tempura," a description and visual information will be displayed immediately, along with related recommendations based on the user's emotions.

[1441] Specific examples

[1442] Specific examples are shown below.

[1443] When a user points their camera at the Japanese menu item "tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text "tempura" from the received image and translates it into "Tempura." The server then uses the Internet to retrieve information related to the tempura image and generates a summary. The summary, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture," is created, encrypted, and sent to the user's device. Furthermore, the emotion engine recognizes interest from the user's facial expression and displays other recommended dishes. This allows the user to view information related to tempura visually and through text, making the best choice.

[1444] Through this system, users can intuitively understand the contents of the dishes and receive personalized information, making it easier to select dishes.

[1445] The processing flow will be explained below.

[1446] Step 1:

[1447] The user points the device's camera at a menu and takes a picture. For example, the user points the camera at the Japanese character string "tempura" on a restaurant menu.

[1448] Step 2:

[1449] The device captures an image of the menu with the camera and temporarily stores it in local storage. The captured image is saved in high resolution, providing optimal conditions for character recognition.

[1450] Step 3:

[1451] The device sends the captured image data to a server in the cloud, using an internet connection and encrypting the data using the appropriate protocol.

[1452] Step 4:

[1453] The server analyzes the received image data and first applies an optical character recognition (OCR) algorithm to extract textual information from the image. For example, it identifies the word "tempura" in the image and converts it into digital text.

[1454] Step 5:

[1455] The server inputs the extracted text information into a translation engine and translates it into the specified language. For example, "tempura" is translated into English as "Tempura."

[1456] Step 6:

[1457] The server performs an internet search based on the translated text. Using APIs and databases, the server retrieves images and descriptions of dishes related to "Tempura." For example, it retrieves an image of tempura and the description, "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture."

[1458] Step 7:

[1459] The server generates a summary from the information it has acquired, making it easy for users to understand. It uses AI technology to integrate data from multiple sources and create a concise and accurate summary.

[1460] Step 8:

[1461] The server then sends the generated image and summary to the device, again using encryption technology to ensure user privacy and data security.

[1462] Step 9:

[1463] The device displays the image and summary received from the server through the application. By checking the screen of the device, the user can obtain specific information related to the "Tempura" menu item.

[1464] Step 10:

[1465] The emotion engine analyzes the user's facial expressions and tone of voice to recognize the user's emotional state. For example, if the user's facial expression indicates joy or interest, that information is recorded.

[1466] Step 11:

[1467] The emotion engine uses the user's emotion information to provide additional information tailored to the user's interests and preferences. For example, if a user is interested in tempura, it will suggest other fried dishes and Japanese food options.

[1468] Step 12:

[1469] The device displays information provided by the emotion engine, and the user can review it and select a dish. For example, in addition to information about tempura, images and descriptions of katsudon and karaage are also displayed.

[1470] This series of steps allows users to intuitively understand the menu contents and select the best dish based on personalized information that responds to their emotions.

[1471] Example 2

[1472] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1473] Conventional menu translation systems have the problem of not providing information that takes into account the user's emotions and interests, and can only display uniform information. As a result, it is difficult for users to quickly obtain the information they really want, and satisfaction cannot be increased. In addition, there are limitations to the accuracy of translation and the reliability of the information, so there is a demand for more accurate information provision.

[1474] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1475] In this invention, the server includes means for capturing menu images with a camera, means for transmitting the images to a remote server, optical character recognition means for extracting text information from the received images, translation means for translating the extracted text information, means for searching a network for food images and descriptions based on the translation results, means for generating the searched food images and descriptions, means for transmitting the food images and descriptions to a user terminal, means for displaying the transmitted information on the user terminal, means for recognizing emotions by analyzing a user's facial expressions and tone of voice, and means for adjusting information based on the emotion recognition, thereby enabling the provision of personalized information that takes into account the user's emotions and interests.

[1476] The "means for capturing an image of a menu using a camera" is a function for taking a digital image of a physical menu using a camera mounted on a user terminal.

[1477] The "means for transmitting the image to a remote server" is a function for transmitting captured image data from a user terminal to a server on the cloud using an internet connection.

[1478] "Optical character recognition means for extracting character information from received images" refers to technology for recognizing character information from transmitted image data and converting it into digital text.

[1479] "Translation means for translating extracted text information" refers to technology for converting text information obtained by OCR into another language.

[1480] "Means for searching the network for food images and descriptions based on the translation results" is a technology that uses the translated text information as input to obtain related images and descriptions using the Internet.

[1481] "Means for generating images and descriptions of searched dishes" is a function that creates images of dishes and related descriptions based on information obtained from the network.

[1482] The "means for transmitting the image and description of the dish to the user terminal" is a technique for transferring the generated image and description of the dish from a remote server to the user terminal.

[1483] The "means for displaying the transmitted information on the user terminal" is a function for visually displaying the image and description of the dish received on the user terminal.

[1484] "Means for recognizing emotions by analyzing a user's facial expressions and tone of voice" refers to technology that uses a camera and microphone to analyze the user's facial expressions and voice to determine the user's emotional state.

[1485] The "means for adjusting information based on emotion recognition" is a technology that dynamically changes the content of the information to be displayed according to the user's recognized emotion.

[1486] The system of the present invention allows users to capture images of menus, extract and translate textual information from the images, search the internet for related food images and descriptions, and even recognize the user's emotions to provide personalized information, providing users with more intuitive and individually tailored information.

[1487] Hardware and software used

[1488] User's device: A smartphone or tablet equipped with a camera, internet connectivity, and a dedicated application. The user captures an image of the menu with the device's camera and sends it to a cloud server. The device is equipped with an emotion engine that can analyze the user's facial expressions and tone of voice.

[1489] Server: The server on the cloud processes the received image data using advanced AI technology and has the following functions:

[1490] 1. OCR (Optical Character Recognition): Uses OCR technology, such as Google Cloud Vision API, to extract text information from received images.

[1491] 2. Translation: To translate the extracted text information into multiple languages, use, for example, the Google Translate API.

[1492] 3. Internet search: Using the translated text, you can search for food images and descriptions on the Internet, for example, using the Bing Search API.

[1493] 4. Summary generation: A generative AI model such as GPT-3 is used to generate a summary from the acquired information.

[1494] 5. Information encryption: The generated information is encrypted and sent to the user terminal using the Advanced Encryption Standard (AES).

[1495] Specific examples

[1496] Consider a scenario where a user points their camera at a Japanese menu item that says "tempura" at a Japanese restaurant. The user's device sends the image to a server in the cloud. The server analyzes the image and extracts the text "tempura" using OCR technology. The server then translates "tempura" to "Tempura" using the Google Translate API. The server then retrieves images and descriptions related to "Tempura" from the Internet using the Bing Search API. Based on the retrieved information, the server uses GPT-3 to generate a summary: "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture." This information is then encrypted using AES and sent to the user's device. The user's device decrypts the received information and displays it to the user through an application. The emotion engine recognizes the user's facial expression as an indication of interest and simultaneously displays other related recommended dishes.

[1497] Prompt Sentence Examples

[1498] "Let the user capture an image of a restaurant menu, extract and translate text from the image, and provide relevant dish images and descriptions. Also, recognize the user's emotions and adjust to provide a personalized information experience."

[1499] This system allows users to intuitively understand the contents of the dish and receive information that is tailored to their individual needs.

[1500] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1501] Step 1:

[1502] The user captures an image of the menu

[1503] The user takes a picture of a restaurant menu using the camera on their smartphone or tablet. This image becomes the initial input for the system. Specifically, the user launches the device's camera application and captures an image of the menu. The image data is then saved via a dedicated application, and the system proceeds to the next step.

[1504] Step 2:

[1505] The device sends the image to the server

[1506] After image data is captured by the camera application, a dedicated application sends the image to a server on the cloud. The input is the captured image data, and the output is the image data sent to the server. Specifically, the application on the device sends the image data to a specified endpoint on the server via an Internet connection.

[1507] Step 3:

[1508] The server performs image recognition and OCR processing

[1509] The server receives image data sent from the device. The input is the sent image data, and the output is the extracted text information. The image recognition module in the server analyzes the image and converts the text information in the image into digital text using OCR (optical character recognition) technology. Specifically, the server uses the "Google Cloud Vision API" to extract the text information in the image.

[1510] Step 4:

[1511] The server translates the extracted text.

[1512] The server takes the character information extracted by the OCR process as input and sends this information to a translation engine. The output is the translated text. The server uses the "Google Translate API" to translate the text into the specified language. Specifically, the server translates the extracted Japanese character information "Tempura" into "Tempura."

[1513] Step 5:

[1514] The server performs an internet search based on the translation results

[1515] Using the translated text as input, the server performs an internet search to gather images and descriptions of related dishes. The output is the retrieved images and descriptions. The server uses the Bing Search API to search for information about "Tempura." Specifically, the server gathers images and descriptions related to "Tempura" from multiple websites.

[1516] Step 6:

[1517] The server generates a summary

[1518] Using the acquired information as input, the server uses a generative AI model to generate a concise summary. The output is summarized text. The server uses GPT-3 to summarize the information, generating a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep-fried. It is known for its crispy texture." Specifically, the server performs a process to summarize the collected descriptions using the generative AI model.

[1519] Step 7:

[1520] The server encrypts the information and sends it to the user's device.

[1521] The server takes the generated image and summary as input and encrypts this information. The output is encrypted data. The server encrypts the data using AES (Advanced Encryption Standard) and sends it to the user's device. Specifically, the server uses AES encryption to keep the data secure and sends it to the device via the Internet.

[1522] Step 8:

[1523] Display information received by the device

[1524] The device receives encrypted information from the server as input, decrypts it, and visually displays it to the user. The output is an image of the displayed dish and a summary. A dedicated application on the device decrypts the encrypted data and displays an image of the dish along with a summary such as "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture." Specifically, the device uses the received key data to decrypt the data and displays it through the application interface.

[1525] Step 9:

[1526] The device recognizes the user's emotions and adjusts the information accordingly.

[1527] Once again, the emotion engine uses the user's facial expression and tone of voice as input to recognize the emotion and adjust the information. The output is additional information based on the user's emotion. The emotion engine uses the "Emotion API" to analyze the user's emotion, and if the user shows interest, it displays additional information about related dishes or recommendations. Specifically, the device uses the camera and microphone to analyze the user's facial expression and voice in real time, dynamically changing the information displayed.

[1528] (Application example 2)

[1529] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1530] In today's increasingly internationalized world, understanding the menu contents of restaurants abroad is an important factor for travelers. However, language barriers and lack of knowledge about cuisine can make it difficult for travelers to accurately understand the menu. Furthermore, conventional menu translation systems lack the ability to provide personalized information that takes into account the user's interests and preferences, making it difficult for them to choose a meal. To resolve this situation, it is necessary to recognize the user's emotions and provide information based on them.

[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1532] In this invention, the server includes optical character recognition means for extracting text information from received images, translation means for translating the extracted text information, and means for searching the Internet for food images and descriptions based on the translation results, thereby making it easier to understand menu contents beyond language barriers and enabling personalized food suggestions.

[1533] A "camera" is a device that captures light and records images or videos.

[1534] A "cloud server" is a remote computing resource accessed via the Internet, a computer system that processes and stores data.

[1535] "Optical character recognition means" is a technology that recognizes character information from image data and extracts it as text data.

[1536] A "translation tool" is a technique for converting text in one language into text in another language.

[1537] "Means of searching the Internet" refers to the technology of obtaining information that is publicly available online through search engines or APIs.

[1538] The "means for generating food images and descriptions" is a technology that constructs visual and text descriptions based on information obtained from search results.

[1539] "Means for transmitting to the user terminal" refers to the technology for transferring data from a server on the cloud to the user's device.

[1540] "Emotion recognition means" is a technology that analyzes and identifies a user's emotional state from facial expressions, tone of voice, etc.

[1541] "Means for making personalized dish suggestions" refers to technology that makes individually optimized dish suggestions based on the user's emotional data and preferences.

[1542] "Means for displaying on the user terminal" refers to a technique for visually displaying the provided information on the user's device.

[1543] This invention combines a system that captures menu images with a camera, extracts and translates text information from the images, and searches the internet for related food images and descriptions to display them, with an emotion engine that recognizes the user's emotions. Specific embodiments of this system are described below.

[1544] System Configuration

[1545] User device:

[1546] The user's device is typically a smartphone or tablet. The device is equipped with a camera and is connected to the Internet. A dedicated application is installed on the device, which has the function of capturing images of menus and sending them to a cloud server. Additionally, the device is equipped with an emotion engine, which can analyze the user's facial expressions and tone of voice.

[1547] Cloud Server:

[1548] The cloud server has high-performance computing resources and performs the following processes:

[1549] 1. Image and Character Recognition (OCR):

[1550] The server receives the menu image sent from the user device and uses optical character recognition software (e.g., pytesseract) to extract the text information within the image. This process results in the names and descriptions of the dishes on the menu being captured as digital text.

[1551] 2. Text Translation:

[1552] The extracted text information is input into a translation engine (e.g., Google Translate API) built into the server and translated into the specified language. For example, "tempura" is translated into English as "Tempura."

[1553] 3. Internet Search:

[1554] The server then performs an internet search based on the translated text to retrieve images and descriptions of dishes from an internet database (e.g., Google search), such as images and descriptions related to "Tempura."

[1555] 4. Information generation:

[1556] Using the acquired information, AI technology generates a concise description, such as "Tempura is a Japanese dish of seafood or vegetables that has been battered and deep-fried. It is known for its crispy texture."

[1557] 5. Encryption and Transmission of Information:

[1558] The generated image and description are encrypted as a security measure and sent to the user's terminal.

[1559] Emotion Engine:

[1560] The emotion engine is integrated into the user's device and recognizes emotions by analyzing the user's facial expressions and tone of voice. It also infers emotions based on the user's input history and behavioral patterns, enabling it to provide information tailored to the user's interests and preferences.

[1561] System Operation

[1562] When a user captures a menu image with their camera, the image is sent to a cloud server. The server extracts text information from the image and translates it into the specified language. Based on the translation results, the server searches the internet for related food images and descriptions, and generates a summary of the information.

[1563] The generated information is tailored based on the user's emotional data. For example, if the emotion engine recognizes that the user is interested, it will provide related information and recommended dishes. This allows users to get personalized information that will help them make better food choices.

[1564] Examples:

[1565] When a user points their camera at the Japanese menu item "Tempura" at a Japanese restaurant, the device captures the image and sends it to the server. The server extracts the text information "Tempura" from the received image and translates it to "Tempura." The server then uses the Internet to obtain information related to the image of tempura and generates a description. The description it creates is "Tempura is a Japanese dish of seafood or vegetables that have been battered and deep fried. It is known for its crispy texture," and the information is encrypted and sent to the user's device. The emotion engine recognizes interest from the user's facial expression and displays other recommended dishes.

[1566] Example prompt sentence:

[1567] "Generate a Python program that captures an image of a dish called tempura on a menu at a Japanese restaurant, extracts and translates the text from the image, analyzes the user's emotions from their facial expressions, and provides the user with the most appropriate food information based on their emotions."

[1568] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1569] Step 1:

[1570] The device captures an image of the menu with its camera.

[1571] Input: Restaurant menu image

[1572] Output: Captured image data

[1573] Specific action: The user takes a photo of the menu using the camera on their smartphone or tablet.

[1574] Step 2:

[1575] The device sends the captured image to a server on the cloud.

[1576] Input: Captured image data

[1577] Output: Image data sent to the server

[1578] Specific operation: The device uses an internet connection to upload image data to a server in the cloud.

[1579] Step 3:

[1580] The server extracts text information from the received image using optical character recognition (OCR).

[1581] Input: Image data sent to the server

[1582] Output: Extracted text information (text data)

[1583] Specific operation: The server uses OCR software (e.g., pytesseract) to analyze the characters in the image and generate text data.

[1584] Step 4:

[1585] The server translates the extracted text information into the specified language using a translation engine.

[1586] Input: Extracted text information (text data)

[1587] Output: Translated text information (text data)

[1588] What happens: The server uses a translation engine, such as the Google Translate API, to translate the text into the specified language.

[1589] Step 5:

[1590] The server searches the internet for images and descriptions of related dishes based on the translation results.

[1591] Input: Translated text information (text data)

[1592] Output: Images and descriptions of the dishes obtained

[1593] Specific operation: The server uses a search engine API to retrieve information about dishes related to the translated text from the Internet.

[1594] Step 6:

[1595] The server generates an image and description of the dish from the information it obtains.

[1596] Input: Image and description of the food obtained

[1597] Output: Generated food images and abbreviated descriptions

[1598] Specific operation: The server uses AI technology to organize the acquired information and generate visually easy-to-understand images and concise explanations.

[1599] Step 7:

[1600] The server encrypts the generated information and transmits it to the user terminal.

[1601] Input: Generated food image and description

[1602] Output: Encrypted dish image and description

[1603] Specific operation: The server encrypts the information as a security measure and sends it to the user's terminal via the Internet.

[1604] Step 8:

[1605] The device analyzes the user's facial expressions and tone of voice to recognize emotions.

[1606] Input: User's facial expression and voice data

[1607] Output: Recognized emotion data

[1608] Specific operation: The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them through an emotion engine.

[1609] Step 9:

[1610] The server makes personalized dish suggestions based on emotion recognition results.

[1611] Input: Recognized emotion data, user behavior history

[1612] Output: Personalized food suggestions

[1613] Specific operation: The server generates individually optimized dish suggestions based on the emotion data and the user's previous history and sends them to the user's device.

[1614] Step 10:

[1615] The terminal displays the transmitted information to the user.

[1616] Input: Encrypted food images, abbreviated descriptions, and personalized food suggestions

[1617] Output: Food images, descriptions, and suggestions displayed on the user's device

[1618] Specific operation: The device decodes the received information and displays it visually to the user.

[1619] The above is a specific processing flow for carrying out the present invention.

[1620] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1621] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1622] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1623] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1624] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1625] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1626] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1627] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1628] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1629] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1630] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1631] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1632] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1633] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1634] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1635] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1636] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1637] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1638] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1639] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1640] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1641] The following is further disclosed regarding the above embodiment.

[1642] (Claim 1)

[1643] a means for capturing an image of the menu with a camera;

[1644] means for transmitting the image to a server on a cloud;

[1645] optical character recognition means for extracting textual information from the received image;

[1646] a translation means for translating the extracted character information;

[1647] A means to search the internet for food images and descriptions based on the translation results;

[1648] means for generating images and descriptions of the searched dishes;

[1649] means for transmitting an image and description of the dish to a user terminal;

[1650] means for displaying the transmitted information on a user terminal;

[1651] A system including:

[1652] (Claim 2)

[1653] 10. The system of claim 1, further comprising means for summarizing the dish description.

[1654] (Claim 3)

[1655] 10. The system of claim 1, further comprising means for using encryption in transmitting the dish images and descriptions.

[1656] "Example 1"

[1657] (Claim 1)

[1658] a means for capturing an image of the menu with a camera;

[1659] means for transmitting the image to a server on a cloud;

[1660] optical character recognition means for extracting textual information from the received image;

[1661] a translation means for translating the extracted character information;

[1662] A means to search the internet for food images and descriptions based on the translation results;

[1663] A means to generate a summary using an AI model that generates images and descriptions of the searched dishes, and

[1664] means for encrypting the image and summary of the dish and transmitting them to a user terminal;

[1665] means for displaying the transmitted information on a user terminal;

[1666] A system including:

[1667] (Claim 2)

[1668] 10. The system of claim 1, further comprising: means for generating a summary of the dish using a generative AI model.

[1669] (Claim 3)

[1670] 10. The system of claim 1, further comprising means for using encryption in transmitting the image and summary of the dish.

[1671] "Application Example 1"

[1672] (Claim 1)

[1673] a means for capturing an image of the menu with a camera;

[1674] means for transmitting the image to a server on a cloud;

[1675] optical character recognition means for extracting textual information from the received image;

[1676] a translation means for translating the extracted character information;

[1677] A means to search the internet for food images and descriptions based on the translation results;

[1678] means for generating images and descriptions of the searched dishes;

[1679] a means for summarizing the generated food images and descriptions as detailed food information for the restaurant;

[1680] means for transmitting the summarized information to a user terminal;

[1681] means for displaying the transmitted information on a user terminal;

[1682] means for allowing a user to check related image data and evaluation information on the terminal;

[1683] A system including:

[1684] (Claim 2)

[1685] 10. The system of claim 1, further comprising means for summarizing the dish description.

[1686] (Claim 3)

[1687] 10. The system of claim 1, further comprising means for using encryption in transmitting the dish images and descriptions.

[1688] "Example 2: Combining Emotion Engines"

[1689] (Claim 1)

[1690] a means for capturing an image of the menu with a camera;

[1691] means for transmitting said image to a remote server;

[1692] optical character recognition means for extracting textual information from the received image;

[1693] a translation means for translating the extracted character information;

[1694] A means for searching the network for food images and descriptions based on the translation results;

[1695] means for generating images and descriptions of the searched dishes;

[1696] means for transmitting an image and description of the dish to a user terminal;

[1697] means for displaying the transmitted information on a user terminal;

[1698] A means of recognizing emotions by analyzing the user's facial expressions and tone of voice;

[1699] A means of tailoring information based on emotion recognition

[1700] A system including:

[1701] (Claim 2)

[1702] 10. The system of claim 1, further comprising means for summarizing the dish description.

[1703] (Claim 3)

[1704] 10. The system of claim 1, further comprising means for using encryption in transmitting the dish images and descriptions.

[1705] "Application example 2 when combining emotion engines"

[1706] (Claim 1)

[1707] a means for capturing an image of the menu with a camera;

[1708] means for transmitting the image to a server on a cloud;

[1709] optical character recognition means for extracting textual information from the received image;

[1710] a translation means for translating the extracted character information;

[1711] A means to search the internet for food images and descriptions based on the translation results;

[1712] means for generating images and descriptions of the searched dishes;

[1713] means for transmitting an image and description of the dish to a user terminal;

[1714] emotion recognition means for recognizing emotions by analyzing a user's facial expression and tone of voice;

[1715] a means for making personalized recipe suggestions based on the emotion data obtained from the emotion recognition means;

[1716] means for displaying the transmitted information on a user terminal;

[1717] A system including:

[1718] (Claim 2)

[1719] 10. The system of claim 1, further comprising means for summarizing the dish description.

[1720] (Claim 3)

[1721] 10. The system of claim 1, further comprising means for using encryption in transmitting the dish images and descriptions. [Explanation of symbols]

[1722] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for capturing an image of the menu with a camera; means for transmitting the image to a server on a cloud; optical character recognition means for extracting textual information from the received image; a translation means for translating the extracted character information; A means to search the internet for food images and descriptions based on the translation results; means for generating images and descriptions of the searched dishes; means for transmitting an image and description of the dish to a user terminal; means for displaying the transmitted information on a user terminal; A system including:

2. The system of claim 1 further comprising means for summarizing the dish description.

3. 10. The system of claim 1, further comprising means for using encryption in transmitting the image and description of the dish.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A