system

A system using image input terminals and AI/OCR technologies to generate ingredient lists and cooking instructions addresses the inefficiency in daily recipe determination, improving ingredient utilization and satisfaction by providing personalized meal suggestions.

JP2026070956APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Efficiently determining daily recipes based on available ingredients and past preferences is a time-consuming and laborious task for many households, leading to inefficient ingredient utilization and reduced food satisfaction due to limited recipe options.

Method used

A system that uses image input terminals to capture ingredient and purchase information, employing AI and OCR technologies to generate an ingredient list, suggest dishes, and provide cooking instructions in both text and video formats, considering past cooking history and user preferences.

Benefits of technology

Streamlines daily meal selection, minimizes food waste, and enhances cooking satisfaction by efficiently utilizing ingredients and providing personalized recipe suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070956000001_ABST
    Figure 2026070956000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] The system receives images of food ingredients or purchase information taken by the user using an image input device. A means for analyzing the image and extracting ingredients or purchase information from the image, A means for generating a list of ingredients based on the extracted information, A means for determining the menu candidates for the day using past cooking history and preference data and the aforementioned list of ingredients, A means of providing users with multiple cooking procedure data related to the candidate dish, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When considering daily menus, efficiently determining recipes based on available ingredients and past preferences is a time-consuming and laborious task for many households. Also, the effective utilization of ingredients may not be fully achieved, resulting in waste of ingredients. In addition, there is a problem that the satisfaction of the family's food decreases due to limited recipe options.

Means for Solving the Problems

[0005] This invention solves this problem by receiving images of ingredients and purchase information taken by the user using an image input terminal, and extracting ingredients or purchase information using artificial intelligence-based image analysis technology. Furthermore, it generates an ingredient list based on the extracted information and determines the suggested dishes for that day by utilizing past cooking history and preference data. This allows the user to select dishes efficiently. In addition, it supports the actual cooking process by providing multiple cooking procedure data for the selected dish candidates. In particular, this invention makes it possible to proceed with cooking in a visually easy-to-understand manner by providing a video format.

[0006] "Images of food ingredients or purchase information taken by the user using an image input device" refers to image data of food ingredients in the refrigerator or products purchased at a store, obtained by the user using a smartphone or tablet.

[0007] "Means for analyzing images and extracting food ingredients or purchase information from them" refers to the process of identifying food ingredients and product information from received image data using AI technology or optical character recognition (OCR) technology and extracting it as text data.

[0008] "Method for generating an ingredient list" refers to the process of creating a list of currently available ingredients based on extracted text data.

[0009] "Past cooking history and preference data" refers to a dataset that includes records of dishes the user has created in the past, as well as taste preferences they have expressed in the past.

[0010] The "method for determining dish candidates" is an algorithm that combines a list of ingredients with user preference data to suggest multiple dishes to be prepared on the day.

[0011] "Means of providing cooking procedure data to users" refers to the process of sending detailed cooking steps for selected dish candidates to the user's device in video or text format.

[0012] "Including video format" refers to content provided in video format to make cooking procedures easier to understand visually.

[0013] "Means of constructing multiple recipe options" refers to the process of presenting users with several choices for a selected dish, including different cooking methods and variations. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system that efficiently utilizes ingredient and purchase information to suggest daily meals. This system employs a method where users can easily provide ingredient and purchase information using an image input terminal, and the system automatically determines the meal based on that information.

[0036] User behavior

[0037] Users first use a smartphone or tablet to take photos of the food in their refrigerator and receipts for purchased items. This provides the system with information about the food items they have available on a daily basis. They can also input photos of dishes they have made in the past and their past preferences.

[0038] Terminal processing

[0039] The device receives images taken by the user and first saves them. Then, it sends these images as data to the server. Communication is conducted using a secure protocol, ensuring user privacy.

[0040] Server-based processing

[0041] The server utilizes AI and optical character recognition (OCR) technologies to analyze image data received from the terminal. This extracts ingredient or purchase information from the image and organizes it as text data. Then, based on this organized information, an ingredient list is generated, and combined with past cooking history and preference data, it determines the menu for the day. This process also takes into account the user's nutritional balance, taste preferences, and even seasonal and special offer information.

[0042] Processing after the menu has been decided.

[0043] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. These instructions include not only text format but also a visually appealing video format, allowing users to easily follow along with the cooking process.

[0044] Specific example

[0045] For example, if a user takes a picture of "chicken, carrots, and onions" as ingredients with their device, the server analyzes the image and builds an ingredient list. Based on the user's preferences, it suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and provides the device with cooking instructions along with video links for each dish. The user can then choose one of these options and proceed with cooking while watching the video.

[0046] In this way, the present invention aims to streamline daily meal selection, minimize food waste in the home, and improve satisfaction with cooking.

[0047] The following describes the processing flow.

[0048] Step 1:

[0049] The user uses their smartphone or tablet to take pictures of food items in the refrigerator or receipts for purchased items. The device acquires this image data using a camera app and temporarily stores it within the app.

[0050] Step 2:

[0051] The device uploads stored image data to a server via a dedicated application. Encrypted communication is used for the upload, ensuring user privacy.

[0052] Step 3:

[0053] The server receives image data and performs image analysis using an AI model. Here, the AI ​​extracts text data about ingredients and purchased items, and reads product information from the receipt using optical character recognition (OCR).

[0054] Step 4:

[0055] Based on the text data extracted by the server, a list of available ingredients is generated in the database. This list is organized by date and ingredient category.

[0056] Step 5:

[0057] The server references past cooking history and preference data, and creates a list of dish suggestions based on the information associated with the ingredient list. Nutritional balance and seasonality are also taken into consideration during this process.

[0058] Step 6:

[0059] The server selects a candidate dish and generates detailed cooking instructions in text or video format. This process involves referencing open recipe databases and information from professional chefs for refinement.

[0060] Step 7:

[0061] The device receives a list of recipe suggestions from the server and notifies the user. The user can then select a recipe from the displayed suggestions and view the cooking instructions.

[0062] Step 8:

[0063] The system checks the necessary ingredients according to the dish selected by the user and begins cooking. Users can efficiently prepare their meals by referring to the cooking instructions provided on the device, which are available as videos and text.

[0064] (Example 1)

[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0066] In modern times, managing ingredients and suggesting meals at home is time-consuming, laborious, and difficult to do efficiently. In particular, there are many challenges in making the most of the ingredients in the refrigerator and suggesting appropriate meals daily that consider individual preferences and nutritional balance. Furthermore, there is a need for cooking instructions to be presented in a detailed, easy-to-understand, and actionable format.

[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0068] In this invention, the server includes means for receiving and analyzing images of items or purchase information captured by an information processing device to extract items or purchase information from the images, means for generating an item list based on the extracted information, means for determining dish candidates using past cooking history and preference data, means for ensuring communication security and protecting data, and means for extracting information using visual recognition technology. This makes it possible to efficiently and safely provide users with optimal dish suggestions and procedures.

[0069] An "information processing device" is an electronic device used by users to take pictures and process them as data.

[0070] "Items" refers to goods or food products that are photographed by an information processing device.

[0071] "Purchase information" refers to information about purchased goods, including receipts and electronic purchase records.

[0072] "Image analysis" is a data processing technique used to extract necessary information from captured image data.

[0073] "Extraction" refers to the process of selecting items and purchase information obtained through image analysis and extracting them as necessary data.

[0074] An "item list" refers to a systematically organized list based on extracted items and purchasing information.

[0075] "Cooking history" refers to records of dishes that have been cooked in the past and the ingredients used.

[0076] "Preference data" refers to information based on users' taste preferences and food choices.

[0077] "Cooking suggestions" refers to a list of possible dishes suggested based on the analyzed information.

[0078] "Communication security" means that information is protected from unauthorized access during data transmission and reception.

[0079] "Data protection" refers to measures taken to safeguard users' personal information and data related to their privacy.

[0080] "Visual recognition technology" is a technology that analyzes image data and enables machines to perform identification equivalent to that of humans.

[0081] This invention is a system that efficiently manages and utilizes goods and purchase information to provide optimal suggestions for users' daily lives. This system utilizes an information processing device to allow users to easily provide goods and purchase information, and then uses this information to determine the most suitable meal.

[0082] User behavior

[0083] Users use information processing devices such as smartphones or tablets to take photos of items in their refrigerator and receipts as purchase information. This operation allows them to provide the system with information about items they can use on a daily basis. They can also input information about meals they have cooked in the past and their personal preferences.

[0084] Terminal processing

[0085] The device receives the image data captured by the user and stores it temporarily. Then, to securely transmit this image data to the server, it uses a protocol that guarantees communication security (e.g., HTTPS).

[0086] Server-based processing

[0087] The server receives image data transmitted from the terminal and uses visual recognition and optical character recognition (OCR) technologies to extract items and purchase information from the image. The extracted data is organized into an item list. The server then compares this item list with past cooking history and preference data, and uses a generative AI model to determine the day's meal options. Nutritional balance, individual preferences, seasonal and special offer information are also taken into consideration.

[0088] Processing after the menu has been decided.

[0089] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. This information is available in text format as well as a visually appealing video format, which the user can use as a guide while cooking.

[0090] Examples of specific cases and prompt statements

[0091] For example, if a user takes a picture of "chicken, carrots, and onions" using their device, the server analyzes the image and builds an item list. Based on the user's preferences, it then suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and generates and displays video links to the cooking procedures for each dish on the device. The user can then choose one of these options and proceed with cooking while referring to the video. An example of a prompt message is, "Please tell me what I can make with the chicken, carrots, and onions I have in the refrigerator. If possible, please also tell me about dishes that are nutritionally balanced and provide detailed cooking instructions."

[0092] In this way, the system aims to improve user satisfaction by streamlining daily inventory management and recipe suggestions, minimizing waste.

[0093] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0094] Step 1:

[0095] The user uses an information processing device to photograph items and receipts inside the refrigerator. Image data is generated as input. This image data records the condition and quantity of the items, as well as purchase information.

[0096] Step 2:

[0097] The terminal temporarily stores image data obtained from the user and then prepares it for transmission to the server. At this stage, a secure transmission protocol is used to ensure the data's security. The output is a data package that can be securely transmitted.

[0098] Step 3:

[0099] The server processes image data received from the terminal. It receives image data as input and analyzes the data using visual recognition technology and optical character recognition (OCR). Specifically, it detects text information within the image and extracts it as item information. The output is data in the form of an item list.

[0100] Step 4:

[0101] The server determines dish candidates based on the generated item list, referencing past cooking history and preference data. This process uses a generative AI model to generate optimal dish suggestions. It uses the item list and preference data as input and creates a list of dish candidates as output.

[0102] Step 5:

[0103] The server creates detailed cooking instructions based on the suggested dishes and sends them to the terminal. The input is a list of suggested dishes, and the output is the cooking instructions for each dish along with a related video link. This allows the user to easily understand how to cook.

[0104] Step 6:

[0105] The user reviews the provided cooking instructions and proceeds with the cooking process accordingly. The user can smoothly complete the cooking process by referring to the videos and text instructions displayed on the device. The output is the finished dish.

[0106] (Application Example 1)

[0107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0108] In modern times, preparing meals at home is often a burden due to busy daily lives. Furthermore, it's difficult to efficiently utilize available ingredients, reduce food waste, and provide appropriate menus. Therefore, there is a growing demand for systems that easily determine appropriate dishes based on ingredient information. Additionally, menu suggestions in food delivery services often lack sufficient utilization of ingredients already available at home, resulting in challenges in improving customer satisfaction.

[0109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0110] In this invention, the server includes means for receiving images of information captured by the user using an image input device, analyzing the images to extract material or purchase information from the images, generating a list of materials based on the extracted information, determining a menu suggestion for the period using past processing history, preference data, and the material list, and suggesting selectable items based on the user's existing materials. This enables the user to make the most of the materials available at that time and prepare a meal efficiently and with high satisfaction. Furthermore, it enables flexible menu suggestions tailored to the user's needs in food delivery services.

[0111] A "user" is an individual or corporation that operates the system and inputs images of food ingredients and purchase information.

[0112] An "image input device" is a device used to input images of information captured by the user into a system, and includes smartphones and tablets.

[0113] "Ingredients" refer to the elements necessary for creating a dish, including food ingredients, seasonings, and other necessities.

[0114] "Purchase information" refers to data about products acquired by the user, including the date and time of purchase, product name, and quantity.

[0115] A "materials list" is a list that organizes and digitizes the extracted materials information.

[0116] "Processing history" refers to records of dishes a user has cooked in the past, and is data used to analyze specific patterns and trends.

[0117] "Preference data" refers to information that indicates a user's taste preferences and dietary preferences, and is used as a reference when suggesting dishes.

[0118] A "recipe suggestion" is a set of dishes or meal options determined based on an ingredient list and other relevant data.

[0119] "Offered items" refer to meals and goods that can be delivered, such as those offered by food delivery services.

[0120] "Demonstration format" refers to a means of visually showing the cooking process to the user, and is expressed through videos or illustrations.

[0121] To realize this invention, the user first uses an image input device such as a smartphone or tablet to take pictures of ingredients in the refrigerator or receipts for purchased items. This provides the system with information on ingredients that are readily available on a daily basis. The user can also input photos of dishes they have made in the past and data on their preferences.

[0122] The image input device prepares these images as data and sends them to the server using a secure protocol. This protocol, such as HTTPS, protects user privacy.

[0123] The server preprocesses images received from terminals using OpenCV, an open-source image processing library. Then, it extracts material and purchase information from the images using AI and optical character recognition (OCR) technologies, such as TENSORFLOW® and Tesseract, and organizes it as text data.

[0124] Furthermore, the server generates a list of ingredients from the extracted information and combines it with past processing history and preference data to determine the optimal dish suggestion for that moment. Based on the suggested dish, the server generates detailed cooking instructions and sends them to the image input device. These cooking instructions are provided in both text and video formats.

[0125] For example, if a user enters "cheese, tomato, and salami" as ingredients, the system analyzes this and suggests dishes such as "homemade pizza" or "toast with lots of toppings." Based on these suggestions, it also presents specific pizza doughs to the user as part of a food delivery menu.

[0126] An example of a prompt for a generative AI model would be, "Based on the ingredients in the refrigerator—cheese, tomatoes, and salami—please suggest a menu of available pizza doughs." This allows for the provision of meals tailored to individual preferences and requests while minimizing food waste.

[0127] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0128] Step 1:

[0129] The user uses an image input device to photograph food items in the refrigerator or receipts for purchased items. This generates image data. The captured images contain information about the available ingredients.

[0130] Step 2:

[0131] The device sends the captured image data to the server. The HTTPS protocol is used for communication, protecting user privacy. At this point, the input is the image data, and the output is the preparation for the server to receive it.

[0132] Step 3:

[0133] The server uses OpenCV to preprocess the received image data. This preprocessing involves denoising the image and extracting necessary regions, preparing the image for analysis. This results in high-quality images for analysis as output.

[0134] Step 4:

[0135] The server uses TensorFlow and Tesseract to extract material and purchase information from pre-processed images. AI model recognition and OCR technology are used to obtain information as text data from the images. The input for this step is a pre-processed image, and the output is text data containing material and purchase information.

[0136] Step 5:

[0137] The server creates an ingredient list based on extracted text data and determines a recipe suggestion by combining it with past processing history and preference data. A generative AI model is used to select a recipe that reflects the user's preferences. The ingredient list and preference data are inputs, and recipe suggestions are output.

[0138] Step 6:

[0139] The server generates detailed cooking instructions based on the selected dish suggestion and provides them to the terminal. The cooking instructions are generated in text and video formats. Specific step-by-step data for realizing the suggested dish is provided as output.

[0140] Step 7:

[0141] Users review suggested dishes and cooking instructions provided through their terminal and order items via delivery service as needed. The terminal receives the suggestions, and the selected information becomes the input for the next action.

[0142] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0143] This invention provides a system that simplifies users' daily food management and meal selection, and further enhances personalized meal suggestions by incorporating an emotion engine. The system has the function of generating meal candidates based on ingredients and purchase information, and the function of recognizing the user's emotions and reflecting them in meal selection.

[0144] User behavior

[0145] The user first uses a smartphone or tablet to take photos of the food items in their refrigerator and receipts for purchased items. In addition, the user communicates their emotions to the device through their own input or an interactive user interface. As a result, data on both food items and emotions is provided to the system.

[0146] Terminal processing

[0147] The device receives images and emotional input data captured by the user and sends them to the server. During this process, image data is compressed if necessary while maintaining high quality before uploading. Emotional data is obtained through methods such as voice, text, and facial expression analysis.

[0148] Server-based processing

[0149] After receiving image data, the server performs image analysis using AI technology and optical character recognition (OCR). This extracts ingredient and purchase information. Next, the emotion engine analyzes the emotion data received from the user to determine the emotional state. Based on this data, it is combined with the user's past cooking history and preference data to generate menu suggestions for the day. By considering the emotional state, the menu selection becomes more tailored to the user's current state.

[0150] Processing after the menu has been decided.

[0151] The server selects a dish and sends it to the user's device in text or video format, along with detailed cooking instructions. The video format is visually easy to understand and serves as a practical cooking guide for the user.

[0152] Specific example

[0153] For example, if a user takes a photo of "cheese, tomatoes, and bread" as ingredients and inputs a happy and relaxed emotional state, the server immediately analyzes this. As a result, it suggests relaxing and enjoyable meal options such as "grilled cheese sandwich, tomato soup, and bruschetta." Cooking instructions for each dish are provided on the device, and the user can choose their desired dish and begin cooking.

[0154] In this way, the present invention aims to improve the cooking experience by considering both ingredient information and emotional state, thereby enabling optimal cooking suggestions for the user.

[0155] The following describes the processing flow.

[0156] Step 1:

[0157] The user uses a smartphone or tablet to take photos of food items in the refrigerator or receipts for purchased items. At the same time, the user inputs their emotions through the device's interface using simple questions and voice guidance.

[0158] Step 2:

[0159] The device transmits captured image data and entered emotion data to the server using a secure protocol. Image data is appropriately compressed, and emotion data is encrypted before transmission.

[0160] Step 3:

[0161] The server receives the image data and performs image analysis using AI technology and optical character recognition (OCR). This extracts information about ingredients and purchases, which are then organized into a database.

[0162] Step 4:

[0163] The server uses an emotion engine to analyze the emotion data sent by the user. This analysis determines the user's current emotional state (e.g., happy, stressed, anxious, etc.).

[0164] Step 5:

[0165] The server combines extracted ingredient lists, past cooking history, preference data, and emotional states to generate meal suggestions suitable for the user on that day. These suggestions take into account nutritional balance, the season, and even the user's emotional state.

[0166] Step 6:

[0167] The server generates detailed cooking instructions for the selected dish in both text and video formats and sends them to the terminal. The videos are visually easy to understand and show the actual cooking process step by step.

[0168] Step 7:

[0169] The terminal displays a list of dish suggestions received from the server to the user and notifies them. The user can then select their preferred dish from the displayed suggestions and view the detailed cooking instructions.

[0170] Step 8:

[0171] The user begins cooking the dish they have selected. Following the instructions provided on the device in text and video formats, they can efficiently proceed with the cooking process. The cooking experience can be enhanced by allowing users to enjoy cooking in a way that aligns with their emotions.

[0172] (Example 2)

[0173] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0174] In modern life, consumers spend a significant amount of time and effort managing their food and choosing meals. Furthermore, because users' emotional states are often overlooked in meal selection, the resulting meal choices frequently fail to align with their desired cooking experience. Therefore, it is necessary to improve the cooking experience by providing more personalized meal suggestions that take into account the user's food availability and emotional state.

[0175] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0176] In this invention, the server includes means for receiving images of ingredients or purchase information captured by the user using an image input device, and emotional state data; means for analyzing the images to extract ingredients and purchase information from the images; and means for analyzing the emotional state data to determine the user's emotional state. This makes it possible to generate optimal dish candidates that take both ingredient information and emotional state into consideration.

[0177] "Users" refer to individuals who operate this system and provide food ingredients and emotional data.

[0178] An "image input device" refers to an electronic device with a camera function that users use to collect information on ingredients and purchases.

[0179] "Emotional state data" refers to information that indicates the user's emotions and is entered into the device in text, audio, or other formats.

[0180] A "generative model" refers to artificial intelligence technology used to generate optimal dish candidates by integrating multiple pieces of information.

[0181] "Visual format" refers to video and image formats that display cooking instructions in a way that makes them easy for users to understand visually.

[0182] A "recipe suggestion" refers to a set of multiple dish options proposed based on the user's available ingredients and emotional state.

[0183] This invention is a system that facilitates users' daily food management and meal selection. Based on the user's food situation and emotional state, the system provides appropriate meal suggestions and primarily consists of an image input device, a terminal, and a server.

[0184] Users take photos of food items in their refrigerator and receipts for purchased items using a smartphone or tablet. They also input their emotional state into the device via text or voice. The food information and emotional state data obtained in this way are then transmitted to a server by the device. Image data is compressed as needed while maintaining high quality.

[0185] The server analyzes the received image data using AI technology and optical character recognition (OCR). This extracts ingredient names and purchase information. Meanwhile, emotional state data is analyzed by an emotion engine to determine the user's current emotional state. The server combines this two pieces of information and uses a generative model to generate optimal dish suggestions. The generative model receives a prompt message and provides new dish suggestions.

[0186] As a concrete example, consider a scenario where a user takes a photo of "cheese, tomatoes, and bread" and inputs a happy and relaxed emotional state. In this case, the server analyzes this and suggests dish options such as "grilled cheese sandwich, tomato soup, and bruschetta." The cooking instructions for the suggested dishes are displayed on the device in a visually easy-to-understand video format.

[0187] An example of a prompt for a generative AI model is: "The user has entered cheese, tomato, and bread as ingredients, and their emotions are happy and relaxed. Please suggest three dishes that meet these conditions."

[0188] Thus, the present invention aims to provide users with appropriate cooking suggestions based on both ingredient information and their emotional state.

[0189] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0190] Step 1:

[0191] The user takes photos of the food items and purchase information inside the refrigerator using an image input device. They also input their emotional state via text and voice through an interactive interface. This process generates food image data and emotional data.

[0192] Step 2:

[0193] The terminal processes food image data and emotion data received from the user. Captured images may be compressed for efficient data transfer while maintaining high quality. Emotion data is converted to a predefined data format and sent to the server. The output consists of compressed image data and formatted emotion data.

[0194] Step 3:

[0195] The server analyzes the image data received from the terminal using AI technology and optical character recognition (OCR). This extracts the names of ingredients and purchase information from the image as text. The input is compressed image data, and the output is text-based ingredient information.

[0196] Step 4:

[0197] The server simultaneously analyzes the received emotion data using an emotion engine to determine the user's emotional state. The input is formatted emotion data, and the output is the identified emotional state. Specifically, the emotion data is classified based on speech and specific word patterns.

[0198] Step 5:

[0199] The server uses a generative AI model to generate optimal dish candidates based on extracted ingredient information and determined emotional states. Prompt messages are input to the generative model, and newly suggested dish ideas are output. Specifically, an evaluation function operates by referencing past cooking history data.

[0200] Step 6:

[0201] The server determines the dish options and sends detailed cooking procedure data to the terminal in a visually easy-to-understand format. The terminal receives this data and presents it to the user as options. The input is the generated dish suggestions, and the output is the cooking procedure data presented to the user. Specific actions include displaying the information in text or video format.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] Modern consumers have a need to manage their meals efficiently and comfortably amidst their busy daily lives, but conventional technology has problems in that it does not adequately offer flexible meal suggestions tailored to emotions and situations, nor does it integrate well with external services. In particular, the challenge lies in simultaneously achieving the effective use of ingredients and providing services that are optimal for each individual user.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means for analyzing visual data and extracting item or purchase information from the visual data, means for generating suggested candidates that reflect the emotional state, and means for making requests from external services based on the suggested candidates. This enables the generation of personalized suggestions that correspond to the user's emotions and situation, and smooth access to external services based on those suggestions.

[0207] An "information input device" is a device used by users to input visual data of goods or purchase information, and includes portable terminals such as smartphones and tablets.

[0208] "Visual data" refers to image data of items and purchase information acquired using an information input device, which can be analyzed using optical character recognition technology.

[0209] The "item list" is a compilation of item information extracted through visual data analysis, and is used to generate subsequent proposal candidates.

[0210] "Past processing history" refers to a record of the choices and actions a user has taken in the past, and serves as the basis for providing personalized suggestions.

[0211] "Preference data" refers to data about users' preferences and tendencies, and is used to generate more appropriate suggestion candidates.

[0212] "Suggestion candidates that reflect emotional state" are recommendations generated based on the user's emotions and serve as information to present the user with multiple options.

[0213] "Means of outsourcing from external services" refers to a function that connects to external food supply services based on proposed options and according to the user's selection, and executes orders and outsourcing.

[0214] The system for implementing this invention allows users to easily manage food ingredients and purchase information, and provides optimal cooking suggestions and connections to external services based on their emotional state. This system includes an information input device such as a smartphone, a server, and a network configuration for connecting to external services.

[0215] First, users use an information input device to photograph food ingredients, receipts, and other purchase information, and send the image data to the server. The server then uses optical character recognition (OCR) technology to analyze the visual data and generate an item list. Possible OCR technologies used in this process include open-source Tesseract and commercial cloud-based services.

[0216] The server then uses an emotion analysis API to evaluate the user's emotional state based on the text or voice input. Along with this emotional data, it references past processing history and preference data to generate suggested options that reflect the emotional state. This part utilizes natural language processing models and machine learning techniques.

[0217] The generated proposal candidates are sent to an information input device and presented to the user in a visual format (e.g., images or videos). The user can select from the proposal candidates, and based on the selected candidate, the server automatically connects with an external food delivery service and executes the order or commission.

[0218] For example, if a user provides image data of "pasta, cheese, and tomatoes" and inputs a mood indicating they want to relax, the server will generate suggestions such as "creamy pasta" or "cheese and tomato salad" and enable them to order from nearby partner restaurants. Examples of prompts include "How are you feeling today?", "Please suggest some dishes to help me relax," and "What pasta dishes do you recommend?".

[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0220] Step 1:

[0221] The user uses an information input device to photograph food items in the refrigerator or receipts for purchased items. This captured image data becomes the input, and the terminal compresses this data as needed while maintaining high quality before sending it to the server. JPEG and PNG are common image data formats.

[0222] Step 2:

[0223] The server uses the received image data as input to extract ingredient and purchase information from the visual data using optical character recognition (OCR) technology. This process analyzes the textual information within the visual data and outputs an item list. OCR technologies used include Tesseract and cloud-based OCR APIs.

[0224] Step 3:

[0225] The user inputs their emotional state as voice or text through an information input device. This emotional data is sent to the server as input. The server uses an emotional analysis API to analyze the input emotional data and outputs the user's emotional state as a numerical value or category.

[0226] Step 4:

[0227] The server receives item lists, emotional states, past processing history, and preference data as input, and uses a generative AI model to generate suggested candidates that reflect the emotional state based on this data. In this process, the generative AI model integrates these multiple data sources to output the optimal suggested candidate.

[0228] Step 5:

[0229] The generated suggestion candidates are sent from the server to the terminal. The terminal receives these outputted suggestion candidates and presents them to the user in a visual format (e.g., images or videos). This makes it easier for the user to visually review the options.

[0230] Step 6:

[0231] The user selects from the suggested options presented on the terminal. This selection is sent as input to the server, which uses this data to place an order or outsource the order for the suggested dish to an external service. The server accesses partner delivery services through the interface and places the order automatically.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention is a system that efficiently utilizes ingredient and purchase information to suggest daily meals. This system employs a method where users can easily provide ingredient and purchase information using an image input terminal, and the system automatically determines the meal based on that information.

[0249] User behavior

[0250] Users first use a smartphone or tablet to take photos of the food in their refrigerator and receipts for purchased items. This provides the system with information about the food items they have available on a daily basis. They can also input photos of dishes they have made in the past and their past preferences.

[0251] Terminal processing

[0252] The device receives images taken by the user and first saves them. Then, it sends these images as data to the server. Communication is conducted using a secure protocol, ensuring user privacy.

[0253] Server-based processing

[0254] The server utilizes AI and optical character recognition (OCR) technologies to analyze image data received from the terminal. This extracts ingredient or purchase information from the image and organizes it as text data. Then, based on this organized information, an ingredient list is generated, and combined with past cooking history and preference data, it determines the menu for the day. This process also takes into account the user's nutritional balance, taste preferences, and even seasonal and special offer information.

[0255] Processing after the menu has been decided.

[0256] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. These instructions include not only text format but also a visually appealing video format, allowing users to easily follow along with the cooking process.

[0257] Specific example

[0258] For example, if a user takes a picture of "chicken, carrots, and onions" as ingredients with their device, the server analyzes the image and builds an ingredient list. Based on the user's preferences, it suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and provides the device with cooking instructions along with video links for each dish. The user can then choose one of these options and proceed with cooking while watching the video.

[0259] In this way, the present invention aims to streamline daily meal selection, minimize food waste in the home, and improve satisfaction with cooking.

[0260] The following describes the processing flow.

[0261] Step 1:

[0262] The user uses their smartphone or tablet to take pictures of food items in the refrigerator or receipts for purchased items. The device acquires this image data using a camera app and temporarily stores it within the app.

[0263] Step 2:

[0264] The device uploads stored image data to a server via a dedicated application. Encrypted communication is used for the upload, ensuring user privacy.

[0265] Step 3:

[0266] The server receives image data and performs image analysis using an AI model. Here, the AI ​​extracts text data about ingredients and purchased items, and reads product information from the receipt using optical character recognition (OCR).

[0267] Step 4:

[0268] Based on the text data extracted by the server, a list of available ingredients is generated in the database. This list is organized by date and ingredient category.

[0269] Step 5:

[0270] The server references past cooking history and preference data, and creates a list of dish suggestions based on the information associated with the ingredient list. Nutritional balance and seasonality are also taken into consideration during this process.

[0271] Step 6:

[0272] The server selects a candidate dish and generates detailed cooking instructions in text or video format. This process involves referencing open recipe databases and information from professional chefs for refinement.

[0273] Step 7:

[0274] The device receives a list of recipe suggestions from the server and notifies the user. The user can then select a recipe from the displayed suggestions and view the cooking instructions.

[0275] Step 8:

[0276] The system checks the necessary ingredients according to the dish selected by the user and begins cooking. Users can efficiently prepare their meals by referring to the cooking instructions provided on the device, which are available as videos and text.

[0277] (Example 1)

[0278] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0279] In modern times, managing ingredients and suggesting meals at home is time-consuming, laborious, and difficult to do efficiently. In particular, there are many challenges in making the most of the ingredients in the refrigerator and suggesting appropriate meals daily that consider individual preferences and nutritional balance. Furthermore, there is a need for cooking instructions to be presented in a detailed, easy-to-understand, and actionable format.

[0280] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0281] In this invention, the server includes means for receiving an image of an article or purchase information photographed by an information processing device, analyzing the image, and extracting the article or purchase information in the image; means for generating an item list based on the extracted information; means for determining cooking candidates using the past cooking history and preference data; means for ensuring communication security and performing data protection; and means for performing information extraction using visual recognition technology. As a result, it becomes possible to efficiently and safely propose the optimal cooking and provide procedures to the user.

[0282] The "information processing device" is an electronic device for a user to take a picture and process it as data.

[0283] The "article" refers to a product or food ingredient that is the object photographed by the information processing device.

[0284] The "purchase information" is information about the purchased product and includes receipts and electronic purchase records.

[0285] "Image analysis" is a data processing method performed to extract necessary information from the photographed image data.

[0286] "Extraction" refers to the operation of selecting the articles and purchase information obtained by image analysis and taking out the necessary data.

[0287] The "item list" refers to a systematically organized list based on the extracted articles and purchase information.

[0288] The "cooking history" means a record of the cooked dishes and used ingredients in the past.

[0289] The "preference data" is information based on the user's taste preferences and ingredient preferences.

[0290] The "cooking candidates" refer to a list of possible dishes proposed based on the analyzed information.

[0291] "Communication security" means that information is protected from unauthorized access during data transmission and reception.

[0292] "Data protection" refers to measures taken to safeguard users' personal information and data related to their privacy.

[0293] "Visual recognition technology" is a technology that analyzes image data and enables machines to perform identification equivalent to that of humans.

[0294] This invention is a system that efficiently manages and utilizes goods and purchase information to provide optimal suggestions for users' daily lives. This system utilizes an information processing device to allow users to easily provide goods and purchase information, and then uses this information to determine the most suitable meal.

[0295] User behavior

[0296] Users use information processing devices such as smartphones or tablets to take photos of items in their refrigerator and receipts as purchase information. This operation allows them to provide the system with information about items they can use on a daily basis. They can also input information about meals they have cooked in the past and their personal preferences.

[0297] Terminal processing

[0298] The device receives the image data captured by the user and stores it temporarily. Then, to securely transmit this image data to the server, it uses a protocol that guarantees communication security (e.g., HTTPS).

[0299] Server-based processing

[0300] The server receives the image data transmitted from the terminal and utilizes visual recognition technology and optical character recognition (OCR) technology to extract items and purchase information from the image. The extracted data is organized as an item list. Subsequently, the server compares this item list with past cooking histories and preference data, and uses a generated AI model to determine the cooking candidates for the day. Nutritional balance, individual preferences, season, and sale information are also considered.

[0301] Processing after cooking decision

[0302] Based on the determined cooking candidates, the server generates detailed cooking procedures and transmits them to the terminal. This information includes not only the text format but also the video format that appeals visually, and the user can refer to this to proceed with cooking.

[0303] Examples of specific cases and prompt sentences

[0304] For example, when the user uses the terminal to photograph "chicken, carrot, and onion", the server analyzes this to construct an item list. Then, based on the user's preference information, cooking candidates such as "chicken curry, stir-fried vegetables, boiled chicken" are proposed, and video links related to the cooking procedures for each are generated and displayed on the terminal. The user can select one from them and proceed with cooking referring to the video. An example of a prompt sentence is "Please tell me the dishes that can be made with the chicken, carrots, and onions in the refrigerator. If possible, also tell me those with a good nutritional balance and the cooking methods in detail."

[0305] In this way, the system aims to streamline daily item management and cooking proposals, minimize waste, and improve user satisfaction.

[0306] The flow of specific processing in Example 1 will be described using FIG. 11.

[0307] Step 1:

[0308] The user uses an information processing device to photograph items and receipts inside the refrigerator. Image data is generated as input. This image data records the condition and quantity of the items, as well as purchase information.

[0309] Step 2:

[0310] The terminal temporarily stores image data obtained from the user and then prepares it for transmission to the server. At this stage, a secure transmission protocol is used to ensure the data's security. The output is a data package that can be securely transmitted.

[0311] Step 3:

[0312] The server processes image data received from the terminal. It receives image data as input and analyzes the data using visual recognition technology and optical character recognition (OCR). Specifically, it detects text information within the image and extracts it as item information. The output is data in the form of an item list.

[0313] Step 4:

[0314] The server determines dish candidates based on the generated item list, referencing past cooking history and preference data. This process uses a generative AI model to generate optimal dish suggestions. It uses the item list and preference data as input and creates a list of dish candidates as output.

[0315] Step 5:

[0316] The server creates detailed cooking instructions based on the suggested dishes and sends them to the terminal. The input is a list of suggested dishes, and the output is the cooking instructions for each dish along with a related video link. This allows the user to easily understand how to cook.

[0317] Step 6:

[0318] The user reviews the provided cooking instructions and proceeds with the cooking process accordingly. The user can smoothly complete the cooking process by referring to the videos and text instructions displayed on the device. The output is the finished dish.

[0319] (Application Example 1)

[0320] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0321] In modern times, preparing meals at home is often a burden due to busy daily lives. Furthermore, it's difficult to efficiently utilize available ingredients, reduce food waste, and provide appropriate menus. Therefore, there is a growing demand for systems that easily determine appropriate dishes based on ingredient information. Additionally, menu suggestions in food delivery services often lack sufficient utilization of ingredients already available at home, resulting in challenges in improving customer satisfaction.

[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0323] In this invention, the server includes means for receiving images of information captured by the user using an image input device, analyzing the images to extract material or purchase information from the images, generating a list of materials based on the extracted information, determining a menu suggestion for the period using past processing history, preference data, and the material list, and suggesting selectable items based on the user's existing materials. This enables the user to make the most of the materials available at that time and prepare a meal efficiently and with high satisfaction. Furthermore, it enables flexible menu suggestions tailored to the user's needs in food delivery services.

[0324] A "user" is an individual or corporation that operates the system and inputs images of food ingredients and purchase information.

[0325] An "image input device" is a device used to input images of information captured by the user into a system, and includes smartphones and tablets.

[0326] "Ingredients" refer to the elements necessary for creating a dish, including food ingredients, seasonings, and other necessities.

[0327] "Purchase information" refers to data about products acquired by the user, including the date and time of purchase, product name, and quantity.

[0328] A "materials list" is a list that organizes and digitizes the extracted materials information.

[0329] "Processing history" refers to records of dishes a user has cooked in the past, and is data used to analyze specific patterns and trends.

[0330] "Preference data" refers to information that indicates a user's taste preferences and dietary preferences, and is used as a reference when suggesting dishes.

[0331] A "recipe suggestion" is a set of dishes or meal options determined based on an ingredient list and other relevant data.

[0332] "Offered items" refer to meals and goods that can be delivered, such as those offered by food delivery services.

[0333] "Demonstration format" refers to a means of visually showing the cooking process to the user, and is expressed through videos or illustrations.

[0334] To realize this invention, the user first uses an image input device such as a smartphone or tablet to take pictures of ingredients in the refrigerator or receipts for purchased items. This provides the system with information on ingredients that are readily available on a daily basis. The user can also input photos of dishes they have made in the past and data on their preferences.

[0335] The image input device prepares these images as data and sends them to the server using a secure protocol. This protocol, such as HTTPS, protects user privacy.

[0336] The server preprocesses images received from terminals using OpenCV, an open-source image processing library. Then, it uses AI and optical character recognition (OCR) technologies, such as TensorFlow and Tesseract, to extract material and purchase information from the images and organize it as text data.

[0337] Furthermore, the server generates a list of ingredients from the extracted information and combines it with past processing history and preference data to determine the optimal dish suggestion for that moment. Based on the suggested dish, the server generates detailed cooking instructions and sends them to the image input device. These cooking instructions are provided in both text and video formats.

[0338] For example, if a user enters "cheese, tomato, and salami" as ingredients, the system analyzes this and suggests dishes such as "homemade pizza" or "toast with lots of toppings." Based on these suggestions, it also presents specific pizza doughs to the user as part of a food delivery menu.

[0339] An example of a prompt for a generative AI model would be, "Based on the ingredients in the refrigerator—cheese, tomatoes, and salami—please suggest a menu of available pizza doughs." This allows for the provision of meals tailored to individual preferences and requests while minimizing food waste.

[0340] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0341] Step 1:

[0342] The user uses an image input device to photograph food items in the refrigerator or receipts for purchased items. This generates image data. The captured images contain information about the available ingredients.

[0343] Step 2:

[0344] The device sends the captured image data to the server. The HTTPS protocol is used for communication, protecting user privacy. At this point, the input is the image data, and the output is the preparation for the server to receive it.

[0345] Step 3:

[0346] The server uses OpenCV to preprocess the received image data. This preprocessing involves denoising the image and extracting necessary regions, preparing the image for analysis. This results in high-quality images for analysis as output.

[0347] Step 4:

[0348] The server uses TensorFlow and Tesseract to extract material and purchase information from pre-processed images. AI model recognition and OCR technology are used to obtain information as text data from the images. The input for this step is a pre-processed image, and the output is text data containing material and purchase information.

[0349] Step 5:

[0350] The server creates an ingredient list based on extracted text data and determines a recipe suggestion by combining it with past processing history and preference data. A generative AI model is used to select a recipe that reflects the user's preferences. The ingredient list and preference data are inputs, and recipe suggestions are output.

[0351] Step 6:

[0352] Based on the selected dish suggestion, the server generates detailed cooking instructions and provides them to the terminal. The cooking instructions are generated in both text and video formats. Specific step-by-step data for realizing the suggested dish is provided as output.

[0353] Step 7:

[0354] Users review suggested dishes and cooking instructions provided through their terminal and order items via delivery service as needed. The terminal receives the suggestions, and the selected information becomes the input for the next action.

[0355] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0356] This invention provides a system that simplifies users' daily food management and meal selection, and further enhances personalized meal suggestions by incorporating an emotion engine. The system has the function of generating meal candidates based on ingredients and purchase information, and the function of recognizing the user's emotions and reflecting them in meal selection.

[0357] User behavior

[0358] The user first uses a smartphone or tablet to take photos of the food items in their refrigerator and receipts for purchased items. In addition, the user communicates their emotions to the device through their own input or an interactive user interface. As a result, data on both food items and emotions is provided to the system.

[0359] Terminal processing

[0360] The device receives images and emotional input data captured by the user and sends them to the server. During this process, image data is compressed if necessary while maintaining high quality before uploading. Emotional data is obtained through methods such as voice, text, and facial expression analysis.

[0361] Server-based processing

[0362] After receiving image data, the server performs image analysis using AI technology and optical character recognition (OCR). This extracts ingredient and purchase information. Next, the emotion engine analyzes the emotion data received from the user to determine the emotional state. Based on this data, it is combined with the user's past cooking history and preference data to generate menu suggestions for the day. By considering the emotional state, the menu selection becomes more tailored to the user's current state.

[0363] Processing after the menu has been decided.

[0364] The server selects a dish and sends it to the user's device in text or video format, along with detailed cooking instructions. The video format is visually easy to understand and serves as a practical cooking guide for the user.

[0365] Specific example

[0366] For example, if a user takes a photo of "cheese, tomatoes, and bread" as ingredients and inputs a happy and relaxed emotional state, the server immediately analyzes this. As a result, it suggests relaxing and enjoyable meal options such as "grilled cheese sandwich, tomato soup, and bruschetta." Cooking instructions for each dish are provided on the device, and the user can choose their desired dish and begin cooking.

[0367] In this way, the present invention aims to improve the cooking experience by considering both ingredient information and emotional state, thereby enabling optimal cooking suggestions for the user.

[0368] The following describes the processing flow.

[0369] Step 1:

[0370] The user uses a smartphone or tablet to take photos of food items in the refrigerator or receipts for purchased items. At the same time, the user inputs their emotions through the device's interface using simple questions and voice guidance.

[0371] Step 2:

[0372] The device transmits captured image data and entered emotion data to the server using a secure protocol. Image data is appropriately compressed, and emotion data is encrypted before transmission.

[0373] Step 3:

[0374] The server receives the image data and performs image analysis using AI technology and optical character recognition (OCR). This extracts information about ingredients and purchases, which are then organized into a database.

[0375] Step 4:

[0376] The server uses an emotion engine to analyze the emotion data sent by the user. This analysis determines the user's current emotional state (e.g., happy, stressed, anxious, etc.).

[0377] Step 5:

[0378] The server combines extracted ingredient lists, past cooking history, preference data, and emotional states to generate meal suggestions suitable for the user on that day. These suggestions take into account nutritional balance, the season, and even the user's emotional state.

[0379] Step 6:

[0380] The server generates detailed cooking instructions for the selected dish in both text and video formats and sends them to the terminal. The videos are visually easy to understand and show the actual cooking process step by step.

[0381] Step 7:

[0382] The terminal displays a list of dish suggestions received from the server to the user and notifies them. The user can then select their preferred dish from the displayed suggestions and view the detailed cooking instructions.

[0383] Step 8:

[0384] The user begins cooking the dish they have selected. Following the instructions provided on the device in text and video formats, they can efficiently proceed with the cooking process. The cooking experience can be enhanced by allowing users to enjoy cooking in a way that aligns with their emotions.

[0385] (Example 2)

[0386] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0387] In modern life, consumers spend a significant amount of time and effort managing their food and choosing meals. Furthermore, because users' emotional states are often overlooked in meal selection, the resulting meal choices frequently fail to align with their desired cooking experience. Therefore, it is necessary to improve the cooking experience by providing more personalized meal suggestions that take into account the user's food availability and emotional state.

[0388] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0389] In this invention, the server includes means for receiving images of ingredients or purchase information captured by the user using an image input device, and emotional state data; means for analyzing the images to extract ingredients and purchase information from the images; and means for analyzing the emotional state data to determine the user's emotional state. This makes it possible to generate optimal dish candidates that take both ingredient information and emotional state into consideration.

[0390] "Users" refer to individuals who operate this system and provide food ingredients and emotional data.

[0391] An "image input device" refers to an electronic device with a camera function that users use to collect information on ingredients and purchases.

[0392] "Emotional state data" refers to information that indicates the user's emotions and is entered into the device in text, audio, or other formats.

[0393] A "generative model" refers to artificial intelligence technology used to generate optimal dish candidates by integrating multiple pieces of information.

[0394] "Visual format" refers to video and image formats that display cooking instructions in a way that makes them easy for users to understand visually.

[0395] A "recipe suggestion" refers to a set of multiple dish options proposed based on the user's available ingredients and emotional state.

[0396] This invention is a system that facilitates users' daily food management and meal selection. Based on the user's food situation and emotional state, the system provides appropriate meal suggestions and primarily consists of an image input device, a terminal, and a server.

[0397] Users take photos of food items in their refrigerator and receipts for purchased items using a smartphone or tablet. They also input their emotional state into the device via text or voice. The food information and emotional state data obtained in this way are then transmitted to a server by the device. Image data is compressed as needed while maintaining high quality.

[0398] The server analyzes the received image data using AI technology and optical character recognition (OCR). This extracts ingredient names and purchase information. Meanwhile, emotional state data is analyzed by an emotion engine to determine the user's current emotional state. The server combines this two pieces of information and uses a generative model to generate optimal dish suggestions. The generative model receives a prompt message and provides new dish suggestions.

[0399] As a concrete example, consider a scenario where a user takes a photo of "cheese, tomatoes, and bread" and inputs a happy and relaxed emotional state. In this case, the server analyzes this and suggests dish options such as "grilled cheese sandwich, tomato soup, and bruschetta." The cooking instructions for the suggested dishes are displayed on the device in a visually easy-to-understand video format.

[0400] An example of a prompt for a generative AI model is: "The user has entered cheese, tomato, and bread as ingredients, and their emotions are happy and relaxed. Please suggest three dishes that meet these conditions."

[0401] Thus, the present invention aims to provide users with appropriate cooking suggestions based on both ingredient information and their emotional state.

[0402] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0403] Step 1:

[0404] The user takes photos of the food items and purchase information inside the refrigerator using an image input device. They also input their emotional state via text and voice through an interactive interface. This process generates food image data and emotional data.

[0405] Step 2:

[0406] The terminal processes food image data and emotion data received from the user. Captured images may be compressed for efficient data transfer while maintaining high quality. Emotion data is converted to a predefined data format and sent to the server. The output consists of compressed image data and formatted emotion data.

[0407] Step 3:

[0408] The server analyzes the image data received from the terminal using AI technology and optical character recognition (OCR). This extracts the names of ingredients and purchase information from the image as text. The input is compressed image data, and the output is text-based ingredient information.

[0409] Step 4:

[0410] The server simultaneously analyzes the received emotion data using an emotion engine to determine the user's emotional state. The input is formatted emotion data, and the output is the identified emotional state. Specifically, the emotion data is classified based on speech and specific word patterns.

[0411] Step 5:

[0412] The server uses a generative AI model to generate optimal dish candidates based on extracted ingredient information and determined emotional states. Prompt messages are input to the generative model, and newly suggested dish ideas are output. Specifically, an evaluation function operates by referencing past cooking history data.

[0413] Step 6:

[0414] The server determines the dish options and sends detailed cooking procedure data to the terminal in a visually easy-to-understand format. The terminal receives this data and presents it to the user as options. The input is the generated dish suggestions, and the output is the cooking procedure data presented to the user. Specific actions include displaying the information in text or video format.

[0415] (Application Example 2)

[0416] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0417] Modern consumers have a need to manage their meals efficiently and comfortably amidst their busy daily lives, but conventional technology has problems in that it does not adequately offer flexible meal suggestions tailored to emotions and situations, nor does it integrate well with external services. In particular, the challenge lies in simultaneously achieving the effective use of ingredients and providing services that are optimal for each individual user.

[0418] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0419] In this invention, the server includes means for analyzing visual data and extracting item or purchase information from the visual data, means for generating suggested candidates that reflect the emotional state, and means for making requests from external services based on the suggested candidates. This enables the generation of personalized suggestions that correspond to the user's emotions and situation, and smooth access to external services based on those suggestions.

[0420] An "information input device" is a device used by users to input visual data of goods or purchase information, and includes portable terminals such as smartphones and tablets.

[0421] "Visual data" refers to image data of items and purchase information acquired using an information input device, which can be analyzed using optical character recognition technology.

[0422] The "item list" is a compilation of item information extracted through visual data analysis, and is used to generate subsequent proposal candidates.

[0423] "Past processing history" refers to a record of the choices and actions a user has taken in the past, and serves as the basis for providing personalized suggestions.

[0424] "Preference data" refers to data about users' preferences and tendencies, and is used to generate more appropriate suggestion candidates.

[0425] "Suggestion candidates that reflect emotional state" are recommendations generated based on the user's emotions and serve as information to present the user with multiple options.

[0426] "Means of outsourcing from external services" refers to a function that connects to external food supply services based on proposed options and according to the user's selection, and executes orders and outsourcing.

[0427] The system for implementing this invention allows users to easily manage food ingredients and purchase information, and provides optimal cooking suggestions and connections to external services based on their emotional state. This system includes an information input device such as a smartphone, a server, and a network configuration for connecting to external services.

[0428] First, users use an information input device to photograph food ingredients, receipts, and other purchase information, and send the image data to the server. The server then uses optical character recognition (OCR) technology to analyze the visual data and generate an item list. Possible OCR technologies used in this process include open-source Tesseract and commercial cloud-based services.

[0429] The server then uses an emotion analysis API to evaluate the user's emotional state based on the text or voice input. Along with this emotional data, it references past processing history and preference data to generate suggested options that reflect the emotional state. This part utilizes natural language processing models and machine learning techniques.

[0430] The generated proposal candidates are sent to an information input device and presented to the user in a visual format (e.g., images or videos). The user can select from the proposal candidates, and based on the selected candidate, the server automatically connects with an external food delivery service and executes the order or commission.

[0431] For example, if a user provides image data of "pasta, cheese, and tomatoes" and inputs a mood indicating they want to relax, the server will generate suggestions such as "creamy pasta" or "cheese and tomato salad" and enable them to order from nearby partner restaurants. Examples of prompts include "How are you feeling today?", "Please suggest some dishes to help me relax," and "What pasta dishes do you recommend?".

[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0433] Step 1:

[0434] The user uses an information input device to photograph food items in the refrigerator or receipts for purchased items. This captured image data becomes the input, and the terminal compresses this data as needed while maintaining high quality before sending it to the server. JPEG and PNG are common image data formats.

[0435] Step 2:

[0436] The server uses the received image data as input to extract ingredient and purchase information from the visual data using optical character recognition (OCR) technology. This process analyzes the textual information within the visual data and outputs an item list. OCR technologies used include Tesseract and cloud-based OCR APIs.

[0437] Step 3:

[0438] The user inputs their emotional state as voice or text through an information input device. This emotional data is sent to the server as input. The server uses an emotional analysis API to analyze the input emotional data and outputs the user's emotional state as a numerical value or category.

[0439] Step 4:

[0440] The server receives item lists, emotional states, past processing history, and preference data as input, and uses a generative AI model to generate suggested candidates that reflect the emotional state based on this data. In this process, the generative AI model integrates these multiple data sources to output the optimal suggested candidate.

[0441] Step 5:

[0442] The generated suggestion candidates are sent from the server to the terminal. The terminal receives these outputted suggestion candidates and presents them to the user in a visual format (e.g., images or videos). This makes it easier for the user to visually review the options.

[0443] Step 6:

[0444] The user selects from the suggested options presented on the terminal. This selection is sent as input to the server, which uses this data to place an order or outsource the order for the suggested dish to an external service. The server accesses partner delivery services through the interface and places the order automatically.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0448] [Third Embodiment]

[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0461] This invention is a system that efficiently utilizes ingredient and purchase information to suggest daily meals. This system employs a method where users can easily provide ingredient and purchase information using an image input terminal, and the system automatically determines the meal based on that information.

[0462] User behavior

[0463] Users first use a smartphone or tablet to take photos of the food in their refrigerator and receipts for purchased items. This provides the system with information about the food items they have available on a daily basis. They can also input photos of dishes they have made in the past and their past preferences.

[0464] Terminal processing

[0465] The device receives images taken by the user and first saves them. Then, it sends these images as data to the server. Communication is conducted using a secure protocol, ensuring user privacy.

[0466] Server-based processing

[0467] The server utilizes AI and optical character recognition (OCR) technologies to analyze image data received from the terminal. This extracts ingredient or purchase information from the image and organizes it as text data. Then, based on this organized information, an ingredient list is generated, and combined with past cooking history and preference data, it determines the menu for the day. This process also takes into account the user's nutritional balance, taste preferences, and even seasonal and special offer information.

[0468] Processing after the menu has been decided.

[0469] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. These instructions include not only text format but also a visually appealing video format, allowing users to easily follow along with the cooking process.

[0470] Specific example

[0471] For example, if a user takes a picture of "chicken, carrots, and onions" as ingredients with their device, the server analyzes the image and builds an ingredient list. Based on the user's preferences, it suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and provides the device with cooking instructions along with video links for each dish. The user can then choose one of these options and proceed with cooking while watching the video.

[0472] In this way, the present invention aims to streamline daily meal selection, minimize food waste in the home, and improve satisfaction with cooking.

[0473] The following describes the processing flow.

[0474] Step 1:

[0475] The user uses their smartphone or tablet to take pictures of food items in the refrigerator or receipts for purchased items. The device acquires this image data using a camera app and temporarily stores it within the app.

[0476] Step 2:

[0477] The device uploads stored image data to a server via a dedicated application. Encrypted communication is used for the upload, ensuring user privacy.

[0478] Step 3:

[0479] The server receives image data and performs image analysis using an AI model. Here, the AI ​​extracts text data about ingredients and purchased items, and reads product information from the receipt using optical character recognition (OCR).

[0480] Step 4:

[0481] Based on the text data extracted by the server, a list of available ingredients is generated in the database. This list is organized by date and ingredient category.

[0482] Step 5:

[0483] The server references past cooking history and preference data, and creates a list of dish suggestions based on the information associated with the ingredient list. Nutritional balance and seasonality are also taken into consideration during this process.

[0484] Step 6:

[0485] The server selects a candidate dish and generates detailed cooking instructions in text or video format. This process involves referencing open recipe databases and information from professional chefs for refinement.

[0486] Step 7:

[0487] The device receives a list of recipe suggestions from the server and notifies the user. The user can then select a recipe from the displayed suggestions and view the cooking instructions.

[0488] Step 8:

[0489] The system checks the necessary ingredients according to the dish selected by the user and begins cooking. Users can efficiently prepare their meals by referring to the cooking instructions provided on the device, which are available as videos and text.

[0490] (Example 1)

[0491] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0492] In modern times, managing ingredients and suggesting meals at home is time-consuming, laborious, and difficult to do efficiently. In particular, there are many challenges in making the most of the ingredients in the refrigerator and suggesting appropriate meals daily that consider individual preferences and nutritional balance. Furthermore, there is a need for cooking instructions to be presented in a detailed, easy-to-understand, and actionable format.

[0493] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0494] In this invention, the server includes means for receiving and analyzing images of items or purchase information captured by an information processing device to extract items or purchase information from the images, means for generating an item list based on the extracted information, means for determining dish candidates using past cooking history and preference data, means for ensuring communication security and protecting data, and means for extracting information using visual recognition technology. This makes it possible to efficiently and safely provide users with optimal dish suggestions and procedures.

[0495] An "information processing device" is an electronic device used by users to take pictures and process them as data.

[0496] "Items" refers to goods or food products that are photographed by an information processing device.

[0497] "Purchase information" refers to information about purchased goods, including receipts and electronic purchase records.

[0498] "Image analysis" is a data processing technique used to extract necessary information from captured image data.

[0499] "Extraction" refers to the process of selecting items and purchase information obtained through image analysis and extracting them as necessary data.

[0500] An "item list" refers to a systematically organized list based on extracted items and purchasing information.

[0501] "Cooking history" refers to records of dishes that have been cooked in the past and the ingredients used.

[0502] "Preference data" refers to information based on users' taste preferences and food choices.

[0503] "Cooking suggestions" refers to a list of possible dishes suggested based on the analyzed information.

[0504] "Communication security" means that information is protected from unauthorized access during data transmission and reception.

[0505] "Data protection" refers to measures taken to safeguard users' personal information and data related to their privacy.

[0506] "Visual recognition technology" is a technology that analyzes image data and enables machines to perform identification equivalent to that of humans.

[0507] This invention is a system that efficiently manages and utilizes goods and purchase information to provide optimal suggestions for users' daily lives. This system utilizes an information processing device to allow users to easily provide goods and purchase information, and then uses this information to determine the most suitable meal.

[0508] User behavior

[0509] Users use information processing devices such as smartphones or tablets to take photos of items in their refrigerator and receipts as purchase information. This operation allows them to provide the system with information about items they can use on a daily basis. They can also input information about meals they have cooked in the past and their personal preferences.

[0510] Terminal processing

[0511] The device receives the image data captured by the user and stores it temporarily. Then, to securely transmit this image data to the server, it uses a protocol that guarantees communication security (e.g., HTTPS).

[0512] Server-based processing

[0513] The server receives image data transmitted from the terminal and uses visual recognition and optical character recognition (OCR) technologies to extract items and purchase information from the image. The extracted data is organized into an item list. The server then compares this item list with past cooking history and preference data, and uses a generative AI model to determine the day's meal options. Nutritional balance, individual preferences, seasonal and special offer information are also taken into consideration.

[0514] Processing after the menu has been decided.

[0515] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. This information is available in text format as well as a visually appealing video format, which the user can use as a guide while cooking.

[0516] Examples of specific cases and prompt statements

[0517] For example, if a user takes a picture of "chicken, carrots, and onions" using their device, the server analyzes the image and builds an item list. Based on the user's preferences, it then suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and generates and displays video links to the cooking procedures for each dish on the device. The user can then choose one of these options and proceed with cooking while referring to the video. An example of a prompt message is, "Please tell me what I can make with the chicken, carrots, and onions I have in the refrigerator. If possible, please also tell me about dishes that are nutritionally balanced and provide detailed cooking instructions."

[0518] In this way, the system aims to improve user satisfaction by streamlining daily inventory management and recipe suggestions, minimizing waste.

[0519] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0520] Step 1:

[0521] The user uses an information processing device to photograph items and receipts inside the refrigerator. Image data is generated as input. This image data records the condition and quantity of the items, as well as purchase information.

[0522] Step 2:

[0523] The terminal temporarily stores image data obtained from the user and then prepares it for transmission to the server. At this stage, a secure transmission protocol is used to ensure the data's security. The output is a data package that can be securely transmitted.

[0524] Step 3:

[0525] The server processes image data received from the terminal. It receives image data as input and analyzes the data using visual recognition technology and optical character recognition (OCR). Specifically, it detects text information within the image and extracts it as item information. The output is data in the form of an item list.

[0526] Step 4:

[0527] The server determines dish candidates based on the generated item list, referencing past cooking history and preference data. This process uses a generative AI model to generate optimal dish suggestions. It uses the item list and preference data as input and creates a list of dish candidates as output.

[0528] Step 5:

[0529] The server creates detailed cooking instructions based on the suggested dishes and sends them to the terminal. The input is a list of suggested dishes, and the output is the cooking instructions for each dish along with a related video link. This allows the user to easily understand how to cook.

[0530] Step 6:

[0531] The user reviews the provided cooking instructions and proceeds with the cooking process accordingly. The user can smoothly complete the cooking process by referring to the videos and text instructions displayed on the device. The output is the finished dish.

[0532] (Application Example 1)

[0533] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0534] In modern times, preparing meals at home is often a burden due to busy daily lives. Furthermore, it's difficult to efficiently utilize available ingredients, reduce food waste, and provide appropriate menus. Therefore, there is a growing demand for systems that easily determine appropriate dishes based on ingredient information. Additionally, menu suggestions in food delivery services often lack sufficient utilization of ingredients already available at home, resulting in challenges in improving customer satisfaction.

[0535] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0536] In this invention, the server includes means for receiving images of information captured by the user using an image input device, analyzing the images to extract material or purchase information from the images, generating a list of materials based on the extracted information, determining a menu suggestion for the period using past processing history, preference data, and the material list, and suggesting selectable items based on the user's existing materials. This enables the user to make the most of the materials available at that time and prepare a meal efficiently and with high satisfaction. Furthermore, it enables flexible menu suggestions tailored to the user's needs in food delivery services.

[0537] A "user" is an individual or corporation that operates the system and inputs images of food ingredients and purchase information.

[0538] An "image input device" is a device used to input images of information captured by the user into a system, and includes smartphones and tablets.

[0539] "Ingredients" refer to the elements necessary for creating a dish, including food ingredients, seasonings, and other necessities.

[0540] "Purchase information" refers to data about products acquired by the user, including the date and time of purchase, product name, and quantity.

[0541] A "materials list" is a list that organizes and digitizes the extracted materials information.

[0542] "Processing history" refers to records of dishes a user has cooked in the past, and is data used to analyze specific patterns and trends.

[0543] "Preference data" refers to information that indicates a user's taste preferences and dietary preferences, and is used as a reference when suggesting dishes.

[0544] A "recipe suggestion" is a set of dishes or meal options determined based on an ingredient list and other relevant data.

[0545] "Offered items" refer to meals and goods that can be delivered, such as those offered by food delivery services.

[0546] "Demonstration format" refers to a means of visually showing the cooking process to the user, and is expressed through videos or illustrations.

[0547] To realize this invention, the user first uses an image input device such as a smartphone or tablet to take pictures of ingredients in the refrigerator or receipts for purchased items. This provides the system with information on ingredients that are readily available on a daily basis. The user can also input photos of dishes they have made in the past and data on their preferences.

[0548] The image input device prepares these images as data and sends them to the server using a secure protocol. This protocol, such as HTTPS, protects user privacy.

[0549] The server preprocesses images received from terminals using OpenCV, an open-source image processing library. Then, it uses AI and optical character recognition (OCR) technologies, such as TensorFlow and Tesseract, to extract material and purchase information from the images and organize it as text data.

[0550] Furthermore, the server generates a list of ingredients from the extracted information and combines it with past processing history and preference data to determine the optimal dish suggestion for that moment. Based on the suggested dish, the server generates detailed cooking instructions and sends them to the image input device. These cooking instructions are provided in both text and video formats.

[0551] For example, if a user enters "cheese, tomato, and salami" as ingredients, the system analyzes this and suggests dishes such as "homemade pizza" or "toast with lots of toppings." Based on these suggestions, it also presents specific pizza doughs to the user as part of a food delivery menu.

[0552] An example of a prompt for a generative AI model would be, "Based on the ingredients in the refrigerator—cheese, tomatoes, and salami—please suggest a menu of available pizza doughs." This allows for the provision of meals tailored to individual preferences and requests while minimizing food waste.

[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0554] Step 1:

[0555] The user uses an image input device to photograph food items in the refrigerator or receipts for purchased items. This generates image data. The captured images contain information about the available ingredients.

[0556] Step 2:

[0557] The device sends the captured image data to the server. The HTTPS protocol is used for communication, protecting user privacy. At this point, the input is the image data, and the output is the preparation for the server to receive it.

[0558] Step 3:

[0559] The server uses OpenCV to preprocess the received image data. This preprocessing involves denoising the image and extracting necessary regions, preparing the image for analysis. This results in high-quality images for analysis as output.

[0560] Step 4:

[0561] The server uses TensorFlow and Tesseract to extract material and purchase information from pre-processed images. AI model recognition and OCR technology are used to obtain information as text data from the images. The input for this step is a pre-processed image, and the output is text data containing material and purchase information.

[0562] Step 5:

[0563] The server creates an ingredient list based on extracted text data and determines a recipe suggestion by combining it with past processing history and preference data. A generative AI model is used to select a recipe that reflects the user's preferences. The ingredient list and preference data are inputs, and recipe suggestions are output.

[0564] Step 6:

[0565] Based on the selected dish suggestion, the server generates detailed cooking instructions and provides them to the terminal. The cooking instructions are generated in both text and video formats. Specific step-by-step data for realizing the suggested dish is provided as output.

[0566] Step 7:

[0567] Users review suggested dishes and cooking instructions provided through their terminal and order items via delivery service as needed. The terminal receives the suggestions, and the selected information becomes the input for the next action.

[0568] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0569] This invention provides a system that simplifies users' daily food management and meal selection, and further enhances personalized meal suggestions by incorporating an emotion engine. The system has the function of generating meal candidates based on ingredients and purchase information, and the function of recognizing the user's emotions and reflecting them in meal selection.

[0570] User behavior

[0571] The user first uses a smartphone or tablet to take photos of the food items in their refrigerator and receipts for purchased items. In addition, the user communicates their emotions to the device through their own input or an interactive user interface. As a result, data on both food items and emotions is provided to the system.

[0572] Terminal processing

[0573] The device receives images and emotional input data captured by the user and sends them to the server. During this process, image data is compressed if necessary while maintaining high quality before uploading. Emotional data is obtained through methods such as voice, text, and facial expression analysis.

[0574] Server-based processing

[0575] After receiving image data, the server performs image analysis using AI technology and optical character recognition (OCR). This extracts ingredient and purchase information. Next, the emotion engine analyzes the emotion data received from the user to determine the emotional state. Based on this data, it is combined with the user's past cooking history and preference data to generate menu suggestions for the day. By considering the emotional state, the menu selection becomes more tailored to the user's current state.

[0576] Processing after the menu has been decided.

[0577] The server selects a dish and sends it to the user's device in text or video format, along with detailed cooking instructions. The video format is visually easy to understand and serves as a practical cooking guide for the user.

[0578] Specific example

[0579] For example, if a user takes a photo of "cheese, tomatoes, and bread" as ingredients and inputs a happy and relaxed emotional state, the server immediately analyzes this. As a result, it suggests relaxing and enjoyable meal options such as "grilled cheese sandwich, tomato soup, and bruschetta." Cooking instructions for each dish are provided on the device, and the user can choose their desired dish and begin cooking.

[0580] In this way, the present invention aims to improve the cooking experience by considering both ingredient information and emotional state, thereby enabling optimal cooking suggestions for the user.

[0581] The following describes the processing flow.

[0582] Step 1:

[0583] The user uses a smartphone or tablet to take photos of food items in the refrigerator or receipts for purchased items. At the same time, the user inputs their emotions through the device's interface using simple questions and voice guidance.

[0584] Step 2:

[0585] The device transmits captured image data and entered emotion data to the server using a secure protocol. Image data is appropriately compressed, and emotion data is encrypted before transmission.

[0586] Step 3:

[0587] The server receives the image data and performs image analysis using AI technology and optical character recognition (OCR). This extracts information about ingredients and purchases, which are then organized into a database.

[0588] Step 4:

[0589] The server uses an emotion engine to analyze the emotion data sent by the user. This analysis determines the user's current emotional state (e.g., happy, stressed, anxious, etc.).

[0590] Step 5:

[0591] The server combines extracted ingredient lists, past cooking history, preference data, and emotional states to generate meal suggestions suitable for the user on that day. These suggestions take into account nutritional balance, the season, and even the user's emotional state.

[0592] Step 6:

[0593] The server generates detailed cooking instructions for the selected dish in both text and video formats and sends them to the terminal. The videos are visually easy to understand and show the actual cooking process step by step.

[0594] Step 7:

[0595] The terminal displays a list of dish suggestions received from the server to the user and notifies them. The user can then select their preferred dish from the displayed suggestions and view the detailed cooking instructions.

[0596] Step 8:

[0597] The user begins cooking the dish they have selected. Following the instructions provided on the device in text and video formats, they can efficiently proceed with the cooking process. The cooking experience can be enhanced by allowing users to enjoy cooking in a way that aligns with their emotions.

[0598] (Example 2)

[0599] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0600] In modern life, consumers spend a significant amount of time and effort managing their food and choosing meals. Furthermore, because users' emotional states are often overlooked in meal selection, the resulting meal choices frequently fail to align with their desired cooking experience. Therefore, it is necessary to improve the cooking experience by providing more personalized meal suggestions that take into account the user's food availability and emotional state.

[0601] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0602] In this invention, the server includes means for receiving images of ingredients or purchase information captured by the user using an image input device, and emotional state data; means for analyzing the images to extract ingredients and purchase information from the images; and means for analyzing the emotional state data to determine the user's emotional state. This makes it possible to generate optimal dish candidates that take both ingredient information and emotional state into consideration.

[0603] "Users" refer to individuals who operate this system and provide food ingredients and emotional data.

[0604] An "image input device" refers to an electronic device with a camera function that users use to collect information on ingredients and purchases.

[0605] "Emotional state data" refers to information that indicates the user's emotions and is entered into the device in text, audio, or other formats.

[0606] A "generative model" refers to artificial intelligence technology used to generate optimal dish candidates by integrating multiple pieces of information.

[0607] "Visual format" refers to video and image formats that display cooking instructions in a way that makes them easy for users to understand visually.

[0608] A "recipe suggestion" refers to a set of multiple dish options proposed based on the user's available ingredients and emotional state.

[0609] This invention is a system that facilitates users' daily food management and meal selection. Based on the user's food situation and emotional state, the system provides appropriate meal suggestions and primarily consists of an image input device, a terminal, and a server.

[0610] Users take photos of food items in their refrigerator and receipts for purchased items using a smartphone or tablet. They also input their emotional state into the device via text or voice. The food information and emotional state data obtained in this way are then transmitted to a server by the device. Image data is compressed as needed while maintaining high quality.

[0611] The server analyzes the received image data using AI technology and optical character recognition (OCR). This extracts ingredient names and purchase information. Meanwhile, emotional state data is analyzed by an emotion engine to determine the user's current emotional state. The server combines this two pieces of information and uses a generative model to generate optimal dish suggestions. The generative model receives a prompt message and provides new dish suggestions.

[0612] As a concrete example, consider a scenario where a user takes a photo of "cheese, tomatoes, and bread" and inputs a happy and relaxed emotional state. In this case, the server analyzes this and suggests dish options such as "grilled cheese sandwich, tomato soup, and bruschetta." The cooking instructions for the suggested dishes are displayed on the device in a visually easy-to-understand video format.

[0613] An example of a prompt for a generative AI model is: "The user has entered cheese, tomato, and bread as ingredients, and their emotions are happy and relaxed. Please suggest three dishes that meet these conditions."

[0614] Thus, the present invention aims to provide users with appropriate cooking suggestions based on both ingredient information and their emotional state.

[0615] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0616] Step 1:

[0617] The user takes photos of the food items and purchase information inside the refrigerator using an image input device. They also input their emotional state via text and voice through an interactive interface. This process generates food image data and emotional data.

[0618] Step 2:

[0619] The terminal processes food image data and emotion data received from the user. Captured images may be compressed for efficient data transfer while maintaining high quality. Emotion data is converted to a predefined data format and sent to the server. The output consists of compressed image data and formatted emotion data.

[0620] Step 3:

[0621] The server analyzes the image data received from the terminal using AI technology and optical character recognition (OCR). This extracts the names of ingredients and purchase information from the image as text. The input is compressed image data, and the output is text-based ingredient information.

[0622] Step 4:

[0623] The server simultaneously analyzes the received emotion data using an emotion engine to determine the user's emotional state. The input is formatted emotion data, and the output is the identified emotional state. Specifically, the emotion data is classified based on speech and specific word patterns.

[0624] Step 5:

[0625] The server uses a generative AI model to generate optimal dish candidates based on extracted ingredient information and determined emotional states. Prompt messages are input to the generative model, and newly suggested dish ideas are output. Specifically, an evaluation function operates by referencing past cooking history data.

[0626] Step 6:

[0627] The server determines the dish options and sends detailed cooking procedure data to the terminal in a visually easy-to-understand format. The terminal receives this data and presents it to the user as options. The input is the generated dish suggestions, and the output is the cooking procedure data presented to the user. Specific actions include displaying the information in text or video format.

[0628] (Application Example 2)

[0629] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0630] Modern consumers have a need to manage their meals efficiently and comfortably amidst their busy daily lives, but conventional technology has problems in that it does not adequately offer flexible meal suggestions tailored to emotions and situations, nor does it integrate well with external services. In particular, the challenge lies in simultaneously achieving the effective use of ingredients and providing services that are optimal for each individual user.

[0631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0632] In this invention, the server includes means for analyzing visual data and extracting item or purchase information from the visual data, means for generating suggested candidates that reflect the emotional state, and means for making requests from external services based on the suggested candidates. This enables the generation of personalized suggestions that correspond to the user's emotions and situation, and smooth access to external services based on those suggestions.

[0633] An "information input device" is a device used by users to input visual data of goods or purchase information, and includes portable terminals such as smartphones and tablets.

[0634] "Visual data" refers to image data of items and purchase information acquired using an information input device, which can be analyzed using optical character recognition technology.

[0635] The "item list" is a compilation of item information extracted through visual data analysis, and is used to generate subsequent proposal candidates.

[0636] "Past processing history" refers to a record of the choices and actions a user has taken in the past, and serves as the basis for providing personalized suggestions.

[0637] "Preference data" refers to data about users' preferences and tendencies, and is used to generate more appropriate suggestion candidates.

[0638] "Suggestion candidates that reflect emotional state" are recommendations generated based on the user's emotions and serve as information to present the user with multiple options.

[0639] "Means of outsourcing from external services" refers to a function that connects to external food supply services based on proposed options and according to the user's selection, and executes orders and outsourcing.

[0640] The system for implementing this invention allows users to easily manage food ingredients and purchase information, and provides optimal cooking suggestions and connections to external services based on their emotional state. This system includes an information input device such as a smartphone, a server, and a network configuration for connecting to external services.

[0641] First, users use an information input device to photograph food ingredients, receipts, and other purchase information, and send the image data to the server. The server then uses optical character recognition (OCR) technology to analyze the visual data and generate an item list. Possible OCR technologies used in this process include open-source Tesseract and commercial cloud-based services.

[0642] The server then uses an emotion analysis API to evaluate the user's emotional state based on the text or voice input. Along with this emotional data, it references past processing history and preference data to generate suggested options that reflect the emotional state. This part utilizes natural language processing models and machine learning techniques.

[0643] The generated proposal candidates are sent to an information input device and presented to the user in a visual format (e.g., images or videos). The user can select from the proposal candidates, and based on the selected candidate, the server automatically connects with an external food delivery service and executes the order or commission.

[0644] For example, if a user provides image data of "pasta, cheese, and tomatoes" and inputs a mood indicating they want to relax, the server will generate suggestions such as "creamy pasta" or "cheese and tomato salad" and enable them to order from nearby partner restaurants. Examples of prompts include "How are you feeling today?", "Please suggest some dishes to help me relax," and "What pasta dishes do you recommend?".

[0645] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0646] Step 1:

[0647] The user uses an information input device to photograph food items in the refrigerator or receipts for purchased items. This captured image data becomes the input, and the terminal compresses this data as needed while maintaining high quality before sending it to the server. JPEG and PNG are common image data formats.

[0648] Step 2:

[0649] The server uses the received image data as input to extract ingredient and purchase information from the visual data using optical character recognition (OCR) technology. This process analyzes the textual information within the visual data and outputs an item list. OCR technologies used include Tesseract and cloud-based OCR APIs.

[0650] Step 3:

[0651] The user inputs their emotional state as voice or text through an information input device. This emotional data is sent to the server as input. The server uses an emotional analysis API to analyze the input emotional data and outputs the user's emotional state as a numerical value or category.

[0652] Step 4:

[0653] The server receives item lists, emotional states, past processing history, and preference data as input, and uses a generative AI model to generate suggested candidates that reflect the emotional state based on this data. In this process, the generative AI model integrates these multiple data sources to output the optimal suggested candidate.

[0654] Step 5:

[0655] The generated suggestion candidates are sent from the server to the terminal. The terminal receives these outputted suggestion candidates and presents them to the user in a visual format (e.g., images or videos). This makes it easier for the user to visually review the options.

[0656] Step 6:

[0657] The user selects from the suggested options presented on the terminal. This selection is sent as input to the server, which uses this data to place an order or outsource the order for the suggested dish to an external service. The server accesses partner delivery services through the interface and places the order automatically.

[0658] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0659] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0660] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0661] [Fourth Embodiment]

[0662] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0663] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0664] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0665] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0666] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0667] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0668] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0669] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0670] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0671] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0672] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0673] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0674] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0675] This invention is a system that efficiently utilizes ingredient and purchase information to suggest daily meals. This system employs a method where users can easily provide ingredient and purchase information using an image input terminal, and the system automatically determines the meal based on that information.

[0676] User behavior

[0677] Users first use a smartphone or tablet to take photos of the food in their refrigerator and receipts for purchased items. This provides the system with information about the food items they have available on a daily basis. They can also input photos of dishes they have made in the past and their past preferences.

[0678] Terminal processing

[0679] The device receives images taken by the user and first saves them. Then, it sends these images as data to the server. Communication is conducted using a secure protocol, ensuring user privacy.

[0680] Server-based processing

[0681] The server utilizes AI and optical character recognition (OCR) technologies to analyze image data received from the terminal. This extracts ingredient or purchase information from the image and organizes it as text data. Then, based on this organized information, an ingredient list is generated, and combined with past cooking history and preference data, it determines the menu for the day. This process also takes into account the user's nutritional balance, taste preferences, and even seasonal and special offer information.

[0682] Processing after the menu has been decided.

[0683] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. These instructions include not only text format but also a visually appealing video format, allowing users to easily follow along with the cooking process.

[0684] Specific example

[0685] For example, if a user takes a picture of "chicken, carrots, and onions" as ingredients with their device, the server analyzes the image and builds an ingredient list. Based on the user's preferences, it suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and provides the device with cooking instructions along with video links for each dish. The user can then choose one of these options and proceed with cooking while watching the video.

[0686] In this way, the present invention aims to streamline daily meal selection, minimize food waste in the home, and improve satisfaction with cooking.

[0687] The following describes the processing flow.

[0688] Step 1:

[0689] The user uses their smartphone or tablet to take pictures of food items in the refrigerator or receipts for purchased items. The device acquires this image data using a camera app and temporarily stores it within the app.

[0690] Step 2:

[0691] The device uploads stored image data to a server via a dedicated application. Encrypted communication is used for the upload, ensuring user privacy.

[0692] Step 3:

[0693] The server receives image data and performs image analysis using an AI model. Here, the AI ​​extracts text data about ingredients and purchased items, and reads product information from the receipt using optical character recognition (OCR).

[0694] Step 4:

[0695] Based on the text data extracted by the server, a list of available ingredients is generated in the database. This list is organized by date and ingredient category.

[0696] Step 5:

[0697] The server references past cooking history and preference data, and creates a list of dish suggestions based on the information associated with the ingredient list. Nutritional balance and seasonality are also taken into consideration during this process.

[0698] Step 6:

[0699] The server selects a candidate dish and generates detailed cooking instructions in text or video format. This process involves referencing open recipe databases and information from professional chefs for refinement.

[0700] Step 7:

[0701] The device receives a list of recipe suggestions from the server and notifies the user. The user can then select a recipe from the displayed suggestions and view the cooking instructions.

[0702] Step 8:

[0703] The system checks the necessary ingredients according to the dish selected by the user and begins cooking. Users can efficiently prepare their meals by referring to the cooking instructions provided on the device, which are available as videos and text.

[0704] (Example 1)

[0705] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0706] In modern times, managing ingredients and suggesting meals at home is time-consuming, laborious, and difficult to do efficiently. In particular, there are many challenges in making the most of the ingredients in the refrigerator and suggesting appropriate meals daily that consider individual preferences and nutritional balance. Furthermore, there is a need for cooking instructions to be presented in a detailed, easy-to-understand, and actionable format.

[0707] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0708] In this invention, the server includes means for receiving and analyzing images of items or purchase information captured by an information processing device to extract items or purchase information from the images, means for generating an item list based on the extracted information, means for determining dish candidates using past cooking history and preference data, means for ensuring communication security and protecting data, and means for extracting information using visual recognition technology. This makes it possible to efficiently and safely provide users with optimal dish suggestions and procedures.

[0709] An "information processing device" is an electronic device used by users to take pictures and process them as data.

[0710] "Items" refers to goods or food products that are photographed by an information processing device.

[0711] "Purchase information" refers to information about purchased goods, including receipts and electronic purchase records.

[0712] "Image analysis" is a data processing technique used to extract necessary information from captured image data.

[0713] "Extraction" refers to the process of selecting items and purchase information obtained through image analysis and extracting them as necessary data.

[0714] An "item list" refers to a systematically organized list based on extracted items and purchasing information.

[0715] "Cooking history" refers to records of dishes that have been cooked in the past and the ingredients used.

[0716] "Preference data" refers to information based on users' taste preferences and food choices.

[0717] "Cooking suggestions" refers to a list of possible dishes suggested based on the analyzed information.

[0718] "Communication security" means that information is protected from unauthorized access during data transmission and reception.

[0719] "Data protection" refers to measures taken to safeguard users' personal information and data related to their privacy.

[0720] "Visual recognition technology" is a technology that analyzes image data and enables machines to perform identification equivalent to that of humans.

[0721] This invention is a system that efficiently manages and utilizes goods and purchase information to provide optimal suggestions for users' daily lives. This system utilizes an information processing device to allow users to easily provide goods and purchase information, and then uses this information to determine the most suitable meal.

[0722] User behavior

[0723] Users use information processing devices such as smartphones or tablets to take photos of items in their refrigerator and receipts as purchase information. This operation allows them to provide the system with information about items they can use on a daily basis. They can also input information about meals they have cooked in the past and their personal preferences.

[0724] Terminal processing

[0725] The device receives the image data captured by the user and stores it temporarily. Then, to securely transmit this image data to the server, it uses a protocol that guarantees communication security (e.g., HTTPS).

[0726] Server-based processing

[0727] The server receives image data transmitted from the terminal and uses visual recognition and optical character recognition (OCR) technologies to extract items and purchase information from the image. The extracted data is organized into an item list. The server then compares this item list with past cooking history and preference data, and uses a generative AI model to determine the day's meal options. Nutritional balance, individual preferences, seasonal and special offer information are also taken into consideration.

[0728] Processing after the menu has been decided.

[0729] Based on the selected dish, the server generates detailed cooking instructions and sends them to the terminal. This information is available in text format as well as a visually appealing video format, which the user can use as a guide while cooking.

[0730] Examples of specific cases and prompt statements

[0731] For example, if a user takes a picture of "chicken, carrots, and onions" using their device, the server analyzes the image and builds an item list. Based on the user's preferences, it then suggests dish options such as "chicken curry, stir-fried vegetables, and braised chicken," and generates and displays video links to the cooking procedures for each dish on the device. The user can then choose one of these options and proceed with cooking while referring to the video. An example of a prompt message is, "Please tell me what I can make with the chicken, carrots, and onions I have in the refrigerator. If possible, please also tell me about dishes that are nutritionally balanced and provide detailed cooking instructions."

[0732] In this way, the system aims to improve user satisfaction by streamlining daily inventory management and recipe suggestions, minimizing waste.

[0733] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0734] Step 1:

[0735] The user uses an information processing device to photograph items and receipts inside the refrigerator. Image data is generated as input. This image data records the condition and quantity of the items, as well as purchase information.

[0736] Step 2:

[0737] The terminal temporarily stores image data obtained from the user and then prepares it for transmission to the server. At this stage, a secure transmission protocol is used to ensure the data's security. The output is a data package that can be securely transmitted.

[0738] Step 3:

[0739] The server processes image data received from the terminal. It receives image data as input and analyzes the data using visual recognition technology and optical character recognition (OCR). Specifically, it detects text information within the image and extracts it as item information. The output is data in the form of an item list.

[0740] Step 4:

[0741] The server determines dish candidates based on the generated item list, referencing past cooking history and preference data. This process uses a generative AI model to generate optimal dish suggestions. It uses the item list and preference data as input and creates a list of dish candidates as output.

[0742] Step 5:

[0743] The server creates detailed cooking instructions based on the suggested dishes and sends them to the terminal. The input is a list of suggested dishes, and the output is the cooking instructions for each dish along with a related video link. This allows the user to easily understand how to cook.

[0744] Step 6:

[0745] The user reviews the provided cooking instructions and proceeds with the cooking process accordingly. The user can smoothly complete the cooking process by referring to the videos and text instructions displayed on the device. The output is the finished dish.

[0746] (Application Example 1)

[0747] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0748] In modern times, preparing meals at home is often a burden due to busy daily lives. Furthermore, it's difficult to efficiently utilize available ingredients, reduce food waste, and provide appropriate menus. Therefore, there is a growing demand for systems that easily determine appropriate dishes based on ingredient information. Additionally, menu suggestions in food delivery services often lack sufficient utilization of ingredients already available at home, resulting in challenges in improving customer satisfaction.

[0749] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0750] In this invention, the server includes means for receiving images of information captured by the user using an image input device, analyzing the images to extract material or purchase information from the images, generating a list of materials based on the extracted information, determining a menu suggestion for the period using past processing history, preference data, and the material list, and suggesting selectable items based on the user's existing materials. This enables the user to make the most of the materials available at that time and prepare a meal efficiently and with high satisfaction. Furthermore, it enables flexible menu suggestions tailored to the user's needs in food delivery services.

[0751] A "user" is an individual or corporation that operates the system and inputs images of food ingredients and purchase information.

[0752] An "image input device" is a device used to input images of information captured by the user into a system, and includes smartphones and tablets.

[0753] "Ingredients" refer to the elements necessary for creating a dish, including food ingredients, seasonings, and other necessities.

[0754] "Purchase information" refers to data about products acquired by the user, including the date and time of purchase, product name, and quantity.

[0755] A "materials list" is a list that organizes and digitizes the extracted materials information.

[0756] "Processing history" refers to records of dishes a user has cooked in the past, and is data used to analyze specific patterns and trends.

[0757] "Preference data" refers to information that indicates a user's taste preferences and dietary preferences, and is used as a reference when suggesting dishes.

[0758] A "recipe suggestion" is a set of dishes or meal options determined based on an ingredient list and other relevant data.

[0759] "Offered items" refer to meals and goods that can be delivered, such as those offered by food delivery services.

[0760] "Demonstration format" refers to a means of visually showing the cooking process to the user, and is expressed through videos or illustrations.

[0761] To realize this invention, the user first uses an image input device such as a smartphone or tablet to take pictures of ingredients in the refrigerator or receipts for purchased items. This provides the system with information on ingredients that are readily available on a daily basis. The user can also input photos of dishes they have made in the past and data on their preferences.

[0762] The image input device prepares these images as data and sends them to the server using a secure protocol. This protocol, such as HTTPS, protects user privacy.

[0763] The server preprocesses images received from terminals using OpenCV, an open-source image processing library. Then, it uses AI and optical character recognition (OCR) technologies, such as TensorFlow and Tesseract, to extract material and purchase information from the images and organize it as text data.

[0764] Furthermore, the server generates a list of ingredients from the extracted information and combines it with past processing history and preference data to determine the optimal dish suggestion for that moment. Based on the suggested dish, the server generates detailed cooking instructions and sends them to the image input device. These cooking instructions are provided in both text and video formats.

[0765] For example, if a user enters "cheese, tomato, and salami" as ingredients, the system analyzes this and suggests dishes such as "homemade pizza" or "toast with lots of toppings." Based on these suggestions, it also presents specific pizza doughs to the user as part of a food delivery menu.

[0766] An example of a prompt for a generative AI model would be, "Based on the ingredients in the refrigerator—cheese, tomatoes, and salami—please suggest a menu of available pizza doughs." This allows for the provision of meals tailored to individual preferences and requests while minimizing food waste.

[0767] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0768] Step 1:

[0769] The user uses an image input device to photograph food items in the refrigerator or receipts for purchased items. This generates image data. The captured images contain information about the available ingredients.

[0770] Step 2:

[0771] The device sends the captured image data to the server. The HTTPS protocol is used for communication, protecting user privacy. At this point, the input is the image data, and the output is the preparation for the server to receive it.

[0772] Step 3:

[0773] The server uses OpenCV to preprocess the received image data. This preprocessing involves denoising the image and extracting necessary regions, preparing the image for analysis. This results in high-quality images for analysis as output.

[0774] Step 4:

[0775] The server uses TensorFlow and Tesseract to extract material and purchase information from pre-processed images. AI model recognition and OCR technology are used to obtain information as text data from the images. The input for this step is a pre-processed image, and the output is text data containing material and purchase information.

[0776] Step 5:

[0777] The server creates an ingredient list based on extracted text data and determines a recipe suggestion by combining it with past processing history and preference data. A generative AI model is used to select a recipe that reflects the user's preferences. The ingredient list and preference data are inputs, and recipe suggestions are output.

[0778] Step 6:

[0779] Based on the selected dish suggestion, the server generates detailed cooking instructions and provides them to the terminal. The cooking instructions are generated in both text and video formats. Specific step-by-step data for realizing the suggested dish is provided as output.

[0780] Step 7:

[0781] Users review suggested dishes and cooking instructions provided through their terminal and order items via delivery service as needed. The terminal receives the suggestions, and the selected information becomes the input for the next action.

[0782] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0783] This invention provides a system that simplifies users' daily food management and meal selection, and further enhances personalized meal suggestions by incorporating an emotion engine. The system has the function of generating meal candidates based on ingredients and purchase information, and the function of recognizing the user's emotions and reflecting them in meal selection.

[0784] User behavior

[0785] The user first uses a smartphone or tablet to take photos of the food items in their refrigerator and receipts for purchased items. In addition, the user communicates their emotions to the device through their own input or an interactive user interface. As a result, data on both food items and emotions is provided to the system.

[0786] Terminal processing

[0787] The device receives images and emotional input data captured by the user and sends them to the server. During this process, image data is compressed if necessary while maintaining high quality before uploading. Emotional data is obtained through methods such as voice, text, and facial expression analysis.

[0788] Server-based processing

[0789] After receiving image data, the server performs image analysis using AI technology and optical character recognition (OCR). This extracts ingredient and purchase information. Next, the emotion engine analyzes the emotion data received from the user to determine the emotional state. Based on this data, it is combined with the user's past cooking history and preference data to generate menu suggestions for the day. By considering the emotional state, the menu selection becomes more tailored to the user's current state.

[0790] Processing after the menu has been decided.

[0791] The server selects a dish and sends it to the user's device in text or video format, along with detailed cooking instructions. The video format is visually easy to understand and serves as a practical cooking guide for the user.

[0792] Specific example

[0793] For example, if a user takes a photo of "cheese, tomatoes, and bread" as ingredients and inputs a happy and relaxed emotional state, the server immediately analyzes this. As a result, it suggests relaxing and enjoyable meal options such as "grilled cheese sandwich, tomato soup, and bruschetta." Cooking instructions for each dish are provided on the device, and the user can choose their desired dish and begin cooking.

[0794] In this way, the present invention aims to improve the cooking experience by considering both ingredient information and emotional state, thereby enabling optimal cooking suggestions for the user.

[0795] The following describes the processing flow.

[0796] Step 1:

[0797] The user uses a smartphone or tablet to take photos of food items in the refrigerator or receipts for purchased items. At the same time, the user inputs their emotions through the device's interface using simple questions and voice guidance.

[0798] Step 2:

[0799] The device transmits captured image data and entered emotion data to the server using a secure protocol. Image data is appropriately compressed, and emotion data is encrypted before transmission.

[0800] Step 3:

[0801] The server receives the image data and performs image analysis using AI technology and optical character recognition (OCR). This extracts information about ingredients and purchases, which are then organized into a database.

[0802] Step 4:

[0803] The server uses an emotion engine to analyze the emotion data sent by the user. This analysis determines the user's current emotional state (e.g., happy, stressed, anxious, etc.).

[0804] Step 5:

[0805] The server combines extracted ingredient lists, past cooking history, preference data, and emotional states to generate meal suggestions suitable for the user on that day. These suggestions take into account nutritional balance, the season, and even the user's emotional state.

[0806] Step 6:

[0807] The server generates detailed cooking instructions for the selected dish in both text and video formats and sends them to the terminal. The videos are visually easy to understand and show the actual cooking process step by step.

[0808] Step 7:

[0809] The terminal displays a list of dish suggestions received from the server to the user and notifies them. The user can then select their preferred dish from the displayed suggestions and view the detailed cooking instructions.

[0810] Step 8:

[0811] The user begins cooking the dish they have selected. Following the instructions provided on the device in text and video formats, they can efficiently proceed with the cooking process. The cooking experience can be enhanced by allowing users to enjoy cooking in a way that aligns with their emotions.

[0812] (Example 2)

[0813] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0814] In modern life, consumers spend a significant amount of time and effort managing their food and choosing meals. Furthermore, because users' emotional states are often overlooked in meal selection, the resulting meal choices frequently fail to align with their desired cooking experience. Therefore, it is necessary to improve the cooking experience by providing more personalized meal suggestions that take into account the user's food availability and emotional state.

[0815] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0816] In this invention, the server includes means for receiving images of ingredients or purchase information captured by the user using an image input device, and emotional state data; means for analyzing the images to extract ingredients and purchase information from the images; and means for analyzing the emotional state data to determine the user's emotional state. This makes it possible to generate optimal dish candidates that take into account both ingredient information and emotional state.

[0817] "Users" refer to individuals who operate this system and provide food ingredients and emotional data.

[0818] An "image input device" refers to an electronic device with a camera function that users use to collect information on ingredients and purchases.

[0819] "Emotional state data" refers to information that indicates the user's emotions and is entered into the device in text, audio, or other formats.

[0820] A "generative model" refers to artificial intelligence technology used to generate optimal dish candidates by integrating multiple pieces of information.

[0821] "Visual format" refers to video and image formats that display cooking instructions in a way that makes them easy for users to understand visually.

[0822] A "recipe suggestion" refers to a set of multiple dish options proposed based on the user's available ingredients and emotional state.

[0823] This invention is a system that facilitates users' daily food management and meal selection. Based on the user's food situation and emotional state, the system provides appropriate meal suggestions and primarily consists of an image input device, a terminal, and a server.

[0824] Users take photos of food items in their refrigerator and receipts for purchased items using a smartphone or tablet. They also input their emotional state into the device via text or voice. The food information and emotional state data obtained in this way are then transmitted to a server by the device. Image data is compressed as needed while maintaining high quality.

[0825] The server analyzes the received image data using AI technology and optical character recognition (OCR). This extracts ingredient names and purchase information. Meanwhile, emotional state data is analyzed by an emotion engine to determine the user's current emotional state. The server combines this two pieces of information and uses a generative model to generate optimal dish suggestions. The generative model receives a prompt message and provides new dish suggestions.

[0826] As a concrete example, consider a scenario where a user takes a photo of "cheese, tomatoes, and bread" and inputs a happy and relaxed emotional state. In this case, the server analyzes this and suggests dish options such as "grilled cheese sandwich, tomato soup, and bruschetta." The cooking instructions for the suggested dishes are displayed on the device in a visually easy-to-understand video format.

[0827] An example of a prompt for a generative AI model is: "The user has entered cheese, tomato, and bread as ingredients, and their emotions are happy and relaxed. Please suggest three dishes that meet these conditions."

[0828] Thus, the present invention aims to provide users with appropriate cooking suggestions based on both ingredient information and their emotional state.

[0829] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0830] Step 1:

[0831] The user takes photos of the food items and purchase information inside the refrigerator using an image input device. They also input their emotional state via text and voice through an interactive interface. This process generates food image data and emotional data.

[0832] Step 2:

[0833] The terminal processes food image data and emotion data received from the user. Captured images may be compressed for efficient data transfer while maintaining high quality. Emotion data is converted to a predefined data format and sent to the server. The output consists of compressed image data and formatted emotion data.

[0834] Step 3:

[0835] The server analyzes the image data received from the terminal using AI technology and optical character recognition (OCR). This extracts the names of ingredients and purchase information from the image as text. The input is compressed image data, and the output is text-based ingredient information.

[0836] Step 4:

[0837] The server simultaneously analyzes the received emotion data using an emotion engine to determine the user's emotional state. The input is formatted emotion data, and the output is the identified emotional state. Specifically, the emotion data is classified based on speech and specific word patterns.

[0838] Step 5:

[0839] The server uses a generative AI model to generate optimal dish candidates based on extracted ingredient information and determined emotional states. Prompt messages are input to the generative model, and newly suggested dish ideas are output. Specifically, an evaluation function operates by referencing past cooking history data.

[0840] Step 6:

[0841] The server determines the dish options and sends detailed cooking procedure data to the terminal in a visually easy-to-understand format. The terminal receives this data and presents it to the user as options. The input is the generated dish suggestions, and the output is the cooking procedure data presented to the user. Specific actions include displaying the information in text or video format.

[0842] (Application Example 2)

[0843] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0844] Modern consumers have a need to manage their meals efficiently and comfortably amidst their busy daily lives, but conventional technology has problems in that it does not adequately offer flexible meal suggestions tailored to emotions and situations, nor does it integrate well with external services. In particular, the challenge lies in simultaneously achieving the effective use of ingredients and providing services that are optimal for each individual user.

[0845] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0846] In this invention, the server includes means for analyzing visual data and extracting item or purchase information from the visual data, means for generating suggested candidates that reflect the emotional state, and means for making requests from external services based on the suggested candidates. This enables the generation of personalized suggestions that correspond to the user's emotions and situation, and smooth access to external services based on those suggestions.

[0847] An "information input device" is a device used by users to input visual data of goods or purchase information, and includes portable terminals such as smartphones and tablets.

[0848] "Visual data" refers to image data of items and purchase information acquired using an information input device, which can be analyzed using optical character recognition technology.

[0849] The "item list" is a compilation of item information extracted through visual data analysis, and is used to generate subsequent proposal candidates.

[0850] "Past processing history" refers to a record of the choices and actions a user has taken in the past, and serves as the basis for providing personalized suggestions.

[0851] "Preference data" refers to data about users' preferences and tendencies, and is used to generate more appropriate suggestion candidates.

[0852] "Suggestion candidates that reflect emotional state" are recommendations generated based on the user's emotions and serve as information to present the user with multiple options.

[0853] "Means of outsourcing from external services" refers to a function that connects to external food supply services based on proposed options and according to the user's selection, and executes orders and outsourcing.

[0854] The system for implementing this invention allows users to easily manage food ingredients and purchase information, and provides optimal cooking suggestions and connections to external services based on their emotional state. This system includes an information input device such as a smartphone, a server, and a network configuration for connecting to external services.

[0855] First, users use an information input device to photograph food ingredients, receipts, and other purchase information, and send the image data to the server. The server then uses optical character recognition (OCR) technology to analyze the visual data and generate an item list. Possible OCR technologies used in this process include open-source Tesseract and commercial cloud-based services.

[0856] The server then uses an emotion analysis API to evaluate the user's emotional state based on the text or voice input. Along with this emotional data, it references past processing history and preference data to generate suggested options that reflect the emotional state. This part utilizes natural language processing models and machine learning techniques.

[0857] The generated proposal candidates are sent to an information input device and presented to the user in a visual format (e.g., images or videos). The user can select from the proposal candidates, and based on the selected candidate, the server automatically connects with an external food delivery service and executes the order or commission.

[0858] For example, if a user provides image data of "pasta, cheese, and tomatoes" and inputs a mood indicating they want to relax, the server will generate suggestions such as "creamy pasta" or "cheese and tomato salad" and enable them to order from nearby partner restaurants. Examples of prompts include "How are you feeling today?", "Please suggest some dishes to help me relax," and "What pasta dishes do you recommend?".

[0859] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0860] Step 1:

[0861] The user uses an information input device to photograph food items in the refrigerator or receipts for purchased items. This captured image data becomes the input, and the terminal compresses this data as needed while maintaining high quality before sending it to the server. JPEG and PNG are common image data formats.

[0862] Step 2:

[0863] The server uses the received image data as input to extract ingredient and purchase information from the visual data using optical character recognition (OCR) technology. This process analyzes the textual information within the visual data and outputs an item list. OCR technologies used include Tesseract and cloud-based OCR APIs.

[0864] Step 3:

[0865] The user inputs their emotional state as voice or text through an information input device. This emotional data is sent to the server as input. The server uses an emotional analysis API to analyze the input emotional data and outputs the user's emotional state as a numerical value or category.

[0866] Step 4:

[0867] The server receives item lists, emotional states, past processing history, and preference data as input, and uses a generative AI model to generate suggested candidates that reflect the emotional state based on this data. In this process, the generative AI model integrates these multiple data sources to output the optimal suggested candidate.

[0868] Step 5:

[0869] The generated suggestion candidates are sent from the server to the terminal. The terminal receives these outputted suggestion candidates and presents them to the user in a visual format (e.g., images or videos). This makes it easier for the user to visually review the options.

[0870] Step 6:

[0871] The user selects from the suggested options presented on the terminal. This selection is sent as input to the server, which uses this data to place an order or outsource the order for the suggested dish to an external service. The server accesses partner delivery services through the interface and places the order automatically.

[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0874] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0893] The following is further disclosed regarding the embodiments described above.

[0894] (Claim 1)

[0895] The system receives images of food ingredients or purchase information taken by the user using an image input device.

[0896] A means for analyzing the image and extracting food ingredients or purchase information from the image,

[0897] A means for generating a list of ingredients based on the extracted information,

[0898] A means for determining the menu candidates for the day using past cooking history and preference data and the aforementioned list of ingredients,

[0899] A means of providing users with multiple cooking procedure data related to the candidate dish,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, wherein the cooking procedure data includes a video format.

[0903] (Claim 3)

[0904] The system according to claim 1, further comprising means for configuring a plurality of recipe options that can be selected by the user.

[0905] "Example 1"

[0906] (Claim 1)

[0907] The system receives images of items or purchase information taken by the user using an information processing device.

[0908] Means for analyzing the image and extracting information about items or purchases within the image,

[0909] A means for generating an item list based on the extracted information,

[0910] A means for determining the menu candidates for the day using past cooking history and preference data and the aforementioned list of items,

[0911] A means of providing users with multiple cooking procedure data related to the candidate dish,

[0912] A means of protecting data using methods that guarantee the security of communications,

[0913] A means of extracting information using visual recognition technology

[0914] A system that includes this.

[0915] (Claim 2)

[0916] The system according to claim 1, wherein the cooking procedure data includes a video format.

[0917] (Claim 3)

[0918] The system according to claim 1, further comprising means for composing a plurality of dish suggestions that the user can select.

[0919] "Application Example 1"

[0920] (Claim 1)

[0921] The system receives images of information captured by the user using an image input device.

[0922] A means for analyzing the image and extracting material or purchase information from the image,

[0923] A means for generating a list of materials based on the extracted information,

[0924] A means for determining a dish suggestion for the period using past processing history and preference data and the aforementioned list of ingredients,

[0925] A means of providing users with multiple cooking process data related to the proposed dish,

[0926] A means of proposing selectable products based on the user's existing materials,

[0927] A system that includes this.

[0928] (Claim 2)

[0929] The system according to claim 1, wherein the cooking process data includes a demonstration format.

[0930] (Claim 3)

[0931] The system according to claim 1, further comprising means for configuring a plurality of delivery method options that can be selected by the user.

[0932] "Example 2 of combining an emotion engine"

[0933] (Claim 1)

[0934] A means for receiving images of food ingredients or purchase information taken by the user using an image input device, and data on emotional state,

[0935] A means for analyzing the image and extracting material and purchase information from the image,

[0936] A means for analyzing the emotional state data to determine the user's emotional state,

[0937] A means for generating appropriate dish candidates using a generative model based on the extracted information and the determined emotional state,

[0938] A means of providing users with multiple cooking procedure data related to the candidate dish,

[0939] A system that includes this.

[0940] (Claim 2)

[0941] The system according to claim 1, wherein the cooking procedure data includes a visual format.

[0942] (Claim 3)

[0943] The system according to claim 1, further comprising means for composing a plurality of recipes that can be selected by the user.

[0944] "Application example 2 when combining with an emotional engine"

[0945] (Claim 1)

[0946] The system receives visual data of items or purchase information photographed by the user using an information input device.

[0947] A means for analyzing the visual data and extracting information about items or purchases within the visual data,

[0948] A means for generating an item list based on the extracted information,

[0949] A means for determining proposed candidates for the relevant day using past processing history and preference data and the aforementioned list of items,

[0950] A means for generating suggested candidates that reflect the emotional state of the user,

[0951] Based on the proposed solution, a means of outsourcing from external services,

[0952] A system that includes this.

[0953] (Claim 2)

[0954] The proposed system according to claim 1, wherein the proposed candidate includes a visual format.

[0955] (Claim 3)

[0956] The system according to claim 1, further comprising means for configuring a plurality of service options that can be selected by the user. [Explanation of Symbols]

[0957] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. The system receives images of food ingredients or purchase information taken by the user using an image input device. A means for analyzing the image and extracting ingredients or purchase information from the image, A means for generating a list of ingredients based on the extracted information, A means for determining the menu candidates for the day using past cooking history and preference data and the aforementioned list of ingredients, A means of providing users with multiple cooking procedure data related to the candidate dish, A system that includes this.

2. The system according to claim 1, wherein the cooking procedure data includes a video format.

3. The system according to claim 1, further comprising means for configuring a plurality of recipe options that can be selected by the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A