System
The cooking assistance system addresses the challenge of determining cooking methods and timing by using a temperature sensor, camera, and analysis means to provide real-time advice, enhancing cooking accuracy through a learning mechanism.
Patent Information
- Application Number
- JP2024137300
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Conventional cooking equipment fails to provide sufficient means for determining appropriate cooking methods and timing, especially for beginners, leading to mistakes in heat and cooking times, resulting in failed dishes.
A cooking assistance system equipped with a temperature sensor to detect ingredient state, a camera to capture images, and an analysis means to evaluate doneness and moisture content, providing real-time cooking advice through a notification system, and incorporating a learning mechanism to improve accuracy based on user feedback.
Enables users, including beginners, to prepare delicious meals accurately by providing real-time cooking guidance that improves with user feedback, ensuring proper cooking results.
Smart Images

Figure 2026034179000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When cooking, many people make mistakes with heat and cooking times, resulting in failed dishes. It is particularly difficult for beginners and those who are not good at cooking to determine the appropriate cooking method and timing. Conventional cooking equipment does not provide sufficient means to solve these problems, so there is a need to provide an environment where anyone can easily prepare delicious meals. [Means for solving the problem]
[0005] The present invention provides an analysis means that includes a temperature sensor that detects the state of ingredients and a camera that captures images of the ingredients, and that analyzes the data obtained from these to evaluate the doneness and moisture content of the ingredients. It also includes a notification means that provides advice on cooking methods and seasonings in natural language based on the analysis results. This allows users to receive accurate cooking advice in real time, creating an environment where mistakes are less likely to occur. It also provides support for creating delicious dishes with even greater accuracy by including a function that acquires ingredient information and recipe information from the user and determines cooking procedures based on this information, as well as a learning means that improves the accuracy of future cooking advice based on user feedback.
[0006] A "temperature sensor" is a device that measures the temperature of ingredients in real time.
[0007] A "camera" is a device for taking images of ingredients.
[0008] The "analysis means for analyzing data to evaluate the doneness and moisture content of ingredients" is a system that processes data obtained from the temperature sensor and camera to determine whether the ingredients are properly cooked.
[0009] The "notification means for providing cooking method and seasoning advice in natural language" is a system that provides cooking instructions and advice to the user in voice or text based on the information obtained from the analysis means.
[0010] "Ingredient information and recipe information acquired from the user" refers to information about the types of ingredients used and cooking procedures that are input by the user when cooking.
[0011] The "learning means" is a system for improving the accuracy of the analysis means based on feedback from users. [Brief explanation of the drawings]
[0012] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0020] [First embodiment]
[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0033] This invention is a cooking assistance system that allows even beginners to easily prepare delicious dishes. This system detects the condition of ingredients in real time and provides appropriate cooking advice. The operation of this system is described in detail below.
[0034] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for obtaining ingredient and recipe information from the user, and a learning means that improves the accuracy of the analysis means based on user feedback.
[0035] Overview of program processing
[0036] System startup and initialization
[0037] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and AI model.
[0038] The device waits for a voice command from the user and displays "Ready. Give me instructions to start cooking."
[0039] Start a session with the user
[0040] The user inputs "I'm going to start cooking" into the terminal by voice.
[0041] The device will ask the user aloud, "Please tell us the name of the dish you would like to cook."
[0042] The user types "steak" and the terminal sends it to the server.
[0043] Enter ingredients and recipe
[0044] The server sends a message to the terminal asking, "Please tell us the type and thickness of meat you would like to use."
[0045] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0046] The user types in "sirloin, 2 cm thick," and the device sends this to the server.
[0047] Sensor data collection and analysis
[0048] The device starts collecting data from the temperature sensor and camera and sends the data to the server.
[0049] The server analyzes the temperature and image data collected in real time to evaluate the doneness and moisture content of the ingredients.
[0050] For example, if it determines that "one side is still raw," the device will notify the user, "One side of the meat is still raw. Please continue cooking."
[0051] Providing advice and coordination
[0052] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal.
[0053] For example, the system may notify the user that "One side is cooked, please turn the meat over," and the user should follow the instructions to turn the meat over.
[0054] The device continues to collect data from the sensors and the server re-analyzes the new data.
[0055] Cooking completion and feedback
[0056] The server determines that the steak is cooked through and sends the message "The steak is done" to the terminal.
[0057] The device notifies the user that "your steak is done."
[0058] The user inputs feedback such as "The steak turned out delicious" into the terminal, and the terminal transmits the feedback to the server.
[0059] The server records user feedback and uses it to improve the accuracy of the analysis method.
[0060] In this way, the cooking assistance system of the present invention analyzes data from temperature sensors and cameras to provide support that allows anyone to easily cook like a pro. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] The server starts the system and establishes the connection between the temperature sensor and the camera, so that the sensor and camera are ready to operate normally.
[0064] Step 2:
[0065] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0066] Step 3:
[0067] The user inputs "I'm going to start cooking" into the terminal by voice.
[0068] Step 4:
[0069] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0070] Step 5:
[0071] The user speaks "steak," which the device recognizes and sends to the server.
[0072] Step 6:
[0073] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0074] Step 7:
[0075] The device asks the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0076] Step 8:
[0077] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0078] Step 9:
[0079] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0080] Step 10:
[0081] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0082] Step 11:
[0083] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking."
[0084] Step 12:
[0085] The device continues to collect data from the temperature sensor and camera.
[0086] Step 13:
[0087] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0088] Step 14:
[0089] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0090] Step 15:
[0091] The user performs the action of turning the meat over.
[0092] Step 16:
[0093] The device continues to collect data and send it to the server.
[0094] Step 17:
[0095] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0096] Step 18:
[0097] The device will notify the user by voice, "Your steak is ready."
[0098] Step 19:
[0099] The user provides voice feedback to the device, such as "The steak turned out delicious."
[0100] Step 20:
[0101] The device sends the provided feedback to the server.
[0102] Step 21:
[0103] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[0104] These are the specific processing steps of the AI-powered cooking assistance system, which enables users to easily prepare delicious meals.
[0105] Example 1
[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0107] Conventional cooking assistance systems have struggled to properly monitor the condition of ingredients and provide real-time advice based on the cooking progress. In particular, accurately measuring the doneness and moisture content of ingredients and providing cooking instructions at the appropriate time have been challenging. Furthermore, there has been a lack of technology to improve the accuracy of the system based on user feedback, leading to a demand for a system that allows even beginners to easily prepare delicious dishes.
[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0109] In this invention, the server includes a temperature sensor that detects the state of the ingredients, a camera that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, notification means that provides advice on cooking methods and seasonings in natural language based on the analysis means, input means that accepts voice commands from the user, means that loads a cooking database and a generative AI model and analyzes based on data from the analysis means, and learning means that improves the accuracy of the analysis means based on feedback from the user. This makes it easy for anyone to cook like a professional.
[0110] A "temperature sensor" is a device that measures the surface and internal temperature of ingredients in real time and provides that data to the system.
[0111] The "camera" is a photographing device that captures images of ingredients and provides the visual data to the system.
[0112] The "analysis means" is a collection of algorithms and programs for evaluating the doneness and moisture content of ingredients based on data obtained from temperature sensors and cameras.
[0113] The "notification means" is a function for conveying the cooking method and seasoning advice obtained by the analysis means to the user in natural language.
[0114] "Input means" refers to an interface through which a user inputs voice commands into the system.
[0115] A "cooking database" is a data storage device that accumulates data and recipes related to various dishes.
[0116] A "generative AI model" is an artificial intelligence algorithm that uses collected data and past feedback to assist analytical methods and optimize cooking instructions.
[0117] The "learning means" is a function for improving the accuracy of the analysis means based on feedback from users.
[0118] "Ingredient information" is data such as the type, amount, and shape of ingredients used in cooking.
[0119] "Recipe information" refers to information such as the steps to make a particular dish, the ingredients needed, and cooking time.
[0120] A "user" is a person who operates the system and cooks according to the advice and instructions provided.
[0121] A "server" is a central processing unit that controls and analyzes the entire system.
[0122] A "terminal" is a device that provides an interface with a user and displays voice commands and notifications.
[0123] This invention is a cooking assistance system that allows even beginners to easily prepare delicious meals. The main components of the system include a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means for analyzing the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, input means for acquiring ingredient information and recipe information from the user, and learning means that improves the accuracy of the analysis means based on user feedback.
[0124] Hardware and Software Configuration
[0125] The server controls the entire system and receives data from temperature sensors (e.g., digital temperature sensors) and cameras (e.g., high-resolution cameras). The server loads a cooking database and generative AI models and uses them to analyze the data in real time. Specific processes performed by the server include collecting temperature data, analyzing image data, and evaluating the doneness and moisture content of ingredients.
[0126] A terminal is a device that provides an interface with the user and waits for voice commands. For example, it is used as a smart speaker or tablet. The terminal receives voice commands from the user and sends them to the server. It also has the function of notifying the user of advice from the server.
[0127] The user is the person who operates the system and proceeds with the cooking. The user inputs voice commands into the terminal and cooks according to the instructions from the system. Furthermore, the user provides feedback to the system after cooking is completed.
[0128] Specific examples
[0129] For example, if a user wants to cook a steak, the following series of processes takes place: The user speaks "I'd like to start cooking" into the device, and the device sends a request to the server. The server then asks the user via the device what dish they want to cook. The user answers "steak," and based on that, the server requests detailed information about the ingredients. The user answers "sirloin, 2 cm thick," and the device sends this to the server.
[0130] The temperature sensor and camera collect the steak's temperature and image data, which are then sent to a server. The server analyzes this data in real time to assess the steak's doneness and moisture content. Based on the results, the server provides advice to the user, such as "One side is still raw. Continue cooking."
[0131] Prompt Sentence Examples
[0132] "I want to grill some sirloin steak. What's the best way to check the doneness and make sure it's delicious?"
[0133] This process allows users to receive accurate cooking advice, enabling even beginners to cook like professionals. Furthermore, as the system continues to learn based on user feedback, it provides increasingly accurate advice.
[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0135] Step 1: Boot and Initialize the System
[0136] The server starts the system. As part of the initialization, the server connects the temperature sensor and camera and checks their operation. Specifically, it acquires test data from the sensors and camera and verifies that they are working properly. The server also loads the cooking database and generative AI model, which will serve as the basis for future data analysis and advice provision.
[0137] Input: None
[0138] Output: System initialization complete, temperature sensor and camera operation confirmation, cooking database and AI model loading status
[0139] Step 2: Start a session with the user
[0140] The user speaks "I'm going to start cooking" into the device. The device recognizes this speech and sends a request to the server. The server then asks the user through the device, "What is the name of the dish you want to cook?" The user responds with "steak," and the device sends this to the server.
[0141] Input: User's voice command "Start cooking"
[0142] Output: Request for confirmation of dish name from server, dish name from user: "Steak"
[0143] Step 3: Enter ingredients and recipe
[0144] The server sends a message to the terminal saying, "Please tell us the type of meat you would like to use and its thickness." The terminal then asks the user aloud, "Please tell us the type of meat you would like to use and its thickness." The user enters "sirloin, 2 cm thick," and the terminal then sends this to the server.
[0145] Input: User's dish name "Steak", User's ingredient information "Sirloin, 2 cm thick"
[0146] Output: Register ingredient information and recipe information on the server
[0147] Step 4: Collect and analyze sensor data
[0148] The device starts collecting data from the temperature sensor and camera. The collected data is sent to the server in real time. The server analyzes the temperature and image data to evaluate the doneness and moisture content of the ingredients. For example, if the server determines that one side is still raw, it sends the result to the device.
[0149] Input: Real-time data from temperature sensors and cameras
[0150] Output: Analysis results, evaluation of doneness and moisture content
[0151] Step 5: Advice and coordination
[0152] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal. Specifically, it notifies the user, "One side is cooked, please turn the meat over." The user then turns the meat over as instructed.
[0153] Input: Analysis results of doneness and moisture content of ingredients
[0154] Output: Specific cooking advice to the user
[0155] Step 6: Cooking complete and feedback
[0156] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal notifies the user "The steak is done." The user enters feedback into the terminal, such as "The steak is delicious," and the terminal sends this to the server. The server records the user's feedback and uses it to improve the accuracy of the analysis method.
[0157] Input: Final analysis result of ingredients, user feedback "The steak turned out delicious."
[0158] Output: Notification of cooking completion, improvement of analytical methods based on feedback
[0159] Through the above process, this system allows even beginners to easily cook like a pro. Furthermore, by learning from user feedback, the system can provide more accurate cooking advice.
[0160] (Application example 1)
[0161] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0162] Maintaining the quality of food during delivery is difficult, especially if temperature control is inadequate, which can lead to loss of flavor and texture. To solve this problem, a system is needed that can monitor the status of food in real time and provide appropriate instructions and advice.
[0163] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0164] In this invention, the server includes a temperature sensor means for detecting the condition of the ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means, and a notification means for evaluating the condition of the food while it is being delivered and notifying the delivery person. This makes it possible to monitor the quality of the food in real time even during delivery, and to provide instructions on the optimal delivery timing and route.
[0165] A "temperature sensor" is a device that detects the temperature of ingredients in real time and provides that data to other analytical means.
[0166] The "camera" is a device for taking an image of the ingredients and transmitting the image data to the analysis means to evaluate the condition of the ingredients.
[0167] The "analysis means" is a means for evaluating the doneness and moisture content of ingredients using data obtained from the temperature sensor and camera.
[0168] The "notification means" is a device or function that provides advice on cooking methods and seasonings in natural language based on the analysis means.
[0169] The "notification means for evaluating the condition of food during delivery" is a function that monitors the condition of food during delivery and notifies the delivery person of the quality of the food and the appropriate delivery timing.
[0170] "Ingredient information" is detailed information such as the type, thickness, and amount of ingredients used in cooking.
[0171] "Recipe information" is information about specific cooking procedures and seasoning methods using ingredients.
[0172] "User feedback" refers to ratings and comments provided by users about the results of cooking and delivery.
[0173] "Learning means" refers to functions and algorithms that improve the accuracy of the analysis means based on user feedback.
[0174] This invention is a cooking and delivery assistance system that provides the delivery person with the appropriate delivery timing and route while maintaining the quality of the food. This system is composed of the following main components.
[0175] The system is equipped with a temperature sensor that detects the condition of ingredients in real time and a camera that captures images of the ingredients. These devices are connected to a smartphone or tablet to monitor the cooking status.
[0176] Next, there is the analytical means for analyzing the data from the temperature sensors and cameras. This analytical means resides on a cloud server and evaluates the doneness and moisture content of ingredients based on the collected temperature and image data. The analytical means used here include machine learning algorithms and generative AI models.
[0177] The system also includes a notification means that provides advice on cooking methods and seasonings in natural language based on the results of the analysis means. This notification means notifies the delivery person in real time via an application installed on their smartphone or tablet.
[0178] Specifically, when the system starts up, the server initializes the temperature sensor and camera, loads the cooking database and AI model, processes the data using analytical means, and sends the results to the delivery person via notification means.
[0179] For example, if a delivery person voice-inputs "I'm starting delivery," the system will ask, "Please tell me the name of the food you're delivering." If the delivery person inputs "pizza," that data will be sent to the server. The server then starts collecting data from the temperature sensor and camera, analyzes it, and evaluates the condition of the food. Based on the evaluation results, it will provide the delivery person with a notification, such as, "The pizza is at the right temperature. Please use the fastest route."
[0180] This method allows the system to provide optimal delivery timing and route information while maintaining the quality of the food. Furthermore, the system also has a learning mechanism for collecting feedback after delivery is completed and improving the accuracy of analysis in the future.
[0181] As a concrete example, the input prompt sentence for the generative AI model is as follows:
[0182] Based on the current temperature of the pizza, assess whether the food is of good quality.
[0183] In this way, it is possible to achieve quality control of food and optimal delivery support in food delivery.
[0184] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0185] Step 1:
[0186] System startup and initialization
[0187] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and generative AI model. After initialization is complete, the terminal displays the message, "Ready. Please give us your instructions to start delivery."
[0188] Step 2:
[0189] Start a session with the user
[0190] The user speaks "I'm starting delivery" into the device. The device converts this voice input into text and sends it to the server. The server receives the data and sends a voice command to the device asking, "Please tell me the name of the dish you want to deliver."
[0191] Step 3:
[0192] Enter dish information
[0193] The user speaks "pizza." The device converts this speech into text and sends it to the server. The server analyzes the received data and sends a voice prompt to the device asking, "Please provide information about the toppings you'll be using."
[0194] Step 4:
[0195] Sensor data collection and analysis
[0196] The device starts collecting data from the temperature sensor and camera. The temperature sensor captures the temperature data of the ingredients, and the camera captures images of the ingredients. This data is sent to the server in real time. The server analyzes the temperature data and image data to evaluate the doneness and moisture content of the ingredients.
[0197] Step 5:
[0198] Cooking status notification and delivery advice
[0199] The server evaluates the current state of the ingredients based on the analysis method. Based on the evaluation result, it notifies the device, "The pizza is at the right temperature. Please use the fastest route." The device then conveys this notification to the user via voice or text.
[0200] Step 6:
[0201] Continuous data collection and reanalysis
[0202] The device continuously collects data via temperature sensors and cameras and sends it to a server, which reanalyzes the new data and provides new advice if the cooking condition changes.
[0203] Step 7:
[0204] Delivery completion and feedback
[0205] Once the delivery is complete, the user voice-inputs "Delivery completed" into the device. The device sends this to the server, which updates the delivery status. The user then inputs feedback such as "The pizza arrived delicious," which the device sends to the server. The server analyzes the user's feedback and uses it as learning data to improve the accuracy of future advice.
[0206] Through these steps, the system can maintain food quality during delivery and notify customers of the optimal delivery time and route.
[0207] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0208] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[0209] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means that analyzes the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means of obtaining ingredient and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion engine that recognizes the user's emotions and adjusts cooking advice based on them.
[0210] Overview of program processing
[0211] System startup and initialization
[0212] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensors, camera, and emotion engine for normal operation. It then loads the cooking database and AI model.
[0213] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0214] Start a session with the user
[0215] The user inputs "I'm going to start cooking" into the terminal by voice.
[0216] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0217] The user speaks "steak," which the device recognizes and sends to the server.
[0218] Enter ingredients and recipe
[0219] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0220] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0221] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0222] Emotion recognition by emotion engine
[0223] The device sends the user's voice and facial image to the emotion engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[0224] Sensor data collection and analysis
[0225] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0226] The server analyzes the data it receives and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. It also adjusts the tone and content of cooking advice based on the emotion engine's evaluation results.
[0227] Providing advice and coordination
[0228] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, it may also provide advice to help them relax, such as "Take a short break."
[0229] Cooking completion and feedback
[0230] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal then notifies the user by voice, "The steak is done."
[0231] The user provides vocal feedback to the device, such as "The steak turned out delicious." The device then sends the provided feedback to the server. The server then uses the received feedback to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[0232] Specific examples
[0233] For example, if a user is cooking a steak and the system determines that the user is tired based on their tone of voice or facial expression, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without pushing yourself."
[0234] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and camera, taking into consideration the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensor, camera, and emotion engine for normal operation.
[0238] Step 2:
[0239] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0240] Step 3:
[0241] The user inputs "I'm going to start cooking" into the terminal by voice.
[0242] Step 4:
[0243] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0244] Step 5:
[0245] The user speaks "steak," which the device recognizes and sends to the server.
[0246] Step 6:
[0247] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0248] Step 7:
[0249] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0250] Step 8:
[0251] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0252] Step 9:
[0253] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0254] Step 10:
[0255] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0256] Step 11:
[0257] The emotion engine analyzes the user's voice and facial image to assess their emotional state, for example determining whether they are feeling stressed or tired.
[0258] Step 12:
[0259] Based on the analysis results and the user's emotional state, the server sends a notification to the device saying, "One side of the meat is still raw. Please continue cooking it." If the user is feeling stressed, the server adds, "It would be good to take a short break."
[0260] Step 13:
[0261] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking. It may be a good idea to take a short break."
[0262] Step 14:
[0263] The device continues to collect data from the temperature sensor and camera.
[0264] Step 15:
[0265] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0266] Step 16:
[0267] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0268] Step 17:
[0269] The user performs the action of turning the meat over.
[0270] Step 18:
[0271] The device continues to collect data and send it to the server.
[0272] Step 19:
[0273] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0274] Step 20:
[0275] The device will notify the user by voice, "Your steak is ready."
[0276] Step 21:
[0277] The user provides voice feedback to the device, such as "The steak turned out delicious."
[0278] Step 22:
[0279] The device sends the provided feedback to the server.
[0280] Step 23:
[0281] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[0282] These are the specific processing steps of the cooking assistance system equipped with AI and an emotion engine, which enables users to easily prepare delicious meals and provides flexible support according to the user's emotional state.
[0283] Example 2
[0284] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0285] Conventional cooking assistance systems provide cooking advice based solely on the state of ingredients, and are unable to take the user's emotional state into account. Furthermore, they lack the support necessary to enable users without specialized cooking knowledge to easily cook like professionals. Furthermore, it is difficult to improve the accuracy of advice through continuous learning based on user feedback.
[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0287] In this invention, the server includes a temperature sensor that detects the state of ingredients, an image sensor that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the image sensor to evaluate the doneness and moisture content of the ingredients, notification means that provides cooking method and seasoning advice in natural language based on the analysis means, an emotion recognition engine that recognizes the user's emotional state based on voice and video, and means for adjusting the tone and content of the cooking advice based on the evaluation of the emotion recognition engine. This makes it possible to provide cooking advice that takes the user's emotional state into consideration, and provides support that allows even users without specialized cooking expertise to easily cook like a professional. It is also possible to improve the accuracy of the advice by utilizing feedback from the user.
[0288] A "temperature sensor" is a device that measures the temperature of ingredients in real time and provides that data to an analytical means.
[0289] An "image sensor" is a device that captures images of ingredients and provides the visual data to an analysis means.
[0290] The "analysis means" is a system that evaluates the doneness and moisture content of ingredients based on data obtained from the temperature sensor and image sensor.
[0291] The "notification means" is a system that provides the user with advice on cooking methods and seasonings in natural language based on the results obtained from the analysis means.
[0292] An "emotion recognition engine" is a combination of software and hardware that analyzes a user's audio and video data and evaluates the user's emotional state.
[0293] The "adjustment means" is a system that appropriately changes the tone and content of cooking advice based on the emotion evaluation results obtained from the emotion recognition engine.
[0294] The "learning means" is a system that incorporates feedback from users into the analysis means to improve the accuracy of future cooking advice.
[0295] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[0296] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, an image sensor that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for acquiring ingredient information and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion recognition engine that recognizes the user's emotions and adjusts cooking advice based on this.
[0297] Hardware and Software Configuration
[0298] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine, ensuring all devices are ready to operate properly. The server also loads the cooking database and generative AI model.
[0299] The device waits for input from the user and displays "Ready to go. Give us your instructions to start cooking." At this point, the user can use voice input.
[0300] When the user speaks "I'm going to start cooking" into the device, the device asks "What is the name of the dish you want to cook?", and the user speaks "steak." The device then sends this information to the server.
[0301] Next, the server receives the name of the dish from the user and sends it to the terminal, asking, "Please tell us the type of meat and thickness you want to use," and the terminal asks the user questions based on that. The user voice-inputs, "Sirloin, 2 cm thick," and the terminal sends that information to the server.
[0302] Emotion Engine
[0303] The device sends the user's voice and facial image to the emotion recognition engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[0304] Sensor data collection and analysis
[0305] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. The server also adjusts the tone and content of the cooking advice based on the evaluation results of the emotion engine.
[0306] Providing advice and feedback
[0307] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, the server will also provide advice to help them relax, such as "It would be good to take a short break."
[0308] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal, which then notifies the user by voice.
[0309] The user provides verbal feedback to the device, such as "The steak turned out delicious," and the device then sends that feedback to the server. The server uses the feedback it receives to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[0310] Specific examples and prompts for generative AI models
[0311] For example, if the emotion recognition engine determines that a user is tired from the tone of their voice and facial expression while cooking a steak, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without overdoing it."
[0312] Example prompt for a generative AI model:
[0313] Describe a situation where the emotion recognition engine determined from audio and video data that the user was stressed while cooking a steak, and what advice should be provided in that situation.
[0314] (Example prompt)
[0315] While the user is cooking a steak, suggest, "Today, let's try a simple steak recipe that doesn't require much effort." Also, give the user advice to relax, such as, "This heat level is fine. Please proceed without pushing yourself."
[0316] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and image sensor, taking into account the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0317] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0318] Step 1: Boot and Initialize the System
[0319] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine. The input is the power supply and communication connection for each device, and the output is that the device is ready to operate normally. This allows the server to load the cooking database and generative AI model. Specifically, it checks the operating status of the device and prepares for data collection.
[0320] Step 2: Waiting for user input
[0321] The device waits for input from the user and displays "Ready to cook. Please give me instructions to start cooking." The input is a voice command from the user, and the output is the display on the device and the activation of the voice assistant. Specifically, the device starts the voice recognition system and waits for the user's voice.
[0322] Step 3: Start a session with the user
[0323] The user speaks to the device, saying, "I'm going to start cooking." The device receives the user's voice input and asks, "Please tell me the name of the dish you want to cook." The input is the user's voice instruction, and the output is the device's question and data transmission to the server. Specifically, the voice recognition system converts the user's instruction into text and sends it to the server.
[0324] Step 4: Enter ingredients and recipe
[0325] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type of meat and thickness you would like to use." The input is the user's dish name information, and the output is a prompt to the terminal. The terminal asks the user aloud, "Please tell us the type of meat and thickness you would like to use." The input is a voice instruction from the user, and the output is data sent to the server. In concrete terms, the terminal uses a voice recognition system to collect user information and sends it to the server.
[0326] Step 5: Emotion Recognition with the Emotion Engine
[0327] The device sends the user's voice and facial image to the emotion recognition engine. The input is the user's voice data and video data, and the output is the emotion engine's analysis results. The emotion recognition engine analyzes this data and evaluates the user's emotional state. For example, if it determines that the user is feeling stressed, it sends this information to the server. Specifically, the emotion engine performs voice and facial expression analysis in real time.
[0328] Step 6: Collect and analyze sensor data
[0329] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The input is temperature data and image data, and the output is the analysis results. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that "one side is still raw." The tone and content of the cooking advice are adjusted taking into account the evaluation results of the emotion engine. Specifically, the server inputs the temperature data and image data into the analysis algorithm and obtains the results.
[0330] Step 7: Advice and coordination
[0331] The server determines the current cooking status based on the analysis results and sends advice such as "One side of the meat is still raw. Please continue cooking" to the terminal. The input is the analysis results and the emotion evaluation results, and the output is customized cooking advice. If the user's emotional state indicates stress, it also provides advice to help them relax, such as "It would be good to take a short break." In concrete terms, the server generates advice based on the cooking status and emotional state.
[0332] Step 8: Cooking complete and feedback
[0333] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The input is the analysis result, and the output is a notification that cooking is complete. The terminal notifies the user by voice that "The steak is done." The user provides feedback to the terminal by voice, saying "The steak turned out delicious." The input is the user's feedback, and the output is the feedback sent to the server. The feedback received by the server is used to improve the accuracy of the analysis method. Specifically, the server uses the feedback data as learning data for the generative AI model, improving the accuracy of advice from the next time onwards.
[0334] (Application example 2)
[0335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0336] Conventional cooking support systems were able to detect the condition of ingredients and provide cooking advice, but they did not provide support that took the user's emotions into consideration. As a result, if the user felt stressed or tired while cooking, appropriate advice was not provided, resulting in a decrease in satisfaction. Furthermore, in the kitchen, it is necessary to detect the condition in real time and respond immediately, so flexible support that adapts to the user's condition is necessary.
[0337] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0338] In this invention, the server includes a temperature sensor means for detecting the state of ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, an emotion analysis means for recognizing the user's emotions and adjusting advice on cooking methods and seasonings based on the emotions, and a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means and emotion analysis means. This enables flexible and appropriate cooking advice to be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[0339] A "temperature sensor that detects the state of ingredients" is a device that detects the temperature of ingredients in real time while they are being cooked.
[0340] The "camera for capturing images of ingredients" is a device for recording and detecting the visual state of ingredients during cooking.
[0341] The "analysis means" refers to devices or software that have the function of analyzing data from the temperature sensor and camera to evaluate the doneness and moisture content of the ingredients.
[0342] The "emotion analysis means" is a device or software that analyzes the user's voice and facial expressions to evaluate their emotional state and adjust the cooking method and seasoning advice.
[0343] The "notification means" is a device or software that has the function of providing the user with advice on cooking methods and seasonings in natural language based on the analysis results and emotion analysis results.
[0344] The "learning means" is a device or software that has the function of receiving feedback from users and storing and analyzing data to improve the accuracy of the analysis means.
[0345] "Real-time" means that processing and analysis are done almost immediately, and results are provided immediately.
[0346] An "emotion engine that recognizes user emotions" is an engine or software that evaluates the user's emotional state from their tone of voice and facial expressions and reflects that in the system.
[0347] The following describes the system configuration and processing details as an embodiment of this invention. The system is composed of a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means that analyzes the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, means for acquiring ingredient information and recipe information from the user, emotion analysis means that recognizes the user's emotions, and learning means that improves the accuracy of the analysis means based on feedback from the user.
[0348] First, when the system starts up, the server establishes connections with the temperature sensor, camera, and sentiment analysis engine, and loads the cooking database and generative AI model. The temperature sensor used is a DHT22, and the camera used is a Raspberry Pi camera module. This prepares the temperature sensor, camera, and sentiment analysis engine to operate normally.
[0349] The server notifies the terminal (a monitor device installed in the kitchen) that preparation is complete, and the terminal enters a state of waiting for input from the user. When the user issues a verbal instruction to start cooking, the terminal responds, listening to the name of the dish to be cooked and the types and amounts of ingredients to be used, and sending this to the server. When the server receives ingredient information and recipe information from the user, an analysis means determines the cooking procedure and generates the necessary advice.
[0350] During cooking, temperature sensors and cameras monitor the condition of the ingredients in real time and send the data to a server. The server analyzes the received data and evaluates the ingredients' doneness and moisture content. At the same time, an emotion analysis unit analyzes the user's emotional state from their voice tone and facial expressions, and if stress or fatigue is detected, the system adjusts the advice accordingly.
[0351] Based on the analysis results, the server generates cooking and seasoning advice and notifies the user via the device. For example, if one side of a steak is still raw while cooking, specific instructions such as "One side of the meat is still raw. Please continue cooking" are provided. Also, if the user shows signs of fatigue, advice encouraging relaxation such as "Take a short break. There is still time."
[0352] Once the cooking is complete, the server sends the results to the device, which notifies the user. When the user provides feedback, the server adds that feedback to its learning curve and uses it to improve the accuracy of future recommendations, allowing the system to provide more accurate recommendations with each use.
[0353] Examples:
[0354] Example prompt: "What cooking advice should you give if the user indicates that the meat is not yet cooked through?"
[0355] Example prompt: "If the user is determined to be tired, what relaxation advice should be offered?"
[0356] In this way, flexible and appropriate cooking advice can be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[0357] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0358] Step 1:
[0359] The server starts the system and establishes connections with the temperature sensor, camera, and sentiment analysis engine. This initializes the temperature sensor, camera, and sentiment analysis engine so that they can operate normally. The input is connection information from the temperature sensor and camera, and based on that, the server checks the device connection status and initializes it. The output is a notification that each device has been successfully connected and is available for use. Specifically, the server checks the response from each device and records the connection status in a log.
[0360] Step 2:
[0361] The user issues a verbal command to the device, saying "Start cooking." The input is the user's voice command, which is converted into text data using speech recognition software. The output is the converted text data, which is sent to the server. Specifically, the device picks up the user's voice with a microphone and converts it into text using speech recognition software (e.g., Google® Cloud Speech-to-Text).
[0362] Step 3:
[0363] The server confirms the "start cooking" instruction received from the user and sends a message to the terminal saying "Please tell us the name of the dish you want to cook." The input is the instruction to start cooking from the user, and the next question is determined based on that. The output is the next question displayed on the terminal and output as voice. In concrete terms, the server sends this question in digital data format to the terminal, and the terminal displays it to the user and outputs it as voice.
[0364] Step 4:
[0365] The user verbally instructs the terminal on the name of the dish to be cooked. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. In concrete terms, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[0366] Step 5:
[0367] The server analyzes the dish name information received from the user, and then sends a question to the terminal to request information on the necessary ingredients. Specifically, the input is text data for the dish name, and based on that, a question is generated to request ingredient information. As an output, this question is displayed and output as voice on the terminal. In concrete terms, the server retrieves the necessary information from a database that stores ingredient information, generates a question based on that, and sends it to the terminal.
[0368] Step 6:
[0369] The user verbally instructs the terminal on ingredient information. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. Specifically, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[0370] Step 7:
[0371] The device sends the user's voice and facial image to an emotion analysis engine. The input is the user's voice and video data, and the emotional state is evaluated based on this. The output is the emotion analysis results sent to the server. Specifically, the device captures the user's video and audio using the camera and microphone, and analyzes them using the emotion analysis engine (e.g., Affectiva SDK).
[0372] Step 8:
[0373] The server starts collecting data from the temperature sensor and camera and analyzes it in real time. The input is data from the temperature sensor and camera, and based on that data, it analyzes the doneness and moisture content of the ingredients. The output is an analysis result that evaluates the cooking status. Specifically, the server acquires data from the temperature sensor and camera and analyzes it using image processing software (e.g., OpenCV) and data analysis software (e.g., TENSORFLOW (registered trademark)).
[0374] Step 9:
[0375] The server determines the current cooking status based on the analysis results and sends cooking advice to the device. If necessary, it also provides advice that takes into account the results of sentiment analysis. The inputs include analysis results from the temperature sensor and camera and sentiment analysis results, and cooking advice is generated based on these. The output is the advice displayed on the device and output as voice. Specifically, the server generates cooking advice using natural language generation software (e.g., Google Cloud Text-to-Speech) and sends it to the device.
[0376] Step 10:
[0377] The server determines whether cooking is complete and sends the result to the terminal. The input is the final analysis results from the temperature sensor and camera, and cooking completion is determined based on these. The output is a cooking completion message that is displayed and output aloud on the terminal. In concrete terms, the server evaluates the final analysis results, generates a message indicating that cooking is complete, and sends it to the terminal.
[0378] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0379] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0380] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0381] [Second embodiment]
[0382] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0383] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0384] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0385] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0386] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0387] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0388] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0389] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0390] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0391] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0392] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0393] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0394] This invention is a cooking assistance system that allows even beginners to easily prepare delicious dishes. This system detects the condition of ingredients in real time and provides appropriate cooking advice. The operation of this system is described in detail below.
[0395] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for obtaining ingredient and recipe information from the user, and a learning means that improves the accuracy of the analysis means based on user feedback.
[0396] Overview of program processing
[0397] System startup and initialization
[0398] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and AI model.
[0399] The device waits for a voice command from the user and displays "Ready. Give me instructions to start cooking."
[0400] Start a session with the user
[0401] The user inputs "I'm going to start cooking" into the terminal by voice.
[0402] The device will ask the user aloud, "Please tell us the name of the dish you would like to cook."
[0403] The user types "steak" and the terminal sends it to the server.
[0404] Enter ingredients and recipe
[0405] The server sends a message to the terminal asking, "Please tell us the type and thickness of meat you would like to use."
[0406] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0407] The user types in "sirloin, 2 cm thick," and the device sends this to the server.
[0408] Sensor data collection and analysis
[0409] The device starts collecting data from the temperature sensor and camera and sends the data to the server.
[0410] The server analyzes the temperature and image data collected in real time to evaluate the doneness and moisture content of the ingredients.
[0411] For example, if it determines that "one side is still raw," the device will notify the user, "One side of the meat is still raw. Please continue cooking."
[0412] Providing advice and coordination
[0413] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal.
[0414] For example, the system may notify the user that "One side is cooked, please turn the meat over," and the user should follow the instructions to turn the meat over.
[0415] The device continues to collect data from the sensors and the server re-analyzes the new data.
[0416] Cooking completion and feedback
[0417] The server determines that the steak is cooked through and sends the message "The steak is done" to the terminal.
[0418] The device notifies the user that "your steak is done."
[0419] The user inputs feedback such as "The steak turned out delicious" into the terminal, and the terminal transmits the feedback to the server.
[0420] The server records user feedback and uses it to improve the accuracy of the analysis method.
[0421] In this way, the cooking assistance system of the present invention analyzes data from temperature sensors and cameras to provide support that allows anyone to easily cook like a pro. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0422] The processing flow will be explained below.
[0423] Step 1:
[0424] The server starts the system and establishes the connection between the temperature sensor and the camera, so that the sensor and camera are ready to operate normally.
[0425] Step 2:
[0426] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0427] Step 3:
[0428] The user inputs "I'm going to start cooking" into the terminal by voice.
[0429] Step 4:
[0430] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0431] Step 5:
[0432] The user speaks "steak," which the device recognizes and sends to the server.
[0433] Step 6:
[0434] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0435] Step 7:
[0436] The device asks the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0437] Step 8:
[0438] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0439] Step 9:
[0440] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0441] Step 10:
[0442] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0443] Step 11:
[0444] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking."
[0445] Step 12:
[0446] The device continues to collect data from the temperature sensor and camera.
[0447] Step 13:
[0448] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0449] Step 14:
[0450] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0451] Step 15:
[0452] The user performs the action of turning the meat over.
[0453] Step 16:
[0454] The device continues to collect data and send it to the server.
[0455] Step 17:
[0456] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0457] Step 18:
[0458] The device will notify the user by voice, "Your steak is ready."
[0459] Step 19:
[0460] The user provides voice feedback to the device, such as "The steak turned out delicious."
[0461] Step 20:
[0462] The device sends the provided feedback to the server.
[0463] Step 21:
[0464] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[0465] These are the specific processing steps of the AI-powered cooking assistance system, which enables users to easily prepare delicious meals.
[0466] Example 1
[0467] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0468] Conventional cooking assistance systems have struggled to properly monitor the condition of ingredients and provide real-time advice based on the cooking progress. In particular, accurately measuring the doneness and moisture content of ingredients and providing cooking instructions at the appropriate time have been challenging. Furthermore, there has been a lack of technology to improve the accuracy of the system based on user feedback, leading to a demand for a system that allows even beginners to easily prepare delicious dishes.
[0469] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0470] In this invention, the server includes a temperature sensor that detects the state of the ingredients, a camera that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, notification means that provides advice on cooking methods and seasonings in natural language based on the analysis means, input means that accepts voice commands from the user, means that loads a cooking database and a generative AI model and analyzes based on data from the analysis means, and learning means that improves the accuracy of the analysis means based on feedback from the user. This makes it easy for anyone to cook like a professional.
[0471] A "temperature sensor" is a device that measures the surface and internal temperature of ingredients in real time and provides that data to the system.
[0472] The "camera" is a photographing device that captures images of ingredients and provides the visual data to the system.
[0473] The "analysis means" is a collection of algorithms and programs for evaluating the doneness and moisture content of ingredients based on data obtained from temperature sensors and cameras.
[0474] The "notification means" is a function for conveying the cooking method and seasoning advice obtained by the analysis means to the user in natural language.
[0475] "Input means" refers to an interface through which a user inputs voice commands into the system.
[0476] A "cooking database" is a data storage device that accumulates data and recipes related to various dishes.
[0477] A "generative AI model" is an artificial intelligence algorithm that uses collected data and past feedback to assist analytical methods and optimize cooking instructions.
[0478] The "learning means" is a function for improving the accuracy of the analysis means based on feedback from users.
[0479] "Ingredient information" is data such as the type, amount, and shape of ingredients used in cooking.
[0480] "Recipe information" refers to information such as the steps to make a particular dish, the ingredients needed, and cooking time.
[0481] A "user" is a person who operates the system and cooks according to the advice and instructions provided.
[0482] A "server" is a central processing unit that controls and analyzes the entire system.
[0483] A "terminal" is a device that provides an interface with a user and displays voice commands and notifications.
[0484] This invention is a cooking assistance system that allows even beginners to easily prepare delicious meals. The main components of the system include a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means for analyzing the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, input means for acquiring ingredient information and recipe information from the user, and learning means that improves the accuracy of the analysis means based on user feedback.
[0485] Hardware and Software Configuration
[0486] The server controls the entire system and receives data from temperature sensors (e.g., digital temperature sensors) and cameras (e.g., high-resolution cameras). The server loads a cooking database and generative AI models and uses them to analyze the data in real time. Specific processes performed by the server include collecting temperature data, analyzing image data, and evaluating the doneness and moisture content of ingredients.
[0487] A terminal is a device that provides an interface with the user and waits for voice commands. For example, it is used as a smart speaker or tablet. The terminal receives voice commands from the user and sends them to the server. It also has the function of notifying the user of advice from the server.
[0488] The user is the person who operates the system and proceeds with the cooking. The user inputs voice commands into the terminal and cooks according to the instructions from the system. Furthermore, the user provides feedback to the system after cooking is completed.
[0489] Specific examples
[0490] For example, if a user wants to cook a steak, the following series of processes takes place: The user speaks "I'd like to start cooking" into the device, and the device sends a request to the server. The server then asks the user via the device what dish they want to cook. The user answers "steak," and based on that, the server requests detailed information about the ingredients. The user answers "sirloin, 2 cm thick," and the device sends this to the server.
[0491] The temperature sensor and camera collect the steak's temperature and image data, which are then sent to a server. The server analyzes this data in real time to assess the steak's doneness and moisture content. Based on the results, the server provides advice to the user, such as "One side is still raw. Continue cooking."
[0492] Prompt Sentence Examples
[0493] "I want to grill some sirloin steak. What's the best way to check the doneness and make sure it's delicious?"
[0494] This process allows users to receive accurate cooking advice, enabling even beginners to cook like professionals. Furthermore, as the system continues to learn based on user feedback, it provides increasingly accurate advice.
[0495] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0496] Step 1: Boot and Initialize the System
[0497] The server starts the system. As part of the initialization, the server connects the temperature sensor and camera and checks their operation. Specifically, it acquires test data from the sensors and camera and verifies that they are working properly. The server also loads the cooking database and generative AI model, which will serve as the basis for future data analysis and advice provision.
[0498] Input: None
[0499] Output: System initialization complete, temperature sensor and camera operation confirmation, cooking database and AI model loading status
[0500] Step 2: Start a session with the user
[0501] The user speaks "I'm going to start cooking" into the device. The device recognizes this speech and sends a request to the server. The server then asks the user through the device, "What is the name of the dish you want to cook?" The user responds with "steak," and the device sends this to the server.
[0502] Input: User's voice command "Start cooking"
[0503] Output: Request for confirmation of dish name from server, dish name from user: "Steak"
[0504] Step 3: Enter ingredients and recipe
[0505] The server sends a message to the terminal saying, "Please tell us the type of meat you would like to use and its thickness." The terminal then asks the user aloud, "Please tell us the type of meat you would like to use and its thickness." The user enters "sirloin, 2 cm thick," and the terminal then sends this to the server.
[0506] Input: User's dish name "Steak", User's ingredient information "Sirloin, 2 cm thick"
[0507] Output: Register ingredient information and recipe information on the server
[0508] Step 4: Collect and analyze sensor data
[0509] The device starts collecting data from the temperature sensor and camera. The collected data is sent to the server in real time. The server analyzes the temperature and image data to evaluate the doneness and moisture content of the ingredients. For example, if the server determines that one side is still raw, it sends the result to the device.
[0510] Input: Real-time data from temperature sensors and cameras
[0511] Output: Analysis results, evaluation of doneness and moisture content
[0512] Step 5: Advice and coordination
[0513] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal. Specifically, it notifies the user, "One side is cooked, please turn the meat over." The user then turns the meat over as instructed.
[0514] Input: Analysis results of doneness and moisture content of ingredients
[0515] Output: Specific cooking advice to the user
[0516] Step 6: Cooking complete and feedback
[0517] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal notifies the user "The steak is done." The user enters feedback into the terminal, such as "The steak is delicious," and the terminal sends this to the server. The server records the user's feedback and uses it to improve the accuracy of the analysis method.
[0518] Input: Final analysis result of ingredients, user feedback "The steak turned out delicious."
[0519] Output: Notification of cooking completion, improvement of analytical methods based on feedback
[0520] Through the above process, this system allows even beginners to easily cook like a pro. Furthermore, by learning from user feedback, the system can provide more accurate cooking advice.
[0521] (Application example 1)
[0522] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0523] Maintaining the quality of food during delivery is difficult, especially if temperature control is inadequate, which can lead to loss of flavor and texture. To solve this problem, a system is needed that can monitor the status of food in real time and provide appropriate instructions and advice.
[0524] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0525] In this invention, the server includes a temperature sensor means for detecting the condition of the ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means, and a notification means for evaluating the condition of the food while it is being delivered and notifying the delivery person. This makes it possible to monitor the quality of the food in real time even during delivery, and to provide instructions on the optimal delivery timing and route.
[0526] A "temperature sensor" is a device that detects the temperature of ingredients in real time and provides that data to other analytical means.
[0527] The "camera" is a device for taking an image of the ingredients and transmitting the image data to the analysis means to evaluate the condition of the ingredients.
[0528] The "analysis means" is a means for evaluating the doneness and moisture content of ingredients using data obtained from the temperature sensor and camera.
[0529] The "notification means" is a device or function that provides advice on cooking methods and seasonings in natural language based on the analysis means.
[0530] The "notification means for evaluating the condition of food during delivery" is a function that monitors the condition of food during delivery and notifies the delivery person of the quality of the food and the appropriate delivery timing.
[0531] "Ingredient information" is detailed information such as the type, thickness, and amount of ingredients used in cooking.
[0532] "Recipe information" is information about specific cooking procedures and seasoning methods using ingredients.
[0533] "User feedback" refers to ratings and comments provided by users about the results of cooking and delivery.
[0534] "Learning means" refers to functions and algorithms that improve the accuracy of the analysis means based on user feedback.
[0535] This invention is a cooking and delivery assistance system that provides the delivery person with the appropriate delivery timing and route while maintaining the quality of the food. This system is composed of the following main components.
[0536] The system is equipped with a temperature sensor that detects the condition of ingredients in real time and a camera that captures images of the ingredients. These devices are connected to a smartphone or tablet to monitor the cooking status.
[0537] Next, there is the analytical means for analyzing the data from the temperature sensors and cameras. This analytical means resides on a cloud server and evaluates the doneness and moisture content of ingredients based on the collected temperature and image data. The analytical means used here include machine learning algorithms and generative AI models.
[0538] The system also includes a notification means that provides advice on cooking methods and seasonings in natural language based on the results of the analysis means. This notification means notifies the delivery person in real time via an application installed on their smartphone or tablet.
[0539] Specifically, when the system starts up, the server initializes the temperature sensor and camera, loads the cooking database and AI model, processes the data using analytical means, and sends the results to the delivery person via notification means.
[0540] For example, if a delivery person voice-inputs "I'm starting delivery," the system will ask, "Please tell me the name of the food you're delivering." If the delivery person inputs "pizza," that data will be sent to the server. The server then starts collecting data from the temperature sensor and camera, analyzes it, and evaluates the condition of the food. Based on the evaluation results, it will provide the delivery person with a notification, such as, "The pizza is at the right temperature. Please use the fastest route."
[0541] This method allows the system to provide optimal delivery timing and route information while maintaining the quality of the food. Furthermore, the system also has a learning mechanism for collecting feedback after delivery is completed and improving the accuracy of analysis in the future.
[0542] As a concrete example, the input prompt sentence for the generative AI model is as follows:
[0543] Based on the current temperature of the pizza, assess whether the food is of good quality.
[0544] In this way, it is possible to achieve quality control of food and optimal delivery support in food delivery.
[0545] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0546] Step 1:
[0547] System startup and initialization
[0548] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and generative AI model. After initialization is complete, the terminal displays the message, "Ready. Please give us your instructions to start delivery."
[0549] Step 2:
[0550] Start a session with the user
[0551] The user speaks "I'm starting delivery" into the device. The device converts this voice input into text and sends it to the server. The server receives the data and sends a voice command to the device asking, "Please tell me the name of the dish you want to deliver."
[0552] Step 3:
[0553] Enter dish information
[0554] The user speaks "pizza." The device converts this speech into text and sends it to the server. The server analyzes the received data and sends a voice prompt to the device asking, "Please provide information about the toppings you'll be using."
[0555] Step 4:
[0556] Sensor data collection and analysis
[0557] The device starts collecting data from the temperature sensor and camera. The temperature sensor captures the temperature data of the ingredients, and the camera captures images of the ingredients. This data is sent to the server in real time. The server analyzes the temperature data and image data to evaluate the doneness and moisture content of the ingredients.
[0558] Step 5:
[0559] Cooking status notification and delivery advice
[0560] The server evaluates the current state of the ingredients based on the analysis method. Based on the evaluation result, it notifies the device, "The pizza is at the right temperature. Please use the fastest route." The device then conveys this notification to the user via voice or text.
[0561] Step 6:
[0562] Continuous data collection and reanalysis
[0563] The device continuously collects data via temperature sensors and cameras and sends it to a server, which reanalyzes the new data and provides new advice if the cooking condition changes.
[0564] Step 7:
[0565] Delivery completion and feedback
[0566] Once the delivery is complete, the user voice-inputs "Delivery completed" into the device. The device sends this to the server, which updates the delivery status. The user then inputs feedback such as "The pizza arrived delicious," which the device sends to the server. The server analyzes the user's feedback and uses it as learning data to improve the accuracy of future advice.
[0567] Through these steps, the system can maintain food quality during delivery and notify customers of the optimal delivery time and route.
[0568] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0569] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[0570] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means that analyzes the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means of obtaining ingredient and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion engine that recognizes the user's emotions and adjusts cooking advice based on them.
[0571] Overview of program processing
[0572] System startup and initialization
[0573] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensors, camera, and emotion engine for normal operation. It then loads the cooking database and AI model.
[0574] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0575] Start a session with the user
[0576] The user inputs "I'm going to start cooking" into the terminal by voice.
[0577] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0578] The user speaks "steak," which the device recognizes and sends to the server.
[0579] Enter ingredients and recipe
[0580] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0581] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0582] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0583] Emotion recognition by emotion engine
[0584] The device sends the user's voice and facial image to the emotion engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[0585] Sensor data collection and analysis
[0586] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0587] The server analyzes the data it receives and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. It also adjusts the tone and content of cooking advice based on the emotion engine's evaluation results.
[0588] Providing advice and coordination
[0589] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, it may also provide advice to help them relax, such as "Take a short break."
[0590] Cooking completion and feedback
[0591] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal then notifies the user by voice, "The steak is done."
[0592] The user provides vocal feedback to the device, such as "The steak turned out delicious." The device then sends the provided feedback to the server. The server then uses the received feedback to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[0593] Specific examples
[0594] For example, if a user is cooking a steak and the system determines that the user is tired based on their tone of voice or facial expression, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without pushing yourself."
[0595] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and camera, taking into consideration the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0596] The processing flow will be explained below.
[0597] Step 1:
[0598] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensor, camera, and emotion engine for normal operation.
[0599] Step 2:
[0600] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0601] Step 3:
[0602] The user inputs "I'm going to start cooking" into the terminal by voice.
[0603] Step 4:
[0604] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0605] Step 5:
[0606] The user speaks "steak," which the device recognizes and sends to the server.
[0607] Step 6:
[0608] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0609] Step 7:
[0610] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0611] Step 8:
[0612] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0613] Step 9:
[0614] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0615] Step 10:
[0616] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0617] Step 11:
[0618] The emotion engine analyzes the user's voice and facial image to assess their emotional state, for example determining whether they are feeling stressed or tired.
[0619] Step 12:
[0620] Based on the analysis results and the user's emotional state, the server sends a notification to the device saying, "One side of the meat is still raw. Please continue cooking it." If the user is feeling stressed, the server adds, "It would be good to take a short break."
[0621] Step 13:
[0622] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking. It may be a good idea to take a short break."
[0623] Step 14:
[0624] The device continues to collect data from the temperature sensor and camera.
[0625] Step 15:
[0626] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0627] Step 16:
[0628] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0629] Step 17:
[0630] The user performs the action of turning the meat over.
[0631] Step 18:
[0632] The device continues to collect data and send it to the server.
[0633] Step 19:
[0634] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0635] Step 20:
[0636] The device will notify the user by voice, "Your steak is ready."
[0637] Step 21:
[0638] The user provides voice feedback to the device, such as "The steak turned out delicious."
[0639] Step 22:
[0640] The device sends the provided feedback to the server.
[0641] Step 23:
[0642] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[0643] These are the specific processing steps of the cooking assistance system equipped with AI and an emotion engine, which enables users to easily prepare delicious meals and provides flexible support according to the user's emotional state.
[0644] Example 2
[0645] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0646] Conventional cooking assistance systems provide cooking advice based solely on the state of ingredients, and are unable to take the user's emotional state into account. Furthermore, they lack the support necessary to enable users without specialized cooking knowledge to easily cook like professionals. Furthermore, it is difficult to improve the accuracy of advice through continuous learning based on user feedback.
[0647] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0648] In this invention, the server includes a temperature sensor that detects the state of ingredients, an image sensor that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the image sensor to evaluate the doneness and moisture content of the ingredients, notification means that provides cooking method and seasoning advice in natural language based on the analysis means, an emotion recognition engine that recognizes the user's emotional state based on voice and video, and means for adjusting the tone and content of the cooking advice based on the evaluation of the emotion recognition engine. This makes it possible to provide cooking advice that takes the user's emotional state into consideration, and provides support that allows even users without specialized cooking expertise to easily cook like a professional. It is also possible to improve the accuracy of the advice by utilizing feedback from the user.
[0649] A "temperature sensor" is a device that measures the temperature of ingredients in real time and provides that data to an analytical means.
[0650] An "image sensor" is a device that captures images of ingredients and provides the visual data to an analysis means.
[0651] The "analysis means" is a system that evaluates the doneness and moisture content of ingredients based on data obtained from the temperature sensor and image sensor.
[0652] The "notification means" is a system that provides the user with advice on cooking methods and seasonings in natural language based on the results obtained from the analysis means.
[0653] An "emotion recognition engine" is a combination of software and hardware that analyzes a user's audio and video data and evaluates the user's emotional state.
[0654] The "adjustment means" is a system that appropriately changes the tone and content of cooking advice based on the emotion evaluation results obtained from the emotion recognition engine.
[0655] The "learning means" is a system that incorporates feedback from users into the analysis means to improve the accuracy of future cooking advice.
[0656] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[0657] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, an image sensor that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for acquiring ingredient information and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion recognition engine that recognizes the user's emotions and adjusts cooking advice based on this.
[0658] Hardware and Software Configuration
[0659] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine, ensuring all devices are ready to operate properly. The server also loads the cooking database and generative AI model.
[0660] The device waits for input from the user and displays "Ready to go. Give us your instructions to start cooking." At this point, the user can use voice input.
[0661] When the user speaks "I'm going to start cooking" into the device, the device asks "What is the name of the dish you want to cook?", and the user speaks "steak." The device then sends this information to the server.
[0662] Next, the server receives the name of the dish from the user and sends it to the terminal, asking, "Please tell us the type of meat and thickness you want to use," and the terminal asks the user questions based on that. The user voice-inputs, "Sirloin, 2 cm thick," and the terminal sends that information to the server.
[0663] Emotion Engine
[0664] The device sends the user's voice and facial image to the emotion recognition engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[0665] Sensor data collection and analysis
[0666] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. The server also adjusts the tone and content of the cooking advice based on the evaluation results of the emotion engine.
[0667] Providing advice and feedback
[0668] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, the server will also provide advice to help them relax, such as "It would be good to take a short break."
[0669] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal, which then notifies the user by voice.
[0670] The user provides verbal feedback to the device, such as "The steak turned out delicious," and the device then sends that feedback to the server. The server uses the feedback it receives to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[0671] Specific examples and prompts for generative AI models
[0672] For example, if the emotion recognition engine determines that a user is tired from the tone of their voice and facial expression while cooking a steak, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without overdoing it."
[0673] Example prompt for a generative AI model:
[0674] Describe a situation where the emotion recognition engine determined from audio and video data that the user was stressed while cooking a steak, and what advice should be provided in that situation.
[0675] (Example prompt)
[0676] While the user is cooking a steak, suggest, "Today, let's try a simple steak recipe that doesn't require much effort." Also, give the user advice to relax, such as, "This heat level is fine. Please proceed without pushing yourself."
[0677] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and image sensor, taking into account the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0678] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0679] Step 1: Boot and Initialize the System
[0680] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine. The input is the power supply and communication connection for each device, and the output is that the device is ready to operate normally. This allows the server to load the cooking database and generative AI model. Specifically, it checks the operating status of the device and prepares for data collection.
[0681] Step 2: Waiting for user input
[0682] The device waits for input from the user and displays "Ready to cook. Please give me instructions to start cooking." The input is a voice command from the user, and the output is the display on the device and the activation of the voice assistant. Specifically, the device starts the voice recognition system and waits for the user's voice.
[0683] Step 3: Start a session with the user
[0684] The user speaks to the device, saying, "I'm going to start cooking." The device receives the user's voice input and asks, "Please tell me the name of the dish you want to cook." The input is the user's voice instruction, and the output is the device's question and data transmission to the server. Specifically, the voice recognition system converts the user's instruction into text and sends it to the server.
[0685] Step 4: Enter ingredients and recipe
[0686] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type of meat and thickness you would like to use." The input is the user's dish name information, and the output is a prompt to the terminal. The terminal asks the user aloud, "Please tell us the type of meat and thickness you would like to use." The input is a voice instruction from the user, and the output is data sent to the server. In concrete terms, the terminal uses a voice recognition system to collect user information and sends it to the server.
[0687] Step 5: Emotion Recognition with the Emotion Engine
[0688] The device sends the user's voice and facial image to the emotion recognition engine. The input is the user's voice data and video data, and the output is the emotion engine's analysis results. The emotion recognition engine analyzes this data and evaluates the user's emotional state. For example, if it determines that the user is feeling stressed, it sends this information to the server. Specifically, the emotion engine performs voice and facial expression analysis in real time.
[0689] Step 6: Collect and analyze sensor data
[0690] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The input is temperature data and image data, and the output is the analysis results. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that "one side is still raw." The tone and content of the cooking advice are adjusted taking into account the evaluation results of the emotion engine. Specifically, the server inputs the temperature data and image data into the analysis algorithm and obtains the results.
[0691] Step 7: Advice and coordination
[0692] The server determines the current cooking status based on the analysis results and sends advice such as "One side of the meat is still raw. Please continue cooking" to the terminal. The input is the analysis results and the emotion evaluation results, and the output is customized cooking advice. If the user's emotional state indicates stress, it also provides advice to help them relax, such as "It would be good to take a short break." In concrete terms, the server generates advice based on the cooking status and emotional state.
[0693] Step 8: Cooking complete and feedback
[0694] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The input is the analysis result, and the output is a notification that cooking is complete. The terminal notifies the user by voice that "The steak is done." The user provides feedback to the terminal by voice, saying "The steak turned out delicious." The input is the user's feedback, and the output is the feedback sent to the server. The feedback received by the server is used to improve the accuracy of the analysis method. Specifically, the server uses the feedback data as learning data for the generative AI model, improving the accuracy of advice from the next time onwards.
[0695] (Application example 2)
[0696] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0697] Conventional cooking support systems were able to detect the condition of ingredients and provide cooking advice, but they did not provide support that took the user's emotions into consideration. As a result, if the user felt stressed or tired while cooking, appropriate advice was not provided, resulting in a decrease in satisfaction. Furthermore, in the kitchen, it is necessary to detect the condition in real time and respond immediately, so flexible support that adapts to the user's condition is necessary.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0699] In this invention, the server includes a temperature sensor means for detecting the state of ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, an emotion analysis means for recognizing the user's emotions and adjusting advice on cooking methods and seasonings based on the emotions, and a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means and emotion analysis means. This enables flexible and appropriate cooking advice to be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[0700] A "temperature sensor that detects the state of ingredients" is a device that detects the temperature of ingredients in real time while they are being cooked.
[0701] The "camera for capturing images of ingredients" is a device for recording and detecting the visual state of ingredients during cooking.
[0702] The "analysis means" refers to devices or software that have the function of analyzing data from the temperature sensor and camera to evaluate the doneness and moisture content of the ingredients.
[0703] The "emotion analysis means" is a device or software that analyzes the user's voice and facial expressions to evaluate their emotional state and adjust the cooking method and seasoning advice.
[0704] The "notification means" is a device or software that has the function of providing the user with advice on cooking methods and seasonings in natural language based on the analysis results and emotion analysis results.
[0705] The "learning means" is a device or software that has the function of receiving feedback from users and storing and analyzing data to improve the accuracy of the analysis means.
[0706] "Real-time" means that processing and analysis are done almost immediately, and results are provided immediately.
[0707] An "emotion engine that recognizes user emotions" is an engine or software that evaluates the user's emotional state from their tone of voice and facial expressions and reflects that in the system.
[0708] The following describes the system configuration and processing details as an embodiment of this invention. The system is composed of a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means that analyzes the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, means for acquiring ingredient information and recipe information from the user, emotion analysis means that recognizes the user's emotions, and learning means that improves the accuracy of the analysis means based on feedback from the user.
[0709] First, when the system starts up, the server establishes connections with the temperature sensor, camera, and sentiment analysis engine, and loads the cooking database and generative AI model. The temperature sensor used is a DHT22, and the camera used is a Raspberry Pi camera module. This prepares the temperature sensor, camera, and sentiment analysis engine to operate normally.
[0710] The server notifies the terminal (a monitor device installed in the kitchen) that preparation is complete, and the terminal enters a state of waiting for input from the user. When the user issues a verbal instruction to start cooking, the terminal responds, listening to the name of the dish to be cooked and the types and amounts of ingredients to be used, and sending this to the server. When the server receives ingredient information and recipe information from the user, an analysis means determines the cooking procedure and generates the necessary advice.
[0711] During cooking, temperature sensors and cameras monitor the condition of the ingredients in real time and send the data to a server. The server analyzes the received data and evaluates the ingredients' doneness and moisture content. At the same time, an emotion analysis unit analyzes the user's emotional state from their voice tone and facial expressions, and if stress or fatigue is detected, the system adjusts the advice accordingly.
[0712] Based on the analysis results, the server generates cooking and seasoning advice and notifies the user via the device. For example, if one side of a steak is still raw while cooking, specific instructions such as "One side of the meat is still raw. Please continue cooking" are provided. Also, if the user shows signs of fatigue, advice encouraging relaxation such as "Take a short break. There is still time."
[0713] Once the cooking is complete, the server sends the results to the device, which notifies the user. When the user provides feedback, the server adds that feedback to its learning curve and uses it to improve the accuracy of future recommendations, allowing the system to provide more accurate recommendations with each use.
[0714] Examples:
[0715] Example prompt: "What cooking advice should you give if the user indicates that the meat is not yet cooked through?"
[0716] Example prompt: "If the user is determined to be tired, what relaxation advice should be offered?"
[0717] In this way, flexible and appropriate cooking advice can be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[0718] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0719] Step 1:
[0720] The server starts the system and establishes connections with the temperature sensor, camera, and sentiment analysis engine. This initializes the temperature sensor, camera, and sentiment analysis engine so that they can operate normally. The input is connection information from the temperature sensor and camera, and based on that, the server checks the device connection status and initializes it. The output is a notification that each device has been successfully connected and is available for use. Specifically, the server checks the response from each device and records the connection status in a log.
[0721] Step 2:
[0722] The user verbally commands the device to "start cooking." The input is the user's voice command, which is converted into text data using speech recognition software. The output is the converted text data, which is sent to the server. Specifically, the device picks up the user's voice with a microphone and converts it into text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[0723] Step 3:
[0724] The server confirms the "start cooking" instruction received from the user and sends a message to the terminal saying "Please tell us the name of the dish you want to cook." The input is the instruction to start cooking from the user, and the next question is determined based on that. The output is the next question displayed on the terminal and output as voice. In concrete terms, the server sends this question in digital data format to the terminal, and the terminal displays it to the user and outputs it as voice.
[0725] Step 4:
[0726] The user verbally instructs the terminal on the name of the dish to be cooked. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. In concrete terms, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[0727] Step 5:
[0728] The server analyzes the dish name information received from the user, and then sends a question to the terminal to request information on the necessary ingredients. Specifically, the input is text data for the dish name, and based on that, a question is generated to request ingredient information. As an output, this question is displayed and output as voice on the terminal. In concrete terms, the server retrieves the necessary information from a database that stores ingredient information, generates a question based on that, and sends it to the terminal.
[0729] Step 6:
[0730] The user verbally instructs the terminal on ingredient information. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. Specifically, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[0731] Step 7:
[0732] The device sends the user's voice and facial image to an emotion analysis engine. The input is the user's voice and video data, and the emotional state is evaluated based on this. The output is the emotion analysis results sent to the server. Specifically, the device captures the user's video and audio using the camera and microphone, and analyzes them using the emotion analysis engine (e.g., Affectiva SDK).
[0733] Step 8:
[0734] The server starts collecting data from the temperature sensor and camera and analyzes it in real time. The input is data from the temperature sensor and camera, and based on that data, it analyzes the doneness and moisture content of the ingredients. The output is an analysis result that evaluates the cooking status. Specifically, the server acquires data from the temperature sensor and camera and analyzes it using image processing software (e.g., OpenCV) and data analysis software (e.g., TensorFlow).
[0735] Step 9:
[0736] The server determines the current cooking status based on the analysis results and sends cooking advice to the device. If necessary, it also provides advice that takes into account the results of sentiment analysis. The inputs include analysis results from the temperature sensor and camera and sentiment analysis results, and cooking advice is generated based on these. The output is the advice displayed on the device and output as voice. Specifically, the server generates cooking advice using natural language generation software (e.g., Google Cloud Text-to-Speech) and sends it to the device.
[0737] Step 10:
[0738] The server determines whether cooking is complete and sends the result to the terminal. The input is the final analysis results from the temperature sensor and camera, and cooking completion is determined based on these. The output is a cooking completion message that is displayed and output aloud on the terminal. In concrete terms, the server evaluates the final analysis results, generates a message indicating that cooking is complete, and sends it to the terminal.
[0739] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0740] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0741] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0742] [Third embodiment]
[0743] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0744] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0745] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0746] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0747] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0748] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0749] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0750] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0751] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0752] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0753] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0754] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0755] This invention is a cooking assistance system that allows even beginners to easily prepare delicious dishes. This system detects the condition of ingredients in real time and provides appropriate cooking advice. The operation of this system is described in detail below.
[0756] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for obtaining ingredient and recipe information from the user, and a learning means that improves the accuracy of the analysis means based on user feedback.
[0757] Overview of program processing
[0758] System startup and initialization
[0759] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and AI model.
[0760] The device waits for a voice command from the user and displays "Ready. Give me instructions to start cooking."
[0761] Start a session with the user
[0762] The user inputs "I'm going to start cooking" into the terminal by voice.
[0763] The device will ask the user aloud, "Please tell us the name of the dish you would like to cook."
[0764] The user types "steak" and the terminal sends it to the server.
[0765] Enter ingredients and recipe
[0766] The server sends a message to the terminal asking, "Please tell us the type and thickness of meat you would like to use."
[0767] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0768] The user types in "sirloin, 2 cm thick," and the device sends this to the server.
[0769] Sensor data collection and analysis
[0770] The device starts collecting data from the temperature sensor and camera and sends the data to the server.
[0771] The server analyzes the temperature and image data collected in real time to evaluate the doneness and moisture content of the ingredients.
[0772] For example, if it determines that "one side is still raw," the device will notify the user, "One side of the meat is still raw. Please continue cooking."
[0773] Providing advice and coordination
[0774] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal.
[0775] For example, the system may notify the user that "One side is cooked, please turn the meat over," and the user should follow the instructions to turn the meat over.
[0776] The device continues to collect data from the sensors and the server re-analyzes the new data.
[0777] Cooking completion and feedback
[0778] The server determines that the steak is cooked through and sends the message "The steak is done" to the terminal.
[0779] The device notifies the user that "your steak is done."
[0780] The user inputs feedback such as "The steak turned out delicious" into the terminal, and the terminal transmits the feedback to the server.
[0781] The server records user feedback and uses it to improve the accuracy of the analysis method.
[0782] In this way, the cooking assistance system of the present invention analyzes data from temperature sensors and cameras to provide support that allows anyone to easily cook like a pro. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0783] The processing flow will be explained below.
[0784] Step 1:
[0785] The server starts the system and establishes the connection between the temperature sensor and the camera, so that the sensor and camera are ready to operate normally.
[0786] Step 2:
[0787] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0788] Step 3:
[0789] The user inputs "I'm going to start cooking" into the terminal by voice.
[0790] Step 4:
[0791] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0792] Step 5:
[0793] The user speaks "steak," which the device recognizes and sends to the server.
[0794] Step 6:
[0795] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0796] Step 7:
[0797] The device asks the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0798] Step 8:
[0799] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0800] Step 9:
[0801] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0802] Step 10:
[0803] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0804] Step 11:
[0805] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking."
[0806] Step 12:
[0807] The device continues to collect data from the temperature sensor and camera.
[0808] Step 13:
[0809] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0810] Step 14:
[0811] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0812] Step 15:
[0813] The user performs the action of turning the meat over.
[0814] Step 16:
[0815] The device continues to collect data and send it to the server.
[0816] Step 17:
[0817] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0818] Step 18:
[0819] The device will notify the user by voice, "Your steak is ready."
[0820] Step 19:
[0821] The user provides voice feedback to the device, such as "The steak turned out delicious."
[0822] Step 20:
[0823] The device sends the provided feedback to the server.
[0824] Step 21:
[0825] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[0826] These are the specific processing steps of the AI-powered cooking assistance system, which enables users to easily prepare delicious meals.
[0827] Example 1
[0828] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0829] Conventional cooking assistance systems have struggled to properly monitor the condition of ingredients and provide real-time advice based on the cooking progress. In particular, accurately measuring the doneness and moisture content of ingredients and providing cooking instructions at the appropriate time have been challenging. Furthermore, there has been a lack of technology to improve the accuracy of the system based on user feedback, leading to a demand for a system that allows even beginners to easily prepare delicious dishes.
[0830] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0831] In this invention, the server includes a temperature sensor that detects the state of the ingredients, a camera that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, notification means that provides advice on cooking methods and seasonings in natural language based on the analysis means, input means that accepts voice commands from the user, means that loads a cooking database and a generative AI model and analyzes based on data from the analysis means, and learning means that improves the accuracy of the analysis means based on feedback from the user. This makes it easy for anyone to cook like a professional.
[0832] A "temperature sensor" is a device that measures the surface and internal temperature of ingredients in real time and provides that data to the system.
[0833] The "camera" is a photographing device that captures images of ingredients and provides the visual data to the system.
[0834] The "analysis means" is a collection of algorithms and programs for evaluating the doneness and moisture content of ingredients based on data obtained from temperature sensors and cameras.
[0835] The "notification means" is a function for conveying the cooking method and seasoning advice obtained by the analysis means to the user in natural language.
[0836] "Input means" refers to an interface through which a user inputs voice commands into the system.
[0837] A "cooking database" is a data storage device that accumulates data and recipes related to various dishes.
[0838] A "generative AI model" is an artificial intelligence algorithm that uses collected data and past feedback to assist analytical methods and optimize cooking instructions.
[0839] The "learning means" is a function for improving the accuracy of the analysis means based on feedback from users.
[0840] "Ingredient information" is data such as the type, amount, and shape of ingredients used in cooking.
[0841] "Recipe information" refers to information such as the steps to make a particular dish, the ingredients needed, and cooking time.
[0842] A "user" is a person who operates the system and cooks according to the advice and instructions provided.
[0843] A "server" is a central processing unit that controls and analyzes the entire system.
[0844] A "terminal" is a device that provides an interface with a user and displays voice commands and notifications.
[0845] This invention is a cooking assistance system that allows even beginners to easily prepare delicious meals. The main components of the system include a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means for analyzing the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, input means for acquiring ingredient information and recipe information from the user, and learning means that improves the accuracy of the analysis means based on user feedback.
[0846] Hardware and Software Configuration
[0847] The server controls the entire system and receives data from temperature sensors (e.g., digital temperature sensors) and cameras (e.g., high-resolution cameras). The server loads a cooking database and generative AI models and uses them to analyze the data in real time. Specific processes performed by the server include collecting temperature data, analyzing image data, and evaluating the doneness and moisture content of ingredients.
[0848] A terminal is a device that provides an interface with the user and waits for voice commands. For example, it is used as a smart speaker or tablet. The terminal receives voice commands from the user and sends them to the server. It also has the function of notifying the user of advice from the server.
[0849] The user is the person who operates the system and proceeds with the cooking. The user inputs voice commands into the terminal and cooks according to the instructions from the system. Furthermore, the user provides feedback to the system after cooking is completed.
[0850] Specific examples
[0851] For example, if a user wants to cook a steak, the following series of processes takes place: The user speaks "I'd like to start cooking" into the device, and the device sends a request to the server. The server then asks the user via the device what dish they want to cook. The user answers "steak," and based on that, the server requests detailed information about the ingredients. The user answers "sirloin, 2 cm thick," and the device sends this to the server.
[0852] The temperature sensor and camera collect the steak's temperature and image data, which are then sent to a server. The server analyzes this data in real time to assess the steak's doneness and moisture content. Based on the results, the server provides advice to the user, such as "One side is still raw. Continue cooking."
[0853] Prompt Sentence Examples
[0854] "I want to grill some sirloin steak. What's the best way to check the doneness and make sure it's delicious?"
[0855] This process allows users to receive accurate cooking advice, enabling even beginners to cook like professionals. Furthermore, as the system continues to learn based on user feedback, it provides increasingly accurate advice.
[0856] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0857] Step 1: Boot and Initialize the System
[0858] The server starts the system. As part of the initialization, the server connects the temperature sensor and camera and checks their operation. Specifically, it acquires test data from the sensors and camera and verifies that they are working properly. The server also loads the cooking database and generative AI model, which will serve as the basis for future data analysis and advice provision.
[0859] Input: None
[0860] Output: System initialization complete, temperature sensor and camera operation confirmation, cooking database and AI model loading status
[0861] Step 2: Start a session with the user
[0862] The user speaks "I'm going to start cooking" into the device. The device recognizes this speech and sends a request to the server. The server then asks the user through the device, "What is the name of the dish you want to cook?" The user responds with "steak," and the device sends this to the server.
[0863] Input: User's voice command "Start cooking"
[0864] Output: Request for confirmation of dish name from server, dish name from user: "Steak"
[0865] Step 3: Enter ingredients and recipe
[0866] The server sends a message to the terminal saying, "Please tell us the type of meat you would like to use and its thickness." The terminal then asks the user aloud, "Please tell us the type of meat you would like to use and its thickness." The user enters "sirloin, 2 cm thick," and the terminal then sends this to the server.
[0867] Input: User's dish name "Steak", User's ingredient information "Sirloin, 2 cm thick"
[0868] Output: Register ingredient information and recipe information on the server
[0869] Step 4: Collect and analyze sensor data
[0870] The device starts collecting data from the temperature sensor and camera. The collected data is sent to the server in real time. The server analyzes the temperature and image data to evaluate the doneness and moisture content of the ingredients. For example, if the server determines that one side is still raw, it sends the result to the device.
[0871] Input: Real-time data from temperature sensors and cameras
[0872] Output: Analysis results, evaluation of doneness and moisture content
[0873] Step 5: Advice and coordination
[0874] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal. Specifically, it notifies the user, "One side is cooked, please turn the meat over." The user then turns the meat over as instructed.
[0875] Input: Analysis results of doneness and moisture content of ingredients
[0876] Output: Specific cooking advice to the user
[0877] Step 6: Cooking complete and feedback
[0878] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal notifies the user "The steak is done." The user enters feedback into the terminal, such as "The steak is delicious," and the terminal sends this to the server. The server records the user's feedback and uses it to improve the accuracy of the analysis method.
[0879] Input: Final analysis result of ingredients, user feedback "The steak turned out delicious."
[0880] Output: Notification of cooking completion, improvement of analytical methods based on feedback
[0881] Through the above process, this system allows even beginners to easily cook like a pro. Furthermore, by learning from user feedback, the system can provide more accurate cooking advice.
[0882] (Application example 1)
[0883] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0884] Maintaining the quality of food during delivery is difficult, especially if temperature control is inadequate, which can lead to loss of flavor and texture. To solve this problem, a system is needed that can monitor the status of food in real time and provide appropriate instructions and advice.
[0885] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0886] In this invention, the server includes a temperature sensor means for detecting the condition of the ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means, and a notification means for evaluating the condition of the food while it is being delivered and notifying the delivery person. This makes it possible to monitor the quality of the food in real time even during delivery, and to provide instructions on the optimal delivery timing and route.
[0887] A "temperature sensor" is a device that detects the temperature of ingredients in real time and provides that data to other analytical means.
[0888] The "camera" is a device for taking an image of the ingredients and transmitting the image data to the analysis means to evaluate the condition of the ingredients.
[0889] The "analysis means" is a means for evaluating the doneness and moisture content of ingredients using data obtained from the temperature sensor and camera.
[0890] The "notification means" is a device or function that provides advice on cooking methods and seasonings in natural language based on the analysis means.
[0891] The "notification means for evaluating the condition of food during delivery" is a function that monitors the condition of food during delivery and notifies the delivery person of the quality of the food and the appropriate delivery timing.
[0892] "Ingredient information" is detailed information such as the type, thickness, and amount of ingredients used in cooking.
[0893] "Recipe information" is information about specific cooking procedures and seasoning methods using ingredients.
[0894] "User feedback" refers to ratings and comments provided by users about the results of cooking and delivery.
[0895] "Learning means" refers to functions and algorithms that improve the accuracy of the analysis means based on user feedback.
[0896] This invention is a cooking and delivery assistance system that provides the delivery person with the appropriate delivery timing and route while maintaining the quality of the food. This system is composed of the following main components.
[0897] The system is equipped with a temperature sensor that detects the condition of ingredients in real time and a camera that captures images of the ingredients. These devices are connected to a smartphone or tablet to monitor the cooking status.
[0898] Next, there is the analytical means for analyzing the data from the temperature sensors and cameras. This analytical means resides on a cloud server and evaluates the doneness and moisture content of ingredients based on the collected temperature and image data. The analytical means used here include machine learning algorithms and generative AI models.
[0899] The system also includes a notification means that provides advice on cooking methods and seasonings in natural language based on the results of the analysis means. This notification means notifies the delivery person in real time via an application installed on their smartphone or tablet.
[0900] Specifically, when the system starts up, the server initializes the temperature sensor and camera, loads the cooking database and AI model, processes the data using analytical means, and sends the results to the delivery person via notification means.
[0901] For example, if a delivery person voice-inputs "I'm starting delivery," the system will ask, "Please tell me the name of the food you're delivering." If the delivery person inputs "pizza," that data will be sent to the server. The server then starts collecting data from the temperature sensor and camera, analyzes it, and evaluates the condition of the food. Based on the evaluation results, it will provide the delivery person with a notification, such as, "The pizza is at the right temperature. Please use the fastest route."
[0902] This method allows the system to provide optimal delivery timing and route information while maintaining the quality of the food. Furthermore, the system also has a learning mechanism for collecting feedback after delivery is completed and improving the accuracy of analysis in the future.
[0903] As a concrete example, the input prompt sentence for the generative AI model is as follows:
[0904] Based on the current temperature of the pizza, assess whether the food is of good quality.
[0905] In this way, it is possible to achieve quality control of food and optimal delivery support in food delivery.
[0906] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0907] Step 1:
[0908] System startup and initialization
[0909] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and generative AI model. After initialization is complete, the terminal displays the message, "Ready. Please give us your instructions to start delivery."
[0910] Step 2:
[0911] Start a session with the user
[0912] The user speaks "I'm starting delivery" into the device. The device converts this voice input into text and sends it to the server. The server receives the data and sends a voice command to the device asking, "Please tell me the name of the dish you want to deliver."
[0913] Step 3:
[0914] Enter dish information
[0915] The user speaks "pizza." The device converts this speech into text and sends it to the server. The server analyzes the received data and sends a voice prompt to the device asking, "Please provide information about the toppings you'll be using."
[0916] Step 4:
[0917] Sensor data collection and analysis
[0918] The device starts collecting data from the temperature sensor and camera. The temperature sensor captures the temperature data of the ingredients, and the camera captures images of the ingredients. This data is sent to the server in real time. The server analyzes the temperature data and image data to evaluate the doneness and moisture content of the ingredients.
[0919] Step 5:
[0920] Cooking status notification and delivery advice
[0921] The server evaluates the current state of the ingredients based on the analysis method. Based on the evaluation result, it notifies the device, "The pizza is at the right temperature. Please use the fastest route." The device then conveys this notification to the user via voice or text.
[0922] Step 6:
[0923] Continuous data collection and reanalysis
[0924] The device continuously collects data via temperature sensors and cameras and sends it to a server, which reanalyzes the new data and provides new advice if the cooking condition changes.
[0925] Step 7:
[0926] Delivery completion and feedback
[0927] Once the delivery is complete, the user voice-inputs "Delivery completed" into the device. The device sends this to the server, which updates the delivery status. The user then inputs feedback such as "The pizza arrived delicious," which the device sends to the server. The server analyzes the user's feedback and uses it as learning data to improve the accuracy of future advice.
[0928] Through these steps, the system can maintain food quality during delivery and notify customers of the optimal delivery time and route.
[0929] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0930] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[0931] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means that analyzes the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means of obtaining ingredient and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion engine that recognizes the user's emotions and adjusts cooking advice based on them.
[0932] Overview of program processing
[0933] System startup and initialization
[0934] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensors, camera, and emotion engine for normal operation. It then loads the cooking database and AI model.
[0935] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0936] Start a session with the user
[0937] The user inputs "I'm going to start cooking" into the terminal by voice.
[0938] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0939] The user speaks "steak," which the device recognizes and sends to the server.
[0940] Enter ingredients and recipe
[0941] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0942] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0943] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0944] Emotion recognition by emotion engine
[0945] The device sends the user's voice and facial image to the emotion engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[0946] Sensor data collection and analysis
[0947] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0948] The server analyzes the data it receives and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. It also adjusts the tone and content of cooking advice based on the emotion engine's evaluation results.
[0949] Providing advice and coordination
[0950] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, it may also provide advice to help them relax, such as "Take a short break."
[0951] Cooking completion and feedback
[0952] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal then notifies the user by voice, "The steak is done."
[0953] The user provides vocal feedback to the device, such as "The steak turned out delicious." The device then sends the provided feedback to the server. The server then uses the received feedback to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[0954] Specific examples
[0955] For example, if a user is cooking a steak and the system determines that the user is tired based on their tone of voice or facial expression, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without pushing yourself."
[0956] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and camera, taking into consideration the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[0957] The processing flow will be explained below.
[0958] Step 1:
[0959] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensor, camera, and emotion engine for normal operation.
[0960] Step 2:
[0961] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[0962] Step 3:
[0963] The user inputs "I'm going to start cooking" into the terminal by voice.
[0964] Step 4:
[0965] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[0966] Step 5:
[0967] The user speaks "steak," which the device recognizes and sends to the server.
[0968] Step 6:
[0969] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[0970] Step 7:
[0971] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[0972] Step 8:
[0973] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[0974] Step 9:
[0975] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[0976] Step 10:
[0977] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[0978] Step 11:
[0979] The emotion engine analyzes the user's voice and facial image to assess their emotional state, for example determining whether they are feeling stressed or tired.
[0980] Step 12:
[0981] Based on the analysis results and the user's emotional state, the server sends a notification to the device saying, "One side of the meat is still raw. Please continue cooking it." If the user is feeling stressed, the server adds, "It would be good to take a short break."
[0982] Step 13:
[0983] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking. It may be a good idea to take a short break."
[0984] Step 14:
[0985] The device continues to collect data from the temperature sensor and camera.
[0986] Step 15:
[0987] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[0988] Step 16:
[0989] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[0990] Step 17:
[0991] The user performs the action of turning the meat over.
[0992] Step 18:
[0993] The device continues to collect data and send it to the server.
[0994] Step 19:
[0995] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[0996] Step 20:
[0997] The device will notify the user by voice, "Your steak is ready."
[0998] Step 21:
[0999] The user provides voice feedback to the device, such as "The steak turned out delicious."
[1000] Step 22:
[1001] The device sends the provided feedback to the server.
[1002] Step 23:
[1003] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[1004] These are the specific processing steps of the cooking assistance system equipped with AI and an emotion engine, which enables users to easily prepare delicious meals and provides flexible support according to the user's emotional state.
[1005] Example 2
[1006] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1007] Conventional cooking assistance systems provide cooking advice based solely on the state of ingredients, and are unable to take the user's emotional state into account. Furthermore, they lack the support necessary to enable users without specialized cooking knowledge to easily cook like professionals. Furthermore, it is difficult to improve the accuracy of advice through continuous learning based on user feedback.
[1008] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1009] In this invention, the server includes a temperature sensor that detects the state of ingredients, an image sensor that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the image sensor to evaluate the doneness and moisture content of the ingredients, notification means that provides cooking method and seasoning advice in natural language based on the analysis means, an emotion recognition engine that recognizes the user's emotional state based on voice and video, and means for adjusting the tone and content of the cooking advice based on the evaluation of the emotion recognition engine. This makes it possible to provide cooking advice that takes the user's emotional state into consideration, and provides support that allows even users without specialized cooking expertise to easily cook like a professional. It is also possible to improve the accuracy of the advice by utilizing feedback from the user.
[1010] A "temperature sensor" is a device that measures the temperature of ingredients in real time and provides that data to an analytical means.
[1011] An "image sensor" is a device that captures images of ingredients and provides the visual data to an analysis means.
[1012] The "analysis means" is a system that evaluates the doneness and moisture content of ingredients based on data obtained from the temperature sensor and image sensor.
[1013] The "notification means" is a system that provides the user with advice on cooking methods and seasonings in natural language based on the results obtained from the analysis means.
[1014] An "emotion recognition engine" is a combination of software and hardware that analyzes a user's audio and video data and evaluates the user's emotional state.
[1015] The "adjustment means" is a system that appropriately changes the tone and content of cooking advice based on the emotion evaluation results obtained from the emotion recognition engine.
[1016] The "learning means" is a system that incorporates feedback from users into the analysis means to improve the accuracy of future cooking advice.
[1017] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[1018] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, an image sensor that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for acquiring ingredient information and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion recognition engine that recognizes the user's emotions and adjusts cooking advice based on this.
[1019] Hardware and Software Configuration
[1020] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine, ensuring all devices are ready to operate properly. The server also loads the cooking database and generative AI model.
[1021] The device waits for input from the user and displays "Ready to go. Give us your instructions to start cooking." At this point, the user can use voice input.
[1022] When the user speaks "I'm going to start cooking" into the device, the device asks "What is the name of the dish you want to cook?", and the user speaks "steak." The device then sends this information to the server.
[1023] Next, the server receives the name of the dish from the user and sends it to the terminal, asking, "Please tell us the type of meat and thickness you want to use," and the terminal asks the user questions based on that. The user voice-inputs, "Sirloin, 2 cm thick," and the terminal sends that information to the server.
[1024] Emotion Engine
[1025] The device sends the user's voice and facial image to the emotion recognition engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[1026] Sensor data collection and analysis
[1027] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. The server also adjusts the tone and content of the cooking advice based on the evaluation results of the emotion engine.
[1028] Providing advice and feedback
[1029] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, the server will also provide advice to help them relax, such as "It would be good to take a short break."
[1030] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal, which then notifies the user by voice.
[1031] The user provides verbal feedback to the device, such as "The steak turned out delicious," and the device then sends that feedback to the server. The server uses the feedback it receives to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[1032] Specific examples and prompts for generative AI models
[1033] For example, if the emotion recognition engine determines that a user is tired from the tone of their voice and facial expression while cooking a steak, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without overdoing it."
[1034] Example prompt for a generative AI model:
[1035] Describe a situation where the emotion recognition engine determined from audio and video data that the user was stressed while cooking a steak, and what advice should be provided in that situation.
[1036] (Example prompt)
[1037] While the user is cooking a steak, suggest, "Today, let's try a simple steak recipe that doesn't require much effort." Also, give the user advice to relax, such as, "This heat level is fine. Please proceed without pushing yourself."
[1038] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and image sensor, taking into account the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[1039] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1040] Step 1: Boot and Initialize the System
[1041] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine. The input is the power supply and communication connection for each device, and the output is that the device is ready to operate normally. This allows the server to load the cooking database and generative AI model. Specifically, it checks the operating status of the device and prepares for data collection.
[1042] Step 2: Waiting for user input
[1043] The device waits for input from the user and displays "Ready to cook. Please give me instructions to start cooking." The input is a voice command from the user, and the output is the display on the device and the activation of the voice assistant. Specifically, the device starts the voice recognition system and waits for the user's voice.
[1044] Step 3: Start a session with the user
[1045] The user speaks to the device, saying, "I'm going to start cooking." The device receives the user's voice input and asks, "Please tell me the name of the dish you want to cook." The input is the user's voice instruction, and the output is the device's question and data transmission to the server. Specifically, the voice recognition system converts the user's instruction into text and sends it to the server.
[1046] Step 4: Enter ingredients and recipe
[1047] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type of meat and thickness you would like to use." The input is the user's dish name information, and the output is a prompt to the terminal. The terminal asks the user aloud, "Please tell us the type of meat and thickness you would like to use." The input is a voice instruction from the user, and the output is data sent to the server. In concrete terms, the terminal uses a voice recognition system to collect user information and sends it to the server.
[1048] Step 5: Emotion Recognition with the Emotion Engine
[1049] The device sends the user's voice and facial image to the emotion recognition engine. The input is the user's voice data and video data, and the output is the emotion engine's analysis results. The emotion recognition engine analyzes this data and evaluates the user's emotional state. For example, if it determines that the user is feeling stressed, it sends this information to the server. Specifically, the emotion engine performs voice and facial expression analysis in real time.
[1050] Step 6: Collect and analyze sensor data
[1051] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The input is temperature data and image data, and the output is the analysis results. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that "one side is still raw." The tone and content of the cooking advice are adjusted taking into account the evaluation results of the emotion engine. Specifically, the server inputs the temperature data and image data into the analysis algorithm and obtains the results.
[1052] Step 7: Advice and coordination
[1053] The server determines the current cooking status based on the analysis results and sends advice such as "One side of the meat is still raw. Please continue cooking" to the terminal. The input is the analysis results and the emotion evaluation results, and the output is customized cooking advice. If the user's emotional state indicates stress, it also provides advice to help them relax, such as "It would be good to take a short break." In concrete terms, the server generates advice based on the cooking status and emotional state.
[1054] Step 8: Cooking complete and feedback
[1055] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The input is the analysis result, and the output is a notification that cooking is complete. The terminal notifies the user by voice that "The steak is done." The user provides feedback to the terminal by voice, saying "The steak turned out delicious." The input is the user's feedback, and the output is the feedback sent to the server. The feedback received by the server is used to improve the accuracy of the analysis method. Specifically, the server uses the feedback data as learning data for the generative AI model, improving the accuracy of advice from the next time onwards.
[1056] (Application example 2)
[1057] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1058] Conventional cooking support systems were able to detect the condition of ingredients and provide cooking advice, but they did not provide support that took the user's emotions into consideration. As a result, if the user felt stressed or tired while cooking, appropriate advice was not provided, resulting in a decrease in satisfaction. Furthermore, in the kitchen, it is necessary to detect the condition in real time and respond immediately, so flexible support that adapts to the user's condition is necessary.
[1059] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1060] In this invention, the server includes a temperature sensor means for detecting the state of ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, an emotion analysis means for recognizing the user's emotions and adjusting advice on cooking methods and seasonings based on the emotions, and a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means and emotion analysis means. This enables flexible and appropriate cooking advice to be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[1061] A "temperature sensor that detects the state of ingredients" is a device that detects the temperature of ingredients in real time while they are being cooked.
[1062] The "camera for capturing images of ingredients" is a device for recording and detecting the visual state of ingredients during cooking.
[1063] The "analysis means" refers to devices or software that have the function of analyzing data from the temperature sensor and camera to evaluate the doneness and moisture content of the ingredients.
[1064] The "emotion analysis means" is a device or software that analyzes the user's voice and facial expressions to evaluate their emotional state and adjust the cooking method and seasoning advice.
[1065] The "notification means" is a device or software that has the function of providing the user with advice on cooking methods and seasonings in natural language based on the analysis results and emotion analysis results.
[1066] The "learning means" is a device or software that has the function of receiving feedback from users and storing and analyzing data to improve the accuracy of the analysis means.
[1067] "Real-time" means that processing and analysis are done almost immediately, and results are provided immediately.
[1068] An "emotion engine that recognizes user emotions" is an engine or software that evaluates the user's emotional state from their tone of voice and facial expressions and reflects that in the system.
[1069] The following describes the system configuration and processing details as an embodiment of this invention. The system is composed of a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means that analyzes the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, means for acquiring ingredient information and recipe information from the user, emotion analysis means that recognizes the user's emotions, and learning means that improves the accuracy of the analysis means based on feedback from the user.
[1070] First, when the system starts up, the server establishes connections with the temperature sensor, camera, and sentiment analysis engine, and loads the cooking database and generative AI model. The temperature sensor used is a DHT22, and the camera used is a Raspberry Pi camera module. This prepares the temperature sensor, camera, and sentiment analysis engine to operate normally.
[1071] The server notifies the terminal (a monitor device installed in the kitchen) that preparation is complete, and the terminal enters a state of waiting for input from the user. When the user issues a verbal instruction to start cooking, the terminal responds, listening to the name of the dish to be cooked and the types and amounts of ingredients to be used, and sending this to the server. When the server receives ingredient information and recipe information from the user, an analysis means determines the cooking procedure and generates the necessary advice.
[1072] During cooking, temperature sensors and cameras monitor the condition of the ingredients in real time and send the data to a server. The server analyzes the received data and evaluates the ingredients' doneness and moisture content. At the same time, an emotion analysis unit analyzes the user's emotional state from their voice tone and facial expressions, and if stress or fatigue is detected, the system adjusts the advice accordingly.
[1073] Based on the analysis results, the server generates cooking and seasoning advice and notifies the user via the device. For example, if one side of a steak is still raw while cooking, specific instructions such as "One side of the meat is still raw. Please continue cooking" are provided. Also, if the user shows signs of fatigue, advice encouraging relaxation such as "Take a short break. There is still time."
[1074] Once the cooking is complete, the server sends the results to the device, which notifies the user. When the user provides feedback, the server adds that feedback to its learning curve and uses it to improve the accuracy of future recommendations, allowing the system to provide more accurate recommendations with each use.
[1075] Examples:
[1076] Example prompt: "What cooking advice should you give if the user indicates that the meat is not yet cooked through?"
[1077] Example prompt: "If the user is determined to be tired, what relaxation advice should be offered?"
[1078] In this way, flexible and appropriate cooking advice can be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[1079] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1080] Step 1:
[1081] The server starts the system and establishes connections with the temperature sensor, camera, and sentiment analysis engine. This initializes the temperature sensor, camera, and sentiment analysis engine so that they can operate normally. The input is connection information from the temperature sensor and camera, and based on that, the server checks the device connection status and initializes it. The output is a notification that each device has been successfully connected and is available for use. Specifically, the server checks the response from each device and records the connection status in a log.
[1082] Step 2:
[1083] The user verbally commands the device to "start cooking." The input is the user's voice command, which is converted into text data using speech recognition software. The output is the converted text data, which is sent to the server. Specifically, the device picks up the user's voice with a microphone and converts it into text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1084] Step 3:
[1085] The server confirms the "start cooking" instruction received from the user and sends a message to the terminal saying "Please tell us the name of the dish you want to cook." The input is the instruction to start cooking from the user, and the next question is determined based on that. The output is the next question displayed on the terminal and output as voice. In concrete terms, the server sends this question in digital data format to the terminal, and the terminal displays it to the user and outputs it as voice.
[1086] Step 4:
[1087] The user verbally instructs the terminal on the name of the dish to be cooked. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. In concrete terms, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[1088] Step 5:
[1089] The server analyzes the dish name information received from the user, and then sends a question to the terminal to request information on the necessary ingredients. Specifically, the input is text data for the dish name, and based on that, a question is generated to request ingredient information. As an output, this question is displayed and output as voice on the terminal. In concrete terms, the server retrieves the necessary information from a database that stores ingredient information, generates a question based on that, and sends it to the terminal.
[1090] Step 6:
[1091] The user verbally instructs the terminal on ingredient information. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. Specifically, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[1092] Step 7:
[1093] The device sends the user's voice and facial image to an emotion analysis engine. The input is the user's voice and video data, and the emotional state is evaluated based on this. The output is the emotion analysis results sent to the server. Specifically, the device captures the user's video and audio using the camera and microphone, and analyzes them using the emotion analysis engine (e.g., Affectiva SDK).
[1094] Step 8:
[1095] The server starts collecting data from the temperature sensor and camera and analyzes it in real time. The input is data from the temperature sensor and camera, and based on that data, it analyzes the doneness and moisture content of the ingredients. The output is an analysis result that evaluates the cooking status. Specifically, the server acquires data from the temperature sensor and camera and analyzes it using image processing software (e.g., OpenCV) and data analysis software (e.g., TensorFlow).
[1096] Step 9:
[1097] The server determines the current cooking status based on the analysis results and sends cooking advice to the device. If necessary, it also provides advice that takes into account the results of sentiment analysis. The inputs include analysis results from the temperature sensor and camera and sentiment analysis results, and cooking advice is generated based on these. The output is the advice displayed on the device and output as voice. Specifically, the server generates cooking advice using natural language generation software (e.g., Google Cloud Text-to-Speech) and sends it to the device.
[1098] Step 10:
[1099] The server determines whether cooking is complete and sends the result to the terminal. The input is the final analysis results from the temperature sensor and camera, and cooking completion is determined based on these. The output is a cooking completion message that is displayed and output aloud on the terminal. In concrete terms, the server evaluates the final analysis results, generates a message indicating that cooking is complete, and sends it to the terminal.
[1100] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1101] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1102] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1103] [Fourth embodiment]
[1104] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1105] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1106] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1107] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1108] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1109] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1110] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1111] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1112] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1113] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1114] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1115] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1116] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1117] This invention is a cooking assistance system that allows even beginners to easily prepare delicious dishes. This system detects the condition of ingredients in real time and provides appropriate cooking advice. The operation of this system is described in detail below.
[1118] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for obtaining ingredient and recipe information from the user, and a learning means that improves the accuracy of the analysis means based on user feedback.
[1119] Overview of program processing
[1120] System startup and initialization
[1121] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and AI model.
[1122] The device waits for a voice command from the user and displays "Ready. Give me instructions to start cooking."
[1123] Start a session with the user
[1124] The user inputs "I'm going to start cooking" into the terminal by voice.
[1125] The device will ask the user aloud, "Please tell us the name of the dish you would like to cook."
[1126] The user types "steak" and the terminal sends it to the server.
[1127] Enter ingredients and recipe
[1128] The server sends a message to the terminal asking, "Please tell us the type and thickness of meat you would like to use."
[1129] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[1130] The user types in "sirloin, 2 cm thick," and the device sends this to the server.
[1131] Sensor data collection and analysis
[1132] The device starts collecting data from the temperature sensor and camera and sends the data to the server.
[1133] The server analyzes the temperature and image data collected in real time to evaluate the doneness and moisture content of the ingredients.
[1134] For example, if it determines that "one side is still raw," the device will notify the user, "One side of the meat is still raw. Please continue cooking."
[1135] Providing advice and coordination
[1136] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal.
[1137] For example, the system may notify the user that "One side is cooked, please turn the meat over," and the user should follow the instructions to turn the meat over.
[1138] The device continues to collect data from the sensors and the server re-analyzes the new data.
[1139] Cooking completion and feedback
[1140] The server determines that the steak is cooked through and sends the message "The steak is done" to the terminal.
[1141] The device notifies the user that "your steak is done."
[1142] The user inputs feedback such as "The steak turned out delicious" into the terminal, and the terminal transmits the feedback to the server.
[1143] The server records user feedback and uses it to improve the accuracy of the analysis method.
[1144] In this way, the cooking assistance system of the present invention analyzes data from temperature sensors and cameras to provide support that allows anyone to easily cook like a pro. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[1145] The processing flow will be explained below.
[1146] Step 1:
[1147] The server starts the system and establishes the connection between the temperature sensor and the camera, so that the sensor and camera are ready to operate normally.
[1148] Step 2:
[1149] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[1150] Step 3:
[1151] The user inputs "I'm going to start cooking" into the terminal by voice.
[1152] Step 4:
[1153] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[1154] Step 5:
[1155] The user speaks "steak," which the device recognizes and sends to the server.
[1156] Step 6:
[1157] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[1158] Step 7:
[1159] The device asks the user aloud, "Please tell us the type and thickness of meat you would like to use."
[1160] Step 8:
[1161] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[1162] Step 9:
[1163] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[1164] Step 10:
[1165] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[1166] Step 11:
[1167] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking."
[1168] Step 12:
[1169] The device continues to collect data from the temperature sensor and camera.
[1170] Step 13:
[1171] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[1172] Step 14:
[1173] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[1174] Step 15:
[1175] The user performs the action of turning the meat over.
[1176] Step 16:
[1177] The device continues to collect data and send it to the server.
[1178] Step 17:
[1179] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[1180] Step 18:
[1181] The device will notify the user by voice, "Your steak is ready."
[1182] Step 19:
[1183] The user provides voice feedback to the device, such as "The steak turned out delicious."
[1184] Step 20:
[1185] The device sends the provided feedback to the server.
[1186] Step 21:
[1187] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[1188] These are the specific processing steps of the AI-powered cooking assistance system, which enables users to easily prepare delicious meals.
[1189] Example 1
[1190] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1191] Conventional cooking assistance systems have struggled to properly monitor the condition of ingredients and provide real-time advice based on the cooking progress. In particular, accurately measuring the doneness and moisture content of ingredients and providing cooking instructions at the appropriate time have been challenging. Furthermore, there has been a lack of technology to improve the accuracy of the system based on user feedback, leading to a demand for a system that allows even beginners to easily prepare delicious dishes.
[1192] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1193] In this invention, the server includes a temperature sensor that detects the state of the ingredients, a camera that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, notification means that provides advice on cooking methods and seasonings in natural language based on the analysis means, input means that accepts voice commands from the user, means that loads a cooking database and a generative AI model and analyzes based on data from the analysis means, and learning means that improves the accuracy of the analysis means based on feedback from the user. This makes it easy for anyone to cook like a professional.
[1194] A "temperature sensor" is a device that measures the surface and internal temperature of ingredients in real time and provides that data to the system.
[1195] The "camera" is a photographing device that captures images of ingredients and provides the visual data to the system.
[1196] The "analysis means" is a collection of algorithms and programs for evaluating the doneness and moisture content of ingredients based on data obtained from temperature sensors and cameras.
[1197] The "notification means" is a function for conveying the cooking method and seasoning advice obtained by the analysis means to the user in natural language.
[1198] "Input means" refers to an interface through which a user inputs voice commands into the system.
[1199] A "cooking database" is a data storage device that accumulates data and recipes related to various dishes.
[1200] A "generative AI model" is an artificial intelligence algorithm that uses collected data and past feedback to assist analytical methods and optimize cooking instructions.
[1201] The "learning means" is a function for improving the accuracy of the analysis means based on feedback from users.
[1202] "Ingredient information" is data such as the type, amount, and shape of ingredients used in cooking.
[1203] "Recipe information" refers to information such as the steps to make a particular dish, the ingredients needed, and cooking time.
[1204] A "user" is a person who operates the system and cooks according to the advice and instructions provided.
[1205] A "server" is a central processing unit that controls and analyzes the entire system.
[1206] A "terminal" is a device that provides an interface with a user and displays voice commands and notifications.
[1207] This invention is a cooking assistance system that allows even beginners to easily prepare delicious meals. The main components of the system include a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means for analyzing the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, input means for acquiring ingredient information and recipe information from the user, and learning means that improves the accuracy of the analysis means based on user feedback.
[1208] Hardware and Software Configuration
[1209] The server controls the entire system and receives data from temperature sensors (e.g., digital temperature sensors) and cameras (e.g., high-resolution cameras). The server loads a cooking database and generative AI models and uses them to analyze the data in real time. Specific processes performed by the server include collecting temperature data, analyzing image data, and evaluating the doneness and moisture content of ingredients.
[1210] A terminal is a device that provides an interface with the user and waits for voice commands. For example, it is used as a smart speaker or tablet. The terminal receives voice commands from the user and sends them to the server. It also has the function of notifying the user of advice from the server.
[1211] The user is the person who operates the system and proceeds with the cooking. The user inputs voice commands into the terminal and cooks according to the instructions from the system. Furthermore, the user provides feedback to the system after cooking is completed.
[1212] Specific examples
[1213] For example, if a user wants to cook a steak, the following series of processes takes place: The user speaks "I'd like to start cooking" into the device, and the device sends a request to the server. The server then asks the user via the device what dish they want to cook. The user answers "steak," and based on that, the server requests detailed information about the ingredients. The user answers "sirloin, 2 cm thick," and the device sends this to the server.
[1214] The temperature sensor and camera collect the steak's temperature and image data, which are then sent to a server. The server analyzes this data in real time to assess the steak's doneness and moisture content. Based on the results, the server provides advice to the user, such as "One side is still raw. Continue cooking."
[1215] Prompt Sentence Examples
[1216] "I want to grill some sirloin steak. What's the best way to check the doneness and make sure it's delicious?"
[1217] This process allows users to receive accurate cooking advice, enabling even beginners to cook like professionals. Furthermore, as the system continues to learn based on user feedback, it provides increasingly accurate advice.
[1218] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1219] Step 1: Boot and Initialize the System
[1220] The server starts the system. As part of the initialization, the server connects the temperature sensor and camera and checks their operation. Specifically, it acquires test data from the sensors and camera and verifies that they are working properly. The server also loads the cooking database and generative AI model, which will serve as the basis for future data analysis and advice provision.
[1221] Input: None
[1222] Output: System initialization complete, temperature sensor and camera operation confirmation, cooking database and AI model loading status
[1223] Step 2: Start a session with the user
[1224] The user speaks "I'm going to start cooking" into the device. The device recognizes this speech and sends a request to the server. The server then asks the user through the device, "What is the name of the dish you want to cook?" The user responds with "steak," and the device sends this to the server.
[1225] Input: User's voice command "Start cooking"
[1226] Output: Request for confirmation of dish name from server, dish name from user: "Steak"
[1227] Step 3: Enter ingredients and recipe
[1228] The server sends a message to the terminal saying, "Please tell us the type of meat you would like to use and its thickness." The terminal then asks the user aloud, "Please tell us the type of meat you would like to use and its thickness." The user enters "sirloin, 2 cm thick," and the terminal then sends this to the server.
[1229] Input: User's dish name "Steak", User's ingredient information "Sirloin, 2 cm thick"
[1230] Output: Register ingredient information and recipe information on the server
[1231] Step 4: Collect and analyze sensor data
[1232] The device starts collecting data from the temperature sensor and camera. The collected data is sent to the server in real time. The server analyzes the temperature and image data to evaluate the doneness and moisture content of the ingredients. For example, if the server determines that one side is still raw, it sends the result to the device.
[1233] Input: Real-time data from temperature sensors and cameras
[1234] Output: Analysis results, evaluation of doneness and moisture content
[1235] Step 5: Advice and coordination
[1236] The server determines the current cooking status based on the analysis results and issues instructions to the user via the terminal. Specifically, it notifies the user, "One side is cooked, please turn the meat over." The user then turns the meat over as instructed.
[1237] Input: Analysis results of doneness and moisture content of ingredients
[1238] Output: Specific cooking advice to the user
[1239] Step 6: Cooking complete and feedback
[1240] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal notifies the user "The steak is done." The user enters feedback into the terminal, such as "The steak is delicious," and the terminal sends this to the server. The server records the user's feedback and uses it to improve the accuracy of the analysis method.
[1241] Input: Final analysis result of ingredients, user feedback "The steak turned out delicious."
[1242] Output: Notification of cooking completion, improvement of analytical methods based on feedback
[1243] Through the above process, this system allows even beginners to easily cook like a pro. Furthermore, by learning from user feedback, the system can provide more accurate cooking advice.
[1244] (Application example 1)
[1245] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1246] Maintaining the quality of food during delivery is difficult, especially if temperature control is inadequate, which can lead to loss of flavor and texture. To solve this problem, a system is needed that can monitor the status of food in real time and provide appropriate instructions and advice.
[1247] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1248] In this invention, the server includes a temperature sensor means for detecting the condition of the ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means, and a notification means for evaluating the condition of the food while it is being delivered and notifying the delivery person. This makes it possible to monitor the quality of the food in real time even during delivery, and to provide instructions on the optimal delivery timing and route.
[1249] A "temperature sensor" is a device that detects the temperature of ingredients in real time and provides that data to other analytical means.
[1250] The "camera" is a device for taking an image of the ingredients and transmitting the image data to the analysis means to evaluate the condition of the ingredients.
[1251] The "analysis means" is a means for evaluating the doneness and moisture content of ingredients using data obtained from the temperature sensor and camera.
[1252] The "notification means" is a device or function that provides advice on cooking methods and seasonings in natural language based on the analysis means.
[1253] The "notification means for evaluating the condition of food during delivery" is a function that monitors the condition of food during delivery and notifies the delivery person of the quality of the food and the appropriate delivery timing.
[1254] "Ingredient information" is detailed information such as the type, thickness, and amount of ingredients used in cooking.
[1255] "Recipe information" is information about specific cooking procedures and seasoning methods using ingredients.
[1256] "User feedback" refers to ratings and comments provided by users about the results of cooking and delivery.
[1257] "Learning means" refers to functions and algorithms that improve the accuracy of the analysis means based on user feedback.
[1258] This invention is a cooking and delivery assistance system that provides the delivery person with the appropriate delivery timing and route while maintaining the quality of the food. This system is composed of the following main components.
[1259] The system is equipped with a temperature sensor that detects the condition of ingredients in real time and a camera that captures images of the ingredients. These devices are connected to a smartphone or tablet to monitor the cooking status.
[1260] Next, there is the analytical means for analyzing the data from the temperature sensors and cameras. This analytical means resides on a cloud server and evaluates the doneness and moisture content of ingredients based on the collected temperature and image data. The analytical means used here include machine learning algorithms and generative AI models.
[1261] The system also includes a notification means that provides advice on cooking methods and seasonings in natural language based on the results of the analysis means. This notification means notifies the delivery person in real time via an application installed on their smartphone or tablet.
[1262] Specifically, when the system starts up, the server initializes the temperature sensor and camera, loads the cooking database and AI model, processes the data using analytical means, and sends the results to the delivery person via notification means.
[1263] For example, if a delivery person voice-inputs "I'm starting delivery," the system will ask, "Please tell me the name of the food you're delivering." If the delivery person inputs "pizza," that data will be sent to the server. The server then starts collecting data from the temperature sensor and camera, analyzes it, and evaluates the condition of the food. Based on the evaluation results, it will provide the delivery person with a notification, such as, "The pizza is at the right temperature. Please use the fastest route."
[1264] This method allows the system to provide optimal delivery timing and route information while maintaining the quality of the food. Furthermore, the system also has a learning mechanism for collecting feedback after delivery is completed and improving the accuracy of analysis in the future.
[1265] As a concrete example, the input prompt sentence for the generative AI model is as follows:
[1266] Based on the current temperature of the pizza, assess whether the food is of good quality.
[1267] In this way, it is possible to achieve quality control of food and optimal delivery support in food delivery.
[1268] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1269] Step 1:
[1270] System startup and initialization
[1271] The server starts the system, initializes the temperature sensor and camera, and loads the cooking database and generative AI model. After initialization is complete, the terminal displays the message, "Ready. Please give us your instructions to start delivery."
[1272] Step 2:
[1273] Start a session with the user
[1274] The user speaks "I'm starting delivery" into the device. The device converts this voice input into text and sends it to the server. The server receives the data and sends a voice command to the device asking, "Please tell me the name of the dish you want to deliver."
[1275] Step 3:
[1276] Enter dish information
[1277] The user speaks "pizza." The device converts this speech into text and sends it to the server. The server analyzes the received data and sends a voice prompt to the device asking, "Please provide information about the toppings you'll be using."
[1278] Step 4:
[1279] Sensor data collection and analysis
[1280] The device starts collecting data from the temperature sensor and camera. The temperature sensor captures the temperature data of the ingredients, and the camera captures images of the ingredients. This data is sent to the server in real time. The server analyzes the temperature data and image data to evaluate the doneness and moisture content of the ingredients.
[1281] Step 5:
[1282] Cooking status notification and delivery advice
[1283] The server evaluates the current state of the ingredients based on the analysis method. Based on the evaluation result, it notifies the device, "The pizza is at the right temperature. Please use the fastest route." The device then conveys this notification to the user via voice or text.
[1284] Step 6:
[1285] Continuous data collection and reanalysis
[1286] The device continuously collects data via temperature sensors and cameras and sends it to a server, which reanalyzes the new data and provides new advice if the cooking condition changes.
[1287] Step 7:
[1288] Delivery completion and feedback
[1289] Once the delivery is complete, the user voice-inputs "Delivery completed" into the device. The device sends this to the server, which updates the delivery status. The user then inputs feedback such as "The pizza arrived delicious," which the device sends to the server. The server analyzes the user's feedback and uses it as learning data to improve the accuracy of future advice.
[1290] Through these steps, the system can maintain food quality during delivery and notify customers of the optimal delivery time and route.
[1291] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1292] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[1293] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, a camera that captures images of the ingredients, an analysis means that analyzes the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means of obtaining ingredient and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion engine that recognizes the user's emotions and adjusts cooking advice based on them.
[1294] Overview of program processing
[1295] System startup and initialization
[1296] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensors, camera, and emotion engine for normal operation. It then loads the cooking database and AI model.
[1297] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[1298] Start a session with the user
[1299] The user inputs "I'm going to start cooking" into the terminal by voice.
[1300] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[1301] The user speaks "steak," which the device recognizes and sends to the server.
[1302] Enter ingredients and recipe
[1303] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[1304] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[1305] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[1306] Emotion recognition by emotion engine
[1307] The device sends the user's voice and facial image to the emotion engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[1308] Sensor data collection and analysis
[1309] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[1310] The server analyzes the data it receives and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. It also adjusts the tone and content of cooking advice based on the emotion engine's evaluation results.
[1311] Providing advice and coordination
[1312] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, it may also provide advice to help them relax, such as "Take a short break."
[1313] Cooking completion and feedback
[1314] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The terminal then notifies the user by voice, "The steak is done."
[1315] The user provides vocal feedback to the device, such as "The steak turned out delicious." The device then sends the provided feedback to the server. The server then uses the received feedback to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[1316] Specific examples
[1317] For example, if a user is cooking a steak and the system determines that the user is tired based on their tone of voice or facial expression, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without pushing yourself."
[1318] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and camera, taking into consideration the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[1319] The processing flow will be explained below.
[1320] Step 1:
[1321] The server starts the system and establishes connections with the temperature sensor, camera, and emotion engine, which then prepares the sensor, camera, and emotion engine for normal operation.
[1322] Step 2:
[1323] The terminal waits for input from the user and displays "Ready, give me instructions to start cooking."
[1324] Step 3:
[1325] The user inputs "I'm going to start cooking" into the terminal by voice.
[1326] Step 4:
[1327] The device receives voice input from the user and asks, "Please tell us the name of the dish you would like to cook."
[1328] Step 5:
[1329] The user speaks "steak," which the device recognizes and sends to the server.
[1330] Step 6:
[1331] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type and thickness of meat you would like to use."
[1332] Step 7:
[1333] The device will ask the user aloud, "Please tell us the type and thickness of meat you would like to use."
[1334] Step 8:
[1335] The user speaks "sirloin, 2 cm thick," and the device sends this to the server.
[1336] Step 9:
[1337] The device starts collecting data from the temperature sensor and camera and sends the data to the server in real time.
[1338] Step 10:
[1339] The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it determines that one side is still raw.
[1340] Step 11:
[1341] The emotion engine analyzes the user's voice and facial image to assess their emotional state, for example determining whether they are feeling stressed or tired.
[1342] Step 12:
[1343] Based on the analysis results and the user's emotional state, the server sends a notification to the device saying, "One side of the meat is still raw. Please continue cooking it." If the user is feeling stressed, the server adds, "It would be good to take a short break."
[1344] Step 13:
[1345] The device will notify the user by voice, "One side of the meat is still raw. Please continue cooking. It may be a good idea to take a short break."
[1346] Step 14:
[1347] The device continues to collect data from the temperature sensor and camera.
[1348] Step 15:
[1349] The server reanalyzes the new data in real time and determines, "One side is cooked, please flip the meat over."
[1350] Step 16:
[1351] The device will notify the user by voice, "One side is cooked, please turn the meat over."
[1352] Step 17:
[1353] The user performs the action of turning the meat over.
[1354] Step 18:
[1355] The device continues to collect data and send it to the server.
[1356] Step 19:
[1357] The server analyzes and determines that the steak is cooked through, and sends the result to the terminal saying, "The steak is done."
[1358] Step 20:
[1359] The device will notify the user by voice, "Your steak is ready."
[1360] Step 21:
[1361] The user provides voice feedback to the device, such as "The steak turned out delicious."
[1362] Step 22:
[1363] The device sends the provided feedback to the server.
[1364] Step 23:
[1365] The server uses the feedback it receives to improve the accuracy of its analysis methods, and the system learns from this process to improve the accuracy of future cooking advice.
[1366] These are the specific processing steps of the cooking assistance system equipped with AI and an emotion engine, which enables users to easily prepare delicious meals and provides flexible support according to the user's emotional state.
[1367] Example 2
[1368] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1369] Conventional cooking assistance systems provide cooking advice based solely on the state of ingredients, and are unable to take the user's emotional state into account. Furthermore, they lack the support necessary to enable users without specialized cooking knowledge to easily cook like professionals. Furthermore, it is difficult to improve the accuracy of advice through continuous learning based on user feedback.
[1370] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1371] In this invention, the server includes a temperature sensor that detects the state of ingredients, an image sensor that captures images of the ingredients, analysis means that analyzes data from the temperature sensor and the image sensor to evaluate the doneness and moisture content of the ingredients, notification means that provides cooking method and seasoning advice in natural language based on the analysis means, an emotion recognition engine that recognizes the user's emotional state based on voice and video, and means for adjusting the tone and content of the cooking advice based on the evaluation of the emotion recognition engine. This makes it possible to provide cooking advice that takes the user's emotional state into consideration, and provides support that allows even users without specialized cooking expertise to easily cook like a professional. It is also possible to improve the accuracy of the advice by utilizing feedback from the user.
[1372] A "temperature sensor" is a device that measures the temperature of ingredients in real time and provides that data to an analytical means.
[1373] An "image sensor" is a device that captures images of ingredients and provides the visual data to an analysis means.
[1374] The "analysis means" is a system that evaluates the doneness and moisture content of ingredients based on data obtained from the temperature sensor and image sensor.
[1375] The "notification means" is a system that provides the user with advice on cooking methods and seasonings in natural language based on the results obtained from the analysis means.
[1376] An "emotion recognition engine" is a combination of software and hardware that analyzes a user's audio and video data and evaluates the user's emotional state.
[1377] The "adjustment means" is a system that appropriately changes the tone and content of cooking advice based on the emotion evaluation results obtained from the emotion recognition engine.
[1378] The "learning means" is a system that incorporates feedback from users into the analysis means to improve the accuracy of future cooking advice.
[1379] This invention realizes user-friendly cooking support by combining a cooking assistance system that detects the state of ingredients in real time and provides appropriate cooking advice with an emotion engine that recognizes the user's emotions. The operation of this system is described in detail below.
[1380] The system consists of the following main components: a temperature sensor that detects the condition of the ingredients, an image sensor that captures images of the ingredients, an analysis means for analyzing the obtained data, a notification means that provides cooking methods and seasonings in natural language based on the analysis results, a means for acquiring ingredient information and recipe information from the user, a learning means that improves the accuracy of the analysis means based on user feedback, and an emotion recognition engine that recognizes the user's emotions and adjusts cooking advice based on this.
[1381] Hardware and Software Configuration
[1382] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine, ensuring all devices are ready to operate properly. The server also loads the cooking database and generative AI model.
[1383] The device waits for input from the user and displays "Ready to go. Give us your instructions to start cooking." At this point, the user can use voice input.
[1384] When the user speaks "I'm going to start cooking" into the device, the device asks "What is the name of the dish you want to cook?", and the user speaks "steak." The device then sends this information to the server.
[1385] Next, the server receives the name of the dish from the user and sends it to the terminal, asking, "Please tell us the type of meat and thickness you want to use," and the terminal asks the user questions based on that. The user voice-inputs, "Sirloin, 2 cm thick," and the terminal sends that information to the server.
[1386] Emotion Engine
[1387] The device sends the user's voice and facial image to the emotion recognition engine, which analyzes this data and evaluates the user's emotional state. For example, if the user is feeling stressed, the emotion engine sends that information to the server.
[1388] Sensor data collection and analysis
[1389] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that one side is still raw. The server also adjusts the tone and content of the cooking advice based on the evaluation results of the emotion engine.
[1390] Providing advice and feedback
[1391] The server determines the current cooking status based on the analysis results and sends advice to the device, such as "One side of the meat is still raw. Please continue cooking." If the user's emotional state indicates stress, the server will also provide advice to help them relax, such as "It would be good to take a short break."
[1392] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal, which then notifies the user by voice.
[1393] The user provides verbal feedback to the device, such as "The steak turned out delicious," and the device then sends that feedback to the server. The server uses the feedback it receives to improve the accuracy of its analysis methods. The system learns from this process and improves the accuracy of future cooking advice.
[1394] Specific examples and prompts for generative AI models
[1395] For example, if the emotion recognition engine determines that a user is tired from the tone of their voice and facial expression while cooking a steak, it will suggest, "Today, try a simple steak recipe that doesn't require much effort." It will also provide advice during the cooking process to encourage relaxation, such as, "This heat level is fine. Please proceed without overdoing it."
[1396] Example prompt for a generative AI model:
[1397] Describe a situation where the emotion recognition engine determined from audio and video data that the user was stressed while cooking a steak, and what advice should be provided in that situation.
[1398] (Example prompt)
[1399] While the user is cooking a steak, suggest, "Today, let's try a simple steak recipe that doesn't require much effort." Also, give the user advice to relax, such as, "This heat level is fine. Please proceed without pushing yourself."
[1400] In this way, the cooking assistance system of the present invention analyzes data from the temperature sensor and image sensor, taking into account the user's emotions, and provides support that allows anyone to easily cook like a professional. Furthermore, the system continues to learn based on user feedback, making it possible to provide advice with even greater accuracy.
[1401] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1402] Step 1: Boot and Initialize the System
[1403] The server starts the system and establishes connections with the temperature sensor, image sensor, and emotion recognition engine. The input is the power supply and communication connection for each device, and the output is that the device is ready to operate normally. This allows the server to load the cooking database and generative AI model. Specifically, it checks the operating status of the device and prepares for data collection.
[1404] Step 2: Waiting for user input
[1405] The device waits for input from the user and displays "Ready to cook. Please give me instructions to start cooking." The input is a voice command from the user, and the output is the display on the device and the activation of the voice assistant. Specifically, the device starts the voice recognition system and waits for the user's voice.
[1406] Step 3: Start a session with the user
[1407] The user speaks to the device, saying, "I'm going to start cooking." The device receives the user's voice input and asks, "Please tell me the name of the dish you want to cook." The input is the user's voice instruction, and the output is the device's question and data transmission to the server. Specifically, the voice recognition system converts the user's instruction into text and sends it to the server.
[1408] Step 4: Enter ingredients and recipe
[1409] The server receives the dish name information from the user and sends it to the terminal, "Please tell us the type of meat and thickness you would like to use." The input is the user's dish name information, and the output is a prompt to the terminal. The terminal asks the user aloud, "Please tell us the type of meat and thickness you would like to use." The input is a voice instruction from the user, and the output is data sent to the server. In concrete terms, the terminal uses a voice recognition system to collect user information and sends it to the server.
[1410] Step 5: Emotion Recognition with the Emotion Engine
[1411] The device sends the user's voice and facial image to the emotion recognition engine. The input is the user's voice data and video data, and the output is the emotion engine's analysis results. The emotion recognition engine analyzes this data and evaluates the user's emotional state. For example, if it determines that the user is feeling stressed, it sends this information to the server. Specifically, the emotion engine performs voice and facial expression analysis in real time.
[1412] Step 6: Collect and analyze sensor data
[1413] The device starts collecting data from the temperature sensor and image sensor, and sends the data to the server in real time. The input is temperature data and image data, and the output is the analysis results. The server analyzes the received data and evaluates the doneness and moisture content of the ingredients. For example, it may determine that "one side is still raw." The tone and content of the cooking advice are adjusted taking into account the evaluation results of the emotion engine. Specifically, the server inputs the temperature data and image data into the analysis algorithm and obtains the results.
[1414] Step 7: Advice and coordination
[1415] The server determines the current cooking status based on the analysis results and sends advice such as "One side of the meat is still raw. Please continue cooking" to the terminal. The input is the analysis results and the emotion evaluation results, and the output is customized cooking advice. If the user's emotional state indicates stress, it also provides advice to help them relax, such as "It would be good to take a short break." In concrete terms, the server generates advice based on the cooking status and emotional state.
[1416] Step 8: Cooking complete and feedback
[1417] The server determines that the steak is cooked through and sends the result "The steak is done" to the terminal. The input is the analysis result, and the output is a notification that cooking is complete. The terminal notifies the user by voice that "The steak is done." The user provides feedback to the terminal by voice, saying "The steak turned out delicious." The input is the user's feedback, and the output is the feedback sent to the server. The feedback received by the server is used to improve the accuracy of the analysis method. Specifically, the server uses the feedback data as learning data for the generative AI model, improving the accuracy of advice from the next time onwards.
[1418] (Application example 2)
[1419] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1420] Conventional cooking support systems were able to detect the condition of ingredients and provide cooking advice, but they did not provide support that took the user's emotions into consideration. As a result, if the user felt stressed or tired while cooking, appropriate advice was not provided, resulting in a decrease in satisfaction. Furthermore, in the kitchen, it is necessary to detect the condition in real time and respond immediately, so flexible support that adapts to the user's condition is necessary.
[1421] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1422] In this invention, the server includes a temperature sensor means for detecting the state of ingredients, a camera means for capturing images of the ingredients, an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of the ingredients, an emotion analysis means for recognizing the user's emotions and adjusting advice on cooking methods and seasonings based on the emotions, and a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means and emotion analysis means. This enables flexible and appropriate cooking advice to be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[1423] A "temperature sensor that detects the state of ingredients" is a device that detects the temperature of ingredients in real time while they are being cooked.
[1424] The "camera for capturing images of ingredients" is a device for recording and detecting the visual state of ingredients during cooking.
[1425] The "analysis means" refers to devices or software that have the function of analyzing data from the temperature sensor and camera to evaluate the doneness and moisture content of the ingredients.
[1426] The "emotion analysis means" is a device or software that analyzes the user's voice and facial expressions to evaluate their emotional state and adjust the cooking method and seasoning advice.
[1427] The "notification means" is a device or software that has the function of providing the user with advice on cooking methods and seasonings in natural language based on the analysis results and emotion analysis results.
[1428] The "learning means" is a device or software that has the function of receiving feedback from users and storing and analyzing data to improve the accuracy of the analysis means.
[1429] "Real-time" means that processing and analysis are done almost immediately, and results are provided immediately.
[1430] An "emotion engine that recognizes user emotions" is an engine or software that evaluates the user's emotional state from their tone of voice and facial expressions and reflects that in the system.
[1431] The following describes the system configuration and processing details as an embodiment of this invention. The system is composed of a temperature sensor that detects the state of ingredients, a camera that captures images of the ingredients, analysis means that analyzes the obtained data, notification means that provides cooking methods and seasonings in natural language based on the analysis results, means for acquiring ingredient information and recipe information from the user, emotion analysis means that recognizes the user's emotions, and learning means that improves the accuracy of the analysis means based on feedback from the user.
[1432] First, when the system starts up, the server establishes connections with the temperature sensor, camera, and sentiment analysis engine, and loads the cooking database and generative AI model. The temperature sensor used is a DHT22, and the camera used is a Raspberry Pi camera module. This prepares the temperature sensor, camera, and sentiment analysis engine to operate normally.
[1433] The server notifies the terminal (a monitor device installed in the kitchen) that preparation is complete, and the terminal enters a state of waiting for input from the user. When the user issues a verbal instruction to start cooking, the terminal responds, listening to the name of the dish to be cooked and the types and amounts of ingredients to be used, and sending this to the server. When the server receives ingredient information and recipe information from the user, an analysis means determines the cooking procedure and generates the necessary advice.
[1434] During cooking, temperature sensors and cameras monitor the condition of the ingredients in real time and send the data to a server. The server analyzes the received data and evaluates the ingredients' doneness and moisture content. At the same time, an emotion analysis unit analyzes the user's emotional state from their voice tone and facial expressions, and if stress or fatigue is detected, the system adjusts the advice accordingly.
[1435] Based on the analysis results, the server generates cooking and seasoning advice and notifies the user via the device. For example, if one side of a steak is still raw while cooking, specific instructions such as "One side of the meat is still raw. Please continue cooking" are provided. Also, if the user shows signs of fatigue, advice encouraging relaxation such as "Take a short break. There is still time."
[1436] Once the cooking is complete, the server sends the results to the device, which notifies the user. When the user provides feedback, the server adds that feedback to its learning curve and uses it to improve the accuracy of future recommendations, allowing the system to provide more accurate recommendations with each use.
[1437] Examples:
[1438] Example prompt: "What cooking advice should you give if the user indicates that the meat is not yet cooked through?"
[1439] Example prompt: "If the user is determined to be tired, what relaxation advice should be offered?"
[1440] In this way, flexible and appropriate cooking advice can be given according to the user's emotional state, which is expected to improve cooking efficiency and satisfaction.
[1441] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1442] Step 1:
[1443] The server starts the system and establishes connections with the temperature sensor, camera, and sentiment analysis engine. This initializes the temperature sensor, camera, and sentiment analysis engine so that they can operate normally. The input is connection information from the temperature sensor and camera, and based on that, the server checks the device connection status and initializes it. The output is a notification that each device has been successfully connected and is available for use. Specifically, the server checks the response from each device and records the connection status in a log.
[1444] Step 2:
[1445] The user verbally commands the device to "start cooking." The input is the user's voice command, which is converted into text data using speech recognition software. The output is the converted text data, which is sent to the server. Specifically, the device picks up the user's voice with a microphone and converts it into text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1446] Step 3:
[1447] The server confirms the "start cooking" instruction received from the user and sends a message to the terminal saying "Please tell us the name of the dish you want to cook." The input is the instruction to start cooking from the user, and the next question is determined based on that. The output is the next question displayed on the terminal and output as voice. In concrete terms, the server sends this question in digital data format to the terminal, and the terminal displays it to the user and outputs it as voice.
[1448] Step 4:
[1449] The user verbally instructs the terminal on the name of the dish to be cooked. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. In concrete terms, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[1450] Step 5:
[1451] The server analyzes the dish name information received from the user, and then sends a question to the terminal to request information on the necessary ingredients. Specifically, the input is text data for the dish name, and based on that, a question is generated to request ingredient information. As an output, this question is displayed and output as voice on the terminal. In concrete terms, the server retrieves the necessary information from a database that stores ingredient information, generates a question based on that, and sends it to the terminal.
[1452] Step 6:
[1453] The user verbally instructs the terminal on ingredient information. The input is the user's voice instruction, which is converted into text data using voice recognition software. The output is the converted text data, which is sent to the server. Specifically, the terminal picks up the user's voice with a microphone and converts it into text using voice recognition software.
[1454] Step 7:
[1455] The device sends the user's voice and facial image to an emotion analysis engine. The input is the user's voice and video data, and the emotional state is evaluated based on this. The output is the emotion analysis results sent to the server. Specifically, the device captures the user's video and audio using the camera and microphone, and analyzes them using the emotion analysis engine (e.g., Affectiva SDK).
[1456] Step 8:
[1457] The server starts collecting data from the temperature sensor and camera and analyzes it in real time. The input is data from the temperature sensor and camera, and based on that data, it analyzes the doneness and moisture content of the ingredients. The output is an analysis result that evaluates the cooking status. Specifically, the server acquires data from the temperature sensor and camera and analyzes it using image processing software (e.g., OpenCV) and data analysis software (e.g., TensorFlow).
[1458] Step 9:
[1459] The server determines the current cooking status based on the analysis results and sends cooking advice to the device. If necessary, it also provides advice that takes into account the results of sentiment analysis. The inputs include analysis results from the temperature sensor and camera and sentiment analysis results, and cooking advice is generated based on these. The output is the advice displayed on the device and output as voice. Specifically, the server generates cooking advice using natural language generation software (e.g., Google Cloud Text-to-Speech) and sends it to the device.
[1460] Step 10:
[1461] The server determines whether cooking is complete and sends the result to the terminal. The input is the final analysis results from the temperature sensor and camera, and cooking completion is determined based on these. The output is a cooking completion message that is displayed and output aloud on the terminal. In concrete terms, the server evaluates the final analysis results, generates a message indicating that cooking is complete, and sends it to the terminal.
[1462] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1463] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1464] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1465] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1466] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1467] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1468] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1469] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1470] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1471] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1472] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1473] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1474] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1475] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1476] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1477] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1478] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1479] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1480] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1481] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1482] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1483] The following is further disclosed regarding the above embodiment.
[1484] (Claim 1)
[1485] A temperature sensor that detects the condition of the ingredients,
[1486] a camera for capturing images of ingredients;
[1487] an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of ingredients;
[1488] a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means;
[1489] A system including:
[1490] (Claim 2)
[1491] 2. The system according to claim 1, wherein ingredient information and recipe information are obtained from a user, and the analysis means determines a cooking procedure based on this information.
[1492] (Claim 3)
[1493] 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice.
[1494] "Example 1"
[1495] (Claim 1)
[1496] A temperature sensor that detects the condition of the ingredients,
[1497] a camera for capturing images of ingredients;
[1498] an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of ingredients;
[1499] a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means;
[1500] input means for accepting voice commands from a user;
[1501] A means for loading a cooking database and a generating AI model and analyzing the data based on the analysis means;
[1502] A learning method to improve the accuracy of the analysis method based on user feedback.
[1503] A system including:
[1504] (Claim 2)
[1505] 2. The system according to claim 1, wherein ingredient information and recipe information are obtained from a user, and the analysis means determines a cooking procedure based on this information.
[1506] (Claim 3)
[1507] 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice.
[1508] "Application Example 1"
[1509] (Claim 1)
[1510] A temperature sensor that detects the condition of the ingredients,
[1511] a camera for capturing images of ingredients;
[1512] an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of ingredients;
[1513] a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means;
[1514] a notification means for evaluating the status of the food during delivery and notifying the delivery person;
[1515] A system including:
[1516] (Claim 2)
[1517] 2. The system according to claim 1, wherein ingredient information and recipe information are obtained from a user, and the analysis means determines a cooking procedure based on this information.
[1518] (Claim 3)
[1519] 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice.
[1520] "Example 2: Combining Emotion Engines"
[1521] (Claim 1)
[1522] A temperature sensor that detects the condition of the ingredients,
[1523] an image sensor for acquiring an image of the ingredients;
[1524] an analysis means for analyzing data from the temperature sensor and the image sensor to evaluate the doneness and moisture content of the ingredients;
[1525] a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means;
[1526] an emotion recognition engine that recognizes the user's emotional state based on their voice and video;
[1527] means for adjusting the tone and content of cooking advice based on the evaluation of the emotion recognition engine;
[1528] A system including:
[1529] (Claim 2)
[1530] 2. The system according to claim 1, wherein ingredient information and recipe information are obtained from a user, and the analysis means determines a cooking procedure based on this information.
[1531] (Claim 3)
[1532] 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice.
[1533] "Application example 2 when combining emotion engines"
[1534] (Claim 1)
[1535] A temperature sensor that detects the condition of the ingredients,
[1536] a camera for capturing images of ingredients;
[1537] an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of ingredients;
[1538] emotion analysis means for recognizing the user's emotions and adjusting cooking method and seasoning advice based on the emotions;
[1539] a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means and the emotion analysis means;
[1540] …
[1541] A system including:
[1542] (Claim 2)
[1543] 2. The system according to claim 1, wherein ingredient information and recipe information are obtained from a user, and the analysis means determines a cooking procedure based on this information.
[1544] (Claim 3)
[1545] 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice. [Explanation of symbols]
[1546] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A temperature sensor that detects the condition of the ingredients, a camera for capturing images of ingredients; an analysis means for analyzing data from the temperature sensor and the camera to evaluate the doneness and moisture content of ingredients; a notification means for providing advice on cooking methods and seasonings in natural language based on the analysis means; A system including:
2. 2. The system according to claim 1, wherein ingredient information and recipe information are acquired from a user, and the analysis means determines a cooking procedure based on this information.
3. 10. The system of claim 1, further comprising: a learning means for receiving feedback from a user and for said analyzing means for improving the accuracy of future cooking advice.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A