System
The smart AI assistant system simplifies barbecuing by using voice recognition and image analysis to provide real-time advice and recipe suggestions, addressing the challenges faced by beginners and busy individuals.
Patent Information
- Application Number
- JP2024120618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Barbecuing is challenging for beginners and busy individuals due to difficulties in adjusting heat and positioning ingredients, lacking real-time monitoring, and requiring significant effort and time.
A smart AI assistant system with voice recognition, image analysis, cooking management, communication, and user interface capabilities that provides real-time advice and recipe suggestions based on user input and grill conditions.
Enables easy and high-quality barbecuing by reducing hassle and effort, allowing users to enjoy the experience without extensive knowledge or experience.
Smart Images

Figure 2026019209000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The purpose of this invention is to reduce the hassle that beginners and busy people face when preparing and cooking a barbecue, so that anyone can easily enjoy a barbecue. In particular, it aims to eliminate the difficulties that are unique to barbecues, such as adjusting the heat and changing the position of the ingredients. [Means for solving the problem]
[0005] The present invention proposes a system that includes a voice recognition means, an image analysis means, a cooking management means, a communication means, and a user interface means. The voice recognition means receives voice commands from the user and converts them into text data, and the image analysis means captures and analyzes images of the barbecue grill in real time. The cooking management means generates advice on adjusting the heat and moving ingredients based on the results of the image analysis, and the communication means sends and receives data bidirectionally with the server. The user interface means provides the user with appropriate advice via voice or on-screen display, significantly reducing the hassle of barbecuing. Furthermore, the cooking management means suggests recipes based on the user's past cooking history and current cooking status, allowing the user to easily try new dishes.
[0006] A "voice recognition means" is a device or system that receives voice commands from a user and converts them into text data.
[0007] "Image analysis means" refers to a device or system that analyzes image data acquired from a photographic device such as a camera and evaluates the state and characteristics of the target.
[0008] A "cooking management means" is a device or system that generates and presents cooking instructions and advice based on acquired data (such as gold analysis and voice recognition results).
[0009] A "communication means" is a device or system that has a network connection function for sending and receiving data.
[0010] "User interface means" refers to an interface for exchanging information between the system and the user, such as an audio output device or a display.
[0011] A "barbecue grill" is a cooking device equipped with a grill and a grill rack used for outdoor cooking.
[0012] "Text data" refers to character information converted by a voice recognition means.
[0013] A "recipe" refers to information about the steps and ingredients for making a particular dish.
[0014] "Real-time" refers to immediate processing or response in real time.
[0015] "Advice" refers to instructions or advice regarding cooking or operation. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a smart AI assistant system that supports the barbecue experience, and includes voice recognition, image analysis, cooking management, communication, and user interface. The program and processing of this system are explained below in natural language with concrete examples.
[0038] System Overview
[0039] When a user starts a barbecue, this system uses a voice recognition means to receive the user's voice commands and convert them into text data. Next, an image analysis means captures and analyzes real-time images of the barbecue grill. Based on the analysis results, a cooking management means generates advice such as adjusting the heat or changing the position of ingredients. The generated advice is sent to the user terminal via a communication means and presented to the user via a user interface means.
[0040] Program processing
[0041] 1. The user starts a barbecue
[0042] The user issues the voice command "Start the barbecue."
[0043] The terminal receives the voice, and the voice recognition means converts it into text data, which is then sent to the server.
[0044] 2. Analysis of audio data
[0045] The server analyzes the text data and determines that the user is about to start a barbecue, and then performs the initial setup.
[0046] 3. Preparatory Instructions
[0047] The server generates preparation instructions for the user (e.g., "Prepare the grill") based on the initial settings and sends them to the terminal.
[0048] The device will provide instructions to the user via voice or visual display.
[0049] 4. Real-time image acquisition
[0050] The user points the device's camera at the barbecue grill.
[0051] The device uses a camera to capture an image of the grill and transmits the image data to a server.
[0052] 5. Analysis of Image Data
[0053] The server analyzes the image data and evaluates the placement of ingredients, the state of the fire, and the degree of doneness.
[0054] The server determines the necessary adjustments based on the analysis results.
[0055] 6. Generating and Providing Advice
[0056] The server generates advice on adjusting the heat and changing the position of ingredients and sends it to the terminal.
[0057] The device presents the generated advice to the user by voice or on-screen display.
[0058] The user makes adjustments according to the advice.
[0059] Specific examples
[0060] The process of starting a barbecue
[0061] 1. The user issues a voice command
[0062] User: "Start a barbecue."
[0063] The terminal converts the voice command into text data using a voice recognition means and transmits it to the server.
[0064] 2. The server analyzes the request and sets up the initial settings
[0065] Server: Parses the user's request, performs initialization, and generates preparation instructions.
[0066] Terminal: Prompts the user with the instruction "Prepare the grill."
[0067] Real-time image analysis and advice
[0068] 1. User takes a picture of the grill
[0069] User: Point your device camera at the grill.
[0070] The device sends image data to a server, which analyzes the image.
[0071] 2. The server analyzes the image and generates advice
[0072] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[0073] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[0074] The system of the present invention allows users to easily enjoy delicious barbecues without detailed knowledge or experience. It also significantly reduces the effort required for cooking, allowing users to spend more time concentrating on conversations and meals with family and friends.
[0075] The processing flow will be explained below.
[0076] Step 1:
[0077] The user issues the voice command "Start the barbecue."
[0078] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[0079] Step 2:
[0080] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[0081] The server performs the initial setup and creates the necessary setup instructions.
[0082] Step 3:
[0083] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[0084] The terminal presents instructions to the user by voice or on-screen display.
[0085] The user follows the instructions to prepare the grill.
[0086] Step 4:
[0087] The user points the device's camera at the grill and places the ingredients on it.
[0088] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[0089] Step 5:
[0090] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[0091] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0092] Step 6:
[0093] The server generates advice (e.g., "Turn up the heat a little," "Move the chicken to the right") and sends it to the device.
[0094] The advice received by the terminal is presented to the user by voice or screen display.
[0095] The user adjusts the heat and position of the ingredients according to the advice.
[0096] Step 7:
[0097] The user asks the device, "What should I make next?"
[0098] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0099] Step 8:
[0100] The server selects the optimal recipe based on the user's past cooking history and current ingredient information.
[0101] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[0102] The terminal presents recipe suggestions to the user by voice or on-screen display.
[0103] Step 9:
[0104] The server periodically analyzes the images to determine whether the food is cooked.
[0105] The server generates a notification saying "The chicken is done" and sends it to the device.
[0106] The terminal notifies the user by voice when the food is done.
[0107] The user removes the cooked chicken and enjoys the meal.
[0108] Example 1
[0109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0110] In traditional barbecue cooking, beginners and those with little experience often have difficulty controlling the heat and arranging ingredients properly. Additionally, the lack of real-time monitoring and proper advice can lead to inconsistent cooking quality. This can result in users spending a lot of time and effort on cooking, preventing them from fully enjoying their barbecue experience.
[0111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0112] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a real-time data collection unit, a data analysis unit, and an advice generation unit. This allows users to start a barbecue using voice commands and manage the cooking appropriately through real-time image data analysis. Furthermore, specific advice based on the analysis results can be received by voice or displayed on the screen, allowing even beginners to easily enjoy a high-quality barbecue experience.
[0113] "Voice recognition means" refers to technology that receives voice commands from a user and converts them into text data.
[0114] "Image analysis means" refers to technology that acquires real-time images of the barbecue grill and analyzes the placement of ingredients, the state of the fire, and the degree of doneness.
[0115] "Cooking management means" refers to technology that manages the cooking process based on image analysis results and audio data.
[0116] "Communication means" refers to the technology for sending and receiving data between a terminal and a server.
[0117] "User interface means" refers to means for providing visual or audio information to a user and accepting input from the user.
[0118] "Real-time data collection means" refers to technology that acquires images and audio data of barbecue grills in real time.
[0119] "Data analysis means" refers to technology that analyzes collected voice and image data to evaluate the user's intentions and the condition of the grill.
[0120] "Advice generation means" refers to technology that generates specific advice necessary for cooking based on the results of data analysis.
[0121] This invention is a smart AI assistant system that supports barbecue experiences, and includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, a real-time data collection means, a data analysis means, and an advice generation means.
[0122] System Overview
[0123] This system helps users start a barbecue. The user starts the system by issuing the voice command "Start a barbecue." Each step and the hardware and software used are explained below.
[0124] Hardware and software used
[0125] The system uses the following major hardware and software:
[0126] Speech recognition method: Google Cloud Speech-to-Text
[0127] Image analysis methods: OpenCV, TensorFlow
[0128] Cooking control means: Control logic within the system
[0129] Communication method: Wi-Fi or Bluetooth
[0130] User interface means: smartphone application
[0131] Real-time data collection method: built-in camera and microphone of smartphone
[0132] Data analysis method: Google Cloud Natural Language API
[0133] Advice generation means: Advice generation engine within the system
[0134] Program processing
[0135] The program begins when the user issues the voice command "Start the barbecue." The device receives this voice, converts it into text data using Google Cloud Speech-to-Text, and sends the data to the server. The server then analyzes the text data using the Google Cloud Natural Language API, understands the user's intent, and sets up the initial barbecue setup.
[0136] Once the initial setup is complete, the server generates preparation instructions for the user, such as "Prepare the grill," and sends them to the device, which then presents the instructions to the user as a voice message or a screen display.
[0137] Next, the user points their smartphone camera at the barbecue grill. The device uses the camera to capture real-time images and sends the image data to a server. The server then uses image analysis technology (OpenCV, TensorFlow) to analyze the images and evaluate the placement of ingredients, the state of the fire, and the degree of doneness. Based on the evaluation, the server determines the necessary adjustments and generates specific advice (e.g., "Turn up the heat a little" or "Move the meat to the right").
[0138] The generated advice is sent to the device, which then presents it to the user via voice or screen display, allowing the user to make adjustments according to the advice.
[0139] Specific examples
[0140] The process of starting a barbecue
[0141] 1. The user issues a voice command
[0142] User: "Start a barbecue."
[0143] The device uses a voice recognition tool (Google Cloud Speech-to-Text) to convert the voice command into text data and send it to the server.
[0144] 2. The server analyzes the request and sets up the initial settings
[0145] Server: Parses the user request (Google Cloud Natural Language API), performs initialization, and generates preparation instructions.
[0146] Terminal: Prompt the user (audio or screen prompt) to "Prepare the grill."
[0147] Real-time image analysis and advice
[0148] 1. User takes a picture of the grill
[0149] User: Point your device camera at the grill.
[0150] The device sends image data to the server (captured by the smartphone's built-in camera), and the server analyzes the image (OpenCV, TensorFlow).
[0151] 2. The server analyzes the image and generates advice
[0152] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[0153] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[0154] Example of input prompt for generative AI model
[0155] 1. Prompts for user command to generate preparation instructions when starting a barbecue:
[0156] After receiving a voice command from the user to "start barbecue," perform the initial setup and generate and present preparation instructions to the user.
[0157] Example: If the user says "I want to start the barbecue," the device might display or speak instructions such as "Prepare the grill."
[0158] 2. Real-time image analysis prompts advice generation:
[0159] Based on the analysis of real-time images of the barbecue grill, generate and present specific advice to the user on adjusting the heat or repositioning the ingredients.
[0160] For example, if the grill is running low, you might offer advice such as, "Please turn up the heat a little."
[0161] This system allows users to easily enjoy delicious food even if they have no knowledge or experience of barbecue. The system provides users with easy-to-understand instructions and helps them cook.
[0162] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0163] Step 1:
[0164] The user issues a voice command such as "Start the barbecue."
[0165] Input: Voice commands from the user.
[0166] Action: The device receives audio through the microphone.
[0167] Output: The captured audio data is generated.
[0168] Step 2:
[0169] The terminal converts the received voice data into text data using a voice recognition means.
[0170] Input: Audio data.
[0171] How it works: Converts audio data into text using Google Cloud Speech-to-Text.
[0172] Output: Text data is generated and sent to the server.
[0173] Step 3:
[0174] The server analyzes the received text data and understands the user's intent.
[0175] Input: Text data.
[0176] How it works: It uses the Google Cloud Natural Language API to analyze text data and recognize the user's intention to start a barbecue.
[0177] Output: Configuration data reflecting the initial settings is generated.
[0178] Step 4:
[0179] The server generates and sends preparation instructions to the user based on the initial settings.
[0180] Input: Configuration data.
[0181] Operation: Using the cooking management means, a preparation instruction such as "Prepare the grill" is generated and sent to the terminal.
[0182] Output: Preparation instruction data is sent to the terminal.
[0183] Step 5:
[0184] The terminal presents the received preparation instructions to the user.
[0185] Input: Preparation instruction data.
[0186] What it does: Uses user interface means to present instructions to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[0187] Output: The user receives the instructions.
[0188] Step 6:
[0189] A user points their smartphone camera at a barbecue grill.
[0190] Input: Instructions from the server.
[0191] How it works: Point your smartphone camera at the barbecue grill to capture real-time images.
[0192] Output: Real-time image data is generated.
[0193] Step 7:
[0194] The real-time images acquired by the terminal are sent to the server.
[0195] Input: Real-time image data.
[0196] Operation: Image data captured using the camera is sent to the server.
[0197] Output: The image data is transferred to the server.
[0198] Step 8:
[0199] The server analyzes the received image data.
[0200] Input: Image data.
[0201] How it works: Image analysis technology (OpenCV, TensorFlow) is used to analyze the placement of ingredients, the state of the fire, and the degree of doneness.
[0202] Output: Analysis result data is generated.
[0203] Step 9:
[0204] The server generates and sends advice based on the analysis results.
[0205] Input: Analysis result data.
[0206] Operation: Using the advice generation means, specific advice regarding adjusting the heat or changing the position of ingredients is generated and sent to the terminal.
[0207] Output: Advice data is sent to the terminal.
[0208] Step 10:
[0209] The device presents the received advice to the user.
[0210] Input: Advice data.
[0211] What it does: Uses user interface means to present advice to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[0212] Output: User receives advice.
[0213] Step 11:
[0214] The user adjusts the barbecue grill according to the provided advice.
[0215] Enter: Advice.
[0216] Action: Making specific adjustments such as adjusting the heat or changing the position of ingredients.
[0217] Output: Adjusted BBQ grill.
[0218] (Application example 1)
[0219] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0220] Traditional barbecue and cooking relied on the knowledge and experience of the chef, making it difficult to ensure consistent quality. It was also difficult to monitor the status of the food in real time and make appropriate adjustments, which required a lot of time and effort. This made it difficult for inexperienced chefs to provide satisfying food.
[0221] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0222] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a unit for evaluating and adjusting cooking quality in real time, and a unit for providing advice using a generative AI model. This makes it possible to provide high-quality, consistent cooking without relying on the chef's experience. Furthermore, by being able to grasp the cooking status in real time and make appropriate adjustments, it is possible to significantly reduce the effort required and provide a satisfying cooking experience.
[0223] "Voice recognition means" refers to any device or technology that receives voice commands from a user and converts them into text data.
[0224] "Image analysis means" refers to devices and technologies in general that analyze image data captured in real time and recognize specific objects or conditions.
[0225] "Cooking management means" refers to all devices and technologies that manage cooking based on the analysis results and generate instructions such as adjusting the heat or changing the position of ingredients.
[0226] "Communication means" refers to all devices and technologies for transmitting and receiving data.
[0227] "User interface means" refers generally to devices and techniques that allow a user to interact with a system and obtain information.
[0228] "Means for evaluating and adjusting food quality in real time" refers to all devices and technologies that analyze data acquired in real time, evaluate food quality, and make appropriate adjustments.
[0229] "Means for providing advice using a generative AI model" refers to all devices and technologies that use a generative AI model to provide appropriate advice to users based on the analysis results.
[0230] This invention is a smart AI cooking assistant system for food delivery, which includes voice recognition means, image analysis means, cooking management means, communication means, user interface means, means for evaluating and adjusting cooking quality in real time, and means for providing advice using a generative AI model, enabling cooks to deliver high-quality, consistent food.
[0231] Specifically, users operate the system using a smartphone or head-mounted display. For example, when a user issues a voice command such as "Start a barbecue," the smartphone's microphone captures the voice and converts it into text data using a speech recognition library. The converted text data is then sent to a server for initial setup.
[0232] Next, the user uses their smartphone camera to capture images of the grill in real time. The image data is captured using the OpenCV library and sent to the server. The server analyzes the images and evaluates the placement of ingredients and the state of the flame. Based on the evaluation, the cooking management tool generates advice on adjusting the heat or repositioning ingredients.
[0233] The generated advice is sent as text data to the smartphone via the communication means. The advice is presented to the user via the user interface means as voice or on-screen display. For example, if the advice is "Turn up the heat a little," the user adjusts the heat according to the instruction.
[0234] The system can also use generative AI models to provide more advanced advice, such as providing tailored advice in response to prompts like "What's the best way to adjust the heat on my barbecue?" or "How do I check if my chicken is done?"
[0235] In this way, the system enables consistent, high-quality food to be served in real time, while significantly reducing the chef's workload and providing a satisfying cooking experience.
[0236] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0237] Step 1:
[0238] The user issues the voice command "Start the barbecue." The device captures the voice command using the built-in microphone and converts the voice into text data using a voice recognition library (e.g., SpeechRecognition library). This text data is sent to the server. The input is the voice command, and the output is text data. Data processing involves analyzing the voice waveform data and converting it into text data.
[0239] Step 2:
[0240] The server receives the text data and recognizes that the user is about to start a barbecue. The server generates a preparation instruction, "Prepare the grill," as the initial setting and sends this instruction to the terminal. The input is the text data, and the output is the initial setting and the preparation instruction. As a data operation, it performs text analysis and sets the initial setting flag.
[0241] Step 3:
[0242] The device receives the preparation instruction from the server and presents it to the user by voice or on-screen display. The input is the preparation instruction from the server, and the output is the instruction presented to the user. Specifically, the instruction is played back by voice using the device's speaker.
[0243] Step 4:
[0244] A user points their smartphone camera at a barbecue grill. The device uses the camera to capture real-time images of the grill and sends the image data to a server. The input is the camera image and the output is image data. Image capture and compression are performed as data processing.
[0245] Step 5:
[0246] The server analyzes the image data it receives and evaluates the placement of ingredients and the state of the fire. This is done using an image analysis library (e.g., OpenCV). The input is the image data, and the output is the analysis results. Data calculations involve extracting basic image features and obtaining information on the heat level and the position of ingredients.
[0247] Step 6:
[0248] The server uses the cooking management means based on the analysis results to generate advice on adjusting the heat or changing the position of ingredients. This advice is sent to the terminal as text data via communication means. The input is the analysis results, and the output is advice. As a data calculation, appropriate cooking instructions are automatically generated based on the analysis results.
[0249] Step 7:
[0250] The device receives advice from the server and presents it to the user by voice or on-screen display. The input is advice from the server, and the output is instructions presented to the user. Specifically, the advice is played back by voice using the device's speaker, or displayed as text on the screen.
[0251] Step 8:
[0252] The user makes adjustments based on the advice from the server. This step acts as a feedback loop, returning to step 4 and beyond to make the next adjustment if necessary. The input is the advice, and the output is the adjustment action taken by the user. Specific actions taken by the user include adjusting the heat or changing the position of ingredients.
[0253] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0254] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. Below, we will explain the program and processing of this system with concrete examples.
[0255] System Overview
[0256] The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, and an emotion engine. The voice recognition unit receives a user's voice command and converts it into text data. The image analysis unit acquires and analyzes real-time images of the barbecue grill. The cooking management unit generates advice based on the results of the image analysis, which is sent to the user terminal via the communication unit and presented to the user via the user interface unit. The emotion engine also analyzes the user's emotions and adjusts the cooking advice based on the analysis.
[0257] Program processing
[0258] 1. The user starts a barbecue
[0259] The user issues the voice command "Start the barbecue."
[0260] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[0261] 2. Analysis of audio data
[0262] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[0263] The server performs the initial setup and creates the necessary setup instructions.
[0264] 3. Preparatory Instructions
[0265] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[0266] The device will provide instructions to the user via voice or on-screen display.
[0267] The user follows the instructions to prepare the grill.
[0268] 4. Real-time image acquisition
[0269] The user points the device's camera at the grill and places the ingredients on it.
[0270] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[0271] 5. Analysis of Image Data
[0272] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[0273] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0274] 6. Emotion Analysis Using an Emotion Engine
[0275] The device captures the user's facial expressions and tone of voice and sends them to the emotion engine.
[0276] The server uses an emotion engine to analyze the user's emotions and adjusts cooking advice based on this.
[0277] 7. Generating and Providing Advice
[0278] The server adjusts the advice based on the emotion data and sends the generated advice (e.g., "Turn up the heat a little," "Move the chicken to the right") to the terminal.
[0279] The advice received by the device is presented to the user by voice or on-screen display.
[0280] The user adjusts the heat and position of the ingredients according to the advice.
[0281] 8. Recipe suggestions
[0282] The user asks the device, "What should I make next?"
[0283] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0284] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[0285] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[0286] The device will present recipe suggestions to the user via voice or on-screen display.
[0287] 9. Emotional Cooking Guidance
[0288] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[0289] The server generates cooking guidance based on the user's emotions (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device.
[0290] The device will notify the user of emotionally sensitive notifications via voice.
[0291] The user removes the cooked chicken and enjoys the meal.
[0292] The system of the present invention allows users to easily enjoy delicious barbecues, even without detailed knowledge or experience. Furthermore, by combining it with an emotion engine, a more personalized and comfortable barbecue experience can be provided. By significantly reducing the effort required for cooking and providing emotionally appropriate encouragement and advice, users can spend more time concentrating on conversations with family and friends and on their meals.
[0293] The processing flow will be explained below.
[0294] Step 1:
[0295] The user issues the voice command "Start the barbecue."
[0296] The device uses a microphone to capture the user's voice.
[0297] Step 2:
[0298] The terminal converts the voice into text data using a voice recognition means.
[0299] The terminal transmits the converted text data to the server.
[0300] Step 3:
[0301] The server receives the text data, analyzes it, and recognizes it as a request to "start barbecue."
[0302] The server performs the initialization and generates the necessary provisioning instructions.
[0303] Step 4:
[0304] The server sends the generated preparation instruction to the terminal.
[0305] The terminal will give the user instructions to prepare the grill by voice or on the screen, such as "Prepare the grill."
[0306] Step 5:
[0307] The user follows the instructions to prepare the grill.
[0308] The user points the device's camera at the grill and captures an image.
[0309] Step 6:
[0310] The device captures an image of the grill and sends the image data to a server.
[0311] The server receives the image data and uses image analysis means to evaluate the placement of ingredients, the state of the fire, and the degree of doneness.
[0312] Step 7:
[0313] Based on the results of image analysis, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0314] The server transmits the generated advice to the terminal.
[0315] Step 8:
[0316] The device displays the received advice to the user (e.g., "Turn up the heat a little").
[0317] The user adjusts the heat and position of the ingredients according to the advice.
[0318] Step 9:
[0319] The device captures the user's facial expressions and voice tone and transmits the emotional data to the server.
[0320] The server analyzes the user's emotions using an emotion engine.
[0321] Step 10:
[0322] The server adjusts cooking advice based on the results of emotion analysis (e.g., if the user seems anxious, give advice in a gentler tone).
[0323] The server sends the adjusted advice to the terminal.
[0324] Step 11:
[0325] The terminal presents the adjusted advice to the user.
[0326] The user proceeds with cooking according to the advice.
[0327] Step 12:
[0328] The user asks the device, "What should I make next?"
[0329] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0330] Step 13:
[0331] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[0332] The server generates recipe suggestions and sends them to the terminal.
[0333] Step 14:
[0334] The device will present recipe suggestions (e.g., "We recommend this chicken marinade") to the user via voice or on-screen display.
[0335] Users can try new dishes by following suggested recipes.
[0336] Step 15:
[0337] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[0338] The server generates a notification saying "The chicken is done" and sends it to the device.
[0339] Step 16:
[0340] The terminal notifies the user by voice when the food is done.
[0341] The user removes the cooked chicken and enjoys the meal.
[0342] Example 2
[0343] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0344] Enjoying a barbecue often requires skilled techniques and a lot of prior knowledge. However, for users who are trying barbecue for the first time or who are unfamiliar with cooking, these requirements are a high hurdle and make it difficult to fully enjoy a barbecue. In addition, there is a problem that the user experience is inconsistent because the cooking status is not grasped in real time or appropriate advice is not provided according to the user's emotions.
[0345] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, an emotion analysis means, and a user interface means. This allows even beginners to easily receive detailed cooking advice on barbecues, and enables personalized support tailored to the user's emotions. Specifically, the voice recognition means recognizes the user's voice commands and converts them into text data. The image analysis means captures real-time images of the grill and analyzes the state of the ingredients and the fire. Furthermore, the emotion analysis means analyzes the user's emotions from their facial expressions and voice tone. The cooking management means integrates these analysis results to generate appropriate cooking advice, which is then provided to the user through the user interface means.
[0346] "Speech recognition means" refers to means having the function of receiving voice commands from a user and converting them into text data.
[0347] The "image analysis means" is a means having the function of analyzing images captured in real time and evaluating the state of the ingredients and the fire.
[0348] The "cooking management means" is a means having a function of generating cooking advice based on the results of the image analysis means and the emotion analysis means.
[0349] "Communication means" refers to a means that utilizes protocols and technologies for sending and receiving data.
[0350] A "user interface means" is a means for presenting information to a user and receiving input from a user.
[0351] The "emotion analysis means" is a means having the function of capturing the user's facial expressions and tone of voice and analyzing their emotions.
[0352] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. This system includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, and an emotion analysis means.
[0353] Hardware and software used
[0354] The system uses the following major hardware and software:
[0355] 1. Speech recognition means: Typically using a cloud-based speech recognition service, e.g., a speech recognition cloud API.
[0356] 2. Image analysis method: Uses the OpenCV library and deep learning models (e.g., TensorFlow).
[0357] 3. Cooking control method: A custom algorithm implemented in Python is used.
[0358] 4. Communication method: Network communication is performed using the HTTP / HTTPS protocol.
[0359] 5. User interface: Implemented through a smartphone app (iOS / Android) or web browser.
[0360] 6. Emotion analysis means: Use an API that performs facial expression analysis and voice tone analysis (e.g., emotion analysis cloud API).
[0361] Processing of this system
[0362] 1. When a user wants to start a barbecue, they say the voice command "Start the barbecue." This voice command is converted into text data by the device using a voice recognition cloud API and sent to the server.
[0363] 2. The server analyzes the received text data and recognizes that the user's request is to "start a barbecue." It uses a natural language processing (NLP) library to process this. The server then performs its initial setup and creates the necessary preparation instructions.
[0364] 3. The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the device. The device then presents the instructions to the user via voice or visual display, and the user prepares the grill accordingly.
[0365] 4. The user points the device's camera at the grill and places the food on it. The device uses the camera to capture real-time images of the grill and sends the image data to the server.
[0366] 5. The server analyzes the received image data using OpenCV and deep learning models. The analysis results include the placement of the ingredients, the state of the fire, and the degree of doneness. Based on this information, the server generates specific advice on adjusting the heat or repositioning the ingredients.
[0367] 6. The device captures the user's facial expressions and voice tone and sends them to the emotion analysis cloud API. The server uses an emotion engine to analyze the user's emotions and adjust the cooking advice accordingly.
[0368] 7. The server generates advice based on the emotion data and sends it to the device. The device then presents the advice to the user via voice or on-screen display. The user adjusts the heat and position of the ingredients according to the advice.
[0369] 8. The user asks the device, "What should I make next?" The device uses a voice recognition cloud API to convert the question into text data and sends it to the server. The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The server generates a recipe suggestion (e.g., "I recommend marinated chicken") and sends it to the device. The device then presents the recipe suggestion to the user by voice or on a screen.
[0370] 9. The server periodically performs image and emotion analysis to determine whether the food is cooked. The server generates cooking guidance based on the user's emotion (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device. The device then notifies the user with voice notifications that take emotion into consideration. The user follows the instructions to remove the cooked chicken and enjoy their meal.
[0371] Specific examples
[0372] Example 1: Your first barbecue
[0373] 1. The user says the voice command "Start the barbecue."
[0374] 2. The device converts the voice into text and sends it to the server.
[0375] 3. The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal.
[0376] 4. The user prepares the grill and points the device camera at the grill.
[0377] 5. The device sends the captured image to the server, which then analyzes the image.
[0378] 6. The server generates advice on adjusting the heat and sends it to the device.
[0379] 7. The device provides advice to the user, who then adjusts the heat.
[0380] Prompt Sentence Examples
[0381] Prompt 1: "Start the BBQ Assistant to provide a set of guidance for first-time barbecuers."
[0382] Prompt 2: "When a user requests a new recipe, recommend a recipe that fits their mood that day."
[0383] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0384] Step 1:
[0385] The user issues the voice command "Start the barbecue."
[0386] The device converts the speech into text data using a speech recognition cloud API. The input is the user's speech, and the output is the converted text data. The device then sends this text data to the server.
[0387] Step 2:
[0388] The server analyzes the received text data. The input is text data sent from the device, and the server uses a natural language processing library to recognize that it is a request to "start a barbecue." The output is the result of recognizing the request as "start a barbecue." The server performs initial configuration and creates the necessary preparation instructions.
[0389] Step 3:
[0390] The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal. The input is the initial setting result, and the output is a text message of the preparation instruction. The terminal receives this instruction and presents it to the user by voice or on-screen display. The user follows the instruction and performs the action of preparing the grill.
[0391] Step 4:
[0392] The user points the device's camera at the grill and places ingredients on it. The input is the camera image. The device uses the camera to capture real-time images of the grill and sends the image data to the server. The output is the captured real-time image data.
[0393] Step 5:
[0394] The server analyzes the received image data. The input is image data sent from the device. The server uses OpenCV and deep learning models (such as TensorFlow) to evaluate the placement of ingredients, the state of the fire, and the degree of doneness. The output is the analysis results, which are summarized as specific advice on adjusting the heat or changing the position of ingredients.
[0395] Step 6:
[0396] The device captures the user's facial expressions and voice tone. The input is the user's facial expression and voice data. The device sends this to the emotion analysis cloud API. The output is the emotion analysis result, which is sent to the server.
[0397] Step 7:
[0398] The server receives the emotion data analysis results and generates adjustment advice. The input is emotion data and image analysis results. The server adjusts the cooking advice based on this data and generates specific advice (e.g., "Turn up the heat a little," "Move the chicken to the right"). The output is the adjusted advice. The server sends this advice to the device.
[0399] Step 8:
[0400] The device provides advice to the user by voice or on-screen display. The input is advice data from the server, and the output is the information presented to the user. The user follows the advice and adjusts the heat and position of the ingredients.
[0401] Step 9:
[0402] The user asks the device, "What should I make next?" The input is the user's voice command. The device again uses the speech recognition cloud API to convert the question into text data and sends that data to the server. The output is the converted text data.
[0403] Step 10:
[0404] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The input is the user's past cooking history, current ingredient information, and emotional data, and the output is the optimal recipe suggestion. The server generates a recipe suggestion (e.g., "We recommend marinated chicken") and sends it to the device.
[0405] Step 11:
[0406] The terminal presents recipe suggestions to the user by voice or on-screen display. The input is the recipe suggestion data from the server, and the output is the suggestion to the user.
[0407] Step 12:
[0408] The server periodically performs image analysis and emotion analysis. The input is real-time images and user emotion data. The server determines whether the ingredients are cooked and generates cooking guidance. The output is cooking guidance (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?"). The server then sends the generated cooking guidance to the device.
[0409] Step 13:
[0410] The device provides emotionally sensitive notifications to the user. The input is cooking guidance data from the server, and the output is a notification to the user. The user follows the instructions, removes the cooked chicken, and enjoys the meal.
[0411] Through the above processing steps, the smart AI assistant system of the present invention can help users enjoy a more enjoyable and easier barbecue. By providing personalized advice based on the cooking progress and the user's emotions, even beginners can enjoy a delicious barbecue.
[0412] (Application example 2)
[0413] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0414] Conventional food delivery systems lack personalized menu suggestions and cooking advice based on the user's emotional state, which hinders the improvement of user experience. Furthermore, when users are cooking, they often do not receive appropriate advice in real time, which leads to a decrease in satisfaction.
[0415] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, an emotion analysis means, and a database means. This makes it possible to propose optimal menus and provide cooking advice according to the user's emotional state, which is expected to improve the user experience.
[0416] "Voice recognition" is a technology that receives a user's voice commands and converts them into text data.
[0417] "Image analysis" is a technology that analyzes the state of ingredients and fire based on image data captured in real time.
[0418] "Cooking management" is a technology that generates cooking advice based on the results of image analysis and presents it to the user.
[0419] "Communication means" refers to the technology used to send and receive voice data, image data, and analysis results between the server and the user terminal.
[0420] A "user interface" is a means for receiving voice commands and image data input from a user, and for providing generated advice and menu suggestions to the user.
[0421] "Emotion analysis" is a technology that analyzes a user's facial expressions and tone of voice to understand their emotional state.
[0422] The "database" is a technology that stores a user's past order history and emotional data, and generates suggestions based on the analysis results as needed.
[0423] The present invention relates to an emotion-responsive food delivery support application, specifically a system that provides optimal menu suggestions and cooking advice based on a user's voice commands and emotional state. The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, an emotion analysis unit, and a database unit.
[0424] System Overview
[0425] The solution begins when a user launches a food delivery application on their smartphone. First, the user issues a voice command such as "Tell me what menu items you want." The smartphone uses a microphone to capture the voice and converts it into text using a speech recognition library called Google Cloud Speech-to-Text. This text is then sent to a server, which analyzes it and understands the user's request.
[0426] Next, an emotion analysis tool analyzes the user's facial expressions and tone of voice to generate emotion data. This uses the Microsoft Azure Emotion API. The emotion data is used to understand the user's emotional state, and appropriate menu suggestions and cooking advice are provided based on that emotional state. Past order history and emotion data are stored in a database (e.g., MySQL), and optimal suggestions are generated based on this data.
[0427] For example, if the user is in a state of mind where they want to relax, the system will suggest "relaxing herbal tea." If the user is cooking at home, the system will perform real-time image analysis and provide appropriate cooking advice. Image analysis uses a library called OpenCV.
[0428] The specific hardware and software used
[0429] Hardware: Smartphone (microphone, camera, speaker)
[0430] software:
[0431] Speech recognition library: Google Cloud Speech-to-Text
[0432] Sentiment analysis library: Microsoft Azure Emotion API
[0433] Image analysis library: OpenCV
[0434] Database: MySQL
[0435] Prompt Sentence Examples
[0436] Specific examples of prompts are as follows:
[0437] User's order history: pizza, pasta, salad
[0438] Emotional state: Relaxed
[0439] User Question: What's your recommended menu item?
[0440] Response: I recommend herbal tea. It's perfect for relaxing.
[0441] In this way, the system of the present invention allows users to enjoy a personalized food delivery experience based on their voice commands and emotional state.
[0442] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0443] Step 1:
[0444] A user launches a food delivery application on their smartphone and says, "Tell me what menu items you recommend." Here, the input is the user's voice command, and the smartphone's microphone is used to capture the voice data. This voice data is converted into text data using the Google Cloud Speech-to-Text library. The converted text data is the output.
[0445] Step 2:
[0446] The speech-recognized text data is sent to the server. Here, the input is the converted text data, and the server receives the data. The server performs text analysis to understand the user's request. As a result of the analysis, data called "recommended menu suggestions" is output.
[0447] Step 3:
[0448] The emotion analyzer captures the user's facial expressions and voice tone and generates emotion data using the Microsoft Azure Emotion API, where the input is the user's real-time facial expression data and voice tone, and the API analyzes the user's emotional state (e.g., relaxed, happy, etc.), and the analyzed emotion data is the output.
[0449] Step 4:
[0450] The server retrieves the generated emotion data and past order history from a database. Here, the input is emotion data and order history data. This data is retrieved through a database query, and the server selects the optimal menu based on this data. The selected menu suggestion is the output.
[0451] Step 5:
[0452] The server generates a selected menu suggestion (e.g., "I recommend herbal tea. It's perfect for relaxing.") and sends it to the smartphone via communication means. Here, the input is the data of the selected menu suggestion, which is sent to the smartphone. The output is the smartphone presenting the suggested menu to the user by voice or on-screen display.
[0453] Step 6:
[0454] When a user cooks at home, they use their smartphone camera to capture ingredients and the cooking process. The input is real-time image data, which is analyzed using an image analysis library (OpenCV). The analysis results are then output as cooking advice (e.g., "Adjust the heat.").
[0455] Step 7:
[0456] The server sends the generated cooking advice to the smartphone and presents it to the user. Here, the input is the analyzed cooking advice data, and the data is sent to the smartphone. The output of the smartphone is to provide advice to the user by voice or on-screen display.
[0457] Step 8:
[0458] The user then asks a follow-up question (e.g., "What's the next recommended menu item?") by voice. The smartphone's microphone is again used to capture the voice data, which is then converted into text data using the Google Cloud Speech-to-Text library. This voice data is the input, and the converted text data is the output.
[0459] Step 9:
[0460] The server analyzes the converted text data and generates the next menu suggestion based on the information from the database and the sentiment data. Here, the input is the text data based on the follow-up question and related data, and the output is the next menu suggestion as the analysis result.
[0461] Step 10:
[0462] The server sends the next menu proposal to the smartphone and presents it to the user. Here, the input is the generated next menu proposal data, which is sent to the smartphone. The output of the smartphone is to present it to the user by voice or on-screen display.
[0463] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0464] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0465] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0466] [Second embodiment]
[0467] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0468] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0469] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0470] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0471] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0472] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0473] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0474] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0475] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0476] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0477] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0478] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0479] This invention is a smart AI assistant system that supports the barbecue experience, and includes voice recognition, image analysis, cooking management, communication, and user interface. The program and processing of this system are explained below in natural language with concrete examples.
[0480] System Overview
[0481] When a user starts a barbecue, this system uses a voice recognition means to receive the user's voice commands and convert them into text data. Next, an image analysis means captures and analyzes real-time images of the barbecue grill. Based on the analysis results, a cooking management means generates advice such as adjusting the heat or changing the position of ingredients. The generated advice is sent to the user terminal via a communication means and presented to the user via a user interface means.
[0482] Program processing
[0483] 1. The user starts a barbecue
[0484] The user issues the voice command "Start the barbecue."
[0485] The terminal receives the voice, and the voice recognition means converts it into text data, which is then sent to the server.
[0486] 2. Analysis of audio data
[0487] The server analyzes the text data and determines that the user is about to start a barbecue, and then performs the initial setup.
[0488] 3. Preparatory Instructions
[0489] The server generates preparation instructions for the user (e.g., "Prepare the grill") based on the initial settings and sends them to the terminal.
[0490] The device will provide instructions to the user via voice or visual display.
[0491] 4. Real-time image acquisition
[0492] The user points the device's camera at the barbecue grill.
[0493] The device uses a camera to capture an image of the grill and transmits the image data to a server.
[0494] 5. Analysis of Image Data
[0495] The server analyzes the image data and evaluates the placement of ingredients, the state of the fire, and the degree of doneness.
[0496] The server determines the necessary adjustments based on the analysis results.
[0497] 6. Generating and Providing Advice
[0498] The server generates advice on adjusting the heat and changing the position of ingredients and sends it to the terminal.
[0499] The device presents the generated advice to the user by voice or on-screen display.
[0500] The user makes adjustments according to the advice.
[0501] Specific examples
[0502] The process of starting a barbecue
[0503] 1. The user issues a voice command
[0504] User: "Start a barbecue."
[0505] The terminal converts the voice command into text data using a voice recognition means and transmits it to the server.
[0506] 2. The server analyzes the request and sets up the initial settings
[0507] Server: Parses the user's request, performs initialization, and generates preparation instructions.
[0508] Terminal: Prompts the user with the instruction "Prepare the grill."
[0509] Real-time image analysis and advice
[0510] 1. User takes a picture of the grill
[0511] User: Point your device camera at the grill.
[0512] The device sends image data to a server, which analyzes the image.
[0513] 2. The server analyzes the image and generates advice
[0514] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[0515] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[0516] The system of the present invention allows users to easily enjoy delicious barbecues without detailed knowledge or experience. It also significantly reduces the effort required for cooking, allowing users to spend more time concentrating on conversations and meals with family and friends.
[0517] The processing flow will be explained below.
[0518] Step 1:
[0519] The user issues the voice command "Start the barbecue."
[0520] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[0521] Step 2:
[0522] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[0523] The server performs the initial setup and creates the necessary setup instructions.
[0524] Step 3:
[0525] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[0526] The terminal presents instructions to the user by voice or on-screen display.
[0527] The user follows the instructions to prepare the grill.
[0528] Step 4:
[0529] The user points the device's camera at the grill and places the ingredients on it.
[0530] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[0531] Step 5:
[0532] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[0533] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0534] Step 6:
[0535] The server generates advice (e.g., "Turn up the heat a little," "Move the chicken to the right") and sends it to the device.
[0536] The advice received by the terminal is presented to the user by voice or screen display.
[0537] The user adjusts the heat and position of the ingredients according to the advice.
[0538] Step 7:
[0539] The user asks the device, "What should I make next?"
[0540] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0541] Step 8:
[0542] The server selects the optimal recipe based on the user's past cooking history and current ingredient information.
[0543] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[0544] The terminal presents recipe suggestions to the user by voice or on-screen display.
[0545] Step 9:
[0546] The server periodically analyzes the images to determine whether the food is cooked.
[0547] The server generates a notification saying "The chicken is done" and sends it to the device.
[0548] The terminal notifies the user by voice when the food is done.
[0549] The user removes the cooked chicken and enjoys the meal.
[0550] Example 1
[0551] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0552] In traditional barbecue cooking, beginners and those with little experience often have difficulty controlling the heat and arranging ingredients properly. Additionally, the lack of real-time monitoring and proper advice can lead to inconsistent cooking quality. This can result in users spending a lot of time and effort on cooking, preventing them from fully enjoying their barbecue experience.
[0553] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0554] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a real-time data collection unit, a data analysis unit, and an advice generation unit. This allows users to start a barbecue using voice commands and manage the cooking appropriately through real-time image data analysis. Furthermore, specific advice based on the analysis results can be received by voice or displayed on the screen, allowing even beginners to easily enjoy a high-quality barbecue experience.
[0555] "Voice recognition means" refers to technology that receives voice commands from a user and converts them into text data.
[0556] "Image analysis means" refers to technology that acquires real-time images of the barbecue grill and analyzes the placement of ingredients, the state of the fire, and the degree of doneness.
[0557] "Cooking management means" refers to technology that manages the cooking process based on image analysis results and audio data.
[0558] "Communication means" refers to the technology for sending and receiving data between a terminal and a server.
[0559] "User interface means" refers to means for providing visual or audio information to a user and accepting input from the user.
[0560] "Real-time data collection means" refers to technology that acquires images and audio data of barbecue grills in real time.
[0561] "Data analysis means" refers to technology that analyzes collected voice and image data to evaluate the user's intentions and the condition of the grill.
[0562] "Advice generation means" refers to technology that generates specific advice necessary for cooking based on the results of data analysis.
[0563] This invention is a smart AI assistant system that supports barbecue experiences, and includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, a real-time data collection means, a data analysis means, and an advice generation means.
[0564] System Overview
[0565] This system helps users start a barbecue. The user starts the system by issuing the voice command "Start a barbecue." Each step and the hardware and software used are explained below.
[0566] Hardware and software used
[0567] The system uses the following major hardware and software:
[0568] Speech recognition method: Google Cloud Speech-to-Text
[0569] Image analysis methods: OpenCV, TensorFlow
[0570] Cooking control means: Control logic within the system
[0571] Communication method: Wi-Fi or Bluetooth
[0572] User interface means: smartphone application
[0573] Real-time data collection method: built-in camera and microphone of smartphone
[0574] Data analysis method: Google Cloud Natural Language API
[0575] Advice generation means: Advice generation engine within the system
[0576] Program processing
[0577] The program begins when the user issues the voice command "Start the barbecue." The device receives this voice, converts it into text data using Google Cloud Speech-to-Text, and sends the data to the server. The server then analyzes the text data using the Google Cloud Natural Language API, understands the user's intent, and sets up the initial barbecue setup.
[0578] Once the initial setup is complete, the server generates preparation instructions for the user, such as "Prepare the grill," and sends them to the device, which then presents the instructions to the user as a voice message or a screen display.
[0579] Next, the user points their smartphone camera at the barbecue grill. The device uses the camera to capture real-time images and sends the image data to a server. The server then uses image analysis technology (OpenCV, TensorFlow) to analyze the images and evaluate the placement of ingredients, the state of the fire, and the degree of doneness. Based on the evaluation, the server determines the necessary adjustments and generates specific advice (e.g., "Turn up the heat a little" or "Move the meat to the right").
[0580] The generated advice is sent to the device, which then presents it to the user via voice or screen display, allowing the user to make adjustments according to the advice.
[0581] Specific examples
[0582] The process of starting a barbecue
[0583] 1. The user issues a voice command
[0584] User: "Start a barbecue."
[0585] The device uses a voice recognition tool (Google Cloud Speech-to-Text) to convert the voice command into text data and send it to the server.
[0586] 2. The server analyzes the request and sets up the initial settings
[0587] Server: Parses the user request (Google Cloud Natural Language API), performs initialization, and generates preparation instructions.
[0588] Terminal: Prompt the user (audio or screen prompt) to "Prepare the grill."
[0589] Real-time image analysis and advice
[0590] 1. User takes a picture of the grill
[0591] User: Point your device camera at the grill.
[0592] The device sends image data to the server (captured by the smartphone's built-in camera), and the server analyzes the image (OpenCV, TensorFlow).
[0593] 2. The server analyzes the image and generates advice
[0594] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[0595] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[0596] Example of input prompt for generative AI model
[0597] 1. Prompts for user command to generate preparation instructions when starting a barbecue:
[0598] After receiving a voice command from the user to "start barbecue," perform the initial setup and generate and present preparation instructions to the user.
[0599] Example: If the user says "I want to start the barbecue," the device might display or speak instructions such as "Prepare the grill."
[0600] 2. Real-time image analysis prompts advice generation:
[0601] Based on the analysis of real-time images of the barbecue grill, generate and present specific advice to the user on adjusting the heat or repositioning the ingredients.
[0602] For example, if the grill is running low, you might offer advice such as, "Please turn up the heat a little."
[0603] This system allows users to easily enjoy delicious food even if they have no knowledge or experience of barbecue. The system provides users with easy-to-understand instructions and helps them cook.
[0604] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0605] Step 1:
[0606] The user issues a voice command such as "Start the barbecue."
[0607] Input: Voice commands from the user.
[0608] Action: The device receives audio through the microphone.
[0609] Output: The captured audio data is generated.
[0610] Step 2:
[0611] The terminal converts the received voice data into text data using a voice recognition means.
[0612] Input: Audio data.
[0613] How it works: Converts audio data into text using Google Cloud Speech-to-Text.
[0614] Output: Text data is generated and sent to the server.
[0615] Step 3:
[0616] The server analyzes the received text data and understands the user's intent.
[0617] Input: Text data.
[0618] How it works: It uses the Google Cloud Natural Language API to analyze text data and recognize the user's intention to start a barbecue.
[0619] Output: Configuration data reflecting the initial settings is generated.
[0620] Step 4:
[0621] The server generates and sends preparation instructions to the user based on the initial settings.
[0622] Input: Configuration data.
[0623] Operation: Using the cooking management means, a preparation instruction such as "Prepare the grill" is generated and sent to the terminal.
[0624] Output: Preparation instruction data is sent to the terminal.
[0625] Step 5:
[0626] The terminal presents the received preparation instructions to the user.
[0627] Input: Preparation instruction data.
[0628] What it does: Uses user interface means to present instructions to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[0629] Output: The user receives the instructions.
[0630] Step 6:
[0631] A user points their smartphone camera at a barbecue grill.
[0632] Input: Instructions from the server.
[0633] How it works: Point your smartphone camera at the barbecue grill to capture real-time images.
[0634] Output: Real-time image data is generated.
[0635] Step 7:
[0636] The real-time images acquired by the terminal are sent to the server.
[0637] Input: Real-time image data.
[0638] Operation: Image data captured using the camera is sent to the server.
[0639] Output: The image data is transferred to the server.
[0640] Step 8:
[0641] The server analyzes the received image data.
[0642] Input: Image data.
[0643] How it works: Image analysis technology (OpenCV, TensorFlow) is used to analyze the placement of ingredients, the state of the fire, and the degree of doneness.
[0644] Output: Analysis result data is generated.
[0645] Step 9:
[0646] The server generates and sends advice based on the analysis results.
[0647] Input: Analysis result data.
[0648] Operation: Using the advice generation means, specific advice regarding adjusting the heat or changing the position of ingredients is generated and sent to the terminal.
[0649] Output: Advice data is sent to the terminal.
[0650] Step 10:
[0651] The device presents the received advice to the user.
[0652] Input: Advice data.
[0653] What it does: Uses user interface means to present advice to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[0654] Output: User receives advice.
[0655] Step 11:
[0656] The user adjusts the barbecue grill according to the provided advice.
[0657] Enter: Advice.
[0658] Action: Making specific adjustments such as adjusting the heat or changing the position of ingredients.
[0659] Output: Adjusted BBQ grill.
[0660] (Application example 1)
[0661] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0662] Traditional barbecue and cooking relied on the knowledge and experience of the chef, making it difficult to ensure consistent quality. It was also difficult to monitor the status of the food in real time and make appropriate adjustments, which required a lot of time and effort. This made it difficult for inexperienced chefs to provide satisfying food.
[0663] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0664] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a unit for evaluating and adjusting cooking quality in real time, and a unit for providing advice using a generative AI model. This makes it possible to provide high-quality, consistent cooking without relying on the chef's experience. Furthermore, by being able to grasp the cooking status in real time and make appropriate adjustments, it is possible to significantly reduce the effort required and provide a satisfying cooking experience.
[0665] "Voice recognition means" refers to any device or technology that receives voice commands from a user and converts them into text data.
[0666] "Image analysis means" refers to devices and technologies in general that analyze image data captured in real time and recognize specific objects or conditions.
[0667] "Cooking management means" refers to all devices and technologies that manage cooking based on the analysis results and generate instructions such as adjusting the heat or changing the position of ingredients.
[0668] "Communication means" refers to all devices and technologies for transmitting and receiving data.
[0669] "User interface means" refers generally to devices and techniques that allow a user to interact with a system and obtain information.
[0670] "Means for evaluating and adjusting food quality in real time" refers to all devices and technologies that analyze data acquired in real time, evaluate food quality, and make appropriate adjustments.
[0671] "Means for providing advice using a generative AI model" refers to all devices and technologies that use a generative AI model to provide appropriate advice to users based on the analysis results.
[0672] This invention is a smart AI cooking assistant system for food delivery, which includes voice recognition means, image analysis means, cooking management means, communication means, user interface means, means for evaluating and adjusting cooking quality in real time, and means for providing advice using a generative AI model, enabling cooks to deliver high-quality, consistent food.
[0673] Specifically, users operate the system using a smartphone or head-mounted display. For example, when a user issues a voice command such as "Start a barbecue," the smartphone's microphone captures the voice and converts it into text data using a speech recognition library. The converted text data is then sent to a server for initial setup.
[0674] Next, the user uses their smartphone camera to capture images of the grill in real time. The image data is captured using the OpenCV library and sent to the server. The server analyzes the images and evaluates the placement of ingredients and the state of the flame. Based on the evaluation, the cooking management tool generates advice on adjusting the heat or repositioning ingredients.
[0675] The generated advice is sent as text data to the smartphone via the communication means. The advice is presented to the user via the user interface means as voice or on-screen display. For example, if the advice is "Turn up the heat a little," the user adjusts the heat according to the instruction.
[0676] The system can also use generative AI models to provide more advanced advice, such as providing tailored advice in response to prompts like "What's the best way to adjust the heat on my barbecue?" or "How do I check if my chicken is done?"
[0677] In this way, the system enables consistent, high-quality food to be served in real time, while significantly reducing the chef's workload and providing a satisfying cooking experience.
[0678] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0679] Step 1:
[0680] The user issues the voice command "Start the barbecue." The device captures the voice command using the built-in microphone and converts the voice into text data using a voice recognition library (e.g., SpeechRecognition library). This text data is sent to the server. The input is the voice command, and the output is text data. Data processing involves analyzing the voice waveform data and converting it into text data.
[0681] Step 2:
[0682] The server receives the text data and recognizes that the user is about to start a barbecue. The server generates a preparation instruction, "Prepare the grill," as the initial setting and sends this instruction to the terminal. The input is the text data, and the output is the initial setting and the preparation instruction. As a data operation, it performs text analysis and sets the initial setting flag.
[0683] Step 3:
[0684] The device receives the preparation instruction from the server and presents it to the user by voice or on-screen display. The input is the preparation instruction from the server, and the output is the instruction presented to the user. Specifically, the instruction is played back by voice using the device's speaker.
[0685] Step 4:
[0686] A user points their smartphone camera at a barbecue grill. The device uses the camera to capture real-time images of the grill and sends the image data to a server. The input is the camera image and the output is image data. Image capture and compression are performed as data processing.
[0687] Step 5:
[0688] The server analyzes the image data it receives and evaluates the placement of ingredients and the state of the fire. This is done using an image analysis library (e.g., OpenCV). The input is the image data, and the output is the analysis results. Data calculations involve extracting basic image features and obtaining information on the heat level and the position of ingredients.
[0689] Step 6:
[0690] The server uses the cooking management means based on the analysis results to generate advice on adjusting the heat or changing the position of ingredients. This advice is sent to the terminal as text data via communication means. The input is the analysis results, and the output is advice. As a data calculation, appropriate cooking instructions are automatically generated based on the analysis results.
[0691] Step 7:
[0692] The device receives advice from the server and presents it to the user by voice or on-screen display. The input is advice from the server, and the output is instructions presented to the user. Specifically, the advice is played back by voice using the device's speaker, or displayed as text on the screen.
[0693] Step 8:
[0694] The user makes adjustments based on the advice from the server. This step acts as a feedback loop, returning to step 4 and beyond to make the next adjustment if necessary. The input is the advice, and the output is the adjustment action taken by the user. Specific actions taken by the user include adjusting the heat or changing the position of ingredients.
[0695] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0696] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. Below, we will explain the program and processing of this system with concrete examples.
[0697] System Overview
[0698] The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, and an emotion engine. The voice recognition unit receives a user's voice command and converts it into text data. The image analysis unit acquires and analyzes real-time images of the barbecue grill. The cooking management unit generates advice based on the results of the image analysis, which is sent to the user terminal via the communication unit and presented to the user via the user interface unit. The emotion engine also analyzes the user's emotions and adjusts the cooking advice based on the analysis.
[0699] Program processing
[0700] 1. The user starts a barbecue
[0701] The user issues the voice command "Start the barbecue."
[0702] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[0703] 2. Analysis of audio data
[0704] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[0705] The server performs the initial setup and creates the necessary setup instructions.
[0706] 3. Preparatory Instructions
[0707] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[0708] The device will provide instructions to the user via voice or on-screen display.
[0709] The user follows the instructions to prepare the grill.
[0710] 4. Real-time image acquisition
[0711] The user points the device's camera at the grill and places the ingredients on it.
[0712] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[0713] 5. Analysis of Image Data
[0714] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[0715] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0716] 6. Emotion Analysis Using an Emotion Engine
[0717] The device captures the user's facial expressions and tone of voice and sends them to the emotion engine.
[0718] The server uses an emotion engine to analyze the user's emotions and adjusts cooking advice based on this.
[0719] 7. Generating and Providing Advice
[0720] The server adjusts the advice based on the emotion data and sends the generated advice (e.g., "Turn up the heat a little," "Move the chicken to the right") to the terminal.
[0721] The advice received by the device is presented to the user by voice or on-screen display.
[0722] The user adjusts the heat and position of the ingredients according to the advice.
[0723] 8. Recipe suggestions
[0724] The user asks the device, "What should I make next?"
[0725] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0726] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[0727] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[0728] The device will present recipe suggestions to the user via voice or on-screen display.
[0729] 9. Emotional Cooking Guidance
[0730] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[0731] The server generates cooking guidance based on the user's emotions (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device.
[0732] The device will notify the user of emotionally sensitive notifications via voice.
[0733] The user removes the cooked chicken and enjoys the meal.
[0734] The system of the present invention allows users to easily enjoy delicious barbecues, even without detailed knowledge or experience. Furthermore, by combining it with an emotion engine, a more personalized and comfortable barbecue experience can be provided. By significantly reducing the effort required for cooking and providing emotionally appropriate encouragement and advice, users can spend more time concentrating on conversations with family and friends and on their meals.
[0735] The processing flow will be explained below.
[0736] Step 1:
[0737] The user issues the voice command "Start the barbecue."
[0738] The device uses a microphone to capture the user's voice.
[0739] Step 2:
[0740] The terminal converts the voice into text data using a voice recognition means.
[0741] The terminal transmits the converted text data to the server.
[0742] Step 3:
[0743] The server receives the text data, analyzes it, and recognizes it as a request to "start barbecue."
[0744] The server performs the initialization and generates the necessary provisioning instructions.
[0745] Step 4:
[0746] The server sends the generated preparation instruction to the terminal.
[0747] The terminal will give the user instructions to prepare the grill by voice or on the screen, such as "Prepare the grill."
[0748] Step 5:
[0749] The user follows the instructions to prepare the grill.
[0750] The user points the device's camera at the grill and captures an image.
[0751] Step 6:
[0752] The device captures an image of the grill and sends the image data to a server.
[0753] The server receives the image data and uses image analysis means to evaluate the placement of ingredients, the state of the fire, and the degree of doneness.
[0754] Step 7:
[0755] Based on the results of image analysis, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0756] The server transmits the generated advice to the terminal.
[0757] Step 8:
[0758] The device displays the received advice to the user (e.g., "Turn up the heat a little").
[0759] The user adjusts the heat and position of the ingredients according to the advice.
[0760] Step 9:
[0761] The device captures the user's facial expressions and voice tone and transmits the emotional data to the server.
[0762] The server analyzes the user's emotions using an emotion engine.
[0763] Step 10:
[0764] The server adjusts cooking advice based on the results of emotion analysis (e.g., if the user seems anxious, give advice in a gentler tone).
[0765] The server sends the adjusted advice to the terminal.
[0766] Step 11:
[0767] The terminal presents the adjusted advice to the user.
[0768] The user proceeds with cooking according to the advice.
[0769] Step 12:
[0770] The user asks the device, "What should I make next?"
[0771] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0772] Step 13:
[0773] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[0774] The server generates recipe suggestions and sends them to the terminal.
[0775] Step 14:
[0776] The device will present recipe suggestions (e.g., "We recommend this chicken marinade") to the user via voice or on-screen display.
[0777] Users can try new dishes by following suggested recipes.
[0778] Step 15:
[0779] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[0780] The server generates a notification saying "The chicken is done" and sends it to the device.
[0781] Step 16:
[0782] The terminal notifies the user by voice when the food is done.
[0783] The user removes the cooked chicken and enjoys the meal.
[0784] Example 2
[0785] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0786] Enjoying a barbecue often requires skilled techniques and a lot of prior knowledge. However, for users who are trying barbecue for the first time or who are unfamiliar with cooking, these requirements are a high hurdle and make it difficult to fully enjoy a barbecue. In addition, there is a problem that the user experience is inconsistent because the cooking status is not grasped in real time or appropriate advice is not provided according to the user's emotions.
[0787] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, an emotion analysis means, and a user interface means. This allows even beginners to easily receive detailed cooking advice on barbecues, and enables personalized support tailored to the user's emotions. Specifically, the voice recognition means recognizes the user's voice commands and converts them into text data. The image analysis means captures real-time images of the grill and analyzes the state of the ingredients and the fire. Furthermore, the emotion analysis means analyzes the user's emotions from their facial expressions and voice tone. The cooking management means integrates these analysis results to generate appropriate cooking advice, which is then provided to the user through the user interface means.
[0788] "Speech recognition means" refers to means having the function of receiving voice commands from a user and converting them into text data.
[0789] The "image analysis means" is a means having the function of analyzing images captured in real time and evaluating the state of the ingredients and the fire.
[0790] The "cooking management means" is a means having a function of generating cooking advice based on the results of the image analysis means and the emotion analysis means.
[0791] "Communication means" refers to a means that utilizes protocols and technologies for sending and receiving data.
[0792] A "user interface means" is a means for presenting information to a user and receiving input from a user.
[0793] The "emotion analysis means" is a means having the function of capturing the user's facial expressions and tone of voice and analyzing their emotions.
[0794] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. This system includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, and an emotion analysis means.
[0795] Hardware and software used
[0796] The system uses the following major hardware and software:
[0797] 1. Speech recognition means: Typically using a cloud-based speech recognition service, e.g., a speech recognition cloud API.
[0798] 2. Image analysis method: Uses the OpenCV library and deep learning models (e.g., TensorFlow).
[0799] 3. Cooking control method: A custom algorithm implemented in Python is used.
[0800] 4. Communication method: Network communication is performed using the HTTP / HTTPS protocol.
[0801] 5. User interface: Implemented through a smartphone app (iOS / Android) or web browser.
[0802] 6. Emotion analysis means: Use an API that performs facial expression analysis and voice tone analysis (e.g., emotion analysis cloud API).
[0803] Processing of this system
[0804] 1. When a user wants to start a barbecue, they say the voice command "Start the barbecue." This voice command is converted into text data by the device using a voice recognition cloud API and sent to the server.
[0805] 2. The server analyzes the received text data and recognizes that the user's request is to "start a barbecue." It uses a natural language processing (NLP) library to process this. The server then performs its initial setup and creates the necessary preparation instructions.
[0806] 3. The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the device. The device then presents the instructions to the user via voice or visual display, and the user prepares the grill accordingly.
[0807] 4. The user points the device's camera at the grill and places the food on it. The device uses the camera to capture real-time images of the grill and sends the image data to the server.
[0808] 5. The server analyzes the received image data using OpenCV and deep learning models. The analysis results include the placement of the ingredients, the state of the fire, and the degree of doneness. Based on this information, the server generates specific advice on adjusting the heat or repositioning the ingredients.
[0809] 6. The device captures the user's facial expressions and voice tone and sends them to the emotion analysis cloud API. The server uses an emotion engine to analyze the user's emotions and adjust the cooking advice accordingly.
[0810] 7. The server generates advice based on the emotion data and sends it to the device. The device then presents the advice to the user via voice or on-screen display. The user adjusts the heat and position of the ingredients according to the advice.
[0811] 8. The user asks the device, "What should I make next?" The device uses a voice recognition cloud API to convert the question into text data and sends it to the server. The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The server generates a recipe suggestion (e.g., "I recommend marinated chicken") and sends it to the device. The device then presents the recipe suggestion to the user by voice or on a screen.
[0812] 9. The server periodically performs image and emotion analysis to determine whether the food is cooked. The server generates cooking guidance based on the user's emotion (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device. The device then notifies the user with voice notifications that take emotion into consideration. The user follows the instructions to remove the cooked chicken and enjoy their meal.
[0813] Specific examples
[0814] Example 1: Your first barbecue
[0815] 1. The user says the voice command "Start the barbecue."
[0816] 2. The device converts the voice into text and sends it to the server.
[0817] 3. The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal.
[0818] 4. The user prepares the grill and points the device camera at the grill.
[0819] 5. The device sends the captured image to the server, which then analyzes the image.
[0820] 6. The server generates advice on adjusting the heat and sends it to the device.
[0821] 7. The device provides advice to the user, who then adjusts the heat.
[0822] Prompt Sentence Examples
[0823] Prompt 1: "Start the BBQ Assistant to provide a set of guidance for first-time barbecuers."
[0824] Prompt 2: "When a user requests a new recipe, recommend a recipe that fits their mood that day."
[0825] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0826] Step 1:
[0827] The user issues the voice command "Start the barbecue."
[0828] The device converts the speech into text data using a speech recognition cloud API. The input is the user's speech, and the output is the converted text data. The device then sends this text data to the server.
[0829] Step 2:
[0830] The server analyzes the received text data. The input is text data sent from the device, and the server uses a natural language processing library to recognize that it is a request to "start a barbecue." The output is the result of recognizing the request as "start a barbecue." The server performs initial configuration and creates the necessary preparation instructions.
[0831] Step 3:
[0832] The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal. The input is the initial setting result, and the output is a text message of the preparation instruction. The terminal receives this instruction and presents it to the user by voice or on-screen display. The user follows the instruction and performs the action of preparing the grill.
[0833] Step 4:
[0834] The user points the device's camera at the grill and places ingredients on it. The input is the camera image. The device uses the camera to capture real-time images of the grill and sends the image data to the server. The output is the captured real-time image data.
[0835] Step 5:
[0836] The server analyzes the received image data. The input is image data sent from the device. The server uses OpenCV and deep learning models (such as TensorFlow) to evaluate the placement of ingredients, the state of the fire, and the degree of doneness. The output is the analysis results, which are summarized as specific advice on adjusting the heat or changing the position of ingredients.
[0837] Step 6:
[0838] The device captures the user's facial expressions and voice tone. The input is the user's facial expression and voice data. The device sends this to the emotion analysis cloud API. The output is the emotion analysis result, which is sent to the server.
[0839] Step 7:
[0840] The server receives the emotion data analysis results and generates adjustment advice. The input is emotion data and image analysis results. The server adjusts the cooking advice based on this data and generates specific advice (e.g., "Turn up the heat a little," "Move the chicken to the right"). The output is the adjusted advice. The server sends this advice to the device.
[0841] Step 8:
[0842] The device provides advice to the user by voice or on-screen display. The input is advice data from the server, and the output is the information presented to the user. The user follows the advice and adjusts the heat and position of the ingredients.
[0843] Step 9:
[0844] The user asks the device, "What should I make next?" The input is the user's voice command. The device again uses the speech recognition cloud API to convert the question into text data and sends that data to the server. The output is the converted text data.
[0845] Step 10:
[0846] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The input is the user's past cooking history, current ingredient information, and emotional data, and the output is the optimal recipe suggestion. The server generates a recipe suggestion (e.g., "We recommend marinated chicken") and sends it to the device.
[0847] Step 11:
[0848] The terminal presents recipe suggestions to the user by voice or on-screen display. The input is the recipe suggestion data from the server, and the output is the suggestion to the user.
[0849] Step 12:
[0850] The server periodically performs image analysis and emotion analysis. The input is real-time images and user emotion data. The server determines whether the ingredients are cooked and generates cooking guidance. The output is cooking guidance (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?"). The server then sends the generated cooking guidance to the device.
[0851] Step 13:
[0852] The device provides emotionally sensitive notifications to the user. The input is cooking guidance data from the server, and the output is a notification to the user. The user follows the instructions, removes the cooked chicken, and enjoys the meal.
[0853] Through the above processing steps, the smart AI assistant system of the present invention can help users enjoy a more enjoyable and easier barbecue. By providing personalized advice based on the cooking progress and the user's emotions, even beginners can enjoy a delicious barbecue.
[0854] (Application example 2)
[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0856] Conventional food delivery systems lack personalized menu suggestions and cooking advice based on the user's emotional state, which hinders the improvement of user experience. Furthermore, when users are cooking, they often do not receive appropriate advice in real time, which leads to a decrease in satisfaction.
[0857] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, an emotion analysis means, and a database means. This makes it possible to propose optimal menus and provide cooking advice according to the user's emotional state, which is expected to improve the user experience.
[0858] "Voice recognition" is a technology that receives a user's voice commands and converts them into text data.
[0859] "Image analysis" is a technology that analyzes the state of ingredients and fire based on image data captured in real time.
[0860] "Cooking management" is a technology that generates cooking advice based on the results of image analysis and presents it to the user.
[0861] "Communication means" refers to the technology used to send and receive voice data, image data, and analysis results between the server and the user terminal.
[0862] A "user interface" is a means for receiving voice commands and image data input from a user, and for providing generated advice and menu suggestions to the user.
[0863] "Emotion analysis" is a technology that analyzes a user's facial expressions and tone of voice to understand their emotional state.
[0864] The "database" is a technology that stores a user's past order history and emotional data, and generates suggestions based on the analysis results as needed.
[0865] The present invention relates to an emotion-responsive food delivery support application, specifically a system that provides optimal menu suggestions and cooking advice based on a user's voice commands and emotional state. The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, an emotion analysis unit, and a database unit.
[0866] System Overview
[0867] The solution begins when a user launches a food delivery application on their smartphone. First, the user issues a voice command such as "Tell me what menu items you want." The smartphone uses a microphone to capture the voice and converts it into text using a speech recognition library called Google Cloud Speech-to-Text. This text is then sent to a server, which analyzes it and understands the user's request.
[0868] Next, an emotion analysis tool analyzes the user's facial expressions and tone of voice to generate emotion data. This uses the Microsoft Azure Emotion API. The emotion data is used to understand the user's emotional state, and appropriate menu suggestions and cooking advice are provided based on that emotional state. Past order history and emotion data are stored in a database (e.g., MySQL), and optimal suggestions are generated based on this data.
[0869] For example, if the user is in a state of mind where they want to relax, the system will suggest "relaxing herbal tea." If the user is cooking at home, the system will perform real-time image analysis and provide appropriate cooking advice. Image analysis uses a library called OpenCV.
[0870] The specific hardware and software used
[0871] Hardware: Smartphone (microphone, camera, speaker)
[0872] software:
[0873] Speech recognition library: Google Cloud Speech-to-Text
[0874] Sentiment analysis library: Microsoft Azure Emotion API
[0875] Image analysis library: OpenCV
[0876] Database: MySQL
[0877] Prompt Sentence Examples
[0878] Specific examples of prompts are as follows:
[0879] User's order history: pizza, pasta, salad
[0880] Emotional state: Relaxed
[0881] User Question: What's your recommended menu item?
[0882] Response: I recommend herbal tea. It's perfect for relaxing.
[0883] In this way, the system of the present invention allows users to enjoy a personalized food delivery experience based on their voice commands and emotional state.
[0884] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0885] Step 1:
[0886] A user launches a food delivery application on their smartphone and says, "Tell me what menu items you recommend." Here, the input is the user's voice command, and the smartphone's microphone is used to capture the voice data. This voice data is converted into text data using the Google Cloud Speech-to-Text library. The converted text data is the output.
[0887] Step 2:
[0888] The speech-recognized text data is sent to the server. Here, the input is the converted text data, and the server receives the data. The server performs text analysis to understand the user's request. As a result of the analysis, data called "recommended menu suggestions" is output.
[0889] Step 3:
[0890] The emotion analyzer captures the user's facial expressions and voice tone and generates emotion data using the Microsoft Azure Emotion API, where the input is the user's real-time facial expression data and voice tone, and the API analyzes the user's emotional state (e.g., relaxed, happy, etc.), and the analyzed emotion data is the output.
[0891] Step 4:
[0892] The server retrieves the generated emotion data and past order history from a database. Here, the input is emotion data and order history data. This data is retrieved through a database query, and the server selects the optimal menu based on this data. The selected menu suggestion is the output.
[0893] Step 5:
[0894] The server generates a selected menu suggestion (e.g., "I recommend herbal tea. It's perfect for relaxing.") and sends it to the smartphone via communication means. Here, the input is the data of the selected menu suggestion, which is sent to the smartphone. The output is the smartphone presenting the suggested menu to the user by voice or on-screen display.
[0895] Step 6:
[0896] When a user cooks at home, they use their smartphone camera to capture ingredients and the cooking process. The input is real-time image data, which is analyzed using an image analysis library (OpenCV). The analysis results are then output as cooking advice (e.g., "Adjust the heat.").
[0897] Step 7:
[0898] The server sends the generated cooking advice to the smartphone and presents it to the user. Here, the input is the analyzed cooking advice data, and the data is sent to the smartphone. The output of the smartphone is to provide advice to the user by voice or on-screen display.
[0899] Step 8:
[0900] The user then asks a follow-up question (e.g., "What's the next recommended menu item?") by voice. The smartphone's microphone is again used to capture the voice data, which is then converted into text data using the Google Cloud Speech-to-Text library. This voice data is the input, and the converted text data is the output.
[0901] Step 9:
[0902] The server analyzes the converted text data and generates the next menu suggestion based on the information from the database and the sentiment data. Here, the input is the text data based on the follow-up question and related data, and the output is the next menu suggestion as the analysis result.
[0903] Step 10:
[0904] The server sends the next menu proposal to the smartphone and presents it to the user. Here, the input is the generated next menu proposal data, which is sent to the smartphone. The output of the smartphone is to present it to the user by voice or on-screen display.
[0905] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0906] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0907] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0908] [Third embodiment]
[0909] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0910] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0911] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0912] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0913] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0914] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0915] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0916] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0917] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0918] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0919] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0920] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0921] This invention is a smart AI assistant system that supports the barbecue experience, and includes voice recognition, image analysis, cooking management, communication, and user interface. The program and processing of this system are explained below in natural language with concrete examples.
[0922] System Overview
[0923] When a user starts a barbecue, this system uses a voice recognition means to receive the user's voice commands and convert them into text data. Next, an image analysis means captures and analyzes real-time images of the barbecue grill. Based on the analysis results, a cooking management means generates advice such as adjusting the heat or changing the position of ingredients. The generated advice is sent to the user terminal via a communication means and presented to the user via a user interface means.
[0924] Program processing
[0925] 1. The user starts a barbecue
[0926] The user issues the voice command "Start the barbecue."
[0927] The terminal receives the voice, and the voice recognition means converts it into text data, which is then sent to the server.
[0928] 2. Analysis of audio data
[0929] The server analyzes the text data and determines that the user is about to start a barbecue, and then performs the initial setup.
[0930] 3. Preparatory Instructions
[0931] The server generates preparation instructions for the user (e.g., "Prepare the grill") based on the initial settings and sends them to the terminal.
[0932] The device will provide instructions to the user via voice or visual display.
[0933] 4. Real-time image acquisition
[0934] The user points the device's camera at the barbecue grill.
[0935] The device uses a camera to capture an image of the grill and transmits the image data to a server.
[0936] 5. Analysis of Image Data
[0937] The server analyzes the image data and evaluates the placement of ingredients, the state of the fire, and the degree of doneness.
[0938] The server determines the necessary adjustments based on the analysis results.
[0939] 6. Generating and Providing Advice
[0940] The server generates advice on adjusting the heat and changing the position of ingredients and sends it to the terminal.
[0941] The device presents the generated advice to the user by voice or on-screen display.
[0942] The user makes adjustments according to the advice.
[0943] Specific examples
[0944] The process of starting a barbecue
[0945] 1. The user issues a voice command
[0946] User: "Start a barbecue."
[0947] The terminal converts the voice command into text data using a voice recognition means and transmits it to the server.
[0948] 2. The server analyzes the request and sets up the initial settings
[0949] Server: Parses the user's request, performs initialization, and generates preparation instructions.
[0950] Terminal: Prompts the user with the instruction "Prepare the grill."
[0951] Real-time image analysis and advice
[0952] 1. User takes a picture of the grill
[0953] User: Point your device camera at the grill.
[0954] The device sends image data to a server, which analyzes the image.
[0955] 2. The server analyzes the image and generates advice
[0956] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[0957] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[0958] The system of the present invention allows users to easily enjoy delicious barbecues without detailed knowledge or experience. It also significantly reduces the effort required for cooking, allowing users to spend more time concentrating on conversations and meals with family and friends.
[0959] The processing flow will be explained below.
[0960] Step 1:
[0961] The user issues the voice command "Start the barbecue."
[0962] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[0963] Step 2:
[0964] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[0965] The server performs the initial setup and creates the necessary setup instructions.
[0966] Step 3:
[0967] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[0968] The terminal presents instructions to the user by voice or on-screen display.
[0969] The user follows the instructions to prepare the grill.
[0970] Step 4:
[0971] The user points the device's camera at the grill and places the ingredients on it.
[0972] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[0973] Step 5:
[0974] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[0975] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[0976] Step 6:
[0977] The server generates advice (e.g., "Turn up the heat a little," "Move the chicken to the right") and sends it to the device.
[0978] The advice received by the terminal is presented to the user by voice or screen display.
[0979] The user adjusts the heat and position of the ingredients according to the advice.
[0980] Step 7:
[0981] The user asks the device, "What should I make next?"
[0982] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[0983] Step 8:
[0984] The server selects the optimal recipe based on the user's past cooking history and current ingredient information.
[0985] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[0986] The terminal presents recipe suggestions to the user by voice or on-screen display.
[0987] Step 9:
[0988] The server periodically analyzes the images to determine whether the food is cooked.
[0989] The server generates a notification saying "The chicken is done" and sends it to the device.
[0990] The terminal notifies the user by voice when the food is done.
[0991] The user removes the cooked chicken and enjoys the meal.
[0992] Example 1
[0993] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0994] In traditional barbecue cooking, beginners and those with little experience often have difficulty controlling the heat and arranging ingredients properly. Additionally, the lack of real-time monitoring and proper advice can lead to inconsistent cooking quality. This can result in users spending a lot of time and effort on cooking, preventing them from fully enjoying their barbecue experience.
[0995] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0996] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a real-time data collection unit, a data analysis unit, and an advice generation unit. This allows users to start a barbecue using voice commands and manage the cooking appropriately through real-time image data analysis. Furthermore, specific advice based on the analysis results can be received by voice or displayed on the screen, allowing even beginners to easily enjoy a high-quality barbecue experience.
[0997] "Voice recognition means" refers to technology that receives voice commands from a user and converts them into text data.
[0998] "Image analysis means" refers to technology that acquires real-time images of the barbecue grill and analyzes the placement of ingredients, the state of the fire, and the degree of doneness.
[0999] "Cooking management means" refers to technology that manages the cooking process based on image analysis results and audio data.
[1000] "Communication means" refers to the technology for sending and receiving data between a terminal and a server.
[1001] "User interface means" refers to means for providing visual or audio information to a user and accepting input from the user.
[1002] "Real-time data collection means" refers to technology that acquires images and audio data of barbecue grills in real time.
[1003] "Data analysis means" refers to technology that analyzes collected voice and image data to evaluate the user's intentions and the condition of the grill.
[1004] "Advice generation means" refers to technology that generates specific advice necessary for cooking based on the results of data analysis.
[1005] This invention is a smart AI assistant system that supports barbecue experiences, and includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, a real-time data collection means, a data analysis means, and an advice generation means.
[1006] System Overview
[1007] This system helps users start a barbecue. The user starts the system by issuing the voice command "Start a barbecue." Each step and the hardware and software used are explained below.
[1008] Hardware and software used
[1009] The system uses the following major hardware and software:
[1010] Speech recognition method: Google Cloud Speech-to-Text
[1011] Image analysis methods: OpenCV, TensorFlow
[1012] Cooking control means: Control logic within the system
[1013] Communication method: Wi-Fi or Bluetooth
[1014] User interface means: smartphone application
[1015] Real-time data collection method: built-in camera and microphone of smartphone
[1016] Data analysis method: Google Cloud Natural Language API
[1017] Advice generation means: Advice generation engine within the system
[1018] Program processing
[1019] The program begins when the user issues the voice command "Start the barbecue." The device receives this voice, converts it into text data using Google Cloud Speech-to-Text, and sends the data to the server. The server then analyzes the text data using the Google Cloud Natural Language API, understands the user's intent, and sets up the initial barbecue setup.
[1020] Once the initial setup is complete, the server generates preparation instructions for the user, such as "Prepare the grill," and sends them to the device, which then presents the instructions to the user as a voice message or a screen display.
[1021] Next, the user points their smartphone camera at the barbecue grill. The device uses the camera to capture real-time images and sends the image data to a server. The server then uses image analysis technology (OpenCV, TensorFlow) to analyze the images and evaluate the placement of ingredients, the state of the fire, and the degree of doneness. Based on the evaluation, the server determines the necessary adjustments and generates specific advice (e.g., "Turn up the heat a little" or "Move the meat to the right").
[1022] The generated advice is sent to the device, which then presents it to the user via voice or screen display, allowing the user to make adjustments according to the advice.
[1023] Specific examples
[1024] The process of starting a barbecue
[1025] 1. The user issues a voice command
[1026] User: "Start a barbecue."
[1027] The device uses a voice recognition tool (Google Cloud Speech-to-Text) to convert the voice command into text data and send it to the server.
[1028] 2. The server analyzes the request and sets up the initial settings
[1029] Server: Parses the user request (Google Cloud Natural Language API), performs initialization, and generates preparation instructions.
[1030] Terminal: Prompt the user (audio or screen prompt) to "Prepare the grill."
[1031] Real-time image analysis and advice
[1032] 1. User takes a picture of the grill
[1033] User: Point your device camera at the grill.
[1034] The device sends image data to the server (captured by the smartphone's built-in camera), and the server analyzes the image (OpenCV, TensorFlow).
[1035] 2. The server analyzes the image and generates advice
[1036] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[1037] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[1038] Example of input prompt for generative AI model
[1039] 1. Prompts for user command to generate preparation instructions when starting a barbecue:
[1040] After receiving a voice command from the user to "start barbecue," perform the initial setup and generate and present preparation instructions to the user.
[1041] Example: If the user says "I want to start the barbecue," the device might display or speak instructions such as "Prepare the grill."
[1042] 2. Real-time image analysis prompts advice generation:
[1043] Based on the analysis of real-time images of the barbecue grill, generate and present specific advice to the user on adjusting the heat or repositioning the ingredients.
[1044] For example, if the grill is running low, you might offer advice such as, "Please turn up the heat a little."
[1045] This system allows users to easily enjoy delicious food even if they have no knowledge or experience of barbecue. The system provides users with easy-to-understand instructions and helps them cook.
[1046] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1047] Step 1:
[1048] The user issues a voice command such as "Start the barbecue."
[1049] Input: Voice commands from the user.
[1050] Action: The device receives audio through the microphone.
[1051] Output: The captured audio data is generated.
[1052] Step 2:
[1053] The terminal converts the received voice data into text data using a voice recognition means.
[1054] Input: Audio data.
[1055] How it works: Converts audio data into text using Google Cloud Speech-to-Text.
[1056] Output: Text data is generated and sent to the server.
[1057] Step 3:
[1058] The server analyzes the received text data and understands the user's intent.
[1059] Input: Text data.
[1060] How it works: It uses the Google Cloud Natural Language API to analyze text data and recognize the user's intention to start a barbecue.
[1061] Output: Configuration data reflecting the initial settings is generated.
[1062] Step 4:
[1063] The server generates and sends preparation instructions to the user based on the initial settings.
[1064] Input: Configuration data.
[1065] Operation: Using the cooking management means, a preparation instruction such as "Prepare the grill" is generated and sent to the terminal.
[1066] Output: Preparation instruction data is sent to the terminal.
[1067] Step 5:
[1068] The terminal presents the received preparation instructions to the user.
[1069] Input: Preparation instruction data.
[1070] What it does: Uses user interface means to present instructions to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[1071] Output: The user receives the instructions.
[1072] Step 6:
[1073] A user points their smartphone camera at a barbecue grill.
[1074] Input: Instructions from the server.
[1075] How it works: Point your smartphone camera at the barbecue grill to capture real-time images.
[1076] Output: Real-time image data is generated.
[1077] Step 7:
[1078] The real-time images acquired by the terminal are sent to the server.
[1079] Input: Real-time image data.
[1080] Operation: Image data captured using the camera is sent to the server.
[1081] Output: The image data is transferred to the server.
[1082] Step 8:
[1083] The server analyzes the received image data.
[1084] Input: Image data.
[1085] How it works: Image analysis technology (OpenCV, TensorFlow) is used to analyze the placement of ingredients, the state of the fire, and the degree of doneness.
[1086] Output: Analysis result data is generated.
[1087] Step 9:
[1088] The server generates and sends advice based on the analysis results.
[1089] Input: Analysis result data.
[1090] Operation: Using the advice generation means, specific advice regarding adjusting the heat or changing the position of ingredients is generated and sent to the terminal.
[1091] Output: Advice data is sent to the terminal.
[1092] Step 10:
[1093] The device presents the received advice to the user.
[1094] Input: Advice data.
[1095] What it does: Uses user interface means to present advice to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[1096] Output: User receives advice.
[1097] Step 11:
[1098] The user adjusts the barbecue grill according to the provided advice.
[1099] Enter: Advice.
[1100] Action: Making specific adjustments such as adjusting the heat or changing the position of ingredients.
[1101] Output: Adjusted BBQ grill.
[1102] (Application example 1)
[1103] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] Traditional barbecue and cooking relied on the knowledge and experience of the chef, making it difficult to ensure consistent quality. It was also difficult to monitor the status of the food in real time and make appropriate adjustments, which required a lot of time and effort. This made it difficult for inexperienced chefs to provide satisfying food.
[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1106] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a unit for evaluating and adjusting cooking quality in real time, and a unit for providing advice using a generative AI model. This makes it possible to provide high-quality, consistent cooking without relying on the chef's experience. Furthermore, by being able to grasp the cooking status in real time and make appropriate adjustments, it is possible to significantly reduce the effort required and provide a satisfying cooking experience.
[1107] "Voice recognition means" refers to any device or technology that receives voice commands from a user and converts them into text data.
[1108] "Image analysis means" refers to devices and technologies in general that analyze image data captured in real time and recognize specific objects or conditions.
[1109] "Cooking management means" refers to all devices and technologies that manage cooking based on the analysis results and generate instructions such as adjusting the heat or changing the position of ingredients.
[1110] "Communication means" refers to all devices and technologies for transmitting and receiving data.
[1111] "User interface means" refers generally to devices and techniques that allow a user to interact with a system and obtain information.
[1112] "Means for evaluating and adjusting food quality in real time" refers to all devices and technologies that analyze data acquired in real time, evaluate food quality, and make appropriate adjustments.
[1113] "Means for providing advice using a generative AI model" refers to all devices and technologies that use a generative AI model to provide appropriate advice to users based on the analysis results.
[1114] This invention is a smart AI cooking assistant system for food delivery, which includes voice recognition means, image analysis means, cooking management means, communication means, user interface means, means for evaluating and adjusting cooking quality in real time, and means for providing advice using a generative AI model, enabling cooks to deliver high-quality, consistent food.
[1115] Specifically, users operate the system using a smartphone or head-mounted display. For example, when a user issues a voice command such as "Start a barbecue," the smartphone's microphone captures the voice and converts it into text data using a speech recognition library. The converted text data is then sent to a server for initial setup.
[1116] Next, the user uses their smartphone camera to capture images of the grill in real time. The image data is captured using the OpenCV library and sent to the server. The server analyzes the images and evaluates the placement of ingredients and the state of the flame. Based on the evaluation, the cooking management tool generates advice on adjusting the heat or repositioning ingredients.
[1117] The generated advice is sent as text data to the smartphone via the communication means. The advice is presented to the user via the user interface means as voice or on-screen display. For example, if the advice is "Turn up the heat a little," the user adjusts the heat according to the instruction.
[1118] The system can also use generative AI models to provide more advanced advice, such as providing tailored advice in response to prompts like "What's the best way to adjust the heat on my barbecue?" or "How do I check if my chicken is done?"
[1119] In this way, the system enables consistent, high-quality food to be served in real time, while significantly reducing the chef's workload and providing a satisfying cooking experience.
[1120] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1121] Step 1:
[1122] The user issues the voice command "Start the barbecue." The device captures the voice command using the built-in microphone and converts the voice into text data using a voice recognition library (e.g., SpeechRecognition library). This text data is sent to the server. The input is the voice command, and the output is text data. Data processing involves analyzing the voice waveform data and converting it into text data.
[1123] Step 2:
[1124] The server receives the text data and recognizes that the user is about to start a barbecue. The server generates a preparation instruction, "Prepare the grill," as the initial setting and sends this instruction to the terminal. The input is the text data, and the output is the initial setting and the preparation instruction. As a data operation, it performs text analysis and sets the initial setting flag.
[1125] Step 3:
[1126] The device receives the preparation instruction from the server and presents it to the user by voice or on-screen display. The input is the preparation instruction from the server, and the output is the instruction presented to the user. Specifically, the instruction is played back by voice using the device's speaker.
[1127] Step 4:
[1128] A user points their smartphone camera at a barbecue grill. The device uses the camera to capture real-time images of the grill and sends the image data to a server. The input is the camera image and the output is image data. Image capture and compression are performed as data processing.
[1129] Step 5:
[1130] The server analyzes the image data it receives and evaluates the placement of ingredients and the state of the fire. This is done using an image analysis library (e.g., OpenCV). The input is the image data, and the output is the analysis results. Data calculations involve extracting basic image features and obtaining information on the heat level and the position of ingredients.
[1131] Step 6:
[1132] The server uses the cooking management means based on the analysis results to generate advice on adjusting the heat or changing the position of ingredients. This advice is sent to the terminal as text data via communication means. The input is the analysis results, and the output is advice. As a data calculation, appropriate cooking instructions are automatically generated based on the analysis results.
[1133] Step 7:
[1134] The device receives advice from the server and presents it to the user by voice or on-screen display. The input is advice from the server, and the output is instructions presented to the user. Specifically, the advice is played back by voice using the device's speaker, or displayed as text on the screen.
[1135] Step 8:
[1136] The user makes adjustments based on the advice from the server. This step acts as a feedback loop, returning to step 4 and beyond to make the next adjustment if necessary. The input is the advice, and the output is the adjustment action taken by the user. Specific actions taken by the user include adjusting the heat or changing the position of ingredients.
[1137] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1138] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. Below, we will explain the program and processing of this system with concrete examples.
[1139] System Overview
[1140] The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, and an emotion engine. The voice recognition unit receives a user's voice command and converts it into text data. The image analysis unit acquires and analyzes real-time images of the barbecue grill. The cooking management unit generates advice based on the results of the image analysis, which is sent to the user terminal via the communication unit and presented to the user via the user interface unit. The emotion engine also analyzes the user's emotions and adjusts the cooking advice based on the analysis.
[1141] Program processing
[1142] 1. The user starts a barbecue
[1143] The user issues the voice command "Start the barbecue."
[1144] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[1145] 2. Analysis of audio data
[1146] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[1147] The server performs the initial setup and creates the necessary setup instructions.
[1148] 3. Preparatory Instructions
[1149] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[1150] The device will provide instructions to the user via voice or on-screen display.
[1151] The user follows the instructions to prepare the grill.
[1152] 4. Real-time image acquisition
[1153] The user points the device's camera at the grill and places the ingredients on it.
[1154] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[1155] 5. Analysis of Image Data
[1156] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[1157] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[1158] 6. Emotion Analysis Using an Emotion Engine
[1159] The device captures the user's facial expressions and tone of voice and sends them to the emotion engine.
[1160] The server uses an emotion engine to analyze the user's emotions and adjusts cooking advice based on this.
[1161] 7. Generating and Providing Advice
[1162] The server adjusts the advice based on the emotion data and sends the generated advice (e.g., "Turn up the heat a little," "Move the chicken to the right") to the terminal.
[1163] The advice received by the device is presented to the user by voice or on-screen display.
[1164] The user adjusts the heat and position of the ingredients according to the advice.
[1165] 8. Recipe suggestions
[1166] The user asks the device, "What should I make next?"
[1167] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[1168] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[1169] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[1170] The device will present recipe suggestions to the user via voice or on-screen display.
[1171] 9. Emotional Cooking Guidance
[1172] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[1173] The server generates cooking guidance based on the user's emotions (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device.
[1174] The device will notify the user of emotionally sensitive notifications via voice.
[1175] The user removes the cooked chicken and enjoys the meal.
[1176] The system of the present invention allows users to easily enjoy delicious barbecues, even without detailed knowledge or experience. Furthermore, by combining it with an emotion engine, a more personalized and comfortable barbecue experience can be provided. By significantly reducing the effort required for cooking and providing emotionally appropriate encouragement and advice, users can spend more time concentrating on conversations with family and friends and on their meals.
[1177] The processing flow will be explained below.
[1178] Step 1:
[1179] The user issues the voice command "Start the barbecue."
[1180] The device uses a microphone to capture the user's voice.
[1181] Step 2:
[1182] The terminal converts the voice into text data using a voice recognition means.
[1183] The terminal transmits the converted text data to the server.
[1184] Step 3:
[1185] The server receives the text data, analyzes it, and recognizes it as a request to "start barbecue."
[1186] The server performs the initialization and generates the necessary provisioning instructions.
[1187] Step 4:
[1188] The server sends the generated preparation instruction to the terminal.
[1189] The terminal will give the user instructions to prepare the grill by voice or on the screen, such as "Prepare the grill."
[1190] Step 5:
[1191] The user follows the instructions to prepare the grill.
[1192] The user points the device's camera at the grill and captures an image.
[1193] Step 6:
[1194] The device captures an image of the grill and sends the image data to a server.
[1195] The server receives the image data and uses image analysis means to evaluate the placement of ingredients, the state of the fire, and the degree of doneness.
[1196] Step 7:
[1197] Based on the results of image analysis, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[1198] The server transmits the generated advice to the terminal.
[1199] Step 8:
[1200] The device displays the received advice to the user (e.g., "Turn up the heat a little").
[1201] The user adjusts the heat and position of the ingredients according to the advice.
[1202] Step 9:
[1203] The device captures the user's facial expressions and voice tone and transmits the emotional data to the server.
[1204] The server analyzes the user's emotions using an emotion engine.
[1205] Step 10:
[1206] The server adjusts cooking advice based on the results of emotion analysis (e.g., if the user seems anxious, give advice in a gentler tone).
[1207] The server sends the adjusted advice to the terminal.
[1208] Step 11:
[1209] The terminal presents the adjusted advice to the user.
[1210] The user proceeds with cooking according to the advice.
[1211] Step 12:
[1212] The user asks the device, "What should I make next?"
[1213] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[1214] Step 13:
[1215] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[1216] The server generates recipe suggestions and sends them to the terminal.
[1217] Step 14:
[1218] The device will present recipe suggestions (e.g., "We recommend this chicken marinade") to the user via voice or on-screen display.
[1219] Users can try new dishes by following suggested recipes.
[1220] Step 15:
[1221] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[1222] The server generates a notification saying "The chicken is done" and sends it to the device.
[1223] Step 16:
[1224] The terminal notifies the user by voice when the food is done.
[1225] The user removes the cooked chicken and enjoys the meal.
[1226] Example 2
[1227] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1228] Enjoying a barbecue often requires skilled techniques and a lot of prior knowledge. However, for users who are trying barbecue for the first time or who are unfamiliar with cooking, these requirements are a high hurdle and make it difficult to fully enjoy a barbecue. In addition, there is a problem that the user experience is inconsistent because the cooking status is not grasped in real time or appropriate advice is not provided according to the user's emotions.
[1229] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, an emotion analysis means, and a user interface means. This allows even beginners to easily receive detailed cooking advice on barbecues, and enables personalized support tailored to the user's emotions. Specifically, the voice recognition means recognizes the user's voice commands and converts them into text data. The image analysis means captures real-time images of the grill and analyzes the state of the ingredients and the fire. Furthermore, the emotion analysis means analyzes the user's emotions from their facial expressions and voice tone. The cooking management means integrates these analysis results to generate appropriate cooking advice, which is then provided to the user through the user interface means.
[1230] "Speech recognition means" refers to means having the function of receiving voice commands from a user and converting them into text data.
[1231] The "image analysis means" is a means having the function of analyzing images captured in real time and evaluating the state of the ingredients and the fire.
[1232] The "cooking management means" is a means having a function of generating cooking advice based on the results of the image analysis means and the emotion analysis means.
[1233] "Communication means" refers to a means that utilizes protocols and technologies for sending and receiving data.
[1234] A "user interface means" is a means for presenting information to a user and receiving input from a user.
[1235] The "emotion analysis means" is a means having the function of capturing the user's facial expressions and tone of voice and analyzing their emotions.
[1236] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. This system includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, and an emotion analysis means.
[1237] Hardware and software used
[1238] The system uses the following major hardware and software:
[1239] 1. Speech recognition means: Typically using a cloud-based speech recognition service, e.g., a speech recognition cloud API.
[1240] 2. Image analysis method: Uses the OpenCV library and deep learning models (e.g., TensorFlow).
[1241] 3. Cooking control method: A custom algorithm implemented in Python is used.
[1242] 4. Communication method: Network communication is performed using the HTTP / HTTPS protocol.
[1243] 5. User interface: Implemented through a smartphone app (iOS / Android) or web browser.
[1244] 6. Emotion analysis means: Use an API that performs facial expression analysis and voice tone analysis (e.g., emotion analysis cloud API).
[1245] Processing of this system
[1246] 1. When a user wants to start a barbecue, they say the voice command "Start the barbecue." This voice command is converted into text data by the device using a voice recognition cloud API and sent to the server.
[1247] 2. The server analyzes the received text data and recognizes that the user's request is to "start a barbecue." It uses a natural language processing (NLP) library to process this. The server then performs its initial setup and creates the necessary preparation instructions.
[1248] 3. The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the device. The device then presents the instructions to the user via voice or visual display, and the user prepares the grill accordingly.
[1249] 4. The user points the device's camera at the grill and places the food on it. The device uses the camera to capture real-time images of the grill and sends the image data to the server.
[1250] 5. The server analyzes the received image data using OpenCV and deep learning models. The analysis results include the placement of the ingredients, the state of the fire, and the degree of doneness. Based on this information, the server generates specific advice on adjusting the heat or repositioning the ingredients.
[1251] 6. The device captures the user's facial expressions and voice tone and sends them to the emotion analysis cloud API. The server uses an emotion engine to analyze the user's emotions and adjust the cooking advice accordingly.
[1252] 7. The server generates advice based on the emotion data and sends it to the device. The device then presents the advice to the user via voice or on-screen display. The user adjusts the heat and position of the ingredients according to the advice.
[1253] 8. The user asks the device, "What should I make next?" The device uses a voice recognition cloud API to convert the question into text data and sends it to the server. The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The server generates a recipe suggestion (e.g., "I recommend marinated chicken") and sends it to the device. The device then presents the recipe suggestion to the user by voice or on a screen.
[1254] 9. The server periodically performs image and emotion analysis to determine whether the food is cooked. The server generates cooking guidance based on the user's emotion (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device. The device then notifies the user with voice notifications that take emotion into consideration. The user follows the instructions to remove the cooked chicken and enjoy their meal.
[1255] Specific examples
[1256] Example 1: Your first barbecue
[1257] 1. The user says the voice command "Start the barbecue."
[1258] 2. The device converts the voice into text and sends it to the server.
[1259] 3. The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal.
[1260] 4. The user prepares the grill and points the device camera at the grill.
[1261] 5. The device sends the captured image to the server, which then analyzes the image.
[1262] 6. The server generates advice on adjusting the heat and sends it to the device.
[1263] 7. The device provides advice to the user, who then adjusts the heat.
[1264] Prompt Sentence Examples
[1265] Prompt 1: "Start the BBQ Assistant to provide a set of guidance for first-time barbecuers."
[1266] Prompt 2: "When a user requests a new recipe, recommend a recipe that fits their mood that day."
[1267] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1268] Step 1:
[1269] The user issues the voice command "Start the barbecue."
[1270] The device converts the speech into text data using a speech recognition cloud API. The input is the user's speech, and the output is the converted text data. The device then sends this text data to the server.
[1271] Step 2:
[1272] The server analyzes the received text data. The input is text data sent from the device, and the server uses a natural language processing library to recognize that it is a request to "start a barbecue." The output is the result of recognizing the request as "start a barbecue." The server performs initial configuration and creates the necessary preparation instructions.
[1273] Step 3:
[1274] The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal. The input is the initial setting result, and the output is a text message of the preparation instruction. The terminal receives this instruction and presents it to the user by voice or on-screen display. The user follows the instruction and performs the action of preparing the grill.
[1275] Step 4:
[1276] The user points the device's camera at the grill and places ingredients on it. The input is the camera image. The device uses the camera to capture real-time images of the grill and sends the image data to the server. The output is the captured real-time image data.
[1277] Step 5:
[1278] The server analyzes the received image data. The input is image data sent from the device. The server uses OpenCV and deep learning models (such as TensorFlow) to evaluate the placement of ingredients, the state of the fire, and the degree of doneness. The output is the analysis results, which are summarized as specific advice on adjusting the heat or changing the position of ingredients.
[1279] Step 6:
[1280] The device captures the user's facial expressions and voice tone. The input is the user's facial expression and voice data. The device sends this to the emotion analysis cloud API. The output is the emotion analysis result, which is sent to the server.
[1281] Step 7:
[1282] The server receives the emotion data analysis results and generates adjustment advice. The input is emotion data and image analysis results. The server adjusts the cooking advice based on this data and generates specific advice (e.g., "Turn up the heat a little," "Move the chicken to the right"). The output is the adjusted advice. The server sends this advice to the device.
[1283] Step 8:
[1284] The device provides advice to the user by voice or on-screen display. The input is advice data from the server, and the output is the information presented to the user. The user follows the advice and adjusts the heat and position of the ingredients.
[1285] Step 9:
[1286] The user asks the device, "What should I make next?" The input is the user's voice command. The device again uses the speech recognition cloud API to convert the question into text data and sends that data to the server. The output is the converted text data.
[1287] Step 10:
[1288] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The input is the user's past cooking history, current ingredient information, and emotional data, and the output is the optimal recipe suggestion. The server generates a recipe suggestion (e.g., "We recommend marinated chicken") and sends it to the device.
[1289] Step 11:
[1290] The terminal presents recipe suggestions to the user by voice or on-screen display. The input is the recipe suggestion data from the server, and the output is the suggestion to the user.
[1291] Step 12:
[1292] The server periodically performs image analysis and emotion analysis. The input is real-time images and user emotion data. The server determines whether the ingredients are cooked and generates cooking guidance. The output is cooking guidance (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?"). The server then sends the generated cooking guidance to the device.
[1293] Step 13:
[1294] The device provides emotionally sensitive notifications to the user. The input is cooking guidance data from the server, and the output is a notification to the user. The user follows the instructions, removes the cooked chicken, and enjoys the meal.
[1295] Through the above processing steps, the smart AI assistant system of the present invention can help users enjoy a more enjoyable and easier barbecue. By providing personalized advice based on the cooking progress and the user's emotions, even beginners can enjoy a delicious barbecue.
[1296] (Application example 2)
[1297] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1298] Conventional food delivery systems lack personalized menu suggestions and cooking advice based on the user's emotional state, which hinders the improvement of user experience. Furthermore, when users are cooking, they often do not receive appropriate advice in real time, which leads to a decrease in satisfaction.
[1299] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, an emotion analysis means, and a database means. This makes it possible to propose optimal menus and provide cooking advice according to the user's emotional state, which is expected to improve the user experience.
[1300] "Voice recognition" is a technology that receives a user's voice commands and converts them into text data.
[1301] "Image analysis" is a technology that analyzes the state of ingredients and fire based on image data captured in real time.
[1302] "Cooking management" is a technology that generates cooking advice based on the results of image analysis and presents it to the user.
[1303] "Communication means" refers to the technology used to send and receive voice data, image data, and analysis results between the server and the user terminal.
[1304] A "user interface" is a means for receiving voice commands and image data input from a user, and for providing generated advice and menu suggestions to the user.
[1305] "Emotion analysis" is a technology that analyzes a user's facial expressions and tone of voice to understand their emotional state.
[1306] The "database" is a technology that stores a user's past order history and emotional data, and generates suggestions based on the analysis results as needed.
[1307] The present invention relates to an emotion-responsive food delivery support application, specifically a system that provides optimal menu suggestions and cooking advice based on a user's voice commands and emotional state. The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, an emotion analysis unit, and a database unit.
[1308] System Overview
[1309] The solution begins when a user launches a food delivery application on their smartphone. First, the user issues a voice command such as "Tell me what menu items you want." The smartphone uses a microphone to capture the voice and converts it into text using a speech recognition library called Google Cloud Speech-to-Text. This text is then sent to a server, which analyzes it and understands the user's request.
[1310] Next, an emotion analysis tool analyzes the user's facial expressions and tone of voice to generate emotion data. This uses the Microsoft Azure Emotion API. The emotion data is used to understand the user's emotional state, and appropriate menu suggestions and cooking advice are provided based on that emotional state. Past order history and emotion data are stored in a database (e.g., MySQL), and optimal suggestions are generated based on this data.
[1311] For example, if the user is in a state of mind where they want to relax, the system will suggest "relaxing herbal tea." If the user is cooking at home, the system will perform real-time image analysis and provide appropriate cooking advice. Image analysis uses a library called OpenCV.
[1312] The specific hardware and software used
[1313] Hardware: Smartphone (microphone, camera, speaker)
[1314] software:
[1315] Speech recognition library: Google Cloud Speech-to-Text
[1316] Sentiment analysis library: Microsoft Azure Emotion API
[1317] Image analysis library: OpenCV
[1318] Database: MySQL
[1319] Prompt Sentence Examples
[1320] Specific examples of prompts are as follows:
[1321] User's order history: pizza, pasta, salad
[1322] Emotional state: Relaxed
[1323] User Question: What's your recommended menu item?
[1324] Response: I recommend herbal tea. It's perfect for relaxing.
[1325] In this way, the system of the present invention allows users to enjoy a personalized food delivery experience based on their voice commands and emotional state.
[1326] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1327] Step 1:
[1328] A user launches a food delivery application on their smartphone and says, "Tell me what menu items you recommend." Here, the input is the user's voice command, and the smartphone's microphone is used to capture the voice data. This voice data is converted into text data using the Google Cloud Speech-to-Text library. The converted text data is the output.
[1329] Step 2:
[1330] The speech-recognized text data is sent to the server. Here, the input is the converted text data, and the server receives the data. The server performs text analysis to understand the user's request. As a result of the analysis, data called "recommended menu suggestions" is output.
[1331] Step 3:
[1332] The emotion analyzer captures the user's facial expressions and voice tone and generates emotion data using the Microsoft Azure Emotion API, where the input is the user's real-time facial expression data and voice tone, and the API analyzes the user's emotional state (e.g., relaxed, happy, etc.), and the analyzed emotion data is the output.
[1333] Step 4:
[1334] The server retrieves the generated emotion data and past order history from a database. Here, the input is emotion data and order history data. This data is retrieved through a database query, and the server selects the optimal menu based on this data. The selected menu suggestion is the output.
[1335] Step 5:
[1336] The server generates a selected menu suggestion (e.g., "I recommend herbal tea. It's perfect for relaxing.") and sends it to the smartphone via communication means. Here, the input is the data of the selected menu suggestion, which is sent to the smartphone. The output is the smartphone presenting the suggested menu to the user by voice or on-screen display.
[1337] Step 6:
[1338] When a user cooks at home, they use their smartphone camera to capture ingredients and the cooking process. The input is real-time image data, which is analyzed using an image analysis library (OpenCV). The analysis results are then output as cooking advice (e.g., "Adjust the heat.").
[1339] Step 7:
[1340] The server sends the generated cooking advice to the smartphone and presents it to the user. Here, the input is the analyzed cooking advice data, and the data is sent to the smartphone. The output of the smartphone is to provide advice to the user by voice or on-screen display.
[1341] Step 8:
[1342] The user then asks a follow-up question (e.g., "What's the next recommended menu item?") by voice. The smartphone's microphone is again used to capture the voice data, which is then converted into text data using the Google Cloud Speech-to-Text library. This voice data is the input, and the converted text data is the output.
[1343] Step 9:
[1344] The server analyzes the converted text data and generates the next menu suggestion based on the information from the database and the sentiment data. Here, the input is the text data based on the follow-up question and related data, and the output is the next menu suggestion as the analysis result.
[1345] Step 10:
[1346] The server sends the next menu proposal to the smartphone and presents it to the user. Here, the input is the generated next menu proposal data, which is sent to the smartphone. The output of the smartphone is to present it to the user by voice or on-screen display.
[1347] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1348] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1349] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1350] [Fourth embodiment]
[1351] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1352] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1353] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1354] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1355] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1356] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1357] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1358] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1359] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1360] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1361] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1362] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1363] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1364] This invention is a smart AI assistant system that supports the barbecue experience, and includes voice recognition, image analysis, cooking management, communication, and user interface. The program and processing of this system are explained below in natural language with concrete examples.
[1365] System Overview
[1366] When a user starts a barbecue, this system uses a voice recognition means to receive the user's voice commands and convert them into text data. Next, an image analysis means captures and analyzes real-time images of the barbecue grill. Based on the analysis results, a cooking management means generates advice such as adjusting the heat or changing the position of ingredients. The generated advice is sent to the user terminal via a communication means and presented to the user via a user interface means.
[1367] Program processing
[1368] 1. The user starts a barbecue
[1369] The user issues the voice command "Start the barbecue."
[1370] The terminal receives the voice, and the voice recognition means converts it into text data, which is then sent to the server.
[1371] 2. Analysis of audio data
[1372] The server analyzes the text data and determines that the user is about to start a barbecue, and then performs the initial setup.
[1373] 3. Preparatory Instructions
[1374] The server generates preparation instructions for the user (e.g., "Prepare the grill") based on the initial settings and sends them to the terminal.
[1375] The device will provide instructions to the user via voice or visual display.
[1376] 4. Real-time image acquisition
[1377] The user points the device's camera at the barbecue grill.
[1378] The device uses a camera to capture an image of the grill and transmits the image data to a server.
[1379] 5. Analysis of Image Data
[1380] The server analyzes the image data and evaluates the placement of ingredients, the state of the fire, and the degree of doneness.
[1381] The server determines the necessary adjustments based on the analysis results.
[1382] 6. Generating and Providing Advice
[1383] The server generates advice on adjusting the heat and changing the position of ingredients and sends it to the terminal.
[1384] The device presents the generated advice to the user by voice or on-screen display.
[1385] The user makes adjustments according to the advice.
[1386] Specific examples
[1387] The process of starting a barbecue
[1388] 1. The user issues a voice command
[1389] User: "Start a barbecue."
[1390] The terminal converts the voice command into text data using a voice recognition means and transmits it to the server.
[1391] 2. The server analyzes the request and sets up the initial settings
[1392] Server: Parses the user's request, performs initialization, and generates preparation instructions.
[1393] Terminal: Prompts the user with the instruction "Prepare the grill."
[1394] Real-time image analysis and advice
[1395] 1. User takes a picture of the grill
[1396] User: Point your device camera at the grill.
[1397] The device sends image data to a server, which analyzes the image.
[1398] 2. The server analyzes the image and generates advice
[1399] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[1400] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[1401] The system of the present invention allows users to easily enjoy delicious barbecues without detailed knowledge or experience. It also significantly reduces the effort required for cooking, allowing users to spend more time concentrating on conversations and meals with family and friends.
[1402] The processing flow will be explained below.
[1403] Step 1:
[1404] The user issues the voice command "Start the barbecue."
[1405] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[1406] Step 2:
[1407] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[1408] The server performs the initial setup and creates the necessary setup instructions.
[1409] Step 3:
[1410] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[1411] The terminal presents instructions to the user by voice or on-screen display.
[1412] The user follows the instructions to prepare the grill.
[1413] Step 4:
[1414] The user points the device's camera at the grill and places the ingredients on it.
[1415] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[1416] Step 5:
[1417] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[1418] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[1419] Step 6:
[1420] The server generates advice (e.g., "Turn up the heat a little," "Move the chicken to the right") and sends it to the device.
[1421] The advice received by the terminal is presented to the user by voice or screen display.
[1422] The user adjusts the heat and position of the ingredients according to the advice.
[1423] Step 7:
[1424] The user asks the device, "What should I make next?"
[1425] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[1426] Step 8:
[1427] The server selects the optimal recipe based on the user's past cooking history and current ingredient information.
[1428] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[1429] The terminal presents recipe suggestions to the user by voice or on-screen display.
[1430] Step 9:
[1431] The server periodically analyzes the images to determine whether the food is cooked.
[1432] The server generates a notification saying "The chicken is done" and sends it to the device.
[1433] The terminal notifies the user by voice when the food is done.
[1434] The user removes the cooked chicken and enjoys the meal.
[1435] Example 1
[1436] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1437] In traditional barbecue cooking, beginners and those with little experience often have difficulty controlling the heat and arranging ingredients properly. Additionally, the lack of real-time monitoring and proper advice can lead to inconsistent cooking quality. This can result in users spending a lot of time and effort on cooking, preventing them from fully enjoying their barbecue experience.
[1438] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1439] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a real-time data collection unit, a data analysis unit, and an advice generation unit. This allows users to start a barbecue using voice commands and manage the cooking appropriately through real-time image data analysis. Furthermore, specific advice based on the analysis results can be received by voice or displayed on the screen, allowing even beginners to easily enjoy a high-quality barbecue experience.
[1440] "Voice recognition means" refers to technology that receives voice commands from a user and converts them into text data.
[1441] "Image analysis means" refers to technology that acquires real-time images of the barbecue grill and analyzes the placement of ingredients, the state of the fire, and the degree of doneness.
[1442] "Cooking management means" refers to technology that manages the cooking process based on image analysis results and audio data.
[1443] "Communication means" refers to the technology for sending and receiving data between a terminal and a server.
[1444] "User interface means" refers to means for providing visual or audio information to a user and accepting input from the user.
[1445] "Real-time data collection means" refers to technology that acquires images and audio data of barbecue grills in real time.
[1446] "Data analysis means" refers to technology that analyzes collected voice and image data to evaluate the user's intentions and the condition of the grill.
[1447] "Advice generation means" refers to technology that generates specific advice necessary for cooking based on the results of data analysis.
[1448] This invention is a smart AI assistant system that supports barbecue experiences, and includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, a real-time data collection means, a data analysis means, and an advice generation means.
[1449] System Overview
[1450] This system helps users start a barbecue. The user starts the system by issuing the voice command "Start a barbecue." Each step and the hardware and software used are explained below.
[1451] Hardware and software used
[1452] The system uses the following major hardware and software:
[1453] Speech recognition method: Google Cloud Speech-to-Text
[1454] Image analysis methods: OpenCV, TensorFlow
[1455] Cooking control means: Control logic within the system
[1456] Communication method: Wi-Fi or Bluetooth
[1457] User interface means: smartphone application
[1458] Real-time data collection method: built-in camera and microphone of smartphone
[1459] Data analysis method: Google Cloud Natural Language API
[1460] Advice generation means: Advice generation engine within the system
[1461] Program processing
[1462] The program begins when the user issues the voice command "Start the barbecue." The device receives this voice, converts it into text data using Google Cloud Speech-to-Text, and sends the data to the server. The server then analyzes the text data using the Google Cloud Natural Language API, understands the user's intent, and sets up the initial barbecue setup.
[1463] Once the initial setup is complete, the server generates preparation instructions for the user, such as "Prepare the grill," and sends them to the device, which then presents the instructions to the user as a voice message or a screen display.
[1464] Next, the user points their smartphone camera at the barbecue grill. The device uses the camera to capture real-time images and sends the image data to a server. The server then uses image analysis technology (OpenCV, TensorFlow) to analyze the images and evaluate the placement of ingredients, the state of the fire, and the degree of doneness. Based on the evaluation, the server determines the necessary adjustments and generates specific advice (e.g., "Turn up the heat a little" or "Move the meat to the right").
[1465] The generated advice is sent to the device, which then presents it to the user via voice or screen display, allowing the user to make adjustments according to the advice.
[1466] Specific examples
[1467] The process of starting a barbecue
[1468] 1. The user issues a voice command
[1469] User: "Start a barbecue."
[1470] The device uses a voice recognition tool (Google Cloud Speech-to-Text) to convert the voice command into text data and send it to the server.
[1471] 2. The server analyzes the request and sets up the initial settings
[1472] Server: Parses the user request (Google Cloud Natural Language API), performs initialization, and generates preparation instructions.
[1473] Terminal: Prompt the user (audio or screen prompt) to "Prepare the grill."
[1474] Real-time image analysis and advice
[1475] 1. User takes a picture of the grill
[1476] User: Point your device camera at the grill.
[1477] The device sends image data to the server (captured by the smartphone's built-in camera), and the server analyzes the image (OpenCV, TensorFlow).
[1478] 2. The server analyzes the image and generates advice
[1479] Server: Based on the results of image analysis, it generates advice on adjusting the heat and changing the position of ingredients.
[1480] Device: Provides specific advice to the user via voice or on-screen display, such as "Turn up the heat a little."
[1481] Example of input prompt for generative AI model
[1482] 1. Prompts for user command to generate preparation instructions when starting a barbecue:
[1483] After receiving a voice command from the user to "start barbecue," perform the initial setup and generate and present preparation instructions to the user.
[1484] Example: If the user says "I want to start the barbecue," the device might display or speak instructions such as "Prepare the grill."
[1485] 2. Real-time image analysis prompts advice generation:
[1486] Based on the analysis of real-time images of the barbecue grill, generate and present specific advice to the user on adjusting the heat or repositioning the ingredients.
[1487] For example, if the grill is running low, you might offer advice such as, "Please turn up the heat a little."
[1488] This system allows users to easily enjoy delicious food even if they have no knowledge or experience of barbecue. The system provides users with easy-to-understand instructions and helps them cook.
[1489] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1490] Step 1:
[1491] The user issues a voice command such as "Start the barbecue."
[1492] Input: Voice commands from the user.
[1493] Action: The device receives audio through the microphone.
[1494] Output: The captured audio data is generated.
[1495] Step 2:
[1496] The terminal converts the received voice data into text data using a voice recognition means.
[1497] Input: Audio data.
[1498] How it works: Converts audio data into text using Google Cloud Speech-to-Text.
[1499] Output: Text data is generated and sent to the server.
[1500] Step 3:
[1501] The server analyzes the received text data and understands the user's intent.
[1502] Input: Text data.
[1503] How it works: It uses the Google Cloud Natural Language API to analyze text data and recognize the user's intention to start a barbecue.
[1504] Output: Configuration data reflecting the initial settings is generated.
[1505] Step 4:
[1506] The server generates and sends preparation instructions to the user based on the initial settings.
[1507] Input: Configuration data.
[1508] Operation: Using the cooking management means, a preparation instruction such as "Prepare the grill" is generated and sent to the terminal.
[1509] Output: Preparation instruction data is sent to the terminal.
[1510] Step 5:
[1511] The terminal presents the received preparation instructions to the user.
[1512] Input: Preparation instruction data.
[1513] What it does: Uses user interface means to present instructions to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[1514] Output: The user receives the instructions.
[1515] Step 6:
[1516] A user points their smartphone camera at a barbecue grill.
[1517] Input: Instructions from the server.
[1518] How it works: Point your smartphone camera at the barbecue grill to capture real-time images.
[1519] Output: Real-time image data is generated.
[1520] Step 7:
[1521] The real-time images acquired by the terminal are sent to the server.
[1522] Input: Real-time image data.
[1523] Operation: Image data captured using the camera is sent to the server.
[1524] Output: The image data is transferred to the server.
[1525] Step 8:
[1526] The server analyzes the received image data.
[1527] Input: Image data.
[1528] How it works: Image analysis technology (OpenCV, TensorFlow) is used to analyze the placement of ingredients, the state of the fire, and the degree of doneness.
[1529] Output: Analysis result data is generated.
[1530] Step 9:
[1531] The server generates and sends advice based on the analysis results.
[1532] Input: Analysis result data.
[1533] Operation: Using the advice generation means, specific advice regarding adjusting the heat or changing the position of ingredients is generated and sent to the terminal.
[1534] Output: Advice data is sent to the terminal.
[1535] Step 10:
[1536] The device presents the received advice to the user.
[1537] Input: Advice data.
[1538] What it does: Uses user interface means to present advice to the user via voice messages (Text-to-Speech engine) or on-screen displays.
[1539] Output: User receives advice.
[1540] Step 11:
[1541] The user adjusts the barbecue grill according to the provided advice.
[1542] Enter: Advice.
[1543] Action: Making specific adjustments such as adjusting the heat or changing the position of ingredients.
[1544] Output: Adjusted BBQ grill.
[1545] (Application example 1)
[1546] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1547] Traditional barbecue and cooking relied on the knowledge and experience of the chef, making it difficult to ensure consistent quality. It was also difficult to monitor the status of the food in real time and make appropriate adjustments, which required a lot of time and effort. This made it difficult for inexperienced chefs to provide satisfying food.
[1548] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1549] In this invention, the server includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, a unit for evaluating and adjusting cooking quality in real time, and a unit for providing advice using a generative AI model. This makes it possible to provide high-quality, consistent cooking without relying on the chef's experience. Furthermore, by being able to grasp the cooking status in real time and make appropriate adjustments, it is possible to significantly reduce the effort required and provide a satisfying cooking experience.
[1550] "Voice recognition means" refers to any device or technology that receives voice commands from a user and converts them into text data.
[1551] "Image analysis means" refers to devices and technologies in general that analyze image data captured in real time and recognize specific objects or conditions.
[1552] "Cooking management means" refers to all devices and technologies that manage cooking based on the analysis results and generate instructions such as adjusting the heat or changing the position of ingredients.
[1553] "Communication means" refers to all devices and technologies for transmitting and receiving data.
[1554] "User interface means" refers generally to devices and techniques that allow a user to interact with a system and obtain information.
[1555] "Means for evaluating and adjusting food quality in real time" refers to all devices and technologies that analyze data acquired in real time, evaluate food quality, and make appropriate adjustments.
[1556] "Means for providing advice using a generative AI model" refers to all devices and technologies that use a generative AI model to provide appropriate advice to users based on the analysis results.
[1557] This invention is a smart AI cooking assistant system for food delivery, which includes voice recognition means, image analysis means, cooking management means, communication means, user interface means, means for evaluating and adjusting cooking quality in real time, and means for providing advice using a generative AI model, enabling cooks to deliver high-quality, consistent food.
[1558] Specifically, users operate the system using a smartphone or head-mounted display. For example, when a user issues a voice command such as "Start a barbecue," the smartphone's microphone captures the voice and converts it into text data using a speech recognition library. The converted text data is then sent to a server for initial setup.
[1559] Next, the user uses their smartphone camera to capture images of the grill in real time. The image data is captured using the OpenCV library and sent to the server. The server analyzes the images and evaluates the placement of ingredients and the state of the flame. Based on the evaluation, the cooking management tool generates advice on adjusting the heat or repositioning ingredients.
[1560] The generated advice is sent as text data to the smartphone via the communication means. The advice is presented to the user via the user interface means as voice or on-screen display. For example, if the advice is "Turn up the heat a little," the user adjusts the heat according to the instruction.
[1561] The system can also use generative AI models to provide more advanced advice, such as providing tailored advice in response to prompts like "What's the best way to adjust the heat on my barbecue?" or "How do I check if my chicken is done?"
[1562] In this way, the system enables consistent, high-quality food to be served in real time, while significantly reducing the chef's workload and providing a satisfying cooking experience.
[1563] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1564] Step 1:
[1565] The user issues the voice command "Start the barbecue." The device captures the voice command using the built-in microphone and converts the voice into text data using a voice recognition library (e.g., SpeechRecognition library). This text data is sent to the server. The input is the voice command, and the output is text data. Data processing involves analyzing the voice waveform data and converting it into text data.
[1566] Step 2:
[1567] The server receives the text data and recognizes that the user is about to start a barbecue. The server generates a preparation instruction, "Prepare the grill," as the initial setting and sends this instruction to the terminal. The input is the text data, and the output is the initial setting and the preparation instruction. As a data operation, it performs text analysis and sets the initial setting flag.
[1568] Step 3:
[1569] The device receives the preparation instruction from the server and presents it to the user by voice or on-screen display. The input is the preparation instruction from the server, and the output is the instruction presented to the user. Specifically, the instruction is played back by voice using the device's speaker.
[1570] Step 4:
[1571] A user points their smartphone camera at a barbecue grill. The device uses the camera to capture real-time images of the grill and sends the image data to a server. The input is the camera image and the output is image data. Image capture and compression are performed as data processing.
[1572] Step 5:
[1573] The server analyzes the image data it receives and evaluates the placement of ingredients and the state of the fire. This is done using an image analysis library (e.g., OpenCV). The input is the image data, and the output is the analysis results. Data calculations involve extracting basic image features and obtaining information on the heat level and the position of ingredients.
[1574] Step 6:
[1575] The server uses the cooking management means based on the analysis results to generate advice on adjusting the heat or changing the position of ingredients. This advice is sent to the terminal as text data via communication means. The input is the analysis results, and the output is advice. As a data calculation, appropriate cooking instructions are automatically generated based on the analysis results.
[1576] Step 7:
[1577] The device receives advice from the server and presents it to the user by voice or on-screen display. The input is advice from the server, and the output is instructions presented to the user. Specifically, the advice is played back by voice using the device's speaker, or displayed as text on the screen.
[1578] Step 8:
[1579] The user makes adjustments based on the advice from the server. This step acts as a feedback loop, returning to step 4 and beyond to make the next adjustment if necessary. The input is the advice, and the output is the adjustment action taken by the user. Specific actions taken by the user include adjusting the heat or changing the position of ingredients.
[1580] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1581] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. Below, we will explain the program and processing of this system with concrete examples.
[1582] System Overview
[1583] The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, and an emotion engine. The voice recognition unit receives a user's voice command and converts it into text data. The image analysis unit acquires and analyzes real-time images of the barbecue grill. The cooking management unit generates advice based on the results of the image analysis, which is sent to the user terminal via the communication unit and presented to the user via the user interface unit. The emotion engine also analyzes the user's emotions and adjusts the cooking advice based on the analysis.
[1584] Program processing
[1585] 1. The user starts a barbecue
[1586] The user issues the voice command "Start the barbecue."
[1587] The terminal converts the voice into text data using a voice recognition means and transmits the data to the server.
[1588] 2. Analysis of audio data
[1589] The server receives the text data, analyzes it, and recognizes that it is a request to "start barbecue."
[1590] The server performs the initial setup and creates the necessary setup instructions.
[1591] 3. Preparatory Instructions
[1592] The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the terminal.
[1593] The device will provide instructions to the user via voice or on-screen display.
[1594] The user follows the instructions to prepare the grill.
[1595] 4. Real-time image acquisition
[1596] The user points the device's camera at the grill and places the ingredients on it.
[1597] The device uses a camera to capture real-time images of the grill and transmits the image data to a server.
[1598] 5. Analysis of Image Data
[1599] The server analyzes the image data and evaluates the placement of the ingredients, the state of the fire, and the degree of doneness.
[1600] Based on the analysis results, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[1601] 6. Emotion Analysis Using an Emotion Engine
[1602] The device captures the user's facial expressions and tone of voice and sends them to the emotion engine.
[1603] The server uses an emotion engine to analyze the user's emotions and adjusts cooking advice based on this.
[1604] 7. Generating and Providing Advice
[1605] The server adjusts the advice based on the emotion data and sends the generated advice (e.g., "Turn up the heat a little," "Move the chicken to the right") to the terminal.
[1606] The advice received by the device is presented to the user by voice or on-screen display.
[1607] The user adjusts the heat and position of the ingredients according to the advice.
[1608] 8. Recipe suggestions
[1609] The user asks the device, "What should I make next?"
[1610] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[1611] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[1612] The server generates a recipe suggestion (e.g., "We recommend a chicken marinade") and sends it to the device.
[1613] The device will present recipe suggestions to the user via voice or on-screen display.
[1614] 9. Emotional Cooking Guidance
[1615] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[1616] The server generates cooking guidance based on the user's emotions (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device.
[1617] The device will notify the user of emotionally sensitive notifications via voice.
[1618] The user removes the cooked chicken and enjoys the meal.
[1619] The system of the present invention allows users to easily enjoy delicious barbecues, even without detailed knowledge or experience. Furthermore, by combining it with an emotion engine, a more personalized and comfortable barbecue experience can be provided. By significantly reducing the effort required for cooking and providing emotionally appropriate encouragement and advice, users can spend more time concentrating on conversations with family and friends and on their meals.
[1620] The processing flow will be explained below.
[1621] Step 1:
[1622] The user issues the voice command "Start the barbecue."
[1623] The device uses a microphone to capture the user's voice.
[1624] Step 2:
[1625] The terminal converts the voice into text data using a voice recognition means.
[1626] The terminal transmits the converted text data to the server.
[1627] Step 3:
[1628] The server receives the text data, analyzes it, and recognizes it as a request to "start barbecue."
[1629] The server performs the initialization and generates the necessary provisioning instructions.
[1630] Step 4:
[1631] The server sends the generated preparation instruction to the terminal.
[1632] The terminal will give the user instructions to prepare the grill by voice or on the screen, such as "Prepare the grill."
[1633] Step 5:
[1634] The user follows the instructions to prepare the grill.
[1635] The user points the device's camera at the grill and captures an image.
[1636] Step 6:
[1637] The device captures an image of the grill and sends the image data to a server.
[1638] The server receives the image data and uses image analysis means to evaluate the placement of ingredients, the state of the fire, and the degree of doneness.
[1639] Step 7:
[1640] Based on the results of image analysis, the server generates specific advice on adjusting the heat and changing the position of ingredients.
[1641] The server transmits the generated advice to the terminal.
[1642] Step 8:
[1643] The device displays the received advice to the user (e.g., "Turn up the heat a little").
[1644] The user adjusts the heat and position of the ingredients according to the advice.
[1645] Step 9:
[1646] The device captures the user's facial expressions and voice tone and transmits the emotional data to the server.
[1647] The server analyzes the user's emotions using an emotion engine.
[1648] Step 10:
[1649] The server adjusts cooking advice based on the results of emotion analysis (e.g., if the user seems anxious, give advice in a gentler tone).
[1650] The server sends the adjusted advice to the terminal.
[1651] Step 11:
[1652] The terminal presents the adjusted advice to the user.
[1653] The user proceeds with cooking according to the advice.
[1654] Step 12:
[1655] The user asks the device, "What should I make next?"
[1656] The terminal converts the question into text data using a voice recognition means and transmits the data to the server.
[1657] Step 13:
[1658] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data.
[1659] The server generates recipe suggestions and sends them to the terminal.
[1660] Step 14:
[1661] The device will present recipe suggestions (e.g., "We recommend this chicken marinade") to the user via voice or on-screen display.
[1662] Users can try new dishes by following suggested recipes.
[1663] Step 15:
[1664] The server periodically performs image and sentiment analysis to determine whether the food is cooked.
[1665] The server generates a notification saying "The chicken is done" and sends it to the device.
[1666] Step 16:
[1667] The terminal notifies the user by voice when the food is done.
[1668] The user removes the cooked chicken and enjoys the meal.
[1669] Example 2
[1670] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1671] Enjoying a barbecue often requires skilled techniques and a lot of prior knowledge. However, for users who are trying barbecue for the first time or who are unfamiliar with cooking, these requirements are a high hurdle and make it difficult to fully enjoy a barbecue. In addition, there is a problem that the user experience is inconsistent because the cooking status is not grasped in real time or appropriate advice is not provided according to the user's emotions.
[1672] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, an emotion analysis means, and a user interface means. This allows even beginners to easily receive detailed cooking advice on barbecues, and enables personalized support tailored to the user's emotions. Specifically, the voice recognition means recognizes the user's voice commands and converts them into text data. The image analysis means captures real-time images of the grill and analyzes the state of the ingredients and the fire. Furthermore, the emotion analysis means analyzes the user's emotions from their facial expressions and voice tone. The cooking management means integrates these analysis results to generate appropriate cooking advice, which is then provided to the user through the user interface means.
[1673] "Speech recognition means" refers to means having the function of receiving voice commands from a user and converting them into text data.
[1674] The "image analysis means" is a means having the function of analyzing images captured in real time and evaluating the state of the ingredients and the fire.
[1675] The "cooking management means" is a means having a function of generating cooking advice based on the results of the image analysis means and the emotion analysis means.
[1676] "Communication means" refers to a means that utilizes protocols and technologies for sending and receiving data.
[1677] A "user interface means" is a means for presenting information to a user and receiving input from a user.
[1678] The "emotion analysis means" is a means having the function of capturing the user's facial expressions and tone of voice and analyzing their emotions.
[1679] This invention is a system that further improves the user experience by combining an emotion engine with a smart AI assistant system that supports the barbecue experience. This system includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, and an emotion analysis means.
[1680] Hardware and software used
[1681] The system uses the following major hardware and software:
[1682] 1. Speech recognition means: Typically using a cloud-based speech recognition service, e.g., a speech recognition cloud API.
[1683] 2. Image analysis method: Uses the OpenCV library and deep learning models (e.g., TensorFlow).
[1684] 3. Cooking control method: A custom algorithm implemented in Python is used.
[1685] 4. Communication method: Network communication is performed using the HTTP / HTTPS protocol.
[1686] 5. User interface: Implemented through a smartphone app (iOS / Android) or web browser.
[1687] 6. Emotion analysis means: Use an API that performs facial expression analysis and voice tone analysis (e.g., emotion analysis cloud API).
[1688] Processing of this system
[1689] 1. When a user wants to start a barbecue, they say the voice command "Start the barbecue." This voice command is converted into text data by the device using a voice recognition cloud API and sent to the server.
[1690] 2. The server analyzes the received text data and recognizes that the user's request is to "start a barbecue." It uses a natural language processing (NLP) library to process this. The server then performs its initial setup and creates the necessary preparation instructions.
[1691] 3. The server generates a preparation instruction (e.g., "Prepare the grill") and sends it to the device. The device then presents the instructions to the user via voice or visual display, and the user prepares the grill accordingly.
[1692] 4. The user points the device's camera at the grill and places the food on it. The device uses the camera to capture real-time images of the grill and sends the image data to the server.
[1693] 5. The server analyzes the received image data using OpenCV and deep learning models. The analysis results include the placement of the ingredients, the state of the fire, and the degree of doneness. Based on this information, the server generates specific advice on adjusting the heat or repositioning the ingredients.
[1694] 6. The device captures the user's facial expressions and voice tone and sends them to the emotion analysis cloud API. The server uses an emotion engine to analyze the user's emotions and adjust the cooking advice accordingly.
[1695] 7. The server generates advice based on the emotion data and sends it to the device. The device then presents the advice to the user via voice or on-screen display. The user adjusts the heat and position of the ingredients according to the advice.
[1696] 8. The user asks the device, "What should I make next?" The device uses a voice recognition cloud API to convert the question into text data and sends it to the server. The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The server generates a recipe suggestion (e.g., "I recommend marinated chicken") and sends it to the device. The device then presents the recipe suggestion to the user by voice or on a screen.
[1697] 9. The server periodically performs image and emotion analysis to determine whether the food is cooked. The server generates cooking guidance based on the user's emotion (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?") and sends it to the device. The device then notifies the user with voice notifications that take emotion into consideration. The user follows the instructions to remove the cooked chicken and enjoy their meal.
[1698] Specific examples
[1699] Example 1: Your first barbecue
[1700] 1. The user says the voice command "Start the barbecue."
[1701] 2. The device converts the voice into text and sends it to the server.
[1702] 3. The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal.
[1703] 4. The user prepares the grill and points the device camera at the grill.
[1704] 5. The device sends the captured image to the server, which then analyzes the image.
[1705] 6. The server generates advice on adjusting the heat and sends it to the device.
[1706] 7. The device provides advice to the user, who then adjusts the heat.
[1707] Prompt Sentence Examples
[1708] Prompt 1: "Start the BBQ Assistant to provide a set of guidance for first-time barbecuers."
[1709] Prompt 2: "When a user requests a new recipe, recommend a recipe that fits their mood that day."
[1710] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1711] Step 1:
[1712] The user issues the voice command "Start the barbecue."
[1713] The device converts the speech into text data using a speech recognition cloud API. The input is the user's speech, and the output is the converted text data. The device then sends this text data to the server.
[1714] Step 2:
[1715] The server analyzes the received text data. The input is text data sent from the device, and the server uses a natural language processing library to recognize that it is a request to "start a barbecue." The output is the result of recognizing the request as "start a barbecue." The server performs initial configuration and creates the necessary preparation instructions.
[1716] Step 3:
[1717] The server generates a preparation instruction such as "Prepare the grill" and sends it to the terminal. The input is the initial setting result, and the output is a text message of the preparation instruction. The terminal receives this instruction and presents it to the user by voice or on-screen display. The user follows the instruction and performs the action of preparing the grill.
[1718] Step 4:
[1719] The user points the device's camera at the grill and places ingredients on it. The input is the camera image. The device uses the camera to capture real-time images of the grill and sends the image data to the server. The output is the captured real-time image data.
[1720] Step 5:
[1721] The server analyzes the received image data. The input is image data sent from the device. The server uses OpenCV and deep learning models (such as TensorFlow) to evaluate the placement of ingredients, the state of the fire, and the degree of doneness. The output is the analysis results, which are summarized as specific advice on adjusting the heat or changing the position of ingredients.
[1722] Step 6:
[1723] The device captures the user's facial expressions and voice tone. The input is the user's facial expression and voice data. The device sends this to the emotion analysis cloud API. The output is the emotion analysis result, which is sent to the server.
[1724] Step 7:
[1725] The server receives the emotion data analysis results and generates adjustment advice. The input is emotion data and image analysis results. The server adjusts the cooking advice based on this data and generates specific advice (e.g., "Turn up the heat a little," "Move the chicken to the right"). The output is the adjusted advice. The server sends this advice to the device.
[1726] Step 8:
[1727] The device provides advice to the user by voice or on-screen display. The input is advice data from the server, and the output is the information presented to the user. The user follows the advice and adjusts the heat and position of the ingredients.
[1728] Step 9:
[1729] The user asks the device, "What should I make next?" The input is the user's voice command. The device again uses the speech recognition cloud API to convert the question into text data and sends that data to the server. The output is the converted text data.
[1730] Step 10:
[1731] The server selects the optimal recipe based on the user's past cooking history, current ingredient information, and emotional data. The input is the user's past cooking history, current ingredient information, and emotional data, and the output is the optimal recipe suggestion. The server generates a recipe suggestion (e.g., "We recommend marinated chicken") and sends it to the device.
[1732] Step 11:
[1733] The terminal presents recipe suggestions to the user by voice or on-screen display. The input is the recipe suggestion data from the server, and the output is the suggestion to the user.
[1734] Step 12:
[1735] The server periodically performs image analysis and emotion analysis. The input is real-time images and user emotion data. The server determines whether the ingredients are cooked and generates cooking guidance. The output is cooking guidance (e.g., "The chicken is cooked perfectly!", "Are you ready to proceed?"). The server then sends the generated cooking guidance to the device.
[1736] Step 13:
[1737] The device provides emotionally sensitive notifications to the user. The input is cooking guidance data from the server, and the output is a notification to the user. The user follows the instructions, removes the cooked chicken, and enjoys the meal.
[1738] Through the above processing steps, the smart AI assistant system of the present invention can help users enjoy a more enjoyable and easier barbecue. By providing personalized advice based on the cooking progress and the user's emotions, even beginners can enjoy a delicious barbecue.
[1739] (Application example 2)
[1740] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1741] Conventional food delivery systems lack personalized menu suggestions and cooking advice based on the user's emotional state, which hinders the improvement of user experience. Furthermore, when users are cooking, they often do not receive appropriate advice in real time, which leads to a decrease in satisfaction.
[1742] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, an image analysis means, a cooking management means, a communication means, a user interface means, an emotion analysis means, and a database means. This makes it possible to propose optimal menus and provide cooking advice according to the user's emotional state, which is expected to improve the user experience.
[1743] "Voice recognition" is a technology that receives a user's voice commands and converts them into text data.
[1744] "Image analysis" is a technology that analyzes the state of ingredients and fire based on image data captured in real time.
[1745] "Cooking management" is a technology that generates cooking advice based on the results of image analysis and presents it to the user.
[1746] "Communication means" refers to the technology used to send and receive voice data, image data, and analysis results between the server and the user terminal.
[1747] A "user interface" is a means for receiving voice commands and image data input from a user, and for providing generated advice and menu suggestions to the user.
[1748] "Emotion analysis" is a technology that analyzes a user's facial expressions and tone of voice to understand their emotional state.
[1749] The "database" is a technology that stores a user's past order history and emotional data, and generates suggestions based on the analysis results as needed.
[1750] The present invention relates to an emotion-responsive food delivery support application, specifically a system that provides optimal menu suggestions and cooking advice based on a user's voice commands and emotional state. The system includes a voice recognition unit, an image analysis unit, a cooking management unit, a communication unit, a user interface unit, an emotion analysis unit, and a database unit.
[1751] System Overview
[1752] The solution begins when a user launches a food delivery application on their smartphone. First, the user issues a voice command such as "Tell me what menu items you want." The smartphone uses a microphone to capture the voice and converts it into text using a speech recognition library called Google Cloud Speech-to-Text. This text is then sent to a server, which analyzes it and understands the user's request.
[1753] Next, an emotion analysis tool analyzes the user's facial expressions and tone of voice to generate emotion data. This uses the Microsoft Azure Emotion API. The emotion data is used to understand the user's emotional state, and appropriate menu suggestions and cooking advice are provided based on that emotional state. Past order history and emotion data are stored in a database (e.g., MySQL), and optimal suggestions are generated based on this data.
[1754] For example, if the user is in a state of mind where they want to relax, the system will suggest "relaxing herbal tea." If the user is cooking at home, the system will perform real-time image analysis and provide appropriate cooking advice. Image analysis uses a library called OpenCV.
[1755] The specific hardware and software used
[1756] Hardware: Smartphone (microphone, camera, speaker)
[1757] software:
[1758] Speech recognition library: Google Cloud Speech-to-Text
[1759] Sentiment analysis library: Microsoft Azure Emotion API
[1760] Image analysis library: OpenCV
[1761] Database: MySQL
[1762] Prompt Sentence Examples
[1763] Specific examples of prompts are as follows:
[1764] User's order history: pizza, pasta, salad
[1765] Emotional state: Relaxed
[1766] User Question: What's your recommended menu item?
[1767] Response: I recommend herbal tea. It's perfect for relaxing.
[1768] In this way, the system of the present invention allows users to enjoy a personalized food delivery experience based on their voice commands and emotional state.
[1769] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1770] Step 1:
[1771] A user launches a food delivery application on their smartphone and says, "Tell me what menu items you recommend." Here, the input is the user's voice command, and the smartphone's microphone is used to capture the voice data. This voice data is converted into text data using the Google Cloud Speech-to-Text library. The converted text data is the output.
[1772] Step 2:
[1773] The speech-recognized text data is sent to the server. Here, the input is the converted text data, and the server receives the data. The server performs text analysis to understand the user's request. As a result of the analysis, data called "recommended menu suggestions" is output.
[1774] Step 3:
[1775] The emotion analyzer captures the user's facial expressions and voice tone and generates emotion data using the Microsoft Azure Emotion API, where the input is the user's real-time facial expression data and voice tone, and the API analyzes the user's emotional state (e.g., relaxed, happy, etc.), and the analyzed emotion data is the output.
[1776] Step 4:
[1777] The server retrieves the generated emotion data and past order history from a database. Here, the input is emotion data and order history data. This data is retrieved through a database query, and the server selects the optimal menu based on this data. The selected menu suggestion is the output.
[1778] Step 5:
[1779] The server generates a selected menu suggestion (e.g., "I recommend herbal tea. It's perfect for relaxing.") and sends it to the smartphone via communication means. Here, the input is the data of the selected menu suggestion, which is sent to the smartphone. The output is the smartphone presenting the suggested menu to the user by voice or on-screen display.
[1780] Step 6:
[1781] When a user cooks at home, they use their smartphone camera to capture ingredients and the cooking process. The input is real-time image data, which is analyzed using an image analysis library (OpenCV). The analysis results are then output as cooking advice (e.g., "Adjust the heat.").
[1782] Step 7:
[1783] The server sends the generated cooking advice to the smartphone and presents it to the user. Here, the input is the analyzed cooking advice data, and the data is sent to the smartphone. The output of the smartphone is to provide advice to the user by voice or on-screen display.
[1784] Step 8:
[1785] The user then asks a follow-up question (e.g., "What's the next recommended menu item?") by voice. The smartphone's microphone is again used to capture the voice data, which is then converted into text data using the Google Cloud Speech-to-Text library. This voice data is the input, and the converted text data is the output.
[1786] Step 9:
[1787] The server analyzes the converted text data and generates the next menu suggestion based on the information from the database and the sentiment data. Here, the input is the text data based on the follow-up question and related data, and the output is the next menu suggestion as the analysis result.
[1788] Step 10:
[1789] The server sends the next menu proposal to the smartphone and presents it to the user. Here, the input is the generated next menu proposal data, which is sent to the smartphone. The output of the smartphone is to present it to the user by voice or on-screen display.
[1790] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1791] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1792] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1793] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1794] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1795] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1796] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1797] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1798] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1799] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1800] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1801] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1802] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1803] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1804] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1805] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1806] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1807] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1808] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1809] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1810] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1811] The following is further disclosed regarding the above embodiment.
[1812] (Claim 1)
[1813] a speech recognition means;
[1814] Image analysis means;
[1815] Cooking management means;
[1816] means of communication;
[1817] user interface means;
[1818] A system including:
[1819] (Claim 2)
[1820] 10. The system of claim 1, wherein the speech recognition means includes means for receiving voice commands from a user and converting them into text data.
[1821] (Claim 3)
[1822] 2. The system of claim 1, wherein the image analysis means includes means for capturing images of the barbecue grill in real time and analyzing the state of the ingredients and the fire.
[1823] (Claim 4)
[1824] 2. The system according to claim 1, wherein the cooking management means includes means for generating advice regarding adjusting heat and moving ingredients based on the results of image analysis.
[1825] (Claim 5)
[1826] 2. The system of claim 1, wherein the communication means includes means for transmitting data from the voice recognition means and the image analysis means to a server and for communicating advice from the server to the user interface means.
[1827] (Claim 6)
[1828] 2. The system according to claim 1, wherein the user interface means includes means for presenting advice to the user by voice or by displaying on a screen.
[1829] (Claim 7)
[1830] The system according to claim 1, wherein the cooking management means includes means for suggesting recipes based on the user's past cooking history and current cooking status.
[1831] "Example 1"
[1832] (Claim 1)
[1833] a speech recognition means;
[1834] Image analysis means;
[1835] Cooking management means;
[1836] means of communication;
[1837] user interface means;
[1838] a real-time data collection means;
[1839] data analysis means;
[1840] an advice generating means;
[1841] A system including:
[1842] (Claim 2)
[1843] 10. The system of claim 1, wherein the speech recognition means includes means for receiving voice commands from a user and converting them into text data.
[1844] (Claim 3)
[1845] 2. The system of claim 1, wherein the image analysis means includes means for capturing images of the barbecue grill in real time and analyzing the state of the ingredients and the fire.
[1846] (Claim 4)
[1847] 10. The system of claim 1, wherein the data analysis means includes means for analyzing user intent using natural language processing techniques.
[1848] (Claim 5)
[1849] 2. The system according to claim 1, wherein the advice generating means includes means for generating specific advice regarding adjusting heat or changing the position of ingredients based on the analysis results.
[1850] (Claim 6)
[1851] 2. The system of claim 1, wherein the user interface means includes means for presenting the generated advice to the user by voice or by displaying it on a screen.
[1852] "Application Example 1"
[1853] (Claim 1)
[1854] a speech recognition means;
[1855] Image analysis means;
[1856] Cooking management means;
[1857] means of communication;
[1858] user interface means;
[1859] A means to assess and adjust cooking quality in real time;
[1860] A means of providing advice using a generative AI model; and
[1861] A system including:
[1862] (Claim 2)
[1863] 10. The system of claim 1, wherein the speech recognition means includes means for receiving voice commands from a user and converting them into text data.
[1864] (Claim 3)
[1865] 2. The system of claim 1, wherein the image analysis means includes means for capturing images of the barbecue grill in real time and analyzing the state of the ingredients and the fire.
[1866] "Example 2: Combining Emotion Engines"
[1867] (Claim 1)
[1868] a speech recognition means;
[1869] Image analysis means;
[1870] Cooking management means;
[1871] means of communication;
[1872] user interface means;
[1873] A sentiment analysis means;
[1874] A system including:
[1875] (Claim 2)
[1876] 10. The system of claim 1, wherein the speech recognition means includes means for receiving voice commands from a user and converting them into text data.
[1877] (Claim 3)
[1878] 2. The system of claim 1, wherein the image analysis means includes means for capturing images of the barbecue equipment in real time and analyzing the state of the food and the fire.
[1879] (Claim 4)
[1880] 2. The system of claim 1, wherein the emotion analysis means includes means for capturing a user's facial expression and tone of voice and analyzing the emotion.
[1881] (Claim 5)
[1882] The system according to claim 1 , wherein the cooking management means includes means for generating cooking advice based on results of the image analysis means and the sentiment analysis means.
[1883] (Claim 6)
[1884] 2. The system according to claim 1, wherein the user interface means includes means for presenting the advice generated by the cooking management means to the user by voice or on-screen display.
[1885] "Application example 2 when combining emotion engines"
[1886] (Claim 1)
[1887] a speech recognition means;
[1888] Image analysis means;
[1889] Cooking management means;
[1890] means of communication;
[1891] user interface means;
[1892] A sentiment analysis means;
[1893] database means;
[1894] A system including:
[1895] (Claim 2)
[1896] 10. The system of claim 1, wherein the speech recognition means includes means for receiving voice commands from a user and converting them into text data.
[1897] (Claim 3)
[1898] The system according to claim 1, wherein the image analysis means includes means for capturing and analyzing the state of ingredients and fire in real time.
[1899] (Claim 4)
[1900] 2. The system of claim 1, wherein the emotion analysis means includes means for analyzing a user's facial expression and tone of voice to generate emotion data.
[1901] (Claim 5)
[1902] 2. The system of claim 1, wherein the database means includes means for storing a user's past order history and emotion data, and for generating suggestions based on the analysis results.
[1903] (Claim 6)
[1904] The system of claim 1 , further comprising means for adjusting cooking advice based on the emotion data. [Explanation of symbols]
[1905] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a voice recognition means; Image analysis means; Cooking control means; means of communication; User interface means; A system including:
2. 2. The system of claim 1, wherein said speech recognition means includes means for receiving voice commands from a user and converting them into text data.
3. The system of claim 1 , wherein the image analysis means includes means for capturing images of the barbecue grill in real time and analyzing the state of the food and the fire.
4. The system according to claim 1 , wherein the cooking management means includes means for generating advice regarding adjustment of heat and movement of ingredients based on the results of image analysis.
5. 2. The system of claim 1, wherein said communication means includes means for transmitting data from the voice recognition means and the image analysis means to a server and means for conveying advice from the server to the user interface means.
6. 2. The system according to claim 1, wherein said user interface means includes means for presenting advice to the user by voice or by display on a screen.
7. The system according to claim 1 , wherein the cooking management means includes means for suggesting recipes based on the user's past cooking history and current cooking status.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A