System
A system that analyzes user cooking videos against professional recipes, generating feedback and advice, addresses the challenge of recreating professional recipes at home by improving cooking skills through detailed comparison and synthesis.
Patent Information
- Application Number
- JP2024116427
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Users face challenges in recreating professional recipes at home due to variations in cooking utensils and environments, leading to difficulty in understanding and accurately following the recipes, which limits the improvement of their cooking skills.
A system that allows users to upload their cooking videos, compare them with professional recipe videos, analyze differences, and generate feedback and advice, synthesizing a composite video for easy understanding and improvement.
Facilitates accurate recreation of professional recipes at home by providing specific feedback and advice, enhancing cooking skills through detailed analysis and comparison with professional methods.
Smart Images

Figure 2026014953000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In conventional cooking classes and recipe videos, users face many challenges when trying to recreate professional recipes at home because the cooking utensils and environments used vary. Specifically, users often have difficulty understanding how to make the recipe in their own environment, even if they understand the recipe. In such situations, it is difficult for users to accurately recreate professional recipes, limiting the improvement of their cooking skills. The present invention aims to solve these challenges and provide technology that allows users to easily recreate professional recipes in their own cooking environment. [Means for solving the problem]
[0005] The present invention solves the above problems by providing a system including the following means.
[0006] A means for uploading cooking videos taken by users;
[0007] A way to import professional recipe videos,
[0008] A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos;
[0009] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences;
[0010] means for generating feedback and advice for a user based on the difference;
[0011] means for generating a composite video including the generated feedback and advice;
[0012] means for providing the composite video to a user;
[0013] It is a system including:
[0014] This system allows users to film and upload cooking videos at home, compare them with professional recipe videos, and receive specific advice suited to their home environment. This makes it easier for users to more accurately recreate professional recipes and improve their cooking skills.
[0015] "User" refers to someone who uses the system to cook at home.
[0016] "Cooking videos" refer to videos of users cooking at home.
[0017] "Professional recipe videos" refer to videos filmed based on recipes provided by professional chefs or cooks.
[0018] "Means for uploading" refers to the functions and methods for sending cooking videos filmed by users to the system.
[0019] "Means of importing" refers to the functions and methods for obtaining professional recipe videos into the system.
[0020] "Means of analysis" refers to the functions and methods for analyzing the content of cooking videos and recipe videos and extracting the necessary information.
[0021] "Cooking procedure" refers to the specific steps and order of operations involved in preparing a dish.
[0022] "Utensils" refers to the tools used in cooking (e.g., pots, frying pans, knives, etc.).
[0023] "Ingredients" refers to the ingredients used in cooking.
[0024] "Means for analyzing differences" refers to functions and methods for comparing a user's cooking video with a professional recipe video to find differences.
[0025] "Means for generating feedback and advice" refers to a function or method for creating specific improvements and advice for the user based on the analyzed differences.
[0026] "Synthetic videos" refer to videos that combine users' cooking videos with professional recipe videos and include feedback and advice.
[0027] "Means for providing" refers to the functions and methods for displaying or transmitting the generated composite video to the user. [Brief explanation of the drawings]
[0028] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0029] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0030] First, the terms used in the following description will be explained.
[0031] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0032] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0033] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0034] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0035] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0036] [First embodiment]
[0037] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0038] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0039] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0040] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0041] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0042] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0043] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0044] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0045] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0046] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0047] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0048] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0049] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0050] System Overview
[0051] This system consists of a user terminal, a server for processing, and a network for exchanging information between them. The system programs are described in detail below.
[0052] User device functions
[0053] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[0054] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[0055] Server Features
[0056] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[0057] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[0058] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[0059] 4. Difference analysis function: The app has a function to compare the user's cooking video with a professional recipe video and analyze the differences. Specifically, it identifies differences in cooking procedures, utensils, and ingredients.
[0060] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as "the amount of salt is too much" or "shorten the heating time."
[0061] 6. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[0062] Overview of operation
[0063] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with information from professional recipe videos, and differences are analyzed. Based on the analysis results, feedback and advice tailored to the user is generated and inserted into the composite video. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[0064] Specific examples
[0065] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. As a result of the analysis, it identifies differences such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, it generates specific advice such as "slicing the bacon thinner," "mixing the eggs early," and "reducing the amount of salt by half." This advice is then combined with a professional recipe video and provided to the user. The user watches the combined video and understands how to make a carbonara that is closer to what a professional would make.
[0066] In this way, the system makes it easy for users to recreate professional recipes at home and helps improve their cooking skills.
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] Users film themselves cooking at home, using the camera function of their smartphones or tablets to record the entire cooking process.
[0070] Step 2:
[0071] Cooking videos taken by users are uploaded to the system using a dedicated application or website.
[0072] Step 3:
[0073] The device receives the uploaded cooking video and sends it to the server.
[0074] Step 4:
[0075] The server stores the received cooking video in a database and begins analyzing it.
[0076] Step 5:
[0077] The server analyzes the video frame by frame and extracts information about cooking steps, utensils, and ingredients. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are extracted.
[0078] Step 6:
[0079] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[0080] Step 7:
[0081] The server also performs video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[0082] Step 8:
[0083] The server compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, specific differences such as "too much salt" or "too short heating time" can be identified.
[0084] Step 9:
[0085] Based on the analysis results, the server generates specific feedback and advice for the user, such as "reduce the amount of salt" or "extend the cooking time."
[0086] Step 10:
[0087] The server generates a composite video for each user, which displays a professional recipe video and the user's cooking video side by side, and inserts feedback and advice based on the differences.
[0088] Step 11:
[0089] The server sends the generated composite video to the terminal.
[0090] Step 12:
[0091] The device then displays the synthesized video to the user, who can then watch the video to understand the differences between the professional's movements and their own, and use the feedback to improve their cooking procedures and techniques.
[0092] Example 1
[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0094] With conventional cooking methods, it was difficult for users to recreate professional recipes, and there were limited ways to receive specific feedback and advice. Furthermore, there was no system that analyzed the user's cooking process in detail and provided specific suggestions for improvement. This meant that users were unable to faithfully recreate professional recipes, making it difficult to efficiently improve their cooking skills.
[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0096] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and the recipe videos using a video analysis library and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences using a deep learning model, means for automatically generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice using video editing software, and means for providing the composite video to the user. This makes it easier for users to recreate professional recipes at home and allows them to efficiently improve their cooking skills by receiving specific feedback and advice.
[0097] 1. "Cooking videos filmed by users" refers to video data that records the process of a user cooking at home.
[0098] 2. "Professional recipe video" refers to video data that records the process of a professional chef preparing a specific dish.
[0099] 3. "Video Analysis Library" means a software library for analyzing video data frame by frame and extracting specific information from each frame.
[0100] 4. A "cooking procedure" is a series of steps for preparing a dish, including the specific tasks and order of the steps.
[0101] 5. "Utensils" are tools and equipment used in cooking, including knives and pots.
[0102] 6. "Ingredients" are the ingredients and seasonings used in making a dish.
[0103] 7. A "deep learning model" is a machine learning algorithm that performs highly accurate analysis and predictions by learning features from large amounts of data.
[0104] 8. "Differences" are the differences between the user's cooking procedures, equipment, and ingredients and those of professional recipe videos.
[0105] 9. "Feedback" means evaluation, comments, or advice provided to the user based on the analysis results.
[0106] 10. "Advice" means specific instructions or advice provided to a user to improve their cooking skills.
[0107] 11. "Synthetic video" is video data that displays a professional recipe video and a user's cooking video side by side, with feedback and advice based on the differences inserted.
[0108] 12. "Video editing software" means software for combining and editing multiple video data to generate new video.
[0109] 13. "Uploading means" refers to a function or device that allows users to send cooking videos they have filmed to the server.
[0110] 14. "Means for capturing" means a function or device that allows the server to receive and store professional recipe videos.
[0111] 15. "Means for providing" refers to a function or device for transmitting the generated composite video to a user so that the video can be viewed.
[0112] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0113] This system consists of a user terminal, a server, and a network that exchanges information between them. Each element of the system is explained in detail below.
[0114] User device functions
[0115] Shooting Function:
[0116] Users use the camera on their smartphone or tablet to film themselves cooking at home. To do so, they launch the device's "camera" app and press the record button to record the entire cooking process from start to finish.
[0117] Upload function:
[0118] The user uploads the video they have taken to the server through a dedicated application (e.g., "CookingPro"). To do this, open the application, go to the upload screen, and select the saved video file.
[0119] Server Features
[0120] Video receiving and storage function:
[0121] The server receives cooking videos uploaded by users and stores them in a storage service (e.g., Amazon S3).
[0122] Video analysis features:
[0123] The server uses software such as OpenCV and ffmpeg to analyze the received video frame by frame, extracting information about the cooking steps, utensils used, and ingredients.
[0124] Database features:
[0125] The server stores the cooking procedures, utensils, and ingredient information extracted through the analysis in a database (e.g., MySQL or PostgreSQL).
[0126] Differential analysis features:
[0127] The server compares the user's cooking video with professional recipe videos using a deep learning model (e.g., TensorFlow or PyTorch) to analyze differences in cooking procedures, utensils used, and ingredients.
[0128] Feedback generation features:
[0129] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "cut the bacon thinner," "add the eggs earlier," or "reduce the amount of salt by half."
[0130] Composite video generation function:
[0131] The server displays the professional recipe video and the user's cooking video side by side, creating a composite video that includes feedback. The composite is created using video editing software (e.g., Adobe Premiere Pro or MoviePy).
[0132] Features provided:
[0133] The server provides the generated composite video to the user via a communication method such as an application or email.
[0134] Specific examples
[0135] For example, a user can film a video of themselves cooking "Carbonara" and upload it to the server via the "CookingPro" app. The server receives the video and begins analyzing it. OpenCV and ffmpeg are used for the analysis, and information such as "how to cut the bacon," "when to mix the eggs," and "how much salt" is extracted from each frame. The information obtained is then stored in a database.
[0136] The server then compares the video with a professional "Carbonara" recipe video and analyzes differences such as the bacon being too thick, the eggs being added too late, or the amount of salt being too high. This is done using TensorFlow. Based on the analysis results, feedback is generated, such as "slicing the bacon thinner," "adding the eggs earlier," or "reducing the amount of salt by half." Adobe Premiere Pro is then used to generate a composite video that displays the professional recipe video and the user's cooking video side by side. Users can watch this composite video and understand specific areas for improvement, allowing them to improve their cooking skills.
[0137] Prompt Sentence Examples
[0138] The system analyzes the cooking video of carbonara filmed by the user, such as how to cut the bacon, when to mix the eggs, and the amount of salt, and provides specific feedback to create a composite video.
[0139] In this way, this system makes it easier for users to recreate professional recipes at home and helps improve their cooking skills.
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Step 1:
[0142] A user shoots a cooking video.
[0143] Specifically, the user launches the camera app on their smartphone or tablet and presses the record button to record the cooking process. Taking detailed photos, including the overall image of the food being cooked and the movements of the hands, improves the accuracy of the analysis.
[0144] The input is a real scene of the cooking process, and the output is a recorded video file.
[0145] Step 2:
[0146] Cooking videos taken by users are uploaded from their terminals to a server.
[0147] Specifically, the user opens a dedicated application (e.g., "CookingPro"), goes to the upload screen, selects the saved video file, and presses the send button.
[0148] The input is a cooking video file stored on the user's terminal, and the output is a video file sent to the server.
[0149] Step 3:
[0150] The server receives and stores the video.
[0151] Specifically, the server receives the uploaded video data, stores it in a temporary directory, and then transfers it to cloud storage (e.g., Amazon S3).
[0152] The input is a video file uploaded by a user, and the output is a video file saved in the server's storage.
[0153] Step 4:
[0154] The server begins analyzing the video.
[0155] Specifically, the server uses OpenCV and ffmpeg to divide the received video into frames and extract information about the cooking steps, utensils, and ingredients from each frame. For example, it performs image recognition on each frame to identify ingredients and cooking utensils.
[0156] The input is a stored video file, and the output is extracted cooking instructions, utensils, and ingredient information for each frame.
[0157] Step 5:
[0158] The server stores information about cooking procedures, equipment, and ingredients in a database.
[0159] Specifically, the server stores the information extracted by the analysis in a database (e.g., MySQL or PostgreSQL) and manages the information in a structured manner.
[0160] The input is the information on cooking procedures, utensils, and ingredients obtained as a result of the analysis, and the output is a database in which this information is stored.
[0161] Step 6:
[0162] The server compares it with professional recipe videos and analyzes the differences.
[0163] Specifically, the server uses deep learning models (e.g., TensorFlow or PyTorch) to analyze and compare the user's cooking video with professional recipe videos, identifying differences in cooking procedures, utensils, and ingredients, for example, detecting discrepancies in the timeline and differences in procedures.
[0164] The input is the analysis results of the user's cooking video and the analysis results of the professional recipe video, and the output is the identified difference information.
[0165] Step 7:
[0166] The server generates specific feedback based on the differences.
[0167] Specifically, the server uses a generative AI model to automatically generate specific feedback and advice for the user based on the difference information, such as "cut the bacon thinner," "add the eggs early," or "reduce the amount of salt."
[0168] The input is the difference information and the output is the generated feedback and advice.
[0169] Step 8:
[0170] The server generates a composite video including the feedback and provides it to the user.
[0171] Specifically, the server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video with feedback inserted using video editing software (e.g., Adobe Premiere Pro or MoviePy). The generated composite video is then provided to the user's device or via email.
[0172] The inputs are professional recipe videos, user cooking videos, and generated feedback, and the output is a composite video file.
[0173] These processing steps make it easier for users to recreate professional recipes at home and improve their cooking skills by receiving specific feedback and advice.
[0174] (Application example 1)
[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0176] Conventional cooking assistance systems lack the ability to provide real-time feedback when users are trying to recreate professional recipes. This makes it difficult for users to notice problems while cooking, resulting in less-than-perfect dishes. Furthermore, there is a lack of a way to specifically analyze differences in cooking procedures and ingredient usage, and provide the results in an easy-to-understand manner.
[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0178] In this invention, the server includes a means for uploading cooking videos filmed by users, a means for importing cooking recipe videos, a means for analyzing the cooking videos and the recipe videos and extracting their respective cooking steps, utensils, and ingredients, a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, a means for generating feedback and advice for the user based on the differences, a means for displaying the generated feedback and advice in real time, a means for generating a composite video including the feedback and advice, and a means for providing the composite video to the user. This allows users to receive feedback in real time while cooking, making it easier to identify specific areas for improvement. Furthermore, the differences between professional recipe videos and the user's cooking video are clearly presented, allowing users to easily recreate recipes.
[0179] "Cooking videos filmed by users" are videos in which users record the cooking process using a home camera or smart device.
[0180] A "cooking recipe video" is a video in which a professional or expert demonstrates a cooking method while explaining the steps.
[0181] A "means" is a method or mechanism for achieving a specific function or purpose.
[0182] "Analysis" means analyzing the video content frame by frame and extracting information such as each cooking step, utensils used, and ingredients.
[0183] A "cooking procedure" is a series of steps or methods for preparing a dish.
[0184] "Utensils" are tools and machines used in cooking.
[0185] "Ingredients" are the ingredients or components used to make a dish.
[0186] "Differences" are differences or discrepancies between the user's cooking procedures, utensils used, and ingredients and the content of professional recipe videos.
[0187] "Feedback" refers to information or advice provided to the user based on the analysis results.
[0188] "Advice" refers to suggestions to the user for specific improvements, such as cooking methods or how to use ingredients.
[0189] "Real-time display" means that users receive instant feedback and advice while cooking.
[0190] A "composite video" is a video that combines a professional recipe video with a user's cooking video, and inserts feedback and advice based on the differences.
[0191] This invention is a system that makes it easy for users to recreate professional recipes when cooking at home. The main components of the system include a user terminal and a server. The system programs and their processing are described in detail below.
[0192] First, the user uses a device such as a smartphone or smart glasses to record the cooking process. The video is then uploaded to a server via a dedicated application. This application has both recording and uploading functions, providing an environment in which users can easily share videos.
[0193] The server receives both cooking videos and cooking recipe videos uploaded by users and has the analysis function to analyze each. For the analysis, OpenCV (a library for video analysis) and TensorFlow (a library for using machine learning models) are used. Specifically, the video is analyzed frame by frame to extract information on cooking steps, utensils used, and ingredients.
[0194] The extracted information is compared with the content of the cooking recipe video. This comparison uses differential analysis to identify the differences between the user's cooking and the professional recipe. The results of the differential analysis generate specific feedback and advice for the user. This feedback and advice is displayed in real time on the user's device.
[0195] In addition, the system synthesizes professional recipe videos with users' cooking videos, generating a composite video that includes feedback and advice based on the differences, making it easier for users to identify specific areas for improvement while cooking.
[0196] For example, a user may film and upload a video of themselves making beef stew. In this case, the server analyzes the video and identifies differences such as "the meat is cut too large" and "the amount of wine is too small." Based on this, specific advice such as "cut the meat into bite-sized pieces" and "add more wine" is generated.
[0197] An example of a prompt for a generative AI model is:
[0198] "Compare the professional recipe video for making beef stew with the user cooking video and provide the following differences and feedback:
[0199] 1. Size of meat cut
[0200] 2. Amount of wine used
[0201] This system makes it easy for users to faithfully recreate professional recipes at home, improving their cooking skills.
[0202] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0203] Step 1:
[0204] The user takes a video of the cooking process using a smartphone or smart glasses camera. The input at this stage is the cooking action and scene. The video is captured and saved as a video file. The output is a cooking video file.
[0205] Step 2:
[0206] The user uploads the cooking video they have filmed to the server using a dedicated application. At this stage, the input is the cooking video file, and the output is a notification that the data has been transferred to the server.
[0207] Step 3:
[0208] The server saves the cooking video received from the user. The input at this stage is the uploaded cooking video file, and the output is the destination directory and file path. Here, the file system is used to save the video in the appropriate directory.
[0209] Step 4:
[0210] The server analyzes stored cooking videos and professional cooking recipe videos. The inputs are cooking video files and professional recipe video files. The videos are decomposed frame by frame, and the OpenCV library is used to extract information about cooking steps, utensils, and ingredients. The output is a list of cooking steps, utensils, and ingredients.
[0211] Step 5:
[0212] The server compares the user's cooking procedures, utensils, and ingredients with the professional recipes based on the extracted information. This comparison uses a difference analysis algorithm. The input is the list of information extracted in step 4 and the professional recipe information list. The output is the specific differences between the user's cooking and the recipes.
[0213] Step 6:
[0214] The server generates specific feedback and advice for the user based on the results of the difference analysis. The input at this stage is the difference information. The generative AI model is used to generate advice. The output is a feedback message and a list of specific advice.
[0215] Step 7:
[0216] The server displays the generated feedback and advice on the user's terminal in real time. The input at this stage is a feedback / advice list, and the output is the feedback / advice displayed on the user's terminal screen.
[0217] Step 8:
[0218] The server synthesizes professional recipe videos and user cooking videos, generating a composite video with feedback and advice based on the differences. The inputs are cooking video files, recipe video files, and a feedback and advice list. The composite video is created using video editing software. The output is a composite video file.
[0219] Step 9:
[0220] The server provides the generated composite video to the user. The input at this stage is the composite video file, and the output is the composite video file sent to the user's device. It is provided through a dedicated application or a web browser.
[0221] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0222] This invention is a system that makes it easy for users to recreate professional recipes at home, and also recognizes the user's emotions and provides feedback and advice accordingly. This system synthesizes a video of the user cooking with a professional recipe video and analyzes the differences to generate specific feedback and advice, and also adjusts the content and tone of this feedback and advice according to the user's emotions.
[0223] System Overview
[0224] This system consists of a user's terminal, a server for processing, and a network for exchanging information between them. It also includes an emotion engine, which adds the ability to recognize the user's emotions. The system's program is described in detail below.
[0225] User device functions
[0226] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[0227] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[0228] 3. Emotion recognition function: Equipped with a function to analyze the user's facial expressions, tone of voice, words, etc. This allows the user's emotions to be recognized in real time.
[0229] Server Features
[0230] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[0231] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[0232] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[0233] 4. Difference analysis function: The function compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[0234] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as reducing the amount of salt or extending the cooking time.
[0235] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state in real time and adjusts the tone and content of feedback and advice accordingly. For example, if the user is feeling stressed, the system may add words of encouragement.
[0236] 7. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[0237] Overview of operation
[0238] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with that of professional recipe videos, and differences are analyzed. Based on the analysis results, personalized feedback and advice is generated and inserted into the composite video. In addition, the system monitors the user's emotional state in real time and adjusts the tone and content of the feedback as needed. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[0239] Specific examples
[0240] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. The analysis identifies discrepancies, such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it adds encouraging messages and advice for relaxation. These pieces of advice are combined with a professional recipe video and provided to the user. Users can watch the combined video, compare their own movements with those of the professional, understand the necessary improvements, and incorporate them into their next cooking experience.
[0241] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills.Furthermore, by providing feedback based on the user's emotions, the system achieves a more comfortable and effective learning experience.
[0242] The processing flow will be explained below.
[0243] Step 1:
[0244] The user films themselves cooking at home using a smartphone or tablet camera, and their facial expressions and tone of voice are also recorded during the filming.
[0245] Step 2:
[0246] Cooking videos taken by users are uploaded to the system via a dedicated application or website.
[0247] Step 3:
[0248] The device receives the uploaded cooking video and sends it to the server.
[0249] Step 4:
[0250] The server stores the received cooking video in a database and begins video analysis.
[0251] Step 5:
[0252] The server analyzes each frame of the video and extracts information about the cooking steps, utensils, and ingredients used. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are analyzed.
[0253] Step 6:
[0254] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[0255] Step 7:
[0256] The server performs similar video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[0257] Step 8:
[0258] The server compares the user's cooking video with a professional recipe video and analyzes the differences in specific cooking steps, utensils, and ingredients. For example, it identifies differences such as "the user added too much salt" or "the heating time was too short."
[0259] Step 9:
[0260] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "reduce the amount of salt to one teaspoon" or "shorten the cooking time from 10 minutes to 5 minutes."
[0261] Step 10:
[0262] The server uses an emotion engine to monitor the user's emotional state in real time, analyzing their facial expressions, tone of voice, and words to determine whether they are feeling stressed.
[0263] Step 11:
[0264] The server adjusts the content and tone of the feedback and advice depending on the user's emotions. For example, if the user is feeling stressed, it adds an encouraging message.
[0265] Step 12:
[0266] The server generates a personalized composite video for each user, which displays a professional recipe video alongside the user's cooking video, and inserts feedback and emotional advice.
[0267] Step 13:
[0268] The server sends the generated composite video to the terminal.
[0269] Step 14:
[0270] The device then displays the synthesized video to the user. The user watches the video, understands the differences between the professional's movements and their own, and can reflect this in their next cooking experience. Receiving feedback based on their emotions makes learning more stress-free and effective.
[0271] Example 2
[0272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0273] Conventional cooking assistance systems have made it difficult for users to accurately recreate professional recipes at home. Furthermore, they generally provide feedback and advice without taking into account the user's individual cooking situation or emotions, making it difficult to provide effective support. Furthermore, there was a lack of a way to accurately analyze the differences between cooking videos and professional recipe videos and provide users with specific areas for improvement. This made it difficult for users to obtain specific guidance on how to improve their cooking skills.
[0274] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0275] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, means for generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice, means for providing the composite video to the user, and means for recognizing the user's emotions in real time and adjusting the content and tone of the feedback and advice. This makes it easier for users to accurately recreate professional recipes at home and to receive specific feedback and advice tailored to their individual cooking situations and emotions.
[0276] A "user" is an individual who cooks at home and inputs the cooking process into the system.
[0277] A "cooking video" is a video file that records the cooking process filmed by a user.
[0278] A "professional recipe video" is a video file in which a professional chef demonstrates cooking steps.
[0279] "Means for uploading" refers to the functions and interfaces that allow users to send cooking videos they have filmed to a server.
[0280] "Means of import" refers to the functionality and interface that allows professional recipe videos to be used within the system.
[0281] "Means of analysis" refers to algorithms and software that analyze cooking and recipe videos and extract the cooking steps, utensils, and ingredients for each.
[0282] "Comparison means" refers to functions and algorithms for comparing extracted cooking steps, utensils, and ingredients between cooking videos and recipe videos.
[0283] "Means for analyzing differences" refers to functions or algorithms for identifying and analyzing the differences between cooking videos and recipe videos.
[0284] The "means for generating feedback and advice" refers to a function or algorithm for generating specific feedback and improvement suggestions for the user based on the results of the differential analysis.
[0285] "Means for generating composite videos" refers to functions and algorithms that align professional recipe videos with users' cooking videos to create videos that include feedback and advice based on the analysis results.
[0286] The "means for providing" refers to the functions and interfaces that allow users to view the generated composite video.
[0287] "Means for recognizing emotions" refers to functions and algorithms that analyze a user's facial expressions, tone of voice, words, etc. to identify the user's emotional state in real time.
[0288] "Means for adjusting the content and tone of feedback and advice" refers to functions and algorithms for changing the content and expression of feedback and advice depending on the user's emotional state.
[0289] This invention is a system that allows users to easily recreate professional recipes when cooking at home, and also recognizes the user's emotions in real time and provides feedback and advice. This system is composed of a user terminal, a server, and a network connecting these. A specific embodiment of this system will be described below.
[0290] User device functions
[0291] 1. Shooting function
[0292] When cooking at home, users use the camera function of their smartphones or tablets to record the cooking process, and these video files can then be analyzed.
[0293] 2. Upload function
[0294] Users upload their cooking videos to the server via a dedicated application or web browser. Uploading requires an internet connection and may take some time depending on the size of the video.
[0295] 3. Emotion recognition function
[0296] This function analyzes the user's facial expressions, tone of voice, words, etc. in real time to recognize their emotions, making it possible to determine the level of stress or joy the user is feeling.
[0297] Server Features
[0298] 1. Video reception and storage function
[0299] The server receives and stores cooking videos uploaded by users. Metadata (e.g., shooting date and time, user ID) is also stored in the video file, making it easier to analyze and search later.
[0300] 2. Video analysis function
[0301] The server analyzes the stored cooking videos frame by frame using computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow). The analysis extracts cooking steps, utensils, and ingredients, and records them in a database.
[0302] 3. Differential analysis function
[0303] The server compares the user's cooking video with professional recipe videos to identify differences in cooking procedures, utensils, and ingredients. This analysis uses comparison algorithms and natural language processing techniques. The identified differences are stored in a database.
[0304] 4. Feedback generation function
[0305] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, sometimes using a generative AI model (e.g., GPT-3) to automatically generate feedback with specific improvements.
[0306] 5. Emotion-adaptive feedback function
[0307] The server monitors the user's emotional state in real time and adjusts the content and tone of the feedback and advice it provides, for example adding an encouraging message if the user is feeling stressed.
[0308] 6. Composite video generation function
[0309] The system displays professional recipe videos and user cooking videos side by side, and generates a composite video incorporating feedback based on the analysis results. Comments about the user's feelings may also be added to this composite video.
[0310] Specific examples
[0311] For example, suppose a user takes a video of themselves making carbonara and uploads it to a server. The server receives the video and begins analyzing it. The analysis results identify the following differences:
[0312] Bacon cut differently
[0313] Mixing the eggs too late
[0314] A lot of salt
[0315] Based on this, the following specific advice is generated:
[0316] Bacon should be thinly sliced
[0317] Mix the eggs quickly
[0318] Reduce the amount of salt by half
[0319] Additionally, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it will add encouraging messages and advice on how to relax.
[0320] Prompt Sentence Examples
[0321] "I'd like specific feedback on my steps in making this carbonara. I'd especially like some advice on when to mix the eggs and how much salt to use."
[0322] In this way, users can watch the composite video, compare their own movements with those of a professional, understand what improvements are needed, and improve their cooking skills at home.
[0323] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0324] Step 1:
[0325] A user shoots a cooking video.
[0326] Specific operation: The user uses the camera on their smartphone or tablet to film themselves cooking at home, adjusting the camera angle and lighting to ensure a clear video.
[0327] Input: The user's cooking scene.
[0328] Output: Cooking video file.
[0329] Step 2:
[0330] The device uploads the cooking video to the server.
[0331] Specific operation: The user instructs the device to upload the video they have taken to the server using a dedicated application or web browser. The device then transfers the video file to the server via the Internet.
[0332] Input: Filmed cooking video files.
[0333] Output: Cooking video files uploaded to the server.
[0334] Step 3:
[0335] The server receives and stores the cooking video.
[0336] Specific operation: The server receives and stores the uploaded video file, along with metadata (e.g., shooting date and time, user ID).
[0337] Input: Uploaded cooking video file and metadata.
[0338] Output: Saved cooking video files and metadata.
[0339] Step 4:
[0340] The server analyzes the video.
[0341] How it works: The server analyzes the stored cooking videos frame by frame, and uses computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow) to extract cooking steps, utensils, and ingredients.
[0342] Input: Saved cooking video file.
[0343] Output: Extracted cooking steps, utensils, and ingredients data.
[0344] Step 5:
[0345] The server analyzes the differences with professional recipe videos.
[0346] Specific operation: The server compares the user's cooking video with professional recipe videos and uses comparison algorithms and natural language processing techniques to identify differences in cooking steps, utensils used, and ingredients.
[0347] Input: Extracted cooking instructions, utensils, and ingredient data; professional recipe videos.
[0348] Output: Differential analysis results.
[0349] Step 6:
[0350] The server generates feedback and advice.
[0351] Specific operation: The server generates feedback and advice for the user based on the results of the differential analysis. It uses a generative AI model (e.g., GPT-3) to automatically create feedback including specific improvements.
[0352] Input: Differential analysis results.
[0353] Output: Generated feedback and advice.
[0354] Step 7:
[0355] The server adjusts according to the user's emotional state.
[0356] Specific operation: The server uses emotion recognition to analyze the user's facial expressions, tone of voice, and words to recognize their emotions. It then adjusts the content and tone of the feedback and advice provided according to their emotional state.
[0357] Input: User facial expressions, tone of voice, and verbal data.
[0358] Output: Tailored feedback and advice.
[0359] Step 8:
[0360] The server generates the composite video.
[0361] How it works: The server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video that includes feedback and advice based on the analysis results. It may also add comments about the user's feelings to the video.
[0362] Inputs: Tailored feedback and advice; professional recipe videos; user cooking videos.
[0363] Output: The generated composite video.
[0364] Step 9:
[0365] The user watches the composite video.
[0366] Specific operation: The user watches a composite video provided by the server on their device. While watching this composite video, they compare their own cooking with that of a professional and understand what improvements are needed.
[0367] Input: The generated synthetic video.
[0368] Output: Improved cooking skills of the user.
[0369] (Application example 2)
[0370] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0371] Conventional cooking feedback systems aim to improve users' cooking skills, but they lack functionality that takes users' emotions into consideration, which often causes stress and makes it difficult for users to continue using the system. Furthermore, they lack specific suggestions for improvements when reproducing professional recipes, making it difficult for users to effectively improve their own cooking skills.
[0372] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: a means for uploading cooking videos filmed by the user; a means for importing professional recipe videos; a means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients; a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences; a means for generating feedback and advice for the user based on the differences; a means for recognizing the user's emotions and adjusting the content and tone of the feedback and advice according to the emotions; a means for generating a composite video including the generated feedback and advice; and a means for providing the composite video to the user. This makes it possible to provide less stressful feedback according to the user's emotional state and present specific improvements.
[0373] A "cooking video filmed by a user" is a video recording of a user cooking using a device such as a smartphone or tablet.
[0374] "Professional recipe videos" are videos containing cooking steps shown by professional chefs or cooking experts, and are intended for users to use as reference.
[0375] "Means for analyzing cooking videos and recipe videos and extracting the respective cooking steps, utensils, and ingredients" refers to a technical means for analyzing uploaded cooking videos and recipe videos and automatically recognizing and extracting the cooking steps, utensils used, and ingredients from each frame.
[0376] The "means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences" refers to a technical means for comparing a user's cooking video with a professional recipe video and identifying differences in cooking procedures, utensils, and ingredients.
[0377] The "means for generating feedback and advice for the user based on the difference" is a technology for automatically generating feedback and specific advice that suggests improvements to the user's cooking method based on the results of the difference analysis.
[0378] "Means for recognizing a user's emotions and adjusting the content and tone of feedback or advice according to those emotions" refers to a technical means for analyzing a user's facial expressions, tone of voice, words, etc. to grasp their emotional state, and adjusting the content and tone of the feedback or advice provided according to those emotions.
[0379] The "means for generating a composite video including generated feedback and advice" is a technology for automatically generating a composite video that inserts comparison results and feedback based on a user's cooking video and a professional recipe video.
[0380] The "means for providing a composite video to a user" refers to a technical means for transmitting the generated composite video to a user's terminal so that the user can view the video.
[0381] This invention is a system that makes it easy for users to recreate professional recipes at home. It recognizes the user's emotions and provides feedback and advice according to those emotions. The system analyzes and synthesizes the user's cooking video and the professional recipe video, and identifies and provides specific differences.
[0382] System configuration
[0383] This system consists of user terminals, servers for processing information, and a network that connects them. The specific functions of each element are explained below.
[0384] User terminal
[0385] User terminals mainly include smartphones and tablets.
[0386] 1. Recording function: Equipped with a camera function that allows users to record videos of themselves cooking at home, for example, using a smartphone camera.
[0387] 2. Upload function: An application or web browser is provided for uploading the recorded cooking videos to the server.
[0388] 3. Emotion recognition function: Using the camera and microphone, the system analyzes the user's facial expressions, tone of voice, and words in real time to recognize the user's emotions.
[0389] server
[0390] The server analyzes the data sent from the user device and generates feedback and advice. Specific functions include:
[0391] 1. Video reception and storage function: Receives cooking videos uploaded by users and stores them on the server.
[0392] 2. Video analysis function: Analyzes cooking videos and professional recipe videos frame by frame to extract information on cooking steps, utensils, and ingredients. OpenCV and dlib are used here.
[0393] 3. Database function: Equipped with a database for managing information on extracted cooking procedures, utensils, and ingredients.
[0394] 4. Difference analysis function: Compares cooking videos with professional recipe videos and analyzes differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[0395] 5. Feedback generation function: Generates specific feedback and advice for users based on differential analysis. The software used is a generative AI model.
[0396] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state and adjusts the tone and content of feedback and advice. For example, if the user is feeling stressed, the system adds advice on how to relax.
[0397] 7. Composite video generation function: Combines the user's cooking video with a professional recipe video to generate a composite video that includes feedback and advice.
[0398] Overview of operation
[0399] The user cooks, films the process with their smartphone, and uploads the video to the server via a dedicated app. The server receives the video and begins analysis. The extracted cooking steps, utensils, and ingredient information are compared with that of professional recipe videos to identify differences. Specific feedback and advice is generated based on the differences, and the system also monitors the user's emotional state and adapts the content and tone of the feedback. The professional recipe video and the user's cooking video are then combined to generate a composite video that includes feedback and advice. The composite video is finally provided to the user, who can watch it to improve their own cooking skills.
[0400] Specific examples
[0401] For example, a user can film themselves making carbonara and upload it via a dedicated app. The server receives the video and begins analyzing it, identifying discrepancies such as the way the bacon was cut, the timing of the eggs being added too late, and the amount of salt being too much. Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the user feels anxious or stressed while cooking, the emotion engine detects this and adds encouraging messages or advice on how to relax. This advice is then combined with a professional recipe video and provided to the user. The user can then watch the combined video, compare their own movements with those of the professional, and use the results to improve their cooking skills.
[0402] Prompt Sentence Examples
[0403] Users film themselves making carbonara and upload it to the app, where it will compare it with professional recipe videos and provide specific advice.
[0404] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills. It also provides emotional feedback to users, making the learning experience more comfortable and effective.
[0405] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0406] Step 1:
[0407] The user takes a photo of themselves cooking using a device (smartphone or tablet).
[0408] Input: User's cooking video (obtained in real time via the device camera)
[0409] Output: Cooking video file
[0410] Step 2:
[0411] Cooking videos taken by the device are uploaded to a server via a dedicated application.
[0412] Input: Cooking video file
[0413] Output: Cooking video data uploaded to the server
[0414] Step 3:
[0415] The server receives the uploaded cooking videos and stores them in a database.
[0416] Input: User's cooking video data
[0417] Output: Cooking video data stored in a database
[0418] Step 4:
[0419] The server retrieves professional recipe videos from a database.
[0420] Input: Professional recipe video data request
[0421] Output: Professional recipe video data extracted from the database
[0422] Step 5:
[0423] The server analyzes cooking and recipe videos frame by frame and extracts information on each cooking step, utensils used, and ingredients.
[0424] Input: Cooking video data and recipe video data
[0425] Data processing: Analyzes video frame by frame using OpenCV and dlib, automatically recognizing cooking steps, utensils, and ingredients
[0426] Output: Extracted cooking instructions, utensils, and ingredient data
[0427] Step 6:
[0428] The server compares the extracted cooking procedures, utensils, and ingredient data and analyzes the differences.
[0429] Input: User cooking instructions, equipment, and ingredient data; Professional recipe instructions, equipment, and ingredient data
[0430] Data calculation: Calculates differences in cooking procedures, utensils, and ingredients
[0431] Output: Difference data of cooking procedures, utensils, and ingredients
[0432] Step 7:
[0433] The server generates feedback and advice for the user based on the difference data.
[0434] Input: Cooking procedure, utensils, ingredient differential data
[0435] Data shaping: Using generative AI models to generate specific advice
[0436] Output: Specific feedback and advice to the user
[0437] Step 8:
[0438] The server recognizes the user's emotions and adjusts the content and tone of the feedback and advice according to the emotions.
[0439] Input: User's emotional data (facial expressions, tone of voice, and word analysis results)
[0440] Data Computation: Feedback Modulation Based on Emotion Engine
[0441] Output: Emotionally tailored feedback and advice
[0442] Step 9:
[0443] The server generates a composite video that includes the generated feedback and advice.
[0444] Input: Adjusted feedback and advice data, user cooking video data, professional recipe video data
[0445] Data processing: Combining cooking videos and recipe videos, and inserting feedback
[0446] Output: Generated composite video
[0447] Step 10:
[0448] The server provides the generated composite video to the user's terminal.
[0449] Input: Synthetic video data
[0450] Output: Composite video sent to the user's device
[0451] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0452] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0453] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0454] [Second embodiment]
[0455] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0456] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0457] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0458] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0459] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0460] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0461] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0462] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0463] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0464] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0465] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0466] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0467] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0468] System Overview
[0469] This system consists of a user terminal, a server for processing, and a network for exchanging information between them. The system programs are described in detail below.
[0470] User device functions
[0471] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[0472] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[0473] Server Features
[0474] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[0475] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[0476] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[0477] 4. Difference analysis function: The app has a function to compare the user's cooking video with a professional recipe video and analyze the differences. Specifically, it identifies differences in cooking procedures, utensils, and ingredients.
[0478] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as "the amount of salt is too much" or "shorten the heating time."
[0479] 6. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[0480] Overview of operation
[0481] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with information from professional recipe videos, and differences are analyzed. Based on the analysis results, feedback and advice tailored to the user is generated and inserted into the composite video. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[0482] Specific examples
[0483] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. As a result of the analysis, it identifies differences such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, it generates specific advice such as "slicing the bacon thinner," "mixing the eggs early," and "reducing the amount of salt by half." This advice is then combined with a professional recipe video and provided to the user. The user watches the combined video and understands how to make a carbonara that is closer to what a professional would make.
[0484] In this way, the system makes it easy for users to recreate professional recipes at home and helps improve their cooking skills.
[0485] The processing flow will be explained below.
[0486] Step 1:
[0487] Users film themselves cooking at home, using the camera function of their smartphones or tablets to record the entire cooking process.
[0488] Step 2:
[0489] Cooking videos taken by users are uploaded to the system using a dedicated application or website.
[0490] Step 3:
[0491] The device receives the uploaded cooking video and sends it to the server.
[0492] Step 4:
[0493] The server stores the received cooking video in a database and begins analyzing it.
[0494] Step 5:
[0495] The server analyzes the video frame by frame and extracts information about cooking steps, utensils, and ingredients. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are extracted.
[0496] Step 6:
[0497] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[0498] Step 7:
[0499] The server also performs video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[0500] Step 8:
[0501] The server compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, specific differences such as "too much salt" or "too short heating time" can be identified.
[0502] Step 9:
[0503] Based on the analysis results, the server generates specific feedback and advice for the user, such as "reduce the amount of salt" or "extend the cooking time."
[0504] Step 10:
[0505] The server generates a composite video for each user, which displays a professional recipe video and the user's cooking video side by side, and inserts feedback and advice based on the differences.
[0506] Step 11:
[0507] The server sends the generated composite video to the terminal.
[0508] Step 12:
[0509] The device then displays the synthesized video to the user, who can then watch the video to understand the differences between the professional's movements and their own, and use the feedback to improve their cooking procedures and techniques.
[0510] Example 1
[0511] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0512] With conventional cooking methods, it was difficult for users to recreate professional recipes, and there were limited ways to receive specific feedback and advice. Furthermore, there was no system that analyzed the user's cooking process in detail and provided specific suggestions for improvement. This meant that users were unable to faithfully recreate professional recipes, making it difficult to efficiently improve their cooking skills.
[0513] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0514] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and the recipe videos using a video analysis library and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences using a deep learning model, means for automatically generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice using video editing software, and means for providing the composite video to the user. This makes it easier for users to recreate professional recipes at home and allows them to efficiently improve their cooking skills by receiving specific feedback and advice.
[0515] 1. "Cooking videos filmed by users" refers to video data that records the process of a user cooking at home.
[0516] 2. "Professional recipe video" refers to video data that records the process of a professional chef preparing a specific dish.
[0517] 3. "Video Analysis Library" means a software library for analyzing video data frame by frame and extracting specific information from each frame.
[0518] 4. A "cooking procedure" is a series of steps for preparing a dish, including the specific tasks and order of the steps.
[0519] 5. "Utensils" are tools and equipment used in cooking, including knives and pots.
[0520] 6. "Ingredients" are the ingredients and seasonings used in making a dish.
[0521] 7. A "deep learning model" is a machine learning algorithm that performs highly accurate analysis and predictions by learning features from large amounts of data.
[0522] 8. "Differences" are the differences between the user's cooking procedures, equipment, and ingredients and those of professional recipe videos.
[0523] 9. "Feedback" means evaluation, comments, or advice provided to the user based on the analysis results.
[0524] 10. "Advice" means specific instructions or advice provided to a user to improve their cooking skills.
[0525] 11. "Synthetic video" is video data that displays a professional recipe video and a user's cooking video side by side, with feedback and advice based on the differences inserted.
[0526] 12. "Video editing software" means software for combining and editing multiple video data to generate new video.
[0527] 13. "Uploading means" refers to a function or device that allows users to send cooking videos they have filmed to the server.
[0528] 14. "Means for capturing" means a function or device that allows the server to receive and store professional recipe videos.
[0529] 15. "Means for providing" refers to a function or device for transmitting the generated composite video to a user so that the video can be viewed.
[0530] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0531] This system consists of a user terminal, a server, and a network that exchanges information between them. Each element of the system is explained in detail below.
[0532] User device functions
[0533] Shooting Function:
[0534] Users use the camera on their smartphone or tablet to film themselves cooking at home. To do so, they launch the device's "camera" app and press the record button to record the entire cooking process from start to finish.
[0535] Upload function:
[0536] The user uploads the video they have taken to the server through a dedicated application (e.g., "CookingPro"). To do this, open the application, go to the upload screen, and select the saved video file.
[0537] Server Features
[0538] Video receiving and storage function:
[0539] The server receives cooking videos uploaded by users and stores them in a storage service (e.g., Amazon S3).
[0540] Video analysis features:
[0541] The server uses software such as OpenCV and ffmpeg to analyze the received video frame by frame, extracting information about the cooking steps, utensils used, and ingredients.
[0542] Database features:
[0543] The server stores the cooking procedures, utensils, and ingredient information extracted through the analysis in a database (e.g., MySQL or PostgreSQL).
[0544] Differential analysis features:
[0545] The server compares the user's cooking video with professional recipe videos using a deep learning model (e.g., TensorFlow or PyTorch) to analyze differences in cooking procedures, utensils used, and ingredients.
[0546] Feedback generation features:
[0547] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "cut the bacon thinner," "add the eggs earlier," or "reduce the amount of salt by half."
[0548] Composite video generation function:
[0549] The server displays the professional recipe video and the user's cooking video side by side, creating a composite video that includes feedback. The composite is created using video editing software (e.g., Adobe Premiere Pro or MoviePy).
[0550] Features provided:
[0551] The server provides the generated composite video to the user via a communication method such as an application or email.
[0552] Specific examples
[0553] For example, a user can film a video of themselves cooking "Carbonara" and upload it to the server via the "CookingPro" app. The server receives the video and begins analyzing it. OpenCV and ffmpeg are used for the analysis, and information such as "how to cut the bacon," "when to mix the eggs," and "how much salt" is extracted from each frame. The information obtained is then stored in a database.
[0554] The server then compares the video with a professional "Carbonara" recipe video and analyzes differences such as the bacon being too thick, the eggs being added too late, or the amount of salt being too high. This is done using TensorFlow. Based on the analysis results, feedback is generated, such as "slicing the bacon thinner," "adding the eggs earlier," or "reducing the amount of salt by half." Adobe Premiere Pro is then used to generate a composite video that displays the professional recipe video and the user's cooking video side by side. Users can watch this composite video and understand specific areas for improvement, allowing them to improve their cooking skills.
[0555] Prompt Sentence Examples
[0556] The system analyzes the cooking video of carbonara filmed by the user, such as how to cut the bacon, when to mix the eggs, and the amount of salt, and provides specific feedback to create a composite video.
[0557] In this way, this system makes it easier for users to recreate professional recipes at home and helps improve their cooking skills.
[0558] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0559] Step 1:
[0560] A user shoots a cooking video.
[0561] Specifically, the user launches the camera app on their smartphone or tablet and presses the record button to record the cooking process. Taking detailed photos, including the overall image of the food being cooked and the movements of the hands, improves the accuracy of the analysis.
[0562] The input is a real scene of the cooking process, and the output is a recorded video file.
[0563] Step 2:
[0564] Cooking videos taken by users are uploaded from their terminals to a server.
[0565] Specifically, the user opens a dedicated application (e.g., "CookingPro"), goes to the upload screen, selects the saved video file, and presses the send button.
[0566] The input is a cooking video file stored on the user's terminal, and the output is a video file sent to the server.
[0567] Step 3:
[0568] The server receives and stores the video.
[0569] Specifically, the server receives the uploaded video data, stores it in a temporary directory, and then transfers it to cloud storage (e.g., Amazon S3).
[0570] The input is a video file uploaded by a user, and the output is a video file saved in the server's storage.
[0571] Step 4:
[0572] The server begins analyzing the video.
[0573] Specifically, the server uses OpenCV and ffmpeg to divide the received video into frames and extract information about the cooking steps, utensils, and ingredients from each frame. For example, it performs image recognition on each frame to identify ingredients and cooking utensils.
[0574] The input is a stored video file, and the output is extracted cooking instructions, utensils, and ingredient information for each frame.
[0575] Step 5:
[0576] The server stores information about cooking procedures, equipment, and ingredients in a database.
[0577] Specifically, the server stores the information extracted by the analysis in a database (e.g., MySQL or PostgreSQL) and manages the information in a structured manner.
[0578] The input is the information on cooking procedures, utensils, and ingredients obtained as a result of the analysis, and the output is a database in which this information is stored.
[0579] Step 6:
[0580] The server compares it with professional recipe videos and analyzes the differences.
[0581] Specifically, the server uses deep learning models (e.g., TensorFlow or PyTorch) to analyze and compare the user's cooking video with professional recipe videos, identifying differences in cooking procedures, utensils, and ingredients, for example, detecting discrepancies in the timeline and differences in procedures.
[0582] The input is the analysis results of the user's cooking video and the analysis results of the professional recipe video, and the output is the identified difference information.
[0583] Step 7:
[0584] The server generates specific feedback based on the differences.
[0585] Specifically, the server uses a generative AI model to automatically generate specific feedback and advice for the user based on the difference information, such as "cut the bacon thinner," "add the eggs early," or "reduce the amount of salt."
[0586] The input is the difference information and the output is the generated feedback and advice.
[0587] Step 8:
[0588] The server generates a composite video including the feedback and provides it to the user.
[0589] Specifically, the server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video with feedback inserted using video editing software (e.g., Adobe Premiere Pro or MoviePy). The generated composite video is then provided to the user's device or via email.
[0590] The inputs are professional recipe videos, user cooking videos, and generated feedback, and the output is a composite video file.
[0591] These processing steps make it easier for users to recreate professional recipes at home and improve their cooking skills by receiving specific feedback and advice.
[0592] (Application example 1)
[0593] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0594] Conventional cooking assistance systems lack the ability to provide real-time feedback when users are trying to recreate professional recipes. This makes it difficult for users to notice problems while cooking, resulting in less-than-perfect dishes. Furthermore, there is a lack of a way to specifically analyze differences in cooking procedures and ingredient usage, and provide the results in an easy-to-understand manner.
[0595] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0596] In this invention, the server includes a means for uploading cooking videos filmed by users, a means for importing cooking recipe videos, a means for analyzing the cooking videos and the recipe videos and extracting their respective cooking steps, utensils, and ingredients, a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, a means for generating feedback and advice for the user based on the differences, a means for displaying the generated feedback and advice in real time, a means for generating a composite video including the feedback and advice, and a means for providing the composite video to the user. This allows users to receive feedback in real time while cooking, making it easier to identify specific areas for improvement. Furthermore, the differences between professional recipe videos and the user's cooking video are clearly presented, allowing users to easily recreate recipes.
[0597] "Cooking videos filmed by users" are videos in which users record the cooking process using a home camera or smart device.
[0598] A "cooking recipe video" is a video in which a professional or expert demonstrates a cooking method while explaining the steps.
[0599] A "means" is a method or mechanism for achieving a specific function or purpose.
[0600] "Analysis" means analyzing the video content frame by frame and extracting information such as each cooking step, utensils used, and ingredients.
[0601] A "cooking procedure" is a series of steps or methods for preparing a dish.
[0602] "Utensils" are tools and machines used in cooking.
[0603] "Ingredients" are the ingredients or components used to make a dish.
[0604] "Differences" are differences or discrepancies between the user's cooking procedures, utensils used, and ingredients and the content of professional recipe videos.
[0605] "Feedback" refers to information or advice provided to the user based on the analysis results.
[0606] "Advice" refers to suggestions to the user for specific improvements, such as cooking methods or how to use ingredients.
[0607] "Real-time display" means that users receive instant feedback and advice while cooking.
[0608] A "composite video" is a video that combines a professional recipe video with a user's cooking video, and inserts feedback and advice based on the differences.
[0609] This invention is a system that makes it easy for users to recreate professional recipes when cooking at home. The main components of the system include a user terminal and a server. The system programs and their processing are described in detail below.
[0610] First, the user uses a device such as a smartphone or smart glasses to record the cooking process. The video is then uploaded to a server via a dedicated application. This application has both recording and uploading functions, providing an environment in which users can easily share videos.
[0611] The server receives both cooking videos and cooking recipe videos uploaded by users and has the analysis function to analyze each. For the analysis, OpenCV (a library for video analysis) and TensorFlow (a library for using machine learning models) are used. Specifically, the video is analyzed frame by frame to extract information on cooking steps, utensils used, and ingredients.
[0612] The extracted information is compared with the content of the cooking recipe video. This comparison uses differential analysis to identify the differences between the user's cooking and the professional recipe. The results of the differential analysis generate specific feedback and advice for the user. This feedback and advice is displayed in real time on the user's device.
[0613] In addition, the system synthesizes professional recipe videos with users' cooking videos, generating a composite video that includes feedback and advice based on the differences, making it easier for users to identify specific areas for improvement while cooking.
[0614] For example, a user may film and upload a video of themselves making beef stew. In this case, the server analyzes the video and identifies differences such as "the meat is cut too large" and "the amount of wine is too small." Based on this, specific advice such as "cut the meat into bite-sized pieces" and "add more wine" is generated.
[0615] An example of a prompt for a generative AI model is:
[0616] "Compare the professional recipe video for making beef stew with the user cooking video and provide the following differences and feedback:
[0617] 1. Size of meat cut
[0618] 2. Amount of wine used
[0619] This system makes it easy for users to faithfully recreate professional recipes at home, improving their cooking skills.
[0620] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0621] Step 1:
[0622] The user takes a video of the cooking process using a smartphone or smart glasses camera. The input at this stage is the cooking action and scene. The video is captured and saved as a video file. The output is a cooking video file.
[0623] Step 2:
[0624] The user uploads the cooking video they have filmed to the server using a dedicated application. At this stage, the input is the cooking video file, and the output is a notification that the data has been transferred to the server.
[0625] Step 3:
[0626] The server saves the cooking video received from the user. The input at this stage is the uploaded cooking video file, and the output is the destination directory and file path. Here, the file system is used to save the video in the appropriate directory.
[0627] Step 4:
[0628] The server analyzes stored cooking videos and professional cooking recipe videos. The inputs are cooking video files and professional recipe video files. The videos are decomposed frame by frame, and the OpenCV library is used to extract information about cooking steps, utensils, and ingredients. The output is a list of cooking steps, utensils, and ingredients.
[0629] Step 5:
[0630] The server compares the user's cooking procedures, utensils, and ingredients with the professional recipes based on the extracted information. This comparison uses a difference analysis algorithm. The input is the list of information extracted in step 4 and the professional recipe information list. The output is the specific differences between the user's cooking and the recipes.
[0631] Step 6:
[0632] The server generates specific feedback and advice for the user based on the results of the difference analysis. The input at this stage is the difference information. The generative AI model is used to generate advice. The output is a feedback message and a list of specific advice.
[0633] Step 7:
[0634] The server displays the generated feedback and advice on the user's terminal in real time. The input at this stage is a feedback / advice list, and the output is the feedback / advice displayed on the user's terminal screen.
[0635] Step 8:
[0636] The server synthesizes professional recipe videos and user cooking videos, generating a composite video with feedback and advice based on the differences. The inputs are cooking video files, recipe video files, and a feedback and advice list. The composite video is created using video editing software. The output is a composite video file.
[0637] Step 9:
[0638] The server provides the generated composite video to the user. The input at this stage is the composite video file, and the output is the composite video file sent to the user's device. It is provided through a dedicated application or a web browser.
[0639] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0640] This invention is a system that makes it easy for users to recreate professional recipes at home, and also recognizes the user's emotions and provides feedback and advice accordingly. This system synthesizes a video of the user cooking with a professional recipe video and analyzes the differences to generate specific feedback and advice, and also adjusts the content and tone of this feedback and advice according to the user's emotions.
[0641] System Overview
[0642] This system consists of a user's terminal, a server for processing, and a network for exchanging information between them. It also includes an emotion engine, which adds the ability to recognize the user's emotions. The system's program is described in detail below.
[0643] User device functions
[0644] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[0645] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[0646] 3. Emotion recognition function: Equipped with a function to analyze the user's facial expressions, tone of voice, words, etc. This allows the user's emotions to be recognized in real time.
[0647] Server Features
[0648] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[0649] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[0650] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[0651] 4. Difference analysis function: The function compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[0652] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as reducing the amount of salt or extending the cooking time.
[0653] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state in real time and adjusts the tone and content of feedback and advice accordingly. For example, if the user is feeling stressed, the system may add words of encouragement.
[0654] 7. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[0655] Overview of operation
[0656] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with that of professional recipe videos, and differences are analyzed. Based on the analysis results, personalized feedback and advice is generated and inserted into the composite video. In addition, the system monitors the user's emotional state in real time and adjusts the tone and content of the feedback as needed. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[0657] Specific examples
[0658] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. The analysis identifies discrepancies, such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it adds encouraging messages and advice for relaxation. These pieces of advice are combined with a professional recipe video and provided to the user. Users can watch the combined video, compare their own movements with those of the professional, understand the necessary improvements, and incorporate them into their next cooking experience.
[0659] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills.Furthermore, by providing feedback based on the user's emotions, the system achieves a more comfortable and effective learning experience.
[0660] The processing flow will be explained below.
[0661] Step 1:
[0662] The user films themselves cooking at home using a smartphone or tablet camera, and their facial expressions and tone of voice are also recorded during the filming.
[0663] Step 2:
[0664] Cooking videos taken by users are uploaded to the system via a dedicated application or website.
[0665] Step 3:
[0666] The device receives the uploaded cooking video and sends it to the server.
[0667] Step 4:
[0668] The server stores the received cooking video in a database and begins video analysis.
[0669] Step 5:
[0670] The server analyzes each frame of the video and extracts information about the cooking steps, utensils, and ingredients used. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are analyzed.
[0671] Step 6:
[0672] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[0673] Step 7:
[0674] The server performs similar video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[0675] Step 8:
[0676] The server compares the user's cooking video with a professional recipe video and analyzes the differences in specific cooking steps, utensils, and ingredients. For example, it identifies differences such as "the user added too much salt" or "the heating time was too short."
[0677] Step 9:
[0678] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "reduce the amount of salt to one teaspoon" or "shorten the cooking time from 10 minutes to 5 minutes."
[0679] Step 10:
[0680] The server uses an emotion engine to monitor the user's emotional state in real time, analyzing their facial expressions, tone of voice, and words to determine whether they are feeling stressed.
[0681] Step 11:
[0682] The server adjusts the content and tone of the feedback and advice depending on the user's emotions. For example, if the user is feeling stressed, it adds an encouraging message.
[0683] Step 12:
[0684] The server generates a personalized composite video for each user, which displays a professional recipe video alongside the user's cooking video, and inserts feedback and emotional advice.
[0685] Step 13:
[0686] The server sends the generated composite video to the terminal.
[0687] Step 14:
[0688] The device then displays the synthesized video to the user. The user watches the video, understands the differences between the professional's movements and their own, and can reflect this in their next cooking experience. Receiving feedback based on their emotions makes learning more stress-free and effective.
[0689] Example 2
[0690] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0691] Conventional cooking assistance systems have made it difficult for users to accurately recreate professional recipes at home. Furthermore, they generally provide feedback and advice without taking into account the user's individual cooking situation or emotions, making it difficult to provide effective support. Furthermore, there was a lack of a way to accurately analyze the differences between cooking videos and professional recipe videos and provide users with specific areas for improvement. This made it difficult for users to obtain specific guidance on how to improve their cooking skills.
[0692] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0693] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, means for generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice, means for providing the composite video to the user, and means for recognizing the user's emotions in real time and adjusting the content and tone of the feedback and advice. This makes it easier for users to accurately recreate professional recipes at home and to receive specific feedback and advice tailored to their individual cooking situations and emotions.
[0694] A "user" is an individual who cooks at home and inputs the cooking process into the system.
[0695] A "cooking video" is a video file that records the cooking process filmed by a user.
[0696] A "professional recipe video" is a video file in which a professional chef demonstrates cooking steps.
[0697] "Means for uploading" refers to the functions and interfaces that allow users to send cooking videos they have filmed to a server.
[0698] "Means of import" refers to the functionality and interface that allows professional recipe videos to be used within the system.
[0699] "Means of analysis" refers to algorithms and software that analyze cooking and recipe videos and extract the cooking steps, utensils, and ingredients for each.
[0700] "Comparison means" refers to functions and algorithms for comparing extracted cooking steps, utensils, and ingredients between cooking videos and recipe videos.
[0701] "Means for analyzing differences" refers to functions or algorithms for identifying and analyzing the differences between cooking videos and recipe videos.
[0702] The "means for generating feedback and advice" refers to a function or algorithm for generating specific feedback and improvement suggestions for the user based on the results of the differential analysis.
[0703] "Means for generating composite videos" refers to functions and algorithms that align professional recipe videos with users' cooking videos to create videos that include feedback and advice based on the analysis results.
[0704] The "means for providing" refers to the functions and interfaces that allow users to view the generated composite video.
[0705] "Means for recognizing emotions" refers to functions and algorithms that analyze a user's facial expressions, tone of voice, words, etc. to identify the user's emotional state in real time.
[0706] "Means for adjusting the content and tone of feedback and advice" refers to functions and algorithms for changing the content and expression of feedback and advice depending on the user's emotional state.
[0707] This invention is a system that allows users to easily recreate professional recipes when cooking at home, and also recognizes the user's emotions in real time and provides feedback and advice. This system is composed of a user terminal, a server, and a network connecting these. A specific embodiment of this system will be described below.
[0708] User device functions
[0709] 1. Shooting function
[0710] When cooking at home, users use the camera function of their smartphones or tablets to record the cooking process, and these video files can then be analyzed.
[0711] 2. Upload function
[0712] Users upload their cooking videos to the server via a dedicated application or web browser. Uploading requires an internet connection and may take some time depending on the size of the video.
[0713] 3. Emotion recognition function
[0714] This function analyzes the user's facial expressions, tone of voice, words, etc. in real time to recognize their emotions, making it possible to determine the level of stress or joy the user is feeling.
[0715] Server Features
[0716] 1. Video reception and storage function
[0717] The server receives and stores cooking videos uploaded by users. Metadata (e.g., shooting date and time, user ID) is also stored in the video file, making it easier to analyze and search later.
[0718] 2. Video analysis function
[0719] The server analyzes the stored cooking videos frame by frame using computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow). The analysis extracts cooking steps, utensils, and ingredients, and records them in a database.
[0720] 3. Differential analysis function
[0721] The server compares the user's cooking video with professional recipe videos to identify differences in cooking procedures, utensils, and ingredients. This analysis uses comparison algorithms and natural language processing techniques. The identified differences are stored in a database.
[0722] 4. Feedback generation function
[0723] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, sometimes using a generative AI model (e.g., GPT-3) to automatically generate feedback with specific improvements.
[0724] 5. Emotion-adaptive feedback function
[0725] The server monitors the user's emotional state in real time and adjusts the content and tone of the feedback and advice it provides, for example adding an encouraging message if the user is feeling stressed.
[0726] 6. Composite video generation function
[0727] The system displays professional recipe videos and user cooking videos side by side, and generates a composite video incorporating feedback based on the analysis results. Comments about the user's feelings may also be added to this composite video.
[0728] Specific examples
[0729] For example, suppose a user takes a video of themselves making carbonara and uploads it to a server. The server receives the video and begins analyzing it. The analysis results identify the following differences:
[0730] Bacon cut differently
[0731] Mixing the eggs too late
[0732] A lot of salt
[0733] Based on this, the following specific advice is generated:
[0734] Bacon should be thinly sliced
[0735] Mix the eggs quickly
[0736] Reduce the amount of salt by half
[0737] Additionally, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it will add encouraging messages and advice on how to relax.
[0738] Prompt Sentence Examples
[0739] "I'd like specific feedback on my steps in making this carbonara. I'd especially like some advice on when to mix the eggs and how much salt to use."
[0740] In this way, users can watch the composite video, compare their own movements with those of a professional, understand what improvements are needed, and improve their cooking skills at home.
[0741] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0742] Step 1:
[0743] A user shoots a cooking video.
[0744] Specific operation: The user uses the camera on their smartphone or tablet to film themselves cooking at home, adjusting the camera angle and lighting to ensure a clear video.
[0745] Input: The user's cooking scene.
[0746] Output: Cooking video file.
[0747] Step 2:
[0748] The device uploads the cooking video to the server.
[0749] Specific operation: The user instructs the device to upload the video they have taken to the server using a dedicated application or web browser. The device then transfers the video file to the server via the Internet.
[0750] Input: Filmed cooking video files.
[0751] Output: Cooking video files uploaded to the server.
[0752] Step 3:
[0753] The server receives and stores the cooking video.
[0754] Specific operation: The server receives and stores the uploaded video file, along with metadata (e.g., shooting date and time, user ID).
[0755] Input: Uploaded cooking video file and metadata.
[0756] Output: Saved cooking video files and metadata.
[0757] Step 4:
[0758] The server analyzes the video.
[0759] How it works: The server analyzes the stored cooking videos frame by frame, and uses computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow) to extract cooking steps, utensils, and ingredients.
[0760] Input: Saved cooking video file.
[0761] Output: Extracted cooking steps, utensils, and ingredients data.
[0762] Step 5:
[0763] The server analyzes the differences with professional recipe videos.
[0764] Specific operation: The server compares the user's cooking video with professional recipe videos and uses comparison algorithms and natural language processing techniques to identify differences in cooking steps, utensils used, and ingredients.
[0765] Input: Extracted cooking instructions, utensils, and ingredient data; professional recipe videos.
[0766] Output: Differential analysis results.
[0767] Step 6:
[0768] The server generates feedback and advice.
[0769] Specific operation: The server generates feedback and advice for the user based on the results of the differential analysis. It uses a generative AI model (e.g., GPT-3) to automatically create feedback including specific improvements.
[0770] Input: Differential analysis results.
[0771] Output: Generated feedback and advice.
[0772] Step 7:
[0773] The server adjusts according to the user's emotional state.
[0774] Specific operation: The server uses emotion recognition to analyze the user's facial expressions, tone of voice, and words to recognize their emotions. It then adjusts the content and tone of the feedback and advice provided according to their emotional state.
[0775] Input: User facial expressions, tone of voice, and verbal data.
[0776] Output: Tailored feedback and advice.
[0777] Step 8:
[0778] The server generates the composite video.
[0779] How it works: The server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video that includes feedback and advice based on the analysis results. It may also add comments about the user's feelings to the video.
[0780] Inputs: Tailored feedback and advice; professional recipe videos; user cooking videos.
[0781] Output: The generated composite video.
[0782] Step 9:
[0783] The user watches the composite video.
[0784] Specific operation: The user watches a composite video provided by the server on their device. While watching this composite video, they compare their own cooking with that of a professional and understand what improvements are needed.
[0785] Input: The generated synthetic video.
[0786] Output: Improved cooking skills of the user.
[0787] (Application example 2)
[0788] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0789] Conventional cooking feedback systems aim to improve users' cooking skills, but they lack functionality that takes users' emotions into consideration, which often causes stress and makes it difficult for users to continue using the system. Furthermore, they lack specific suggestions for improvements when reproducing professional recipes, making it difficult for users to effectively improve their own cooking skills.
[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: a means for uploading cooking videos filmed by the user; a means for importing professional recipe videos; a means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients; a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences; a means for generating feedback and advice for the user based on the differences; a means for recognizing the user's emotions and adjusting the content and tone of the feedback and advice according to the emotions; a means for generating a composite video including the generated feedback and advice; and a means for providing the composite video to the user. This makes it possible to provide less stressful feedback according to the user's emotional state and present specific improvements.
[0791] A "cooking video filmed by a user" is a video recording of a user cooking using a device such as a smartphone or tablet.
[0792] "Professional recipe videos" are videos containing cooking steps shown by professional chefs or cooking experts, and are intended for users to use as reference.
[0793] "Means for analyzing cooking videos and recipe videos and extracting the respective cooking steps, utensils, and ingredients" refers to a technical means for analyzing uploaded cooking videos and recipe videos and automatically recognizing and extracting the cooking steps, utensils used, and ingredients from each frame.
[0794] The "means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences" refers to a technical means for comparing a user's cooking video with a professional recipe video and identifying differences in cooking procedures, utensils, and ingredients.
[0795] The "means for generating feedback and advice for the user based on the difference" is a technology for automatically generating feedback and specific advice that suggests improvements to the user's cooking method based on the results of the difference analysis.
[0796] "Means for recognizing a user's emotions and adjusting the content and tone of feedback or advice according to those emotions" refers to a technical means for analyzing a user's facial expressions, tone of voice, words, etc. to grasp their emotional state, and adjusting the content and tone of the feedback or advice provided according to those emotions.
[0797] The "means for generating a composite video including generated feedback and advice" is a technology for automatically generating a composite video that inserts comparison results and feedback based on a user's cooking video and a professional recipe video.
[0798] The "means for providing a composite video to a user" refers to a technical means for transmitting the generated composite video to a user's terminal so that the user can view the video.
[0799] This invention is a system that makes it easy for users to recreate professional recipes at home. It recognizes the user's emotions and provides feedback and advice according to those emotions. The system analyzes and synthesizes the user's cooking video and the professional recipe video, and identifies and provides specific differences.
[0800] System configuration
[0801] This system consists of user terminals, servers for processing information, and a network that connects them. The specific functions of each element are explained below.
[0802] User terminal
[0803] User terminals mainly include smartphones and tablets.
[0804] 1. Recording function: Equipped with a camera function that allows users to record videos of themselves cooking at home, for example, using a smartphone camera.
[0805] 2. Upload function: An application or web browser is provided for uploading the recorded cooking videos to the server.
[0806] 3. Emotion recognition function: Using the camera and microphone, the system analyzes the user's facial expressions, tone of voice, and words in real time to recognize the user's emotions.
[0807] server
[0808] The server analyzes the data sent from the user device and generates feedback and advice. Specific functions include:
[0809] 1. Video reception and storage function: Receives cooking videos uploaded by users and stores them on the server.
[0810] 2. Video analysis function: Analyzes cooking videos and professional recipe videos frame by frame to extract information on cooking steps, utensils, and ingredients. OpenCV and dlib are used here.
[0811] 3. Database function: Equipped with a database for managing information on extracted cooking procedures, utensils, and ingredients.
[0812] 4. Difference analysis function: Compares cooking videos with professional recipe videos and analyzes differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[0813] 5. Feedback generation function: Generates specific feedback and advice for users based on differential analysis. The software used is a generative AI model.
[0814] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state and adjusts the tone and content of feedback and advice. For example, if the user is feeling stressed, the system adds advice on how to relax.
[0815] 7. Composite video generation function: Combines the user's cooking video with a professional recipe video to generate a composite video that includes feedback and advice.
[0816] Overview of operation
[0817] The user cooks, films the process with their smartphone, and uploads the video to the server via a dedicated app. The server receives the video and begins analysis. The extracted cooking steps, utensils, and ingredient information are compared with that of professional recipe videos to identify differences. Specific feedback and advice is generated based on the differences, and the system also monitors the user's emotional state and adapts the content and tone of the feedback. The professional recipe video and the user's cooking video are then combined to generate a composite video that includes feedback and advice. The composite video is finally provided to the user, who can watch it to improve their own cooking skills.
[0818] Specific examples
[0819] For example, a user can film themselves making carbonara and upload it via a dedicated app. The server receives the video and begins analyzing it, identifying discrepancies such as the way the bacon was cut, the timing of the eggs being added too late, and the amount of salt being too much. Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the user feels anxious or stressed while cooking, the emotion engine detects this and adds encouraging messages or advice on how to relax. This advice is then combined with a professional recipe video and provided to the user. The user can then watch the combined video, compare their own movements with those of the professional, and use the results to improve their cooking skills.
[0820] Prompt Sentence Examples
[0821] Users film themselves making carbonara and upload it to the app, where it will compare it with professional recipe videos and provide specific advice.
[0822] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills. It also provides emotional feedback to users, making the learning experience more comfortable and effective.
[0823] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0824] Step 1:
[0825] The user takes a photo of themselves cooking using a device (smartphone or tablet).
[0826] Input: User's cooking video (obtained in real time via the device camera)
[0827] Output: Cooking video file
[0828] Step 2:
[0829] Cooking videos taken by the device are uploaded to a server via a dedicated application.
[0830] Input: Cooking video file
[0831] Output: Cooking video data uploaded to the server
[0832] Step 3:
[0833] The server receives the uploaded cooking videos and stores them in a database.
[0834] Input: User's cooking video data
[0835] Output: Cooking video data stored in a database
[0836] Step 4:
[0837] The server retrieves professional recipe videos from a database.
[0838] Input: Professional recipe video data request
[0839] Output: Professional recipe video data extracted from the database
[0840] Step 5:
[0841] The server analyzes cooking and recipe videos frame by frame and extracts information on each cooking step, utensils used, and ingredients.
[0842] Input: Cooking video data and recipe video data
[0843] Data processing: Analyzes video frame by frame using OpenCV and dlib, automatically recognizing cooking steps, utensils, and ingredients
[0844] Output: Extracted cooking instructions, utensils, and ingredient data
[0845] Step 6:
[0846] The server compares the extracted cooking procedures, utensils, and ingredient data and analyzes the differences.
[0847] Input: User cooking instructions, equipment, and ingredient data; Professional recipe instructions, equipment, and ingredient data
[0848] Data calculation: Calculates differences in cooking procedures, utensils, and ingredients
[0849] Output: Difference data of cooking procedures, utensils, and ingredients
[0850] Step 7:
[0851] The server generates feedback and advice for the user based on the difference data.
[0852] Input: Cooking procedure, utensils, ingredient differential data
[0853] Data shaping: Using generative AI models to generate specific advice
[0854] Output: Specific feedback and advice to the user
[0855] Step 8:
[0856] The server recognizes the user's emotions and adjusts the content and tone of the feedback and advice according to the emotions.
[0857] Input: User's emotional data (facial expressions, tone of voice, and word analysis results)
[0858] Data Computation: Feedback Modulation Based on Emotion Engine
[0859] Output: Emotionally tailored feedback and advice
[0860] Step 9:
[0861] The server generates a composite video that includes the generated feedback and advice.
[0862] Input: Adjusted feedback and advice data, user cooking video data, professional recipe video data
[0863] Data processing: Combining cooking videos and recipe videos, and inserting feedback
[0864] Output: Generated composite video
[0865] Step 10:
[0866] The server provides the generated composite video to the user's terminal.
[0867] Input: Synthetic video data
[0868] Output: Composite video sent to the user's device
[0869] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0870] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0871] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0872] [Third embodiment]
[0873] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0874] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0875] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0876] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0877] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0878] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0879] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0880] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0881] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0882] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0883] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0884] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0885] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0886] System Overview
[0887] This system consists of a user terminal, a server for processing, and a network for exchanging information between them. The system programs are described in detail below.
[0888] User device functions
[0889] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[0890] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[0891] Server Features
[0892] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[0893] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[0894] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[0895] 4. Difference analysis function: The app has a function to compare the user's cooking video with a professional recipe video and analyze the differences. Specifically, it identifies differences in cooking procedures, utensils, and ingredients.
[0896] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as "the amount of salt is too much" or "shorten the heating time."
[0897] 6. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[0898] Overview of operation
[0899] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with information from professional recipe videos, and differences are analyzed. Based on the analysis results, feedback and advice tailored to the user is generated and inserted into the composite video. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[0900] Specific examples
[0901] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. As a result of the analysis, it identifies differences such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, it generates specific advice such as "slicing the bacon thinner," "mixing the eggs early," and "reducing the amount of salt by half." This advice is then combined with a professional recipe video and provided to the user. The user watches the combined video and understands how to make a carbonara that is closer to what a professional would make.
[0902] In this way, the system makes it easy for users to recreate professional recipes at home and helps improve their cooking skills.
[0903] The processing flow will be explained below.
[0904] Step 1:
[0905] Users film themselves cooking at home, using the camera function of their smartphones or tablets to record the entire cooking process.
[0906] Step 2:
[0907] Cooking videos taken by users are uploaded to the system using a dedicated application or website.
[0908] Step 3:
[0909] The device receives the uploaded cooking video and sends it to the server.
[0910] Step 4:
[0911] The server stores the received cooking video in a database and begins analyzing it.
[0912] Step 5:
[0913] The server analyzes the video frame by frame and extracts information about cooking steps, utensils, and ingredients. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are extracted.
[0914] Step 6:
[0915] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[0916] Step 7:
[0917] The server also performs video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[0918] Step 8:
[0919] The server compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, specific differences such as "too much salt" or "too short heating time" can be identified.
[0920] Step 9:
[0921] Based on the analysis results, the server generates specific feedback and advice for the user, such as "reduce the amount of salt" or "extend the cooking time."
[0922] Step 10:
[0923] The server generates a composite video for each user, which displays a professional recipe video and the user's cooking video side by side, and inserts feedback and advice based on the differences.
[0924] Step 11:
[0925] The server sends the generated composite video to the terminal.
[0926] Step 12:
[0927] The device then displays the synthesized video to the user, who can then watch the video to understand the differences between the professional's movements and their own, and use the feedback to improve their cooking procedures and techniques.
[0928] Example 1
[0929] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0930] With conventional cooking methods, it was difficult for users to recreate professional recipes, and there were limited ways to receive specific feedback and advice. Furthermore, there was no system that analyzed the user's cooking process in detail and provided specific suggestions for improvement. This meant that users were unable to faithfully recreate professional recipes, making it difficult to efficiently improve their cooking skills.
[0931] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0932] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and the recipe videos using a video analysis library and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences using a deep learning model, means for automatically generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice using video editing software, and means for providing the composite video to the user. This makes it easier for users to recreate professional recipes at home and allows them to efficiently improve their cooking skills by receiving specific feedback and advice.
[0933] 1. "Cooking videos filmed by users" refers to video data that records the process of a user cooking at home.
[0934] 2. "Professional recipe video" refers to video data that records the process of a professional chef preparing a specific dish.
[0935] 3. "Video Analysis Library" means a software library for analyzing video data frame by frame and extracting specific information from each frame.
[0936] 4. A "cooking procedure" is a series of steps for preparing a dish, including the specific tasks and order of the steps.
[0937] 5. "Utensils" are tools and equipment used in cooking, including knives and pots.
[0938] 6. "Ingredients" are the ingredients and seasonings used in making a dish.
[0939] 7. A "deep learning model" is a machine learning algorithm that performs highly accurate analysis and predictions by learning features from large amounts of data.
[0940] 8. "Differences" are the differences between the user's cooking procedures, equipment, and ingredients and those of professional recipe videos.
[0941] 9. "Feedback" means evaluation, comments, or advice provided to the user based on the analysis results.
[0942] 10. "Advice" means specific instructions or advice provided to a user to improve their cooking skills.
[0943] 11. "Synthetic video" is video data that displays a professional recipe video and a user's cooking video side by side, with feedback and advice based on the differences inserted.
[0944] 12. "Video editing software" means software for combining and editing multiple video data to generate new video.
[0945] 13. "Uploading means" refers to a function or device that allows users to send cooking videos they have filmed to the server.
[0946] 14. "Means for capturing" means a function or device that allows the server to receive and store professional recipe videos.
[0947] 15. "Means for providing" refers to a function or device for transmitting the generated composite video to a user so that the video can be viewed.
[0948] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[0949] This system consists of a user terminal, a server, and a network that exchanges information between them. Each element of the system is explained in detail below.
[0950] User device functions
[0951] Shooting Function:
[0952] Users use the camera on their smartphone or tablet to film themselves cooking at home. To do so, they launch the device's "camera" app and press the record button to record the entire cooking process from start to finish.
[0953] Upload function:
[0954] The user uploads the video they have taken to the server through a dedicated application (e.g., "CookingPro"). To do this, open the application, go to the upload screen, and select the saved video file.
[0955] Server Features
[0956] Video receiving and storage function:
[0957] The server receives cooking videos uploaded by users and stores them in a storage service (e.g., Amazon S3).
[0958] Video analysis features:
[0959] The server uses software such as OpenCV and ffmpeg to analyze the received video frame by frame, extracting information about the cooking steps, utensils used, and ingredients.
[0960] Database features:
[0961] The server stores the cooking procedures, utensils, and ingredient information extracted through the analysis in a database (e.g., MySQL or PostgreSQL).
[0962] Differential analysis features:
[0963] The server compares the user's cooking video with professional recipe videos using a deep learning model (e.g., TensorFlow or PyTorch) to analyze differences in cooking procedures, utensils used, and ingredients.
[0964] Feedback generation features:
[0965] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "cut the bacon thinner," "add the eggs earlier," or "reduce the amount of salt by half."
[0966] Composite video generation function:
[0967] The server displays the professional recipe video and the user's cooking video side by side, creating a composite video that includes feedback. The composite is created using video editing software (e.g., Adobe Premiere Pro or MoviePy).
[0968] Features provided:
[0969] The server provides the generated composite video to the user via a communication method such as an application or email.
[0970] Specific examples
[0971] For example, a user can film a video of themselves cooking "Carbonara" and upload it to the server via the "CookingPro" app. The server receives the video and begins analyzing it. OpenCV and ffmpeg are used for the analysis, and information such as "how to cut the bacon," "when to mix the eggs," and "how much salt" is extracted from each frame. The information obtained is then stored in a database.
[0972] The server then compares the video with a professional "Carbonara" recipe video and analyzes differences such as the bacon being too thick, the eggs being added too late, or the amount of salt being too high. This is done using TensorFlow. Based on the analysis results, feedback is generated, such as "slicing the bacon thinner," "adding the eggs earlier," or "reducing the amount of salt by half." Adobe Premiere Pro is then used to generate a composite video that displays the professional recipe video and the user's cooking video side by side. Users can watch this composite video and understand specific areas for improvement, allowing them to improve their cooking skills.
[0973] Prompt Sentence Examples
[0974] The system analyzes the cooking video of carbonara filmed by the user, such as how to cut the bacon, when to mix the eggs, and the amount of salt, and provides specific feedback to create a composite video.
[0975] In this way, this system makes it easier for users to recreate professional recipes at home and helps improve their cooking skills.
[0976] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0977] Step 1:
[0978] A user shoots a cooking video.
[0979] Specifically, the user launches the camera app on their smartphone or tablet and presses the record button to record the cooking process. Taking detailed photos, including the overall image of the food being cooked and the movements of the hands, improves the accuracy of the analysis.
[0980] The input is a real scene of the cooking process, and the output is a recorded video file.
[0981] Step 2:
[0982] Cooking videos taken by users are uploaded from their terminals to a server.
[0983] Specifically, the user opens a dedicated application (e.g., "CookingPro"), goes to the upload screen, selects the saved video file, and presses the send button.
[0984] The input is a cooking video file stored on the user's terminal, and the output is a video file sent to the server.
[0985] Step 3:
[0986] The server receives and stores the video.
[0987] Specifically, the server receives the uploaded video data, stores it in a temporary directory, and then transfers it to cloud storage (e.g., Amazon S3).
[0988] The input is a video file uploaded by a user, and the output is a video file saved in the server's storage.
[0989] Step 4:
[0990] The server begins analyzing the video.
[0991] Specifically, the server uses OpenCV and ffmpeg to divide the received video into frames and extract information about the cooking steps, utensils, and ingredients from each frame. For example, it performs image recognition on each frame to identify ingredients and cooking utensils.
[0992] The input is a stored video file, and the output is extracted cooking instructions, utensils, and ingredient information for each frame.
[0993] Step 5:
[0994] The server stores information about cooking procedures, equipment, and ingredients in a database.
[0995] Specifically, the server stores the information extracted by the analysis in a database (e.g., MySQL or PostgreSQL) and manages the information in a structured manner.
[0996] The input is the information on cooking procedures, utensils, and ingredients obtained as a result of the analysis, and the output is a database in which this information is stored.
[0997] Step 6:
[0998] The server compares it with professional recipe videos and analyzes the differences.
[0999] Specifically, the server uses deep learning models (e.g., TensorFlow or PyTorch) to analyze and compare the user's cooking video with professional recipe videos, identifying differences in cooking procedures, utensils, and ingredients, for example, detecting discrepancies in the timeline and differences in procedures.
[1000] The input is the analysis results of the user's cooking video and the analysis results of the professional recipe video, and the output is the identified difference information.
[1001] Step 7:
[1002] The server generates specific feedback based on the differences.
[1003] Specifically, the server uses a generative AI model to automatically generate specific feedback and advice for the user based on the difference information, such as "cut the bacon thinner," "add the eggs early," or "reduce the amount of salt."
[1004] The input is the difference information and the output is the generated feedback and advice.
[1005] Step 8:
[1006] The server generates a composite video including the feedback and provides it to the user.
[1007] Specifically, the server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video with feedback inserted using video editing software (e.g., Adobe Premiere Pro or MoviePy). The generated composite video is then provided to the user's device or via email.
[1008] The inputs are professional recipe videos, user cooking videos, and generated feedback, and the output is a composite video file.
[1009] These processing steps make it easier for users to recreate professional recipes at home and improve their cooking skills by receiving specific feedback and advice.
[1010] (Application example 1)
[1011] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1012] Conventional cooking assistance systems lack the ability to provide real-time feedback when users are trying to recreate professional recipes. This makes it difficult for users to notice problems while cooking, resulting in less-than-perfect dishes. Furthermore, there is a lack of a way to specifically analyze differences in cooking procedures and ingredient usage, and provide the results in an easy-to-understand manner.
[1013] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1014] In this invention, the server includes a means for uploading cooking videos filmed by users, a means for importing cooking recipe videos, a means for analyzing the cooking videos and the recipe videos and extracting their respective cooking steps, utensils, and ingredients, a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, a means for generating feedback and advice for the user based on the differences, a means for displaying the generated feedback and advice in real time, a means for generating a composite video including the feedback and advice, and a means for providing the composite video to the user. This allows users to receive feedback in real time while cooking, making it easier to identify specific areas for improvement. Furthermore, the differences between professional recipe videos and the user's cooking video are clearly presented, allowing users to easily recreate recipes.
[1015] "Cooking videos filmed by users" are videos in which users record the cooking process using a home camera or smart device.
[1016] A "cooking recipe video" is a video in which a professional or expert demonstrates a cooking method while explaining the steps.
[1017] A "means" is a method or mechanism for achieving a specific function or purpose.
[1018] "Analysis" means analyzing the video content frame by frame and extracting information such as each cooking step, utensils used, and ingredients.
[1019] A "cooking procedure" is a series of steps or methods for preparing a dish.
[1020] "Utensils" are tools and machines used in cooking.
[1021] "Ingredients" are the ingredients or components used to make a dish.
[1022] "Differences" are differences or discrepancies between the user's cooking procedures, utensils used, and ingredients and the content of professional recipe videos.
[1023] "Feedback" refers to information or advice provided to the user based on the analysis results.
[1024] "Advice" refers to suggestions to the user for specific improvements, such as cooking methods or how to use ingredients.
[1025] "Real-time display" means that users receive instant feedback and advice while cooking.
[1026] A "composite video" is a video that combines a professional recipe video with a user's cooking video, and inserts feedback and advice based on the differences.
[1027] This invention is a system that makes it easy for users to recreate professional recipes when cooking at home. The main components of the system include a user terminal and a server. The system programs and their processing are described in detail below.
[1028] First, the user uses a device such as a smartphone or smart glasses to record the cooking process. The video is then uploaded to a server via a dedicated application. This application has both recording and uploading functions, providing an environment in which users can easily share videos.
[1029] The server receives both cooking videos and cooking recipe videos uploaded by users and has the analysis function to analyze each. For the analysis, OpenCV (a library for video analysis) and TensorFlow (a library for using machine learning models) are used. Specifically, the video is analyzed frame by frame to extract information on cooking steps, utensils used, and ingredients.
[1030] The extracted information is compared with the content of the cooking recipe video. This comparison uses differential analysis to identify the differences between the user's cooking and the professional recipe. The results of the differential analysis generate specific feedback and advice for the user. This feedback and advice is displayed in real time on the user's device.
[1031] In addition, the system synthesizes professional recipe videos with users' cooking videos, generating a composite video that includes feedback and advice based on the differences, making it easier for users to identify specific areas for improvement while cooking.
[1032] For example, a user may film and upload a video of themselves making beef stew. In this case, the server analyzes the video and identifies differences such as "the meat is cut too large" and "the amount of wine is too small." Based on this, specific advice such as "cut the meat into bite-sized pieces" and "add more wine" is generated.
[1033] An example of a prompt for a generative AI model is:
[1034] "Compare the professional recipe video for making beef stew with the user cooking video and provide the following differences and feedback:
[1035] 1. Size of meat cut
[1036] 2. Amount of wine used
[1037] This system makes it easy for users to faithfully recreate professional recipes at home, improving their cooking skills.
[1038] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1039] Step 1:
[1040] The user takes a video of the cooking process using a smartphone or smart glasses camera. The input at this stage is the cooking action and scene. The video is captured and saved as a video file. The output is a cooking video file.
[1041] Step 2:
[1042] The user uploads the cooking video they have filmed to the server using a dedicated application. At this stage, the input is the cooking video file, and the output is a notification that the data has been transferred to the server.
[1043] Step 3:
[1044] The server saves the cooking video received from the user. The input at this stage is the uploaded cooking video file, and the output is the destination directory and file path. Here, the file system is used to save the video in the appropriate directory.
[1045] Step 4:
[1046] The server analyzes stored cooking videos and professional cooking recipe videos. The inputs are cooking video files and professional recipe video files. The videos are decomposed frame by frame, and the OpenCV library is used to extract information about cooking steps, utensils, and ingredients. The output is a list of cooking steps, utensils, and ingredients.
[1047] Step 5:
[1048] The server compares the user's cooking procedures, utensils, and ingredients with the professional recipes based on the extracted information. This comparison uses a difference analysis algorithm. The input is the list of information extracted in step 4 and the professional recipe information list. The output is the specific differences between the user's cooking and the recipes.
[1049] Step 6:
[1050] The server generates specific feedback and advice for the user based on the results of the difference analysis. The input at this stage is the difference information. The generative AI model is used to generate advice. The output is a feedback message and a list of specific advice.
[1051] Step 7:
[1052] The server displays the generated feedback and advice on the user's terminal in real time. The input at this stage is a feedback / advice list, and the output is the feedback / advice displayed on the user's terminal screen.
[1053] Step 8:
[1054] The server synthesizes professional recipe videos and user cooking videos, generating a composite video with feedback and advice based on the differences. The inputs are cooking video files, recipe video files, and a feedback and advice list. The composite video is created using video editing software. The output is a composite video file.
[1055] Step 9:
[1056] The server provides the generated composite video to the user. The input at this stage is the composite video file, and the output is the composite video file sent to the user's device. It is provided through a dedicated application or a web browser.
[1057] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1058] This invention is a system that makes it easy for users to recreate professional recipes at home, and also recognizes the user's emotions and provides feedback and advice accordingly. This system synthesizes a video of the user cooking with a professional recipe video and analyzes the differences to generate specific feedback and advice, and also adjusts the content and tone of this feedback and advice according to the user's emotions.
[1059] System Overview
[1060] This system consists of a user's terminal, a server for processing, and a network for exchanging information between them. It also includes an emotion engine, which adds the ability to recognize the user's emotions. The system's program is described in detail below.
[1061] User device functions
[1062] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[1063] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[1064] 3. Emotion recognition function: Equipped with a function to analyze the user's facial expressions, tone of voice, words, etc. This allows the user's emotions to be recognized in real time.
[1065] Server Features
[1066] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[1067] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[1068] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[1069] 4. Difference analysis function: The function compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[1070] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as reducing the amount of salt or extending the cooking time.
[1071] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state in real time and adjusts the tone and content of feedback and advice accordingly. For example, if the user is feeling stressed, the system may add words of encouragement.
[1072] 7. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[1073] Overview of operation
[1074] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with that of professional recipe videos, and differences are analyzed. Based on the analysis results, personalized feedback and advice is generated and inserted into the composite video. In addition, the system monitors the user's emotional state in real time and adjusts the tone and content of the feedback as needed. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[1075] Specific examples
[1076] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. The analysis identifies discrepancies, such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it adds encouraging messages and advice for relaxation. These pieces of advice are combined with a professional recipe video and provided to the user. Users can watch the combined video, compare their own movements with those of the professional, understand the necessary improvements, and incorporate them into their next cooking experience.
[1077] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills.Furthermore, by providing feedback based on the user's emotions, the system achieves a more comfortable and effective learning experience.
[1078] The processing flow will be explained below.
[1079] Step 1:
[1080] The user films themselves cooking at home using a smartphone or tablet camera, and their facial expressions and tone of voice are also recorded during the filming.
[1081] Step 2:
[1082] Cooking videos taken by users are uploaded to the system via a dedicated application or website.
[1083] Step 3:
[1084] The device receives the uploaded cooking video and sends it to the server.
[1085] Step 4:
[1086] The server stores the received cooking video in a database and begins video analysis.
[1087] Step 5:
[1088] The server analyzes each frame of the video and extracts information about the cooking steps, utensils, and ingredients used. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are analyzed.
[1089] Step 6:
[1090] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[1091] Step 7:
[1092] The server performs similar video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[1093] Step 8:
[1094] The server compares the user's cooking video with a professional recipe video and analyzes the differences in specific cooking steps, utensils, and ingredients. For example, it identifies differences such as "the user added too much salt" or "the heating time was too short."
[1095] Step 9:
[1096] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "reduce the amount of salt to one teaspoon" or "shorten the cooking time from 10 minutes to 5 minutes."
[1097] Step 10:
[1098] The server uses an emotion engine to monitor the user's emotional state in real time, analyzing their facial expressions, tone of voice, and words to determine whether they are feeling stressed.
[1099] Step 11:
[1100] The server adjusts the content and tone of the feedback and advice depending on the user's emotions. For example, if the user is feeling stressed, it adds an encouraging message.
[1101] Step 12:
[1102] The server generates a personalized composite video for each user, which displays a professional recipe video alongside the user's cooking video, and inserts feedback and emotional advice.
[1103] Step 13:
[1104] The server sends the generated composite video to the terminal.
[1105] Step 14:
[1106] The device then displays the synthesized video to the user. The user watches the video, understands the differences between the professional's movements and their own, and can reflect this in their next cooking experience. Receiving feedback based on their emotions makes learning more stress-free and effective.
[1107] Example 2
[1108] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1109] Conventional cooking assistance systems have made it difficult for users to accurately recreate professional recipes at home. Furthermore, they generally provide feedback and advice without taking into account the user's individual cooking situation or emotions, making it difficult to provide effective support. Furthermore, there was a lack of a way to accurately analyze the differences between cooking videos and professional recipe videos and provide users with specific areas for improvement. This made it difficult for users to obtain specific guidance on how to improve their cooking skills.
[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1111] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, means for generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice, means for providing the composite video to the user, and means for recognizing the user's emotions in real time and adjusting the content and tone of the feedback and advice. This makes it easier for users to accurately recreate professional recipes at home and to receive specific feedback and advice tailored to their individual cooking situations and emotions.
[1112] A "user" is an individual who cooks at home and inputs the cooking process into the system.
[1113] A "cooking video" is a video file that records the cooking process filmed by a user.
[1114] A "professional recipe video" is a video file in which a professional chef demonstrates cooking steps.
[1115] "Means for uploading" refers to the functions and interfaces that allow users to send cooking videos they have filmed to a server.
[1116] "Means of import" refers to the functionality and interface that allows professional recipe videos to be used within the system.
[1117] "Means of analysis" refers to algorithms and software that analyze cooking and recipe videos and extract the cooking steps, utensils, and ingredients for each.
[1118] "Comparison means" refers to functions and algorithms for comparing extracted cooking steps, utensils, and ingredients between cooking videos and recipe videos.
[1119] "Means for analyzing differences" refers to functions or algorithms for identifying and analyzing the differences between cooking videos and recipe videos.
[1120] The "means for generating feedback and advice" refers to a function or algorithm for generating specific feedback and improvement suggestions for the user based on the results of the differential analysis.
[1121] "Means for generating composite videos" refers to functions and algorithms that align professional recipe videos with users' cooking videos to create videos that include feedback and advice based on the analysis results.
[1122] The "means for providing" refers to the functions and interfaces that allow users to view the generated composite video.
[1123] "Means for recognizing emotions" refers to functions and algorithms that analyze a user's facial expressions, tone of voice, words, etc. to identify the user's emotional state in real time.
[1124] "Means for adjusting the content and tone of feedback and advice" refers to functions and algorithms for changing the content and expression of feedback and advice depending on the user's emotional state.
[1125] This invention is a system that allows users to easily recreate professional recipes when cooking at home, and also recognizes the user's emotions in real time and provides feedback and advice. This system is composed of a user terminal, a server, and a network connecting these. A specific embodiment of this system will be described below.
[1126] User device functions
[1127] 1. Shooting function
[1128] When cooking at home, users use the camera function of their smartphones or tablets to record the cooking process, and these video files can then be analyzed.
[1129] 2. Upload function
[1130] Users upload their cooking videos to the server via a dedicated application or web browser. Uploading requires an internet connection and may take some time depending on the size of the video.
[1131] 3. Emotion recognition function
[1132] This function analyzes the user's facial expressions, tone of voice, words, etc. in real time to recognize their emotions, making it possible to determine the level of stress or joy the user is feeling.
[1133] Server Features
[1134] 1. Video reception and storage function
[1135] The server receives and stores cooking videos uploaded by users. Metadata (e.g., shooting date and time, user ID) is also stored in the video file, making it easier to analyze and search later.
[1136] 2. Video analysis function
[1137] The server analyzes the stored cooking videos frame by frame using computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow). The analysis extracts cooking steps, utensils, and ingredients, and records them in a database.
[1138] 3. Differential analysis function
[1139] The server compares the user's cooking video with professional recipe videos to identify differences in cooking procedures, utensils, and ingredients. This analysis uses comparison algorithms and natural language processing techniques. The identified differences are stored in a database.
[1140] 4. Feedback generation function
[1141] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, sometimes using a generative AI model (e.g., GPT-3) to automatically generate feedback with specific improvements.
[1142] 5. Emotion-adaptive feedback function
[1143] The server monitors the user's emotional state in real time and adjusts the content and tone of the feedback and advice it provides, for example adding an encouraging message if the user is feeling stressed.
[1144] 6. Composite video generation function
[1145] The system displays professional recipe videos and user cooking videos side by side, and generates a composite video incorporating feedback based on the analysis results. Comments about the user's feelings may also be added to this composite video.
[1146] Specific examples
[1147] For example, suppose a user takes a video of themselves making carbonara and uploads it to a server. The server receives the video and begins analyzing it. The analysis results identify the following differences:
[1148] Bacon cut differently
[1149] Mixing the eggs too late
[1150] A lot of salt
[1151] Based on this, the following specific advice is generated:
[1152] Bacon should be thinly sliced
[1153] Mix the eggs quickly
[1154] Reduce the amount of salt by half
[1155] Additionally, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it will add encouraging messages and advice on how to relax.
[1156] Prompt Sentence Examples
[1157] "I'd like specific feedback on my steps in making this carbonara. I'd especially like some advice on when to mix the eggs and how much salt to use."
[1158] In this way, users can watch the composite video, compare their own movements with those of a professional, understand what improvements are needed, and improve their cooking skills at home.
[1159] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1160] Step 1:
[1161] A user shoots a cooking video.
[1162] Specific operation: The user uses the camera on their smartphone or tablet to film themselves cooking at home, adjusting the camera angle and lighting to ensure a clear video.
[1163] Input: The user's cooking scene.
[1164] Output: Cooking video file.
[1165] Step 2:
[1166] The device uploads the cooking video to the server.
[1167] Specific operation: The user instructs the device to upload the video they have taken to the server using a dedicated application or web browser. The device then transfers the video file to the server via the Internet.
[1168] Input: Filmed cooking video files.
[1169] Output: Cooking video files uploaded to the server.
[1170] Step 3:
[1171] The server receives and stores the cooking video.
[1172] Specific operation: The server receives and stores the uploaded video file, along with metadata (e.g., shooting date and time, user ID).
[1173] Input: Uploaded cooking video file and metadata.
[1174] Output: Saved cooking video files and metadata.
[1175] Step 4:
[1176] The server analyzes the video.
[1177] How it works: The server analyzes the stored cooking videos frame by frame, and uses computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow) to extract cooking steps, utensils, and ingredients.
[1178] Input: Saved cooking video file.
[1179] Output: Extracted cooking steps, utensils, and ingredients data.
[1180] Step 5:
[1181] The server analyzes the differences with professional recipe videos.
[1182] Specific operation: The server compares the user's cooking video with professional recipe videos and uses comparison algorithms and natural language processing techniques to identify differences in cooking steps, utensils used, and ingredients.
[1183] Input: Extracted cooking instructions, utensils, and ingredient data; professional recipe videos.
[1184] Output: Differential analysis results.
[1185] Step 6:
[1186] The server generates feedback and advice.
[1187] Specific operation: The server generates feedback and advice for the user based on the results of the differential analysis. It uses a generative AI model (e.g., GPT-3) to automatically create feedback including specific improvements.
[1188] Input: Differential analysis results.
[1189] Output: Generated feedback and advice.
[1190] Step 7:
[1191] The server adjusts according to the user's emotional state.
[1192] Specific operation: The server uses emotion recognition to analyze the user's facial expressions, tone of voice, and words to recognize their emotions. It then adjusts the content and tone of the feedback and advice provided according to their emotional state.
[1193] Input: User facial expressions, tone of voice, and verbal data.
[1194] Output: Tailored feedback and advice.
[1195] Step 8:
[1196] The server generates the composite video.
[1197] How it works: The server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video that includes feedback and advice based on the analysis results. It may also add comments about the user's feelings to the video.
[1198] Inputs: Tailored feedback and advice; professional recipe videos; user cooking videos.
[1199] Output: The generated composite video.
[1200] Step 9:
[1201] The user watches the composite video.
[1202] Specific operation: The user watches a composite video provided by the server on their device. While watching this composite video, they compare their own cooking with that of a professional and understand what improvements are needed.
[1203] Input: The generated synthetic video.
[1204] Output: Improved cooking skills of the user.
[1205] (Application example 2)
[1206] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1207] Conventional cooking feedback systems aim to improve users' cooking skills, but they lack functionality that takes users' emotions into consideration, which often causes stress and makes it difficult for users to continue using the system. Furthermore, they lack specific suggestions for improvements when reproducing professional recipes, making it difficult for users to effectively improve their own cooking skills.
[1208] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: a means for uploading cooking videos filmed by the user; a means for importing professional recipe videos; a means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients; a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences; a means for generating feedback and advice for the user based on the differences; a means for recognizing the user's emotions and adjusting the content and tone of the feedback and advice according to the emotions; a means for generating a composite video including the generated feedback and advice; and a means for providing the composite video to the user. This makes it possible to provide less stressful feedback according to the user's emotional state and present specific improvements.
[1209] A "cooking video filmed by a user" is a video recording of a user cooking using a device such as a smartphone or tablet.
[1210] "Professional recipe videos" are videos containing cooking steps shown by professional chefs or cooking experts, and are intended for users to use as reference.
[1211] "Means for analyzing cooking videos and recipe videos and extracting the respective cooking steps, utensils, and ingredients" refers to a technical means for analyzing uploaded cooking videos and recipe videos and automatically recognizing and extracting the cooking steps, utensils used, and ingredients from each frame.
[1212] The "means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences" refers to a technical means for comparing a user's cooking video with a professional recipe video and identifying differences in cooking procedures, utensils, and ingredients.
[1213] The "means for generating feedback and advice for the user based on the difference" is a technology for automatically generating feedback and specific advice that suggests improvements to the user's cooking method based on the results of the difference analysis.
[1214] "Means for recognizing a user's emotions and adjusting the content and tone of feedback or advice according to those emotions" refers to a technical means for analyzing a user's facial expressions, tone of voice, words, etc. to grasp their emotional state, and adjusting the content and tone of the feedback or advice provided according to those emotions.
[1215] The "means for generating a composite video including generated feedback and advice" is a technology for automatically generating a composite video that inserts comparison results and feedback based on a user's cooking video and a professional recipe video.
[1216] The "means for providing a composite video to a user" refers to a technical means for transmitting the generated composite video to a user's terminal so that the user can view the video.
[1217] This invention is a system that makes it easy for users to recreate professional recipes at home. It recognizes the user's emotions and provides feedback and advice according to those emotions. The system analyzes and synthesizes the user's cooking video and the professional recipe video, and identifies and provides specific differences.
[1218] System configuration
[1219] This system consists of user terminals, servers for processing information, and a network that connects them. The specific functions of each element are explained below.
[1220] User terminal
[1221] User terminals mainly include smartphones and tablets.
[1222] 1. Recording function: Equipped with a camera function that allows users to record videos of themselves cooking at home, for example, using a smartphone camera.
[1223] 2. Upload function: An application or web browser is provided for uploading the recorded cooking videos to the server.
[1224] 3. Emotion recognition function: Using the camera and microphone, the system analyzes the user's facial expressions, tone of voice, and words in real time to recognize the user's emotions.
[1225] server
[1226] The server analyzes the data sent from the user device and generates feedback and advice. Specific functions include:
[1227] 1. Video reception and storage function: Receives cooking videos uploaded by users and stores them on the server.
[1228] 2. Video analysis function: Analyzes cooking videos and professional recipe videos frame by frame to extract information on cooking steps, utensils, and ingredients. OpenCV and dlib are used here.
[1229] 3. Database function: Equipped with a database for managing information on extracted cooking procedures, utensils, and ingredients.
[1230] 4. Difference analysis function: Compares cooking videos with professional recipe videos and analyzes differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[1231] 5. Feedback generation function: Generates specific feedback and advice for users based on differential analysis. The software used is a generative AI model.
[1232] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state and adjusts the tone and content of feedback and advice. For example, if the user is feeling stressed, the system adds advice on how to relax.
[1233] 7. Composite video generation function: Combines the user's cooking video with a professional recipe video to generate a composite video that includes feedback and advice.
[1234] Overview of operation
[1235] The user cooks, films the process with their smartphone, and uploads the video to the server via a dedicated app. The server receives the video and begins analysis. The extracted cooking steps, utensils, and ingredient information are compared with that of professional recipe videos to identify differences. Specific feedback and advice is generated based on the differences, and the system also monitors the user's emotional state and adapts the content and tone of the feedback. The professional recipe video and the user's cooking video are then combined to generate a composite video that includes feedback and advice. The composite video is finally provided to the user, who can watch it to improve their own cooking skills.
[1236] Specific examples
[1237] For example, a user can film themselves making carbonara and upload it via a dedicated app. The server receives the video and begins analyzing it, identifying discrepancies such as the way the bacon was cut, the timing of the eggs being added too late, and the amount of salt being too much. Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the user feels anxious or stressed while cooking, the emotion engine detects this and adds encouraging messages or advice on how to relax. This advice is then combined with a professional recipe video and provided to the user. The user can then watch the combined video, compare their own movements with those of the professional, and use the results to improve their cooking skills.
[1238] Prompt Sentence Examples
[1239] Users film themselves making carbonara and upload it to the app, where it will compare it with professional recipe videos and provide specific advice.
[1240] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills. It also provides emotional feedback to users, making the learning experience more comfortable and effective.
[1241] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1242] Step 1:
[1243] The user takes a photo of themselves cooking using a device (smartphone or tablet).
[1244] Input: User's cooking video (obtained in real time via the device camera)
[1245] Output: Cooking video file
[1246] Step 2:
[1247] Cooking videos taken by the device are uploaded to a server via a dedicated application.
[1248] Input: Cooking video file
[1249] Output: Cooking video data uploaded to the server
[1250] Step 3:
[1251] The server receives the uploaded cooking videos and stores them in a database.
[1252] Input: User's cooking video data
[1253] Output: Cooking video data stored in a database
[1254] Step 4:
[1255] The server retrieves professional recipe videos from a database.
[1256] Input: Professional recipe video data request
[1257] Output: Professional recipe video data extracted from the database
[1258] Step 5:
[1259] The server analyzes cooking and recipe videos frame by frame and extracts information on each cooking step, utensils used, and ingredients.
[1260] Input: Cooking video data and recipe video data
[1261] Data processing: Analyzes video frame by frame using OpenCV and dlib, automatically recognizing cooking steps, utensils, and ingredients
[1262] Output: Extracted cooking instructions, utensils, and ingredient data
[1263] Step 6:
[1264] The server compares the extracted cooking procedures, utensils, and ingredient data and analyzes the differences.
[1265] Input: User cooking instructions, equipment, and ingredient data; Professional recipe instructions, equipment, and ingredient data
[1266] Data calculation: Calculates differences in cooking procedures, utensils, and ingredients
[1267] Output: Difference data of cooking procedures, utensils, and ingredients
[1268] Step 7:
[1269] The server generates feedback and advice for the user based on the difference data.
[1270] Input: Cooking procedure, utensils, ingredient differential data
[1271] Data shaping: Using generative AI models to generate specific advice
[1272] Output: Specific feedback and advice to the user
[1273] Step 8:
[1274] The server recognizes the user's emotions and adjusts the content and tone of the feedback and advice according to the emotions.
[1275] Input: User's emotional data (facial expressions, tone of voice, and word analysis results)
[1276] Data Computation: Feedback Modulation Based on Emotion Engine
[1277] Output: Emotionally tailored feedback and advice
[1278] Step 9:
[1279] The server generates a composite video that includes the generated feedback and advice.
[1280] Input: Adjusted feedback and advice data, user cooking video data, professional recipe video data
[1281] Data processing: Combining cooking videos and recipe videos, and inserting feedback
[1282] Output: Generated composite video
[1283] Step 10:
[1284] The server provides the generated composite video to the user's terminal.
[1285] Input: Synthetic video data
[1286] Output: Composite video sent to the user's device
[1287] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1288] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1289] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1290] [Fourth embodiment]
[1291] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1292] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1293] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1294] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1295] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1296] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1297] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1298] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1299] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1300] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1301] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1302] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1303] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1304] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[1305] System Overview
[1306] This system consists of a user terminal, a server for processing, and a network for exchanging information between them. The system programs are described in detail below.
[1307] User device functions
[1308] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[1309] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[1310] Server Features
[1311] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[1312] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[1313] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[1314] 4. Difference analysis function: The app has a function to compare the user's cooking video with a professional recipe video and analyze the differences. Specifically, it identifies differences in cooking procedures, utensils, and ingredients.
[1315] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as "the amount of salt is too much" or "shorten the heating time."
[1316] 6. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[1317] Overview of operation
[1318] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with information from professional recipe videos, and differences are analyzed. Based on the analysis results, feedback and advice tailored to the user is generated and inserted into the composite video. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[1319] Specific examples
[1320] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. As a result of the analysis, it identifies differences such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, it generates specific advice such as "slicing the bacon thinner," "mixing the eggs early," and "reducing the amount of salt by half." This advice is then combined with a professional recipe video and provided to the user. The user watches the combined video and understands how to make a carbonara that is closer to what a professional would make.
[1321] In this way, the system makes it easy for users to recreate professional recipes at home and helps improve their cooking skills.
[1322] The processing flow will be explained below.
[1323] Step 1:
[1324] Users film themselves cooking at home, using the camera function of their smartphones or tablets to record the entire cooking process.
[1325] Step 2:
[1326] Cooking videos taken by users are uploaded to the system using a dedicated application or website.
[1327] Step 3:
[1328] The device receives the uploaded cooking video and sends it to the server.
[1329] Step 4:
[1330] The server stores the received cooking video in a database and begins analyzing it.
[1331] Step 5:
[1332] The server analyzes the video frame by frame and extracts information about cooking steps, utensils, and ingredients. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are extracted.
[1333] Step 6:
[1334] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[1335] Step 7:
[1336] The server also performs video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[1337] Step 8:
[1338] The server compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, specific differences such as "too much salt" or "too short heating time" can be identified.
[1339] Step 9:
[1340] Based on the analysis results, the server generates specific feedback and advice for the user, such as "reduce the amount of salt" or "extend the cooking time."
[1341] Step 10:
[1342] The server generates a composite video for each user, which displays a professional recipe video and the user's cooking video side by side, and inserts feedback and advice based on the differences.
[1343] Step 11:
[1344] The server sends the generated composite video to the terminal.
[1345] Step 12:
[1346] The device then displays the synthesized video to the user, who can then watch the video to understand the differences between the professional's movements and their own, and use the feedback to improve their cooking procedures and techniques.
[1347] Example 1
[1348] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1349] With conventional cooking methods, it was difficult for users to recreate professional recipes, and there were limited ways to receive specific feedback and advice. Furthermore, there was no system that analyzed the user's cooking process in detail and provided specific suggestions for improvement. This meant that users were unable to faithfully recreate professional recipes, making it difficult to efficiently improve their cooking skills.
[1350] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1351] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and the recipe videos using a video analysis library and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences using a deep learning model, means for automatically generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice using video editing software, and means for providing the composite video to the user. This makes it easier for users to recreate professional recipes at home and allows them to efficiently improve their cooking skills by receiving specific feedback and advice.
[1352] 1. "Cooking videos filmed by users" refers to video data that records the process of a user cooking at home.
[1353] 2. "Professional recipe video" refers to video data that records the process of a professional chef preparing a specific dish.
[1354] 3. "Video Analysis Library" means a software library for analyzing video data frame by frame and extracting specific information from each frame.
[1355] 4. A "cooking procedure" is a series of steps for preparing a dish, including the specific tasks and order of the steps.
[1356] 5. "Utensils" are tools and equipment used in cooking, including knives and pots.
[1357] 6. "Ingredients" are the ingredients and seasonings used in making a dish.
[1358] 7. A "deep learning model" is a machine learning algorithm that performs highly accurate analysis and predictions by learning features from large amounts of data.
[1359] 8. "Differences" are the differences between the user's cooking procedures, equipment, and ingredients and those of professional recipe videos.
[1360] 9. "Feedback" means evaluation, comments, or advice provided to the user based on the analysis results.
[1361] 10. "Advice" means specific instructions or advice provided to a user to improve their cooking skills.
[1362] 11. "Synthetic video" is video data that displays a professional recipe video and a user's cooking video side by side, with feedback and advice based on the differences inserted.
[1363] 12. "Video editing software" means software for combining and editing multiple video data to generate new video.
[1364] 13. "Uploading means" refers to a function or device that allows users to send cooking videos they have filmed to the server.
[1365] 14. "Means for capturing" means a function or device that allows the server to receive and store professional recipe videos.
[1366] 15. "Means for providing" refers to a function or device for transmitting the generated composite video to a user so that the video can be viewed.
[1367] This invention is a system that makes it easy for users to recreate professional recipes at home. The system synthesizes a video of the user cooking with a professional recipe video, analyzes the differences, and provides specific feedback and advice.
[1368] This system consists of a user terminal, a server, and a network that exchanges information between them. Each element of the system is explained in detail below.
[1369] User device functions
[1370] Shooting Function:
[1371] Users use the camera on their smartphone or tablet to film themselves cooking at home. To do so, they launch the device's "camera" app and press the record button to record the entire cooking process from start to finish.
[1372] Upload function:
[1373] The user uploads the video they have taken to the server through a dedicated application (e.g., "CookingPro"). To do this, open the application, go to the upload screen, and select the saved video file.
[1374] Server Features
[1375] Video receiving and storage function:
[1376] The server receives cooking videos uploaded by users and stores them in a storage service (e.g., Amazon S3).
[1377] Video analysis features:
[1378] The server uses software such as OpenCV and ffmpeg to analyze the received video frame by frame, extracting information about the cooking steps, utensils used, and ingredients.
[1379] Database features:
[1380] The server stores the cooking procedures, utensils, and ingredient information extracted through the analysis in a database (e.g., MySQL or PostgreSQL).
[1381] Differential analysis features:
[1382] The server compares the user's cooking video with professional recipe videos using a deep learning model (e.g., TensorFlow or PyTorch) to analyze differences in cooking procedures, utensils used, and ingredients.
[1383] Feedback generation features:
[1384] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "cut the bacon thinner," "add the eggs earlier," or "reduce the amount of salt by half."
[1385] Composite video generation function:
[1386] The server displays the professional recipe video and the user's cooking video side by side, creating a composite video that includes feedback. The composite is created using video editing software (e.g., Adobe Premiere Pro or MoviePy).
[1387] Features provided:
[1388] The server provides the generated composite video to the user via a communication method such as an application or email.
[1389] Specific examples
[1390] For example, a user can film a video of themselves cooking "Carbonara" and upload it to the server via the "CookingPro" app. The server receives the video and begins analyzing it. OpenCV and ffmpeg are used for the analysis, and information such as "how to cut the bacon," "when to mix the eggs," and "how much salt" is extracted from each frame. The information obtained is then stored in a database.
[1391] The server then compares the video with a professional "Carbonara" recipe video and analyzes differences such as the bacon being too thick, the eggs being added too late, or the amount of salt being too high. This is done using TensorFlow. Based on the analysis results, feedback is generated, such as "slicing the bacon thinner," "adding the eggs earlier," or "reducing the amount of salt by half." Adobe Premiere Pro is then used to generate a composite video that displays the professional recipe video and the user's cooking video side by side. Users can watch this composite video and understand specific areas for improvement, allowing them to improve their cooking skills.
[1392] Prompt Sentence Examples
[1393] The system analyzes the cooking video of carbonara filmed by the user, such as how to cut the bacon, when to mix the eggs, and the amount of salt, and provides specific feedback to create a composite video.
[1394] In this way, this system makes it easier for users to recreate professional recipes at home and helps improve their cooking skills.
[1395] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1396] Step 1:
[1397] A user shoots a cooking video.
[1398] Specifically, the user launches the camera app on their smartphone or tablet and presses the record button to record the cooking process. Taking detailed photos, including the overall image of the food being cooked and the movements of the hands, improves the accuracy of the analysis.
[1399] The input is a real scene of the cooking process, and the output is a recorded video file.
[1400] Step 2:
[1401] Cooking videos taken by users are uploaded from their terminals to a server.
[1402] Specifically, the user opens a dedicated application (e.g., "CookingPro"), goes to the upload screen, selects the saved video file, and presses the send button.
[1403] The input is a cooking video file stored on the user's terminal, and the output is a video file sent to the server.
[1404] Step 3:
[1405] The server receives and stores the video.
[1406] Specifically, the server receives the uploaded video data, stores it in a temporary directory, and then transfers it to cloud storage (e.g., Amazon S3).
[1407] The input is a video file uploaded by a user, and the output is a video file saved in the server's storage.
[1408] Step 4:
[1409] The server begins analyzing the video.
[1410] Specifically, the server uses OpenCV and ffmpeg to divide the received video into frames and extract information about the cooking steps, utensils, and ingredients from each frame. For example, it performs image recognition on each frame to identify ingredients and cooking utensils.
[1411] The input is a stored video file, and the output is extracted cooking instructions, utensils, and ingredient information for each frame.
[1412] Step 5:
[1413] The server stores information about cooking procedures, equipment, and ingredients in a database.
[1414] Specifically, the server stores the information extracted by the analysis in a database (e.g., MySQL or PostgreSQL) and manages the information in a structured manner.
[1415] The input is the information on cooking procedures, utensils, and ingredients obtained as a result of the analysis, and the output is a database in which this information is stored.
[1416] Step 6:
[1417] The server compares it with professional recipe videos and analyzes the differences.
[1418] Specifically, the server uses deep learning models (e.g., TensorFlow or PyTorch) to analyze and compare the user's cooking video with professional recipe videos, identifying differences in cooking procedures, utensils, and ingredients, for example, detecting discrepancies in the timeline and differences in procedures.
[1419] The input is the analysis results of the user's cooking video and the analysis results of the professional recipe video, and the output is the identified difference information.
[1420] Step 7:
[1421] The server generates specific feedback based on the differences.
[1422] Specifically, the server uses a generative AI model to automatically generate specific feedback and advice for the user based on the difference information, such as "cut the bacon thinner," "add the eggs early," or "reduce the amount of salt."
[1423] The input is the difference information and the output is the generated feedback and advice.
[1424] Step 8:
[1425] The server generates a composite video including the feedback and provides it to the user.
[1426] Specifically, the server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video with feedback inserted using video editing software (e.g., Adobe Premiere Pro or MoviePy). The generated composite video is then provided to the user's device or via email.
[1427] The inputs are professional recipe videos, user cooking videos, and generated feedback, and the output is a composite video file.
[1428] These processing steps make it easier for users to recreate professional recipes at home and improve their cooking skills by receiving specific feedback and advice.
[1429] (Application example 1)
[1430] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1431] Conventional cooking assistance systems lack the ability to provide real-time feedback when users are trying to recreate professional recipes. This makes it difficult for users to notice problems while cooking, resulting in less-than-perfect dishes. Furthermore, there is a lack of a way to specifically analyze differences in cooking procedures and ingredient usage, and provide the results in an easy-to-understand manner.
[1432] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1433] In this invention, the server includes a means for uploading cooking videos filmed by users, a means for importing cooking recipe videos, a means for analyzing the cooking videos and the recipe videos and extracting their respective cooking steps, utensils, and ingredients, a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, a means for generating feedback and advice for the user based on the differences, a means for displaying the generated feedback and advice in real time, a means for generating a composite video including the feedback and advice, and a means for providing the composite video to the user. This allows users to receive feedback in real time while cooking, making it easier to identify specific areas for improvement. Furthermore, the differences between professional recipe videos and the user's cooking video are clearly presented, allowing users to easily recreate recipes.
[1434] "Cooking videos filmed by users" are videos in which users record the cooking process using a home camera or smart device.
[1435] A "cooking recipe video" is a video in which a professional or expert demonstrates a cooking method while explaining the steps.
[1436] A "means" is a method or mechanism for achieving a specific function or purpose.
[1437] "Analysis" means analyzing the video content frame by frame and extracting information such as each cooking step, utensils used, and ingredients.
[1438] A "cooking procedure" is a series of steps or methods for preparing a dish.
[1439] "Utensils" are tools and machines used in cooking.
[1440] "Ingredients" are the ingredients or components used to make a dish.
[1441] "Differences" are differences or discrepancies between the user's cooking procedures, utensils used, and ingredients and the content of professional recipe videos.
[1442] "Feedback" refers to information or advice provided to the user based on the analysis results.
[1443] "Advice" refers to suggestions to the user for specific improvements, such as cooking methods or how to use ingredients.
[1444] "Real-time display" means that users receive instant feedback and advice while cooking.
[1445] A "composite video" is a video that combines a professional recipe video with a user's cooking video, and inserts feedback and advice based on the differences.
[1446] This invention is a system that makes it easy for users to recreate professional recipes when cooking at home. The main components of the system include a user terminal and a server. The system programs and their processing are described in detail below.
[1447] First, the user uses a device such as a smartphone or smart glasses to record the cooking process. The video is then uploaded to a server via a dedicated application. This application has both recording and uploading functions, providing an environment in which users can easily share videos.
[1448] The server receives both cooking videos and cooking recipe videos uploaded by users and has the analysis function to analyze each. For the analysis, OpenCV (a library for video analysis) and TensorFlow (a library for using machine learning models) are used. Specifically, the video is analyzed frame by frame to extract information on cooking steps, utensils used, and ingredients.
[1449] The extracted information is compared with the content of the cooking recipe video. This comparison uses differential analysis to identify the differences between the user's cooking and the professional recipe. The results of the differential analysis generate specific feedback and advice for the user. This feedback and advice is displayed in real time on the user's device.
[1450] In addition, the system synthesizes professional recipe videos with users' cooking videos, generating a composite video that includes feedback and advice based on the differences, making it easier for users to identify specific areas for improvement while cooking.
[1451] For example, a user may film and upload a video of themselves making beef stew. In this case, the server analyzes the video and identifies differences such as "the meat is cut too large" and "the amount of wine is too small." Based on this, specific advice such as "cut the meat into bite-sized pieces" and "add more wine" is generated.
[1452] An example of a prompt for a generative AI model is:
[1453] "Compare the professional recipe video for making beef stew with the user cooking video and provide the following differences and feedback:
[1454] 1. Size of meat cut
[1455] 2. Amount of wine used
[1456] This system makes it easy for users to faithfully recreate professional recipes at home, improving their cooking skills.
[1457] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1458] Step 1:
[1459] The user takes a video of the cooking process using a smartphone or smart glasses camera. The input at this stage is the cooking action and scene. The video is captured and saved as a video file. The output is a cooking video file.
[1460] Step 2:
[1461] The user uploads the cooking video they have filmed to the server using a dedicated application. At this stage, the input is the cooking video file, and the output is a notification that the data has been transferred to the server.
[1462] Step 3:
[1463] The server saves the cooking video received from the user. The input at this stage is the uploaded cooking video file, and the output is the destination directory and file path. Here, the file system is used to save the video in the appropriate directory.
[1464] Step 4:
[1465] The server analyzes stored cooking videos and professional cooking recipe videos. The inputs are cooking video files and professional recipe video files. The videos are decomposed frame by frame, and the OpenCV library is used to extract information about cooking steps, utensils, and ingredients. The output is a list of cooking steps, utensils, and ingredients.
[1466] Step 5:
[1467] The server compares the user's cooking procedures, utensils, and ingredients with the professional recipes based on the extracted information. This comparison uses a difference analysis algorithm. The input is the list of information extracted in step 4 and the professional recipe information list. The output is the specific differences between the user's cooking and the recipes.
[1468] Step 6:
[1469] The server generates specific feedback and advice for the user based on the results of the difference analysis. The input at this stage is the difference information. The generative AI model is used to generate advice. The output is a feedback message and a list of specific advice.
[1470] Step 7:
[1471] The server displays the generated feedback and advice on the user's terminal in real time. The input at this stage is a feedback / advice list, and the output is the feedback / advice displayed on the user's terminal screen.
[1472] Step 8:
[1473] The server synthesizes professional recipe videos and user cooking videos, generating a composite video with feedback and advice based on the differences. The inputs are cooking video files, recipe video files, and a feedback and advice list. The composite video is created using video editing software. The output is a composite video file.
[1474] Step 9:
[1475] The server provides the generated composite video to the user. The input at this stage is the composite video file, and the output is the composite video file sent to the user's device. It is provided through a dedicated application or a web browser.
[1476] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1477] This invention is a system that makes it easy for users to recreate professional recipes at home, and also recognizes the user's emotions and provides feedback and advice accordingly. This system synthesizes a video of the user cooking with a professional recipe video and analyzes the differences to generate specific feedback and advice, and also adjusts the content and tone of this feedback and advice according to the user's emotions.
[1478] System Overview
[1479] This system consists of a user's terminal, a server for processing, and a network for exchanging information between them. It also includes an emotion engine, which adds the ability to recognize the user's emotions. The system's program is described in detail below.
[1480] User device functions
[1481] 1. Camera function: The device has a camera function that allows users to take pictures of themselves cooking at home. Typical examples are smartphones and tablets.
[1482] 2. Upload function: The system has a function that allows users to upload cooking videos they have taken to the server. This is done through a dedicated application or a web browser.
[1483] 3. Emotion recognition function: Equipped with a function to analyze the user's facial expressions, tone of voice, words, etc. This allows the user's emotions to be recognized in real time.
[1484] Server Features
[1485] 1. Video reception and storage function: The app has the function to receive and store cooking videos uploaded by users.
[1486] 2. Video analysis function: This function analyzes cooking videos and professional recipe videos. This function analyzes the video frame by frame and extracts information about cooking steps, utensils used, and ingredients.
[1487] 3. Database function: Equipped with a database for storing and managing information on extracted cooking procedures, utensils, and ingredients.
[1488] 4. Difference analysis function: The function compares the user's cooking video with a professional recipe video and analyzes the differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[1489] 5. Feedback generation function: Based on the results of differential analysis, the system generates specific feedback and advice for the user, such as reducing the amount of salt or extending the cooking time.
[1490] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state in real time and adjusts the tone and content of feedback and advice accordingly. For example, if the user is feeling stressed, the system may add words of encouragement.
[1491] 7. Composite video generation function: This function generates a composite video specifically for the user. This video displays a professional recipe video and the user's cooking video side by side, and provides feedback based on the differences.
[1492] Overview of operation
[1493] The user cooks at home and films the cooking process. The video is then uploaded to the system via the application. The received video is analyzed by the server, and information on cooking steps, utensils, and ingredients is extracted. The extracted information is compared with that of professional recipe videos, and differences are analyzed. Based on the analysis results, personalized feedback and advice is generated and inserted into the composite video. In addition, the system monitors the user's emotional state in real time and adjusts the tone and content of the feedback as needed. By watching this composite video, the user can understand specific improvements that need to be made in their home environment and improve their cooking skills.
[1494] Specific examples
[1495] For example, suppose a user films and uploads a video of themselves making "carbonara." The server receives the video and begins analyzing it. The analysis identifies discrepancies, such as "the bacon was cut differently," "the timing of adding the eggs was too late," and "the amount of salt was too high." Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it adds encouraging messages and advice for relaxation. These pieces of advice are combined with a professional recipe video and provided to the user. Users can watch the combined video, compare their own movements with those of the professional, understand the necessary improvements, and incorporate them into their next cooking experience.
[1496] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills.Furthermore, by providing feedback based on the user's emotions, the system achieves a more comfortable and effective learning experience.
[1497] The processing flow will be explained below.
[1498] Step 1:
[1499] The user films themselves cooking at home using a smartphone or tablet camera, and their facial expressions and tone of voice are also recorded during the filming.
[1500] Step 2:
[1501] Cooking videos taken by users are uploaded to the system via a dedicated application or website.
[1502] Step 3:
[1503] The device receives the uploaded cooking video and sends it to the server.
[1504] Step 4:
[1505] The server stores the received cooking video in a database and begins video analysis.
[1506] Step 5:
[1507] The server analyzes each frame of the video and extracts information about the cooking steps, utensils, and ingredients used. For example, steps such as "put water in the pot," "add salt," and "boil the pasta" are analyzed.
[1508] Step 6:
[1509] The server retrieves professional recipe videos from the Internet or retrieves already stored professional recipe videos from a database.
[1510] Step 7:
[1511] The server performs similar video analysis on professional recipe videos to extract cooking steps, utensils, and ingredients, such as measuring the amount of water, adding the appropriate amount of salt, and maintaining the appropriate cooking time.
[1512] Step 8:
[1513] The server compares the user's cooking video with a professional recipe video and analyzes the differences in specific cooking steps, utensils, and ingredients. For example, it identifies differences such as "the user added too much salt" or "the heating time was too short."
[1514] Step 9:
[1515] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, such as "reduce the amount of salt to one teaspoon" or "shorten the cooking time from 10 minutes to 5 minutes."
[1516] Step 10:
[1517] The server uses an emotion engine to monitor the user's emotional state in real time, analyzing their facial expressions, tone of voice, and words to determine whether they are feeling stressed.
[1518] Step 11:
[1519] The server adjusts the content and tone of the feedback and advice depending on the user's emotions. For example, if the user is feeling stressed, it adds an encouraging message.
[1520] Step 12:
[1521] The server generates a personalized composite video for each user, which displays a professional recipe video alongside the user's cooking video, and inserts feedback and emotional advice.
[1522] Step 13:
[1523] The server sends the generated composite video to the terminal.
[1524] Step 14:
[1525] The device then displays the synthesized video to the user. The user watches the video, understands the differences between the professional's movements and their own, and can reflect this in their next cooking experience. Receiving feedback based on their emotions makes learning more stress-free and effective.
[1526] Example 2
[1527] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1528] Conventional cooking assistance systems have made it difficult for users to accurately recreate professional recipes at home. Furthermore, they generally provide feedback and advice without taking into account the user's individual cooking situation or emotions, making it difficult to provide effective support. Furthermore, there was a lack of a way to accurately analyze the differences between cooking videos and professional recipe videos and provide users with specific areas for improvement. This made it difficult for users to obtain specific guidance on how to improve their cooking skills.
[1529] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1530] In this invention, the server includes means for uploading cooking videos filmed by users, means for importing professional recipe videos, means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients, means for comparing the cooking steps, utensils, and ingredients and analyzing the differences, means for generating feedback and advice for the user based on the differences, means for generating a composite video including the generated feedback and advice, means for providing the composite video to the user, and means for recognizing the user's emotions in real time and adjusting the content and tone of the feedback and advice. This makes it easier for users to accurately recreate professional recipes at home and to receive specific feedback and advice tailored to their individual cooking situations and emotions.
[1531] A "user" is an individual who cooks at home and inputs the cooking process into the system.
[1532] A "cooking video" is a video file that records the cooking process filmed by a user.
[1533] A "professional recipe video" is a video file in which a professional chef demonstrates cooking steps.
[1534] "Means for uploading" refers to the functions and interfaces that allow users to send cooking videos they have filmed to a server.
[1535] "Means of import" refers to the functionality and interface that allows professional recipe videos to be used within the system.
[1536] "Means of analysis" refers to algorithms and software that analyze cooking and recipe videos and extract the cooking steps, utensils, and ingredients for each.
[1537] "Comparison means" refers to functions and algorithms for comparing extracted cooking steps, utensils, and ingredients between cooking videos and recipe videos.
[1538] "Means for analyzing differences" refers to functions or algorithms for identifying and analyzing the differences between cooking videos and recipe videos.
[1539] The "means for generating feedback and advice" refers to a function or algorithm for generating specific feedback and improvement suggestions for the user based on the results of the differential analysis.
[1540] "Means for generating composite videos" refers to functions and algorithms that align professional recipe videos with users' cooking videos to create videos that include feedback and advice based on the analysis results.
[1541] The "means for providing" refers to the functions and interfaces that allow users to view the generated composite video.
[1542] "Means for recognizing emotions" refers to functions and algorithms that analyze a user's facial expressions, tone of voice, words, etc. to identify the user's emotional state in real time.
[1543] "Means for adjusting the content and tone of feedback and advice" refers to functions and algorithms for changing the content and expression of feedback and advice depending on the user's emotional state.
[1544] This invention is a system that allows users to easily recreate professional recipes when cooking at home, and also recognizes the user's emotions in real time and provides feedback and advice. This system is composed of a user terminal, a server, and a network connecting these. A specific embodiment of this system will be described below.
[1545] User device functions
[1546] 1. Shooting function
[1547] When cooking at home, users use the camera function of their smartphones or tablets to record the cooking process, and these video files can then be analyzed.
[1548] 2. Upload function
[1549] Users upload their cooking videos to the server via a dedicated application or web browser. Uploading requires an internet connection and may take some time depending on the size of the video.
[1550] 3. Emotion recognition function
[1551] This function analyzes the user's facial expressions, tone of voice, words, etc. in real time to recognize their emotions, making it possible to determine the level of stress or joy the user is feeling.
[1552] Server Features
[1553] 1. Video reception and storage function
[1554] The server receives and stores cooking videos uploaded by users. Metadata (e.g., shooting date and time, user ID) is also stored in the video file, making it easier to analyze and search later.
[1555] 2. Video analysis function
[1556] The server analyzes the stored cooking videos frame by frame using computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow). The analysis extracts cooking steps, utensils, and ingredients, and records them in a database.
[1557] 3. Differential analysis function
[1558] The server compares the user's cooking video with professional recipe videos to identify differences in cooking procedures, utensils, and ingredients. This analysis uses comparison algorithms and natural language processing techniques. The identified differences are stored in a database.
[1559] 4. Feedback generation function
[1560] Based on the results of the differential analysis, the server generates specific feedback and advice for the user, sometimes using a generative AI model (e.g., GPT-3) to automatically generate feedback with specific improvements.
[1561] 5. Emotion-adaptive feedback function
[1562] The server monitors the user's emotional state in real time and adjusts the content and tone of the feedback and advice it provides, for example adding an encouraging message if the user is feeling stressed.
[1563] 6. Composite video generation function
[1564] The system displays professional recipe videos and user cooking videos side by side, and generates a composite video incorporating feedback based on the analysis results. Comments about the user's feelings may also be added to this composite video.
[1565] Specific examples
[1566] For example, suppose a user takes a video of themselves making carbonara and uploads it to a server. The server receives the video and begins analyzing it. The analysis results identify the following differences:
[1567] Bacon cut differently
[1568] Mixing the eggs too late
[1569] A lot of salt
[1570] Based on this, the following specific advice is generated:
[1571] Bacon should be thinly sliced
[1572] Mix the eggs quickly
[1573] Reduce the amount of salt by half
[1574] Additionally, if the emotion engine determines that the user is feeling anxious or stressed while cooking, it will add encouraging messages and advice on how to relax.
[1575] Prompt Sentence Examples
[1576] "I'd like specific feedback on my steps in making this carbonara. I'd especially like some advice on when to mix the eggs and how much salt to use."
[1577] In this way, users can watch the composite video, compare their own movements with those of a professional, understand what improvements are needed, and improve their cooking skills at home.
[1578] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1579] Step 1:
[1580] A user shoots a cooking video.
[1581] Specific operation: The user uses the camera on their smartphone or tablet to film themselves cooking at home, adjusting the camera angle and lighting to ensure a clear video.
[1582] Input: The user's cooking scene.
[1583] Output: Cooking video file.
[1584] Step 2:
[1585] The device uploads the cooking video to the server.
[1586] Specific operation: The user instructs the device to upload the video they have taken to the server using a dedicated application or web browser. The device then transfers the video file to the server via the Internet.
[1587] Input: Filmed cooking video files.
[1588] Output: Cooking video files uploaded to the server.
[1589] Step 3:
[1590] The server receives and stores the cooking video.
[1591] Specific operation: The server receives and stores the uploaded video file, along with metadata (e.g., shooting date and time, user ID).
[1592] Input: Uploaded cooking video file and metadata.
[1593] Output: Saved cooking video files and metadata.
[1594] Step 4:
[1595] The server analyzes the video.
[1596] How it works: The server analyzes the stored cooking videos frame by frame, and uses computer vision techniques and machine learning algorithms (e.g., OpenCV and TensorFlow) to extract cooking steps, utensils, and ingredients.
[1597] Input: Saved cooking video file.
[1598] Output: Extracted cooking steps, utensils, and ingredients data.
[1599] Step 5:
[1600] The server analyzes the differences with professional recipe videos.
[1601] Specific operation: The server compares the user's cooking video with professional recipe videos and uses comparison algorithms and natural language processing techniques to identify differences in cooking steps, utensils used, and ingredients.
[1602] Input: Extracted cooking instructions, utensils, and ingredient data; professional recipe videos.
[1603] Output: Differential analysis results.
[1604] Step 6:
[1605] The server generates feedback and advice.
[1606] Specific operation: The server generates feedback and advice for the user based on the results of the differential analysis. It uses a generative AI model (e.g., GPT-3) to automatically create feedback including specific improvements.
[1607] Input: Differential analysis results.
[1608] Output: Generated feedback and advice.
[1609] Step 7:
[1610] The server adjusts according to the user's emotional state.
[1611] Specific operation: The server uses emotion recognition to analyze the user's facial expressions, tone of voice, and words to recognize their emotions. It then adjusts the content and tone of the feedback and advice provided according to their emotional state.
[1612] Input: User facial expressions, tone of voice, and verbal data.
[1613] Output: Tailored feedback and advice.
[1614] Step 8:
[1615] The server generates the composite video.
[1616] How it works: The server displays professional recipe videos and the user's cooking videos side by side, and generates a composite video that includes feedback and advice based on the analysis results. It may also add comments about the user's feelings to the video.
[1617] Inputs: Tailored feedback and advice; professional recipe videos; user cooking videos.
[1618] Output: The generated composite video.
[1619] Step 9:
[1620] The user watches the composite video.
[1621] Specific operation: The user watches a composite video provided by the server on their device. While watching this composite video, they compare their own cooking with that of a professional and understand what improvements are needed.
[1622] Input: The generated synthetic video.
[1623] Output: Improved cooking skills of the user.
[1624] (Application example 2)
[1625] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1626] Conventional cooking feedback systems aim to improve users' cooking skills, but they lack functionality that takes users' emotions into consideration, which often causes stress and makes it difficult for users to continue using the system. Furthermore, they lack specific suggestions for improvements when reproducing professional recipes, making it difficult for users to effectively improve their own cooking skills.
[1627] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: a means for uploading cooking videos filmed by the user; a means for importing professional recipe videos; a means for analyzing the cooking videos and recipe videos and extracting their respective cooking steps, utensils, and ingredients; a means for comparing the cooking steps, utensils, and ingredients and analyzing the differences; a means for generating feedback and advice for the user based on the differences; a means for recognizing the user's emotions and adjusting the content and tone of the feedback and advice according to the emotions; a means for generating a composite video including the generated feedback and advice; and a means for providing the composite video to the user. This makes it possible to provide less stressful feedback according to the user's emotional state and present specific improvements.
[1628] A "cooking video filmed by a user" is a video recording of a user cooking using a device such as a smartphone or tablet.
[1629] "Professional recipe videos" are videos containing cooking steps shown by professional chefs or cooking experts, and are intended for users to use as reference.
[1630] "Means for analyzing cooking videos and recipe videos and extracting the respective cooking steps, utensils, and ingredients" refers to a technical means for analyzing uploaded cooking videos and recipe videos and automatically recognizing and extracting the cooking steps, utensils used, and ingredients from each frame.
[1631] The "means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences" refers to a technical means for comparing a user's cooking video with a professional recipe video and identifying differences in cooking procedures, utensils, and ingredients.
[1632] The "means for generating feedback and advice for the user based on the difference" is a technology for automatically generating feedback and specific advice that suggests improvements to the user's cooking method based on the results of the difference analysis.
[1633] "Means for recognizing a user's emotions and adjusting the content and tone of feedback or advice according to those emotions" refers to a technical means for analyzing a user's facial expressions, tone of voice, words, etc. to grasp their emotional state, and adjusting the content and tone of the feedback or advice provided according to those emotions.
[1634] The "means for generating a composite video including generated feedback and advice" is a technology for automatically generating a composite video that inserts comparison results and feedback based on a user's cooking video and a professional recipe video.
[1635] The "means for providing a composite video to a user" refers to a technical means for transmitting the generated composite video to a user's terminal so that the user can view the video.
[1636] This invention is a system that makes it easy for users to recreate professional recipes at home. It recognizes the user's emotions and provides feedback and advice according to those emotions. The system analyzes and synthesizes the user's cooking video and the professional recipe video, and identifies and provides specific differences.
[1637] System configuration
[1638] This system consists of user terminals, servers for processing information, and a network that connects them. The specific functions of each element are explained below.
[1639] User terminal
[1640] User terminals mainly include smartphones and tablets.
[1641] 1. Recording function: Equipped with a camera function that allows users to record videos of themselves cooking at home, for example, using a smartphone camera.
[1642] 2. Upload function: An application or web browser is provided for uploading the recorded cooking videos to the server.
[1643] 3. Emotion recognition function: Using the camera and microphone, the system analyzes the user's facial expressions, tone of voice, and words in real time to recognize the user's emotions.
[1644] server
[1645] The server analyzes the data sent from the user device and generates feedback and advice. Specific functions include:
[1646] 1. Video reception and storage function: Receives cooking videos uploaded by users and stores them on the server.
[1647] 2. Video analysis function: Analyzes cooking videos and professional recipe videos frame by frame to extract information on cooking steps, utensils, and ingredients. OpenCV and dlib are used here.
[1648] 3. Database function: Equipped with a database for managing information on extracted cooking procedures, utensils, and ingredients.
[1649] 4. Difference analysis function: Compares cooking videos with professional recipe videos and analyzes differences in cooking procedures, utensils, and ingredients. For example, it identifies specific differences such as the amount of salt being too high or the heating time being too short.
[1650] 5. Feedback generation function: Generates specific feedback and advice for users based on differential analysis. The software used is a generative AI model.
[1651] 6. Emotion-adaptive feedback: Using an emotion engine, the system monitors the user's emotional state and adjusts the tone and content of feedback and advice. For example, if the user is feeling stressed, the system adds advice on how to relax.
[1652] 7. Composite video generation function: Combines the user's cooking video with a professional recipe video to generate a composite video that includes feedback and advice.
[1653] Overview of operation
[1654] The user cooks, films the process with their smartphone, and uploads the video to the server via a dedicated app. The server receives the video and begins analysis. The extracted cooking steps, utensils, and ingredient information are compared with that of professional recipe videos to identify differences. Specific feedback and advice is generated based on the differences, and the system also monitors the user's emotional state and adapts the content and tone of the feedback. The professional recipe video and the user's cooking video are then combined to generate a composite video that includes feedback and advice. The composite video is finally provided to the user, who can watch it to improve their own cooking skills.
[1655] Specific examples
[1656] For example, a user can film themselves making carbonara and upload it via a dedicated app. The server receives the video and begins analyzing it, identifying discrepancies such as the way the bacon was cut, the timing of the eggs being added too late, and the amount of salt being too much. Based on this, specific advice is generated, such as "slicing the bacon thinner," "mixing the eggs earlier," and "reducing the amount of salt by half." Furthermore, if the user feels anxious or stressed while cooking, the emotion engine detects this and adds encouraging messages or advice on how to relax. This advice is then combined with a professional recipe video and provided to the user. The user can then watch the combined video, compare their own movements with those of the professional, and use the results to improve their cooking skills.
[1657] Prompt Sentence Examples
[1658] Users film themselves making carbonara and upload it to the app, where it will compare it with professional recipe videos and provide specific advice.
[1659] In this way, the system helps users easily recreate professional recipes at home and improve their cooking skills. It also provides emotional feedback to users, making the learning experience more comfortable and effective.
[1660] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1661] Step 1:
[1662] The user takes a photo of themselves cooking using a device (smartphone or tablet).
[1663] Input: User's cooking video (obtained in real time via the device camera)
[1664] Output: Cooking video file
[1665] Step 2:
[1666] Cooking videos taken by the device are uploaded to a server via a dedicated application.
[1667] Input: Cooking video file
[1668] Output: Cooking video data uploaded to the server
[1669] Step 3:
[1670] The server receives the uploaded cooking videos and stores them in a database.
[1671] Input: User's cooking video data
[1672] Output: Cooking video data stored in a database
[1673] Step 4:
[1674] The server retrieves professional recipe videos from a database.
[1675] Input: Professional recipe video data request
[1676] Output: Professional recipe video data extracted from the database
[1677] Step 5:
[1678] The server analyzes cooking and recipe videos frame by frame and extracts information on each cooking step, utensils used, and ingredients.
[1679] Input: Cooking video data and recipe video data
[1680] Data processing: Analyzes video frame by frame using OpenCV and dlib, automatically recognizing cooking steps, utensils, and ingredients
[1681] Output: Extracted cooking instructions, utensils, and ingredient data
[1682] Step 6:
[1683] The server compares the extracted cooking procedures, utensils, and ingredient data and analyzes the differences.
[1684] Input: User cooking instructions, equipment, and ingredient data; Professional recipe instructions, equipment, and ingredient data
[1685] Data calculation: Calculates differences in cooking procedures, utensils, and ingredients
[1686] Output: Difference data of cooking procedures, utensils, and ingredients
[1687] Step 7:
[1688] The server generates feedback and advice for the user based on the difference data.
[1689] Input: Cooking procedure, utensils, ingredient differential data
[1690] Data shaping: Using generative AI models to generate specific advice
[1691] Output: Specific feedback and advice to the user
[1692] Step 8:
[1693] The server recognizes the user's emotions and adjusts the content and tone of the feedback and advice according to the emotions.
[1694] Input: User's emotional data (facial expressions, tone of voice, and word analysis results)
[1695] Data Computation: Feedback Modulation Based on Emotion Engine
[1696] Output: Emotionally tailored feedback and advice
[1697] Step 9:
[1698] The server generates a composite video that includes the generated feedback and advice.
[1699] Input: Adjusted feedback and advice data, user cooking video data, professional recipe video data
[1700] Data processing: Combining cooking videos and recipe videos, and inserting feedback
[1701] Output: Generated composite video
[1702] Step 10:
[1703] The server provides the generated composite video to the user's terminal.
[1704] Input: Synthetic video data
[1705] Output: Composite video sent to the user's device
[1706] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1707] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1708] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1709] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1710] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1711] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1712] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1713] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1714] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1715] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1716] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1717] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1718] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1719] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1720] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1721] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1722] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1723] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1724] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1725] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1726] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1727] The following is further disclosed regarding the above embodiment.
[1728] (Claim 1)
[1729] A means for uploading cooking videos taken by users;
[1730] A way to import professional recipe videos,
[1731] A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos;
[1732] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences;
[1733] means for generating feedback and advice for a user based on the difference;
[1734] means for generating a composite video including the generated feedback and advice;
[1735] means for providing the composite video to a user;
[1736] A system including:
[1737] (Claim 2)
[1738] 2. The system according to claim 1, wherein the analyzing means analyzes the video frame by frame and extracts cooking steps, utensils, and ingredients.
[1739] (Claim 3)
[1740] 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements tailored to the user.
[1741] "Example 1"
[1742] (Claim 1)
[1743] A means for uploading cooking videos taken by users;
[1744] A way to import professional recipe videos,
[1745] A means for analyzing the cooking videos and the recipe videos using a video analysis library and extracting the cooking steps, utensils, and ingredients of each of the cooking videos;
[1746] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences using a deep learning model;
[1747] means for automatically generating feedback and advice for a user based on the difference;
[1748] means for generating a composite video including the generated feedback and advice using video editing software;
[1749] means for providing the composite video to a user;
[1750] A system including:
[1751] (Claim 2)
[1752] 2. The system according to claim 1, wherein the analyzing means analyzes the video frame by frame and extracts cooking steps, utensils, and ingredients.
[1753] (Claim 3)
[1754] 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements tailored to the user using a generative AI model.
[1755] "Application Example 1"
[1756] (Claim 1)
[1757] A means for uploading cooking videos taken by users;
[1758] A way to import cooking recipe videos,
[1759] A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos;
[1760] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences;
[1761] means for generating feedback and advice for a user based on the difference;
[1762] means for displaying the generated feedback and advice in real time;
[1763] means for generating a composite video including said feedback and advice;
[1764] means for providing the composite video to a user;
[1765] A system including:
[1766] (Claim 2)
[1767] 2. The system according to claim 1, wherein the analyzing means analyzes the video frame by frame and extracts cooking steps, utensils, and ingredients.
[1768] (Claim 3)
[1769] 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements tailored to the user.
[1770] "Example 2: Combining Emotion Engines"
[1771] (Claim 1)
[1772] A means for uploading cooking videos taken by users;
[1773] A way to import professional recipe videos,
[1774] A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos;
[1775] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences;
[1776] means for generating feedback and advice for a user based on the difference;
[1777] means for generating a composite video including the generated feedback and advice;
[1778] means for providing the composite video to a user;
[1779] A means to recognize user emotions in real time and adjust the content and tone of feedback and advice;
[1780] A system including:
[1781] (Claim 2)
[1782] 2. The system according to claim 1, wherein the analyzing means analyzes the video frame by frame and extracts cooking steps, utensils, and ingredients.
[1783] (Claim 3)
[1784] 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements tailored to the user.
[1785] (Claim 4)
[1786] 2. The system of claim 1, wherein the emotion recognition means analyzes the user's facial expressions, tone of voice, and words.
[1787] (Claim 5)
[1788] 10. The system of claim 1, wherein said emotionally adaptive feedback means monitors the user's emotional state and adjusts the tone and content of the feedback and advice.
[1789] (Claim 6)
[1790] The system according to claim 1, wherein the composite video generation means displays a professional recipe video and a user's cooking video side by side and generates a composite video in which feedback according to the difference is inserted.
[1791] "Application example 2 when combining emotion engines"
[1792] (Claim 1)
[1793] A means for uploading cooking videos taken by users;
[1794] A way to import professional recipe videos,
[1795] A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos;
[1796] means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences;
[1797] means for generating feedback and advice for a user based on the difference;
[1798] A means for recognizing a user's emotions and adjusting the content and tone of feedback and advice according to the emotions;
[1799] means for generating a composite video including the generated feedback and advice;
[1800] means for providing the composite video to a user;
[1801] A system including:
[1802] (Claim 2)
[1803] 2. The system according to claim 1, wherein the analyzing means analyzes the video frame by frame and extracts cooking steps, utensils, and ingredients.
[1804] (Claim 3)
[1805] 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements tailored to the user, and adjusts the tone and content of the feedback according to emotions. [Explanation of symbols]
[1806] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for uploading cooking videos taken by users; A way to import professional recipe videos, A means for analyzing the cooking video and the recipe video and extracting cooking procedures, utensils, and ingredients for each of the cooking videos; means for comparing the cooking procedures, utensils, and ingredients and analyzing the differences; means for generating feedback and advice for a user based on the difference; means for generating a composite video including the generated feedback and advice; means for providing the composite video to a user; A system including:
2. 2. The system according to claim 1, wherein said analyzing means analyzes the video frame by frame to extract cooking steps, utensils and ingredients.
3. 2. The system of claim 1, wherein the feedback and advice generating means automatically generates feedback including specific improvements specific to the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A