system

The system addresses the inefficiencies in cooking by analyzing images to suggest recipes and automate cooking processes, improving efficiency and motivation through personalized guidance.

JP2026068446APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing cooking systems lack effective utilization of available ingredients, guidance for combining ingredients, and automation of cooking equipment, leading to decreased motivation and efficiency, especially for beginners.

Method used

A system that analyzes user images to identify ingredients, suggests culinary foods, guides cooking procedures through audio and video, and automates cooking processes with compatible equipment.

Benefits of technology

Enhances cooking efficiency, increases motivation, and maximizes the functionality of cooking equipment by effectively utilizing available materials and providing personalized guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068446000001_ABST
    Figure 2026068446000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for analyzing images sent by users to identify materials contained within those images, A means for proposing a cookable food based on the aforementioned materials, A means for providing audio and video guidance on the preparation procedure of the aforementioned food, A system that includes means for performing automated cooking in conjunction with cooking equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In cooking, there is a problem in that there is a lack of proposals for recipes that effectively utilize available ingredients and specific guides for supporting the cooking process. Also, especially for beginners, it is often unclear how to combine ingredients and which procedures to follow for cooking, which leads to a decrease in the motivation for cooking. Furthermore, while the automation of cooking equipment is progressing, its potential capabilities are often not fully utilized. It is desired to solve these problems.

Means for Solving the Problems

[0005] This invention solves this problem by providing a means for analyzing images transmitted by the user and identifying materials contained within those images. Furthermore, it provides a means for suggesting culinary foods based on the identified materials and a means for guiding the user through the cooking procedure of the food selected by the user using voice and video, thereby supporting the cooking process. In addition, it improves cooking efficiency by providing a means for automatic cooking in cooperation with cooking equipment. In this way, it achieves effective utilization of available materials, increases motivation to cook, and maximizes the functionality of cooking equipment.

[0006] "User" refers to an individual or group that uses the system to receive cooking assistance.

[0007] "Image" refers to visual information that a user sends to the system, and is data used to identify materials.

[0008] "Ingredients" refers to the elements used in cooking, such as food ingredients and seasonings.

[0009] "Analysis" refers to the process of information processing performed to identify the materials within an image.

[0010] "Food" refers to food that is prepared through cooking, and means what is obtained as a result of combining ingredients.

[0011] "Suggestions" refer to the menu options presented to the user based on the analysis results.

[0012] "Cooking procedure" refers to the sequence of operations and steps necessary to complete a food product.

[0013] "Audio and video" guidance refers to providing auditory and visual information to help users understand the cooking procedure.

[0014] "Cooking equipment" refers to electrical appliances and machinery used for cooking.

[0015] "Automatic cooking" refers to the system controlling cooking equipment to automatically perform the cooking process.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is a system that uses images to provide support to users in cooking by effectively utilizing available materials. An embodiment thereof is specifically described below.

[0038] This system involves a series of steps where the user sends an image to a server via a terminal, the image is analyzed, suggestions are made, and cooking is guided.

[0039] First, the user takes pictures of ingredients and seasonings they have on hand using the device's camera and uploads the images to the server via the application. After receiving the images, the server uses deep learning technology to analyze them and identify the ingredients. The analyzed ingredient information is then compared with a food database, and based on the user's preferences and past history, the system suggests foods that can be cooked.

[0040] The user selects a menu item from several food options displayed on the device. This selection information is sent back to the server, which retrieves the cooking instructions for the selected food and downloads them to the device as audio and video content. The device then plays the received audio and video, guiding the user to intuitively follow the cooking process.

[0041] Furthermore, the server works in conjunction with cooking appliances to automatically adjust the temperature and time according to pre-set recipes, streamlining the cooking process. For example, when a user is making pasta, the smart oven automatically starts heating to the appropriate temperature, and the timer function is integrated. This is especially helpful when trying out new recipes.

[0042] This system also incorporates an advertising display function, allowing it to provide users with product information and promotions related to the suggested recipes. By analyzing users' cooking history and preferences, it improves the accuracy of advertisements and enables more personalized recommendations. Furthermore, future plans include integration with health management programs, aiming to provide suggestions that contribute to users' health maintenance.

[0043] For example, if a user takes photos of tomatoes, carrots, and onions at home and uploads them, the server analyzes them and suggests dishes like minestrone or vegetable soup. If the user selects minestrone, detailed recipe instructions, along with voice guidance, are provided on the device, making cooking smooth. Furthermore, if compatible cooking equipment is available, the heating time is automatically set, minimizing user intervention.

[0044] In this way, the present invention provides users with a new cooking experience and achieves increased efficiency and convenience in cooking.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The user uses their device's camera to photograph ingredients and seasonings they have on hand. They then review the captured images within the app and prepare them for uploading to the server.

[0048] Step 2:

[0049] The device sends the uploaded image data to the server. During this process, the image data undergoes format conversion and compression as needed.

[0050] Step 3:

[0051] The server inputs the received image data into an image analysis module to identify the types and characteristics of ingredients and seasonings. This process utilizes machine learning algorithms.

[0052] Step 4:

[0053] Based on the analysis results, the server consults a food database and identifies multiple recipes that can be generated from the available ingredients. In doing so, it also takes into account the user's past selection history, preferences, and allergy information.

[0054] Step 5:

[0055] The server sends a list of suggested dishes to the terminal. The terminal displays this information on its user interface, making it available for the user to select.

[0056] Step 6:

[0057] The user selects the dish they want to make from the dish options displayed on their device. Once the selection is complete, that information is sent back to the server.

[0058] Step 7:

[0059] The server retrieves audio guides and video content, including cooking instructions for the selected dish, and sends them to the terminal.

[0060] Step 8:

[0061] The device plays back the received audio and video, guiding the user through the cooking process using both audio and video. The guidance is timed to each step.

[0062] Step 9:

[0063] The server works in conjunction with smart cooking appliances, sending commands to automate operations necessary for the cooking process—such as when to start and stop heating the oven.

[0064] Step 10:

[0065] After cooking is complete, the server updates the user's cooking history and saves it as data to help improve future services and the accuracy of ad displays.

[0066] (Example 1)

[0067] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0068] In cooking, it is essential to effectively utilize the ingredients available to the user and to easily provide a variety of cooking options. Furthermore, it is desirable to streamline the cooking process so that even beginners or users with limited time can cook intuitively. Additionally, providing food suggestions tailored to individual preferences and health conditions is a crucial challenge.

[0069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0070] In this invention, the server includes means for analyzing visual information transmitted by the user to identify materials contained in the visual information, means for suggesting culinary foods based on the materials, and means for providing auditory and visual guidance on the cooking procedure of the foods. This allows the user to enjoy a variety of dishes without wasting any of the ingredients they have on hand, maximizing the efficiency and effectiveness of cooking. Furthermore, this also includes providing personalized food suggestions based on the user's preferences and health condition, supporting optimal cooking and health management tailored to individual needs.

[0071] "Visual information" refers to image data and other visual data captured by the user.

[0072] "Ingredients" refers to substances, including food and seasonings, used in cooking.

[0073] "Cookable food" refers to food that is generated based on analyzed ingredients and suggested to users so that they can actually cook and consume it.

[0074] "Auditory and visual guidance" refers to providing users with information and instructions through audio guides and video content.

[0075] "Cooking equipment" refers to devices or equipment used for cooking, including those with automatic temperature and time control functions.

[0076] "Automated cooking" refers to a process in which cooking equipment performs cooking according to a set program, without requiring direct human intervention.

[0077] "Comparing analyzed visual information with stored data" refers to the process of comparing information obtained through image analysis with existing databases to confirm the consistency and relevance of the information.

[0078] "Personalized product and promotional information" refers to advertisements and sales information provided based on the user's preferences and usage history.

[0079] "Health management information" refers to data such as the user's health status and nutritional balance, which is taken into consideration when making cooking suggestions.

[0080] "Automatic control of specified temperature and time" refers to a function in which the cooking device autonomously proceeds with cooking based on the set temperature and time.

[0081] The system of the present invention streamlines cooking based on the ingredients the user possesses, providing a richer dining experience. The embodiments are described in detail below.

[0082] The user uses their device's camera to photograph materials they have on hand. For example, if using materials such as tomatoes, carrots, and onions, the user photographs them and sends the image data to the server through a pre-configured application. This application runs on smartphones and tablets.

[0083] The server uses deep learning technology to analyze the received image data. Frameworks such as TENSORFLOW® and PyTorch are employed for this process, and calculations are performed to identify the materials within the image. Subsequently, the identified materials are compared with a food database, and several cookable food items are suggested, taking into account the user's preferences and past transaction history.

[0084] The terminal notifies the user of suggestions received from the server and visually displays the recipe. When the user selects a menu item of interest, the server generates detailed cooking instructions for that menu item and provides them as audio guides and video content. This content is delivered in streaming format using libraries such as FFmpeg.

[0085] Furthermore, the cooking device and server are linked via communication, allowing for automatic control of the temperature and cooking time required for each dish. For example, when cooking minestrone, the smart oven is set to the appropriate temperature and heating continues for the specified time. This significantly reduces the effort users spend on cooking.

[0086] This system also has the function of providing personalized advertisements and promotional information to individual users. By analyzing the user's cooking history and preference information, it suggests highly relevant products and services.

[0087] As a concrete example, a possible scenario is where, based on images of tomatoes, carrots, and onions taken by the user, the server suggests minestrone or vegetable soup, and if the user selects minestrone, the terminal provides a detailed recipe.

[0088] An example of a prompt message would be, "Upload images of tomatoes, carrots, and onions taken with your device, and then display the recipe suggestions after the server has analyzed them."

[0089] In this way, the present invention can support the user's cooking activities and provide a new cooking experience.

[0090] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0091] Step 1:

[0092] The user takes a picture of an object in their possession using the device's camera. Inputs include image files of items such as tomatoes, carrots, and onions. The device then sends this image data to the server via the application. The output is the image data sent to the server.

[0093] Step 2:

[0094] The server analyzes the received image data. The input is image data sent by the user. The server uses a generative AI model to analyze the image and perform data processing to identify the material. Deep learning technology is applied in this process. The output is the identified material information.

[0095] Step 3:

[0096] The server compares the identified material information with the stored database. The input is the material information as a result of the analysis, and the server performs an operation to compare it with the existing food database. The output is a list of related foods.

[0097] Step 4:

[0098] The server suggests foods that can be cooked based on the user's preferences and past history. The input consists of a list of ingredients and user history data. The server analyzes the information and performs data calculations using a recommendation algorithm. The output is a list of suggested recipes.

[0099] Step 5:

[0100] The terminal displays the suggested recipes in the user interface. The input is recipe information provided by the server. The terminal performs specific actions to visualize the recipes and notify the user. The output is the recipe options displayed to the user.

[0101] Step 6:

[0102] The user selects a menu item of interest from the presented recipes. The input is the recipe options displayed on the screen. Based on the user's selection, the output will be the information for the specific recipe selected.

[0103] Step 7:

[0104] The server generates detailed cooking instructions based on the selected recipe. The input is the recipe information selected by the user. The server processes the data to create the cooking instructions as audio guides and video content. The output is the generated cooking instruction content data.

[0105] Step 8:

[0106] The terminal plays audio guides and video content received from the server. The input is cooking procedure data sent from the server. The terminal activates its playback function and provides guidance to the user. The output is visual and auditory information of the cooking instructions provided to the user.

[0107] Step 9:

[0108] The server communicates with the cooking device to automatically adjust the temperature and time. The input consists of the cooking parameters included in the selected recipe. The server then sends configuration information to the cooking device and performs specific actions to enable automatic control. The output is the adjusted heating settings and time management.

[0109] (Application Example 1)

[0110] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0111] Modern consumers struggle to find the optimal combination of ingredients for cooking from the various ingredients available in stores, and to devise dishes that efficiently utilize those ingredients. Furthermore, they lack opportunities to obtain useful product information related to the ingredients, and opportunities to receive easy-to-understand guidance on cooking procedures. This results in increased time and effort spent on shopping and cooking, reducing convenience. Therefore, there is a need for a means to simultaneously provide an environment where customers can easily receive cooking suggestions in stores, and to provide advertising information related to the ingredients.

[0112] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0113] In this invention, the server includes means for analyzing images transmitted by the user to identify materials contained in the images, means for suggesting cookable foods based on the materials, and means for displaying food suggestions based on the materials and advertisements for cooking equipment via electronic devices installed in the store. This makes it possible to provide in-store material selection and optimal cooking suggestions, as well as display advertisements for related products on the spot.

[0114] An "image" is a digital recording of visual information and serves as a source of information for visually identifying materials.

[0115] "Materials" refer to the raw ingredients used in food preparation, and are the subjects of analysis from the captured images.

[0116] "Cookable foods" refer to dishes and food / drinks that can be made using specific ingredients, and are the subject of the proposal.

[0117] A "suggestion" is the act of presenting users with a selection of foods that can be cooked, and is a process that encourages users to make choices.

[0118] "Audio and video" refers to functions that convey cooking procedures to the user, and is a means of providing information through both sight and sound.

[0119] An "electronic device" is a device that has the ability to input, process, and display information, and functions as an interface between the user and the system.

[0120] "Advertising" is a means of providing information aimed at informing users about related products and services and increasing their desire to purchase them.

[0121] "Automatic cooking" refers to a function in which cooking appliances independently execute cooking processes according to a pre-set program.

[0122] This system is designed to enhance the user's in-store shopping experience. It consists of servers, terminals, and related electronic devices, allowing users to receive on-the-spot guidance on everything from ingredient selection to cooking suggestions.

[0123] The server uses a software environment implementing deep learning technology (specifically TensorFlow and PyTorch) to analyze images captured by the user's device camera. This allows the server to identify the captured material and compare it with a food database to suggest foods that can be cooked. The suggested foods are then sent to the user's device as audio and video content, providing clear and easy-to-understand guidance. This guidance utilizes speech synthesis and video streaming technologies.

[0124] Furthermore, sales promotion will be achieved by displaying advertisements for products related to the suggested food items to users via electronic devices. AdTech platforms, such as Google® AdMob, will be used for this ad delivery.

[0125] For example, if a user takes a picture of chicken breast and broccoli in a store, the system will analyze the image and suggest a creamy chicken and broccoli dish. In this case, a prompt message such as, "Recognize the ingredients in the following image, generate a recipe for chicken breast and broccoli based on that, and also display advertisements for the necessary cooking utensils," will be sent to the server, enabling simultaneous display of cooking suggestions and advertisements related to the ingredients.

[0126] In this way, the in-store shopping and cooking experience becomes intuitive and efficient, resulting in a system that can improve user satisfaction.

[0127] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0128] Step 1:

[0129] The user uses their device to photograph items they are considering purchasing in the store. During this process, the device's camera acquires high-resolution image data, which is then sent to a server as image data via the application.

[0130] Step 2:

[0131] The server analyzes the received images using a deep learning model. The input images are preprocessed using frameworks such as TensorFlow and PyTorch to identify the source material. As a result of the analysis, source information is generated and compared with a database.

[0132] Step 3:

[0133] Based on the matching results, the server searches the food database for relevant cookable foods and generates suggestions. The input is the ingredient information obtained in step 2, and the output is a list of relevant recipes. This is sent to the terminal in a format that the user can select.

[0134] Step 4:

[0135] The user selects their preferred recipe from several displayed on the device. The selected recipe name is resent to the server and used for the next process.

[0136] Step 5:

[0137] The server generates audio and video guides for the cooking process based on the selected recipe. The input is the selected recipe information, and the output is audio and video files. These are downloaded to the terminal and can be viewed by the user. Speech synthesis and video streaming technologies are used.

[0138] Step 6:

[0139] The device plays back received audio and video data, providing users with easy-to-understand cooking information. It is designed to be intuitively usable through its interface.

[0140] Step 7:

[0141] The server provides a mechanism to link suggested recipes with related product advertisements and display the most relevant advertisements to the user. The input is promotional information related to cookable foods, and the output is the content of the displayed advertisements.

[0142] This series of steps creates a system that allows users to smoothly proceed through the entire process, from selecting ingredients to cooking and purchasing necessary products.

[0143] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0144] This invention combines a system that analyzes images provided by the user to suggest cooking methods with an emotion engine that recognizes the user's emotions. This system enables the provision of more appropriate food suggestions and cooking guides based on the user's emotional state.

[0145] First, the user takes a picture of the ingredients or seasonings they want to use using the device's camera. The captured image is sent to the server via the app. The device is equipped with emotion recognition capabilities, which analyze the user's conversation, facial expressions, and voice to identify their emotions. This emotion information is also sent to the server along with the image data.

[0146] When the server receives an image, it uses an image analysis module to identify the material. In addition, the emotion engine evaluates the received emotional information to determine the user's current emotional state. This information is referenced when making food recommendations. For example, if the emotional state is positive, a new cooking method will be suggested; if it is negative, a recipe prioritizing convenience will be presented.

[0147] Multiple suggested cooking options are sent to the device, and the user makes a selection on the device. Depending on the selected menu, the server sends audio and video guidance to the device, providing the user with visual and auditory instructions. The device plays this back, providing the user with both visual and auditory guidance. The emotion engine can also continuously obtain user feedback during cooking and provide advice as needed.

[0148] Furthermore, the server works in conjunction with smart cooking devices to automatically adjust cooking parameters according to the progress of each cooking step. In this way, a cooking experience tailored to the user's emotional state is provided, supporting a more satisfying cooking experience. For example, if the user's emotions suggest stress, a simple and relaxing cooking method will be suggested.

[0149] This system also records the user's emotions and cooking data, which will be used as training data to improve the accuracy of food recommendations and emotion recognition in the future. As a result, it will be possible to provide a personalized experience for each user, reducing the burden of cooking and increasing the enjoyment.

[0150] In this way, the present invention realizes cooking support that takes emotional states into account, and provides new added value to the user.

[0151] The following describes the processing flow.

[0152] Step 1:

[0153] The user uses their device's camera to photograph the currently available ingredients and seasonings. The captured images are then prepared to be sent to the server via the app's interface.

[0154] Step 2:

[0155] The device analyzes the user's facial expressions and voice using an emotion recognition sensor. The obtained emotion data is transmitted to the server in parallel with the image data.

[0156] Step 3:

[0157] The server inputs the received image data into an image analysis module to identify ingredients and seasonings. This information is used as basic data to suggest foods that can be cooked.

[0158] Step 4:

[0159] The server uses an emotion engine to analyze the received emotion data and identify the user's current emotional state. Based on this emotional state, it adjusts food recommendations.

[0160] Step 5:

[0161] The server matches the user's emotional state with ingredient information to generate suitable dish suggestions. When making suggestions, it takes into account the difficulty level and cooking time based on the user's emotional state.

[0162] Step 6:

[0163] The terminal displays a list of recipe suggestions sent from the server on the user interface. The user selects the dish they want to make from this list.

[0164] Step 7:

[0165] Based on the user's selection, the server retrieves the cooking instructions for the selected dish and prepares the associated audio guide and visual instructional video.

[0166] Step 8:

[0167] The terminal plays audio guides and videos provided by the server, guiding the user through the cooking steps. These guides include emotionally responsive tips.

[0168] Step 9:

[0169] The server receives emotional feedback from the user and controls the smart cooking equipment according to the progress of cooking, optimizing the automated cooking process.

[0170] Step 10:

[0171] Once cooking is complete, the server records the user's emotional data and cooking data, which are then used as training data to improve the accuracy of future suggestions.

[0172] (Example 2)

[0173] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0174] There is a need for a system that can improve user satisfaction by providing optimal food suggestions and cooking experiences tailored to the user's emotional state when preparing food. Existing technologies lack the flexibility to consider changes in the user's emotions and individual needs when providing food suggestions and cooking guides, making it difficult to address the individual emotions of users.

[0175] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0176] In this invention, the server includes means for analyzing visual data transmitted by the user to identify ingredients, means for recognizing the user's emotional state and suggesting food, means for guiding the user through the food preparation process using voice and video, means for performing automatic cooking in cooperation with a cooking device, and means for recording the user's emotional information and cooking data to improve the accuracy of future food suggestions. This makes it possible to suggest food that is appropriate to the user's emotional state, thereby providing a highly satisfying cooking experience.

[0177] "Visual data" refers to visual information such as images and videos, which is information provided by the user through the camera of their device.

[0178] "Ingredients" refers to raw materials used in food preparation, and are food components identified through specific image analysis.

[0179] "Emotional state" refers to the user's current psychological state and feelings, and is information revealed by emotion recognition technology.

[0180] "Food suggestions" refers to the act of providing appropriate cooking ideas and menus based on the user's emotional state and the ingredients available.

[0181] "Audio" and "video" refer to means of transmitting information in auditory and visual forms, respectively, and are used as means of guiding the cooking procedure.

[0182] A "cooking device" refers to a machine or technology that automatically prepares food, and is controlled in conjunction with a server.

[0183] "Recommendation accuracy" refers to the degree to which the food and cooking method suggestions made to the user accurately match the user's needs and circumstances.

[0184] This invention combines a system that analyzes visual data provided by the user to suggest cooking methods with a function that recognizes the user's emotional state. The user takes pictures of the ingredients needed for cooking using the camera on their device. The device uses an emotion recognition engine to identify the user's emotional state by analyzing voice, facial expressions, and conversation data. This information is transmitted to the server along with the visual data. The server is equipped with an advanced image analysis module and an emotion recognition engine, which use these tools to identify the ingredients and evaluate the user's emotional state.

[0185] Specifically, the user will use an electronic terminal equipped with a camera. This terminal will be equipped with an application that utilizes an AI model to perform real-time image processing and emotion recognition. On the server side, a powerful computer will run an image analysis module and an emotion engine, and perform data processing using AI algorithms. The software will utilize AI technologies, including image recognition and natural language processing, to enable food recommendations tailored to the user's needs.

[0186] A concrete example demonstrating the advantages of this system is when a user takes a picture of tomatoes and basil and the system determines that the user is feeling relaxed; in such a case, a simple tomato and basil pasta recipe is suggested. The user selects the menu on their device, and the server guides them through the cooking process using voice and video. During this process, advice and adjustments are made based on the progress of the cooking and the user's emotional feedback.

[0187] Furthermore, user emotional information and cooking data are used to improve the accuracy of future suggestions. In this way, a personalized experience is provided, resulting in a less stressful cooking method for the user.

[0188] An example of a prompt message is, "Suggest a dish based on the image and provide a cooking guide that takes the user's emotions into consideration." This is how you can instruct the system. Based on this prompt, the generative AI model will provide the most suitable cooking suggestion for the user.

[0189] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0190] Step 1:

[0191] The user uses the device's camera to capture visual data of the ingredients to be used in cooking. At this stage, the input is physical material (e.g., tomatoes or basil), which is captured as a digital image on the device. The device then saves the captured image within the application.

[0192] Step 2:

[0193] The device uses a built-in emotion recognition engine to analyze the user's voice and facial expressions in real time. In this process, voice data and facial expression data are used as input, and the emotional state (for example, an emotion indicating stress) is output. Emotion recognition technology quantifies and categorizes emotions for recording.

[0194] Step 3:

[0195] The device transmits captured visual data and emotional state data to the server. The input for this communication consists of image data and emotional data, which are passed to the server via the network. The transmitted data serves as the basis for analysis processing on the server.

[0196] Step 4:

[0197] The server inputs the received visual data into an image analysis module to identify the materials. For example, it uses image recognition technology to determine if the material is a tomato or basil. At this stage, the input is image data, and the output is information about the identified materials. The analysis results are used as a basis for deciding which dishes to suggest to the user.

[0198] Step 5:

[0199] The server analyzes the emotional information received by the emotion recognition engine and evaluates the user's emotional state. The input is emotional data, and the output is the specific emotional category the user is currently experiencing (e.g., "I want to relax"). This information is used to select what kind of meal to suggest.

[0200] Step 6:

[0201] The server generates optimal food suggestions based on information about the ingredients and the evaluation results of the emotional state. For example, if the emotion of wanting to relax is detected, a simple tomato and basil pasta dish will be suggested. The input is the analysis results of ingredient data and emotional data, and the output is a cooking suggestion.

[0202] Step 7:

[0203] The server prepares the suggested food preparation steps as audio and video guides and sends them to the terminal. In this step, the specific preparation steps based on the suggested dish are input and output as visual and auditory instructions to the user. The terminal receives this and provides the guide to the user.

[0204] Step 8:

[0205] The device continuously acquires user feedback and additional emotional data during cooking. This enables real-time advice and adjustments. Inputs are user feedback and new emotional data, while outputs are further advice and procedural modifications as needed.

[0206] (Application Example 2)

[0207] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0208] Modern consumers are increasingly seeking personalized suggestions for meals and food delivery that reflect their individual emotional states and preferences. Traditional systems fail to consider the user's emotional state when suggesting food, resulting in a lack of appropriate choices. Furthermore, cooking instructions and automated cooking processes are not optimized based on the user's emotional state, limiting the improvement in satisfaction.

[0209] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0210] In this invention, the server includes means for analyzing images transmitted by the user to identify objects contained within the images, means for analyzing emotions and optimizing food suggestions based on the emotional state, and means for providing audio and video guidance on the cooking procedure for the items. This enables personalized food suggestions and appropriate cooking guidance tailored to the user's emotional state.

[0211] A "user" is an individual or group that uses this system to send images and receive food suggestions and cooking assistance.

[0212] An "image" is data that records visual information containing an object, transmitted by the user.

[0213] "Analysis" refers to a processing method that extracts specific information from transmitted images or data and performs recognition or classification.

[0214] "Object" refers to cooking ingredients or related items that are included in the image taken by the user.

[0215] "Emotion" is a concept that encompasses elements used to analyze and identify a user's emotional state.

[0216] "Emotional state" refers to analyzed information that reflects the user's inner mood and psychological state.

[0217] "Food suggestion" is the process of presenting users with selectable, culinary items based on the identification of objects and emotional states.

[0218] "Providing guidance through audio and video" means providing cooking instructions to the user using audio guides and visual presentations.

[0219] "Automated cooking" is a process in which food is cooked automatically in conjunction with cooking equipment, based on pre-set procedures.

[0220] This food delivery system uses programs that perform emotion analysis and image analysis to suggest food items based on the user's emotional state. The system primarily operates on devices such as smartphones and smart glasses, as well as on servers running on cloud services.

[0221] The server receives image and audio data captured by the user using the device's camera. This image data is used to identify objects through an analysis module. An example of a module used is an image recognition algorithm that utilizes TensorFlow.

[0222] Furthermore, the server uses an emotion analysis engine to understand the user's emotional state from voice and facial expression data. Emotion recognition APIs such as Microsoft® Azure® Cognitive Services can be used here. The analysis results reflect the user's current psychological state, and the suggested food items are optimized based on this data.

[0223] The terminal functions as a device for receiving food suggestions and providing cooking instructions via audio and video. This allows users to proceed with cooking using visual and auditory guidance. Furthermore, communication with the server enables the system to provide additional feedback when the user's emotional state changes and update suggestions as needed.

[0224] For example, if a user is feeling like "I want a relaxing meal today," the system might suggest a salad suitable as a light snack. An example of inputting such prompts into an AI model might be: "Based on the user's voice recording and photo data, identify the user's current mood. Then, generate a list of food recommendations based on that mood."

[0225] In this way, the system can understand the user's emotional state, make food suggestions accordingly, and provide a personalized dining experience for each individual user.

[0226] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0227] Step 1:

[0228] The user takes pictures of ingredients and seasonings using the device's camera.

[0229] Image data is generated as input, and the terminal sends this image data to the server.

[0230] The transmitted image data will serve as input for the next analysis.

[0231] Step 2:

[0232] The server receives the image data and uses an image analysis module to identify objects within the image.

[0233] The input is image data, and object identification is performed using image recognition algorithms such as TensorFlow.

[0234] The output is a list of identified objects.

[0235] Step 3:

[0236] The server receives the user's voice data or facial expression data acquired by the camera, and analyzes their emotional state through an emotion analysis engine.

[0237] The input is either audio data or facial expression images, which are analyzed using emotion recognition APIs such as Microsoft Azure Cognitive Services.

[0238] The output is data that indicates the user's emotional state.

[0239] Step 4:

[0240] The server uses a food suggestion algorithm to propose the best food based on the identified object and emotional state.

[0241] The input consists of a list of objects and emotional state data, and an AI model generates behavioral recommendations based on this data.

[0242] The output is a list of food product suggestions.

[0243] Step 5:

[0244] The terminal receives a list of suggestions from the server and presents the suggestions to the user visually and audibly.

[0245] The input is a list of food suggestions, and the output is a selection of options displayed to the user.

[0246] In terms of specific actions, a menu is displayed on the screen, and suggestions are explained with voice guidance.

[0247] Step 6:

[0248] When the user makes a selection from the suggested food items, the terminal sends the selection information to the server.

[0249] The input is the user's selection information, and the selected food items are sent to the server as output.

[0250] Step 7:

[0251] The server generates cooking instructions based on the selected food item and sends specific audio and video guides to the terminal.

[0252] The input is the selected food item, and the output is audio and video data of the cooking procedure.

[0253] This process helps the user easily perform the next cooking step.

[0254] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0255] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0256] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0257] [Second Embodiment]

[0258] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0259] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0260] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0261] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0262] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0263] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0264] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0265] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0266] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0267] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0268] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0269] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0270] This invention is a system that uses images to provide support to users in cooking by effectively utilizing available materials. An embodiment thereof is specifically described below.

[0271] This system involves a series of steps where the user sends an image to a server via a terminal, the image is analyzed, suggestions are made, and cooking is guided.

[0272] First, the user takes pictures of ingredients and seasonings they have on hand using the device's camera and uploads the images to the server via the application. After receiving the images, the server uses deep learning technology to analyze them and identify the ingredients. The analyzed ingredient information is then compared with a food database, and based on the user's preferences and past history, the system suggests foods that can be cooked.

[0273] The user selects a menu item from several food options displayed on the device. This selection information is sent back to the server, which retrieves the cooking instructions for the selected food and downloads them to the device as audio and video content. The device then plays the received audio and video, guiding the user to intuitively follow the cooking process.

[0274] Furthermore, the server works in conjunction with cooking appliances to automatically adjust the temperature and time according to pre-set recipes, streamlining the cooking process. For example, when a user is making pasta, the smart oven automatically starts heating to the appropriate temperature, and the timer function is integrated. This is especially helpful when trying out new recipes.

[0275] This system also incorporates an advertising display function, allowing it to provide users with product information and promotions related to the suggested recipes. By analyzing users' cooking history and preferences, it improves the accuracy of advertisements and enables more personalized recommendations. Furthermore, future plans include integration with health management programs, aiming to provide suggestions that contribute to users' health maintenance.

[0276] For example, if a user takes photos of tomatoes, carrots, and onions at home and uploads them, the server analyzes them and suggests dishes like minestrone or vegetable soup. If the user selects minestrone, detailed recipe instructions, along with voice guidance, are provided on the device, making cooking smooth. Furthermore, if compatible cooking equipment is available, the heating time is automatically set, minimizing user intervention.

[0277] In this way, the present invention provides a new cooking experience to the user and realizes the improvement of cooking efficiency and convenience.

[0278] The following describes the process flow.

[0279] Step 1:

[0280] The user uses the camera of the terminal to take pictures of the materials and seasonings at hand. Check the taken image on the app and get ready to upload it to the server.

[0281] Step 2:

[0282] The terminal sends the uploaded image data to the server. At this time, the image data is format-converted and compressed as necessary.

[0283] Step 3:

[0284] The server inputs the received image data into the image analysis module to identify the types and characteristics of the materials and seasonings. A machine learning algorithm is used for this process.

[0285] Step 4:

[0286] Based on the analysis results, the server refers to the food database and selects a plurality of cooking recipes that can be generated from the available materials. At this time, the user's past selection history, preferences, and allergy information are also considered.

[0287] Step 5:

[0288] The server sends the list of proposed dishes to the terminal. The terminal displays this information on the user interface and makes it selectable by the user.

[0289] Step 6:

[0290] The user selects the dish they want to make from the dish options displayed on their device. Once the selection is complete, that information is sent back to the server.

[0291] Step 7:

[0292] The server retrieves audio guides and video content, including cooking instructions for the selected dish, and sends them to the terminal.

[0293] Step 8:

[0294] The device plays back the received audio and video, guiding the user through the cooking process using both audio and video. The guidance is timed to each step.

[0295] Step 9:

[0296] The server works in conjunction with smart cooking appliances, sending commands to automate operations necessary for the cooking process—such as when to start and stop heating the oven.

[0297] Step 10:

[0298] After cooking is complete, the server updates the user's cooking history and saves it as data to help improve future services and the accuracy of ad displays.

[0299] (Example 1)

[0300] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0301] In cooking, it is essential to effectively utilize the ingredients available to the user and to easily provide a variety of cooking options. Furthermore, it is desirable to streamline the cooking process so that even beginners or users with limited time can cook intuitively. Additionally, providing food suggestions tailored to individual preferences and health conditions is a crucial challenge.

[0302] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in the first embodiment is realized by the following means.

[0303] In this invention, the server includes means for analyzing visual information transmitted from a user and identifying materials contained in the visual information, means for proposing cookable foods based on the materials, and means for guiding the cooking procedures of the foods both audibly and visually. Thereby, the user can enjoy various dishes without wasting the materials at hand, and can maximize the efficiency and effect of cooking. Furthermore, this also includes providing individualized food proposals based on the user's preferences and health conditions, and supports optimal cooking and health management according to each demand.

[0304] "Visual information" refers to image data captured by the user or other visual data.

[0305] "Materials" refer to substances including foods and seasonings used in cooking.

[0306] "Cookable foods" refer to foods that are generated based on the analyzed materials and are proposed so that the user can actually cook and consume them.

[0307] "Guiding both audibly and visually" refers to providing information and instructions to the user through voice guides and video contents.

[0308] "Cooking apparatus" refers to equipment or devices for cooking, including those having an automatic control function for temperature and time.

[0309] "Automated cooking" refers to a process in which a cooking apparatus cooks according to a set program without the need for direct human operation.

[0310] "Comparing analyzed visual information with stored data" refers to the process of comparing information obtained through image analysis with existing databases to confirm the consistency and relevance of the information.

[0311] "Personalized product and promotional information" refers to advertisements and sales information provided based on the user's preferences and usage history.

[0312] "Health management information" refers to data such as the user's health status and nutritional balance, which is taken into consideration when making cooking suggestions.

[0313] "Automatic control of specified temperature and time" refers to a function in which the cooking device autonomously proceeds with cooking based on the set temperature and time.

[0314] The system of the present invention streamlines cooking based on the ingredients the user possesses, providing a richer dining experience. The embodiments are described in detail below.

[0315] The user uses their device's camera to photograph materials they have on hand. For example, if using materials such as tomatoes, carrots, and onions, the user photographs them and sends the image data to the server through a pre-configured application. This application runs on smartphones and tablets.

[0316] The server uses deep learning technology to analyze the received image data. Frameworks such as TensorFlow and PyTorch are employed for this process, and calculations are performed to identify the materials within the image. Subsequently, the identified materials are compared with a food database, and several cookable food items are suggested, taking into account the user's preferences and past transaction history.

[0317] The terminal notifies the user of suggestions received from the server and visually displays the recipe. When the user selects a menu item of interest, the server generates detailed cooking instructions for that menu item and provides them as audio guides and video content. This content is delivered in streaming format using libraries such as FFmpeg.

[0318] Furthermore, the cooking device and server are linked via communication, allowing for automatic control of the temperature and cooking time required for each dish. For example, when cooking minestrone, the smart oven is set to the appropriate temperature and heating continues for the specified time. This significantly reduces the effort users spend on cooking.

[0319] This system also has the function of providing personalized advertisements and promotional information to individual users. By analyzing the user's cooking history and preference information, it suggests highly relevant products and services.

[0320] As a concrete example, a possible scenario is where, based on images of tomatoes, carrots, and onions taken by the user, the server suggests minestrone or vegetable soup, and if the user selects minestrone, the terminal provides a detailed recipe.

[0321] An example of a prompt message would be, "Upload images of tomatoes, carrots, and onions taken with your device, and then display the recipe suggestions after the server has analyzed them."

[0322] In this way, the present invention can support the user's cooking activities and provide a new cooking experience.

[0323] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0324] Step 1:

[0325] The user takes a picture of an object in their possession using the device's camera. Inputs include image files of items such as tomatoes, carrots, and onions. The device then sends this image data to the server via the application. The output is the image data sent to the server.

[0326] Step 2:

[0327] The server analyzes the received image data. The input is image data sent by the user. The server uses a generative AI model to analyze the image and perform data processing to identify the material. Deep learning technology is applied in this process. The output is the identified material information.

[0328] Step 3:

[0329] The server compares the identified material information with the stored database. The input is the material information as a result of the analysis, and the server performs an operation to compare it with the existing food database. The output is a list of related foods.

[0330] Step 4:

[0331] The server suggests foods that can be cooked based on the user's preferences and past history. The input consists of a list of ingredients and user history data. The server analyzes the information and performs data calculations using a recommendation algorithm. The output is a list of suggested recipes.

[0332] Step 5:

[0333] The terminal displays the suggested recipes in the user interface. The input is recipe information provided by the server. The terminal performs specific actions to visualize the recipes and notify the user. The output is the recipe options displayed to the user.

[0334] Step 6:

[0335] The user selects a menu item of interest from the presented recipes. The input is the recipe options displayed on the screen. Based on the user's selection, the output will be the information for the specific recipe selected.

[0336] Step 7:

[0337] The server generates detailed cooking instructions based on the selected recipe. The input is the recipe information selected by the user. The server processes the data to create the cooking instructions as audio guides and video content. The output is the generated cooking instruction content data.

[0338] Step 8:

[0339] The terminal plays audio guides and video content received from the server. The input is cooking procedure data sent from the server. The terminal activates its playback function and provides guidance to the user. The output is visual and auditory information of the cooking instructions provided to the user.

[0340] Step 9:

[0341] The server communicates with the cooking device to automatically adjust the temperature and time. The input consists of the cooking parameters included in the selected recipe. The server then sends configuration information to the cooking device and performs specific actions to enable automatic control. The output is the adjusted heating settings and time management.

[0342] (Application Example 1)

[0343] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0344] Modern consumers struggle to find the optimal combination of ingredients for cooking from the various ingredients available in stores, and to devise dishes that efficiently utilize those ingredients. Furthermore, they lack opportunities to obtain useful product information related to the ingredients, and opportunities to receive easy-to-understand guidance on cooking procedures. This results in increased time and effort spent on shopping and cooking, reducing convenience. Therefore, there is a need for a means to simultaneously provide an environment where customers can easily receive cooking suggestions in stores, and to provide advertising information related to the ingredients.

[0345] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0346] In this invention, the server includes means for analyzing images transmitted by the user to identify materials contained in the images, means for suggesting cookable foods based on the materials, and means for displaying food suggestions based on the materials and advertisements for cooking equipment via electronic devices installed in the store. This makes it possible to provide in-store material selection and optimal cooking suggestions, as well as display advertisements for related products on the spot.

[0347] An "image" is a digital recording of visual information and serves as a source of information for visually identifying materials.

[0348] "Materials" refer to the raw ingredients used in food preparation, and are the subjects of analysis from the captured images.

[0349] "Cookable foods" refer to dishes and food / drinks that can be made using specific ingredients, and are the subject of the proposal.

[0350] A "suggestion" is the act of presenting users with a selection of foods that can be cooked, and is a process that encourages users to make choices.

[0351] "Audio and video" refers to functions that convey cooking procedures to the user, and is a means of providing information through both sight and sound.

[0352] An "electronic device" is a device that has the ability to input, process, and display information, and functions as an interface between the user and the system.

[0353] "Advertising" is a means of providing information aimed at informing users about related products and services and increasing their desire to purchase them.

[0354] "Automatic cooking" refers to a function in which cooking appliances independently execute cooking processes according to a pre-set program.

[0355] This system is designed to enhance the user's in-store shopping experience. It consists of servers, terminals, and related electronic devices, allowing users to receive on-the-spot guidance on everything from ingredient selection to cooking suggestions.

[0356] The server uses a software environment implementing deep learning technology (specifically TensorFlow and PyTorch) to analyze images captured by the user's device camera. This allows the server to identify the captured material and compare it with a food database to suggest foods that can be cooked. The suggested foods are then sent to the user's device as audio and video content, providing clear and easy-to-understand guidance. This guidance utilizes speech synthesis and video streaming technologies.

[0357] Furthermore, sales promotion will be achieved by displaying advertisements for products related to the suggested food items to users via electronic devices. AdTech platforms, such as Google AdMob, will be used for this ad delivery.

[0358] For example, if a user takes a picture of chicken breast and broccoli in a store, the system will analyze the image and suggest a creamy chicken and broccoli dish. In this case, a prompt message such as, "Recognize the ingredients in the following image, generate a recipe for chicken breast and broccoli based on that, and also display advertisements for the necessary cooking utensils," will be sent to the server, enabling simultaneous display of cooking suggestions and advertisements related to the ingredients.

[0359] In this way, the in-store shopping and cooking experience becomes intuitive and efficient, resulting in a system that can improve user satisfaction.

[0360] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0361] Step 1:

[0362] The user uses their device to photograph items they are considering purchasing in the store. During this process, the device's camera acquires high-resolution image data, which is then sent to a server as image data via the application.

[0363] Step 2:

[0364] The server analyzes the received images using a deep learning model. The input images are preprocessed using frameworks such as TensorFlow and PyTorch to identify the source material. As a result of the analysis, source information is generated and compared with a database.

[0365] Step 3:

[0366] Based on the matching results, the server searches the food database for relevant cookable foods and generates suggestions. The input is the ingredient information obtained in step 2, and the output is a list of relevant recipes. This is sent to the terminal in a format that the user can select.

[0367] Step 4:

[0368] The user selects their preferred recipe from several displayed on the device. The selected recipe name is resent to the server and used for the next process.

[0369] Step 5:

[0370] The server generates audio and video guides for the cooking process based on the selected recipe. The input is the selected recipe information, and the output is audio and video files. These are downloaded to the terminal and can be viewed by the user. Speech synthesis and video streaming technologies are used.

[0371] Step 6:

[0372] The device plays back received audio and video data, providing users with easy-to-understand cooking information. It is designed to be intuitively usable through its interface.

[0373] Step 7:

[0374] The server provides a mechanism to link suggested recipes with related product advertisements and display the most relevant advertisements to the user. The input is promotional information related to cookable foods, and the output is the content of the displayed advertisements.

[0375] This series of steps creates a system that allows users to smoothly proceed through the entire process, from selecting ingredients to cooking and purchasing necessary products.

[0376] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0377] This invention combines a system that analyzes images provided by the user to suggest cooking methods with an emotion engine that recognizes the user's emotions. This system enables the provision of more appropriate food suggestions and cooking guides based on the user's emotional state.

[0378] First, the user takes a picture of the ingredients or seasonings they want to use using the device's camera. The captured image is sent to the server via the app. The device is equipped with emotion recognition capabilities, which analyze the user's conversation, facial expressions, and voice to identify their emotions. This emotion information is also sent to the server along with the image data.

[0379] When the server receives an image, it uses an image analysis module to identify the material. In addition, the emotion engine evaluates the received emotional information to determine the user's current emotional state. This information is referenced when making food recommendations. For example, if the emotional state is positive, a new cooking method will be suggested; if it is negative, a recipe prioritizing convenience will be presented.

[0380] Multiple suggested cooking options are sent to the device, and the user makes a selection on the device. Depending on the selected menu, the server sends audio and video guidance to the device, providing the user with visual and auditory instructions. The device plays this back, providing the user with both visual and auditory guidance. The emotion engine can also continuously obtain user feedback during cooking and provide advice as needed.

[0381] Furthermore, the server works in conjunction with smart cooking devices to automatically adjust cooking parameters according to the progress of each cooking step. In this way, a cooking experience tailored to the user's emotional state is provided, supporting a more satisfying cooking experience. For example, if the user's emotions suggest stress, a simple and relaxing cooking method will be suggested.

[0382] This system also records the user's emotions and cooking data, which will be used as training data to improve the accuracy of food recommendations and emotion recognition in the future. As a result, it will be possible to provide a personalized experience for each user, reducing the burden of cooking and increasing the enjoyment.

[0383] In this way, the present invention realizes cooking support that takes emotional states into account, and provides new added value to the user.

[0384] The following describes the processing flow.

[0385] Step 1:

[0386] The user uses their device's camera to photograph the currently available ingredients and seasonings. The captured images are then prepared to be sent to the server via the app's interface.

[0387] Step 2:

[0388] The device analyzes the user's facial expressions and voice using an emotion recognition sensor. The obtained emotion data is transmitted to the server in parallel with the image data.

[0389] Step 3:

[0390] The server inputs the received image data into an image analysis module to identify ingredients and seasonings. This information is used as basic data to suggest foods that can be cooked.

[0391] Step 4:

[0392] The server uses an emotion engine to analyze the received emotion data and identify the user's current emotional state. Based on this emotional state, it adjusts food recommendations.

[0393] Step 5:

[0394] The server matches the user's emotional state with ingredient information to generate suitable dish suggestions. When making suggestions, it takes into account the difficulty level and cooking time based on the user's emotional state.

[0395] Step 6:

[0396] The terminal displays a list of recipe suggestions sent from the server on the user interface. The user selects the dish they want to make from this list.

[0397] Step 7:

[0398] Based on the user's selection, the server retrieves the cooking instructions for the selected dish and prepares the associated audio guide and visual instructional video.

[0399] Step 8:

[0400] The terminal plays audio guides and videos provided by the server, guiding the user through the cooking steps. These guides include emotionally responsive tips.

[0401] Step 9:

[0402] The server receives emotional feedback from the user and controls the smart cooking equipment according to the progress of cooking, optimizing the automated cooking process.

[0403] Step 10:

[0404] Once cooking is complete, the server records the user's emotional data and cooking data, which are then used as training data to improve the accuracy of future suggestions.

[0405] (Example 2)

[0406] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0407] There is a need for a system that can improve user satisfaction by providing optimal food suggestions and cooking experiences tailored to the user's emotional state when preparing food. Existing technologies lack the flexibility to consider changes in the user's emotions and individual needs when providing food suggestions and cooking guides, making it difficult to address the individual emotions of users.

[0408] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0409] In this invention, the server includes means for analyzing visual data transmitted by the user to identify ingredients, means for recognizing the user's emotional state and suggesting food, means for guiding the user through the food preparation process using voice and video, means for performing automatic cooking in cooperation with a cooking device, and means for recording the user's emotional information and cooking data to improve the accuracy of future food suggestions. This makes it possible to suggest food that is appropriate to the user's emotional state, thereby providing a highly satisfying cooking experience.

[0410] "Visual data" refers to visual information such as images and videos, which is information provided by the user through the camera of their device.

[0411] "Ingredients" refers to raw materials used in food preparation, and are food components identified through specific image analysis.

[0412] "Emotional state" refers to the user's current psychological state and feelings, and is information revealed by emotion recognition technology.

[0413] "Food suggestions" refers to the act of providing appropriate cooking ideas and menus based on the user's emotional state and the ingredients available.

[0414] "Audio" and "video" refer to means of transmitting information in auditory and visual forms, respectively, and are used as means of guiding the cooking procedure.

[0415] A "cooking device" refers to a machine or technology that automatically prepares food, and is controlled in conjunction with a server.

[0416] "Recommendation accuracy" refers to the degree to which the food and cooking method suggestions made to the user accurately match the user's needs and circumstances.

[0417] This invention combines a system that analyzes visual data provided by the user to suggest cooking methods with a function that recognizes the user's emotional state. The user takes pictures of the ingredients needed for cooking using the camera on their device. The device uses an emotion recognition engine to identify the user's emotional state by analyzing voice, facial expressions, and conversation data. This information is transmitted to the server along with the visual data. The server is equipped with an advanced image analysis module and an emotion recognition engine, which use these tools to identify the ingredients and evaluate the user's emotional state.

[0418] Specifically, the user will use an electronic terminal equipped with a camera. This terminal will be equipped with an application that utilizes an AI model to perform real-time image processing and emotion recognition. On the server side, a powerful computer will run an image analysis module and an emotion engine, and perform data processing using AI algorithms. The software will utilize AI technologies, including image recognition and natural language processing, to enable food recommendations tailored to the user's needs.

[0419] A concrete example demonstrating the advantages of this system is when a user takes a picture of tomatoes and basil and the system determines that the user is feeling relaxed; in such a case, a simple tomato and basil pasta recipe is suggested. The user selects the menu on their device, and the server guides them through the cooking process using voice and video. During this process, advice and adjustments are made based on the progress of the cooking and the user's emotional feedback.

[0420] Furthermore, user emotional information and cooking data are used to improve the accuracy of future suggestions. In this way, a personalized experience is provided, resulting in a less stressful cooking method for the user.

[0421] An example of a prompt message is, "Suggest a dish based on the image and provide a cooking guide that takes the user's emotions into consideration." This is how you can instruct the system. Based on this prompt, the generative AI model will provide the most suitable cooking suggestion for the user.

[0422] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0423] Step 1:

[0424] The user uses the device's camera to capture visual data of the ingredients to be used in cooking. At this stage, the input is physical material (e.g., tomatoes or basil), which is captured as a digital image on the device. The device then saves the captured image within the application.

[0425] Step 2:

[0426] The device uses a built-in emotion recognition engine to analyze the user's voice and facial expressions in real time. In this process, voice data and facial expression data are used as input, and the emotional state (for example, an emotion indicating stress) is output. Emotion recognition technology quantifies and categorizes emotions for recording.

[0427] Step 3:

[0428] The device transmits captured visual data and emotional state data to the server. The input for this communication consists of image data and emotional data, which are passed to the server via the network. The transmitted data serves as the basis for analysis processing on the server.

[0429] Step 4:

[0430] The server inputs the received visual data into an image analysis module to identify the materials. For example, it uses image recognition technology to determine if the material is a tomato or basil. At this stage, the input is image data, and the output is information about the identified materials. The analysis results are used as a basis for deciding which dishes to suggest to the user.

[0431] Step 5:

[0432] The server analyzes the emotional information received by the emotion recognition engine and evaluates the user's emotional state. The input is emotional data, and the output is the specific emotional category the user is currently experiencing (e.g., "I want to relax"). This information is used to select what kind of meal to suggest.

[0433] Step 6:

[0434] The server generates optimal food suggestions based on information about the ingredients and the evaluation results of the emotional state. For example, if the emotion of wanting to relax is detected, a simple tomato and basil pasta dish will be suggested. The input is the analysis results of ingredient data and emotional data, and the output is a cooking suggestion.

[0435] Step 7:

[0436] The server prepares the suggested food preparation steps as audio and video guides and sends them to the terminal. In this step, the specific preparation steps based on the suggested dish are input and output as visual and auditory instructions to the user. The terminal receives this and provides the guide to the user.

[0437] Step 8:

[0438] The device continuously acquires user feedback and additional emotional data during cooking. This enables real-time advice and adjustments. Inputs are user feedback and new emotional data, while outputs are further advice and procedural modifications as needed.

[0439] (Application Example 2)

[0440] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0441] Modern consumers are increasingly seeking personalized suggestions for meals and food delivery that reflect their individual emotional states and preferences. Traditional systems fail to consider the user's emotional state when suggesting food, resulting in a lack of appropriate choices. Furthermore, cooking instructions and automated cooking processes are not optimized based on the user's emotional state, limiting the improvement in satisfaction.

[0442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0443] In this invention, the server includes means for analyzing images transmitted by the user to identify objects contained within the images, means for analyzing emotions and optimizing food suggestions based on the emotional state, and means for providing audio and video guidance on the cooking procedure for the items. This enables personalized food suggestions and appropriate cooking guidance tailored to the user's emotional state.

[0444] A "user" is an individual or group that uses this system to send images and receive food suggestions and cooking assistance.

[0445] An "image" is data that records visual information containing an object, transmitted by the user.

[0446] "Analysis" refers to a processing method that extracts specific information from transmitted images or data and performs recognition or classification.

[0447] "Object" refers to cooking ingredients or related items that are included in the image taken by the user.

[0448] "Emotion" is a concept that encompasses elements used to analyze and identify a user's emotional state.

[0449] "Emotional state" refers to analyzed information that reflects the user's inner mood and psychological state.

[0450] "Food suggestion" is the process of presenting users with selectable, culinary items based on the identification of objects and emotional states.

[0451] "Providing guidance through audio and video" means providing cooking instructions to the user using audio guides and visual presentations.

[0452] "Automated cooking" is a process in which food is cooked automatically in conjunction with cooking equipment, based on pre-set procedures.

[0453] This food delivery system uses programs that perform emotion analysis and image analysis to suggest food items based on the user's emotional state. The system primarily operates on devices such as smartphones and smart glasses, as well as on servers running on cloud services.

[0454] The server receives image and audio data captured by the user using the device's camera. This image data is used to identify objects through an analysis module. An example of a module used is an image recognition algorithm that utilizes TensorFlow.

[0455] Furthermore, the server uses an emotion analysis engine to understand the user's emotional state from voice and facial expression data. Emotion recognition APIs such as Microsoft Azure Cognitive Services can be used here. The analysis results reflect the user's current psychological state, and the suggested food items are optimized based on this data.

[0456] The terminal functions as a device for receiving food suggestions and providing cooking instructions via audio and video. This allows users to proceed with cooking using visual and auditory guidance. Furthermore, communication with the server enables the system to provide additional feedback when the user's emotional state changes and update suggestions as needed.

[0457] For example, if a user is feeling like "I want a relaxing meal today," the system might suggest a salad suitable as a light snack. An example of inputting such prompts into an AI model might be: "Based on the user's voice recording and photo data, identify the user's current mood. Then, generate a list of food recommendations based on that mood."

[0458] In this way, the system can understand the user's emotional state, make food suggestions accordingly, and provide a personalized dining experience for each individual user.

[0459] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0460] Step 1:

[0461] The user takes pictures of ingredients and seasonings using the device's camera.

[0462] Image data is generated as input, and the terminal sends this image data to the server.

[0463] The transmitted image data will serve as input for the next analysis.

[0464] Step 2:

[0465] The server receives the image data and uses an image analysis module to identify objects within the image.

[0466] The input is image data, and object identification is performed using image recognition algorithms such as TensorFlow.

[0467] The output is a list of identified objects.

[0468] Step 3:

[0469] The server receives the user's voice data or facial expression data acquired by the camera, and analyzes their emotional state through an emotion analysis engine.

[0470] The input is either audio data or facial expression images, which are analyzed using emotion recognition APIs such as Microsoft Azure Cognitive Services.

[0471] The output is data that indicates the user's emotional state.

[0472] Step 4:

[0473] The server uses a food suggestion algorithm to propose the best food based on the identified object and emotional state.

[0474] The input consists of a list of objects and emotional state data, and an AI model generates behavioral recommendations based on this data.

[0475] The output is a list of food product suggestions.

[0476] Step 5:

[0477] The terminal receives a list of suggestions from the server and presents the suggestions to the user visually and audibly.

[0478] The input is a list of food suggestions, and the output is a selection of options displayed to the user.

[0479] In terms of specific actions, the system displays a menu on the screen and explains the suggestions with voice guidance.

[0480] Step 6:

[0481] When the user makes a selection from the suggested food items, the terminal sends the selection information to the server.

[0482] The input is the user's selection information, and the selected food items are sent to the server as output.

[0483] Step 7:

[0484] The server generates cooking instructions based on the selected food item and sends specific audio and video guides to the terminal.

[0485] The input is the selected food item, and the output is audio and video data of the cooking procedure.

[0486] This process helps the user easily perform the next cooking step.

[0487] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0488] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0489] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0490] [Third Embodiment]

[0491] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0492] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0493] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0494] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0495] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0496] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0497] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0498] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0499] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0500] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0501] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0502] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0503] This invention is a system that uses images to provide support to users in cooking by effectively utilizing available materials. An embodiment thereof is specifically described below.

[0504] This system involves a series of steps where the user sends an image to a server via a terminal, the image is analyzed, suggestions are made, and cooking is guided.

[0505] First, the user takes pictures of ingredients and seasonings they have on hand using the device's camera and uploads the images to the server via the application. After receiving the images, the server uses deep learning technology to analyze them and identify the ingredients. The analyzed ingredient information is then compared with a food database, and based on the user's preferences and past history, the system suggests foods that can be cooked.

[0506] The user selects a menu item from several food options displayed on the device. This selection information is sent back to the server, which retrieves the cooking instructions for the selected food and downloads them to the device as audio and video content. The device then plays the received audio and video, guiding the user to intuitively follow the cooking process.

[0507] Furthermore, the server works in conjunction with cooking appliances to automatically adjust the temperature and time according to pre-set recipes, streamlining the cooking process. For example, when a user is making pasta, the smart oven automatically starts heating to the appropriate temperature, and the timer function is integrated. This is especially helpful when trying out new recipes.

[0508] This system also incorporates an advertising display function, allowing it to provide users with product information and promotions related to the suggested recipes. By analyzing users' cooking history and preferences, it improves the accuracy of advertisements and enables more personalized recommendations. Furthermore, future plans include integration with health management programs, aiming to provide suggestions that contribute to users' health maintenance.

[0509] For example, if a user takes photos of tomatoes, carrots, and onions at home and uploads them, the server analyzes them and suggests dishes like minestrone or vegetable soup. If the user selects minestrone, detailed recipe instructions, along with voice guidance, are provided on the device, making cooking smooth. Furthermore, if compatible cooking equipment is available, the heating time is automatically set, minimizing user intervention.

[0510] In this way, the present invention provides users with a new cooking experience and achieves increased efficiency and convenience in cooking.

[0511] The following describes the processing flow.

[0512] Step 1:

[0513] The user uses their device's camera to photograph ingredients and seasonings they have on hand. They then review the captured images within the app and prepare them for uploading to the server.

[0514] Step 2:

[0515] The device sends the uploaded image data to the server. During this process, the image data undergoes format conversion and compression as needed.

[0516] Step 3:

[0517] The server inputs the received image data into an image analysis module to identify the types and characteristics of ingredients and seasonings. This process utilizes machine learning algorithms.

[0518] Step 4:

[0519] Based on the analysis results, the server consults a food database and identifies multiple recipes that can be generated from the available ingredients. In doing so, it also takes into account the user's past selection history, preferences, and allergy information.

[0520] Step 5:

[0521] The server sends a list of suggested dishes to the terminal. The terminal displays this information on its user interface, making it available for the user to select.

[0522] Step 6:

[0523] The user selects the dish they want to make from the dish options displayed on their device. Once the selection is complete, that information is sent back to the server.

[0524] Step 7:

[0525] The server retrieves audio guides and video content, including cooking instructions for the selected dish, and sends them to the terminal.

[0526] Step 8:

[0527] The device plays back the received audio and video, guiding the user through the cooking process using both audio and video. The guidance is timed to each step.

[0528] Step 9:

[0529] The server works in conjunction with smart cooking appliances, sending commands to automate operations necessary for the cooking process—such as when to start and stop heating the oven.

[0530] Step 10:

[0531] After cooking is complete, the server updates the user's cooking history and saves it as data to help improve future services and the accuracy of ad displays.

[0532] (Example 1)

[0533] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0534] In cooking, it is essential to effectively utilize the ingredients available to the user and to easily provide a variety of cooking options. Furthermore, it is desirable to streamline the cooking process so that even beginners or users with limited time can cook intuitively. Additionally, providing food suggestions tailored to individual preferences and health conditions is a crucial challenge.

[0535] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0536] In this invention, the server includes means for analyzing visual information transmitted by the user to identify materials contained in the visual information, means for suggesting culinary foods based on the materials, and means for providing auditory and visual guidance on the cooking procedure of the foods. This allows the user to enjoy a variety of dishes without wasting any of the ingredients they have on hand, maximizing the efficiency and effectiveness of cooking. Furthermore, this also includes providing personalized food suggestions based on the user's preferences and health condition, supporting optimal cooking and health management tailored to individual needs.

[0537] "Visual information" refers to image data and other visual data captured by the user.

[0538] "Ingredients" refers to substances, including food and seasonings, used in cooking.

[0539] "Cookable food" refers to food that is generated based on analyzed ingredients and suggested to users so that they can actually cook and consume it.

[0540] "Auditory and visual guidance" refers to providing users with information and instructions through audio guides and video content.

[0541] "Cooking equipment" refers to devices or equipment used for cooking, including those with automatic temperature and time control functions.

[0542] "Automated cooking" refers to a process in which cooking equipment performs cooking according to a set program, without requiring direct human intervention.

[0543] "Comparing analyzed visual information with stored data" refers to the process of comparing information obtained through image analysis with existing databases to confirm the consistency and relevance of the information.

[0544] "Personalized product and promotional information" refers to advertisements and sales information provided based on the user's preferences and usage history.

[0545] "Health management information" refers to data such as the user's health status and nutritional balance, which is taken into consideration when making cooking suggestions.

[0546] "Automatic control of specified temperature and time" refers to a function in which the cooking device autonomously proceeds with cooking based on the set temperature and time.

[0547] The system of the present invention streamlines cooking based on the ingredients the user possesses, providing a richer dining experience. The embodiments are described in detail below.

[0548] The user uses their device's camera to photograph materials they have on hand. For example, if using materials such as tomatoes, carrots, and onions, the user photographs them and sends the image data to the server through a pre-configured application. This application runs on smartphones and tablets.

[0549] The server uses deep learning technology to analyze the received image data. Frameworks such as TensorFlow and PyTorch are employed for this process, and calculations are performed to identify the materials within the image. Subsequently, the identified materials are compared with a food database, and several cookable food items are suggested, taking into account the user's preferences and past transaction history.

[0550] The terminal notifies the user of suggestions received from the server and visually displays the recipe. When the user selects a menu item of interest, the server generates detailed cooking instructions for that menu item and provides them as audio guides and video content. This content is delivered in streaming format using libraries such as FFmpeg.

[0551] Furthermore, the cooking device and server are linked via communication, allowing for automatic control of the temperature and cooking time required for each dish. For example, when cooking minestrone, the smart oven is set to the appropriate temperature and heating continues for the specified time. This significantly reduces the effort users spend on cooking.

[0552] This system also has the function of providing personalized advertisements and promotional information to individual users. By analyzing the user's cooking history and preference information, it suggests highly relevant products and services.

[0553] As a concrete example, a possible scenario is where, based on images of tomatoes, carrots, and onions taken by the user, the server suggests minestrone or vegetable soup, and if the user selects minestrone, the terminal provides a detailed recipe.

[0554] An example of a prompt message would be, "Upload images of tomatoes, carrots, and onions taken with your device, and then display the recipe suggestions after the server has analyzed them."

[0555] In this way, the present invention can support the user's cooking activities and provide a new cooking experience.

[0556] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0557] Step 1:

[0558] The user takes a picture of an object in their possession using the device's camera. Inputs include image files of items such as tomatoes, carrots, and onions. The device then sends this image data to the server via the application. The output is the image data sent to the server.

[0559] Step 2:

[0560] The server analyzes the received image data. The input is image data sent by the user. The server uses a generative AI model to analyze the image and perform data processing to identify the material. Deep learning technology is applied in this process. The output is the identified material information.

[0561] Step 3:

[0562] The server compares the identified material information with the stored database. The input is the material information as a result of the analysis, and the server performs an operation to compare it with the existing food database. The output is a list of related foods.

[0563] Step 4:

[0564] The server suggests foods that can be cooked based on the user's preferences and past history. The input consists of a list of ingredients and user history data. The server analyzes the information and performs data calculations using a recommendation algorithm. The output is a list of suggested recipes.

[0565] Step 5:

[0566] The terminal displays the suggested recipes in the user interface. The input is recipe information provided by the server. The terminal performs specific actions to visualize the recipes and notify the user. The output is the recipe options displayed to the user.

[0567] Step 6:

[0568] The user selects a menu item of interest from the presented recipes. The input is the recipe options displayed on the screen. Based on the user's selection, the output will be the information for the specific recipe selected.

[0569] Step 7:

[0570] The server generates detailed cooking instructions based on the selected recipe. The input is the recipe information selected by the user. The server processes the data to create the cooking instructions as audio guides and video content. The output is the generated cooking instruction content data.

[0571] Step 8:

[0572] The terminal plays audio guides and video content received from the server. The input is cooking procedure data sent from the server. The terminal activates its playback function and provides guidance to the user. The output is visual and auditory information of the cooking instructions provided to the user.

[0573] Step 9:

[0574] The server communicates with the cooking device to automatically adjust the temperature and time. The input consists of the cooking parameters included in the selected recipe. The server then sends configuration information to the cooking device and performs specific actions to enable automatic control. The output is the adjusted heating settings and time management.

[0575] (Application Example 1)

[0576] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0577] Modern consumers struggle to find the optimal combination of ingredients for cooking from the various ingredients available in stores, and to devise dishes that efficiently utilize those ingredients. Furthermore, they lack opportunities to obtain useful product information related to the ingredients, and opportunities to receive easy-to-understand guidance on cooking procedures. This results in increased time and effort spent on shopping and cooking, reducing convenience. Therefore, there is a need for a means to simultaneously provide an environment where customers can easily receive cooking suggestions in stores, and to provide advertising information related to the ingredients.

[0578] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0579] In this invention, the server includes means for analyzing images transmitted by the user to identify materials contained in the images, means for suggesting cookable foods based on the materials, and means for displaying food suggestions based on the materials and advertisements for cooking equipment via electronic devices installed in the store. This makes it possible to provide in-store material selection and optimal cooking suggestions, as well as display advertisements for related products on the spot.

[0580] An "image" is a digital recording of visual information and serves as a source of information for visually identifying materials.

[0581] "Materials" refer to the raw ingredients used in food preparation, and are the subjects of analysis from the captured images.

[0582] "Cookable foods" refer to dishes and food / drinks that can be made using specific ingredients, and are the subject of the proposal.

[0583] A "suggestion" is the act of presenting users with a selection of foods that can be cooked, and is a process that encourages users to make choices.

[0584] "Audio and video" refers to functions that convey cooking procedures to the user, and is a means of providing information through both sight and sound.

[0585] An "electronic device" is a device that has the ability to input, process, and display information, and functions as an interface between the user and the system.

[0586] "Advertising" is a means of providing information aimed at informing users about related products and services and increasing their desire to purchase them.

[0587] "Automatic cooking" refers to a function in which cooking appliances independently execute cooking processes according to a pre-set program.

[0588] This system is designed to enhance the user's in-store shopping experience. It consists of servers, terminals, and related electronic devices, allowing users to receive on-the-spot guidance on everything from ingredient selection to cooking suggestions.

[0589] The server uses a software environment implementing deep learning technology (specifically TensorFlow and PyTorch) to analyze images captured by the user's device camera. This allows the server to identify the captured material and compare it with a food database to suggest foods that can be cooked. The suggested foods are then sent to the user's device as audio and video content, providing clear and easy-to-understand guidance. This guidance utilizes speech synthesis and video streaming technologies.

[0590] Furthermore, sales promotion will be achieved by displaying advertisements for products related to the suggested food items to users via electronic devices. AdTech platforms, such as Google AdMob, will be used for this ad delivery.

[0591] For example, if a user takes a picture of chicken breast and broccoli in a store, the system will analyze the image and suggest a creamy chicken and broccoli dish. In this case, a prompt message such as, "Recognize the ingredients in the following image, generate a recipe for chicken breast and broccoli based on that, and also display advertisements for the necessary cooking utensils," will be sent to the server, enabling simultaneous display of cooking suggestions and advertisements related to the ingredients.

[0592] In this way, the in-store shopping and cooking experience becomes intuitive and efficient, resulting in a system that can improve user satisfaction.

[0593] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0594] Step 1:

[0595] The user uses their device to photograph items they are considering purchasing in the store. During this process, the device's camera acquires high-resolution image data, which is then sent to a server as image data via the application.

[0596] Step 2:

[0597] The server analyzes the received images using a deep learning model. The input images are preprocessed using frameworks such as TensorFlow and PyTorch to identify the source material. As a result of the analysis, source information is generated and compared with a database.

[0598] Step 3:

[0599] Based on the matching results, the server searches the food database for relevant cookable foods and generates suggestions. The input is the ingredient information obtained in step 2, and the output is a list of relevant recipes. This is sent to the terminal in a format that the user can select.

[0600] Step 4:

[0601] The user selects their preferred recipe from several displayed on the device. The selected recipe name is resent to the server and used for the next process.

[0602] Step 5:

[0603] The server generates audio and video guides for the cooking process based on the selected recipe. The input is the selected recipe information, and the output is audio and video files. These are downloaded to the terminal and can be viewed by the user. Speech synthesis and video streaming technologies are used.

[0604] Step 6:

[0605] The device plays back received audio and video data, providing users with easy-to-understand cooking information. It is designed to be intuitively usable through its interface.

[0606] Step 7:

[0607] The server provides a mechanism to link suggested recipes with related product advertisements and display the most relevant advertisements to the user. The input is promotional information related to cookable foods, and the output is the content of the displayed advertisements.

[0608] This series of steps creates a system that allows users to smoothly proceed through the entire process, from selecting ingredients to cooking and purchasing necessary products.

[0609] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0610] This invention combines a system that analyzes images provided by the user to suggest cooking methods with an emotion engine that recognizes the user's emotions. This system enables the provision of more appropriate food suggestions and cooking guides based on the user's emotional state.

[0611] First, the user takes a picture of the ingredients or seasonings they want to use using the device's camera. The captured image is sent to the server via the app. The device is equipped with emotion recognition capabilities, which analyze the user's conversation, facial expressions, and voice to identify their emotions. This emotion information is also sent to the server along with the image data.

[0612] When the server receives an image, it uses an image analysis module to identify the material. In addition, the emotion engine evaluates the received emotional information to determine the user's current emotional state. This information is referenced when making food recommendations. For example, if the emotional state is positive, a new cooking method will be suggested; if it is negative, a recipe prioritizing convenience will be presented.

[0613] Multiple suggested cooking options are sent to the device, and the user makes a selection on the device. Depending on the selected menu, the server sends audio and video guidance to the device, providing the user with visual and auditory instructions. The device plays this back, providing the user with both visual and auditory guidance. The emotion engine can also continuously obtain user feedback during cooking and provide advice as needed.

[0614] Furthermore, the server works in conjunction with smart cooking devices to automatically adjust cooking parameters according to the progress of each cooking step. In this way, a cooking experience tailored to the user's emotional state is provided, supporting a more satisfying cooking experience. For example, if the user's emotions suggest stress, a simple and relaxing cooking method will be suggested.

[0615] This system also records the user's emotions and cooking data, which will be used as training data to improve the accuracy of food recommendations and emotion recognition in the future. As a result, it will be possible to provide a personalized experience for each user, reducing the burden of cooking and increasing the enjoyment.

[0616] In this way, the present invention realizes cooking support that takes emotional states into account, and provides new added value to the user.

[0617] The following describes the processing flow.

[0618] Step 1:

[0619] The user uses their device's camera to photograph the currently available ingredients and seasonings. The captured images are then prepared to be sent to the server via the app's interface.

[0620] Step 2:

[0621] The device analyzes the user's facial expressions and voice using an emotion recognition sensor. The obtained emotion data is transmitted to the server in parallel with the image data.

[0622] Step 3:

[0623] The server inputs the received image data into an image analysis module to identify ingredients and seasonings. This information is used as basic data to suggest foods that can be cooked.

[0624] Step 4:

[0625] The server uses an emotion engine to analyze the received emotion data and identify the user's current emotional state. Based on this emotional state, it adjusts food recommendations.

[0626] Step 5:

[0627] The server matches the user's emotional state with ingredient information to generate suitable dish suggestions. When making suggestions, it takes into account the difficulty level and cooking time based on the user's emotional state.

[0628] Step 6:

[0629] The terminal displays a list of recipe suggestions sent from the server on the user interface. The user selects the dish they want to make from this list.

[0630] Step 7:

[0631] Based on the user's selection, the server retrieves the cooking instructions for the selected dish and prepares the associated audio guide and visual instructional video.

[0632] Step 8:

[0633] The terminal plays audio guides and videos provided by the server, guiding the user through the cooking steps. These guides include emotionally responsive tips.

[0634] Step 9:

[0635] The server receives emotional feedback from the user and controls the smart cooking equipment according to the progress of cooking, optimizing the automated cooking process.

[0636] Step 10:

[0637] Once cooking is complete, the server records the user's emotional data and cooking data, which are then used as training data to improve the accuracy of future suggestions.

[0638] (Example 2)

[0639] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0640] There is a need for a system that can improve user satisfaction by providing optimal food suggestions and cooking experiences tailored to the user's emotional state when preparing food. Existing technologies lack the flexibility to consider changes in the user's emotions and individual needs when providing food suggestions and cooking guides, making it difficult to address the individual emotions of users.

[0641] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0642] In this invention, the server includes means for analyzing visual data transmitted by the user to identify ingredients, means for recognizing the user's emotional state and suggesting food, means for guiding the user through the food preparation process using voice and video, means for performing automatic cooking in cooperation with a cooking device, and means for recording the user's emotional information and cooking data to improve the accuracy of future food suggestions. This makes it possible to suggest food that is appropriate to the user's emotional state, thereby providing a highly satisfying cooking experience.

[0643] "Visual data" refers to visual information such as images and videos, which is information provided by the user through the camera of their device.

[0644] "Ingredients" refers to raw materials used in food preparation, and are food components identified through specific image analysis.

[0645] "Emotional state" refers to the user's current psychological state and feelings, and is information revealed by emotion recognition technology.

[0646] "Food suggestions" refers to the act of providing appropriate cooking ideas and menus based on the user's emotional state and the ingredients available.

[0647] "Audio" and "video" refer to means of transmitting information in auditory and visual forms, respectively, and are used as means of guiding the cooking procedure.

[0648] A "cooking device" refers to a machine or technology that automatically prepares food, and is controlled in conjunction with a server.

[0649] "Recommendation accuracy" refers to the degree to which the food and cooking method suggestions made to the user accurately match the user's needs and circumstances.

[0650] This invention combines a system that analyzes visual data provided by the user to suggest cooking methods with a function that recognizes the user's emotional state. The user takes pictures of the ingredients needed for cooking using the camera on their device. The device uses an emotion recognition engine to identify the user's emotional state by analyzing voice, facial expressions, and conversation data. This information is transmitted to the server along with the visual data. The server is equipped with an advanced image analysis module and an emotion recognition engine, which use these tools to identify the ingredients and evaluate the user's emotional state.

[0651] Specifically, the user will use an electronic terminal equipped with a camera. This terminal will be equipped with an application that utilizes an AI model to perform real-time image processing and emotion recognition. On the server side, a powerful computer will run an image analysis module and an emotion engine, and perform data processing using AI algorithms. The software will utilize AI technologies, including image recognition and natural language processing, to enable food recommendations tailored to the user's needs.

[0652] A concrete example demonstrating the advantages of this system is when a user takes a picture of tomatoes and basil and the system determines that the user is feeling relaxed; in such a case, a simple tomato and basil pasta recipe is suggested. The user selects the menu on their device, and the server guides them through the cooking process using voice and video. During this process, advice and adjustments are made based on the progress of the cooking and the user's emotional feedback.

[0653] Furthermore, user emotional information and cooking data are used to improve the accuracy of future suggestions. In this way, a personalized experience is provided, resulting in a less stressful cooking method for the user.

[0654] An example of a prompt message is, "Suggest a dish based on the image and provide a cooking guide that takes the user's emotions into consideration." This is how you can instruct the system. Based on this prompt, the generative AI model will provide the most suitable cooking suggestion for the user.

[0655] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0656] Step 1:

[0657] The user uses the device's camera to capture visual data of the ingredients to be used in cooking. At this stage, the input is physical material (e.g., tomatoes or basil), which is captured as a digital image on the device. The device then saves the captured image within the application.

[0658] Step 2:

[0659] The device uses a built-in emotion recognition engine to analyze the user's voice and facial expressions in real time. In this process, voice data and facial expression data are used as input, and the emotional state (for example, an emotion indicating stress) is output. Emotion recognition technology quantifies and categorizes emotions for recording.

[0660] Step 3:

[0661] The device transmits captured visual data and emotional state data to the server. The input for this communication consists of image data and emotional data, which are passed to the server via the network. The transmitted data serves as the basis for analysis processing on the server.

[0662] Step 4:

[0663] The server inputs the received visual data into an image analysis module to identify the materials. For example, it uses image recognition technology to determine if the material is a tomato or basil. At this stage, the input is image data, and the output is information about the identified materials. The analysis results are used as a basis for deciding which dishes to suggest to the user.

[0664] Step 5:

[0665] The server analyzes the emotional information received by the emotion recognition engine and evaluates the user's emotional state. The input is emotional data, and the output is the specific emotional category the user is currently experiencing (e.g., "I want to relax"). This information is used to select what kind of meal to suggest.

[0666] Step 6:

[0667] The server generates optimal food suggestions based on information about the ingredients and the evaluation results of the emotional state. For example, if the emotion of wanting to relax is detected, a simple tomato and basil pasta dish will be suggested. The input is the analysis results of ingredient data and emotional data, and the output is a cooking suggestion.

[0668] Step 7:

[0669] The server prepares the suggested food preparation steps as audio and video guides and sends them to the terminal. In this step, the specific preparation steps based on the suggested dish are input and output as visual and auditory instructions to the user. The terminal receives this and provides the guide to the user.

[0670] Step 8:

[0671] The device continuously acquires user feedback and additional emotional data during cooking. This enables real-time advice and adjustments. Inputs are user feedback and new emotional data, while outputs are further advice and procedural modifications as needed.

[0672] (Application Example 2)

[0673] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0674] Modern consumers are increasingly seeking personalized suggestions for meals and food delivery that reflect their individual emotional states and preferences. Traditional systems fail to consider the user's emotional state when suggesting food, resulting in a lack of appropriate choices. Furthermore, cooking instructions and automated cooking processes are not optimized based on the user's emotional state, limiting the improvement in satisfaction.

[0675] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0676] In this invention, the server includes means for analyzing images transmitted by the user to identify objects contained within the images, means for analyzing emotions and optimizing food suggestions based on the emotional state, and means for providing audio and video guidance on the cooking procedure for the items. This enables personalized food suggestions and appropriate cooking guidance tailored to the user's emotional state.

[0677] A "user" is an individual or group that uses this system to send images and receive food suggestions and cooking assistance.

[0678] An "image" is data that records visual information containing an object, transmitted by the user.

[0679] "Analysis" refers to a processing method that extracts specific information from transmitted images or data and performs recognition or classification.

[0680] "Object" refers to cooking ingredients or related items that are included in the image taken by the user.

[0681] "Emotion" is a concept that encompasses elements used to analyze and identify a user's emotional state.

[0682] "Emotional state" refers to analyzed information that reflects the user's inner mood and psychological state.

[0683] "Food suggestion" is the process of presenting users with selectable, culinary items based on the identification of objects and emotional states.

[0684] "Providing guidance through audio and video" means providing cooking instructions to the user using audio guides and visual presentations.

[0685] "Automated cooking" is a process in which food is cooked automatically in conjunction with cooking equipment, based on pre-set procedures.

[0686] This food delivery system uses programs that perform emotion analysis and image analysis to suggest food items based on the user's emotional state. The system primarily operates on devices such as smartphones and smart glasses, as well as on servers running on cloud services.

[0687] The server receives image and audio data captured by the user using the device's camera. This image data is used to identify objects through an analysis module. An example of a module used is an image recognition algorithm that utilizes TensorFlow.

[0688] Furthermore, the server uses an emotion analysis engine to understand the user's emotional state from voice and facial expression data. Emotion recognition APIs such as Microsoft Azure Cognitive Services can be used here. The analysis results reflect the user's current psychological state, and the suggested food items are optimized based on this data.

[0689] The terminal functions as a device for receiving food suggestions and providing cooking instructions via audio and video. This allows users to proceed with cooking using visual and auditory guidance. Furthermore, communication with the server enables the system to provide additional feedback when the user's emotional state changes and update suggestions as needed.

[0690] For example, if a user is feeling like "I want a relaxing meal today," the system might suggest a salad suitable as a light snack. An example of inputting such prompts into an AI model might be: "Based on the user's voice recording and photo data, identify the user's current mood. Then, generate a list of food recommendations based on that mood."

[0691] In this way, the system can understand the user's emotional state, make food suggestions accordingly, and provide a personalized dining experience for each individual user.

[0692] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0693] Step 1:

[0694] The user takes pictures of ingredients and seasonings using the device's camera.

[0695] Image data is generated as input, and the terminal sends this image data to the server.

[0696] The transmitted image data will serve as input for the next analysis.

[0697] Step 2:

[0698] The server receives the image data and uses an image analysis module to identify objects within the image.

[0699] The input is image data, and object identification is performed using image recognition algorithms such as TensorFlow.

[0700] The output is a list of identified objects.

[0701] Step 3:

[0702] The server receives the user's voice data or facial expression data acquired by the camera, and analyzes their emotional state through an emotion analysis engine.

[0703] The input is either audio data or facial expression images, which are analyzed using emotion recognition APIs such as Microsoft Azure Cognitive Services.

[0704] The output is data that indicates the user's emotional state.

[0705] Step 4:

[0706] The server uses a food suggestion algorithm to propose the best food based on the identified object and emotional state.

[0707] The input consists of a list of objects and emotional state data, and an AI model generates behavioral recommendations based on this data.

[0708] The output is a list of food product suggestions.

[0709] Step 5:

[0710] The terminal receives a list of suggestions from the server and presents the suggestions to the user visually and audibly.

[0711] The input is a list of food suggestions, and the output is a selection of options displayed to the user.

[0712] In terms of specific actions, a menu is displayed on the screen, and suggestions are explained with voice guidance.

[0713] Step 6:

[0714] When the user makes a selection from the suggested food items, the terminal sends the selection information to the server.

[0715] The input is the user's selection information, and the selected food items are sent to the server as output.

[0716] Step 7:

[0717] The server generates cooking instructions based on the selected food item and sends specific audio and video guides to the terminal.

[0718] The input is the selected food item, and the output is audio and video data of the cooking procedure.

[0719] This process helps the user easily perform the next cooking step.

[0720] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0721] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0722] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0723] [Fourth Embodiment]

[0724] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0725] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0726] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0727] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0728] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0729] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0730] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0731] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0732] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0733] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0734] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0735] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0736] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0737] This invention is a system that uses images to provide support to users in cooking by effectively utilizing available materials. An embodiment thereof is specifically described below.

[0738] This system involves a series of steps where the user sends an image to a server via a terminal, the image is analyzed, suggestions are made, and cooking is guided.

[0739] First, the user takes pictures of ingredients and seasonings they have on hand using the device's camera and uploads the images to the server via the application. After receiving the images, the server uses deep learning technology to analyze them and identify the ingredients. The analyzed ingredient information is then compared with a food database, and based on the user's preferences and past history, the system suggests foods that can be cooked.

[0740] The user selects a menu item from several food options displayed on the device. This selection information is sent back to the server, which retrieves the cooking instructions for the selected food and downloads them to the device as audio and video content. The device then plays the received audio and video, guiding the user to intuitively follow the cooking process.

[0741] Furthermore, the server works in conjunction with cooking appliances to automatically adjust the temperature and time according to pre-set recipes, streamlining the cooking process. For example, when a user is making pasta, the smart oven automatically starts heating to the appropriate temperature, and the timer function is integrated. This is especially helpful when trying out new recipes.

[0742] This system also incorporates an advertising display function, allowing it to provide users with product information and promotions related to the suggested recipes. By analyzing users' cooking history and preferences, it improves the accuracy of advertisements and enables more personalized recommendations. Furthermore, future plans include integration with health management programs, aiming to provide suggestions that contribute to users' health maintenance.

[0743] For example, if a user takes photos of tomatoes, carrots, and onions at home and uploads them, the server analyzes them and suggests dishes like minestrone or vegetable soup. If the user selects minestrone, detailed recipe instructions, along with voice guidance, are provided on the device, making cooking smooth. Furthermore, if compatible cooking equipment is available, the heating time is automatically set, minimizing user intervention.

[0744] In this way, the present invention provides users with a new cooking experience and achieves increased efficiency and convenience in cooking.

[0745] The following describes the processing flow.

[0746] Step 1:

[0747] The user uses their device's camera to photograph ingredients and seasonings they have on hand. They then review the captured images within the app and prepare them for uploading to the server.

[0748] Step 2:

[0749] The device sends the uploaded image data to the server. During this process, the image data undergoes format conversion and compression as needed.

[0750] Step 3:

[0751] The server inputs the received image data into an image analysis module to identify the types and characteristics of ingredients and seasonings. This process utilizes machine learning algorithms.

[0752] Step 4:

[0753] Based on the analysis results, the server consults a food database and identifies multiple recipes that can be generated from the available ingredients. In doing so, it also takes into account the user's past selection history, preferences, and allergy information.

[0754] Step 5:

[0755] The server sends a list of suggested dishes to the terminal. The terminal displays this information on its user interface, making it available for the user to select.

[0756] Step 6:

[0757] The user selects the dish they want to make from the dish options displayed on their device. Once the selection is complete, that information is sent back to the server.

[0758] Step 7:

[0759] The server retrieves audio guides and video content, including cooking instructions for the selected dish, and sends them to the terminal.

[0760] Step 8:

[0761] The device plays back the received audio and video, guiding the user through the cooking process using both audio and video. The guidance is timed to each step.

[0762] Step 9:

[0763] The server works in conjunction with smart cooking appliances, sending commands to automate operations necessary for the cooking process—such as when to start and stop heating the oven.

[0764] Step 10:

[0765] After cooking is complete, the server updates the user's cooking history and saves it as data to help improve future services and the accuracy of ad displays.

[0766] (Example 1)

[0767] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0768] In cooking, it is essential to effectively utilize the ingredients available to the user and to easily provide a variety of cooking options. Furthermore, it is desirable to streamline the cooking process so that even beginners or users with limited time can cook intuitively. Additionally, providing food suggestions tailored to individual preferences and health conditions is a crucial challenge.

[0769] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0770] In this invention, the server includes means for analyzing visual information transmitted by the user to identify materials contained in the visual information, means for suggesting culinary foods based on the materials, and means for providing auditory and visual guidance on the cooking procedure of the foods. This allows the user to enjoy a variety of dishes without wasting any of the ingredients they have on hand, maximizing the efficiency and effectiveness of cooking. Furthermore, this also includes providing personalized food suggestions based on the user's preferences and health condition, supporting optimal cooking and health management tailored to individual needs.

[0771] "Visual information" refers to image data and other visual data captured by the user.

[0772] "Ingredients" refers to substances, including food and seasonings, used in cooking.

[0773] "Cookable food" refers to food that is generated based on analyzed ingredients and suggested to users so that they can actually cook and consume it.

[0774] "Auditory and visual guidance" refers to providing users with information and instructions through audio guides and video content.

[0775] "Cooking equipment" refers to devices or equipment used for cooking, including those with automatic temperature and time control functions.

[0776] "Automated cooking" refers to a process in which cooking equipment performs cooking according to a set program, without requiring direct human intervention.

[0777] "Comparing analyzed visual information with stored data" refers to the process of comparing information obtained through image analysis with existing databases to confirm the consistency and relevance of the information.

[0778] "Personalized product and promotional information" refers to advertisements and sales information provided based on the user's preferences and usage history.

[0779] "Health management information" refers to data such as the user's health status and nutritional balance, which is taken into consideration when making cooking suggestions.

[0780] "Automatic control of specified temperature and time" refers to a function in which the cooking device autonomously proceeds with cooking based on the set temperature and time.

[0781] The system of the present invention streamlines cooking based on the ingredients the user possesses, providing a richer dining experience. The embodiments are described in detail below.

[0782] The user uses their device's camera to photograph materials they have on hand. For example, if using materials such as tomatoes, carrots, and onions, the user photographs them and sends the image data to the server through a pre-configured application. This application runs on smartphones and tablets.

[0783] The server uses deep learning technology to analyze the received image data. Frameworks such as TensorFlow and PyTorch are employed for this process, and calculations are performed to identify the materials within the image. Subsequently, the identified materials are compared with a food database, and several cookable food items are suggested, taking into account the user's preferences and past transaction history.

[0784] The terminal notifies the user of suggestions received from the server and visually displays the recipe. When the user selects a menu item of interest, the server generates detailed cooking instructions for that menu item and provides them as audio guides and video content. This content is delivered in streaming format using libraries such as FFmpeg.

[0785] Furthermore, the cooking device and server are linked via communication, allowing for automatic control of the temperature and cooking time required for each dish. For example, when cooking minestrone, the smart oven is set to the appropriate temperature and heating continues for the specified time. This significantly reduces the effort users spend on cooking.

[0786] This system also has the function of providing personalized advertisements and promotional information to individual users. By analyzing the user's cooking history and preference information, it suggests highly relevant products and services.

[0787] As a concrete example, a possible scenario is where, based on images of tomatoes, carrots, and onions taken by the user, the server suggests minestrone or vegetable soup, and if the user selects minestrone, the terminal provides a detailed recipe.

[0788] An example of a prompt message would be, "Upload images of tomatoes, carrots, and onions taken with your device, and then display the recipe suggestions after the server has analyzed them."

[0789] In this way, the present invention can support the user's cooking activities and provide a new cooking experience.

[0790] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0791] Step 1:

[0792] The user takes a picture of an object in their possession using the device's camera. Inputs include image files of items such as tomatoes, carrots, and onions. The device then sends this image data to the server via the application. The output is the image data sent to the server.

[0793] Step 2:

[0794] The server analyzes the received image data. The input is image data sent by the user. The server uses a generative AI model to analyze the image and perform data processing to identify the material. Deep learning technology is applied in this process. The output is the identified material information.

[0795] Step 3:

[0796] The server compares the identified material information with the stored database. The input is the material information as a result of the analysis, and the server performs an operation to compare it with the existing food database. The output is a list of related foods.

[0797] Step 4:

[0798] The server suggests foods that can be cooked based on the user's preferences and past history. The input consists of a list of ingredients and user history data. The server analyzes the information and performs data calculations using a recommendation algorithm. The output is a list of suggested recipes.

[0799] Step 5:

[0800] The terminal displays the suggested recipes in the user interface. The input is recipe information provided by the server. The terminal performs specific actions to visualize the recipes and notify the user. The output is the recipe options displayed to the user.

[0801] Step 6:

[0802] The user selects a menu item of interest from the presented recipes. The input is the recipe options displayed on the screen. Based on the user's selection, the output will be the information for the specific recipe selected.

[0803] Step 7:

[0804] The server generates detailed cooking instructions based on the selected recipe. The input is the recipe information selected by the user. The server processes the data to create the cooking instructions as audio guides and video content. The output is the generated cooking instruction content data.

[0805] Step 8:

[0806] The terminal plays audio guides and video content received from the server. The input is cooking procedure data sent from the server. The terminal activates its playback function and provides guidance to the user. The output is visual and auditory information of the cooking instructions provided to the user.

[0807] Step 9:

[0808] The server communicates with the cooking device to automatically adjust the temperature and time. The input consists of the cooking parameters included in the selected recipe. The server then sends configuration information to the cooking device and performs specific actions to enable automatic control. The output is the adjusted heating settings and time management.

[0809] (Application Example 1)

[0810] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0811] Modern consumers struggle to find the optimal combination of ingredients for cooking from the various ingredients available in stores, and to devise dishes that efficiently utilize those ingredients. Furthermore, they lack opportunities to obtain useful product information related to the ingredients, and opportunities to receive easy-to-understand guidance on cooking procedures. This results in increased time and effort spent on shopping and cooking, reducing convenience. Therefore, there is a need for a means to simultaneously provide an environment where customers can easily receive cooking suggestions in stores, and to provide advertising information related to the ingredients.

[0812] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0813] In this invention, the server includes means for analyzing images transmitted by the user to identify materials contained in the images, means for suggesting cookable foods based on the materials, and means for displaying food suggestions based on the materials and advertisements for cooking equipment via electronic devices installed in the store. This makes it possible to provide in-store material selection and optimal cooking suggestions, as well as display advertisements for related products on the spot.

[0814] An "image" is a digital recording of visual information and serves as a source of information for visually identifying materials.

[0815] "Materials" refer to the raw ingredients used in food preparation, and are the subjects of analysis from the captured images.

[0816] "Cookable foods" refer to dishes and food / drinks that can be made using specific ingredients, and are the subject of the proposal.

[0817] A "suggestion" is the act of presenting users with a selection of foods that can be cooked, and is a process that encourages users to make choices.

[0818] "Audio and video" refers to functions that convey cooking procedures to the user, and is a means of providing information through both sight and sound.

[0819] An "electronic device" is a device that has the ability to input, process, and display information, and functions as an interface between the user and the system.

[0820] "Advertising" is a means of providing information aimed at informing users about related products and services and increasing their desire to purchase them.

[0821] "Automatic cooking" refers to a function in which cooking appliances independently execute cooking processes according to a pre-set program.

[0822] This system is designed to enhance the user's in-store shopping experience. It consists of servers, terminals, and related electronic devices, allowing users to receive on-the-spot guidance on everything from ingredient selection to cooking suggestions.

[0823] The server uses a software environment implementing deep learning technology (specifically TensorFlow and PyTorch) to analyze images captured by the user's device camera. This allows the server to identify the captured material and compare it with a food database to suggest foods that can be cooked. The suggested foods are then sent to the user's device as audio and video content, providing clear and easy-to-understand guidance. This guidance utilizes speech synthesis and video streaming technologies.

[0824] Furthermore, sales promotion will be achieved by displaying advertisements for products related to the suggested food items to users via electronic devices. AdTech platforms, such as Google AdMob, will be used for this ad delivery.

[0825] For example, if a user takes a picture of chicken breast and broccoli in a store, the system will analyze the image and suggest a creamy chicken and broccoli dish. In this case, a prompt message such as, "Recognize the ingredients in the following image, generate a recipe for chicken breast and broccoli based on that, and also display advertisements for the necessary cooking utensils," will be sent to the server, enabling simultaneous display of cooking suggestions and advertisements related to the ingredients.

[0826] In this way, the in-store shopping and cooking experience becomes intuitive and efficient, resulting in a system that can improve user satisfaction.

[0827] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0828] Step 1:

[0829] The user uses their device to photograph items they are considering purchasing in the store. During this process, the device's camera acquires high-resolution image data, which is then sent to a server as image data via the application.

[0830] Step 2:

[0831] The server analyzes the received images using a deep learning model. The input images are preprocessed using frameworks such as TensorFlow and PyTorch to identify the source material. As a result of the analysis, source information is generated and compared with a database.

[0832] Step 3:

[0833] Based on the matching results, the server searches the food database for relevant cookable foods and generates suggestions. The input is the ingredient information obtained in step 2, and the output is a list of relevant recipes. This is sent to the terminal in a format that the user can select.

[0834] Step 4:

[0835] The user selects their preferred recipe from several displayed on the device. The selected recipe name is resent to the server and used for the next process.

[0836] Step 5:

[0837] The server generates audio and video guides for the cooking process based on the selected recipe. The input is the selected recipe information, and the output is audio and video files. These are downloaded to the terminal and can be viewed by the user. Speech synthesis and video streaming technologies are used.

[0838] Step 6:

[0839] The device plays back received audio and video data, providing users with easy-to-understand cooking information. It is designed to be intuitively usable through its interface.

[0840] Step 7:

[0841] The server provides a mechanism to link suggested recipes with related product advertisements and display the most relevant advertisements to the user. The input is promotional information related to cookable foods, and the output is the content of the displayed advertisements.

[0842] This series of steps creates a system that allows users to smoothly proceed through the entire process, from selecting ingredients to cooking and purchasing necessary products.

[0843] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0844] This invention combines a system that analyzes images provided by the user to suggest cooking methods with an emotion engine that recognizes the user's emotions. This system enables the provision of more appropriate food suggestions and cooking guides based on the user's emotional state.

[0845] First, the user takes a picture of the ingredients or seasonings they want to use using the device's camera. The captured image is sent to the server via the app. The device is equipped with emotion recognition capabilities, which analyze the user's conversation, facial expressions, and voice to identify their emotions. This emotion information is also sent to the server along with the image data.

[0846] When the server receives an image, it uses an image analysis module to identify the material. In addition, the emotion engine evaluates the received emotional information to determine the user's current emotional state. This information is referenced when making food recommendations. For example, if the emotional state is positive, a new cooking method will be suggested; if it is negative, a recipe prioritizing convenience will be presented.

[0847] Multiple suggested cooking options are sent to the device, and the user makes a selection on the device. Depending on the selected menu, the server sends audio and video guidance to the device, providing the user with visual and auditory instructions. The device plays this back, providing the user with both visual and auditory guidance. The emotion engine can also continuously obtain user feedback during cooking and provide advice as needed.

[0848] Furthermore, the server works in conjunction with smart cooking devices to automatically adjust cooking parameters according to the progress of each cooking step. In this way, a cooking experience tailored to the user's emotional state is provided, supporting a more satisfying cooking experience. For example, if the user's emotions suggest stress, a simple and relaxing cooking method will be suggested.

[0849] This system also records the user's emotions and cooking data, which will be used as training data to improve the accuracy of food recommendations and emotion recognition in the future. As a result, it will be possible to provide a personalized experience for each user, reducing the burden of cooking and increasing the enjoyment.

[0850] In this way, the present invention realizes cooking support that takes emotional states into account, and provides new added value to the user.

[0851] The following describes the processing flow.

[0852] Step 1:

[0853] The user uses their device's camera to photograph the currently available ingredients and seasonings. The captured images are then prepared to be sent to the server via the app's interface.

[0854] Step 2:

[0855] The device analyzes the user's facial expressions and voice using an emotion recognition sensor. The obtained emotion data is transmitted to the server in parallel with the image data.

[0856] Step 3:

[0857] The server inputs the received image data into an image analysis module to identify ingredients and seasonings. This information is used as basic data to suggest foods that can be cooked.

[0858] Step 4:

[0859] The server uses an emotion engine to analyze the received emotion data and identify the user's current emotional state. Based on this emotional state, it adjusts food recommendations.

[0860] Step 5:

[0861] The server matches the user's emotional state with ingredient information to generate suitable dish suggestions. When making suggestions, it takes into account the difficulty level and cooking time based on the user's emotional state.

[0862] Step 6:

[0863] The terminal displays a list of recipe suggestions sent from the server on the user interface. The user selects the dish they want to make from this list.

[0864] Step 7:

[0865] Based on the user's selection, the server retrieves the cooking instructions for the selected dish and prepares the associated audio guide and visual instructional video.

[0866] Step 8:

[0867] The terminal plays audio guides and videos provided by the server, guiding the user through the cooking steps. These guides include emotionally responsive tips.

[0868] Step 9:

[0869] The server receives emotional feedback from the user and controls the smart cooking equipment according to the progress of cooking, optimizing the automated cooking process.

[0870] Step 10:

[0871] Once cooking is complete, the server records the user's emotional data and cooking data, which are then used as training data to improve the accuracy of future suggestions.

[0872] (Example 2)

[0873] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0874] There is a need for a system that can improve user satisfaction by providing optimal food suggestions and cooking experiences tailored to the user's emotional state when preparing food. Existing technologies lack the flexibility to consider changes in the user's emotions and individual needs when providing food suggestions and cooking guides, making it difficult to address the individual emotions of users.

[0875] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0876] In this invention, the server includes means for analyzing visual data transmitted by the user to identify ingredients, means for recognizing the user's emotional state and suggesting food, means for guiding the user through the food preparation process using voice and video, means for performing automatic cooking in cooperation with a cooking device, and means for recording the user's emotional information and cooking data to improve the accuracy of future food suggestions. This makes it possible to suggest food that is appropriate to the user's emotional state, thereby providing a highly satisfying cooking experience.

[0877] "Visual data" refers to visual information such as images and videos, which is information provided by the user through the camera of their device.

[0878] "Ingredients" refers to raw materials used in food preparation, and are food components identified through specific image analysis.

[0879] "Emotional state" refers to the user's current psychological state and feelings, and is information revealed by emotion recognition technology.

[0880] "Food suggestions" refers to the act of providing appropriate cooking ideas and menus based on the user's emotional state and the ingredients available.

[0881] "Audio" and "video" refer to means of transmitting information in auditory and visual forms, respectively, and are used as means of guiding the cooking procedure.

[0882] A "cooking device" refers to a machine or technology that automatically prepares food, and is controlled in conjunction with a server.

[0883] "Recommendation accuracy" refers to the degree to which the food and cooking method suggestions made to the user accurately match the user's needs and circumstances.

[0884] This invention combines a system that analyzes visual data provided by the user to suggest cooking methods with a function that recognizes the user's emotional state. The user takes pictures of the ingredients needed for cooking using the camera on their device. The device uses an emotion recognition engine to identify the user's emotional state by analyzing voice, facial expressions, and conversation data. This information is transmitted to the server along with the visual data. The server is equipped with an advanced image analysis module and an emotion recognition engine, which use these tools to identify the ingredients and evaluate the user's emotional state.

[0885] Specifically, the user will use an electronic terminal equipped with a camera. This terminal will be equipped with an application that utilizes an AI model to perform real-time image processing and emotion recognition. On the server side, a powerful computer will run an image analysis module and an emotion engine, and perform data processing using AI algorithms. The software will utilize AI technologies, including image recognition and natural language processing, to enable food recommendations tailored to the user's needs.

[0886] A concrete example demonstrating the advantages of this system is when a user takes a picture of tomatoes and basil and the system determines that the user is feeling relaxed; in such a case, a simple tomato and basil pasta recipe is suggested. The user selects the menu on their device, and the server guides them through the cooking process using voice and video. During this process, advice and adjustments are made based on the progress of the cooking and the user's emotional feedback.

[0887] Furthermore, user emotional information and cooking data are used to improve the accuracy of future suggestions. In this way, a personalized experience is provided, resulting in a less stressful cooking method for the user.

[0888] An example of a prompt message is, "Suggest a dish based on the image and provide a cooking guide that takes the user's emotions into consideration." This is how you can instruct the system. Based on this prompt, the generative AI model will provide the most suitable cooking suggestion for the user.

[0889] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0890] Step 1:

[0891] The user uses the device's camera to capture visual data of the ingredients to be used in cooking. At this stage, the input is physical material (e.g., tomatoes or basil), which is captured as a digital image on the device. The device then saves the captured image within the application.

[0892] Step 2:

[0893] The device uses a built-in emotion recognition engine to analyze the user's voice and facial expressions in real time. In this process, voice data and facial expression data are used as input, and the emotional state (for example, an emotion indicating stress) is output. Emotion recognition technology quantifies and categorizes emotions for recording.

[0894] Step 3:

[0895] The device transmits captured visual data and emotional state data to the server. The input for this communication consists of image data and emotional data, which are passed to the server via the network. The transmitted data serves as the basis for analysis processing on the server.

[0896] Step 4:

[0897] The server inputs the received visual data into an image analysis module to identify the materials. For example, it uses image recognition technology to determine if the material is a tomato or basil. At this stage, the input is image data, and the output is information about the identified materials. The analysis results are used as a basis for deciding which dishes to suggest to the user.

[0898] Step 5:

[0899] The server analyzes the emotional information received by the emotion recognition engine and evaluates the user's emotional state. The input is emotional data, and the output is the specific emotional category the user is currently experiencing (e.g., "I want to relax"). This information is used to select what kind of meal to suggest.

[0900] Step 6:

[0901] The server generates optimal food suggestions based on information about the ingredients and the evaluation results of the emotional state. For example, if the emotion of wanting to relax is detected, a simple tomato and basil pasta dish will be suggested. The input is the analysis results of ingredient data and emotional data, and the output is a cooking suggestion.

[0902] Step 7:

[0903] The server prepares the suggested food preparation steps as audio and video guides and sends them to the terminal. In this step, the specific preparation steps based on the suggested dish are input and output as visual and auditory instructions to the user. The terminal receives this and provides the guide to the user.

[0904] Step 8:

[0905] The device continuously acquires user feedback and additional emotional data during cooking. This enables real-time advice and adjustments. Inputs are user feedback and new emotional data, while outputs are further advice and procedural modifications as needed.

[0906] (Application Example 2)

[0907] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0908] Modern consumers are increasingly seeking personalized suggestions for meals and food delivery that reflect their individual emotional states and preferences. Traditional systems fail to consider the user's emotional state when suggesting food, resulting in a lack of appropriate choices. Furthermore, cooking instructions and automated cooking processes are not optimized based on the user's emotional state, limiting the improvement in satisfaction.

[0909] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0910] In this invention, the server includes means for analyzing images transmitted by the user to identify objects contained within the images, means for analyzing emotions and optimizing food suggestions based on the emotional state, and means for providing audio and video guidance on the cooking procedure for the items. This enables personalized food suggestions and appropriate cooking guidance tailored to the user's emotional state.

[0911] A "user" is an individual or group that uses this system to send images and receive food suggestions and cooking assistance.

[0912] An "image" is data that records visual information containing an object, transmitted by the user.

[0913] "Analysis" refers to a processing method that extracts specific information from transmitted images or data and performs recognition or classification.

[0914] "Object" refers to cooking ingredients or related items that are included in the image taken by the user.

[0915] "Emotion" is a concept that encompasses elements used to analyze and identify a user's emotional state.

[0916] "Emotional state" refers to analyzed information that reflects the user's inner mood and psychological state.

[0917] "Food suggestion" is the process of presenting users with selectable, culinary items based on the identification of objects and emotional states.

[0918] "Providing guidance through audio and video" means providing cooking instructions to the user using audio guides and visual presentations.

[0919] "Automated cooking" is a process in which food is cooked automatically in conjunction with cooking equipment, based on pre-set procedures.

[0920] This food delivery system uses programs that perform emotion analysis and image analysis to suggest food items based on the user's emotional state. The system primarily operates on devices such as smartphones and smart glasses, as well as on servers running on cloud services.

[0921] The server receives image and audio data captured by the user using the device's camera. This image data is used to identify objects through an analysis module. An example of a module used is an image recognition algorithm that utilizes TensorFlow.

[0922] Furthermore, the server uses an emotion analysis engine to understand the user's emotional state from voice and facial expression data. Emotion recognition APIs such as Microsoft Azure Cognitive Services can be used here. The analysis results reflect the user's current psychological state, and the suggested food items are optimized based on this data.

[0923] The terminal functions as a device for receiving food suggestions and providing cooking instructions via audio and video. This allows users to proceed with cooking using visual and auditory guidance. Furthermore, communication with the server enables the system to provide additional feedback when the user's emotional state changes and update suggestions as needed.

[0924] For example, if a user is feeling like "I want a relaxing meal today," the system might suggest a salad suitable as a light snack. An example of inputting such prompts into an AI model might be: "Based on the user's voice recording and photo data, identify the user's current mood. Then, generate a list of food recommendations based on that mood."

[0925] In this way, the system can understand the user's emotional state, make food suggestions accordingly, and provide a personalized dining experience for each individual user.

[0926] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0927] Step 1:

[0928] The user takes pictures of ingredients and seasonings using the device's camera.

[0929] Image data is generated as input, and the terminal sends this image data to the server.

[0930] The transmitted image data will serve as input for the next analysis.

[0931] Step 2:

[0932] The server receives the image data and uses an image analysis module to identify objects within the image.

[0933] The input is image data, and object identification is performed using image recognition algorithms such as TensorFlow.

[0934] The output is a list of identified objects.

[0935] Step 3:

[0936] The server receives the user's voice data or facial expression data acquired by the camera, and analyzes their emotional state through an emotion analysis engine.

[0937] The input is either audio data or facial expression images, which are analyzed using emotion recognition APIs such as Microsoft Azure Cognitive Services.

[0938] The output is data that indicates the user's emotional state.

[0939] Step 4:

[0940] The server uses a food suggestion algorithm to propose the best food based on the identified object and emotional state.

[0941] The input consists of a list of objects and emotional state data, and an AI model generates behavioral recommendations based on this data.

[0942] The output is a list of food product suggestions.

[0943] Step 5:

[0944] The terminal receives a list of suggestions from the server and presents the suggestions to the user visually and audibly.

[0945] The input is a list of food suggestions, and the output is a selection of options displayed to the user.

[0946] In terms of specific actions, a menu is displayed on the screen, and suggestions are explained with voice guidance.

[0947] Step 6:

[0948] When the user makes a selection from the suggested food items, the terminal sends the selection information to the server.

[0949] The input is the user's selection information, and the selected food items are sent to the server as output.

[0950] Step 7:

[0951] The server generates cooking instructions based on the selected food item and sends specific audio and video guides to the terminal.

[0952] The input is the selected food item, and the output is audio and video data of the cooking procedure.

[0953] This process helps the user easily perform the next cooking step.

[0954] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0955] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0956] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0957] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0958] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0959] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0960] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0961] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0962] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0963] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0964] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0965] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0966] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0967] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0968] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0969] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0970] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0971] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0972] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0973] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0974] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0975] The following is further disclosed regarding the embodiments described above.

[0976] (Claim 1)

[0977] A means for analyzing images sent by users to identify materials contained within those images,

[0978] A means for proposing a cookable food based on the aforementioned materials,

[0979] A means for providing audio and video guidance on the preparation procedure of the aforementioned food,

[0980] A system that includes means for performing automated cooking in conjunction with cooking equipment.

[0981] (Claim 2)

[0982] The system according to claim 1, which provides instructions on the cooking procedure of the food via an electronic device.

[0983] (Claim 3)

[0984] The system according to claim 1, which uses information to provide personalized food recommendations.

[0985] "Example 1"

[0986] (Claim 1)

[0987] A means for analyzing visual information transmitted by the user and identifying the materials contained within that visual information,

[0988] A means for proposing a cookable food based on the aforementioned materials,

[0989] Means for guiding the cooking procedure of the aforementioned food product through auditory and visual means,

[0990] A means of performing automated cooking in conjunction with cooking equipment,

[0991] A means for comparing the analyzed visual information with stored data,

[0992] Means of providing personalized product information and promotional information,

[0993] A means of providing meal suggestions that take health management information into consideration,

[0994] Means for automatically controlling the specified temperature and time,

[0995] A system that includes this.

[0996] (Claim 2)

[0997] The system according to claim 1, which provides instructions for the cooking procedure of the food through an electronic device.

[0998] (Claim 3)

[0999] The system according to claim 1, which uses information to provide personalized food recommendations.

[1000] "Application Example 1"

[1001] (Claim 1)

[1002] A means for analyzing images sent by users to identify materials contained within those images,

[1003] A means for proposing a cookable food based on the aforementioned materials,

[1004] A means for providing audio and video guidance on the preparation procedure of the aforementioned food,

[1005] A means of presenting food suggestions based on ingredients and advertisements for cooking equipment via electronic devices installed in the store,

[1006] A means of displaying advertisements for products related to the proposed food,

[1007] A system that includes means for performing automated cooking in conjunction with cooking equipment.

[1008] (Claim 2)

[1009] The system according to claim 1, which provides instructions for cooking food and advertisements for related products via electronic devices.

[1010] (Claim 3)

[1011] The system according to claim 1, which uses information to provide personalized food suggestions and advertisements for related products.

[1012] "Example 2 of combining an emotion engine"

[1013] (Claim 1)

[1014] A means for analyzing visual data transmitted by the user to identify the materials contained within the visual data,

[1015] A means of recognizing the user's emotional state and making food recommendations based on that information,

[1016] A means for providing audio and video guidance on the preparation procedure of the aforementioned food,

[1017] A means of performing automatic cooking in conjunction with a cooking device,

[1018] A system that includes means to record user emotional information and cooking data to improve the accuracy of future food recommendations.

[1019] (Claim 2)

[1020] The system according to claim 1, which provides instructions on the cooking procedure of the food via an electronic terminal.

[1021] (Claim 3)

[1022] The system according to claim 1, which uses acquired information to provide personalized food recommendations.

[1023] "Application example 2 when combining with an emotional engine"

[1024] (Claim 1)

[1025] A means for analyzing images sent by the user to identify objects contained within the images,

[1026] Means for proposing a cookable article based on the aforementioned object,

[1027] A means for analyzing emotions and optimizing food recommendations based on emotional states,

[1028] A means for providing audio and video guidance on the cooking procedure of the aforementioned item,

[1029] A system that includes means for performing automated cooking in conjunction with cooking equipment.

[1030] (Claim 2)

[1031] The system according to claim 1, which provides instructions for the cooking procedure of the article through an electronic device.

[1032] (Claim 3)

[1033] The system according to claim 1, which uses emotional information to provide personalized product suggestions. [Explanation of Symbols]

[1034] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for analyzing images sent by users to identify materials contained within those images, A means for proposing a cookable food based on the aforementioned materials, A means for providing audio and video guidance on the preparation procedure of the aforementioned food, A system that includes means for performing automated cooking in conjunction with cooking equipment.

2. The system according to claim 1, which provides instructions on the cooking procedure of the food via an electronic device.

3. The system according to claim 1, which uses information to provide personalized food recommendations.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A