System

A digital system digitizes and personalizes mother's cooking recipes, providing voice-guided instructions in her voice, addressing the challenge of recreating flavors and accommodating dietary restrictions, thereby strengthening family bonds.

JP2026023995APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126316
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

It is difficult to accurately recreate a mother's cooking methods and flavors, especially when there are dietary restrictions or taste preferences, which can lead to a loss of special family memories and bonds.

Method used

A digital system that digitizes mother's cooking recipes, analyzes them for ingredients and steps, learns her voice, and provides audio guidance tailored to individual preferences, allowing for customized cooking instructions in her voice, with feedback loops to improve recipes over time.

Benefits of technology

Enables accurate passing of mother's cooking techniques to the next generation, accommodating individual needs and deepening family bonds through personalized and continuously improved culinary experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023995000001_ABST
    Figure 2026023995000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for digitizing a mother's cooking recipe; means for analyzing the digitized recipe and extracting ingredients, quantities, and cooking procedures; means for storing the extracted data in a database; means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning techniques; and means for audibly guiding the user through the cooking procedures in the mother's voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] While a mother's cooking holds special meaning for a family, it is difficult to accurately recreate the cooking method and flavor. Some families want to pass on their mother's cooking to the next generation, but some find it difficult to master the techniques or have dietary restrictions. This invention solves these challenges, passing on the family's special flavors to the next generation and deepening family bonds. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for digitizing mother's cooking recipes, a means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps, a means for storing the extracted data in a database, a means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning technology, and a means for providing audio guidance of cooking steps to a user in the mother's voice. Furthermore, by including a means for customizing recipes based on the user's dietary restrictions and taste preferences, and a means for collecting feedback data from users and updating the recipe database, it is possible to recreate the taste of mother's cooking in a way that satisfies the whole family.

[0006] "Digitization" refers to the general process of converting physical documents, audio, video, etc. into electronic form.

[0007] "Means for analysis" refers to the technology used to analyze collected digital data and extract necessary information.

[0008] "Ingredients" refers to the food ingredients and seasonings used to make a meal.

[0009] "Quantity" refers to the specific amount of each ingredient or seasoning needed in a dish.

[0010] A "cooking procedure" refers to a series of specific operations or steps required to complete a dish.

[0011] "Database storage means" refers to the techniques and devices used to organize the extracted information in a certain format and store and manage it in a digital database.

[0012] "Audio data" refers to a digital file of audio recordings of the mother's cooking procedures and other information.

[0013] "Speech recognition" refers to the process of extracting text data from speech.

[0014] "Voice cloning technology" refers to technology that learns the characteristics of a specific person's voice and generates a new voice using that person's voice.

[0015] "Voice guidance means" refers to technology and devices for providing instructions and information to users through voice.

[0016] "Dietary restrictions" refer to ingredients and seasonings that should be avoided due to specific health conditions or personal preferences.

[0017] "Means of customization" refers to the technology of adjusting and modifying recipes and procedures according to the user's specific requests and conditions.

[0018] "Feedback data" refers to information such as opinions, evaluations, and areas for improvement provided by users.

[0019] "Means for updating the recipe database" refers to the technology for modifying or adding stored recipe information based on collected feedback. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] This invention is a digital system for accurately passing on mother's cooking recipes to the next generation, aiming to deepen special family memories and bonds while recreating the flavor of mother's cooking. The program of this system includes functions for data collection, analysis, audio guidance, customization, and feedback.

[0042] 1. Data Collection

[0043] User: First, the user takes a note of his mother's cooking recipes and a video of his mother cooking on his smartphone.

[0044] Device: Upload the photos and videos you have taken to the server via the application.

[0045] Server: Receives the uploaded data and converts the recipe into text using OCR technology.

[0046] 2. Recipe analysis and database storage

[0047] Server: Recipe data converted to text using OCR is fed into the AI ​​model, and information such as ingredients, quantities, and cooking steps is extracted.

[0048] Server: Structures the extracted data and stores it in a database.

[0049] 3. Collecting and Learning Audio Data

[0050] User: Records mother explaining cooking steps on smartphone.

[0051] Device: Upload the recording data from the application to the server.

[0052] Server: The uploaded voice data is analyzed through a voice recognition process and converted into text. At the same time, voice cloning technology is used to learn the mother's voice and generate a voice model.

[0053] 4. Providing cooking guides

[0054] User: Selects the recipe they want to cook within the application.

[0055] Terminal: Sends a request for the selected recipe to the server.

[0056] Server: Retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[0057] Device: The acquired cooking instructions are displayed to the user, and instructions such as "First, cut the chicken into bite-sized pieces" are also provided in a mother's voice.

[0058] 5. User Customization

[0059] User: Enter allergies and individual taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0060] Terminal: Sends the entered customization information to the server.

[0061] Server: Adjusts the recipe based on the customization information and generates the appropriate instructions.

[0062] Device: Displays the customized recipe and provides audio guidance if needed.

[0063] Specific examples

[0064] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is given audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also making dishes that accommodate their own dietary restrictions.

[0065] In addition, after cooking, users can enter feedback into the app, which is analyzed by the server and the recipe database is constantly updated to provide more accurate guidance the next time the user cooks.

[0066] By combining the above functions, the system of the present invention can accurately pass on the taste of mother's cooking to the next generation, further deepening family ties.

[0067] The processing flow will be explained below.

[0068] Step 1: Data collection

[0069] User: Takes notes of his mother's cooking recipes and videos of himself cooking on his smartphone.

[0070] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[0071] Server: Receives the uploaded data and temporarily stores it in a database.

[0072] Step 2: Digitize your recipes

[0073] Server: Using OCR technology, extracts text information from photos and videos and converts recipes into text format.

[0074] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[0075] Step 3: Collecting audio data

[0076] User: Records her mother explaining the recipe steps on her smartphone.

[0077] Terminal: Upload the recorded audio data to the server using a dedicated application.

[0078] Server: Receives the uploaded audio data and stores it in a database.

[0079] Step 4: Analyze and train audio data

[0080] Server: Extracts text data from the voice data using voice recognition technology.

[0081] Server: Organizes the text data into cooking instructions and stores them in a database.

[0082] Server: Using voice cloning technology, learns the characteristics of the mother's voice and generates a voice model.

[0083] Step 5: Parse and save the recipe

[0084] Server: The text data obtained by OCR is fed into the AI ​​model, which analyzes ingredients, quantities, and cooking procedures.

[0085] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0086] Step 6: User selects recipe

[0087] User: Select the recipe they want to cook within the dedicated application.

[0088] Terminal: Sends a request for the selected recipe to the server.

[0089] Step 7: Cooking instructions generation and audio guidance

[0090] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0091] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0092] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0093] Step 8: Enter and update customization information

[0094] User: Enters specific dietary restrictions and taste preferences (e.g., "low salt," "medium spicy," etc.) into a dedicated application.

[0095] Terminal: Sends the entered customization information to the server.

[0096] Server: Adjusts the recipe based on the user's customization information.

[0097] Device: Provides the adjusted recipe to the user, and also provides audio guidance if necessary.

[0098] Step 9: Gather and incorporate feedback

[0099] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[0100] Terminal: Sends the input feedback data to the server.

[0101] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0102] Example 1

[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0104] It is not easy to accurately pass on the flavors and steps of a mother's cooking to the next generation. If recipes are not accurately digitized, there is a high risk of losing special family memories and bonds. It is also difficult to satisfy the need to receive detailed cooking instructions in the mother's voice. Furthermore, it is difficult to accommodate the different dietary restrictions and taste preferences of each family. To solve these problems, accurately pass on the flavors of a mother's cooking to the next generation, and deepen family bonds, an efficient and accurate digital system is needed.

[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0106] In this invention, the server includes means for digitizing mother's recipes, means for analyzing the digitized recipes and extracting ingredients, portions, and cooking methods, means for storing the extracted data in a data repository, means for collecting mother's voice data and learning mother's voice using voice recognition and voice synthesis technology, and means for providing audio guidance on cooking methods to users in mother's voice, thereby enabling mother's recipes to be accurately passed down to future generations and providing cooking guides customized to the specific needs of each household.

[0107] "My mother's cooking methods" refers to the cooking methods, steps, and recipes that my mother uses.

[0108] "Digitalization" refers to the process of converting analog information into electronic data.

[0109] "Ingredients" refers to the various foods and ingredients used in cooking.

[0110] "Amount" refers to the amount of each ingredient used when cooking.

[0111] "Cooking method" refers to a series of processes or steps to complete a dish using ingredients.

[0112] A "data repository" refers to a database or storage system for efficiently storing and managing collected data.

[0113] "Audio data" refers to digital data that contains recorded audio information such as human voices and spoken words.

[0114] "Speech recognition technology" refers to technology that generates text data from voice data.

[0115] "Speech synthesis technology" refers to technology that generates new voices based on certain voice samples.

[0116] "Voice guidance" refers to a method of conveying specific procedures or information to a user by providing voice guidance.

[0117] This invention is a digital system that aims to pass on mother's cooking techniques to the next generation and deepen family memories and bonds. The system has functions for data collection, analysis, audio guidance, customization, and feedback, and is designed to make it easier for users to recreate the taste of their mother's cooking.

[0118] Data collection

[0119] The first thing a user does is digitize their mother's cooking methods. They take photos of their mother's recipe notebooks and videos of their mother cooking with their smartphone. They then upload these photos and videos to a server using a dedicated application.

[0120] The server receives the uploaded photos and videos and converts the recipes into text data using OCR (Optical Character Recognition) technology, a process that converts analog information into electronic data.

[0121] Recipe analysis and database storage

[0122] The server then inputs the recipe data, converted to text using OCR technology, into an AI model to extract information such as ingredients, quantities, cooking methods, etc. Specifically, it uses natural language processing technology to analyze the recipe text and automatically identify the necessary information.

[0123] The extracted data is structured by the server and stored in a data repository, making it easy for users to search and browse later.

[0124] Audio data collection and learning

[0125] The user records their mother explaining the cooking steps on their smartphone, and the recording is then uploaded to the server via the application.

[0126] The server analyzes the uploaded voice data and converts it into text using speech recognition technology, while simultaneously learning the mother's voice using speech synthesis technology to generate a voice model.

[0127] Providing cooking guides

[0128] When cooking, the user selects the desired recipe within the application. The selected recipe request is sent from the device to the server. The server retrieves the corresponding recipe from the recipe database and generates an audio guide in the mother's voice along with cooking instructions.

[0129] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, allowing the user to confirm the cooking instructions both visually and audibly.

[0130] User Customization

[0131] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application, which then sends the information from the device to the server, which then adjusts the recipe accordingly.

[0132] The adjusted recipe is then sent back to the device and displayed to the user, with audio guidance provided if needed, allowing the user to create a dish tailored to their own preferences.

[0133] Specific examples

[0134] For example, consider a case where a user selects "curry" and prefers it "low salt" and "medium spicy." The system digitizes the mother's curry recipe using OCR technology and AI analysis. It then generates a recipe with adjusted salt content based on the user's customization information. Finally, the mother's voice provides audio guidance, saying, "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also creating a dish that suits their own preferences.

[0135] Examples of prompt statements

[0136] "Please digitize my mother's homemade curry recipe using OCR technology and AI analysis. Then please guide me through the recipe, which is low in salt and medium in spiciness, using my mother's voice."

[0137] This system makes it possible to accurately pass on the flavor of a mother's cooking to the next generation, further strengthening family bonds.

[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0139] Step 1: Data collection

[0140] User: Takes a photo of his mother's recipe notebook or a video of his mother cooking on his smartphone. The input is the image or video of the recipe.

[0141] Specific operation: The user launches the application and uses the camera function to take notes or record videos.

[0142] Terminal: Uploads captured photos and videos to the server via the application. The output is the transmission of image and video data to the server.

[0143] Specific operation: When you press the "Upload" button within the application, the device will transfer the captured data to the server.

[0144] Step 2: Convert the recipe to text

[0145] Server: Receives uploaded photos and videos and converts them into text data using OCR technology. The input is image or video data, and the output is text data.

[0146] Specific operation: The server uses OCR (Optical Character Recognition) software to convert the received image data into text.

[0147] Step 3: Recipe analysis and database storage

[0148] Server: Text-based recipe data is fed into the AI ​​model to extract information such as ingredients, quantities, cooking steps, etc. The input is text data, and the output is the extracted structured data.

[0149] How it works: The server inputs text data into the AI ​​model and executes a script that automatically analyzes and extracts ingredients and procedures.

[0150] Server: The extracted data is structured and stored in a data repository. The output is structured data stored in a database.

[0151] What it does: The server stores structured data in a relational database, allowing for efficient searching and management.

[0152] Step 4: Collecting and training audio data

[0153] User: Records the voice of the mother explaining the cooking procedure on a smartphone. The input is the mother's voice.

[0154] Specific actions: Use the application's voice recording function to record the mother's explanation.

[0155] Terminal: Uploads recorded data from the application to the server. The output is sending audio data to the server.

[0156] Specific operation: After recording is finished, press the "Upload" button and the audio data will be transferred to the server.

[0157] Server: Converts voice data into text using speech recognition technology, and learns the mother's voice using speech synthesis technology. The input is voice data, and the output is a voice model.

[0158] How it works: The server uses speech recognition software to convert speech to text, and then trains speech synthesis software to learn the mother's voice.

[0159] Step 5: Provide cooking guidance

[0160] User: Selects the recipe they want to cook within the application. Input is a recipe selection request.

[0161] Specific operation: Open "Recipe List" from the application menu and tap to select the desired recipe.

[0162] Terminal: Sends a request for the selected recipe to the server. The output is the request to the server.

[0163] Specific operation: When a user selects a recipe, a request is sent to the server.

[0164] Server: Retrieves the relevant recipe from the database and generates an audio guide in a mother's voice along with cooking instructions. The input is a recipe request, and the output is the audio guide data.

[0165] Specific operation: The server retrieves the relevant recipe from the database and generates a guide voice using the voice model.

[0166] Terminal: Displays cooking instructions to the user and provides audio guidance. Output is display data and audio guidance to the user.

[0167] Specific operation: The cooking steps are presented to the user visually and audibly. For example, the device may say, "First, cut the chicken into bite-sized pieces."

[0168] Step 6: Customizing Users

[0169] User: Enter allergies and taste preferences into the app. Input is customization information.

[0170] Specific operation: Access the form for entering customization information from the application's "Settings" screen and enter the required information.

[0171] Terminal: Sends input customization information to the server. Output: Sends customization information to the server.

[0172] Specific operation: When you press the "Save" button, the input information is sent to the server.

[0173] Server: Adjusts the recipe based on the customization information and generates appropriate instructions. The input is the customization information and the output is the adjusted recipe.

[0174] Specific operation: The server adjusts the amounts of ingredients and seasonings in the recipe based on the customization information.

[0175] Terminal: Displays the customized recipe and provides voice guidance if necessary. The output is the display data and voice guidance for the user.

[0176] What it does: It shows the user the adjusted recipe and provides audio guidance such as "Add a little salt to the chicken."

[0177] (Application example 1)

[0178] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0179] Traditionally, it has been difficult to accurately pass on family recipes to the next generation, especially when it comes to recreating the flavor of a mother's cooking. Enjoying these dishes through delivery services is also uncommon, making it even more challenging to accommodate individual taste preferences and dietary restrictions. This limits opportunities to deepen family memories and bonds. Continuously improving recipes based on feedback is also a challenge.

[0180] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0181] In this invention, the server includes: means for digitizing mother's cooking recipes; means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps; means for storing the extracted data in a database; means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology; means for providing audio guidance to the user on cooking steps in mother's voice; means for cooking dishes based on the recipes and delivering them to the user via a delivery service; and means for collecting user feedback data and updating the recipe database. This allows mother's special recipes to be faithfully passed on to the next generation and for the user to enjoy the dishes through delivery. Furthermore, recipes can be continuously improved based on user feedback, thereby providing a consistently high-quality cooking experience.

[0182] The "recipe digitization method" is a method for collecting mother's recipes as photos and videos and converting them into digital format.

[0183] The "recipe analysis means" is a means for extracting ingredients, quantities, and cooking steps from digitized recipe data.

[0184] The "database storage means" is a means for structuring the analyzed recipe data and storing it in a database.

[0185] The "audio collection means" is a means for collecting and recording the audio of the mother explaining the cooking steps.

[0186] A "speech recognition means" is a means that uses technology to convert collected voice data into text.

[0187] The "voice cloning technology learning means" is a means for learning the mother's voice and reproducing the mother's voice based on the voice data.

[0188] The "audio guidance means" is a means for providing the user with audio guidance on cooking procedures in a mother's voice.

[0189] A "cooking tool" is a tool for cooking a dish based on a recipe.

[0190] "Delivery service means" refers to a means for delivering cooked food to a user through a delivery service.

[0191] A "feedback collection means" is a means for collecting feedback data from users.

[0192] The "recipe database update means" is a means for updating the recipe database based on the collected feedback data.

[0193] A "generative AI model" refers to an artificial intelligence model that extracts specific information from text data and performs analysis.

[0194] A "prompt sentence" is an input sentence for a generative AI model, and is text that contains instructions for analysis.

[0195] This invention is a system that digitizes mother's cooking recipes, accurately passes them on to the next generation, and provides customized meals through a delivery service. The system aims to allow users to recreate their mother's special dishes and deepen family ties.

[0196] Hardware and software used

[0197] Uses hardware and software that integrates servers, smartphones, and delivery services, including:

[0198] Smartphone: Used as data collection and operation interface.

[0199] Server: A central processing unit that performs data analysis, voice recognition, voice cloning, recipe generation, and delivery management.

[0200] Google Cloud Vision API: Uses OCR technology to convert digitized recipes into text.

[0201] OpenAI GPT-4: A generative AI model that analyzes recipe data and extracts ingredients and cooking steps.

[0202] Google Speech-to-Text API: Converts audio data into text.

[0203] Lyrebird or Descript: A voice cloning technology that recreates the mother's voice.

[0204] Uber Eats API: Manages delivery services to deliver prepared meals to users.

[0205] Data processing and operation procedures

[0206] 1. Data Collection

[0207] The user uses their smartphone to record their mother's recipe notes and videos of their mother cooking.

[0208] Photos and videos taken by the device are uploaded to the server via the application.

[0209] 2. Recipe analysis and database storage

[0210] The server receives the uploaded data and converts it into text using OCR technology using the Google Cloud Vision API.

[0211] The server inputs the recipe data, converted into text using OCR, into OpenAI GPT-4 to extract information such as ingredients, quantities, and cooking steps.

[0212] The server structures the extracted data and stores it in a database.

[0213] 3. Collecting and Learning Speech Data

[0214] The user records the audio of their mother explaining the cooking steps on their smartphone.

[0215] The device uploads the recording data from the application to the server.

[0216] The server analyzes the uploaded voice data using the Google Speech-to-Text API and converts it into text, while simultaneously training the mother's voice using Lyrebird or Descript to generate a voice model.

[0217] 4. Providing cooking guides

[0218] The user selects the recipe they want to cook within the application.

[0219] The device sends a request for the selected recipe to the server.

[0220] The server retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[0221] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, such as "First, cut the chicken into bite-sized pieces."

[0222] 5. Food preparation and delivery

[0223] The server arranges for the food to be cooked based on the generated recipe and cooking instructions.

[0224] Once cooked, the food is delivered to the user using the Uber Eats API.

[0225] 6. User Feedback

[0226] The user enters post-cooking feedback into the application.

[0227] The server analyzes the feedback data and updates the recipe database to improve the accuracy of future cooking guides.

[0228] Specific examples

[0229] For example, if a user selects "Mom's Special Curry" and requests "low salt" and "medium spicy," the system analyzes a photo of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. A recipe with adjusted salt content is generated based on the user's low-salt specifications. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." After cooking, the food is delivered to the user via Uber Eats.

[0230] Prompt Sentence Examples

[0231] Analyze photos, videos, and audio data of Mom's recipes uploaded by users to extract ingredients and cooking steps, clone Mom's voice from the audio data, and generate an audio guide that provides cooking instructions in Mom's voice along with the user's desired customization options. Once cooking is complete, deliver the food using the Uber Eats API.

[0232] As a result, the present invention can provide a high-quality cooking experience by faithfully passing on mother's special recipes to the next generation and customizing them according to the user's wishes.

[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0234] Step 1:

[0235] Data collection

[0236] The user takes photos of their mother's cooking recipe notes or videos of their mother cooking with their smartphone. The device then uploads the photos and videos to a server via an application. The input is the photos and videos, and the output is the data sent to the server.

[0237] Step 2:

[0238] Text conversion using OCR

[0239] The server receives the uploaded photos and videos and converts them into text using OCR technology using the Google Cloud Vision API. The input is the photo or video data, and the output is the recipe data in text format. This extracts text information from the photos and videos.

[0240] Step 3:

[0241] Recipe analysis and database storage

[0242] The server inputs the recipe data converted to text using OCR into OpenAI GPT-4, and uses a generative AI model to extract information such as ingredients, quantities, and cooking steps. The input is the converted recipe data, and the output is the extracted structured data. The server then structures the extracted data and stores it in a database. This allows the recipe's digital information to be organized and stored.

[0243] Step 4:

[0244] Audio data collection

[0245] The user records the voice of his mother explaining the cooking procedure on his smartphone. The device uploads the recorded data to the server via the application. The input is the recorded voice data, and the output is the voice data sent to the server.

[0246] Step 5:

[0247] Speech Recognition and Speech Clone Training

[0248] The server analyzes the uploaded audio data using the Google Speech-to-Text API and converts the audio to text. At the same time, it uses Lyrebird or Descript to train the mother's voice and generate a voice model. The input is the audio data, and the output is the converted audio data and a voice model of the mother's voice.

[0249] Step 6:

[0250] Generate cooking guide

[0251] The user selects the recipe they want to cook within the application. The device sends a request for the selected recipe to the server. The server retrieves the recipe from the recipe database and generates audio instructions for the cooking process in a mother's voice. The input is the selected recipe request, and the output is the cooking process with audio instructions.

[0252] Step 7:

[0253] Food preparation and delivery

[0254] The server arranges for the food to be cooked based on the generated recipe and cooking instructions. The cooked food is delivered to the user using the Uber Eats API. The input is the cooking instructions and delivery address information, and the output is the delivered food.

[0255] Step 8:

[0256] Feedback collection and database updates

[0257] Users input their post-cooking feedback into the application. The server analyzes the feedback data and updates the recipe database. The input is the feedback data, and the output is an updated recipe database. This improves the accuracy of cooking guidance from the next time onwards.

[0258] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0259] This invention is a system that digitizes mother's cooking recipes to pass them on to the next generation, and combines them with an emotion engine that recognizes the user's emotions to provide greater emotional connection and satisfaction. The program of this system includes functions for data collection, analysis, voice guidance, customization, feedback, and emotion recognition.

[0260] 1. Data Collection

[0261] User: First, the user takes a photo of the notebook containing his mother's cooking recipes and a video of his mother cooking with his smartphone.

[0262] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[0263] Server: Receives the uploaded data and temporarily stores it in a database.

[0264] 2. Digitizing recipes

[0265] Server: Uses OCR technology to extract text information from photos and videos and convert recipes into text format.

[0266] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[0267] 3. Collecting and Learning Audio Data

[0268] User: Records mother explaining cooking steps on smartphone.

[0269] Terminal: Upload the recorded audio data to the server using a dedicated application.

[0270] Server: Receives the uploaded audio data and stores it in a database.

[0271] Server: Extracts text data from the voice data using voice recognition technology, learns the mother's voice using voice cloning technology, and generates a voice model.

[0272] 4. Parse and save the recipe

[0273] Server: The text data obtained by OCR is input into an AI model to analyze information such as ingredients, quantities, and cooking procedures.

[0274] Server: The analyzed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0275] 5. User-selected and customized recipes

[0276] User: Select the recipe they want to cook within the dedicated application.

[0277] Terminal: Sends a request for the selected recipe to the server.

[0278] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0279] Terminal: Sends the entered customization information to the server.

[0280] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0281] 6. Providing cooking guides

[0282] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0283] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0284] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0285] 7. User Emotion Recognition

[0286] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone.

[0287] Terminal: The emotion engine sends the user's emotion data to the server.

[0288] Server: Adjusts the tone and content of the audio guidance during cooking based on emotional data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[0289] Server: It can analyze the user's emotional data and suggest recipes and cooking procedures that suit their emotions.

[0290] 8. Gather and incorporate feedback

[0291] User: After cooking is complete, enter feedback into the app about the taste and process of the dish.

[0292] Terminal: Sends the input feedback data to the server.

[0293] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0294] Specific examples

[0295] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as "It's okay to take it slowly."

[0296] By combining the above functions, the system of the present invention can accurately convey the taste of mother's cooking to the next generation, further deepening emotional connections and satisfaction.

[0297] The processing flow will be explained below.

[0298] Step 1: Data collection

[0299] User: Records his mother's cooking recipes in a notebook and takes photos of his mother cooking with his smartphone.

[0300] Terminal: Uploads captured photos and videos to the server using a dedicated application.

[0301] Server: Receives the uploaded data and temporarily stores it in a database.

[0302] Step 2: Digitize your recipes

[0303] Server: Extracts text information from uploaded images and videos using OCR technology.

[0304] Server: Analyzes the extracted text data and classifies and organizes it into ingredients, quantities, and cooking steps in the recipe.

[0305] Step 3: Collecting and training audio data

[0306] User: Records mother explaining cooking steps on smartphone.

[0307] Terminal: Uploads recorded audio data to the server via a dedicated application.

[0308] Server: Receives the uploaded audio data and stores it in a database.

[0309] Server: Uses voice recognition technology to convert voice data into text data.

[0310] Server: Using voice cloning technology, the characteristics of the mother's voice are learned and a voice model is generated.

[0311] Step 4: Parse and save the recipe

[0312] Server: The text data obtained by OCR is input into an AI model to analyze ingredients, quantities, and cooking procedures.

[0313] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0314] Step 5: User Recipe Selection and Customization

[0315] User: Select the recipe they want to cook within the dedicated application.

[0316] Terminal: Sends a request for the selected recipe to the server.

[0317] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0318] Terminal: Sends the entered customization information to the server.

[0319] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0320] Step 6: Provide cooking guidance

[0321] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0322] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0323] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0324] Step 7: Recognizing User Emotions

[0325] Device: The emotion engine analyzes the user's facial expressions and voice using the camera and microphone.

[0326] Terminal: Transmits the analyzed emotion data to the server.

[0327] Server: Adjusts the tone and content of the audio guide during cooking based on emotional data. For example, if it detects that the user is feeling stressed, it will soften the tone of the audio guide and insert encouraging words such as "It's okay to take it slowly."

[0328] Step 8: Suggest an emotionally appropriate recipe

[0329] Server: Analyzes the user's emotional data and suggests recipes and cooking procedures that suit their emotional state. For example, if the user is having fun, it can suggest more challenging recipes.

[0330] Terminal: Notifies users of emotion-based recipe suggestions.

[0331] Step 9: Gather and incorporate feedback

[0332] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[0333] Terminal: Sends feedback data to the server.

[0334] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0335] Example 2

[0336] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0337] Conventional systems offer limited ways to pass on mother's cooking recipes to the next generation and lack mechanisms for providing emotional connection and satisfaction. They also struggle to customize to accommodate users' dietary restrictions and individual preferences, and are unable to recognize users' emotions and provide appropriate assistance while cooking. This increases the likelihood of users feeling stressed or creating unsatisfying dishes.

[0338] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0339] In this invention, the server includes means for digitizing mother's cooking recipes, means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in an information management system, means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in mother's voice, and means for recognizing the user's emotions and adjusting the tone and content of the audio guidance based on those emotions. This makes it possible to accurately pass on mother's cooking recipes to the next generation and provide cooking support that takes the user's emotions into consideration, providing a high level of satisfaction.

[0340] "Digitization" is the process of converting analog information into digital form.

[0341] "Analysis" is the process of examining data in detail to understand its structure and meaning.

[0342] "Ingredients" refers to the food ingredients used to make a dish.

[0343] "Quantity" refers to the specific amount of an ingredient used in a dish.

[0344] A "cooking procedure" is a sequence of steps for cooking a dish.

[0345] An "information management system" is a system for storing, managing, retrieving, and using data.

[0346] "Audio data" refers to digital information that records sound.

[0347] "Speech recognition" is a technology that converts voice data into text data.

[0348] "Voice cloning technology" is a technology for reproducing a specific voice.

[0349] "Audio guide" is a means of giving instructions or explanations using voice.

[0350] "User" refers to a person who uses this system.

[0351] "Customization" refers to modifying a system or service to meet a user's specific needs or requirements.

[0352] "Emotion recognition" is a technology that analyzes changes in a user's facial expressions and voice to determine their emotions.

[0353] "Feedback Data" refers to the evaluations and opinions provided by users after using the system.

[0354] This invention is a system that digitizes mother's cooking recipes, passes them on to the next generation, and recognizes the user's emotions. This system includes the following functions: data collection, digitization, voice data collection and learning, recipe analysis and storage, user recipe selection and customization, cooking guide provision, user emotion recognition, and feedback collection and reflection.

[0355] Data collection

[0356] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. For example, he takes photos of each page of the notebook and records a video of his mother cooking in the kitchen.

[0357] Device: Photos and videos are uploaded to the server using a dedicated application. The device checks the data format and network connection status to upload the data appropriately.

[0358] Server: Receives the uploaded data and temporarily stores it in the database. At this time, it checks the format and size of the data and converts it into the appropriate format.

[0359] Recipe digitization

[0360] Server: Using OCR technology such as Google Cloud Vision API, text information is extracted from photos and videos and converted into text format.

[0361] Server: Analyzes text data and classifies and organizes it into recipe ingredients, quantities, and steps. For example, converts "200g of chicken" and "2 potatoes" in an image into text data.

[0362] Audio data collection and learning

[0363] User: Record your mother explaining cooking steps on your smartphone. For example, record your mother explaining, "Cut the chicken into bite-sized pieces."

[0364] Terminal: The recorded audio data is uploaded to the server using a dedicated application. At this time, the format of the audio data is checked and converted to the appropriate format.

[0365] Server: Receives the uploaded voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[0366] Recipe analysis and saving

[0367] Server: The text data obtained by OCR is input into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking procedures.

[0368] Server: The analyzed data is structured and stored in a database such as MongoDB. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0369] User-selected and customized recipes

[0370] User: Selects the recipe they want to cook within the dedicated application, for example, curry or stew.

[0371] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[0372] Terminal: Sends the input customization information to the server. The server adjusts the recipe based on the conditions entered by the user.

[0373] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0374] Providing cooking guides

[0375] Server: Retrieves the selected recipe from the database and generates cooking instructions. Specific steps are clearly displayed and presented to the user in an easy-to-understand manner.

[0376] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0377] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[0378] User Emotion Recognition

[0379] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone. For example, it monitors the user's face in real time and analyzes their emotions.

[0380] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[0381] Server: Adjusts the tone and content of the audio guide based on emotional data. For example, if the user is feeling stressed, the tone of the guide can be changed to a gentler tone. It can also suggest recipes and cooking procedures that are suitable for the user.

[0382] Gathering and implementing feedback

[0383] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user comments on the strength of the flavor and the ease of understanding the process.

[0384] Terminal: Sends feedback data to the server. Collects user evaluations as digital data.

[0385] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time, for example, adjusting the taste based on the feedback.

[0386] Specific examples

[0387] For example, if a user selects "curry" and prefers low salt and medium spiciness, the system will convert their mother's curry recipe into text using OCR and AI technology. Based on the user's customized information, a recipe with adjusted salt content will be generated. A voice guide will be provided in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." If the user becomes stressed, the emotion engine will detect this and change the tone of the voice guide to a gentler tone, encouraging the user by saying, "It's okay to take it easy."

[0388] Example prompts for generative AI models

[0389] "Digitalize your mother's recipes and design a cooking support system that combines voice guidance and emotion recognition."

[0390] "Explain OCR and AI techniques used to analyze photos of cooking recipes and extract ingredients and steps."

[0391] The system of the present invention not only passes on mother's cooking recipes to the next generation, but also provides cooking support that takes the user's emotions into consideration, thereby providing greater satisfaction and emotional connection.

[0392] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0393] Step 1: Data collection

[0394] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. Specifically, he carefully photographs each page of the notebook and records a video of his mother cooking in the kitchen.

[0395] Input: Photos and videos taken

[0396] Output: Captured data is saved on your smartphone (image and video files saved in local storage)

[0397] Device: Uploads photos and videos to the server using a dedicated application. During upload, the network connection status and data integrity are checked.

[0398] Input: Image and video data stored on a smartphone

[0399] Output: Data upload process to the server is completed, and data saved on the server

[0400] Step 2: Digitize your recipes

[0401] Server: Using OCR technology such as Google Cloud Vision API, character information is extracted from photos and videos and converted into text. For example, handwritten or printed characters in an image are converted into text.

[0402] Input: Uploaded photo and video data

[0403] Output: Text data (e.g., "200g chicken, 2 potatoes, cut the chicken into bite-sized pieces")

[0404] Server: Analyzes text data and classifies it into ingredients, quantities, and cooking procedures. Using natural language processing technology, the text is analyzed by phrase and classified into categories.

[0405] Input: Extracted text data

[0406] Output: Structured data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Steps: Cut the chicken into bite-sized pieces")

[0407] Server: Stores organized data in an information management system.

[0408] Input: Structured data

[0409] Output: Recipe data stored in the database

[0410] Step 3: Collecting and training audio data

[0411] User: Record your mother explaining the cooking steps on your smartphone. Specifically, record your mother explaining, "Cut the chicken into bite-sized pieces."

[0412] Input: Recorded audio data

[0413] Output: Audio file saved on your smartphone (e.g. mp3 or wav format)

[0414] Device: The recorded audio data is uploaded to the server using a dedicated application. During the upload, the format and integrity of the audio file are checked.

[0415] Input: Audio file

[0416] Output: Audio data sent to the server

[0417] Server: Receives the voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[0418] Input: Uploaded audio data

[0419] Output: Text data and speech model

[0420] Step 4: Parse and save the recipe

[0421] Server: The text data acquired by OCR is fed into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking steps. The AI ​​model analyzes the context and meaning, clearly distinguishing between ingredients and steps.

[0422] Input: Text data obtained by OCR

[0423] Output: Analyzed data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Procedure: Cut the chicken into bite-sized pieces")

[0424] Server: The parsed data is structured and stored in a database such as MongoDB. The structured data is saved in an easy-to-search format.

[0425] Input: Parsed data

[0426] Output: Recipe information stored in the database

[0427] Step 5: User Recipe Selection and Customization

[0428] User: Selects the recipe they want to cook within the dedicated application. For example, they can select a dish such as curry or stew.

[0429] Input: Recipe selection by user

[0430] Output: The selected recipe information is displayed in the application.

[0431] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[0432] Input: User preferences and conditions

[0433] Output: Customization information sent to the server

[0434] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions. Adjusts the recipe based on the conditions specified by the user.

[0435] Input: Customization information

[0436] Output: Customized recipe information

[0437] Step 6: Provide cooking guidance

[0438] Server: Retrieves the selected recipe from the database and generates cooking instructions. It clearly shows the specific steps and provides a unified guide for the user.

[0439] Input: Recipe information selected by the user

[0440] Output: Cooking instructions

[0441] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0442] Input: cooking instructions, speech model

[0443] Output: Audio guide data

[0444] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[0445] Input: Audio guide data

[0446] Output: Text display and voice guide

[0447] Step 7: Recognizing User Emotions

[0448] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone, for example, to check the user's stress level in real time.

[0449] Input: User's facial expression and voice data

[0450] Output: Parsed emotion data

[0451] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[0452] Input: Parsed emotion data

[0453] Output: Emotion data sent to the server

[0454] Server: Adjust the tone and content of the audio guide based on the emotional data. For example, if the user is feeling stressed, change the tone of the guide to a gentler tone.

[0455] Input: Emotion data

[0456] Output: Adjusted audio description data

[0457] Step 8: Gather and incorporate feedback

[0458] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user enters the difficulty of cooking and their impressions of the taste.

[0459] Input: User feedback

[0460] Output: Feedback data is saved in the app

[0461] Terminal: Sends feedback data to the server. Collects user ratings and comments digitally.

[0462] Input: User feedback data

[0463] Output: Feedback data sent to the server

[0464] Server: Analyzes the feedback data and updates the database to improve the cooking procedures and recipes for the next time. For example, improve a process that received poor taste ratings.

[0465] Input: Feedback data

[0466] Output: Updated recipe database

[0467] (Application example 2)

[0468] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0469] Traditionally, it has been difficult to accurately pass on a mother's cooking recipes to the next generation, and there is a need for a way to preserve those flavors for future generations, especially in an aging society. Furthermore, existing digital recipe systems lack guidance that takes into consideration the user's emotions, making it difficult to reduce stress and improve satisfaction while cooking. Furthermore, when considering use in brick-and-mortar stores, a system that can handle on-site challenges is required.

[0470] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0471] In this invention, the server includes means for digitizing mother's cooking manual, means for analyzing the digitized recipe data and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in a database, means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in the mother's voice, and means for recognizing the user's emotions using an emotion engine and adjusting the tone of the audio guidance. This makes it possible to accurately convey the taste of mother's cooking to the next generation and provide guidance that takes user emotions into consideration, making it possible to provide high satisfaction even when cooking in a physical store.

[0472] A "cooking manual" is a document and video data that contains recipes and procedures for the dishes that a mother would make on a daily basis.

[0473] "Digitalization" means converting information in analog form into digital data.

[0474] "Recipe data" is information that describes the ingredients, quantities, and cooking steps for making a dish.

[0475] A "database" is a digital information storage area that systematically organizes multiple data sets and makes them easy to manage and search.

[0476] "Speech recognition" is a technology that converts human speech into data in a format that a computer can understand.

[0477] "Voice cloning technology" is a technology that learns the characteristics of a specific person's voice and can read any text in that person's voice.

[0478] "Audio guide" is a function that uses voice to give instructions and explanations to users.

[0479] An "emotion engine" is a system that uses sensors such as cameras and microphones to analyze and recognize the user's emotional state and then respond appropriately based on that.

[0480] The system for carrying out this invention digitizes a mother's cooking manual and provides audio guidance of cooking procedures in a manner that takes into consideration the user's feelings. The following is a specific embodiment of this system.

[0481] The system begins with the user taking a photo of their mother's cooking manual. Using a smartphone or smart glasses, the user takes photos of the notes in the cooking manual and their mother cooking. The photos and videos are then uploaded to a server using a dedicated application. The server receives the uploaded data and temporarily stores it in a database.

[0482] The server then uses OCR technology to extract text from photos and videos, converting the recipe data into text format, which is then categorized and organized into recipe ingredients, quantities, and cooking steps.

[0483] When the user provides a voice description of their mother's cooking steps, the voice data recorded on their smartphone is also uploaded to the server using a dedicated application. The server receives the voice data and stores it in a database. Furthermore, it uses voice recognition technology to extract text data from the voice data, and uses voice cloning technology to learn the mother's voice and generate a voice model.

[0484] The server inputs the text data acquired by OCR into an AI model and analyzes information such as ingredients, quantities, cooking procedures, etc. The analyzed data is then structured and stored in a database.

[0485] When a user selects a recipe they want to cook within the application, the request is sent to the server. Based on the user's input of allergies and taste preferences (e.g., "low salt" or "medium spicy"), the server adjusts the recipe and generates appropriate instructions. The adjusted recipe and cooking steps are then audio-guided using the mother's voice.

[0486] While the user is cooking, the device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion engine, and sends the emotion data to the server. The server then adjusts the tone and content of the audio guidance during cooking based on the emotion data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[0487] After cooking, the user enters feedback about the taste and cooking procedure into the app. The device sends the feedback data to the server, which analyzes it and updates the database to improve the cooking procedure and recipe for the next time.

[0488] As a specific example, let's assume that the user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes the mother's video of the curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the low-salt specifications selected by the user, a recipe is generated with the salt amount adjusted. Based on this customized information, the mother's voice provides audio guidance such as, "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as, "It's okay to take it slowly."

[0489] Here are some example prompts using a generative AI model:

[0490] "Digitize your mom's curry recipe, customize it to low salt and medium spice, and provide audio guidance in your mom's voice. If the user is stressed, make the guidance tone gentler."

[0491] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0492] Step 1:

[0493] The user uses a smartphone or smart glasses to take pictures of their mother's cooking manual notes or the mother cooking. The input data is photos or videos, which are uploaded by the application and sent to the server. The server receives the uploaded images and videos and temporarily stores them in a database. This saves the physical recipe information as digital data.

[0494] Step 2:

[0495] The server uses OCR technology to extract text information from the received images and videos. The input is the image or video saved in step 1, and the output is the extracted text data. Specifically, the server uses the pytesseract library to perform character recognition and convert it into text recipe data.

[0496] Step 3:

[0497] The server categorizes and organizes the text data extracted by OCR into categories such as ingredients, quantities, and cooking procedures. The input is the text data obtained from OCR, and the output is structured recipe data. The text is analyzed using an AI model and stored in a database. This allows for efficient management of recipe information.

[0498] Step 4:

[0499] The user records their mother's cooking instructions on their smartphone. The data input is voice data, which is uploaded using a dedicated application and sent to the server. The server receives the voice data and stores it in a database.

[0500] Step 5:

[0501] The server uses speech recognition technology to extract text data from the uploaded audio data. The input is the audio data saved in step 4, and the output is the extracted text data. It then uses voice cloning technology to learn the mother's voice and generate a voice model. This makes it possible to read any text aloud in the mother's voice.

[0502] Step 6:

[0503] The user selects the recipe they want to cook within the dedicated application. The input is the user's recipe selection information, which is sent from the terminal to the server. Based on this information, the server retrieves the corresponding recipe data from the database.

[0504] Step 7:

[0505] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application. The input is customization information, which is sent from the device to the server. The server adjusts the recipe based on the customization information and generates new cooking instructions.

[0506] Step 8:

[0507] The server uses the trained voice model to generate audio guidance based on the adjusted recipe. The input is the adjusted recipe data, and the output is audio guidance. The audio guidance instructs the user on the cooking steps in a mother's voice.

[0508] Step 9:

[0509] While cooking, the device (smart glasses or head-mounted display) uses a camera and microphone to analyze the user's facial expressions and voice with an emotion engine. The input is the user's facial expressions and voice data, and the emotion data is sent from the device to the server.

[0510] Step 10:

[0511] The server adjusts the tone and content of the audio guide based on the emotion data. The input is emotion data, and the output is the adjusted audio guide. For example, if the user is feeling stressed, the tone of the guide can be made gentler and words of encouragement can be added.

[0512] Step 11:

[0513] After cooking is complete, the user enters feedback about the taste and cooking procedure in the app. The input is feedback data, which is sent from the device to the server. The server analyzes the feedback data and updates the database to improve the cooking procedure and recipe for the next time.

[0514] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0515] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0516] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0517] [Second embodiment]

[0518] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0519] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0520] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0521] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0522] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0523] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0524] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0525] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0526] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0527] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0528] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0529] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0530] This invention is a digital system for accurately passing on mother's cooking recipes to the next generation, aiming to deepen special family memories and bonds while recreating the flavor of mother's cooking. The program of this system includes functions for data collection, analysis, audio guidance, customization, and feedback.

[0531] 1. Data Collection

[0532] User: First, the user takes a note of his mother's cooking recipes and a video of his mother cooking on his smartphone.

[0533] Device: Upload the photos and videos you have taken to the server via the application.

[0534] Server: Receives the uploaded data and converts the recipe into text using OCR technology.

[0535] 2. Recipe analysis and database storage

[0536] Server: Recipe data converted to text using OCR is fed into the AI ​​model, and information such as ingredients, quantities, and cooking steps is extracted.

[0537] Server: Structures the extracted data and stores it in a database.

[0538] 3. Collecting and Learning Audio Data

[0539] User: Records mother explaining cooking steps on smartphone.

[0540] Device: Upload the recording data from the application to the server.

[0541] Server: The uploaded voice data is analyzed through a voice recognition process and converted into text. At the same time, voice cloning technology is used to learn the mother's voice and generate a voice model.

[0542] 4. Providing cooking guides

[0543] User: Selects the recipe they want to cook within the application.

[0544] Terminal: Sends a request for the selected recipe to the server.

[0545] Server: Retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[0546] Device: The acquired cooking instructions are displayed to the user, and instructions such as "First, cut the chicken into bite-sized pieces" are also provided in a mother's voice.

[0547] 5. User Customization

[0548] User: Enter allergies and individual taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0549] Terminal: Sends the entered customization information to the server.

[0550] Server: Adjusts the recipe based on the customization information and generates the appropriate instructions.

[0551] Device: Displays the customized recipe and provides audio guidance if needed.

[0552] Specific examples

[0553] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is given audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also making dishes that accommodate their own dietary restrictions.

[0554] In addition, after cooking, users can enter feedback into the app, which is analyzed by the server and the recipe database is constantly updated to provide more accurate guidance the next time the user cooks.

[0555] By combining the above functions, the system of the present invention can accurately pass on the taste of mother's cooking to the next generation, further deepening family ties.

[0556] The processing flow will be explained below.

[0557] Step 1: Data collection

[0558] User: Takes notes of his mother's cooking recipes and videos of himself cooking on his smartphone.

[0559] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[0560] Server: Receives the uploaded data and temporarily stores it in a database.

[0561] Step 2: Digitize your recipes

[0562] Server: Using OCR technology, extracts text information from photos and videos and converts recipes into text format.

[0563] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[0564] Step 3: Collecting audio data

[0565] User: Records her mother explaining the recipe steps on her smartphone.

[0566] Terminal: Upload the recorded audio data to the server using a dedicated application.

[0567] Server: Receives the uploaded audio data and stores it in a database.

[0568] Step 4: Analyze and train audio data

[0569] Server: Extracts text data from the voice data using voice recognition technology.

[0570] Server: Organizes the text data into cooking instructions and stores them in a database.

[0571] Server: Using voice cloning technology, learns the characteristics of the mother's voice and generates a voice model.

[0572] Step 5: Parse and save the recipe

[0573] Server: The text data obtained by OCR is fed into the AI ​​model, which analyzes ingredients, quantities, and cooking procedures.

[0574] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0575] Step 6: User selects recipe

[0576] User: Select the recipe they want to cook within the dedicated application.

[0577] Terminal: Sends a request for the selected recipe to the server.

[0578] Step 7: Cooking instructions generation and audio guidance

[0579] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0580] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0581] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0582] Step 8: Enter and update customization information

[0583] User: Enters specific dietary restrictions and taste preferences (e.g., "low salt," "medium spicy," etc.) into a dedicated application.

[0584] Terminal: Sends the entered customization information to the server.

[0585] Server: Adjusts the recipe based on the user's customization information.

[0586] Device: Provides the adjusted recipe to the user, and also provides audio guidance if necessary.

[0587] Step 9: Gather and incorporate feedback

[0588] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[0589] Terminal: Sends the input feedback data to the server.

[0590] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0591] Example 1

[0592] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0593] It is not easy to accurately pass on the flavors and steps of a mother's cooking to the next generation. If recipes are not accurately digitized, there is a high risk of losing special family memories and bonds. It is also difficult to satisfy the need to receive detailed cooking instructions in the mother's voice. Furthermore, it is difficult to accommodate the different dietary restrictions and taste preferences of each family. To solve these problems, accurately pass on the flavors of a mother's cooking to the next generation, and deepen family bonds, an efficient and accurate digital system is needed.

[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0595] In this invention, the server includes means for digitizing mother's recipes, means for analyzing the digitized recipes and extracting ingredients, portions, and cooking methods, means for storing the extracted data in a data repository, means for collecting mother's voice data and learning mother's voice using voice recognition and voice synthesis technology, and means for providing audio guidance on cooking methods to users in mother's voice, thereby enabling mother's recipes to be accurately passed down to future generations and providing cooking guides customized to the specific needs of each household.

[0596] "My mother's cooking methods" refers to the cooking methods, steps, and recipes that my mother uses.

[0597] "Digitalization" refers to the process of converting analog information into electronic data.

[0598] "Ingredients" refers to the various foods and ingredients used in cooking.

[0599] "Amount" refers to the amount of each ingredient used when cooking.

[0600] "Cooking method" refers to a series of processes or steps to complete a dish using ingredients.

[0601] A "data repository" refers to a database or storage system for efficiently storing and managing collected data.

[0602] "Audio data" refers to digital data that contains recorded audio information such as human voices and spoken words.

[0603] "Speech recognition technology" refers to technology that generates text data from voice data.

[0604] "Speech synthesis technology" refers to technology that generates new voices based on certain voice samples.

[0605] "Voice guidance" refers to a method of conveying specific procedures or information to a user by providing voice guidance.

[0606] This invention is a digital system that aims to pass on mother's cooking techniques to the next generation and deepen family memories and bonds. The system has functions for data collection, analysis, audio guidance, customization, and feedback, and is designed to make it easier for users to recreate the taste of their mother's cooking.

[0607] Data collection

[0608] The first thing a user does is digitize their mother's cooking methods. They take photos of their mother's recipe notebooks and videos of their mother cooking with their smartphone. They then upload these photos and videos to a server using a dedicated application.

[0609] The server receives the uploaded photos and videos and converts the recipes into text data using OCR (Optical Character Recognition) technology, a process that converts analog information into electronic data.

[0610] Recipe analysis and database storage

[0611] The server then inputs the recipe data, converted to text using OCR technology, into an AI model to extract information such as ingredients, quantities, cooking methods, etc. Specifically, it uses natural language processing technology to analyze the recipe text and automatically identify the necessary information.

[0612] The extracted data is structured by the server and stored in a data repository, making it easy for users to search and browse later.

[0613] Audio data collection and learning

[0614] The user records their mother explaining the cooking steps on their smartphone, and the recording is then uploaded to the server via the application.

[0615] The server analyzes the uploaded voice data and converts it into text using speech recognition technology, while simultaneously learning the mother's voice using speech synthesis technology to generate a voice model.

[0616] Providing cooking guides

[0617] When cooking, the user selects the desired recipe within the application. The selected recipe request is sent from the device to the server. The server retrieves the corresponding recipe from the recipe database and generates an audio guide in the mother's voice along with cooking instructions.

[0618] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, allowing the user to confirm the cooking instructions both visually and audibly.

[0619] User Customization

[0620] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application, which then sends the information from the device to the server, which then adjusts the recipe accordingly.

[0621] The adjusted recipe is then sent back to the device and displayed to the user, with audio guidance provided if needed, allowing the user to create a dish tailored to their own preferences.

[0622] Specific examples

[0623] For example, consider a case where a user selects "curry" and prefers it "low salt" and "medium spicy." The system digitizes the mother's curry recipe using OCR technology and AI analysis. It then generates a recipe with adjusted salt content based on the user's customization information. Finally, the mother's voice provides audio guidance, saying, "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also creating a dish that suits their own preferences.

[0624] Examples of prompt statements

[0625] "Please digitize my mother's homemade curry recipe using OCR technology and AI analysis. Then please guide me through the recipe, which is low in salt and medium in spiciness, using my mother's voice."

[0626] This system makes it possible to accurately pass on the flavor of a mother's cooking to the next generation, further strengthening family bonds.

[0627] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0628] Step 1: Data collection

[0629] User: Takes a photo of his mother's recipe notebook or a video of his mother cooking on his smartphone. The input is the image or video of the recipe.

[0630] Specific operation: The user launches the application and uses the camera function to take notes or record videos.

[0631] Terminal: Uploads captured photos and videos to the server via the application. The output is the transmission of image and video data to the server.

[0632] Specific operation: When you press the "Upload" button within the application, the device will transfer the captured data to the server.

[0633] Step 2: Convert the recipe to text

[0634] Server: Receives uploaded photos and videos and converts them into text data using OCR technology. The input is image or video data, and the output is text data.

[0635] Specific operation: The server uses OCR (Optical Character Recognition) software to convert the received image data into text.

[0636] Step 3: Recipe analysis and database storage

[0637] Server: Text-based recipe data is fed into the AI ​​model to extract information such as ingredients, quantities, cooking steps, etc. The input is text data, and the output is the extracted structured data.

[0638] How it works: The server inputs text data into the AI ​​model and executes a script that automatically analyzes and extracts ingredients and procedures.

[0639] Server: The extracted data is structured and stored in a data repository. The output is structured data stored in a database.

[0640] What it does: The server stores structured data in a relational database, allowing for efficient searching and management.

[0641] Step 4: Collecting and training audio data

[0642] User: Records the voice of the mother explaining the cooking procedure on a smartphone. The input is the mother's voice.

[0643] Specific actions: Use the application's voice recording function to record the mother's explanation.

[0644] Terminal: Uploads recorded data from the application to the server. The output is sending audio data to the server.

[0645] Specific operation: After recording is finished, press the "Upload" button and the audio data will be transferred to the server.

[0646] Server: Converts voice data into text using speech recognition technology, and learns the mother's voice using speech synthesis technology. The input is voice data, and the output is a voice model.

[0647] How it works: The server uses speech recognition software to convert speech to text, and then trains speech synthesis software to learn the mother's voice.

[0648] Step 5: Provide cooking guidance

[0649] User: Selects the recipe they want to cook within the application. Input is a recipe selection request.

[0650] Specific operation: Open "Recipe List" from the application menu and tap to select the desired recipe.

[0651] Terminal: Sends a request for the selected recipe to the server. The output is the request to the server.

[0652] Specific operation: When a user selects a recipe, a request is sent to the server.

[0653] Server: Retrieves the relevant recipe from the database and generates an audio guide in a mother's voice along with cooking instructions. The input is a recipe request, and the output is the audio guide data.

[0654] Specific operation: The server retrieves the relevant recipe from the database and generates a guide voice using the voice model.

[0655] Terminal: Displays cooking instructions to the user and provides audio guidance. Output is display data and audio guidance to the user.

[0656] Specific operation: The cooking steps are presented to the user visually and audibly. For example, the device may say, "First, cut the chicken into bite-sized pieces."

[0657] Step 6: Customizing Users

[0658] User: Enter allergies and taste preferences into the app. Input is customization information.

[0659] Specific operation: Access the form for entering customization information from the application's "Settings" screen and enter the required information.

[0660] Terminal: Sends input customization information to the server. Output: Sends customization information to the server.

[0661] Specific operation: When you press the "Save" button, the input information is sent to the server.

[0662] Server: Adjusts the recipe based on the customization information and generates appropriate instructions. The input is the customization information and the output is the adjusted recipe.

[0663] Specific operation: The server adjusts the amounts of ingredients and seasonings in the recipe based on the customization information.

[0664] Terminal: Displays the customized recipe and provides voice guidance if necessary. The output is the display data and voice guidance for the user.

[0665] What it does: It shows the user the adjusted recipe and provides audio guidance such as "Add a little salt to the chicken."

[0666] (Application example 1)

[0667] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0668] Traditionally, it has been difficult to accurately pass on family recipes to the next generation, especially when it comes to recreating the flavor of a mother's cooking. Enjoying these dishes through delivery services is also uncommon, making it even more challenging to accommodate individual taste preferences and dietary restrictions. This limits opportunities to deepen family memories and bonds. Continuously improving recipes based on feedback is also a challenge.

[0669] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0670] In this invention, the server includes: means for digitizing mother's cooking recipes; means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps; means for storing the extracted data in a database; means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology; means for providing audio guidance to the user on cooking steps in mother's voice; means for cooking dishes based on the recipes and delivering them to the user via a delivery service; and means for collecting user feedback data and updating the recipe database. This allows mother's special recipes to be faithfully passed on to the next generation and for the user to enjoy the dishes through delivery. Furthermore, recipes can be continuously improved based on user feedback, thereby providing a consistently high-quality cooking experience.

[0671] The "recipe digitization method" is a method for collecting mother's recipes as photos and videos and converting them into digital format.

[0672] The "recipe analysis means" is a means for extracting ingredients, quantities, and cooking steps from digitized recipe data.

[0673] The "database storage means" is a means for structuring the analyzed recipe data and storing it in a database.

[0674] The "audio collection means" is a means for collecting and recording the audio of the mother explaining the cooking steps.

[0675] A "speech recognition means" is a means that uses technology to convert collected voice data into text.

[0676] The "voice cloning technology learning means" is a means for learning the mother's voice and reproducing the mother's voice based on the voice data.

[0677] The "audio guidance means" is a means for providing the user with audio guidance on cooking procedures in a mother's voice.

[0678] A "cooking tool" is a tool for cooking a dish based on a recipe.

[0679] "Delivery service means" refers to a means for delivering cooked food to a user through a delivery service.

[0680] A "feedback collection means" is a means for collecting feedback data from users.

[0681] The "recipe database update means" is a means for updating the recipe database based on the collected feedback data.

[0682] A "generative AI model" refers to an artificial intelligence model that extracts specific information from text data and performs analysis.

[0683] A "prompt sentence" is an input sentence for a generative AI model, and is text that contains instructions for analysis.

[0684] This invention is a system that digitizes mother's cooking recipes, accurately passes them on to the next generation, and provides customized meals through a delivery service. The system aims to allow users to recreate their mother's special dishes and deepen family ties.

[0685] Hardware and software used

[0686] Uses hardware and software that integrates servers, smartphones, and delivery services, including:

[0687] Smartphone: Used as data collection and operation interface.

[0688] Server: A central processing unit that performs data analysis, voice recognition, voice cloning, recipe generation, and delivery management.

[0689] Google Cloud Vision API: Uses OCR technology to convert digitized recipes into text.

[0690] OpenAI GPT-4: A generative AI model that analyzes recipe data and extracts ingredients and cooking steps.

[0691] Google Speech-to-Text API: Converts audio data into text.

[0692] Lyrebird or Descript: A voice cloning technology that recreates the mother's voice.

[0693] Uber Eats API: Manages delivery services to deliver prepared meals to users.

[0694] Data processing and operation procedures

[0695] 1. Data Collection

[0696] The user uses their smartphone to record their mother's recipe notes and videos of their mother cooking.

[0697] Photos and videos taken by the device are uploaded to the server via the application.

[0698] 2. Recipe analysis and database storage

[0699] The server receives the uploaded data and converts it into text using OCR technology using the Google Cloud Vision API.

[0700] The server inputs the recipe data, converted into text using OCR, into OpenAI GPT-4 to extract information such as ingredients, quantities, and cooking steps.

[0701] The server structures the extracted data and stores it in a database.

[0702] 3. Collecting and Learning Speech Data

[0703] The user records the audio of their mother explaining the cooking steps on their smartphone.

[0704] The device uploads the recording data from the application to the server.

[0705] The server analyzes the uploaded voice data using the Google Speech-to-Text API and converts it into text, while simultaneously training the mother's voice using Lyrebird or Descript to generate a voice model.

[0706] 4. Providing cooking guides

[0707] The user selects the recipe they want to cook within the application.

[0708] The device sends a request for the selected recipe to the server.

[0709] The server retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[0710] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, such as "First, cut the chicken into bite-sized pieces."

[0711] 5. Food preparation and delivery

[0712] The server arranges for the food to be cooked based on the generated recipe and cooking instructions.

[0713] Once cooked, the food is delivered to the user using the Uber Eats API.

[0714] 6. User Feedback

[0715] The user enters post-cooking feedback into the application.

[0716] The server analyzes the feedback data and updates the recipe database to improve the accuracy of future cooking guides.

[0717] Specific examples

[0718] For example, if a user selects "Mom's Special Curry" and requests "low salt" and "medium spicy," the system analyzes a photo of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. A recipe with adjusted salt content is generated based on the user's low-salt specifications. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." After cooking, the food is delivered to the user via Uber Eats.

[0719] Prompt Sentence Examples

[0720] Analyze photos, videos, and audio data of Mom's recipes uploaded by users to extract ingredients and cooking steps, clone Mom's voice from the audio data, and generate an audio guide that provides cooking instructions in Mom's voice along with the user's desired customization options. Once cooking is complete, deliver the food using the Uber Eats API.

[0721] As a result, the present invention can provide a high-quality cooking experience by faithfully passing on mother's special recipes to the next generation and customizing them according to the user's wishes.

[0722] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0723] Step 1:

[0724] Data collection

[0725] The user takes photos of their mother's cooking recipe notes or videos of their mother cooking with their smartphone. The device then uploads the photos and videos to a server via an application. The input is the photos and videos, and the output is the data sent to the server.

[0726] Step 2:

[0727] Text conversion using OCR

[0728] The server receives the uploaded photos and videos and converts them into text using OCR technology using the Google Cloud Vision API. The input is the photo or video data, and the output is the recipe data in text format. This extracts text information from the photos and videos.

[0729] Step 3:

[0730] Recipe analysis and database storage

[0731] The server inputs the recipe data converted to text using OCR into OpenAI GPT-4, and uses a generative AI model to extract information such as ingredients, quantities, and cooking steps. The input is the converted recipe data, and the output is the extracted structured data. The server then structures the extracted data and stores it in a database. This allows the recipe's digital information to be organized and stored.

[0732] Step 4:

[0733] Audio data collection

[0734] The user records the voice of his mother explaining the cooking procedure on his smartphone. The device uploads the recorded data to the server via the application. The input is the recorded voice data, and the output is the voice data sent to the server.

[0735] Step 5:

[0736] Speech Recognition and Speech Clone Training

[0737] The server analyzes the uploaded audio data using the Google Speech-to-Text API and converts the audio to text. At the same time, it uses Lyrebird or Descript to train the mother's voice and generate a voice model. The input is the audio data, and the output is the converted audio data and a voice model of the mother's voice.

[0738] Step 6:

[0739] Generate cooking guide

[0740] The user selects the recipe they want to cook within the application. The device sends a request for the selected recipe to the server. The server retrieves the recipe from the recipe database and generates audio instructions for the cooking process in a mother's voice. The input is the selected recipe request, and the output is the cooking process with audio instructions.

[0741] Step 7:

[0742] Food preparation and delivery

[0743] The server arranges for the food to be cooked based on the generated recipe and cooking instructions. The cooked food is delivered to the user using the Uber Eats API. The input is the cooking instructions and delivery address information, and the output is the delivered food.

[0744] Step 8:

[0745] Feedback collection and database updates

[0746] Users input their post-cooking feedback into the application. The server analyzes the feedback data and updates the recipe database. The input is the feedback data, and the output is an updated recipe database. This improves the accuracy of cooking guidance from the next time onwards.

[0747] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0748] This invention is a system that digitizes mother's cooking recipes to pass them on to the next generation, and combines them with an emotion engine that recognizes the user's emotions to provide greater emotional connection and satisfaction. The program of this system includes functions for data collection, analysis, voice guidance, customization, feedback, and emotion recognition.

[0749] 1. Data Collection

[0750] User: First, the user takes a photo of the notebook containing his mother's cooking recipes and a video of his mother cooking with his smartphone.

[0751] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[0752] Server: Receives the uploaded data and temporarily stores it in a database.

[0753] 2. Digitizing recipes

[0754] Server: Uses OCR technology to extract text information from photos and videos and convert recipes into text format.

[0755] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[0756] 3. Collecting and Learning Audio Data

[0757] User: Records mother explaining cooking steps on smartphone.

[0758] Terminal: Upload the recorded audio data to the server using a dedicated application.

[0759] Server: Receives the uploaded audio data and stores it in a database.

[0760] Server: Extracts text data from the voice data using voice recognition technology, learns the mother's voice using voice cloning technology, and generates a voice model.

[0761] 4. Parse and save the recipe

[0762] Server: The text data obtained by OCR is input into an AI model to analyze information such as ingredients, quantities, and cooking procedures.

[0763] Server: The analyzed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0764] 5. User-selected and customized recipes

[0765] User: Select the recipe they want to cook within the dedicated application.

[0766] Terminal: Sends a request for the selected recipe to the server.

[0767] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0768] Terminal: Sends the entered customization information to the server.

[0769] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0770] 6. Providing cooking guides

[0771] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0772] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0773] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0774] 7. User Emotion Recognition

[0775] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone.

[0776] Terminal: The emotion engine sends the user's emotion data to the server.

[0777] Server: Adjusts the tone and content of the audio guidance during cooking based on emotional data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[0778] Server: It can analyze the user's emotional data and suggest recipes and cooking procedures that suit their emotions.

[0779] 8. Gather and incorporate feedback

[0780] User: After cooking is complete, enter feedback into the app about the taste and process of the dish.

[0781] Terminal: Sends the input feedback data to the server.

[0782] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0783] Specific examples

[0784] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as "It's okay to take it slowly."

[0785] By combining the above functions, the system of the present invention can accurately convey the taste of mother's cooking to the next generation, further deepening emotional connections and satisfaction.

[0786] The processing flow will be explained below.

[0787] Step 1: Data collection

[0788] User: Records his mother's cooking recipes in a notebook and takes photos of his mother cooking with his smartphone.

[0789] Terminal: Uploads captured photos and videos to the server using a dedicated application.

[0790] Server: Receives the uploaded data and temporarily stores it in a database.

[0791] Step 2: Digitize your recipes

[0792] Server: Extracts text information from uploaded images and videos using OCR technology.

[0793] Server: Analyzes the extracted text data and classifies and organizes it into ingredients, quantities, and cooking steps in the recipe.

[0794] Step 3: Collecting and training audio data

[0795] User: Records mother explaining cooking steps on smartphone.

[0796] Terminal: Uploads recorded audio data to the server via a dedicated application.

[0797] Server: Receives the uploaded audio data and stores it in a database.

[0798] Server: Uses voice recognition technology to convert voice data into text data.

[0799] Server: Using voice cloning technology, the characteristics of the mother's voice are learned and a voice model is generated.

[0800] Step 4: Parse and save the recipe

[0801] Server: The text data obtained by OCR is input into an AI model to analyze ingredients, quantities, and cooking procedures.

[0802] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0803] Step 5: User Recipe Selection and Customization

[0804] User: Select the recipe they want to cook within the dedicated application.

[0805] Terminal: Sends a request for the selected recipe to the server.

[0806] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[0807] Terminal: Sends the entered customization information to the server.

[0808] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0809] Step 6: Provide cooking guidance

[0810] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[0811] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0812] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[0813] Step 7: Recognizing User Emotions

[0814] Device: The emotion engine analyzes the user's facial expressions and voice using the camera and microphone.

[0815] Terminal: Transmits the analyzed emotion data to the server.

[0816] Server: Adjusts the tone and content of the audio guide during cooking based on emotional data. For example, if it detects that the user is feeling stressed, it will soften the tone of the audio guide and insert encouraging words such as "It's okay to take it slowly."

[0817] Step 8: Suggest an emotionally appropriate recipe

[0818] Server: Analyzes the user's emotional data and suggests recipes and cooking procedures that suit their emotional state. For example, if the user is having fun, it can suggest more challenging recipes.

[0819] Terminal: Notifies users of emotion-based recipe suggestions.

[0820] Step 9: Gather and incorporate feedback

[0821] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[0822] Terminal: Sends feedback data to the server.

[0823] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[0824] Example 2

[0825] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0826] Conventional systems offer limited ways to pass on mother's cooking recipes to the next generation and lack mechanisms for providing emotional connection and satisfaction. They also struggle to customize to accommodate users' dietary restrictions and individual preferences, and are unable to recognize users' emotions and provide appropriate assistance while cooking. This increases the likelihood of users feeling stressed or creating unsatisfying dishes.

[0827] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0828] In this invention, the server includes means for digitizing mother's cooking recipes, means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in an information management system, means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in mother's voice, and means for recognizing the user's emotions and adjusting the tone and content of the audio guidance based on those emotions. This makes it possible to accurately pass on mother's cooking recipes to the next generation and provide cooking support that takes the user's emotions into consideration, providing a high level of satisfaction.

[0829] "Digitization" is the process of converting analog information into digital form.

[0830] "Analysis" is the process of examining data in detail to understand its structure and meaning.

[0831] "Ingredients" refers to the food ingredients used to make a dish.

[0832] "Quantity" refers to the specific amount of an ingredient used in a dish.

[0833] A "cooking procedure" is a sequence of steps for cooking a dish.

[0834] An "information management system" is a system for storing, managing, retrieving, and using data.

[0835] "Audio data" refers to digital information that records sound.

[0836] "Speech recognition" is a technology that converts voice data into text data.

[0837] "Voice cloning technology" is a technology for reproducing a specific voice.

[0838] "Audio guide" is a means of giving instructions or explanations using voice.

[0839] "User" refers to a person who uses this system.

[0840] "Customization" refers to modifying a system or service to meet a user's specific needs or requirements.

[0841] "Emotion recognition" is a technology that analyzes changes in a user's facial expressions and voice to determine their emotions.

[0842] "Feedback Data" refers to the evaluations and opinions provided by users after using the system.

[0843] This invention is a system that digitizes mother's cooking recipes, passes them on to the next generation, and recognizes the user's emotions. This system includes the following functions: data collection, digitization, voice data collection and learning, recipe analysis and storage, user recipe selection and customization, cooking guide provision, user emotion recognition, and feedback collection and reflection.

[0844] Data collection

[0845] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. For example, he takes photos of each page of the notebook and records a video of his mother cooking in the kitchen.

[0846] Device: Photos and videos are uploaded to the server using a dedicated application. The device checks the data format and network connection status to upload the data appropriately.

[0847] Server: Receives the uploaded data and temporarily stores it in the database. At this time, it checks the format and size of the data and converts it into the appropriate format.

[0848] Recipe digitization

[0849] Server: Using OCR technology such as Google Cloud Vision API, text information is extracted from photos and videos and converted into text format.

[0850] Server: Analyzes text data and classifies and organizes it into recipe ingredients, quantities, and steps. For example, converts "200g of chicken" and "2 potatoes" in an image into text data.

[0851] Audio data collection and learning

[0852] User: Record your mother explaining cooking steps on your smartphone. For example, record your mother explaining, "Cut the chicken into bite-sized pieces."

[0853] Terminal: The recorded audio data is uploaded to the server using a dedicated application. At this time, the format of the audio data is checked and converted to the appropriate format.

[0854] Server: Receives the uploaded voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[0855] Recipe analysis and saving

[0856] Server: The text data obtained by OCR is input into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking procedures.

[0857] Server: The analyzed data is structured and stored in a database such as MongoDB. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[0858] User-selected and customized recipes

[0859] User: Selects the recipe they want to cook within the dedicated application, for example, curry or stew.

[0860] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[0861] Terminal: Sends the input customization information to the server. The server adjusts the recipe based on the conditions entered by the user.

[0862] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[0863] Providing cooking guides

[0864] Server: Retrieves the selected recipe from the database and generates cooking instructions. Specific steps are clearly displayed and presented to the user in an easy-to-understand manner.

[0865] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0866] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[0867] User Emotion Recognition

[0868] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone. For example, it monitors the user's face in real time and analyzes their emotions.

[0869] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[0870] Server: Adjusts the tone and content of the audio guide based on emotional data. For example, if the user is feeling stressed, the tone of the guide can be changed to a gentler tone. It can also suggest recipes and cooking procedures that are suitable for the user.

[0871] Gathering and implementing feedback

[0872] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user comments on the strength of the flavor and the ease of understanding the process.

[0873] Terminal: Sends feedback data to the server. Collects user evaluations as digital data.

[0874] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time, for example, adjusting the taste based on the feedback.

[0875] Specific examples

[0876] For example, if a user selects "curry" and prefers low salt and medium spiciness, the system will convert their mother's curry recipe into text using OCR and AI technology. Based on the user's customized information, a recipe with adjusted salt content will be generated. A voice guide will be provided in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." If the user becomes stressed, the emotion engine will detect this and change the tone of the voice guide to a gentler tone, encouraging the user by saying, "It's okay to take it easy."

[0877] Example prompts for generative AI models

[0878] "Digitalize your mother's recipes and design a cooking support system that combines voice guidance and emotion recognition."

[0879] "Explain OCR and AI techniques used to analyze photos of cooking recipes and extract ingredients and steps."

[0880] The system of the present invention not only passes on mother's cooking recipes to the next generation, but also provides cooking support that takes the user's emotions into consideration, thereby providing greater satisfaction and emotional connection.

[0881] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0882] Step 1: Data collection

[0883] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. Specifically, he carefully photographs each page of the notebook and records a video of his mother cooking in the kitchen.

[0884] Input: Photos and videos taken

[0885] Output: Captured data is saved on your smartphone (image and video files saved in local storage)

[0886] Device: Uploads photos and videos to the server using a dedicated application. During upload, the network connection status and data integrity are checked.

[0887] Input: Image and video data stored on a smartphone

[0888] Output: Data upload process to the server is completed, and data saved on the server

[0889] Step 2: Digitize your recipes

[0890] Server: Using OCR technology such as Google Cloud Vision API, character information is extracted from photos and videos and converted into text. For example, handwritten or printed characters in an image are converted into text.

[0891] Input: Uploaded photo and video data

[0892] Output: Text data (e.g., "200g chicken, 2 potatoes, cut the chicken into bite-sized pieces")

[0893] Server: Analyzes text data and classifies it into ingredients, quantities, and cooking procedures. Using natural language processing technology, the text is analyzed by phrase and classified into categories.

[0894] Input: Extracted text data

[0895] Output: Structured data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Steps: Cut the chicken into bite-sized pieces")

[0896] Server: Stores organized data in an information management system.

[0897] Input: Structured data

[0898] Output: Recipe data stored in the database

[0899] Step 3: Collecting and training audio data

[0900] User: Record your mother explaining the cooking steps on your smartphone. Specifically, record your mother explaining, "Cut the chicken into bite-sized pieces."

[0901] Input: Recorded audio data

[0902] Output: Audio file saved on your smartphone (e.g. mp3 or wav format)

[0903] Device: The recorded audio data is uploaded to the server using a dedicated application. During the upload, the format and integrity of the audio file are checked.

[0904] Input: Audio file

[0905] Output: Audio data sent to the server

[0906] Server: Receives the voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[0907] Input: Uploaded audio data

[0908] Output: Text data and speech model

[0909] Step 4: Parse and save the recipe

[0910] Server: The text data acquired by OCR is fed into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking steps. The AI ​​model analyzes the context and meaning, clearly distinguishing between ingredients and steps.

[0911] Input: Text data obtained by OCR

[0912] Output: Analyzed data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Procedure: Cut the chicken into bite-sized pieces")

[0913] Server: The parsed data is structured and stored in a database such as MongoDB. The structured data is saved in an easy-to-search format.

[0914] Input: Parsed data

[0915] Output: Recipe information stored in the database

[0916] Step 5: User Recipe Selection and Customization

[0917] User: Selects the recipe they want to cook within the dedicated application. For example, they can select a dish such as curry or stew.

[0918] Input: Recipe selection by user

[0919] Output: The selected recipe information is displayed in the application.

[0920] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[0921] Input: User preferences and conditions

[0922] Output: Customization information sent to the server

[0923] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions. Adjusts the recipe based on the conditions specified by the user.

[0924] Input: Customization information

[0925] Output: Customized recipe information

[0926] Step 6: Provide cooking guidance

[0927] Server: Retrieves the selected recipe from the database and generates cooking instructions. It clearly shows the specific steps and provides a unified guide for the user.

[0928] Input: Recipe information selected by the user

[0929] Output: Cooking instructions

[0930] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[0931] Input: cooking instructions, speech model

[0932] Output: Audio guide data

[0933] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[0934] Input: Audio guide data

[0935] Output: Text display and voice guide

[0936] Step 7: Recognizing User Emotions

[0937] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone, for example, to check the user's stress level in real time.

[0938] Input: User's facial expression and voice data

[0939] Output: Parsed emotion data

[0940] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[0941] Input: Parsed emotion data

[0942] Output: Emotion data sent to the server

[0943] Server: Adjust the tone and content of the audio guide based on the emotional data. For example, if the user is feeling stressed, change the tone of the guide to a gentler tone.

[0944] Input: Emotion data

[0945] Output: Adjusted audio description data

[0946] Step 8: Gather and incorporate feedback

[0947] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user enters the difficulty of cooking and their impressions of the taste.

[0948] Input: User feedback

[0949] Output: Feedback data is saved in the app

[0950] Terminal: Sends feedback data to the server. Collects user ratings and comments digitally.

[0951] Input: User feedback data

[0952] Output: Feedback data sent to the server

[0953] Server: Analyzes the feedback data and updates the database to improve the cooking procedures and recipes for the next time. For example, improve a process that received poor taste ratings.

[0954] Input: Feedback data

[0955] Output: Updated recipe database

[0956] (Application example 2)

[0957] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0958] Traditionally, it has been difficult to accurately pass on a mother's cooking recipes to the next generation, and there is a need for a way to preserve those flavors for future generations, especially in an aging society. Furthermore, existing digital recipe systems lack guidance that takes into consideration the user's emotions, making it difficult to reduce stress and improve satisfaction while cooking. Furthermore, when considering use in brick-and-mortar stores, a system that can handle on-site challenges is required.

[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0960] In this invention, the server includes means for digitizing mother's cooking manual, means for analyzing the digitized recipe data and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in a database, means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in the mother's voice, and means for recognizing the user's emotions using an emotion engine and adjusting the tone of the audio guidance. This makes it possible to accurately convey the taste of mother's cooking to the next generation and provide guidance that takes user emotions into consideration, making it possible to provide high satisfaction even when cooking in a physical store.

[0961] A "cooking manual" is a document and video data that contains recipes and procedures for the dishes that a mother would make on a daily basis.

[0962] "Digitalization" means converting information in analog form into digital data.

[0963] "Recipe data" is information that describes the ingredients, quantities, and cooking steps for making a dish.

[0964] A "database" is a digital information storage area that systematically organizes multiple data sets and makes them easy to manage and search.

[0965] "Speech recognition" is a technology that converts human speech into data in a format that a computer can understand.

[0966] "Voice cloning technology" is a technology that learns the characteristics of a specific person's voice and can read any text in that person's voice.

[0967] "Audio guide" is a function that uses voice to give instructions and explanations to users.

[0968] An "emotion engine" is a system that uses sensors such as cameras and microphones to analyze and recognize the user's emotional state and then respond appropriately based on that.

[0969] The system for carrying out this invention digitizes a mother's cooking manual and provides audio guidance of cooking procedures in a manner that takes into consideration the user's feelings. The following is a specific embodiment of this system.

[0970] The system begins with the user taking a photo of their mother's cooking manual. Using a smartphone or smart glasses, the user takes photos of the notes in the cooking manual and their mother cooking. The photos and videos are then uploaded to a server using a dedicated application. The server receives the uploaded data and temporarily stores it in a database.

[0971] The server then uses OCR technology to extract text from photos and videos, converting the recipe data into text format, which is then categorized and organized into recipe ingredients, quantities, and cooking steps.

[0972] When the user provides a voice description of their mother's cooking steps, the voice data recorded on their smartphone is also uploaded to the server using a dedicated application. The server receives the voice data and stores it in a database. Furthermore, it uses voice recognition technology to extract text data from the voice data, and uses voice cloning technology to learn the mother's voice and generate a voice model.

[0973] The server inputs the text data acquired by OCR into an AI model and analyzes information such as ingredients, quantities, cooking procedures, etc. The analyzed data is then structured and stored in a database.

[0974] When a user selects a recipe they want to cook within the application, the request is sent to the server. Based on the user's input of allergies and taste preferences (e.g., "low salt" or "medium spicy"), the server adjusts the recipe and generates appropriate instructions. The adjusted recipe and cooking steps are then audio-guided using the mother's voice.

[0975] While the user is cooking, the device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion engine, and sends the emotion data to the server. The server then adjusts the tone and content of the audio guidance during cooking based on the emotion data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[0976] After cooking, the user enters feedback about the taste and cooking procedure into the app. The device sends the feedback data to the server, which analyzes it and updates the database to improve the cooking procedure and recipe for the next time.

[0977] As a specific example, let's assume that the user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes the mother's video of the curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the low-salt specifications selected by the user, a recipe is generated with the salt amount adjusted. Based on this customized information, the mother's voice provides audio guidance such as, "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as, "It's okay to take it slowly."

[0978] Here are some example prompts using a generative AI model:

[0979] "Digitize your mom's curry recipe, customize it to low salt and medium spice, and provide audio guidance in your mom's voice. If the user is stressed, make the guidance tone gentler."

[0980] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0981] Step 1:

[0982] The user uses a smartphone or smart glasses to take pictures of their mother's cooking manual notes or the mother cooking. The input data is photos or videos, which are uploaded by the application and sent to the server. The server receives the uploaded images and videos and temporarily stores them in a database. This saves the physical recipe information as digital data.

[0983] Step 2:

[0984] The server uses OCR technology to extract text information from the received images and videos. The input is the image or video saved in step 1, and the output is the extracted text data. Specifically, the server uses the pytesseract library to perform character recognition and convert it into text recipe data.

[0985] Step 3:

[0986] The server categorizes and organizes the text data extracted by OCR into categories such as ingredients, quantities, and cooking procedures. The input is the text data obtained from OCR, and the output is structured recipe data. The text is analyzed using an AI model and stored in a database. This allows for efficient management of recipe information.

[0987] Step 4:

[0988] The user records their mother's cooking instructions on their smartphone. The data input is voice data, which is uploaded using a dedicated application and sent to the server. The server receives the voice data and stores it in a database.

[0989] Step 5:

[0990] The server uses speech recognition technology to extract text data from the uploaded audio data. The input is the audio data saved in step 4, and the output is the extracted text data. It then uses voice cloning technology to learn the mother's voice and generate a voice model. This makes it possible to read any text aloud in the mother's voice.

[0991] Step 6:

[0992] The user selects the recipe they want to cook within the dedicated application. The input is the user's recipe selection information, which is sent from the terminal to the server. Based on this information, the server retrieves the corresponding recipe data from the database.

[0993] Step 7:

[0994] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application. The input is customization information, which is sent from the device to the server. The server adjusts the recipe based on the customization information and generates new cooking instructions.

[0995] Step 8:

[0996] The server uses the trained voice model to generate audio guidance based on the adjusted recipe. The input is the adjusted recipe data, and the output is audio guidance. The audio guidance instructs the user on the cooking steps in a mother's voice.

[0997] Step 9:

[0998] While cooking, the device (smart glasses or head-mounted display) uses a camera and microphone to analyze the user's facial expressions and voice with an emotion engine. The input is the user's facial expressions and voice data, and the emotion data is sent from the device to the server.

[0999] Step 10:

[1000] The server adjusts the tone and content of the audio guide based on the emotion data. The input is emotion data, and the output is the adjusted audio guide. For example, if the user is feeling stressed, the tone of the guide can be made gentler and words of encouragement can be added.

[1001] Step 11:

[1002] After cooking is complete, the user enters feedback about the taste and cooking procedure in the app. The input is feedback data, which is sent from the device to the server. The server analyzes the feedback data and updates the database to improve the cooking procedure and recipe for the next time.

[1003] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1004] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1005] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1006] [Third embodiment]

[1007] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1008] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1009] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1010] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1011] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1012] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1013] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1014] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1015] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1016] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1017] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1018] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1019] This invention is a digital system for accurately passing on mother's cooking recipes to the next generation, aiming to deepen special family memories and bonds while recreating the flavor of mother's cooking. The program of this system includes functions for data collection, analysis, audio guidance, customization, and feedback.

[1020] 1. Data Collection

[1021] User: First, the user takes a note of his mother's cooking recipes and a video of his mother cooking on his smartphone.

[1022] Device: Upload the photos and videos you have taken to the server via the application.

[1023] Server: Receives the uploaded data and converts the recipe into text using OCR technology.

[1024] 2. Recipe analysis and database storage

[1025] Server: Recipe data converted to text using OCR is fed into the AI ​​model, and information such as ingredients, quantities, and cooking steps is extracted.

[1026] Server: Structures the extracted data and stores it in a database.

[1027] 3. Collecting and Learning Audio Data

[1028] User: Records mother explaining cooking steps on smartphone.

[1029] Device: Upload the recording data from the application to the server.

[1030] Server: The uploaded voice data is analyzed through a voice recognition process and converted into text. At the same time, voice cloning technology is used to learn the mother's voice and generate a voice model.

[1031] 4. Providing cooking guides

[1032] User: Selects the recipe they want to cook within the application.

[1033] Terminal: Sends a request for the selected recipe to the server.

[1034] Server: Retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[1035] Device: The acquired cooking instructions are displayed to the user, and instructions such as "First, cut the chicken into bite-sized pieces" are also provided in a mother's voice.

[1036] 5. User Customization

[1037] User: Enter allergies and individual taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1038] Terminal: Sends the entered customization information to the server.

[1039] Server: Adjusts the recipe based on the customization information and generates the appropriate instructions.

[1040] Device: Displays the customized recipe and provides audio guidance if needed.

[1041] Specific examples

[1042] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is given audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also making dishes that accommodate their own dietary restrictions.

[1043] In addition, after cooking, users can enter feedback into the app, which is analyzed by the server and the recipe database is constantly updated to provide more accurate guidance the next time the user cooks.

[1044] By combining the above functions, the system of the present invention can accurately pass on the taste of mother's cooking to the next generation, further deepening family ties.

[1045] The processing flow will be explained below.

[1046] Step 1: Data collection

[1047] User: Takes notes of his mother's cooking recipes and videos of himself cooking on his smartphone.

[1048] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[1049] Server: Receives the uploaded data and temporarily stores it in a database.

[1050] Step 2: Digitize your recipes

[1051] Server: Using OCR technology, extracts text information from photos and videos and converts recipes into text format.

[1052] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[1053] Step 3: Collecting audio data

[1054] User: Records her mother explaining the recipe steps on her smartphone.

[1055] Terminal: Upload the recorded audio data to the server using a dedicated application.

[1056] Server: Receives the uploaded audio data and stores it in a database.

[1057] Step 4: Analyze and train audio data

[1058] Server: Extracts text data from the voice data using voice recognition technology.

[1059] Server: Organizes the text data into cooking instructions and stores them in a database.

[1060] Server: Using voice cloning technology, learns the characteristics of the mother's voice and generates a voice model.

[1061] Step 5: Parse and save the recipe

[1062] Server: The text data obtained by OCR is fed into the AI ​​model, which analyzes ingredients, quantities, and cooking procedures.

[1063] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1064] Step 6: User selects recipe

[1065] User: Select the recipe they want to cook within the dedicated application.

[1066] Terminal: Sends a request for the selected recipe to the server.

[1067] Step 7: Cooking instructions generation and audio guidance

[1068] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1069] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1070] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1071] Step 8: Enter and update customization information

[1072] User: Enters specific dietary restrictions and taste preferences (e.g., "low salt," "medium spicy," etc.) into a dedicated application.

[1073] Terminal: Sends the entered customization information to the server.

[1074] Server: Adjusts the recipe based on the user's customization information.

[1075] Device: Provides the adjusted recipe to the user, and also provides audio guidance if necessary.

[1076] Step 9: Gather and incorporate feedback

[1077] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[1078] Terminal: Sends the input feedback data to the server.

[1079] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1080] Example 1

[1081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1082] It is not easy to accurately pass on the flavors and steps of a mother's cooking to the next generation. If recipes are not accurately digitized, there is a high risk of losing special family memories and bonds. It is also difficult to satisfy the need to receive detailed cooking instructions in the mother's voice. Furthermore, it is difficult to accommodate the different dietary restrictions and taste preferences of each family. To solve these problems, accurately pass on the flavors of a mother's cooking to the next generation, and deepen family bonds, an efficient and accurate digital system is needed.

[1083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1084] In this invention, the server includes means for digitizing mother's recipes, means for analyzing the digitized recipes and extracting ingredients, portions, and cooking methods, means for storing the extracted data in a data repository, means for collecting mother's voice data and learning mother's voice using voice recognition and voice synthesis technology, and means for providing audio guidance on cooking methods to users in mother's voice, thereby enabling mother's recipes to be accurately passed down to future generations and providing cooking guides customized to the specific needs of each household.

[1085] "My mother's cooking methods" refers to the cooking methods, steps, and recipes that my mother uses.

[1086] "Digitalization" refers to the process of converting analog information into electronic data.

[1087] "Ingredients" refers to the various foods and ingredients used in cooking.

[1088] "Amount" refers to the amount of each ingredient used when cooking.

[1089] "Cooking method" refers to a series of processes or steps to complete a dish using ingredients.

[1090] A "data repository" refers to a database or storage system for efficiently storing and managing collected data.

[1091] "Audio data" refers to digital data that contains recorded audio information such as human voices and spoken words.

[1092] "Speech recognition technology" refers to technology that generates text data from voice data.

[1093] "Speech synthesis technology" refers to technology that generates new voices based on certain voice samples.

[1094] "Voice guidance" refers to a method of conveying specific procedures or information to a user by providing voice guidance.

[1095] This invention is a digital system that aims to pass on mother's cooking techniques to the next generation and deepen family memories and bonds. The system has functions for data collection, analysis, audio guidance, customization, and feedback, and is designed to make it easier for users to recreate the taste of their mother's cooking.

[1096] Data collection

[1097] The first thing a user does is digitize their mother's cooking methods. They take photos of their mother's recipe notebooks and videos of their mother cooking with their smartphone. They then upload these photos and videos to a server using a dedicated application.

[1098] The server receives the uploaded photos and videos and converts the recipes into text data using OCR (Optical Character Recognition) technology, a process that converts analog information into electronic data.

[1099] Recipe analysis and database storage

[1100] The server then inputs the recipe data, converted to text using OCR technology, into an AI model to extract information such as ingredients, quantities, cooking methods, etc. Specifically, it uses natural language processing technology to analyze the recipe text and automatically identify the necessary information.

[1101] The extracted data is structured by the server and stored in a data repository, making it easy for users to search and browse later.

[1102] Audio data collection and learning

[1103] The user records their mother explaining the cooking steps on their smartphone, and the recording is then uploaded to the server via the application.

[1104] The server analyzes the uploaded voice data and converts it into text using speech recognition technology, while simultaneously learning the mother's voice using speech synthesis technology to generate a voice model.

[1105] Providing cooking guides

[1106] When cooking, the user selects the desired recipe within the application. The selected recipe request is sent from the device to the server. The server retrieves the corresponding recipe from the recipe database and generates an audio guide in the mother's voice along with cooking instructions.

[1107] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, allowing the user to confirm the cooking instructions both visually and audibly.

[1108] User Customization

[1109] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application, which then sends the information from the device to the server, which then adjusts the recipe accordingly.

[1110] The adjusted recipe is then sent back to the device and displayed to the user, with audio guidance provided if needed, allowing the user to create a dish tailored to their own preferences.

[1111] Specific examples

[1112] For example, consider a case where a user selects "curry" and prefers it "low salt" and "medium spicy." The system digitizes the mother's curry recipe using OCR technology and AI analysis. It then generates a recipe with adjusted salt content based on the user's customization information. Finally, the mother's voice provides audio guidance, saying, "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also creating a dish that suits their own preferences.

[1113] Examples of prompt statements

[1114] "Please digitize my mother's homemade curry recipe using OCR technology and AI analysis. Then please guide me through the recipe, which is low in salt and medium in spiciness, using my mother's voice."

[1115] This system makes it possible to accurately pass on the flavor of a mother's cooking to the next generation, further strengthening family bonds.

[1116] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1117] Step 1: Data collection

[1118] User: Takes a photo of his mother's recipe notebook or a video of his mother cooking on his smartphone. The input is the image or video of the recipe.

[1119] Specific operation: The user launches the application and uses the camera function to take notes or record videos.

[1120] Terminal: Uploads captured photos and videos to the server via the application. The output is the transmission of image and video data to the server.

[1121] Specific operation: When you press the "Upload" button within the application, the device will transfer the captured data to the server.

[1122] Step 2: Convert the recipe to text

[1123] Server: Receives uploaded photos and videos and converts them into text data using OCR technology. The input is image or video data, and the output is text data.

[1124] Specific operation: The server uses OCR (Optical Character Recognition) software to convert the received image data into text.

[1125] Step 3: Recipe analysis and database storage

[1126] Server: Text-based recipe data is fed into the AI ​​model to extract information such as ingredients, quantities, cooking steps, etc. The input is text data, and the output is the extracted structured data.

[1127] How it works: The server inputs text data into the AI ​​model and executes a script that automatically analyzes and extracts ingredients and procedures.

[1128] Server: The extracted data is structured and stored in a data repository. The output is structured data stored in a database.

[1129] What it does: The server stores structured data in a relational database, allowing for efficient searching and management.

[1130] Step 4: Collecting and training audio data

[1131] User: Records the voice of the mother explaining the cooking procedure on a smartphone. The input is the mother's voice.

[1132] Specific actions: Use the application's voice recording function to record the mother's explanation.

[1133] Terminal: Uploads recorded data from the application to the server. The output is sending audio data to the server.

[1134] Specific operation: After recording is finished, press the "Upload" button and the audio data will be transferred to the server.

[1135] Server: Converts voice data into text using speech recognition technology, and learns the mother's voice using speech synthesis technology. The input is voice data, and the output is a voice model.

[1136] How it works: The server uses speech recognition software to convert speech to text, and then trains speech synthesis software to learn the mother's voice.

[1137] Step 5: Provide cooking guidance

[1138] User: Selects the recipe they want to cook within the application. Input is a recipe selection request.

[1139] Specific operation: Open "Recipe List" from the application menu and tap to select the desired recipe.

[1140] Terminal: Sends a request for the selected recipe to the server. The output is the request to the server.

[1141] Specific operation: When a user selects a recipe, a request is sent to the server.

[1142] Server: Retrieves the relevant recipe from the database and generates an audio guide in a mother's voice along with cooking instructions. The input is a recipe request, and the output is the audio guide data.

[1143] Specific operation: The server retrieves the relevant recipe from the database and generates a guide voice using the voice model.

[1144] Terminal: Displays cooking instructions to the user and provides audio guidance. Output is display data and audio guidance to the user.

[1145] Specific operation: The cooking steps are presented to the user visually and audibly. For example, the device may say, "First, cut the chicken into bite-sized pieces."

[1146] Step 6: Customizing Users

[1147] User: Enter allergies and taste preferences into the app. Input is customization information.

[1148] Specific operation: Access the form for entering customization information from the application's "Settings" screen and enter the required information.

[1149] Terminal: Sends input customization information to the server. Output: Sends customization information to the server.

[1150] Specific operation: When you press the "Save" button, the input information is sent to the server.

[1151] Server: Adjusts the recipe based on the customization information and generates appropriate instructions. The input is the customization information and the output is the adjusted recipe.

[1152] Specific operation: The server adjusts the amounts of ingredients and seasonings in the recipe based on the customization information.

[1153] Terminal: Displays the customized recipe and provides voice guidance if necessary. The output is the display data and voice guidance for the user.

[1154] What it does: It shows the user the adjusted recipe and provides audio guidance such as "Add a little salt to the chicken."

[1155] (Application example 1)

[1156] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1157] Traditionally, it has been difficult to accurately pass on family recipes to the next generation, especially when it comes to recreating the flavor of a mother's cooking. Enjoying these dishes through delivery services is also uncommon, making it even more challenging to accommodate individual taste preferences and dietary restrictions. This limits opportunities to deepen family memories and bonds. Continuously improving recipes based on feedback is also a challenge.

[1158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1159] In this invention, the server includes: means for digitizing mother's cooking recipes; means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps; means for storing the extracted data in a database; means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology; means for providing audio guidance to the user on cooking steps in mother's voice; means for cooking dishes based on the recipes and delivering them to the user via a delivery service; and means for collecting user feedback data and updating the recipe database. This allows mother's special recipes to be faithfully passed on to the next generation and for the user to enjoy the dishes through delivery. Furthermore, recipes can be continuously improved based on user feedback, thereby providing a consistently high-quality cooking experience.

[1160] The "recipe digitization method" is a method for collecting mother's recipes as photos and videos and converting them into digital format.

[1161] The "recipe analysis means" is a means for extracting ingredients, quantities, and cooking steps from digitized recipe data.

[1162] The "database storage means" is a means for structuring the analyzed recipe data and storing it in a database.

[1163] The "audio collection means" is a means for collecting and recording the audio of the mother explaining the cooking steps.

[1164] A "speech recognition means" is a means that uses technology to convert collected voice data into text.

[1165] The "voice cloning technology learning means" is a means for learning the mother's voice and reproducing the mother's voice based on the voice data.

[1166] The "audio guidance means" is a means for providing the user with audio guidance on cooking procedures in a mother's voice.

[1167] A "cooking tool" is a tool for cooking a dish based on a recipe.

[1168] "Delivery service means" refers to a means for delivering cooked food to a user through a delivery service.

[1169] A "feedback collection means" is a means for collecting feedback data from users.

[1170] The "recipe database update means" is a means for updating the recipe database based on the collected feedback data.

[1171] A "generative AI model" refers to an artificial intelligence model that extracts specific information from text data and performs analysis.

[1172] A "prompt sentence" is an input sentence for a generative AI model, and is text that contains instructions for analysis.

[1173] This invention is a system that digitizes mother's cooking recipes, accurately passes them on to the next generation, and provides customized meals through a delivery service. The system aims to allow users to recreate their mother's special dishes and deepen family ties.

[1174] Hardware and software used

[1175] Uses hardware and software that integrates servers, smartphones, and delivery services, including:

[1176] Smartphone: Used as data collection and operation interface.

[1177] Server: A central processing unit that performs data analysis, voice recognition, voice cloning, recipe generation, and delivery management.

[1178] Google Cloud Vision API: Uses OCR technology to convert digitized recipes into text.

[1179] OpenAI GPT-4: A generative AI model that analyzes recipe data and extracts ingredients and cooking steps.

[1180] Google Speech-to-Text API: Converts audio data into text.

[1181] Lyrebird or Descript: A voice cloning technology that recreates the mother's voice.

[1182] Uber Eats API: Manages delivery services to deliver prepared meals to users.

[1183] Data processing and operation procedures

[1184] 1. Data Collection

[1185] The user uses their smartphone to record their mother's recipe notes and videos of their mother cooking.

[1186] Photos and videos taken by the device are uploaded to the server via the application.

[1187] 2. Recipe analysis and database storage

[1188] The server receives the uploaded data and converts it into text using OCR technology using the Google Cloud Vision API.

[1189] The server inputs the recipe data, converted into text using OCR, into OpenAI GPT-4 to extract information such as ingredients, quantities, and cooking steps.

[1190] The server structures the extracted data and stores it in a database.

[1191] 3. Collecting and Learning Speech Data

[1192] The user records the audio of their mother explaining the cooking steps on their smartphone.

[1193] The device uploads the recording data from the application to the server.

[1194] The server analyzes the uploaded voice data using the Google Speech-to-Text API and converts it into text, while simultaneously training the mother's voice using Lyrebird or Descript to generate a voice model.

[1195] 4. Providing cooking guides

[1196] The user selects the recipe they want to cook within the application.

[1197] The device sends a request for the selected recipe to the server.

[1198] The server retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[1199] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, such as "First, cut the chicken into bite-sized pieces."

[1200] 5. Food preparation and delivery

[1201] The server arranges for the food to be cooked based on the generated recipe and cooking instructions.

[1202] Once cooked, the food is delivered to the user using the Uber Eats API.

[1203] 6. User Feedback

[1204] The user enters post-cooking feedback into the application.

[1205] The server analyzes the feedback data and updates the recipe database to improve the accuracy of future cooking guides.

[1206] Specific examples

[1207] For example, if a user selects "Mom's Special Curry" and requests "low salt" and "medium spicy," the system analyzes a photo of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. A recipe with adjusted salt content is generated based on the user's low-salt specifications. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." After cooking, the food is delivered to the user via Uber Eats.

[1208] Prompt Sentence Examples

[1209] Analyze photos, videos, and audio data of Mom's recipes uploaded by users to extract ingredients and cooking steps, clone Mom's voice from the audio data, and generate an audio guide that provides cooking instructions in Mom's voice along with the user's desired customization options. Once cooking is complete, deliver the food using the Uber Eats API.

[1210] As a result, the present invention can provide a high-quality cooking experience by faithfully passing on mother's special recipes to the next generation and customizing them according to the user's wishes.

[1211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1212] Step 1:

[1213] Data collection

[1214] The user takes photos of their mother's cooking recipe notes or videos of their mother cooking with their smartphone. The device then uploads the photos and videos to a server via an application. The input is the photos and videos, and the output is the data sent to the server.

[1215] Step 2:

[1216] Text conversion using OCR

[1217] The server receives the uploaded photos and videos and converts them into text using OCR technology using the Google Cloud Vision API. The input is the photo or video data, and the output is the recipe data in text format. This extracts text information from the photos and videos.

[1218] Step 3:

[1219] Recipe analysis and database storage

[1220] The server inputs the recipe data converted to text using OCR into OpenAI GPT-4, and uses a generative AI model to extract information such as ingredients, quantities, and cooking steps. The input is the converted recipe data, and the output is the extracted structured data. The server then structures the extracted data and stores it in a database. This allows the recipe's digital information to be organized and stored.

[1221] Step 4:

[1222] Audio data collection

[1223] The user records the voice of his mother explaining the cooking procedure on his smartphone. The device uploads the recorded data to the server via the application. The input is the recorded voice data, and the output is the voice data sent to the server.

[1224] Step 5:

[1225] Speech Recognition and Speech Clone Training

[1226] The server analyzes the uploaded audio data using the Google Speech-to-Text API and converts the audio to text. At the same time, it uses Lyrebird or Descript to train the mother's voice and generate a voice model. The input is the audio data, and the output is the converted audio data and a voice model of the mother's voice.

[1227] Step 6:

[1228] Generate cooking guide

[1229] The user selects the recipe they want to cook within the application. The device sends a request for the selected recipe to the server. The server retrieves the recipe from the recipe database and generates audio instructions for the cooking process in a mother's voice. The input is the selected recipe request, and the output is the cooking process with audio instructions.

[1230] Step 7:

[1231] Food preparation and delivery

[1232] The server arranges for the food to be cooked based on the generated recipe and cooking instructions. The cooked food is delivered to the user using the Uber Eats API. The input is the cooking instructions and delivery address information, and the output is the delivered food.

[1233] Step 8:

[1234] Feedback collection and database updates

[1235] Users input their post-cooking feedback into the application. The server analyzes the feedback data and updates the recipe database. The input is the feedback data, and the output is an updated recipe database. This improves the accuracy of cooking guidance from the next time onwards.

[1236] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1237] This invention is a system that digitizes mother's cooking recipes to pass them on to the next generation, and combines them with an emotion engine that recognizes the user's emotions to provide greater emotional connection and satisfaction. The program of this system includes functions for data collection, analysis, voice guidance, customization, feedback, and emotion recognition.

[1238] 1. Data Collection

[1239] User: First, the user takes a photo of the notebook containing his mother's cooking recipes and a video of his mother cooking with his smartphone.

[1240] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[1241] Server: Receives the uploaded data and temporarily stores it in a database.

[1242] 2. Digitizing recipes

[1243] Server: Uses OCR technology to extract text information from photos and videos and convert recipes into text format.

[1244] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[1245] 3. Collecting and Learning Audio Data

[1246] User: Records mother explaining cooking steps on smartphone.

[1247] Terminal: Upload the recorded audio data to the server using a dedicated application.

[1248] Server: Receives the uploaded audio data and stores it in a database.

[1249] Server: Extracts text data from the voice data using voice recognition technology, learns the mother's voice using voice cloning technology, and generates a voice model.

[1250] 4. Parse and save the recipe

[1251] Server: The text data obtained by OCR is input into an AI model to analyze information such as ingredients, quantities, and cooking procedures.

[1252] Server: The analyzed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1253] 5. User-selected and customized recipes

[1254] User: Select the recipe they want to cook within the dedicated application.

[1255] Terminal: Sends a request for the selected recipe to the server.

[1256] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1257] Terminal: Sends the entered customization information to the server.

[1258] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1259] 6. Providing cooking guides

[1260] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1261] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1262] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1263] 7. User Emotion Recognition

[1264] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone.

[1265] Terminal: The emotion engine sends the user's emotion data to the server.

[1266] Server: Adjusts the tone and content of the audio guidance during cooking based on emotional data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[1267] Server: It can analyze the user's emotional data and suggest recipes and cooking procedures that suit their emotions.

[1268] 8. Gather and incorporate feedback

[1269] User: After cooking is complete, enter feedback into the app about the taste and process of the dish.

[1270] Terminal: Sends the input feedback data to the server.

[1271] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1272] Specific examples

[1273] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as "It's okay to take it slowly."

[1274] By combining the above functions, the system of the present invention can accurately convey the taste of mother's cooking to the next generation, further deepening emotional connections and satisfaction.

[1275] The processing flow will be explained below.

[1276] Step 1: Data collection

[1277] User: Records his mother's cooking recipes in a notebook and takes photos of his mother cooking with his smartphone.

[1278] Terminal: Uploads captured photos and videos to the server using a dedicated application.

[1279] Server: Receives the uploaded data and temporarily stores it in a database.

[1280] Step 2: Digitize your recipes

[1281] Server: Extracts text information from uploaded images and videos using OCR technology.

[1282] Server: Analyzes the extracted text data and classifies and organizes it into ingredients, quantities, and cooking steps in the recipe.

[1283] Step 3: Collecting and training audio data

[1284] User: Records mother explaining cooking steps on smartphone.

[1285] Terminal: Uploads recorded audio data to the server via a dedicated application.

[1286] Server: Receives the uploaded audio data and stores it in a database.

[1287] Server: Uses voice recognition technology to convert voice data into text data.

[1288] Server: Using voice cloning technology, the characteristics of the mother's voice are learned and a voice model is generated.

[1289] Step 4: Parse and save the recipe

[1290] Server: The text data obtained by OCR is input into an AI model to analyze ingredients, quantities, and cooking procedures.

[1291] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1292] Step 5: User Recipe Selection and Customization

[1293] User: Select the recipe they want to cook within the dedicated application.

[1294] Terminal: Sends a request for the selected recipe to the server.

[1295] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1296] Terminal: Sends the entered customization information to the server.

[1297] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1298] Step 6: Provide cooking guidance

[1299] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1300] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1301] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1302] Step 7: Recognizing User Emotions

[1303] Device: The emotion engine analyzes the user's facial expressions and voice using the camera and microphone.

[1304] Terminal: Transmits the analyzed emotion data to the server.

[1305] Server: Adjusts the tone and content of the audio guide during cooking based on emotional data. For example, if it detects that the user is feeling stressed, it will soften the tone of the audio guide and insert encouraging words such as "It's okay to take it slowly."

[1306] Step 8: Suggest an emotionally appropriate recipe

[1307] Server: Analyzes the user's emotional data and suggests recipes and cooking procedures that suit their emotional state. For example, if the user is having fun, it can suggest more challenging recipes.

[1308] Terminal: Notifies users of emotion-based recipe suggestions.

[1309] Step 9: Gather and incorporate feedback

[1310] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[1311] Terminal: Sends feedback data to the server.

[1312] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1313] Example 2

[1314] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1315] Conventional systems offer limited ways to pass on mother's cooking recipes to the next generation and lack mechanisms for providing emotional connection and satisfaction. They also struggle to customize to accommodate users' dietary restrictions and individual preferences, and are unable to recognize users' emotions and provide appropriate assistance while cooking. This increases the likelihood of users feeling stressed or creating unsatisfying dishes.

[1316] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1317] In this invention, the server includes means for digitizing mother's cooking recipes, means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in an information management system, means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in mother's voice, and means for recognizing the user's emotions and adjusting the tone and content of the audio guidance based on those emotions. This makes it possible to accurately pass on mother's cooking recipes to the next generation and provide cooking support that takes the user's emotions into consideration, providing a high level of satisfaction.

[1318] "Digitization" is the process of converting analog information into digital form.

[1319] "Analysis" is the process of examining data in detail to understand its structure and meaning.

[1320] "Ingredients" refers to the food ingredients used to make a dish.

[1321] "Quantity" refers to the specific amount of an ingredient used in a dish.

[1322] A "cooking procedure" is a sequence of steps for cooking a dish.

[1323] An "information management system" is a system for storing, managing, retrieving, and using data.

[1324] "Audio data" refers to digital information that records sound.

[1325] "Speech recognition" is a technology that converts voice data into text data.

[1326] "Voice cloning technology" is a technology for reproducing a specific voice.

[1327] "Audio guide" is a means of giving instructions or explanations using voice.

[1328] "User" refers to a person who uses this system.

[1329] "Customization" refers to modifying a system or service to meet a user's specific needs or requirements.

[1330] "Emotion recognition" is a technology that analyzes changes in a user's facial expressions and voice to determine their emotions.

[1331] "Feedback Data" refers to the evaluations and opinions provided by users after using the system.

[1332] This invention is a system that digitizes mother's cooking recipes, passes them on to the next generation, and recognizes the user's emotions. This system includes the following functions: data collection, digitization, voice data collection and learning, recipe analysis and storage, user recipe selection and customization, cooking guide provision, user emotion recognition, and feedback collection and reflection.

[1333] Data collection

[1334] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. For example, he takes photos of each page of the notebook and records a video of his mother cooking in the kitchen.

[1335] Device: Photos and videos are uploaded to the server using a dedicated application. The device checks the data format and network connection status to upload the data appropriately.

[1336] Server: Receives the uploaded data and temporarily stores it in the database. At this time, it checks the format and size of the data and converts it into the appropriate format.

[1337] Recipe digitization

[1338] Server: Using OCR technology such as Google Cloud Vision API, text information is extracted from photos and videos and converted into text format.

[1339] Server: Analyzes text data and classifies and organizes it into recipe ingredients, quantities, and steps. For example, converts "200g of chicken" and "2 potatoes" in an image into text data.

[1340] Audio data collection and learning

[1341] User: Record your mother explaining cooking steps on your smartphone. For example, record your mother explaining, "Cut the chicken into bite-sized pieces."

[1342] Terminal: The recorded audio data is uploaded to the server using a dedicated application. At this time, the format of the audio data is checked and converted to the appropriate format.

[1343] Server: Receives the uploaded voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[1344] Recipe analysis and saving

[1345] Server: The text data obtained by OCR is input into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking procedures.

[1346] Server: The analyzed data is structured and stored in a database such as MongoDB. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1347] User-selected and customized recipes

[1348] User: Selects the recipe they want to cook within the dedicated application, for example, curry or stew.

[1349] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[1350] Terminal: Sends the input customization information to the server. The server adjusts the recipe based on the conditions entered by the user.

[1351] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1352] Providing cooking guides

[1353] Server: Retrieves the selected recipe from the database and generates cooking instructions. Specific steps are clearly displayed and presented to the user in an easy-to-understand manner.

[1354] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1355] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[1356] User Emotion Recognition

[1357] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone. For example, it monitors the user's face in real time and analyzes their emotions.

[1358] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[1359] Server: Adjusts the tone and content of the audio guide based on emotional data. For example, if the user is feeling stressed, the tone of the guide can be changed to a gentler tone. It can also suggest recipes and cooking procedures that are suitable for the user.

[1360] Gathering and implementing feedback

[1361] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user comments on the strength of the flavor and the ease of understanding the process.

[1362] Terminal: Sends feedback data to the server. Collects user evaluations as digital data.

[1363] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time, for example, adjusting the taste based on the feedback.

[1364] Specific examples

[1365] For example, if a user selects "curry" and prefers low salt and medium spiciness, the system will convert their mother's curry recipe into text using OCR and AI technology. Based on the user's customized information, a recipe with adjusted salt content will be generated. A voice guide will be provided in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." If the user becomes stressed, the emotion engine will detect this and change the tone of the voice guide to a gentler tone, encouraging the user by saying, "It's okay to take it easy."

[1366] Example prompts for generative AI models

[1367] "Digitalize your mother's recipes and design a cooking support system that combines voice guidance and emotion recognition."

[1368] "Explain OCR and AI techniques used to analyze photos of cooking recipes and extract ingredients and steps."

[1369] The system of the present invention not only passes on mother's cooking recipes to the next generation, but also provides cooking support that takes the user's emotions into consideration, thereby providing greater satisfaction and emotional connection.

[1370] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1371] Step 1: Data collection

[1372] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. Specifically, he carefully photographs each page of the notebook and records a video of his mother cooking in the kitchen.

[1373] Input: Photos and videos taken

[1374] Output: Captured data is saved on your smartphone (image and video files saved in local storage)

[1375] Device: Uploads photos and videos to the server using a dedicated application. During upload, the network connection status and data integrity are checked.

[1376] Input: Image and video data stored on a smartphone

[1377] Output: Data upload process to the server is completed, and data saved on the server

[1378] Step 2: Digitize your recipes

[1379] Server: Using OCR technology such as Google Cloud Vision API, character information is extracted from photos and videos and converted into text. For example, handwritten or printed characters in an image are converted into text.

[1380] Input: Uploaded photo and video data

[1381] Output: Text data (e.g., "200g chicken, 2 potatoes, cut the chicken into bite-sized pieces")

[1382] Server: Analyzes text data and classifies it into ingredients, quantities, and cooking procedures. Using natural language processing technology, the text is analyzed by phrase and classified into categories.

[1383] Input: Extracted text data

[1384] Output: Structured data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Steps: Cut the chicken into bite-sized pieces")

[1385] Server: Stores organized data in an information management system.

[1386] Input: Structured data

[1387] Output: Recipe data stored in the database

[1388] Step 3: Collecting and training audio data

[1389] User: Record your mother explaining the cooking steps on your smartphone. Specifically, record your mother explaining, "Cut the chicken into bite-sized pieces."

[1390] Input: Recorded audio data

[1391] Output: Audio file saved on your smartphone (e.g. mp3 or wav format)

[1392] Device: The recorded audio data is uploaded to the server using a dedicated application. During the upload, the format and integrity of the audio file are checked.

[1393] Input: Audio file

[1394] Output: Audio data sent to the server

[1395] Server: Receives the voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[1396] Input: Uploaded audio data

[1397] Output: Text data and speech model

[1398] Step 4: Parse and save the recipe

[1399] Server: The text data acquired by OCR is fed into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking steps. The AI ​​model analyzes the context and meaning, clearly distinguishing between ingredients and steps.

[1400] Input: Text data obtained by OCR

[1401] Output: Analyzed data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Procedure: Cut the chicken into bite-sized pieces")

[1402] Server: The parsed data is structured and stored in a database such as MongoDB. The structured data is saved in an easy-to-search format.

[1403] Input: Parsed data

[1404] Output: Recipe information stored in the database

[1405] Step 5: User Recipe Selection and Customization

[1406] User: Selects the recipe they want to cook within the dedicated application. For example, they can select a dish such as curry or stew.

[1407] Input: Recipe selection by user

[1408] Output: The selected recipe information is displayed in the application.

[1409] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[1410] Input: User preferences and conditions

[1411] Output: Customization information sent to the server

[1412] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions. Adjusts the recipe based on the conditions specified by the user.

[1413] Input: Customization information

[1414] Output: Customized recipe information

[1415] Step 6: Provide cooking guidance

[1416] Server: Retrieves the selected recipe from the database and generates cooking instructions. It clearly shows the specific steps and provides a unified guide for the user.

[1417] Input: Recipe information selected by the user

[1418] Output: Cooking instructions

[1419] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1420] Input: cooking instructions, speech model

[1421] Output: Audio guide data

[1422] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[1423] Input: Audio guide data

[1424] Output: Text display and voice guide

[1425] Step 7: Recognizing User Emotions

[1426] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone, for example, to check the user's stress level in real time.

[1427] Input: User's facial expression and voice data

[1428] Output: Parsed emotion data

[1429] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[1430] Input: Parsed emotion data

[1431] Output: Emotion data sent to the server

[1432] Server: Adjust the tone and content of the audio guide based on the emotional data. For example, if the user is feeling stressed, change the tone of the guide to a gentler tone.

[1433] Input: Emotion data

[1434] Output: Adjusted audio description data

[1435] Step 8: Gather and incorporate feedback

[1436] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user enters the difficulty of cooking and their impressions of the taste.

[1437] Input: User feedback

[1438] Output: Feedback data is saved in the app

[1439] Terminal: Sends feedback data to the server. Collects user ratings and comments digitally.

[1440] Input: User feedback data

[1441] Output: Feedback data sent to the server

[1442] Server: Analyzes the feedback data and updates the database to improve the cooking procedures and recipes for the next time. For example, improve a process that received poor taste ratings.

[1443] Input: Feedback data

[1444] Output: Updated recipe database

[1445] (Application example 2)

[1446] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1447] Traditionally, it has been difficult to accurately pass on a mother's cooking recipes to the next generation, and there is a need for a way to preserve those flavors for future generations, especially in an aging society. Furthermore, existing digital recipe systems lack guidance that takes into consideration the user's emotions, making it difficult to reduce stress and improve satisfaction while cooking. Furthermore, when considering use in brick-and-mortar stores, a system that can handle on-site challenges is required.

[1448] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1449] In this invention, the server includes means for digitizing mother's cooking manual, means for analyzing the digitized recipe data and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in a database, means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in the mother's voice, and means for recognizing the user's emotions using an emotion engine and adjusting the tone of the audio guidance. This makes it possible to accurately convey the taste of mother's cooking to the next generation and provide guidance that takes user emotions into consideration, making it possible to provide high satisfaction even when cooking in a physical store.

[1450] A "cooking manual" is a document and video data that contains recipes and procedures for the dishes that a mother would make on a daily basis.

[1451] "Digitalization" means converting information in analog form into digital data.

[1452] "Recipe data" is information that describes the ingredients, quantities, and cooking steps for making a dish.

[1453] A "database" is a digital information storage area that systematically organizes multiple data sets and makes them easy to manage and search.

[1454] "Speech recognition" is a technology that converts human speech into data in a format that a computer can understand.

[1455] "Voice cloning technology" is a technology that learns the characteristics of a specific person's voice and can read any text in that person's voice.

[1456] "Audio guide" is a function that uses voice to give instructions and explanations to users.

[1457] An "emotion engine" is a system that uses sensors such as cameras and microphones to analyze and recognize the user's emotional state and then respond appropriately based on that.

[1458] The system for carrying out this invention digitizes a mother's cooking manual and provides audio guidance of cooking procedures in a manner that takes into consideration the user's feelings. The following is a specific embodiment of this system.

[1459] The system begins with the user taking a photo of their mother's cooking manual. Using a smartphone or smart glasses, the user takes photos of the notes in the cooking manual and their mother cooking. The photos and videos are then uploaded to a server using a dedicated application. The server receives the uploaded data and temporarily stores it in a database.

[1460] The server then uses OCR technology to extract text from photos and videos, converting the recipe data into text format, which is then categorized and organized into recipe ingredients, quantities, and cooking steps.

[1461] When the user provides a voice description of their mother's cooking steps, the voice data recorded on their smartphone is also uploaded to the server using a dedicated application. The server receives the voice data and stores it in a database. Furthermore, it uses voice recognition technology to extract text data from the voice data, and uses voice cloning technology to learn the mother's voice and generate a voice model.

[1462] The server inputs the text data acquired by OCR into an AI model and analyzes information such as ingredients, quantities, cooking procedures, etc. The analyzed data is then structured and stored in a database.

[1463] When a user selects a recipe they want to cook within the application, the request is sent to the server. Based on the user's input of allergies and taste preferences (e.g., "low salt" or "medium spicy"), the server adjusts the recipe and generates appropriate instructions. The adjusted recipe and cooking steps are then audio-guided using the mother's voice.

[1464] While the user is cooking, the device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion engine, and sends the emotion data to the server. The server then adjusts the tone and content of the audio guidance during cooking based on the emotion data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[1465] After cooking, the user enters feedback about the taste and cooking procedure into the app. The device sends the feedback data to the server, which analyzes it and updates the database to improve the cooking procedure and recipe for the next time.

[1466] As a specific example, let's assume that the user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes the mother's video of the curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the low-salt specifications selected by the user, a recipe is generated with the salt amount adjusted. Based on this customized information, the mother's voice provides audio guidance such as, "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as, "It's okay to take it slowly."

[1467] Here are some example prompts using a generative AI model:

[1468] "Digitize your mom's curry recipe, customize it to low salt and medium spice, and provide audio guidance in your mom's voice. If the user is stressed, make the guidance tone gentler."

[1469] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1470] Step 1:

[1471] The user uses a smartphone or smart glasses to take pictures of their mother's cooking manual notes or the mother cooking. The input data is photos or videos, which are uploaded by the application and sent to the server. The server receives the uploaded images and videos and temporarily stores them in a database. This saves the physical recipe information as digital data.

[1472] Step 2:

[1473] The server uses OCR technology to extract text information from the received images and videos. The input is the image or video saved in step 1, and the output is the extracted text data. Specifically, the server uses the pytesseract library to perform character recognition and convert it into text recipe data.

[1474] Step 3:

[1475] The server categorizes and organizes the text data extracted by OCR into categories such as ingredients, quantities, and cooking procedures. The input is the text data obtained from OCR, and the output is structured recipe data. The text is analyzed using an AI model and stored in a database. This allows for efficient management of recipe information.

[1476] Step 4:

[1477] The user records their mother's cooking instructions on their smartphone. The data input is voice data, which is uploaded using a dedicated application and sent to the server. The server receives the voice data and stores it in a database.

[1478] Step 5:

[1479] The server uses speech recognition technology to extract text data from the uploaded audio data. The input is the audio data saved in step 4, and the output is the extracted text data. It then uses voice cloning technology to learn the mother's voice and generate a voice model. This makes it possible to read any text aloud in the mother's voice.

[1480] Step 6:

[1481] The user selects the recipe they want to cook within the dedicated application. The input is the user's recipe selection information, which is sent from the terminal to the server. Based on this information, the server retrieves the corresponding recipe data from the database.

[1482] Step 7:

[1483] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application. The input is customization information, which is sent from the device to the server. The server adjusts the recipe based on the customization information and generates new cooking instructions.

[1484] Step 8:

[1485] The server uses the trained voice model to generate audio guidance based on the adjusted recipe. The input is the adjusted recipe data, and the output is audio guidance. The audio guidance instructs the user on the cooking steps in a mother's voice.

[1486] Step 9:

[1487] While cooking, the device (smart glasses or head-mounted display) uses a camera and microphone to analyze the user's facial expressions and voice with an emotion engine. The input is the user's facial expressions and voice data, and the emotion data is sent from the device to the server.

[1488] Step 10:

[1489] The server adjusts the tone and content of the audio guide based on the emotion data. The input is emotion data, and the output is the adjusted audio guide. For example, if the user is feeling stressed, the tone of the guide can be made gentler and words of encouragement can be added.

[1490] Step 11:

[1491] After cooking is complete, the user enters feedback about the taste and cooking procedure in the app. The input is feedback data, which is sent from the device to the server. The server analyzes the feedback data and updates the database to improve the cooking procedure and recipe for the next time.

[1492] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1493] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1494] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1495] [Fourth embodiment]

[1496] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1497] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1498] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1499] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1500] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1501] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1502] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1503] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1504] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1505] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1506] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1507] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1508] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1509] This invention is a digital system for accurately passing on mother's cooking recipes to the next generation, aiming to deepen special family memories and bonds while recreating the flavor of mother's cooking. The program of this system includes functions for data collection, analysis, audio guidance, customization, and feedback.

[1510] 1. Data Collection

[1511] User: First, the user takes a note of his mother's cooking recipes and a video of his mother cooking on his smartphone.

[1512] Device: Upload the photos and videos you have taken to the server via the application.

[1513] Server: Receives the uploaded data and converts the recipe into text using OCR technology.

[1514] 2. Recipe analysis and database storage

[1515] Server: Recipe data converted to text using OCR is fed into the AI ​​model, and information such as ingredients, quantities, and cooking steps is extracted.

[1516] Server: Structures the extracted data and stores it in a database.

[1517] 3. Collecting and Learning Audio Data

[1518] User: Records mother explaining cooking steps on smartphone.

[1519] Device: Upload the recording data from the application to the server.

[1520] Server: The uploaded voice data is analyzed through a voice recognition process and converted into text. At the same time, voice cloning technology is used to learn the mother's voice and generate a voice model.

[1521] 4. Providing cooking guides

[1522] User: Selects the recipe they want to cook within the application.

[1523] Terminal: Sends a request for the selected recipe to the server.

[1524] Server: Retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[1525] Device: The acquired cooking instructions are displayed to the user, and instructions such as "First, cut the chicken into bite-sized pieces" are also provided in a mother's voice.

[1526] 5. User Customization

[1527] User: Enter allergies and individual taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1528] Terminal: Sends the entered customization information to the server.

[1529] Server: Adjusts the recipe based on the customization information and generates the appropriate instructions.

[1530] Device: Displays the customized recipe and provides audio guidance if needed.

[1531] Specific examples

[1532] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is given audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also making dishes that accommodate their own dietary restrictions.

[1533] In addition, after cooking, users can enter feedback into the app, which is analyzed by the server and the recipe database is constantly updated to provide more accurate guidance the next time the user cooks.

[1534] By combining the above functions, the system of the present invention can accurately pass on the taste of mother's cooking to the next generation, further deepening family ties.

[1535] The processing flow will be explained below.

[1536] Step 1: Data collection

[1537] User: Takes notes of his mother's cooking recipes and videos of himself cooking on his smartphone.

[1538] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[1539] Server: Receives the uploaded data and temporarily stores it in a database.

[1540] Step 2: Digitize your recipes

[1541] Server: Using OCR technology, extracts text information from photos and videos and converts recipes into text format.

[1542] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[1543] Step 3: Collecting audio data

[1544] User: Records her mother explaining the recipe steps on her smartphone.

[1545] Terminal: Upload the recorded audio data to the server using a dedicated application.

[1546] Server: Receives the uploaded audio data and stores it in a database.

[1547] Step 4: Analyze and train audio data

[1548] Server: Extracts text data from the voice data using voice recognition technology.

[1549] Server: Organizes the text data into cooking instructions and stores them in a database.

[1550] Server: Using voice cloning technology, learns the characteristics of the mother's voice and generates a voice model.

[1551] Step 5: Parse and save the recipe

[1552] Server: The text data obtained by OCR is fed into the AI ​​model, which analyzes ingredients, quantities, and cooking procedures.

[1553] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1554] Step 6: User selects recipe

[1555] User: Select the recipe they want to cook within the dedicated application.

[1556] Terminal: Sends a request for the selected recipe to the server.

[1557] Step 7: Cooking instructions generation and audio guidance

[1558] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1559] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1560] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1561] Step 8: Enter and update customization information

[1562] User: Enters specific dietary restrictions and taste preferences (e.g., "low salt," "medium spicy," etc.) into a dedicated application.

[1563] Terminal: Sends the entered customization information to the server.

[1564] Server: Adjusts the recipe based on the user's customization information.

[1565] Device: Provides the adjusted recipe to the user, and also provides audio guidance if necessary.

[1566] Step 9: Gather and incorporate feedback

[1567] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[1568] Terminal: Sends the input feedback data to the server.

[1569] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1570] Example 1

[1571] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1572] It is not easy to accurately pass on the flavors and steps of a mother's cooking to the next generation. If recipes are not accurately digitized, there is a high risk of losing special family memories and bonds. It is also difficult to satisfy the need to receive detailed cooking instructions in the mother's voice. Furthermore, it is difficult to accommodate the different dietary restrictions and taste preferences of each family. To solve these problems, accurately pass on the flavors of a mother's cooking to the next generation, and deepen family bonds, an efficient and accurate digital system is needed.

[1573] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1574] In this invention, the server includes means for digitizing mother's recipes, means for analyzing the digitized recipes and extracting ingredients, portions, and cooking methods, means for storing the extracted data in a data repository, means for collecting mother's voice data and learning mother's voice using voice recognition and voice synthesis technology, and means for providing audio guidance on cooking methods to users in mother's voice, thereby enabling mother's recipes to be accurately passed down to future generations and providing cooking guides customized to the specific needs of each household.

[1575] "My mother's cooking methods" refers to the cooking methods, steps, and recipes that my mother uses.

[1576] "Digitalization" refers to the process of converting analog information into electronic data.

[1577] "Ingredients" refers to the various foods and ingredients used in cooking.

[1578] "Amount" refers to the amount of each ingredient used when cooking.

[1579] "Cooking method" refers to a series of processes or steps to complete a dish using ingredients.

[1580] A "data repository" refers to a database or storage system for efficiently storing and managing collected data.

[1581] "Audio data" refers to digital data that contains recorded audio information such as human voices and spoken words.

[1582] "Speech recognition technology" refers to technology that generates text data from voice data.

[1583] "Speech synthesis technology" refers to technology that generates new voices based on certain voice samples.

[1584] "Voice guidance" refers to a method of conveying specific procedures or information to a user by providing voice guidance.

[1585] This invention is a digital system that aims to pass on mother's cooking techniques to the next generation and deepen family memories and bonds. The system has functions for data collection, analysis, audio guidance, customization, and feedback, and is designed to make it easier for users to recreate the taste of their mother's cooking.

[1586] Data collection

[1587] The first thing a user does is digitize their mother's cooking methods. They take photos of their mother's recipe notebooks and videos of their mother cooking with their smartphone. They then upload these photos and videos to a server using a dedicated application.

[1588] The server receives the uploaded photos and videos and converts the recipes into text data using OCR (Optical Character Recognition) technology, a process that converts analog information into electronic data.

[1589] Recipe analysis and database storage

[1590] The server then inputs the recipe data, converted to text using OCR technology, into an AI model to extract information such as ingredients, quantities, cooking methods, etc. Specifically, it uses natural language processing technology to analyze the recipe text and automatically identify the necessary information.

[1591] The extracted data is structured by the server and stored in a data repository, making it easy for users to search and browse later.

[1592] Audio data collection and learning

[1593] The user records their mother explaining the cooking steps on their smartphone, and the recording is then uploaded to the server via the application.

[1594] The server analyzes the uploaded voice data and converts it into text using speech recognition technology, while simultaneously learning the mother's voice using speech synthesis technology to generate a voice model.

[1595] Providing cooking guides

[1596] When cooking, the user selects the desired recipe within the application. The selected recipe request is sent from the device to the server. The server retrieves the corresponding recipe from the recipe database and generates an audio guide in the mother's voice along with cooking instructions.

[1597] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, allowing the user to confirm the cooking instructions both visually and audibly.

[1598] User Customization

[1599] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application, which then sends the information from the device to the server, which then adjusts the recipe accordingly.

[1600] The adjusted recipe is then sent back to the device and displayed to the user, with audio guidance provided if needed, allowing the user to create a dish tailored to their own preferences.

[1601] Specific examples

[1602] For example, consider a case where a user selects "curry" and prefers it "low salt" and "medium spicy." The system digitizes the mother's curry recipe using OCR technology and AI analysis. It then generates a recipe with adjusted salt content based on the user's customization information. Finally, the mother's voice provides audio guidance, saying, "Cut the chicken into bite-sized pieces and add a little salt." This allows the user to recreate the mother's unique flavor while also creating a dish that suits their own preferences.

[1603] Examples of prompt statements

[1604] "Please digitize my mother's homemade curry recipe using OCR technology and AI analysis. Then please guide me through the recipe, which is low in salt and medium in spiciness, using my mother's voice."

[1605] This system makes it possible to accurately pass on the flavor of a mother's cooking to the next generation, further strengthening family bonds.

[1606] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1607] Step 1: Data collection

[1608] User: Takes a photo of his mother's recipe notebook or a video of his mother cooking on his smartphone. The input is the image or video of the recipe.

[1609] Specific operation: The user launches the application and uses the camera function to take notes or record videos.

[1610] Terminal: Uploads captured photos and videos to the server via the application. The output is the transmission of image and video data to the server.

[1611] Specific operation: When you press the "Upload" button within the application, the device will transfer the captured data to the server.

[1612] Step 2: Convert the recipe to text

[1613] Server: Receives uploaded photos and videos and converts them into text data using OCR technology. The input is image or video data, and the output is text data.

[1614] Specific operation: The server uses OCR (Optical Character Recognition) software to convert the received image data into text.

[1615] Step 3: Recipe analysis and database storage

[1616] Server: Text-based recipe data is fed into the AI ​​model to extract information such as ingredients, quantities, cooking steps, etc. The input is text data, and the output is the extracted structured data.

[1617] How it works: The server inputs text data into the AI ​​model and executes a script that automatically analyzes and extracts ingredients and procedures.

[1618] Server: The extracted data is structured and stored in a data repository. The output is structured data stored in a database.

[1619] What it does: The server stores structured data in a relational database, allowing for efficient searching and management.

[1620] Step 4: Collecting and training audio data

[1621] User: Records the voice of the mother explaining the cooking procedure on a smartphone. The input is the mother's voice.

[1622] Specific actions: Use the application's voice recording function to record the mother's explanation.

[1623] Terminal: Uploads recorded data from the application to the server. The output is sending audio data to the server.

[1624] Specific operation: After recording is finished, press the "Upload" button and the audio data will be transferred to the server.

[1625] Server: Converts voice data into text using speech recognition technology, and learns the mother's voice using speech synthesis technology. The input is voice data, and the output is a voice model.

[1626] How it works: The server uses speech recognition software to convert speech to text, and then trains speech synthesis software to learn the mother's voice.

[1627] Step 5: Provide cooking guidance

[1628] User: Selects the recipe they want to cook within the application. Input is a recipe selection request.

[1629] Specific operation: Open "Recipe List" from the application menu and tap to select the desired recipe.

[1630] Terminal: Sends a request for the selected recipe to the server. The output is the request to the server.

[1631] Specific operation: When a user selects a recipe, a request is sent to the server.

[1632] Server: Retrieves the relevant recipe from the database and generates an audio guide in a mother's voice along with cooking instructions. The input is a recipe request, and the output is the audio guide data.

[1633] Specific operation: The server retrieves the relevant recipe from the database and generates a guide voice using the voice model.

[1634] Terminal: Displays cooking instructions to the user and provides audio guidance. Output is display data and audio guidance to the user.

[1635] Specific operation: The cooking steps are presented to the user visually and audibly. For example, the device may say, "First, cut the chicken into bite-sized pieces."

[1636] Step 6: Customizing Users

[1637] User: Enter allergies and taste preferences into the app. Input is customization information.

[1638] Specific operation: Access the form for entering customization information from the application's "Settings" screen and enter the required information.

[1639] Terminal: Sends input customization information to the server. Output: Sends customization information to the server.

[1640] Specific operation: When you press the "Save" button, the input information is sent to the server.

[1641] Server: Adjusts the recipe based on the customization information and generates appropriate instructions. The input is the customization information and the output is the adjusted recipe.

[1642] Specific operation: The server adjusts the amounts of ingredients and seasonings in the recipe based on the customization information.

[1643] Terminal: Displays the customized recipe and provides voice guidance if necessary. The output is the display data and voice guidance for the user.

[1644] What it does: It shows the user the adjusted recipe and provides audio guidance such as "Add a little salt to the chicken."

[1645] (Application example 1)

[1646] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1647] Traditionally, it has been difficult to accurately pass on family recipes to the next generation, especially when it comes to recreating the flavor of a mother's cooking. Enjoying these dishes through delivery services is also uncommon, making it even more challenging to accommodate individual taste preferences and dietary restrictions. This limits opportunities to deepen family memories and bonds. Continuously improving recipes based on feedback is also a challenge.

[1648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1649] In this invention, the server includes: means for digitizing mother's cooking recipes; means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps; means for storing the extracted data in a database; means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology; means for providing audio guidance to the user on cooking steps in mother's voice; means for cooking dishes based on the recipes and delivering them to the user via a delivery service; and means for collecting user feedback data and updating the recipe database. This allows mother's special recipes to be faithfully passed on to the next generation and for the user to enjoy the dishes through delivery. Furthermore, recipes can be continuously improved based on user feedback, thereby providing a consistently high-quality cooking experience.

[1650] The "recipe digitization method" is a method for collecting mother's recipes as photos and videos and converting them into digital format.

[1651] The "recipe analysis means" is a means for extracting ingredients, quantities, and cooking steps from digitized recipe data.

[1652] The "database storage means" is a means for structuring the analyzed recipe data and storing it in a database.

[1653] The "audio collection means" is a means for collecting and recording the audio of the mother explaining the cooking steps.

[1654] A "speech recognition means" is a means that uses technology to convert collected voice data into text.

[1655] The "voice cloning technology learning means" is a means for learning the mother's voice and reproducing the mother's voice based on the voice data.

[1656] The "audio guidance means" is a means for providing the user with audio guidance on cooking procedures in a mother's voice.

[1657] A "cooking tool" is a tool for cooking a dish based on a recipe.

[1658] "Delivery service means" refers to a means for delivering cooked food to a user through a delivery service.

[1659] A "feedback collection means" is a means for collecting feedback data from users.

[1660] The "recipe database update means" is a means for updating the recipe database based on the collected feedback data.

[1661] A "generative AI model" refers to an artificial intelligence model that extracts specific information from text data and performs analysis.

[1662] A "prompt sentence" is an input sentence for a generative AI model, and is text that contains instructions for analysis.

[1663] This invention is a system that digitizes mother's cooking recipes, accurately passes them on to the next generation, and provides customized meals through a delivery service. The system aims to allow users to recreate their mother's special dishes and deepen family ties.

[1664] Hardware and software used

[1665] Uses hardware and software that integrates servers, smartphones, and delivery services, including:

[1666] Smartphone: Used as data collection and operation interface.

[1667] Server: A central processing unit that performs data analysis, voice recognition, voice cloning, recipe generation, and delivery management.

[1668] Google Cloud Vision API: Uses OCR technology to convert digitized recipes into text.

[1669] OpenAI GPT-4: A generative AI model that analyzes recipe data and extracts ingredients and cooking steps.

[1670] Google Speech-to-Text API: Converts audio data into text.

[1671] Lyrebird or Descript: A voice cloning technology that recreates the mother's voice.

[1672] Uber Eats API: Manages delivery services to deliver prepared meals to users.

[1673] Data processing and operation procedures

[1674] 1. Data Collection

[1675] The user uses their smartphone to record their mother's recipe notes and videos of their mother cooking.

[1676] Photos and videos taken by the device are uploaded to the server via the application.

[1677] 2. Recipe analysis and database storage

[1678] The server receives the uploaded data and converts it into text using OCR technology using the Google Cloud Vision API.

[1679] The server inputs the recipe data, converted into text using OCR, into OpenAI GPT-4 to extract information such as ingredients, quantities, and cooking steps.

[1680] The server structures the extracted data and stores it in a database.

[1681] 3. Collecting and Learning Speech Data

[1682] The user records the audio of their mother explaining the cooking steps on their smartphone.

[1683] The device uploads the recording data from the application to the server.

[1684] The server analyzes the uploaded voice data using the Google Speech-to-Text API and converts it into text, while simultaneously training the mother's voice using Lyrebird or Descript to generate a voice model.

[1685] 4. Providing cooking guides

[1686] The user selects the recipe they want to cook within the application.

[1687] The device sends a request for the selected recipe to the server.

[1688] The server retrieves the relevant recipe from the recipe database and generates a guide in a mother's voice along with cooking instructions.

[1689] The device displays the cooking instructions to the user and also provides audio guidance in the mother's voice, such as "First, cut the chicken into bite-sized pieces."

[1690] 5. Food preparation and delivery

[1691] The server arranges for the food to be cooked based on the generated recipe and cooking instructions.

[1692] Once cooked, the food is delivered to the user using the Uber Eats API.

[1693] 6. User Feedback

[1694] The user enters post-cooking feedback into the application.

[1695] The server analyzes the feedback data and updates the recipe database to improve the accuracy of future cooking guides.

[1696] Specific examples

[1697] For example, if a user selects "Mom's Special Curry" and requests "low salt" and "medium spicy," the system analyzes a photo of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. A recipe with adjusted salt content is generated based on the user's low-salt specifications. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." After cooking, the food is delivered to the user via Uber Eats.

[1698] Prompt Sentence Examples

[1699] Analyze photos, videos, and audio data of Mom's recipes uploaded by users to extract ingredients and cooking steps, clone Mom's voice from the audio data, and generate an audio guide that provides cooking instructions in Mom's voice along with the user's desired customization options. Once cooking is complete, deliver the food using the Uber Eats API.

[1700] As a result, the present invention can provide a high-quality cooking experience by faithfully passing on mother's special recipes to the next generation and customizing them according to the user's wishes.

[1701] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1702] Step 1:

[1703] Data collection

[1704] The user takes photos of their mother's cooking recipe notes or videos of their mother cooking with their smartphone. The device then uploads the photos and videos to a server via an application. The input is the photos and videos, and the output is the data sent to the server.

[1705] Step 2:

[1706] Text conversion using OCR

[1707] The server receives the uploaded photos and videos and converts them into text using OCR technology using the Google Cloud Vision API. The input is the photo or video data, and the output is the recipe data in text format. This extracts text information from the photos and videos.

[1708] Step 3:

[1709] Recipe analysis and database storage

[1710] The server inputs the recipe data converted to text using OCR into OpenAI GPT-4, and uses a generative AI model to extract information such as ingredients, quantities, and cooking steps. The input is the converted recipe data, and the output is the extracted structured data. The server then structures the extracted data and stores it in a database. This allows the recipe's digital information to be organized and stored.

[1711] Step 4:

[1712] Audio data collection

[1713] The user records the voice of his mother explaining the cooking procedure on his smartphone. The device uploads the recorded data to the server via the application. The input is the recorded voice data, and the output is the voice data sent to the server.

[1714] Step 5:

[1715] Speech Recognition and Speech Clone Training

[1716] The server analyzes the uploaded audio data using the Google Speech-to-Text API and converts the audio to text. At the same time, it uses Lyrebird or Descript to train the mother's voice and generate a voice model. The input is the audio data, and the output is the converted audio data and a voice model of the mother's voice.

[1717] Step 6:

[1718] Generate cooking guide

[1719] The user selects the recipe they want to cook within the application. The device sends a request for the selected recipe to the server. The server retrieves the recipe from the recipe database and generates audio instructions for the cooking process in a mother's voice. The input is the selected recipe request, and the output is the cooking process with audio instructions.

[1720] Step 7:

[1721] Food preparation and delivery

[1722] The server arranges for the food to be cooked based on the generated recipe and cooking instructions. The cooked food is delivered to the user using the Uber Eats API. The input is the cooking instructions and delivery address information, and the output is the delivered food.

[1723] Step 8:

[1724] Feedback collection and database updates

[1725] Users input their post-cooking feedback into the application. The server analyzes the feedback data and updates the recipe database. The input is the feedback data, and the output is an updated recipe database. This improves the accuracy of cooking guidance from the next time onwards.

[1726] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1727] This invention is a system that digitizes mother's cooking recipes to pass them on to the next generation, and combines them with an emotion engine that recognizes the user's emotions to provide greater emotional connection and satisfaction. The program of this system includes functions for data collection, analysis, voice guidance, customization, feedback, and emotion recognition.

[1728] 1. Data Collection

[1729] User: First, the user takes a photo of the notebook containing his mother's cooking recipes and a video of his mother cooking with his smartphone.

[1730] Device: Upload the photos and videos you have taken to the server using a dedicated application.

[1731] Server: Receives the uploaded data and temporarily stores it in a database.

[1732] 2. Digitizing recipes

[1733] Server: Uses OCR technology to extract text information from photos and videos and convert recipes into text format.

[1734] Server: The extracted text data is classified and organized into recipe ingredients, quantities, and steps.

[1735] 3. Collecting and Learning Audio Data

[1736] User: Records mother explaining cooking steps on smartphone.

[1737] Terminal: Upload the recorded audio data to the server using a dedicated application.

[1738] Server: Receives the uploaded audio data and stores it in a database.

[1739] Server: Extracts text data from the voice data using voice recognition technology, learns the mother's voice using voice cloning technology, and generates a voice model.

[1740] 4. Parse and save the recipe

[1741] Server: The text data obtained by OCR is input into an AI model to analyze information such as ingredients, quantities, and cooking procedures.

[1742] Server: The analyzed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1743] 5. User-selected and customized recipes

[1744] User: Select the recipe they want to cook within the dedicated application.

[1745] Terminal: Sends a request for the selected recipe to the server.

[1746] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1747] Terminal: Sends the entered customization information to the server.

[1748] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1749] 6. Providing cooking guides

[1750] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1751] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1752] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1753] 7. User Emotion Recognition

[1754] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone.

[1755] Terminal: The emotion engine sends the user's emotion data to the server.

[1756] Server: Adjusts the tone and content of the audio guidance during cooking based on emotional data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[1757] Server: It can analyze the user's emotional data and suggest recipes and cooking procedures that suit their emotions.

[1758] 8. Gather and incorporate feedback

[1759] User: After cooking is complete, enter feedback into the app about the taste and process of the dish.

[1760] Terminal: Sends the input feedback data to the server.

[1761] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1762] Specific examples

[1763] For example, let's say a user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes a photograph of the mother's curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the user's low-salt specifications, a recipe is generated with the salt amount adjusted. Based on this customized information, the user is provided with audio guidance in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as "It's okay to take it slowly."

[1764] By combining the above functions, the system of the present invention can accurately convey the taste of mother's cooking to the next generation, further deepening emotional connections and satisfaction.

[1765] The processing flow will be explained below.

[1766] Step 1: Data collection

[1767] User: Records his mother's cooking recipes in a notebook and takes photos of his mother cooking with his smartphone.

[1768] Terminal: Uploads captured photos and videos to the server using a dedicated application.

[1769] Server: Receives the uploaded data and temporarily stores it in a database.

[1770] Step 2: Digitize your recipes

[1771] Server: Extracts text information from uploaded images and videos using OCR technology.

[1772] Server: Analyzes the extracted text data and classifies and organizes it into ingredients, quantities, and cooking steps in the recipe.

[1773] Step 3: Collecting and training audio data

[1774] User: Records mother explaining cooking steps on smartphone.

[1775] Terminal: Uploads recorded audio data to the server via a dedicated application.

[1776] Server: Receives the uploaded audio data and stores it in a database.

[1777] Server: Uses voice recognition technology to convert voice data into text data.

[1778] Server: Using voice cloning technology, the characteristics of the mother's voice are learned and a voice model is generated.

[1779] Step 4: Parse and save the recipe

[1780] Server: The text data obtained by OCR is input into an AI model to analyze ingredients, quantities, and cooking procedures.

[1781] Server: The parsed data is structured and stored in a database. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1782] Step 5: User Recipe Selection and Customization

[1783] User: Select the recipe they want to cook within the dedicated application.

[1784] Terminal: Sends a request for the selected recipe to the server.

[1785] User: Enter allergies and taste preferences (e.g., "low salt," "medium spicy," etc.) into the app.

[1786] Terminal: Sends the entered customization information to the server.

[1787] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1788] Step 6: Provide cooking guidance

[1789] Server: Retrieves the selected recipe from the database and generates cooking instructions.

[1790] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1791] Device: The cooking instructions are displayed to the user in text, and a mother's voice provides audio guidance saying, "First, cut the chicken into bite-sized pieces."

[1792] Step 7: Recognizing User Emotions

[1793] Device: The emotion engine analyzes the user's facial expressions and voice using the camera and microphone.

[1794] Terminal: Transmits the analyzed emotion data to the server.

[1795] Server: Adjusts the tone and content of the audio guide during cooking based on emotional data. For example, if it detects that the user is feeling stressed, it will soften the tone of the audio guide and insert encouraging words such as "It's okay to take it slowly."

[1796] Step 8: Suggest an emotionally appropriate recipe

[1797] Server: Analyzes the user's emotional data and suggests recipes and cooking procedures that suit their emotional state. For example, if the user is having fun, it can suggest more challenging recipes.

[1798] Terminal: Notifies users of emotion-based recipe suggestions.

[1799] Step 9: Gather and incorporate feedback

[1800] User: After cooking is complete, the user enters feedback about the taste of the finished dish and the cooking procedure into a dedicated application.

[1801] Terminal: Sends feedback data to the server.

[1802] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time.

[1803] Example 2

[1804] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1805] Conventional systems offer limited ways to pass on mother's cooking recipes to the next generation and lack mechanisms for providing emotional connection and satisfaction. They also struggle to customize to accommodate users' dietary restrictions and individual preferences, and are unable to recognize users' emotions and provide appropriate assistance while cooking. This increases the likelihood of users feeling stressed or creating unsatisfying dishes.

[1806] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1807] In this invention, the server includes means for digitizing mother's cooking recipes, means for analyzing the digitized recipes and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in an information management system, means for collecting mother's voice data and learning mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in mother's voice, and means for recognizing the user's emotions and adjusting the tone and content of the audio guidance based on those emotions. This makes it possible to accurately pass on mother's cooking recipes to the next generation and provide cooking support that takes the user's emotions into consideration, providing a high level of satisfaction.

[1808] "Digitization" is the process of converting analog information into digital form.

[1809] "Analysis" is the process of examining data in detail to understand its structure and meaning.

[1810] "Ingredients" refers to the food ingredients used to make a dish.

[1811] "Quantity" refers to the specific amount of an ingredient used in a dish.

[1812] A "cooking procedure" is a sequence of steps for cooking a dish.

[1813] An "information management system" is a system for storing, managing, retrieving, and using data.

[1814] "Audio data" refers to digital information that records sound.

[1815] "Speech recognition" is a technology that converts voice data into text data.

[1816] "Voice cloning technology" is a technology for reproducing a specific voice.

[1817] "Audio guide" is a means of giving instructions or explanations using voice.

[1818] "User" refers to a person who uses this system.

[1819] "Customization" refers to modifying a system or service to meet a user's specific needs or requirements.

[1820] "Emotion recognition" is a technology that analyzes changes in a user's facial expressions and voice to determine their emotions.

[1821] "Feedback Data" refers to the evaluations and opinions provided by users after using the system.

[1822] This invention is a system that digitizes mother's cooking recipes, passes them on to the next generation, and recognizes the user's emotions. This system includes the following functions: data collection, digitization, voice data collection and learning, recipe analysis and storage, user recipe selection and customization, cooking guide provision, user emotion recognition, and feedback collection and reflection.

[1823] Data collection

[1824] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. For example, he takes photos of each page of the notebook and records a video of his mother cooking in the kitchen.

[1825] Device: Photos and videos are uploaded to the server using a dedicated application. The device checks the data format and network connection status to upload the data appropriately.

[1826] Server: Receives the uploaded data and temporarily stores it in the database. At this time, it checks the format and size of the data and converts it into the appropriate format.

[1827] Recipe digitization

[1828] Server: Using OCR technology such as Google Cloud Vision API, text information is extracted from photos and videos and converted into text format.

[1829] Server: Analyzes text data and classifies and organizes it into recipe ingredients, quantities, and steps. For example, converts "200g of chicken" and "2 potatoes" in an image into text data.

[1830] Audio data collection and learning

[1831] User: Record your mother explaining cooking steps on your smartphone. For example, record your mother explaining, "Cut the chicken into bite-sized pieces."

[1832] Terminal: The recorded audio data is uploaded to the server using a dedicated application. At this time, the format of the audio data is checked and converted to the appropriate format.

[1833] Server: Receives the uploaded voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[1834] Recipe analysis and saving

[1835] Server: The text data obtained by OCR is input into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking procedures.

[1836] Server: The analyzed data is structured and stored in a database such as MongoDB. For example, it can be saved in a format such as "Ingredients: 200g chicken, 2 potatoes" and "Steps: Cut the chicken into bite-sized pieces, peel the potatoes."

[1837] User-selected and customized recipes

[1838] User: Selects the recipe they want to cook within the dedicated application, for example, curry or stew.

[1839] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[1840] Terminal: Sends the input customization information to the server. The server adjusts the recipe based on the conditions entered by the user.

[1841] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions.

[1842] Providing cooking guides

[1843] Server: Retrieves the selected recipe from the database and generates cooking instructions. Specific steps are clearly displayed and presented to the user in an easy-to-understand manner.

[1844] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1845] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[1846] User Emotion Recognition

[1847] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone. For example, it monitors the user's face in real time and analyzes their emotions.

[1848] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[1849] Server: Adjusts the tone and content of the audio guide based on emotional data. For example, if the user is feeling stressed, the tone of the guide can be changed to a gentler tone. It can also suggest recipes and cooking procedures that are suitable for the user.

[1850] Gathering and implementing feedback

[1851] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user comments on the strength of the flavor and the ease of understanding the process.

[1852] Terminal: Sends feedback data to the server. Collects user evaluations as digital data.

[1853] Server: Analyzes the feedback data and updates the database to improve cooking procedures and recipes for the next time, for example, adjusting the taste based on the feedback.

[1854] Specific examples

[1855] For example, if a user selects "curry" and prefers low salt and medium spiciness, the system will convert their mother's curry recipe into text using OCR and AI technology. Based on the user's customized information, a recipe with adjusted salt content will be generated. A voice guide will be provided in the mother's voice, such as "Cut the chicken into bite-sized pieces and add a little salt." If the user becomes stressed, the emotion engine will detect this and change the tone of the voice guide to a gentler tone, encouraging the user by saying, "It's okay to take it easy."

[1856] Example prompts for generative AI models

[1857] "Digitalize your mother's recipes and design a cooking support system that combines voice guidance and emotion recognition."

[1858] "Explain OCR and AI techniques used to analyze photos of cooking recipes and extract ingredients and steps."

[1859] The system of the present invention not only passes on mother's cooking recipes to the next generation, but also provides cooking support that takes the user's emotions into consideration, thereby providing greater satisfaction and emotional connection.

[1860] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1861] Step 1: Data collection

[1862] User: Takes photos of his mother's notebook containing recipes and of her cooking with his smartphone. Specifically, he carefully photographs each page of the notebook and records a video of his mother cooking in the kitchen.

[1863] Input: Photos and videos taken

[1864] Output: Captured data is saved on your smartphone (image and video files saved in local storage)

[1865] Device: Uploads photos and videos to the server using a dedicated application. During upload, the network connection status and data integrity are checked.

[1866] Input: Image and video data stored on a smartphone

[1867] Output: Data upload process to the server is completed, and data saved on the server

[1868] Step 2: Digitize your recipes

[1869] Server: Using OCR technology such as Google Cloud Vision API, character information is extracted from photos and videos and converted into text. For example, handwritten or printed characters in an image are converted into text.

[1870] Input: Uploaded photo and video data

[1871] Output: Text data (e.g., "200g chicken, 2 potatoes, cut the chicken into bite-sized pieces")

[1872] Server: Analyzes text data and classifies it into ingredients, quantities, and cooking procedures. Using natural language processing technology, the text is analyzed by phrase and classified into categories.

[1873] Input: Extracted text data

[1874] Output: Structured data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Steps: Cut the chicken into bite-sized pieces")

[1875] Server: Stores organized data in an information management system.

[1876] Input: Structured data

[1877] Output: Recipe data stored in the database

[1878] Step 3: Collecting and training audio data

[1879] User: Record your mother explaining the cooking steps on your smartphone. Specifically, record your mother explaining, "Cut the chicken into bite-sized pieces."

[1880] Input: Recorded audio data

[1881] Output: Audio file saved on your smartphone (e.g. mp3 or wav format)

[1882] Device: The recorded audio data is uploaded to the server using a dedicated application. During the upload, the format and integrity of the audio file are checked.

[1883] Input: Audio file

[1884] Output: Audio data sent to the server

[1885] Server: Receives the voice data and converts it into text data using the Google Speech-to-Text API. It also generates a voice model of the mother's voice using voice cloning technology.

[1886] Input: Uploaded audio data

[1887] Output: Text data and speech model

[1888] Step 4: Parse and save the recipe

[1889] Server: The text data acquired by OCR is fed into an AI model (e.g., OpenAI's GPT-3) to analyze information such as ingredients, quantities, and cooking steps. The AI ​​model analyzes the context and meaning, clearly distinguishing between ingredients and steps.

[1890] Input: Text data obtained by OCR

[1891] Output: Analyzed data (e.g., "Ingredients: 200g chicken, 2 potatoes" "Procedure: Cut the chicken into bite-sized pieces")

[1892] Server: The parsed data is structured and stored in a database such as MongoDB. The structured data is saved in an easy-to-search format.

[1893] Input: Parsed data

[1894] Output: Recipe information stored in the database

[1895] Step 5: User Recipe Selection and Customization

[1896] User: Selects the recipe they want to cook within the dedicated application. For example, they can select a dish such as curry or stew.

[1897] Input: Recipe selection by user

[1898] Output: The selected recipe information is displayed in the application.

[1899] Terminal: Sends a request for the selected recipe to the server. In addition, the user inputs allergies and taste preferences (e.g., "low salt" or "medium spicy").

[1900] Input: User preferences and conditions

[1901] Output: Customization information sent to the server

[1902] Server: Adjusts the recipe based on the user's customization information and generates appropriate instructions. Adjusts the recipe based on the conditions specified by the user.

[1903] Input: Customization information

[1904] Output: Customized recipe information

[1905] Step 6: Provide cooking guidance

[1906] Server: Retrieves the selected recipe from the database and generates cooking instructions. It clearly shows the specific steps and provides a unified guide for the user.

[1907] Input: Recipe information selected by the user

[1908] Output: Cooking instructions

[1909] Server: Uses a trained speech model to provide audio guidance in a mother's voice for the generated cooking instructions.

[1910] Input: cooking instructions, speech model

[1911] Output: Audio guide data

[1912] Device: The cooking instructions are displayed to the user in text form, and audio guidance is provided in a mother's voice. For example, instructions such as "First, cut the chicken into bite-sized pieces" are provided.

[1913] Input: Audio guide data

[1914] Output: Text display and voice guide

[1915] Step 7: Recognizing User Emotions

[1916] Device: The emotion engine analyzes the user's facial expressions and voice through the camera and microphone, for example, to check the user's stress level in real time.

[1917] Input: User's facial expression and voice data

[1918] Output: Parsed emotion data

[1919] Terminal: Sends emotional data to the server. The user's emotional state, such as the level of stress or joy, is quantified and sent.

[1920] Input: Parsed emotion data

[1921] Output: Emotion data sent to the server

[1922] Server: Adjust the tone and content of the audio guide based on the emotional data. For example, if the user is feeling stressed, change the tone of the guide to a gentler tone.

[1923] Input: Emotion data

[1924] Output: Adjusted audio description data

[1925] Step 8: Gather and incorporate feedback

[1926] User: After cooking is complete, the user enters feedback into the app about the taste and process of the dish. For example, the user enters the difficulty of cooking and their impressions of the taste.

[1927] Input: User feedback

[1928] Output: Feedback data is saved in the app

[1929] Terminal: Sends feedback data to the server. Collects user ratings and comments digitally.

[1930] Input: User feedback data

[1931] Output: Feedback data sent to the server

[1932] Server: Analyzes the feedback data and updates the database to improve the cooking procedures and recipes for the next time. For example, improve a process that received poor taste ratings.

[1933] Input: Feedback data

[1934] Output: Updated recipe database

[1935] (Application example 2)

[1936] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1937] Traditionally, it has been difficult to accurately pass on a mother's cooking recipes to the next generation, and there is a need for a way to preserve those flavors for future generations, especially in an aging society. Furthermore, existing digital recipe systems lack guidance that takes into consideration the user's emotions, making it difficult to reduce stress and improve satisfaction while cooking. Furthermore, when considering use in brick-and-mortar stores, a system that can handle on-site challenges is required.

[1938] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1939] In this invention, the server includes means for digitizing mother's cooking manual, means for analyzing the digitized recipe data and extracting ingredients, quantities, and cooking steps, means for storing the extracted data in a database, means for collecting mother's voice data and learning the mother's voice using voice recognition and voice cloning technology, means for providing audio guidance of cooking steps to the user in the mother's voice, and means for recognizing the user's emotions using an emotion engine and adjusting the tone of the audio guidance. This makes it possible to accurately convey the taste of mother's cooking to the next generation and provide guidance that takes user emotions into consideration, making it possible to provide high satisfaction even when cooking in a physical store.

[1940] A "cooking manual" is a document and video data that contains recipes and procedures for the dishes that a mother would make on a daily basis.

[1941] "Digitalization" means converting information in analog form into digital data.

[1942] "Recipe data" is information that describes the ingredients, quantities, and cooking steps for making a dish.

[1943] A "database" is a digital information storage area that systematically organizes multiple data sets and makes them easy to manage and search.

[1944] "Speech recognition" is a technology that converts human speech into data in a format that a computer can understand.

[1945] "Voice cloning technology" is a technology that learns the characteristics of a specific person's voice and can read any text in that person's voice.

[1946] "Audio guide" is a function that uses voice to give instructions and explanations to users.

[1947] An "emotion engine" is a system that uses sensors such as cameras and microphones to analyze and recognize the user's emotional state and then respond appropriately based on that.

[1948] The system for carrying out this invention digitizes a mother's cooking manual and provides audio guidance of cooking procedures in a manner that takes into consideration the user's feelings. The following is a specific embodiment of this system.

[1949] The system begins with the user taking a photo of their mother's cooking manual. Using a smartphone or smart glasses, the user takes photos of the notes in the cooking manual and their mother cooking. The photos and videos are then uploaded to a server using a dedicated application. The server receives the uploaded data and temporarily stores it in a database.

[1950] The server then uses OCR technology to extract text from photos and videos, converting the recipe data into text format, which is then categorized and organized into recipe ingredients, quantities, and cooking steps.

[1951] When the user provides a voice description of their mother's cooking steps, the voice data recorded on their smartphone is also uploaded to the server using a dedicated application. The server receives the voice data and stores it in a database. Furthermore, it uses voice recognition technology to extract text data from the voice data, and uses voice cloning technology to learn the mother's voice and generate a voice model.

[1952] The server inputs the text data acquired by OCR into an AI model and analyzes information such as ingredients, quantities, cooking procedures, etc. The analyzed data is then structured and stored in a database.

[1953] When a user selects a recipe they want to cook within the application, the request is sent to the server. Based on the user's input of allergies and taste preferences (e.g., "low salt" or "medium spicy"), the server adjusts the recipe and generates appropriate instructions. The adjusted recipe and cooking steps are then audio-guided using the mother's voice.

[1954] While the user is cooking, the device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion engine, and sends the emotion data to the server. The server then adjusts the tone and content of the audio guidance during cooking based on the emotion data. For example, if the user is feeling stressed, the audio guidance can be changed to a gentler tone.

[1955] After cooking, the user enters feedback about the taste and cooking procedure into the app. The device sends the feedback data to the server, which analyzes it and updates the database to improve the cooking procedure and recipe for the next time.

[1956] As a specific example, let's assume that the user selects "curry" and prefers it low-salt and medium-spicy. The system analyzes the mother's video of the curry recipe and digitizes the ingredients and steps using OCR and AI technology. Based on the low-salt specifications selected by the user, a recipe is generated with the salt amount adjusted. Based on this customized information, the mother's voice provides audio guidance such as, "Cut the chicken into bite-sized pieces and add a little salt." Furthermore, if the emotion engine detects the user's stress, it can soften the tone of the audio guidance and add encouraging words such as, "It's okay to take it slowly."

[1957] Here are some example prompts using a generative AI model:

[1958] "Digitize your mom's curry recipe, customize it to low salt and medium spice, and provide audio guidance in your mom's voice. If the user is stressed, make the guidance tone gentler."

[1959] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1960] Step 1:

[1961] The user uses a smartphone or smart glasses to take pictures of their mother's cooking manual notes or the mother cooking. The input data is photos or videos, which are uploaded by the application and sent to the server. The server receives the uploaded images and videos and temporarily stores them in a database. This saves the physical recipe information as digital data.

[1962] Step 2:

[1963] The server uses OCR technology to extract text information from the received images and videos. The input is the image or video saved in step 1, and the output is the extracted text data. Specifically, the server uses the pytesseract library to perform character recognition and convert it into text recipe data.

[1964] Step 3:

[1965] The server categorizes and organizes the text data extracted by OCR into categories such as ingredients, quantities, and cooking procedures. The input is the text data obtained from OCR, and the output is structured recipe data. The text is analyzed using an AI model and stored in a database. This allows for efficient management of recipe information.

[1966] Step 4:

[1967] The user records their mother's cooking instructions on their smartphone. The data input is voice data, which is uploaded using a dedicated application and sent to the server. The server receives the voice data and stores it in a database.

[1968] Step 5:

[1969] The server uses speech recognition technology to extract text data from the uploaded audio data. The input is the audio data saved in step 4, and the output is the extracted text data. It then uses voice cloning technology to learn the mother's voice and generate a voice model. This makes it possible to read any text aloud in the mother's voice.

[1970] Step 6:

[1971] The user selects the recipe they want to cook within the dedicated application. The input is the user's recipe selection information, which is sent from the terminal to the server. Based on this information, the server retrieves the corresponding recipe data from the database.

[1972] Step 7:

[1973] Users input their dietary restrictions and taste preferences (e.g., "low salt" or "medium spicy") into the application. The input is customization information, which is sent from the device to the server. The server adjusts the recipe based on the customization information and generates new cooking instructions.

[1974] Step 8:

[1975] The server uses the trained voice model to generate audio guidance based on the adjusted recipe. The input is the adjusted recipe data, and the output is audio guidance. The audio guidance instructs the user on the cooking steps in a mother's voice.

[1976] Step 9:

[1977] While cooking, the device (smart glasses or head-mounted display) uses a camera and microphone to analyze the user's facial expressions and voice with an emotion engine. The input is the user's facial expressions and voice data, and the emotion data is sent from the device to the server.

[1978] Step 10:

[1979] The server adjusts the tone and content of the audio guide based on the emotion data. The input is emotion data, and the output is the adjusted audio guide. For example, if the user is feeling stressed, the tone of the guide can be made gentler and words of encouragement can be added.

[1980] Step 11:

[1981] After cooking is complete, the user enters feedback about the taste and cooking procedure in the app. The input is feedback data, which is sent from the device to the server. The server analyzes the feedback data and updates the database to improve the cooking procedure and recipe for the next time.

[1982] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1983] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1984] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1985] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1986] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and me...

Claims

1. A way to digitize my mother's cooking recipes, A means of analyzing the digitized recipe and extracting ingredients, quantities, and cooking steps; means for storing the extracted data in a database; A means for collecting voice data of the mother and learning the mother's voice using voice recognition and voice cloning technology; A means for providing audio guidance to the user on cooking procedures in a mother's voice; A system including:

2. 10. The system of claim 1, further comprising means for customizing recipes based on a user's dietary restrictions and taste preferences.

3. 10. The system of claim 1, further comprising means for collecting feedback data from users and updating the recipe database.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A