System
The system addresses the challenge of finding suitable recipes by converting user voice to text, analyzing preferences, and refining recipes with generative AI, providing voice feedback and improving accuracy through user feedback, enhancing the cooking experience.
Patent Information
- Application Number
- JP2024125341
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Users struggle to find suitable recipes from numerous options and face inconvenience while cooking, especially when operating a smartphone, and there is a lack of systems that efficiently manage user preferences, allergy information, and recipe posting.
A system that captures user voice, converts it to text, analyzes preferences, selects recipes, and provides voice feedback, while storing preferences and allergy information, refining recipes with generative AI, and improving accuracy through user feedback.
Enables convenient recipe selection and posting, reflecting user preferences and improving recipe quality by excluding low-rated options, thus enhancing the cooking experience.
Smart Images

Figure 2026023406000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern recipe provision systems, users struggle to find the recipe that best suits them from the numerous recipes available. It is also inconvenient to operate a smartphone while cooking, and users have to work hard to post recipes. The present invention aims to solve these problems and improve users' cooking experience by realizing a system that provides convenient and effective recipe selection. [Means for solving the problem]
[0005] The present invention is a system including the following means: a means for capturing a user's voice and converting the voice data into text data; a means for analyzing the text data, providing a means for selecting a recipe based on the user's request, converting the selected recipe into voice data, and providing the voice data to the user; a means for storing the user's food preferences and allergy information in a database and recommending recipes based on the stored information; a means for improving the recipes entered by the user and providing an automatically generated posting page; a means for collecting feedback data and regularly updating recipe ratings; and a means for excluding low-rated recipes from the learning data based on the feedback data to improve the accuracy of the system's algorithm. This configuration allows users to easily find recipes that suit them, simplifies operations while cooking, and efficiently posts and rates recipes.
[0006] "Audio capture means" refers to devices or software that capture the user's voice as digital data.
[0007] "Means for converting voice data into text data" refers to the technology or algorithm that analyzes acquired voice data and converts it into corresponding text data.
[0008] "Means for analyzing text data" refers to information processing technology for understanding a user's intentions and requests based on converted text data.
[0009] "Means for selecting recipes" refers to the algorithms and software used to select the most suitable recipes from the analyzed user requests.
[0010] "Means for converting into audio data" refers to the technology or algorithm that generates the recipe information from the selected text data as audio data.
[0011] "Means for storing user's food preferences and allergy information in a database" refers to a database system for managing and storing individual preferences and allergy information entered by users.
[0012] "Means for recommending recipes" refers to algorithms or software that suggest optimal recipes based on stored user preferences and allergy information.
[0013] "Means for improving cooking recipes" refers to generative AI technology that formats recipe information entered by users into a more attractive and consistent format.
[0014] "Means for providing an automatically generated posting page" refers to a system for automatically creating and publishing a posting page based on polished recipe information.
[0015] "Means for collecting feedback data" refers to a system for obtaining and managing ratings and comments on recipes provided by users.
[0016] "Means for continuously updating recipe ratings" refers to the technology and algorithms that keep each recipe's rating information up to date based on collected feedback data.
[0017] "Means for removing low-rated recipes from the training data" refers to algorithms or processes that analyze collected feedback data and remove recipes that are determined to have low ratings from the system's training data.
[0018] "Means to improve the accuracy of the system's algorithm" refers to technologies and processes for improving the accuracy of the system's recommendation algorithm and generative AI by excluding recipes with low ratings. [Brief explanation of the drawings]
[0019] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[0041] 1. Implementing a voice-activated UI
[0042] Program Overview:
[0043] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0044] Processing Description:
[0045] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[0046] Examples:
[0047] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[0048] 2. Register user preferences and introduce the most suitable menu
[0049] Program Overview:
[0050] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0051] Processing Description:
[0052] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[0053] Examples:
[0054] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[0055] 3. Recipe posting function to the community
[0056] Program Overview:
[0057] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0058] Processing Description:
[0059] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[0060] Examples:
[0061] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[0062] 4. Data refinement based on user feedback
[0063] Program Overview:
[0064] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0065] Processing Description:
[0066] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[0067] Examples:
[0068] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[0069] In this way, the system of the present invention provides cooking recipes based on voice interaction and reflects user preferences and community participation, allowing users to enjoy a more comfortable and convenient cooking experience.
[0070] The processing flow will be explained below.
[0071] 1. Implementing a voice-activated UI
[0072] Program Overview:
[0073] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0074] Processing Steps:
[0075] Step 1:
[0076] Users speak into the device to ask questions or make requests about food, such as "What would you like for dinner tonight?"
[0077] Step 2:
[0078] The device captures the user's voice. The device uses a built-in microphone to obtain the voice data.
[0079] Step 3:
[0080] The device transmits the captured audio data to the server, and the device transmits the audio data via the Internet.
[0081] Step 4:
[0082] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[0083] Step 5:
[0084] The server analyzes the text data to understand the user's intent, and then searches for appropriate recipe information based on the analyzed data.
[0085] Step 6:
[0086] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[0087] Step 7:
[0088] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[0089] Step 8:
[0090] The server transmits the generated voice data to the terminal, and the voice data is transmitted via the Internet.
[0091] Step 9:
[0092] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[0093] 2. Register user preferences and introduce the most suitable menu
[0094] Program Overview:
[0095] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0096] Processing Steps:
[0097] Step 1:
[0098] The user inputs their food preferences and allergy information through the terminal, such as "vegetarian" or "nut allergy."
[0099] Step 2:
[0100] The device sends the input preference information to the server, and the input data is sent to the server via the Internet.
[0101] Step 3:
[0102] The server stores the received information in a database as a user profile. Profile data is created and stored for each user.
[0103] Step 4:
[0104] The user inputs a specific food request into the device, for example, asking "What do you recommend for lunch today?"
[0105] Step 5:
[0106] The device captures the audio and sends it to a server, which then sends the audio data over the internet.
[0107] Step 6:
[0108] The server passes the received voice data to the voice recognition API and converts it into text data. The text data obtained from the voice recognition API is analyzed.
[0109] Step 7:
[0110] The server references the user profile and uses a filtering algorithm to select the best recipes, based on the user's preferences and allergies.
[0111] Step 8:
[0112] The selected recipe information is passed to the speech synthesis API and converted into voice data, which is then returned to the server.
[0113] Step 9:
[0114] The server sends the audio data to the terminal. The audio data is sent via the Internet.
[0115] Step 10:
[0116] The terminal plays the audio data to provide information to the user. The audio data is played through the terminal's speaker.
[0117] 3. Recipe posting function to the community
[0118] Program Overview:
[0119] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0120] Processing Steps:
[0121] Step 1:
[0122] The user inputs new cooking recipe information into the terminal, for example, "original cookie recipe."
[0123] Step 2:
[0124] The terminal sends the input recipe information to the server, and the input data is sent to the server via the Internet.
[0125] Step 3:
[0126] The server passes the received recipe information to the generation AI model for refinement. The generation AI analyzes the recipe information and generates a more attractive title and description.
[0127] Step 4:
[0128] The server saves the generated submission page data, including the title and description, in a database. The generated submission page is then prepared for public viewing.
[0129] Step 5:
[0130] The server generates a URL for publishing and configures the publishing settings. The submission page will then be viewable by the general public.
[0131] Step 6:
[0132] General users can view published recipe pages through their devices. They can also view the posting page using the device's browser.
[0133] 4. Data refinement based on user feedback
[0134] Program Overview:
[0135] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0136] Processing Steps:
[0137] Step 1:
[0138] The user enters feedback about the recipe they tried into the device, for example, "This recipe was bland."
[0139] Step 2:
[0140] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[0141] Step 3:
[0142] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[0143] Step 4:
[0144] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[0145] Step 5:
[0146] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[0147] Step 6:
[0148] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[0149] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is capable of always providing the latest, high-quality recipe information while reflecting the user's preferences and community participation.
[0150] Example 1
[0151] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0152] Conventional cooking assistant systems are limited to analyzing users' voice input and providing cooking recipes, but lack the ability to reflect users' preferences and allergy information or filter recipes based on user feedback. Furthermore, they lack the ability to automatically refine recipe posts when users post new recipes. This has made improving the user experience a challenge.
[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0154] In this invention, the server includes means for capturing a user's voice, means for converting the voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for collecting user feedback and storing it in a database, means for excluding low-rated recipes from the training data, means for polishing the recipe entered by the user using a generative AI model, and means for inputting prompt sentences into the generative AI model. This makes it possible to provide optimal recipes that reflect the user's preferences and feedback, and to automatically polish new recipes.
[0155] A "means for capturing voice" is a device or method for capturing a user's spoken voice as digital data.
[0156] The "means for converting voice data into text data" is a process for converting captured voice data into text information using voice recognition technology.
[0157] "Means for analyzing text data and selecting cooking recipes that meet user requests" refers to a method that uses natural language processing technology to understand user requests from text data and select the optimal cooking recipe to meet those requests.
[0158] The "means for converting the selected recipe into voice data" is a method for converting text-format recipe information into a format that can be output as voice using voice synthesis technology.
[0159] A "means for providing audio data to a user" is a method for transmitting audio data to a user's device and playing it back.
[0160] The "means for collecting user feedback and storing it in a database" refers to a method for collecting, organizing, and storing the ratings and comments provided by users in a database.
[0161] "Means for excluding low-rated recipes from learning data" refers to the process of analyzing user rating data and removing recipes that have received a certain level of low rating from the system's recommendation candidates.
[0162] "Means for brushing up recipes entered by users using generative AI models" refers to a method of improving recipe information provided by users using generative AI technology to make the content more appealing.
[0163] A "means for inputting prompt text into a generative AI model" is a method for inputting text containing specific instructions or information into a generative AI model and obtaining output based on that text.
[0164] MODE FOR CARRYING OUT THE INVENTION
[0165] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[0166] 1. System Configuration
[0167] This system consists of a device used by the user and a central server. The device has a microphone to capture the user's voice and a speaker to play back the voice data. The server uses a speech recognition API, a speech synthesis API, a generative AI model, and a database to select and respond to the user's request with the optimal recipe.
[0168] 2. Implementing a Voice-Based Interactive UI
[0169] When a user asks, "What would you like for dinner today?", the device captures the voice and sends the voice data to the server. The server uses a speech recognition API such as Amazon Transcribe to convert the voice data into text data. The converted text data is then analyzed to select a cooking recipe that meets the user's request. The results of this selection are then prepared as text data and converted into voice data using a speech synthesis API such as Amazon Polly. The voice data is then sent to the device and provided to the user through the speaker.
[0170] Examples:
[0171] When a user asks, "What's your recommended dish for tonight?" the system will respond aloud with, "How about spaghetti arrabiata? It's easy and delicious."
[0172] 3. Recipe selection that reflects user preferences
[0173] Users enter their food preferences and allergy information into their device and send it to the server. The server stores this information in a database and manages it as a user profile. When a user requests recipe suggestions, the server references this user profile and selects the most suitable recipe.
[0174] Examples:
[0175] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a nut-free vegetarian curry."
[0176] 4. Automatically improve recipe posts
[0177] When a user enters a new recipe into their device and posts it, the server passes the information to a generative AI model to generate an appealing title and description, and the resulting post page is stored in a database and made available for other users to view.
[0178] Examples:
[0179] When a user enters and posts a "new pancake recipe," the system automatically generates a title such as "Easy recipe for fluffy pancakes" and an introductory text, creating a posting page. For example, the prompt for the generative AI model could be "Generate an attractive cooking recipe posting page based on the following content:"
[0180] 5. Data refinement based on user feedback
[0181] The feedback provided by the user is sent from the device to the server and stored in a database. The server periodically aggregates the feedback data and removes low-rated recipes from the learning data. This allows the system to provide only highly rated recipes to users in the future.
[0182] Examples:
[0183] If a user provides feedback such as "this recipe was bland," the server will filter out that recipe if similar feedback continues, improving the accuracy of the system's recipe selection algorithm.
[0184] As described above, the voice-activated cooking assistant system of the present invention provides optimal recipes for users by combining voice input analysis, user profile utilization, recipe refinement using generative AI models, and feedback collection and analysis, thereby enabling users to enjoy a more comfortable and convenient cooking experience.
[0185] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0186] Step 1:
[0187] The user provides voice input. The user says, "What would you like for dinner tonight?"
[0188] Step 2:
[0189] The device captures voice data and sends it to the server. The voice data is obtained in digital form through a microphone and sent to the server using an HTTP request. The input is the user's voice and the output is digital voice data sent to the server.
[0190] Step 3:
[0191] The server receives the voice data and passes it to the voice recognition API. The voice data is stored in cloud storage, and its URL is used as input for the voice recognition API. The input is the URL of the voice data, and the output is text data.
[0192] Step 4:
[0193] The server converts the audio data into text data using a speech recognition API. The text is extracted from the JSON format data returned as an API response. The input is the URL of the audio data and the API response, and the output is the extracted text data.
[0194] Step 5:
[0195] The server analyzes the text data, understands the user's request, and selects the most suitable recipe. It uses natural language processing technology to extract the "dish name" and "conditions" from the text data, and queries the database to search for matching recipes. The input is text data, and the output is information about the selected recipe.
[0196] Step 6:
[0197] The server prepares the selected recipe as text data, structuring it as "Spaghetti Arrabbiata is recommended." The input is the selected recipe information, and the output is the response text data.
[0198] Step 7:
[0199] The server uses a speech synthesis API such as Amazon Polly to convert text data into speech data. The request to the API includes text data and speech settings. The input is text data, and the output is the generated speech data.
[0200] Step 8:
[0201] The server sends the generated voice data to the device and provides it to the user through the device's speaker. The voice data is sent to the device as an HTTP response, and the device plays the data. The input is the voice data, and the output is the voice response provided to the user.
[0202] Step 9:
[0203] The user inputs their food preferences and allergy information into the device, which then sends it to the server. The input information is sent in JSON format. The input is the user's preferences and allergy information, and the output is the user profile data sent to the server.
[0204] Step 10:
[0205] The server stores the received information in a database. It associates the information with the user ID and adds it to the corresponding record in the database. The input is the user profile data, and the output is the user information stored in the database.
[0206] Step 11:
[0207] The user inputs a new cooking recipe and the device sends the information to the server. The input information is sent in JSON format. The input is the cooking recipe entered by the user, and the output is the recipe data sent to the server.
[0208] Step 12:
[0209] The server inputs the received recipe information into the generative AI model. The generative AI model then inputs text data containing specific prompts. The input is the user's recipe information and prompt, and the output is the generated title and description.
[0210] Step 13:
[0211] The server saves the generated post, including the title and description, in a database where it can be viewed by other users. The input is the generated post data, and the output is the post saved in the database.
[0212] Step 14:
[0213] The user inputs feedback for the recipes they have tried, and the terminal sends this information to the server. The input is the user's feedback, and the output is the feedback data sent to the server.
[0214] Step 15:
[0215] The server saves the feedback data in a database. It saves the rating scores and comments in the appropriate fields. The input is the feedback data, and the output is the rating information saved in the database.
[0216] Step 16:
[0217] The server periodically aggregates the feedback data and removes poorly rated recipes from the training data. It then runs a data analysis process to identify and remove poorly rated recipes. The input is the feedback data, and the output is updated training data.
[0218] The above are the specific processing steps of the program of this system.
[0219] (Application example 1)
[0220] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0221] When users use food delivery services in their daily lives, they face challenges such as the time it takes to select the optimal menu and the need to individually confirm preferences and allergy information. Furthermore, when users select a delivery menu, they are unable to make the optimal selection based on their past order history and preferences. Recipe posting and rating functions for sharing with other users are also time-consuming and labor-intensive, as they must be done manually.
[0222] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0223] In this invention, the server includes a means for capturing the user's voice, a means for converting the voice data into text data, and a means for analyzing the text data and selecting a recipe or delivery menu item according to the user's request. This allows the user to quickly request a delivery menu item by voice, streamlining the ordering process. The server also includes a function for storing user preferences and allergy information in a database and recommending delivery menu items based on that information, enabling optimal suggestions for each individual user. Furthermore, by adding a function for refining and automatically generating recipes entered by the user, users can efficiently post and rate recipes for sharing with other users.
[0224] "User" refers to a person who uses the system to make a voice request and receive a cooking recipe or delivery menu.
[0225] "Means for capturing audio" refers to a mechanism for collecting the user's audio using a device such as a microphone.
[0226] "Voice data" refers to data that expresses the user's voice as digital information.
[0227] "Means for converting into text data" refers to a mechanism for converting voice data into text information using voice recognition technology.
[0228] "Text data" refers to digital information that has been converted from audio data into text form.
[0229] "User request" refers to a request or question uttered by a user.
[0230] "Cooking recipes or delivery menus" refers to the food preparation methods and delivery options suggested by the system.
[0231] "Means for selection" refers to a mechanism for selecting the most suitable cooking recipe or delivery menu based on the user's request.
[0232] "Means for converting into voice data" refers to a mechanism that uses voice synthesis technology to convey the selected recipe or delivery menu to the user in voice form.
[0233] "Means of providing" refers to devices such as speakers and headphones used to deliver audio data to users.
[0234] "User's food preferences and allergy information" refers to information about dietary preferences and ingredients to avoid that the user has registered in the system.
[0235] A "database" refers to a system for centrally managing user information and data processed by the system.
[0236] "Recommendation means" refers to a mechanism for suggesting optimal recipes or delivery menus by taking into consideration the user's food preferences and allergy information.
[0237] "Inputted recipe" refers to information on food cooking methods and ingredients that a user has registered or provided to the system.
[0238] "Means of polishing and automatically generating a post page" refers to a mechanism that uses a generative AI model to format information entered by a user into an attractive form and automatically convert it into a publishable format.
[0239] To implement this invention, it is necessary to build a system that proposes optimal cooking recipes or delivery menus based on a user's voice request. This system is configured around a voice-interactive user interface.
[0240] System configuration
[0241] 1. How to capture the user's voice:
[0242] Hardware: A smartphone or tablet with a built-in microphone.
[0243] Role: Collects user voice in real time.
[0244] 2. How to convert audio data to text data:
[0245] Software: speech_recognition library.
[0246] Role: Converts voice data into text.
[0247] 3. A method for analyzing text data and selecting cooking recipes or delivery menus according to user requests:
[0248] Technologies used: Natural Language Processing (NLP) and databases.
[0249] Role: Analyzes user requests and selects the most suitable cooking recipe or delivery menu.
[0250] 4. Means for converting selected information into audio data:
[0251] Software: pyttsx3 library.
[0252] Role: Converts selected recipes or delivery menus into audio data.
[0253] 5. Means of providing audio data to the user:
[0254] Hardware: Smartphone speakers and headphones.
[0255] Role: Communicates information to the user audibly.
[0256] 6. A means to store user's food preferences and allergy information in a database and make recommendations:
[0257] Database: A specialized user profile database.
[0258] Role: Stores user preferences and allergy information to improve future recommendations.
[0259] 7. How to polish a recipe entered by a user and automatically generate a posting page:
[0260] Technology used: Content generation using generative AI models.
[0261] Role: To polish recipe information provided by users to make it more appealing and publish it as an automatically generated posting page.
[0262] Program operation explanation
[0263] The server captures voice data input by the user via a smartphone or tablet and converts the voice data into text using the speech_recognition library. It then analyzes the text data using natural language processing technology to select the optimal recipe or delivery menu item for the user's request. This selection process takes into account the user's preferences and allergies. The selected information is then converted back into voice data using the pyttsx3 library and delivered to the user via the smartphone's speakers or headphones.
[0264] Additionally, when a user enters a new recipe, the recipe information is refined using a generative AI model and automatically generated into an attractive posting page, which is then saved in a database and made publicly available for other users to view.
[0265] Specific examples
[0266] When a user speaks to their smartphone, "What delivery do you recommend for dinner tonight?", the system responds by voice, "Would you like pizza or sushi? Which would you like to order?" If the user selects pizza, the system will suggest the best pizza delivery service.
[0267] Prompt Sentence Examples
[0268] "The following input should contain a list of restaurants that offer pizza and sushi delivery. If the user selects pizza, output the names of the top 5 restaurants listed.
[0269] Input: I want pizza.
[0270] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0271] Step 1:
[0272] The user issues a voice request to the smartphone. For example, say, "What delivery do you recommend for dinner tonight?" The voice is captured by the smartphone's microphone.
[0273] Step 2:
[0274] The device sends the captured audio data to the server. The input is audio data, which is then sent to the server.
[0275] Step 3:
[0276] The server converts the received voice data into text data using the speech_recognition library. The input is voice data, and the output is the converted text data. This conversion uses speech recognition technology.
[0277] Step 4:
[0278] The server analyzes the text data using natural language processing (NLP) technology to understand the user's request. The input is text data, and the output is the analyzed request data. As a concrete example, it identifies that the user's request is for "dinner delivery."
[0279] Step 5:
[0280] The server retrieves the user's preferences and allergy information from the database, compares it with the analysis results, and recommends the most suitable recipe or delivery menu. The input is the analyzed request content and user profile data, and the output is the selected recipe or delivery menu.
[0281] Step 6:
[0282] The server converts the selected recipe or delivery menu into voice data using the pyttsx3 library. The input is text-based recipe or menu information, and the output is voice data.
[0283] Step 7:
[0284] The server sends the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal.
[0285] Step 8:
[0286] The device then provides the received voice data to the user through a speaker or headphones. The input is voice data, and the output is audible audio for the user. Specifically, the smartphone responds with a voice message asking, "Would you like pizza or sushi? Which would you like to order?"
[0287] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0288] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0289] 1. Implementing a voice-activated UI
[0290] Program Overview:
[0291] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0292] Processing Description:
[0293] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[0294] Examples:
[0295] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[0296] 2. Register user preferences and introduce the most suitable menu
[0297] Program Overview:
[0298] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0299] Processing Description:
[0300] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[0301] Examples:
[0302] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[0303] 3. Recipe posting function to the community
[0304] Program Overview:
[0305] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0306] Processing Description:
[0307] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[0308] Examples:
[0309] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[0310] 4. Data refinement based on user feedback
[0311] Program Overview:
[0312] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0313] Processing Description:
[0314] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[0315] Examples:
[0316] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[0317] 5. Recognizing and responding to user emotions using an emotion engine
[0318] Program Overview:
[0319] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[0320] Processing Description:
[0321] When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server converts the voice data into text using a speech recognition API and then analyzes the user's emotions using an emotion engine. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking recipe that best suits the user's emotions is selected. In addition, emotional information is added to the feedback data to help with future recipe suggestions.
[0322] Examples:
[0323] If a user requests in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine will recognize the user's tired emotions and suggest a simple omelet that requires little effort to make."
[0324] Through the above process, the system of the present invention can provide cooking recipes through voice dialogue, always providing the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0325] The processing flow will be explained below.
[0326] Recognizing and responding to user emotions using an emotion engine
[0327] Program Overview:
[0328] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[0329] Processing Steps:
[0330] Step 1:
[0331] The user speaks into the device to make a cooking request, such as, "I'm tired. Do you have a quick dinner?"
[0332] Step 2:
[0333] The device captures the user's voice. The device's microphone is used to obtain the voice data.
[0334] Step 3:
[0335] The device sends the captured audio data to a server, which then transmits the audio data over the Internet.
[0336] Step 4:
[0337] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[0338] Step 5:
[0339] The server passes the text data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the text data and the intonation and tone of the voice to recognize the user's emotions.
[0340] Step 6:
[0341] The server reflects the analyzed emotional information in the user profile, which is then added to the user profile and used to select the next recipe.
[0342] Step 7:
[0343] The server selects the optimal recipe based on the user's request and emotional information. The selection algorithm takes the user's emotional information into account.
[0344] Step 8:
[0345] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[0346] Step 9:
[0347] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[0348] Step 10:
[0349] The server transmits the generated voice data to the terminal, which then transmits the voice data via the Internet.
[0350] Step 11:
[0351] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[0352] Step 12:
[0353] The user tries the recipe and enters feedback into the device, for example, "This recipe was very easy and delicious."
[0354] Step 13:
[0355] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[0356] Step 14:
[0357] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[0358] Step 15:
[0359] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[0360] Step 16:
[0361] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[0362] Step 17:
[0363] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[0364] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is able to provide the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0365] Example 2
[0366] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0367] Conventional cooking assistant systems have difficulty reflecting users' preferences and allergy information when providing appropriate cooking instructions in response to a user's voice request. They also lack the ability to post the cooking instructions entered by the user in an appealing format, or the ability to improve the system's accuracy based on user feedback. Furthermore, they are unable to suggest recipes that take the user's emotions into account, making it difficult to provide personalized services.
[0368] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's voice, means for converting voice data into text data, means for analyzing the text data and selecting cooking procedures according to the user's request, means for converting the selected cooking procedures into voice data, means for providing the voice data to the user, means for storing the user's cooking preferences and allergy information in a database and recommending cooking procedures based on the stored information, means for improving the cooking procedures entered by the user using a generative AI model and automatically generating a posting page, means for collecting user feedback, updating the ratings of the cooking procedures, and excluding low-rated cooking procedures from the system's learning data, and means for recognizing emotions from the user's voice and suggesting optimal cooking procedures based on the emotions, and means for adding user emotion information to the feedback data. This makes it possible to always provide the latest, high-quality cooking procedures taking into account the user's preferences and emotions.
[0369] "User" refers to any individual or group of people who operate the system and use the cooking recipe suggestions and voice interface.
[0370] "Voice capture means" refers to a device or software mechanism for capturing a user's speech as digital data.
[0371] "Voice data" refers to data that is a digital representation of a user's voice.
[0372] "Text data" refers to data in the form of a string of characters obtained by analyzing voice data.
[0373] "Means for analyzing text data" refers to software or algorithms that analyze text data converted from speech and understand the user's intent and request.
[0374] A "cooking recipe" refers to a document that lists the steps or methods for preparing a dish.
[0375] "Means for converting into voice data" refers to software or devices that convert text data back into voice format using voice synthesis technology.
[0376] A "database" refers to a system for efficiently storing and managing various data such as user preferences and allergy information.
[0377] "Recommendation tools" refers to algorithms and software functions that select optimal cooking instructions based on a user's profile information and past data.
[0378] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generates text, images, etc. based on user input.
[0379] "Posting Page" refers to a web page or part of an application where a user can share and publish cooking instructions.
[0380] "Feedback" refers to opinions such as ratings and comments provided by users to the system.
[0381] An "emotion engine" refers to a technology or software component that analyzes emotions from a user's voice or text and understands the user's emotional state.
[0382] "Information processing device" refers to the entire system including a computer and peripheral devices for inputting, processing, and outputting data.
[0383] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking procedures, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0384] Implementing a voice-interactive UI
[0385] The program for this system primarily uses speech recognition and speech synthesis APIs to build a mechanism for analyzing voice input from users in real time and generating appropriate responses. Specifically, it uses the Google Cloud Speech-to-Text API as the speech recognition API and the Google Cloud Text-to-Speech API as the speech synthesis API. When a user asks a question out loud, such as "What would you like for dinner tonight?", the device captures this speech and sends it to the server as audio data. The server passes the received audio data to the speech recognition API, converts it into text data, and analyzes the user's request. After selecting the appropriate cooking instructions, it converts it into audio data using the speech synthesis API and sends it to the device, where it provides the user with a spoken response.
[0386] Register user preferences and introduce the most suitable menu
[0387] The system implements an algorithm that recommends cooking steps based on a database of cooking preferences and allergy information preset by the user. The user enters their preferences and allergy information into the device and sends it to the server. The server stores the received information in a database and creates a user profile. When the user makes a request for cooking suggestions, the server refers to the user profile and selects appropriate cooking steps. These are then provided to the user via voice or text.
[0388] Recipe posting function to the community
[0389] We provide a mechanism to automatically generate posting pages using a generative AI model so that users can easily post cooking instructions. When a user enters new cooking instructions, the device sends the information to a server. The server then uses the generative AI model to refine the received instructions and automatically generate an attractive title and description. The generated posting page is saved in a database and made publicly available for other users to view.
[0390] Data refinement based on user feedback
[0391] We implement a mechanism to collect feedback from users, update the ratings of cooking steps, and exclude low-rated steps from the system's learning data. When a user enters feedback for a cooking step they have tried, the device sends the information to a server. The server stores the feedback data in a database and periodically aggregates and analyzes it. By excluding low-rated steps from the learning data in this way, the accuracy of the algorithm is improved.
[0392] Recognizing and responding to user emotions using an emotion engine
[0393] The system recognizes emotions from the user's voice and suggests optimal cooking steps based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy. When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server uses a speech recognition API to convert it into text data and then uses an emotion engine to analyze the user's emotions. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking steps that best suit the user's emotions are selected. For example, if a user requests, in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine recognizes the user's fatigue and suggests a simple omelet that requires little effort to make."
[0394] Through the above process, the system of the present invention can provide cooking instructions through voice interaction, always providing the latest, high-quality cooking instructions while reflecting the user's preferences and emotions, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0395] Prompt Sentence Examples
[0396] When a user asks, "What's your recommended dish for tonight?", the system responds, "How about spaghetti arrabiata? It's easy and delicious." In this way, the system smoothly suggests the appropriate cooking procedure for a specific situation.
[0397] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0398] Implementing a voice-interactive UI
[0399] (Processing flow)
[0400] Step 1: User makes a voice request
[0401] Step 2: The device captures the audio data and sends it to the server
[0402] Step 3: The server uses a speech recognition API to convert the data into text.
[0403] Step 4: The server analyzes the request
[0404] Step 5: The server selects the cooking instructions based on the request.
[0405] Step 6: The server converts the selected recipe into text data.
[0406] Step 7: The server uses the speech synthesis API to convert the data into audio.
[0407] Step 8: Respond to the user by sending voice data to the device
[0408] (Specific explanation)
[0409] Step 1:
[0410] The user makes a voice request. For example, the user says, "What would you like for dinner tonight?"
[0411] Step 2:
[0412] The device captures the voice data and sends it to the server. Specifically, it uses the device's microphone to capture the user's voice as digital data. The device then sends the captured voice data to the server via the Internet. (Input) Voice data (Output) Data sent to the server
[0413] Step 3:
[0414] The server converts the received voice data into text data using a speech recognition API. For example, the server analyzes the received voice data using a speech recognition API such as Google Cloud Speech-to-Text API. This converts the voice into text data. (Input) Voice data (Output) Text data
[0415] Step 4:
[0416] The server analyzes the request. The server passes the text data to a natural language processing engine for analysis and understands the user's request. For example, it extracts keywords such as "dinner," "recommendations," and "tell me." (Input) Text data (Output) Analysis results
[0417] Step 5:
[0418] The server selects the cooking procedure according to the request. Based on the analysis results, the server refers to the user profile and database to select the appropriate cooking procedure. For example, "Spaghetti Arrabbiata" is selected. (Input) Analysis results (Output) Selected cooking procedure
[0419] Step 6:
[0420] The server converts the selected cooking instructions into text data. The selected cooking instructions are formatted for presentation to the user. For example, a sentence such as "Today's recommended dish is spaghetti arrabbiata" is generated. (Input) Selected cooking instructions (Output) Text data
[0421] Step 7:
[0422] The server converts the text data into audio data using a speech synthesis API. The server converts the text data into audio data using a speech synthesis API such as Google Cloud Text-to-Speech API. (Input) Text data (Output) Audio data
[0423] Step 8:
[0424] The server responds to the user by transmitting the generated voice data to the terminal, and the terminal plays back the received voice data to provide a response to the user.
[0425] (Input) Voice data (Output) Voice response provided to the user
[0426] Register user preferences and introduce the most suitable menu
[0427] (Processing flow)
[0428] Step 1: User registers preferences and allergy information
[0429] Step 2: The device sends the information to the server
[0430] Step 3: The server saves the information to a database
[0431] Step 4: User requests dish suggestions
[0432] Step 5: The server looks up the user profile and selects the best recipe
[0433] Step 6: The server provides the recipe via voice or text
[0434] (Specific explanation)
[0435] Step 1:
[0436] The user registers their preferences and allergy information, for example, by typing "I'm a vegetarian and have a nut allergy" into the terminal.
[0437] Step 2:
[0438] The device sends information to the server. Specifically, the device converts the input information into data packets and sends them to the server. (Input) Preference and allergy information (Output) Data sent to the server
[0439] Step 3:
[0440] The server saves the information in a database. The server analyzes the received user information and saves it in a database as a user profile. For example, it classifies the information by adding tags such as "vegetarian" or "nut allergy." (Input) Data sent (Output) User profile saved in the database
[0441] Step 4:
[0442] The user requests food suggestions, for example, "Tell me what's good for dinner today."
[0443] Step 5:
[0444] The server refers to the user profile and selects the most suitable cooking procedure. For example, "Select vegetarian curry without nuts." (Input) User profile (Output) Selected cooking procedure
[0445] Step 6:
[0446] The server provides cooking instructions by voice or text. The server generates the selected cooking instructions as text data and converts it into voice data using a speech synthesis API. It is played on the device. (Input) Selected cooking instructions (Output) Voice response provided to the user
[0447] Recipe posting function to the community
[0448] (Processing flow)
[0449] Step 1: User enters new cooking instructions
[0450] Step 2: The device sends the recipe information to the server
[0451] Step 3: The server refines the recipe information using the generated AI model
[0452] Step 4: The server automatically generates the submission page
[0453] Step 5: Save the post to the database and publish it
[0454] (Specific explanation)
[0455] Step 1:
[0456] A user inputs new cooking instructions, for example, "new pancake recipe" into the terminal and clicks the submit button.
[0457] Step 2:
[0458] The terminal sends recipe information to the server. The terminal sends the input recipe information to the server as a data packet. (Input) Cooking procedure information (Output) Data sent to the server
[0459] Step 3:
[0460] The server refines the recipe information using a generative AI model. The server analyzes the received step-by-step information using the generative AI model and generates an appealing title and description. For example, it generates a title such as "Easy recipe for fluffy pancakes." (Input) Received step-by-step information (Output) Generated title and description
[0461] Step 4:
[0462] The server automatically generates a posting page. Based on the generated title and description, the server automatically generates a recipe posting page. (Input) Generated title and description (Output) Posting page
[0463] Step 5:
[0464] Save the submitted page in the database and make it public. Save the generated submitted page in the database and make it public. Display it in the web interface so that other users can view it. (Input) Submitted page (Output) Saved in the database and published
[0465] Data refinement based on user feedback
[0466] (Processing flow)
[0467] Step 1: User provides feedback on cooking instructions
[0468] Step 2: The device sends the feedback information to the server
[0469] Step 3: The server stores the feedback data in a database
[0470] Step 4: The server periodically aggregates and analyzes the feedback.
[0471] Step 5: The server removes cooking instructions with low ratings from the training data
[0472] (Specific explanation)
[0473] Step 1:
[0474] The user inputs feedback on the cooking procedure, for example, "This recipe was bland" into the terminal.
[0475] Step 2:
[0476] The terminal sends feedback information to the server. The terminal converts the input feedback information into data packets and sends them to the server. (Input) Feedback information (Output) Data sent to the server
[0477] Step 3:
[0478] The server saves the feedback data in a database. For example, it saves data such as "Restaurant Egg Curry", "Bland taste", and "Rating: 2". (Input) Feedback information sent. (Output) Feedback saved in the database.
[0479] Step 4:
[0480] The server periodically collects and analyzes feedback. Feedback data is periodically collected, and the ratings from many users are analyzed to determine which recipes are highly rated and which are poorly rated. (Input) Feedback stored in the database (Output) Analysis results
[0481] Step 5:
[0482] The server removes cooking steps with low ratings from the training data. By removing recipes with low ratings based on the aggregation results from the training data, the system improves its suggestion accuracy. (Input) Analysis results (Output) Updated training data
[0483] Recognizing and responding to user emotions using an emotion engine
[0484] (Processing flow)
[0485] Step 1: User makes a voice request
[0486] Step 2: The device captures the audio data and sends it to the server
[0487] Step 3: The server uses a speech recognition API to convert the data into text.
[0488] Step 4: The server analyzes the emotion information using the emotion engine
[0489] Step 5: The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal recipe.
[0490] Step 6: The server adds emotion information to the feedback data
[0491] (Specific explanation)
[0492] Step 1:
[0493] The user makes a request by voice. For example, the user may request in a tired voice, "I want a quick dinner."
[0494] Step 2:
[0495] The device captures the audio data and sends it to the server. The device captures the audio data and sends it to the server. (Input) Audio data (Output) Data sent to the server
[0496] Step 3:
[0497] The server uses a speech recognition API to convert the voice data into text data. The server uses a speech recognition API to convert the voice data into text data. (Input) Voice data (Output) Text data
[0498] Step 4:
[0499] The server uses an emotion engine to analyze emotional information. The server uses an emotion engine to analyze emotional information from the tone and pace of the user's voice. For example, emotional information indicating "fatigue" is extracted. (Input) Text data (Output) Emotional information
[0500] Step 5:
[0501] The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal cooking procedure. The server selects the optimal recipe for the user based on the emotional information. For example, it selects a "simple omelet" that can be made with minimal effort. (Input) Emotional information (Output) Selected cooking procedure
[0502] Step 6:
[0503] The server adds the emotional information to the feedback data. The emotional information is saved as feedback data along with the suggested cooking steps, and will be used for future suggestions. (Input) Emotional information (Output) Feedback data
[0504] (Application example 2)
[0505] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0506] Conventional recipe provision systems suggest recipes by analyzing users' voice requests, but they have the problem of being unable to provide recipes that reflect the user's emotions or the situation of the day. This means that users cannot receive suggestions for dishes that suit their physical condition or mood, resulting in a lack of convenience. In addition, there is a lack of a mechanism for improving the accuracy of the system by reflecting user feedback.
[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0508] In this invention, the server includes means for capturing user voice, means for converting voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for analyzing the user's emotions, means for recommending recipes based on the emotion information, and means for collecting user feedback and improving the accuracy of the system. This makes it possible to provide individualized recipes according to each user's emotions, thereby realizing more convenient recipe suggestions.
[0509] The "means for capturing the user's voice" is a mechanism by which the electronic device receives the voice uttered by the user and captures it as digital data.
[0510] The "means for converting voice data into text data" is a mechanism that analyzes the captured voice data and performs a process to express the contents as a string of characters.
[0511] The "means for analyzing text data and selecting cooking recipes that meet the user's requests" refers to an algorithm and system that understands the user's requests based on the converted text data and selects appropriate cooking recipes.
[0512] The "means for converting the selected recipe into audio data" is a mechanism that performs the process of converting the selected recipe from a character string into an audio signal.
[0513] The "means for providing audio data to the user" refers to a device or program for playing back the converted audio data to the user and conveying information.
[0514] "Means for analyzing user emotions" refers to algorithms or systems that determine and analyze the user's emotional state from the tone and content of the voice.
[0515] "Means for recommending cooking recipes based on emotional information" is a system that selects and recommends the most appropriate cooking recipe based on the user's current mood and state based on the results of emotional analysis.
[0516] "Means for collecting user feedback and improving the accuracy of the system" refers to the process and mechanisms for collecting user evaluations and opinions and using them to improve the system's algorithms and database.
[0517] This invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. In particular, it aims to provide more personalized services by combining it with an emotion engine that recognizes the user's emotions.
[0518] System Configuration
[0519] Hardware
[0520] Smartphone: A device that captures the user's voice and receives the analysis results.
[0521] Server: Analyzes voice data, analyzes emotions, selects recipes, and aggregates feedback data
[0522] software
[0523] SpeechRecognition API: Converting voice data into text data
[0524] Emotion Analysis API: Analyzing user emotions from text data
[0525] Text to Speech API: Convert text data into audio data
[0526] Database: Stores user profile information, cooking recipes, and feedback data
[0527] Operation overview
[0528] 1. Audio Capture
[0529] A user speaks to their smartphone, asking, "What would you like for dinner tonight?" The smartphone captures the voice through its built-in microphone and sends this voice data to the server.
[0530] 2. Voice Recognition
[0531] When the server receives the voice data, it uses the SpeechRecognition API to convert the voice data into text data.
[0532] 3. Emotion analysis
[0533] The server then sends the converted text data to the Emotion Analysis API, which analyzes the user's emotions. The analysis result can be, for example, "I'm tired."
[0534] 4. Recipe Selection
[0535] The server selects the most suitable recipe for the user's request based on the results of sentiment analysis and the user profile database. For tired users, it selects easy-to-make dishes.
[0536] 5. Speech Synthesis
[0537] The selected cooking recipe is again treated as text data and converted into audio data using the Text to Speech API.
[0538] 6. Provide answers
[0539] The smartphone receives the converted voice data and provides the user with a spoken response such as, "You seem tired today, how about a quick sandwich?"
[0540] 7. Feedback Collection
[0541] After users try the suggested recipes, they provide feedback. The server collects this feedback data and stores it in a database. In the future, the feedback will be used to improve the accuracy of the algorithm.
[0542] Specific examples
[0543] For example, when a user makes a voice request such as "Tell me something easy to eat," the system captures the voice and the server analyzes the voice data. If the emotion engine recognizes the emotion "tired" from the user's voice, the system will suggest a quick and easy-to-prepare dish such as a "sandwich." The following is an example of a prompt sentence that can be input to the generative AI model:
[0544] Convert the user's voice into text, feed that text into a sentiment analysis engine, select the sentiment, and output a food recommendation accordingly.
[0545] Input: I'm tired today and want a quick meal.
[0546] Emotions: Tired
[0547] Output: You seem tired, would you like some lunch?
[0548] These procedures provide personalized information in response to user requests, improving user convenience.
[0549] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0550] Step 1:
[0551] The user speaks a question into the smartphone. The user's voice data is captured by the smartphone's microphone. The input is the user's voice, and the output is the captured voice data. The voice data is converted into a digital format and sent to the server.
[0552] Step 2:
[0553] The server receives the audio data. The input is digital audio data, and the output is text data. The server converts the audio data into text data using the SpeechRecognition API. In this process, the audio content is converted into a string of characters.
[0554] Step 3:
[0555] The server sends the text data to the Emotion Analysis API and analyzes the user's emotions. The input is text data and the output is emotional information. The server obtains the analysis result, for example, an emotion such as "tired."
[0556] Step 4:
[0557] The server selects the optimal recipe for the user's request based on the emotional information and the user profile database. The input is the emotional analysis results and the user profile, and the output is the selected recipe. In particular, if the user is tired, an easy-to-make dish is selected.
[0558] Step 5:
[0559] The server converts the selected recipe into audio data using the Text to Speech API. The input is the text data of the selected recipe, and the output is audio data. In this process, the recipe in text format is converted back into audio format.
[0560] Step 6:
[0561] The server sends the voice data to the smartphone. The smartphone receives this data and suggests recipes to the user by voice. The input is the converted voice data, and the output is the voice output to the user.
[0562] Step 7:
[0563] After the user has tried the suggested recipe, they input their feedback into their smartphone. The input is the user's feedback, and the output is the collected feedback data. The feedback data is sent to the server and stored in a database.
[0564] Step 8:
[0565] The server analyzes the feedback data and updates the algorithm to improve the system's accuracy. The input is the accumulated feedback data, and the output is an improved recommendation algorithm. This will enable more personalized and accurate recipe suggestions in the future.
[0566] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0567] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0568] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0569] [Second embodiment]
[0570] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0571] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0572] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0573] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0574] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0575] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0576] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0577] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0578] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0579] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0580] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0581] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0582] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[0583] 1. Implementing a voice-activated UI
[0584] Program Overview:
[0585] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0586] Processing Description:
[0587] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[0588] Examples:
[0589] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[0590] 2. Register user preferences and introduce the most suitable menu
[0591] Program Overview:
[0592] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0593] Processing Description:
[0594] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[0595] Examples:
[0596] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[0597] 3. Recipe posting function to the community
[0598] Program Overview:
[0599] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0600] Processing Description:
[0601] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[0602] Examples:
[0603] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[0604] 4. Data refinement based on user feedback
[0605] Program Overview:
[0606] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0607] Processing Description:
[0608] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[0609] Examples:
[0610] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[0611] In this way, the system of the present invention provides cooking recipes based on voice interaction and reflects user preferences and community participation, allowing users to enjoy a more comfortable and convenient cooking experience.
[0612] The processing flow will be explained below.
[0613] 1. Implementing a voice-activated UI
[0614] Program Overview:
[0615] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0616] Processing Steps:
[0617] Step 1:
[0618] Users speak into the device to ask questions or make requests about food, such as "What would you like for dinner tonight?"
[0619] Step 2:
[0620] The device captures the user's voice. The device uses a built-in microphone to obtain the voice data.
[0621] Step 3:
[0622] The device transmits the captured audio data to the server, and the device transmits the audio data via the Internet.
[0623] Step 4:
[0624] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[0625] Step 5:
[0626] The server analyzes the text data to understand the user's intent, and then searches for appropriate recipe information based on the analyzed data.
[0627] Step 6:
[0628] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[0629] Step 7:
[0630] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[0631] Step 8:
[0632] The server transmits the generated voice data to the terminal, and the voice data is transmitted via the Internet.
[0633] Step 9:
[0634] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[0635] 2. Register user preferences and introduce the most suitable menu
[0636] Program Overview:
[0637] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0638] Processing Steps:
[0639] Step 1:
[0640] The user inputs their food preferences and allergy information through the terminal, such as "vegetarian" or "nut allergy."
[0641] Step 2:
[0642] The device sends the input preference information to the server, and the input data is sent to the server via the Internet.
[0643] Step 3:
[0644] The server stores the received information in a database as a user profile. Profile data is created and stored for each user.
[0645] Step 4:
[0646] The user inputs a specific food request into the device, for example, asking "What do you recommend for lunch today?"
[0647] Step 5:
[0648] The device captures the audio and sends it to a server, which then sends the audio data over the internet.
[0649] Step 6:
[0650] The server passes the received voice data to the voice recognition API and converts it into text data. The text data obtained from the voice recognition API is analyzed.
[0651] Step 7:
[0652] The server references the user profile and uses a filtering algorithm to select the best recipes, based on the user's preferences and allergies.
[0653] Step 8:
[0654] The selected recipe information is passed to the speech synthesis API and converted into voice data, which is then returned to the server.
[0655] Step 9:
[0656] The server sends the audio data to the terminal. The audio data is sent via the Internet.
[0657] Step 10:
[0658] The terminal plays the audio data to provide information to the user. The audio data is played through the terminal's speaker.
[0659] 3. Recipe posting function to the community
[0660] Program Overview:
[0661] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0662] Processing Steps:
[0663] Step 1:
[0664] The user inputs new cooking recipe information into the terminal, for example, "original cookie recipe."
[0665] Step 2:
[0666] The terminal sends the input recipe information to the server, and the input data is sent to the server via the Internet.
[0667] Step 3:
[0668] The server passes the received recipe information to the generation AI model for refinement. The generation AI analyzes the recipe information and generates a more attractive title and description.
[0669] Step 4:
[0670] The server saves the generated submission page data, including the title and description, in a database. The generated submission page is then prepared for public viewing.
[0671] Step 5:
[0672] The server generates a URL for publishing and configures the publishing settings. The submission page will then be viewable by the general public.
[0673] Step 6:
[0674] General users can view published recipe pages through their devices. They can also view the posting page using the device's browser.
[0675] 4. Data refinement based on user feedback
[0676] Program Overview:
[0677] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0678] Processing Steps:
[0679] Step 1:
[0680] The user enters feedback about the recipe they tried into the device, for example, "This recipe was bland."
[0681] Step 2:
[0682] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[0683] Step 3:
[0684] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[0685] Step 4:
[0686] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[0687] Step 5:
[0688] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[0689] Step 6:
[0690] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[0691] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is capable of always providing the latest, high-quality recipe information while reflecting the user's preferences and community participation.
[0692] Example 1
[0693] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0694] Conventional cooking assistant systems are limited to analyzing users' voice input and providing cooking recipes, but lack the ability to reflect users' preferences and allergy information or filter recipes based on user feedback. Furthermore, they lack the ability to automatically refine recipe posts when users post new recipes. This has made improving the user experience a challenge.
[0695] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0696] In this invention, the server includes means for capturing a user's voice, means for converting the voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for collecting user feedback and storing it in a database, means for excluding low-rated recipes from the training data, means for polishing the recipe entered by the user using a generative AI model, and means for inputting prompt sentences into the generative AI model. This makes it possible to provide optimal recipes that reflect the user's preferences and feedback, and to automatically polish new recipes.
[0697] A "means for capturing voice" is a device or method for capturing a user's spoken voice as digital data.
[0698] The "means for converting voice data into text data" is a process for converting captured voice data into text information using voice recognition technology.
[0699] "Means for analyzing text data and selecting cooking recipes that meet user requests" refers to a method that uses natural language processing technology to understand user requests from text data and select the optimal cooking recipe to meet those requests.
[0700] The "means for converting the selected recipe into voice data" is a method for converting text-format recipe information into a format that can be output as voice using voice synthesis technology.
[0701] A "means for providing audio data to a user" is a method for transmitting audio data to a user's device and playing it back.
[0702] The "means for collecting user feedback and storing it in a database" refers to a method for collecting, organizing, and storing the ratings and comments provided by users in a database.
[0703] "Means for excluding low-rated recipes from learning data" refers to the process of analyzing user rating data and removing recipes that have received a certain level of low rating from the system's recommendation candidates.
[0704] "Means for brushing up recipes entered by users using generative AI models" refers to a method of improving recipe information provided by users using generative AI technology to make the content more appealing.
[0705] A "means for inputting prompt text into a generative AI model" is a method for inputting text containing specific instructions or information into a generative AI model and obtaining output based on that text.
[0706] MODE FOR CARRYING OUT THE INVENTION
[0707] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[0708] 1. System Configuration
[0709] This system consists of a device used by the user and a central server. The device has a microphone to capture the user's voice and a speaker to play back the voice data. The server uses a speech recognition API, a speech synthesis API, a generative AI model, and a database to select and respond to the user's request with the optimal recipe.
[0710] 2. Implementing a Voice-Based Interactive UI
[0711] When a user asks, "What would you like for dinner today?", the device captures the voice and sends the voice data to the server. The server uses a speech recognition API such as Amazon Transcribe to convert the voice data into text data. The converted text data is then analyzed to select a cooking recipe that meets the user's request. The results of this selection are then prepared as text data and converted into voice data using a speech synthesis API such as Amazon Polly. The voice data is then sent to the device and provided to the user through the speaker.
[0712] Examples:
[0713] When a user asks, "What's your recommended dish for tonight?" the system will respond aloud with, "How about spaghetti arrabiata? It's easy and delicious."
[0714] 3. Recipe selection that reflects user preferences
[0715] Users enter their food preferences and allergy information into their device and send it to the server. The server stores this information in a database and manages it as a user profile. When a user requests recipe suggestions, the server references this user profile and selects the most suitable recipe.
[0716] Examples:
[0717] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a nut-free vegetarian curry."
[0718] 4. Automatically improve recipe posts
[0719] When a user enters a new recipe into their device and posts it, the server passes the information to a generative AI model to generate an appealing title and description, and the resulting post page is stored in a database and made available for other users to view.
[0720] Examples:
[0721] When a user enters and posts a "new pancake recipe," the system automatically generates a title such as "Easy recipe for fluffy pancakes" and an introductory text, creating a posting page. For example, the prompt for the generative AI model could be "Generate an attractive cooking recipe posting page based on the following content:"
[0722] 5. Data refinement based on user feedback
[0723] The feedback provided by the user is sent from the device to the server and stored in a database. The server periodically aggregates the feedback data and removes low-rated recipes from the learning data. This allows the system to provide only highly rated recipes to users in the future.
[0724] Examples:
[0725] If a user provides feedback such as "this recipe was bland," the server will filter out that recipe if similar feedback continues, improving the accuracy of the system's recipe selection algorithm.
[0726] As described above, the voice-activated cooking assistant system of the present invention provides optimal recipes for users by combining voice input analysis, user profile utilization, recipe refinement using generative AI models, and feedback collection and analysis, thereby enabling users to enjoy a more comfortable and convenient cooking experience.
[0727] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0728] Step 1:
[0729] The user provides voice input. The user says, "What would you like for dinner tonight?"
[0730] Step 2:
[0731] The device captures voice data and sends it to the server. The voice data is obtained in digital form through a microphone and sent to the server using an HTTP request. The input is the user's voice and the output is digital voice data sent to the server.
[0732] Step 3:
[0733] The server receives the voice data and passes it to the voice recognition API. The voice data is stored in cloud storage, and its URL is used as input for the voice recognition API. The input is the URL of the voice data, and the output is text data.
[0734] Step 4:
[0735] The server converts the audio data into text data using a speech recognition API. The text is extracted from the JSON format data returned as an API response. The input is the URL of the audio data and the API response, and the output is the extracted text data.
[0736] Step 5:
[0737] The server analyzes the text data, understands the user's request, and selects the most suitable recipe. It uses natural language processing technology to extract the "dish name" and "conditions" from the text data, and queries the database to search for matching recipes. The input is text data, and the output is information about the selected recipe.
[0738] Step 6:
[0739] The server prepares the selected recipe as text data, structuring it as "Spaghetti Arrabbiata is recommended." The input is the selected recipe information, and the output is the response text data.
[0740] Step 7:
[0741] The server uses a speech synthesis API such as Amazon Polly to convert text data into speech data. The request to the API includes text data and speech settings. The input is text data, and the output is the generated speech data.
[0742] Step 8:
[0743] The server sends the generated voice data to the device and provides it to the user through the device's speaker. The voice data is sent to the device as an HTTP response, and the device plays the data. The input is the voice data, and the output is the voice response provided to the user.
[0744] Step 9:
[0745] The user inputs their food preferences and allergy information into the device, which then sends it to the server. The input information is sent in JSON format. The input is the user's preferences and allergy information, and the output is the user profile data sent to the server.
[0746] Step 10:
[0747] The server stores the received information in a database. It associates the information with the user ID and adds it to the corresponding record in the database. The input is the user profile data, and the output is the user information stored in the database.
[0748] Step 11:
[0749] The user inputs a new cooking recipe and the device sends the information to the server. The input information is sent in JSON format. The input is the cooking recipe entered by the user, and the output is the recipe data sent to the server.
[0750] Step 12:
[0751] The server inputs the received recipe information into the generative AI model. The generative AI model then inputs text data containing specific prompts. The input is the user's recipe information and prompt, and the output is the generated title and description.
[0752] Step 13:
[0753] The server saves the generated post, including the title and description, in a database where it can be viewed by other users. The input is the generated post data, and the output is the post saved in the database.
[0754] Step 14:
[0755] The user inputs feedback for the recipes they have tried, and the terminal sends this information to the server. The input is the user's feedback, and the output is the feedback data sent to the server.
[0756] Step 15:
[0757] The server saves the feedback data in a database. It saves the rating scores and comments in the appropriate fields. The input is the feedback data, and the output is the rating information saved in the database.
[0758] Step 16:
[0759] The server periodically aggregates the feedback data and removes poorly rated recipes from the training data. It then runs a data analysis process to identify and remove poorly rated recipes. The input is the feedback data, and the output is updated training data.
[0760] The above are the specific processing steps of the program of this system.
[0761] (Application example 1)
[0762] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0763] When users use food delivery services in their daily lives, they face challenges such as the time it takes to select the optimal menu and the need to individually confirm preferences and allergy information. Furthermore, when users select a delivery menu, they are unable to make the optimal selection based on their past order history and preferences. Recipe posting and rating functions for sharing with other users are also time-consuming and labor-intensive, as they must be done manually.
[0764] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0765] In this invention, the server includes a means for capturing the user's voice, a means for converting the voice data into text data, and a means for analyzing the text data and selecting a recipe or delivery menu item according to the user's request. This allows the user to quickly request a delivery menu item by voice, streamlining the ordering process. The server also includes a function for storing user preferences and allergy information in a database and recommending delivery menu items based on that information, enabling optimal suggestions for each individual user. Furthermore, by adding a function for refining and automatically generating recipes entered by the user, users can efficiently post and rate recipes for sharing with other users.
[0766] "User" refers to a person who uses the system to make a voice request and receive a cooking recipe or delivery menu.
[0767] "Means for capturing audio" refers to a mechanism for collecting the user's audio using a device such as a microphone.
[0768] "Voice data" refers to data that expresses the user's voice as digital information.
[0769] "Means for converting into text data" refers to a mechanism for converting voice data into text information using voice recognition technology.
[0770] "Text data" refers to digital information that has been converted from audio data into text form.
[0771] "User request" refers to a request or question uttered by a user.
[0772] "Cooking recipes or delivery menus" refers to the food preparation methods and delivery options suggested by the system.
[0773] "Means for selection" refers to a mechanism for selecting the most suitable cooking recipe or delivery menu based on the user's request.
[0774] "Means for converting into voice data" refers to a mechanism that uses voice synthesis technology to convey the selected recipe or delivery menu to the user in voice form.
[0775] "Means of providing" refers to devices such as speakers and headphones used to deliver audio data to users.
[0776] "User's food preferences and allergy information" refers to information about dietary preferences and ingredients to avoid that the user has registered in the system.
[0777] A "database" refers to a system for centrally managing user information and data processed by the system.
[0778] "Recommendation means" refers to a mechanism for suggesting optimal recipes or delivery menus by taking into consideration the user's food preferences and allergy information.
[0779] "Inputted recipe" refers to information on food cooking methods and ingredients that a user has registered or provided to the system.
[0780] "Means of polishing and automatically generating a post page" refers to a mechanism that uses a generative AI model to format information entered by a user into an attractive form and automatically convert it into a publishable format.
[0781] To implement this invention, it is necessary to build a system that proposes optimal cooking recipes or delivery menus based on a user's voice request. This system is configured around a voice-interactive user interface.
[0782] System configuration
[0783] 1. How to capture the user's voice:
[0784] Hardware: A smartphone or tablet with a built-in microphone.
[0785] Role: Collects user voice in real time.
[0786] 2. How to convert audio data to text data:
[0787] Software: speech_recognition library.
[0788] Role: Converts voice data into text.
[0789] 3. A method for analyzing text data and selecting cooking recipes or delivery menus according to user requests:
[0790] Technologies used: Natural Language Processing (NLP) and databases.
[0791] Role: Analyzes user requests and selects the most suitable cooking recipe or delivery menu.
[0792] 4. Means for converting selected information into audio data:
[0793] Software: pyttsx3 library.
[0794] Role: Converts selected recipes or delivery menus into audio data.
[0795] 5. Means of providing audio data to the user:
[0796] Hardware: Smartphone speakers and headphones.
[0797] Role: Communicates information to the user audibly.
[0798] 6. A means to store user's food preferences and allergy information in a database and make recommendations:
[0799] Database: A specialized user profile database.
[0800] Role: Stores user preferences and allergy information to improve future recommendations.
[0801] 7. How to polish a recipe entered by a user and automatically generate a posting page:
[0802] Technology used: Content generation using generative AI models.
[0803] Role: To polish recipe information provided by users to make it more appealing and publish it as an automatically generated posting page.
[0804] Program operation explanation
[0805] The server captures voice data input by the user via a smartphone or tablet and converts the voice data into text using the speech_recognition library. It then analyzes the text data using natural language processing technology to select the optimal recipe or delivery menu item for the user's request. This selection process takes into account the user's preferences and allergies. The selected information is then converted back into voice data using the pyttsx3 library and delivered to the user via the smartphone's speakers or headphones.
[0806] Additionally, when a user enters a new recipe, the recipe information is refined using a generative AI model and automatically generated into an attractive posting page, which is then saved in a database and made publicly available for other users to view.
[0807] Specific examples
[0808] When a user speaks to their smartphone, "What delivery do you recommend for dinner tonight?", the system responds by voice, "Would you like pizza or sushi? Which would you like to order?" If the user selects pizza, the system will suggest the best pizza delivery service.
[0809] Prompt Sentence Examples
[0810] "The following input should contain a list of restaurants that offer pizza and sushi delivery. If the user selects pizza, output the names of the top 5 restaurants listed.
[0811] Input: I want pizza.
[0812] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0813] Step 1:
[0814] The user issues a voice request to the smartphone. For example, say, "What delivery do you recommend for dinner tonight?" The voice is captured by the smartphone's microphone.
[0815] Step 2:
[0816] The device sends the captured audio data to the server. The input is audio data, which is then sent to the server.
[0817] Step 3:
[0818] The server converts the received voice data into text data using the speech_recognition library. The input is voice data, and the output is the converted text data. This conversion uses speech recognition technology.
[0819] Step 4:
[0820] The server analyzes the text data using natural language processing (NLP) technology to understand the user's request. The input is text data, and the output is the analyzed request data. As a concrete example, it identifies that the user's request is for "dinner delivery."
[0821] Step 5:
[0822] The server retrieves the user's preferences and allergy information from the database, compares it with the analysis results, and recommends the most suitable recipe or delivery menu. The input is the analyzed request content and user profile data, and the output is the selected recipe or delivery menu.
[0823] Step 6:
[0824] The server converts the selected recipe or delivery menu into voice data using the pyttsx3 library. The input is text-based recipe or menu information, and the output is voice data.
[0825] Step 7:
[0826] The server sends the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal.
[0827] Step 8:
[0828] The device then provides the received voice data to the user through a speaker or headphones. The input is voice data, and the output is audible audio for the user. Specifically, the smartphone responds with a voice message asking, "Would you like pizza or sushi? Which would you like to order?"
[0829] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0830] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0831] 1. Implementing a voice-activated UI
[0832] Program Overview:
[0833] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[0834] Processing Description:
[0835] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[0836] Examples:
[0837] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[0838] 2. Register user preferences and introduce the most suitable menu
[0839] Program Overview:
[0840] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[0841] Processing Description:
[0842] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[0843] Examples:
[0844] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[0845] 3. Recipe posting function to the community
[0846] Program Overview:
[0847] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[0848] Processing Description:
[0849] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[0850] Examples:
[0851] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[0852] 4. Data refinement based on user feedback
[0853] Program Overview:
[0854] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[0855] Processing Description:
[0856] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[0857] Examples:
[0858] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[0859] 5. Recognizing and responding to user emotions using an emotion engine
[0860] Program Overview:
[0861] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[0862] Processing Description:
[0863] When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server converts the voice data into text using a speech recognition API and then analyzes the user's emotions using an emotion engine. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking recipe that best suits the user's emotions is selected. In addition, emotional information is added to the feedback data to help with future recipe suggestions.
[0864] Examples:
[0865] If a user requests in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine will recognize the user's tired emotions and suggest a simple omelet that requires little effort to make."
[0866] Through the above process, the system of the present invention can provide cooking recipes through voice dialogue, always providing the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0867] The processing flow will be explained below.
[0868] Recognizing and responding to user emotions using an emotion engine
[0869] Program Overview:
[0870] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[0871] Processing Steps:
[0872] Step 1:
[0873] The user speaks into the device to make a cooking request, such as, "I'm tired. Do you have a quick dinner?"
[0874] Step 2:
[0875] The device captures the user's voice. The device's microphone is used to obtain the voice data.
[0876] Step 3:
[0877] The device sends the captured audio data to a server, which then transmits the audio data over the Internet.
[0878] Step 4:
[0879] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[0880] Step 5:
[0881] The server passes the text data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the text data and the intonation and tone of the voice to recognize the user's emotions.
[0882] Step 6:
[0883] The server reflects the analyzed emotional information in the user profile, which is then added to the user profile and used to select the next recipe.
[0884] Step 7:
[0885] The server selects the optimal recipe based on the user's request and emotional information. The selection algorithm takes the user's emotional information into account.
[0886] Step 8:
[0887] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[0888] Step 9:
[0889] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[0890] Step 10:
[0891] The server transmits the generated voice data to the terminal, which then transmits the voice data via the Internet.
[0892] Step 11:
[0893] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[0894] Step 12:
[0895] The user tries the recipe and enters feedback into the device, for example, "This recipe was very easy and delicious."
[0896] Step 13:
[0897] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[0898] Step 14:
[0899] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[0900] Step 15:
[0901] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[0902] Step 16:
[0903] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[0904] Step 17:
[0905] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[0906] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is able to provide the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0907] Example 2
[0908] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0909] Conventional cooking assistant systems have difficulty reflecting users' preferences and allergy information when providing appropriate cooking instructions in response to a user's voice request. They also lack the ability to post the cooking instructions entered by the user in an appealing format, or the ability to improve the system's accuracy based on user feedback. Furthermore, they are unable to suggest recipes that take the user's emotions into account, making it difficult to provide personalized services.
[0910] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's voice, means for converting voice data into text data, means for analyzing the text data and selecting cooking procedures according to the user's request, means for converting the selected cooking procedures into voice data, means for providing the voice data to the user, means for storing the user's cooking preferences and allergy information in a database and recommending cooking procedures based on the stored information, means for improving the cooking procedures entered by the user using a generative AI model and automatically generating a posting page, means for collecting user feedback, updating the ratings of the cooking procedures, and excluding low-rated cooking procedures from the system's learning data, and means for recognizing emotions from the user's voice and suggesting optimal cooking procedures based on the emotions, and means for adding user emotion information to the feedback data. This makes it possible to always provide the latest, high-quality cooking procedures taking into account the user's preferences and emotions.
[0911] "User" refers to any individual or group of people who operate the system and use the cooking recipe suggestions and voice interface.
[0912] "Voice capture means" refers to a device or software mechanism for capturing a user's speech as digital data.
[0913] "Voice data" refers to data that is a digital representation of a user's voice.
[0914] "Text data" refers to data in the form of a string of characters obtained by analyzing voice data.
[0915] "Means for analyzing text data" refers to software or algorithms that analyze text data converted from speech and understand the user's intent and request.
[0916] A "cooking recipe" refers to a document that lists the steps or methods for preparing a dish.
[0917] "Means for converting into voice data" refers to software or devices that convert text data back into voice format using voice synthesis technology.
[0918] A "database" refers to a system for efficiently storing and managing various data such as user preferences and allergy information.
[0919] "Recommendation tools" refers to algorithms and software functions that select optimal cooking instructions based on a user's profile information and past data.
[0920] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generates text, images, etc. based on user input.
[0921] "Posting Page" refers to a web page or part of an application where a user can share and publish cooking instructions.
[0922] "Feedback" refers to opinions such as ratings and comments provided by users to the system.
[0923] An "emotion engine" refers to a technology or software component that analyzes emotions from a user's voice or text and understands the user's emotional state.
[0924] "Information processing device" refers to the entire system including a computer and peripheral devices for inputting, processing, and outputting data.
[0925] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking procedures, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[0926] Implementing a voice-interactive UI
[0927] The program for this system primarily uses speech recognition and speech synthesis APIs to build a mechanism for analyzing voice input from users in real time and generating appropriate responses. Specifically, it uses the Google Cloud Speech-to-Text API as the speech recognition API and the Google Cloud Text-to-Speech API as the speech synthesis API. When a user asks a question out loud, such as "What would you like for dinner tonight?", the device captures this speech and sends it to the server as audio data. The server passes the received audio data to the speech recognition API, converts it into text data, and analyzes the user's request. After selecting the appropriate cooking instructions, it converts it into audio data using the speech synthesis API and sends it to the device, where it provides the user with a spoken response.
[0928] Register user preferences and introduce the most suitable menu
[0929] The system implements an algorithm that recommends cooking steps based on a database of cooking preferences and allergy information preset by the user. The user enters their preferences and allergy information into the device and sends it to the server. The server stores the received information in a database and creates a user profile. When the user makes a request for cooking suggestions, the server refers to the user profile and selects appropriate cooking steps. These are then provided to the user via voice or text.
[0930] Recipe posting function to the community
[0931] We provide a mechanism to automatically generate posting pages using a generative AI model so that users can easily post cooking instructions. When a user enters new cooking instructions, the device sends the information to a server. The server then uses the generative AI model to refine the received instructions and automatically generate an attractive title and description. The generated posting page is saved in a database and made publicly available for other users to view.
[0932] Data refinement based on user feedback
[0933] We implement a mechanism to collect feedback from users, update the ratings of cooking steps, and exclude low-rated steps from the system's learning data. When a user enters feedback for a cooking step they have tried, the device sends the information to a server. The server stores the feedback data in a database and periodically aggregates and analyzes it. By excluding low-rated steps from the learning data in this way, the accuracy of the algorithm is improved.
[0934] Recognizing and responding to user emotions using an emotion engine
[0935] The system recognizes emotions from the user's voice and suggests optimal cooking steps based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy. When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server uses a speech recognition API to convert it into text data and then uses an emotion engine to analyze the user's emotions. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking steps that best suit the user's emotions are selected. For example, if a user requests, in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine recognizes the user's fatigue and suggests a simple omelet that requires little effort to make."
[0936] Through the above process, the system of the present invention can provide cooking instructions through voice interaction, always providing the latest, high-quality cooking instructions while reflecting the user's preferences and emotions, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[0937] Prompt Sentence Examples
[0938] When a user asks, "What's your recommended dish for tonight?", the system responds, "How about spaghetti arrabiata? It's easy and delicious." In this way, the system smoothly suggests the appropriate cooking procedure for a specific situation.
[0939] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0940] Implementing a voice-interactive UI
[0941] (Processing flow)
[0942] Step 1: User makes a voice request
[0943] Step 2: The device captures the audio data and sends it to the server
[0944] Step 3: The server uses a speech recognition API to convert the data into text.
[0945] Step 4: The server analyzes the request
[0946] Step 5: The server selects the cooking instructions based on the request.
[0947] Step 6: The server converts the selected recipe into text data.
[0948] Step 7: The server uses the speech synthesis API to convert the data into audio.
[0949] Step 8: Respond to the user by sending voice data to the device
[0950] (Specific explanation)
[0951] Step 1:
[0952] The user makes a voice request. For example, the user says, "What would you like for dinner tonight?"
[0953] Step 2:
[0954] The device captures the voice data and sends it to the server. Specifically, it uses the device's microphone to capture the user's voice as digital data. The device then sends the captured voice data to the server via the Internet. (Input) Voice data (Output) Data sent to the server
[0955] Step 3:
[0956] The server converts the received voice data into text data using a speech recognition API. For example, the server analyzes the received voice data using a speech recognition API such as Google Cloud Speech-to-Text API. This converts the voice into text data. (Input) Voice data (Output) Text data
[0957] Step 4:
[0958] The server analyzes the request. The server passes the text data to a natural language processing engine for analysis and understands the user's request. For example, it extracts keywords such as "dinner," "recommendations," and "tell me." (Input) Text data (Output) Analysis results
[0959] Step 5:
[0960] The server selects the cooking procedure according to the request. Based on the analysis results, the server refers to the user profile and database to select the appropriate cooking procedure. For example, "Spaghetti Arrabbiata" is selected. (Input) Analysis results (Output) Selected cooking procedure
[0961] Step 6:
[0962] The server converts the selected cooking instructions into text data. The selected cooking instructions are formatted for presentation to the user. For example, a sentence such as "Today's recommended dish is spaghetti arrabbiata" is generated. (Input) Selected cooking instructions (Output) Text data
[0963] Step 7:
[0964] The server converts the text data into audio data using a speech synthesis API. The server converts the text data into audio data using a speech synthesis API such as Google Cloud Text-to-Speech API. (Input) Text data (Output) Audio data
[0965] Step 8:
[0966] The server responds to the user by transmitting the generated voice data to the terminal, and the terminal plays back the received voice data to provide a response to the user.
[0967] (Input) Voice data (Output) Voice response provided to the user
[0968] Register user preferences and introduce the most suitable menu
[0969] (Processing flow)
[0970] Step 1: User registers preferences and allergy information
[0971] Step 2: The device sends the information to the server
[0972] Step 3: The server saves the information to a database
[0973] Step 4: User requests dish suggestions
[0974] Step 5: The server looks up the user profile and selects the best recipe
[0975] Step 6: The server provides the recipe via voice or text
[0976] (Specific explanation)
[0977] Step 1:
[0978] The user registers their preferences and allergy information, for example, by typing "I'm a vegetarian and have a nut allergy" into the terminal.
[0979] Step 2:
[0980] The device sends information to the server. Specifically, the device converts the input information into data packets and sends them to the server. (Input) Preference and allergy information (Output) Data sent to the server
[0981] Step 3:
[0982] The server saves the information in a database. The server analyzes the received user information and saves it in a database as a user profile. For example, it classifies the information by adding tags such as "vegetarian" or "nut allergy." (Input) Data sent (Output) User profile saved in the database
[0983] Step 4:
[0984] The user requests food suggestions, for example, "Tell me what's good for dinner today."
[0985] Step 5:
[0986] The server refers to the user profile and selects the most suitable cooking procedure. For example, "Select vegetarian curry without nuts." (Input) User profile (Output) Selected cooking procedure
[0987] Step 6:
[0988] The server provides cooking instructions by voice or text. The server generates the selected cooking instructions as text data and converts it into voice data using a speech synthesis API. It is played on the device. (Input) Selected cooking instructions (Output) Voice response provided to the user
[0989] Recipe posting function to the community
[0990] (Processing flow)
[0991] Step 1: User enters new cooking instructions
[0992] Step 2: The device sends the recipe information to the server
[0993] Step 3: The server refines the recipe information using the generated AI model
[0994] Step 4: The server automatically generates the submission page
[0995] Step 5: Save the post to the database and publish it
[0996] (Specific explanation)
[0997] Step 1:
[0998] A user inputs new cooking instructions, for example, "new pancake recipe" into the terminal and clicks the submit button.
[0999] Step 2:
[1000] The terminal sends recipe information to the server. The terminal sends the input recipe information to the server as a data packet. (Input) Cooking procedure information (Output) Data sent to the server
[1001] Step 3:
[1002] The server refines the recipe information using a generative AI model. The server analyzes the received step-by-step information using the generative AI model and generates an appealing title and description. For example, it generates a title such as "Easy recipe for fluffy pancakes." (Input) Received step-by-step information (Output) Generated title and description
[1003] Step 4:
[1004] The server automatically generates a posting page. Based on the generated title and description, the server automatically generates a recipe posting page. (Input) Generated title and description (Output) Posting page
[1005] Step 5:
[1006] Save the submitted page in the database and make it public. Save the generated submitted page in the database and make it public. Display it in the web interface so that other users can view it. (Input) Submitted page (Output) Saved in the database and published
[1007] Data refinement based on user feedback
[1008] (Processing flow)
[1009] Step 1: User provides feedback on cooking instructions
[1010] Step 2: The device sends the feedback information to the server
[1011] Step 3: The server stores the feedback data in a database
[1012] Step 4: The server periodically aggregates and analyzes the feedback.
[1013] Step 5: The server removes cooking instructions with low ratings from the training data
[1014] (Specific explanation)
[1015] Step 1:
[1016] The user inputs feedback on the cooking procedure, for example, "This recipe was bland" into the terminal.
[1017] Step 2:
[1018] The terminal sends feedback information to the server. The terminal converts the input feedback information into data packets and sends them to the server. (Input) Feedback information (Output) Data sent to the server
[1019] Step 3:
[1020] The server saves the feedback data in a database. For example, it saves data such as "Restaurant Egg Curry", "Bland taste", and "Rating: 2". (Input) Feedback information sent. (Output) Feedback saved in the database.
[1021] Step 4:
[1022] The server periodically collects and analyzes feedback. Feedback data is periodically collected, and the ratings from many users are analyzed to determine which recipes are highly rated and which are poorly rated. (Input) Feedback stored in the database (Output) Analysis results
[1023] Step 5:
[1024] The server removes cooking steps with low ratings from the training data. By removing recipes with low ratings based on the aggregation results from the training data, the system improves its suggestion accuracy. (Input) Analysis results (Output) Updated training data
[1025] Recognizing and responding to user emotions using an emotion engine
[1026] (Processing flow)
[1027] Step 1: User makes a voice request
[1028] Step 2: The device captures the audio data and sends it to the server
[1029] Step 3: The server uses a speech recognition API to convert the data into text.
[1030] Step 4: The server analyzes the emotion information using the emotion engine
[1031] Step 5: The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal recipe.
[1032] Step 6: The server adds emotion information to the feedback data
[1033] (Specific explanation)
[1034] Step 1:
[1035] The user makes a request by voice. For example, the user may request in a tired voice, "I want a quick dinner."
[1036] Step 2:
[1037] The device captures the audio data and sends it to the server. The device captures the audio data and sends it to the server. (Input) Audio data (Output) Data sent to the server
[1038] Step 3:
[1039] The server uses a speech recognition API to convert the voice data into text data. The server uses a speech recognition API to convert the voice data into text data. (Input) Voice data (Output) Text data
[1040] Step 4:
[1041] The server uses an emotion engine to analyze emotional information. The server uses an emotion engine to analyze emotional information from the tone and pace of the user's voice. For example, emotional information indicating "fatigue" is extracted. (Input) Text data (Output) Emotional information
[1042] Step 5:
[1043] The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal cooking procedure. The server selects the optimal recipe for the user based on the emotional information. For example, it selects a "simple omelet" that can be made with minimal effort. (Input) Emotional information (Output) Selected cooking procedure
[1044] Step 6:
[1045] The server adds the emotional information to the feedback data. The emotional information is saved as feedback data along with the suggested cooking steps, and will be used for future suggestions. (Input) Emotional information (Output) Feedback data
[1046] (Application example 2)
[1047] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1048] Conventional recipe provision systems suggest recipes by analyzing users' voice requests, but they have the problem of being unable to provide recipes that reflect the user's emotions or the situation of the day. This means that users cannot receive suggestions for dishes that suit their physical condition or mood, resulting in a lack of convenience. In addition, there is a lack of a mechanism for improving the accuracy of the system by reflecting user feedback.
[1049] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1050] In this invention, the server includes means for capturing user voice, means for converting voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for analyzing the user's emotions, means for recommending recipes based on the emotion information, and means for collecting user feedback and improving the accuracy of the system. This makes it possible to provide individualized recipes according to each user's emotions, thereby realizing more convenient recipe suggestions.
[1051] The "means for capturing the user's voice" is a mechanism by which the electronic device receives the voice uttered by the user and captures it as digital data.
[1052] The "means for converting voice data into text data" is a mechanism that analyzes the captured voice data and performs a process to express the contents as a string of characters.
[1053] The "means for analyzing text data and selecting cooking recipes that meet the user's requests" refers to an algorithm and system that understands the user's requests based on the converted text data and selects appropriate cooking recipes.
[1054] The "means for converting the selected recipe into audio data" is a mechanism that performs the process of converting the selected recipe from a character string into an audio signal.
[1055] The "means for providing audio data to the user" refers to a device or program for playing back the converted audio data to the user and conveying information.
[1056] "Means for analyzing user emotions" refers to algorithms or systems that determine and analyze the user's emotional state from the tone and content of the voice.
[1057] "Means for recommending cooking recipes based on emotional information" is a system that selects and recommends the most appropriate cooking recipe based on the user's current mood and state based on the results of emotional analysis.
[1058] "Means for collecting user feedback and improving the accuracy of the system" refers to the process and mechanisms for collecting user evaluations and opinions and using them to improve the system's algorithms and database.
[1059] This invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. In particular, it aims to provide more personalized services by combining it with an emotion engine that recognizes the user's emotions.
[1060] System Configuration
[1061] Hardware
[1062] Smartphone: A device that captures the user's voice and receives the analysis results.
[1063] Server: Analyzes voice data, analyzes emotions, selects recipes, and aggregates feedback data
[1064] software
[1065] SpeechRecognition API: Converting voice data into text data
[1066] Emotion Analysis API: Analyzing user emotions from text data
[1067] Text to Speech API: Convert text data into audio data
[1068] Database: Stores user profile information, cooking recipes, and feedback data
[1069] Operation overview
[1070] 1. Audio Capture
[1071] A user speaks to their smartphone, asking, "What would you like for dinner tonight?" The smartphone captures the voice through its built-in microphone and sends this voice data to the server.
[1072] 2. Voice Recognition
[1073] When the server receives the voice data, it uses the SpeechRecognition API to convert the voice data into text data.
[1074] 3. Emotion analysis
[1075] The server then sends the converted text data to the Emotion Analysis API, which analyzes the user's emotions. The analysis result can be, for example, "I'm tired."
[1076] 4. Recipe Selection
[1077] The server selects the most suitable recipe for the user's request based on the results of sentiment analysis and the user profile database. For tired users, it selects easy-to-make dishes.
[1078] 5. Speech Synthesis
[1079] The selected cooking recipe is again treated as text data and converted into audio data using the Text to Speech API.
[1080] 6. Provide answers
[1081] The smartphone receives the converted voice data and provides the user with a spoken response such as, "You seem tired today, how about a quick sandwich?"
[1082] 7. Feedback Collection
[1083] After users try the suggested recipes, they provide feedback. The server collects this feedback data and stores it in a database. In the future, the feedback will be used to improve the accuracy of the algorithm.
[1084] Specific examples
[1085] For example, when a user makes a voice request such as "Tell me something easy to eat," the system captures the voice and the server analyzes the voice data. If the emotion engine recognizes the emotion "tired" from the user's voice, the system will suggest a quick and easy-to-prepare dish such as a "sandwich." The following is an example of a prompt sentence that can be input to the generative AI model:
[1086] Convert the user's voice into text, feed that text into a sentiment analysis engine, select the sentiment, and output a food recommendation accordingly.
[1087] Input: I'm tired today and want a quick meal.
[1088] Emotions: Tired
[1089] Output: You seem tired, would you like some lunch?
[1090] These procedures provide personalized information in response to user requests, improving user convenience.
[1091] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1092] Step 1:
[1093] The user speaks a question into the smartphone. The user's voice data is captured by the smartphone's microphone. The input is the user's voice, and the output is the captured voice data. The voice data is converted into a digital format and sent to the server.
[1094] Step 2:
[1095] The server receives the audio data. The input is digital audio data, and the output is text data. The server converts the audio data into text data using the SpeechRecognition API. In this process, the audio content is converted into a string of characters.
[1096] Step 3:
[1097] The server sends the text data to the Emotion Analysis API and analyzes the user's emotions. The input is text data and the output is emotional information. The server obtains the analysis result, for example, an emotion such as "tired."
[1098] Step 4:
[1099] The server selects the optimal recipe for the user's request based on the emotional information and the user profile database. The input is the emotional analysis results and the user profile, and the output is the selected recipe. In particular, if the user is tired, an easy-to-make dish is selected.
[1100] Step 5:
[1101] The server converts the selected recipe into audio data using the Text to Speech API. The input is the text data of the selected recipe, and the output is audio data. In this process, the recipe in text format is converted back into audio format.
[1102] Step 6:
[1103] The server sends the voice data to the smartphone. The smartphone receives this data and suggests recipes to the user by voice. The input is the converted voice data, and the output is the voice output to the user.
[1104] Step 7:
[1105] After the user has tried the suggested recipe, they input their feedback into their smartphone. The input is the user's feedback, and the output is the collected feedback data. The feedback data is sent to the server and stored in a database.
[1106] Step 8:
[1107] The server analyzes the feedback data and updates the algorithm to improve the system's accuracy. The input is the accumulated feedback data, and the output is an improved recommendation algorithm. This will enable more personalized and accurate recipe suggestions in the future.
[1108] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1109] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1110] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1111] [Third embodiment]
[1112] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1113] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1114] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1115] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1116] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1117] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1118] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1119] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1120] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1121] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1122] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1123] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1124] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[1125] 1. Implementing a voice-activated UI
[1126] Program Overview:
[1127] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1128] Processing Description:
[1129] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[1130] Examples:
[1131] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[1132] 2. Register user preferences and introduce the most suitable menu
[1133] Program Overview:
[1134] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1135] Processing Description:
[1136] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[1137] Examples:
[1138] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[1139] 3. Recipe posting function to the community
[1140] Program Overview:
[1141] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1142] Processing Description:
[1143] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[1144] Examples:
[1145] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[1146] 4. Data refinement based on user feedback
[1147] Program Overview:
[1148] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1149] Processing Description:
[1150] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[1151] Examples:
[1152] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[1153] In this way, the system of the present invention provides cooking recipes based on voice interaction and reflects user preferences and community participation, allowing users to enjoy a more comfortable and convenient cooking experience.
[1154] The processing flow will be explained below.
[1155] 1. Implementing a voice-activated UI
[1156] Program Overview:
[1157] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1158] Processing Steps:
[1159] Step 1:
[1160] Users speak into the device to ask questions or make requests about food, such as "What would you like for dinner tonight?"
[1161] Step 2:
[1162] The device captures the user's voice. The device uses a built-in microphone to obtain the voice data.
[1163] Step 3:
[1164] The device transmits the captured audio data to the server, and the device transmits the audio data via the Internet.
[1165] Step 4:
[1166] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[1167] Step 5:
[1168] The server analyzes the text data to understand the user's intent, and then searches for appropriate recipe information based on the analyzed data.
[1169] Step 6:
[1170] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[1171] Step 7:
[1172] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[1173] Step 8:
[1174] The server transmits the generated voice data to the terminal, and the voice data is transmitted via the Internet.
[1175] Step 9:
[1176] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[1177] 2. Register user preferences and introduce the most suitable menu
[1178] Program Overview:
[1179] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1180] Processing Steps:
[1181] Step 1:
[1182] The user inputs their food preferences and allergy information through the terminal, such as "vegetarian" or "nut allergy."
[1183] Step 2:
[1184] The device sends the input preference information to the server, and the input data is sent to the server via the Internet.
[1185] Step 3:
[1186] The server stores the received information in a database as a user profile. Profile data is created and stored for each user.
[1187] Step 4:
[1188] The user inputs a specific food request into the device, for example, asking "What do you recommend for lunch today?"
[1189] Step 5:
[1190] The device captures the audio and sends it to a server, which then sends the audio data over the internet.
[1191] Step 6:
[1192] The server passes the received voice data to the voice recognition API and converts it into text data. The text data obtained from the voice recognition API is analyzed.
[1193] Step 7:
[1194] The server references the user profile and uses a filtering algorithm to select the best recipes, based on the user's preferences and allergies.
[1195] Step 8:
[1196] The selected recipe information is passed to the speech synthesis API and converted into voice data, which is then returned to the server.
[1197] Step 9:
[1198] The server sends the audio data to the terminal. The audio data is sent via the Internet.
[1199] Step 10:
[1200] The terminal plays the audio data to provide information to the user. The audio data is played through the terminal's speaker.
[1201] 3. Recipe posting function to the community
[1202] Program Overview:
[1203] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1204] Processing Steps:
[1205] Step 1:
[1206] The user inputs new cooking recipe information into the terminal, for example, "original cookie recipe."
[1207] Step 2:
[1208] The terminal sends the input recipe information to the server, and the input data is sent to the server via the Internet.
[1209] Step 3:
[1210] The server passes the received recipe information to the generation AI model for refinement. The generation AI analyzes the recipe information and generates a more attractive title and description.
[1211] Step 4:
[1212] The server saves the generated submission page data, including the title and description, in a database. The generated submission page is then prepared for public viewing.
[1213] Step 5:
[1214] The server generates a URL for publishing and configures the publishing settings. The submission page will then be viewable by the general public.
[1215] Step 6:
[1216] General users can view published recipe pages through their devices. They can also view the posting page using the device's browser.
[1217] 4. Data refinement based on user feedback
[1218] Program Overview:
[1219] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1220] Processing Steps:
[1221] Step 1:
[1222] The user enters feedback about the recipe they tried into the device, for example, "This recipe was bland."
[1223] Step 2:
[1224] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[1225] Step 3:
[1226] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[1227] Step 4:
[1228] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[1229] Step 5:
[1230] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[1231] Step 6:
[1232] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[1233] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is capable of always providing the latest, high-quality recipe information while reflecting the user's preferences and community participation.
[1234] Example 1
[1235] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1236] Conventional cooking assistant systems are limited to analyzing users' voice input and providing cooking recipes, but lack the ability to reflect users' preferences and allergy information or filter recipes based on user feedback. Furthermore, they lack the ability to automatically refine recipe posts when users post new recipes. This has made improving the user experience a challenge.
[1237] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1238] In this invention, the server includes means for capturing a user's voice, means for converting the voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for collecting user feedback and storing it in a database, means for excluding low-rated recipes from the training data, means for polishing the recipe entered by the user using a generative AI model, and means for inputting prompt sentences into the generative AI model. This makes it possible to provide optimal recipes that reflect the user's preferences and feedback, and to automatically polish new recipes.
[1239] A "means for capturing voice" is a device or method for capturing a user's spoken voice as digital data.
[1240] The "means for converting voice data into text data" is a process for converting captured voice data into text information using voice recognition technology.
[1241] "Means for analyzing text data and selecting cooking recipes that meet user requests" refers to a method that uses natural language processing technology to understand user requests from text data and select the optimal cooking recipe to meet those requests.
[1242] The "means for converting the selected recipe into voice data" is a method for converting text-format recipe information into a format that can be output as voice using voice synthesis technology.
[1243] A "means for providing audio data to a user" is a method for transmitting audio data to a user's device and playing it back.
[1244] The "means for collecting user feedback and storing it in a database" refers to a method for collecting, organizing, and storing the ratings and comments provided by users in a database.
[1245] "Means for excluding low-rated recipes from learning data" refers to the process of analyzing user rating data and removing recipes that have received a certain level of low rating from the system's recommendation candidates.
[1246] "Means for brushing up recipes entered by users using generative AI models" refers to a method of improving recipe information provided by users using generative AI technology to make the content more appealing.
[1247] A "means for inputting prompt text into a generative AI model" is a method for inputting text containing specific instructions or information into a generative AI model and obtaining output based on that text.
[1248] MODE FOR CARRYING OUT THE INVENTION
[1249] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[1250] 1. System Configuration
[1251] This system consists of a device used by the user and a central server. The device has a microphone to capture the user's voice and a speaker to play back the voice data. The server uses a speech recognition API, a speech synthesis API, a generative AI model, and a database to select and respond to the user's request with the optimal recipe.
[1252] 2. Implementing a Voice-Based Interactive UI
[1253] When a user asks, "What would you like for dinner today?", the device captures the voice and sends the voice data to the server. The server uses a speech recognition API such as Amazon Transcribe to convert the voice data into text data. The converted text data is then analyzed to select a cooking recipe that meets the user's request. The results of this selection are then prepared as text data and converted into voice data using a speech synthesis API such as Amazon Polly. The voice data is then sent to the device and provided to the user through the speaker.
[1254] Examples:
[1255] When a user asks, "What's your recommended dish for tonight?" the system will respond aloud with, "How about spaghetti arrabiata? It's easy and delicious."
[1256] 3. Recipe selection that reflects user preferences
[1257] Users enter their food preferences and allergy information into their device and send it to the server. The server stores this information in a database and manages it as a user profile. When a user requests recipe suggestions, the server references this user profile and selects the most suitable recipe.
[1258] Examples:
[1259] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a nut-free vegetarian curry."
[1260] 4. Automatically improve recipe posts
[1261] When a user enters a new recipe into their device and posts it, the server passes the information to a generative AI model to generate an appealing title and description, and the resulting post page is stored in a database and made available for other users to view.
[1262] Examples:
[1263] When a user enters and posts a "new pancake recipe," the system automatically generates a title such as "Easy recipe for fluffy pancakes" and an introductory text, creating a posting page. For example, the prompt for the generative AI model could be "Generate an attractive cooking recipe posting page based on the following content:"
[1264] 5. Data refinement based on user feedback
[1265] The feedback provided by the user is sent from the device to the server and stored in a database. The server periodically aggregates the feedback data and removes low-rated recipes from the learning data. This allows the system to provide only highly rated recipes to users in the future.
[1266] Examples:
[1267] If a user provides feedback such as "this recipe was bland," the server will filter out that recipe if similar feedback continues, improving the accuracy of the system's recipe selection algorithm.
[1268] As described above, the voice-activated cooking assistant system of the present invention provides optimal recipes for users by combining voice input analysis, user profile utilization, recipe refinement using generative AI models, and feedback collection and analysis, thereby enabling users to enjoy a more comfortable and convenient cooking experience.
[1269] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1270] Step 1:
[1271] The user provides voice input. The user says, "What would you like for dinner tonight?"
[1272] Step 2:
[1273] The device captures voice data and sends it to the server. The voice data is obtained in digital form through a microphone and sent to the server using an HTTP request. The input is the user's voice and the output is digital voice data sent to the server.
[1274] Step 3:
[1275] The server receives the voice data and passes it to the voice recognition API. The voice data is stored in cloud storage, and its URL is used as input for the voice recognition API. The input is the URL of the voice data, and the output is text data.
[1276] Step 4:
[1277] The server converts the audio data into text data using a speech recognition API. The text is extracted from the JSON format data returned as an API response. The input is the URL of the audio data and the API response, and the output is the extracted text data.
[1278] Step 5:
[1279] The server analyzes the text data, understands the user's request, and selects the most suitable recipe. It uses natural language processing technology to extract the "dish name" and "conditions" from the text data, and queries the database to search for matching recipes. The input is text data, and the output is information about the selected recipe.
[1280] Step 6:
[1281] The server prepares the selected recipe as text data, structuring it as "Spaghetti Arrabbiata is recommended." The input is the selected recipe information, and the output is the response text data.
[1282] Step 7:
[1283] The server uses a speech synthesis API such as Amazon Polly to convert text data into speech data. The request to the API includes text data and speech settings. The input is text data, and the output is the generated speech data.
[1284] Step 8:
[1285] The server sends the generated voice data to the device and provides it to the user through the device's speaker. The voice data is sent to the device as an HTTP response, and the device plays the data. The input is the voice data, and the output is the voice response provided to the user.
[1286] Step 9:
[1287] The user inputs their food preferences and allergy information into the device, which then sends it to the server. The input information is sent in JSON format. The input is the user's preferences and allergy information, and the output is the user profile data sent to the server.
[1288] Step 10:
[1289] The server stores the received information in a database. It associates the information with the user ID and adds it to the corresponding record in the database. The input is the user profile data, and the output is the user information stored in the database.
[1290] Step 11:
[1291] The user inputs a new cooking recipe and the device sends the information to the server. The input information is sent in JSON format. The input is the cooking recipe entered by the user, and the output is the recipe data sent to the server.
[1292] Step 12:
[1293] The server inputs the received recipe information into the generative AI model. The generative AI model then inputs text data containing specific prompts. The input is the user's recipe information and prompt, and the output is the generated title and description.
[1294] Step 13:
[1295] The server saves the generated post, including the title and description, in a database where it can be viewed by other users. The input is the generated post data, and the output is the post saved in the database.
[1296] Step 14:
[1297] The user inputs feedback for the recipes they have tried, and the terminal sends this information to the server. The input is the user's feedback, and the output is the feedback data sent to the server.
[1298] Step 15:
[1299] The server saves the feedback data in a database. It saves the rating scores and comments in the appropriate fields. The input is the feedback data, and the output is the rating information saved in the database.
[1300] Step 16:
[1301] The server periodically aggregates the feedback data and removes poorly rated recipes from the training data. It then runs a data analysis process to identify and remove poorly rated recipes. The input is the feedback data, and the output is updated training data.
[1302] The above are the specific processing steps of the program of this system.
[1303] (Application example 1)
[1304] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1305] When users use food delivery services in their daily lives, they face challenges such as the time it takes to select the optimal menu and the need to individually confirm preferences and allergy information. Furthermore, when users select a delivery menu, they are unable to make the optimal selection based on their past order history and preferences. Recipe posting and rating functions for sharing with other users are also time-consuming and labor-intensive, as they must be done manually.
[1306] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1307] In this invention, the server includes a means for capturing the user's voice, a means for converting the voice data into text data, and a means for analyzing the text data and selecting a recipe or delivery menu item according to the user's request. This allows the user to quickly request a delivery menu item by voice, streamlining the ordering process. The server also includes a function for storing user preferences and allergy information in a database and recommending delivery menu items based on that information, enabling optimal suggestions for each individual user. Furthermore, by adding a function for refining and automatically generating recipes entered by the user, users can efficiently post and rate recipes for sharing with other users.
[1308] "User" refers to a person who uses the system to make a voice request and receive a cooking recipe or delivery menu.
[1309] "Means for capturing audio" refers to a mechanism for collecting the user's audio using a device such as a microphone.
[1310] "Voice data" refers to data that expresses the user's voice as digital information.
[1311] "Means for converting into text data" refers to a mechanism for converting voice data into text information using voice recognition technology.
[1312] "Text data" refers to digital information that has been converted from audio data into text form.
[1313] "User request" refers to a request or question uttered by a user.
[1314] "Cooking recipes or delivery menus" refers to the food preparation methods and delivery options suggested by the system.
[1315] "Means for selection" refers to a mechanism for selecting the most suitable cooking recipe or delivery menu based on the user's request.
[1316] "Means for converting into voice data" refers to a mechanism that uses voice synthesis technology to convey the selected recipe or delivery menu to the user in voice form.
[1317] "Means of providing" refers to devices such as speakers and headphones used to deliver audio data to users.
[1318] "User's food preferences and allergy information" refers to information about dietary preferences and ingredients to avoid that the user has registered in the system.
[1319] A "database" refers to a system for centrally managing user information and data processed by the system.
[1320] "Recommendation means" refers to a mechanism for suggesting optimal recipes or delivery menus by taking into consideration the user's food preferences and allergy information.
[1321] "Inputted recipe" refers to information on food cooking methods and ingredients that a user has registered or provided to the system.
[1322] "Means of polishing and automatically generating a post page" refers to a mechanism that uses a generative AI model to format information entered by a user into an attractive form and automatically convert it into a publishable format.
[1323] To implement this invention, it is necessary to build a system that proposes optimal cooking recipes or delivery menus based on a user's voice request. This system is configured around a voice-interactive user interface.
[1324] System configuration
[1325] 1. How to capture the user's voice:
[1326] Hardware: A smartphone or tablet with a built-in microphone.
[1327] Role: Collects user voice in real time.
[1328] 2. How to convert audio data to text data:
[1329] Software: speech_recognition library.
[1330] Role: Converts voice data into text.
[1331] 3. A method for analyzing text data and selecting cooking recipes or delivery menus according to user requests:
[1332] Technologies used: Natural Language Processing (NLP) and databases.
[1333] Role: Analyzes user requests and selects the most suitable cooking recipe or delivery menu.
[1334] 4. Means for converting selected information into audio data:
[1335] Software: pyttsx3 library.
[1336] Role: Converts selected recipes or delivery menus into audio data.
[1337] 5. Means of providing audio data to the user:
[1338] Hardware: Smartphone speakers and headphones.
[1339] Role: Communicates information to the user audibly.
[1340] 6. A means to store user's food preferences and allergy information in a database and make recommendations:
[1341] Database: A specialized user profile database.
[1342] Role: Stores user preferences and allergy information to improve future recommendations.
[1343] 7. How to polish a recipe entered by a user and automatically generate a posting page:
[1344] Technology used: Content generation using generative AI models.
[1345] Role: To polish recipe information provided by users to make it more appealing and publish it as an automatically generated posting page.
[1346] Program operation explanation
[1347] The server captures voice data input by the user via a smartphone or tablet and converts the voice data into text using the speech_recognition library. It then analyzes the text data using natural language processing technology to select the optimal recipe or delivery menu item for the user's request. This selection process takes into account the user's preferences and allergies. The selected information is then converted back into voice data using the pyttsx3 library and delivered to the user via the smartphone's speakers or headphones.
[1348] Additionally, when a user enters a new recipe, the recipe information is refined using a generative AI model and automatically generated into an attractive posting page, which is then saved in a database and made publicly available for other users to view.
[1349] Specific examples
[1350] When a user speaks to their smartphone, "What delivery do you recommend for dinner tonight?", the system responds by voice, "Would you like pizza or sushi? Which would you like to order?" If the user selects pizza, the system will suggest the best pizza delivery service.
[1351] Prompt Sentence Examples
[1352] "The following input should contain a list of restaurants that offer pizza and sushi delivery. If the user selects pizza, output the names of the top 5 restaurants listed.
[1353] Input: I want pizza.
[1354] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1355] Step 1:
[1356] The user issues a voice request to the smartphone. For example, say, "What delivery do you recommend for dinner tonight?" The voice is captured by the smartphone's microphone.
[1357] Step 2:
[1358] The device sends the captured audio data to the server. The input is audio data, which is then sent to the server.
[1359] Step 3:
[1360] The server converts the received voice data into text data using the speech_recognition library. The input is voice data, and the output is the converted text data. This conversion uses speech recognition technology.
[1361] Step 4:
[1362] The server analyzes the text data using natural language processing (NLP) technology to understand the user's request. The input is text data, and the output is the analyzed request data. As a concrete example, it identifies that the user's request is for "dinner delivery."
[1363] Step 5:
[1364] The server retrieves the user's preferences and allergy information from the database, compares it with the analysis results, and recommends the most suitable recipe or delivery menu. The input is the analyzed request content and user profile data, and the output is the selected recipe or delivery menu.
[1365] Step 6:
[1366] The server converts the selected recipe or delivery menu into voice data using the pyttsx3 library. The input is text-based recipe or menu information, and the output is voice data.
[1367] Step 7:
[1368] The server sends the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal.
[1369] Step 8:
[1370] The device then provides the received voice data to the user through a speaker or headphones. The input is voice data, and the output is audible audio for the user. Specifically, the smartphone responds with a voice message asking, "Would you like pizza or sushi? Which would you like to order?"
[1371] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1372] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[1373] 1. Implementing a voice-activated UI
[1374] Program Overview:
[1375] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1376] Processing Description:
[1377] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[1378] Examples:
[1379] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[1380] 2. Register user preferences and introduce the most suitable menu
[1381] Program Overview:
[1382] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1383] Processing Description:
[1384] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[1385] Examples:
[1386] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[1387] 3. Recipe posting function to the community
[1388] Program Overview:
[1389] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1390] Processing Description:
[1391] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[1392] Examples:
[1393] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[1394] 4. Data refinement based on user feedback
[1395] Program Overview:
[1396] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1397] Processing Description:
[1398] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[1399] Examples:
[1400] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[1401] 5. Recognizing and responding to user emotions using an emotion engine
[1402] Program Overview:
[1403] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[1404] Processing Description:
[1405] When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server converts the voice data into text using a speech recognition API and then analyzes the user's emotions using an emotion engine. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking recipe that best suits the user's emotions is selected. In addition, emotional information is added to the feedback data to help with future recipe suggestions.
[1406] Examples:
[1407] If a user requests in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine will recognize the user's tired emotions and suggest a simple omelet that requires little effort to make."
[1408] Through the above process, the system of the present invention can provide cooking recipes through voice dialogue, always providing the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[1409] The processing flow will be explained below.
[1410] Recognizing and responding to user emotions using an emotion engine
[1411] Program Overview:
[1412] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[1413] Processing Steps:
[1414] Step 1:
[1415] The user speaks into the device to make a cooking request, such as, "I'm tired. Do you have a quick dinner?"
[1416] Step 2:
[1417] The device captures the user's voice. The device's microphone is used to obtain the voice data.
[1418] Step 3:
[1419] The device sends the captured audio data to a server, which then transmits the audio data over the Internet.
[1420] Step 4:
[1421] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[1422] Step 5:
[1423] The server passes the text data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the text data and the intonation and tone of the voice to recognize the user's emotions.
[1424] Step 6:
[1425] The server reflects the analyzed emotional information in the user profile, which is then added to the user profile and used to select the next recipe.
[1426] Step 7:
[1427] The server selects the optimal recipe based on the user's request and emotional information. The selection algorithm takes the user's emotional information into account.
[1428] Step 8:
[1429] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[1430] Step 9:
[1431] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[1432] Step 10:
[1433] The server transmits the generated voice data to the terminal, which then transmits the voice data via the Internet.
[1434] Step 11:
[1435] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[1436] Step 12:
[1437] The user tries the recipe and enters feedback into the device, for example, "This recipe was very easy and delicious."
[1438] Step 13:
[1439] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[1440] Step 14:
[1441] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[1442] Step 15:
[1443] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[1444] Step 16:
[1445] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[1446] Step 17:
[1447] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[1448] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is able to provide the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[1449] Example 2
[1450] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1451] Conventional cooking assistant systems have difficulty reflecting users' preferences and allergy information when providing appropriate cooking instructions in response to a user's voice request. They also lack the ability to post the cooking instructions entered by the user in an appealing format, or the ability to improve the system's accuracy based on user feedback. Furthermore, they are unable to suggest recipes that take the user's emotions into account, making it difficult to provide personalized services.
[1452] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's voice, means for converting voice data into text data, means for analyzing the text data and selecting cooking procedures according to the user's request, means for converting the selected cooking procedures into voice data, means for providing the voice data to the user, means for storing the user's cooking preferences and allergy information in a database and recommending cooking procedures based on the stored information, means for improving the cooking procedures entered by the user using a generative AI model and automatically generating a posting page, means for collecting user feedback, updating the ratings of the cooking procedures, and excluding low-rated cooking procedures from the system's learning data, and means for recognizing emotions from the user's voice and suggesting optimal cooking procedures based on the emotions, and means for adding user emotion information to the feedback data. This makes it possible to always provide the latest, high-quality cooking procedures taking into account the user's preferences and emotions.
[1453] "User" refers to any individual or group of people who operate the system and use the cooking recipe suggestions and voice interface.
[1454] "Voice capture means" refers to a device or software mechanism for capturing a user's speech as digital data.
[1455] "Voice data" refers to data that is a digital representation of a user's voice.
[1456] "Text data" refers to data in the form of a string of characters obtained by analyzing voice data.
[1457] "Means for analyzing text data" refers to software or algorithms that analyze text data converted from speech and understand the user's intent and request.
[1458] A "cooking recipe" refers to a document that lists the steps or methods for preparing a dish.
[1459] "Means for converting into voice data" refers to software or devices that convert text data back into voice format using voice synthesis technology.
[1460] A "database" refers to a system for efficiently storing and managing various data such as user preferences and allergy information.
[1461] "Recommendation tools" refers to algorithms and software functions that select optimal cooking instructions based on a user's profile information and past data.
[1462] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generates text, images, etc. based on user input.
[1463] "Posting Page" refers to a web page or part of an application where a user can share and publish cooking instructions.
[1464] "Feedback" refers to opinions such as ratings and comments provided by users to the system.
[1465] An "emotion engine" refers to a technology or software component that analyzes emotions from a user's voice or text and understands the user's emotional state.
[1466] "Information processing device" refers to the entire system including a computer and peripheral devices for inputting, processing, and outputting data.
[1467] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking procedures, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[1468] Implementing a voice-interactive UI
[1469] The program for this system primarily uses speech recognition and speech synthesis APIs to build a mechanism for analyzing voice input from users in real time and generating appropriate responses. Specifically, it uses the Google Cloud Speech-to-Text API as the speech recognition API and the Google Cloud Text-to-Speech API as the speech synthesis API. When a user asks a question out loud, such as "What would you like for dinner tonight?", the device captures this speech and sends it to the server as audio data. The server passes the received audio data to the speech recognition API, converts it into text data, and analyzes the user's request. After selecting the appropriate cooking instructions, it converts it into audio data using the speech synthesis API and sends it to the device, where it provides the user with a spoken response.
[1470] Register user preferences and introduce the most suitable menu
[1471] The system implements an algorithm that recommends cooking steps based on a database of cooking preferences and allergy information preset by the user. The user enters their preferences and allergy information into the device and sends it to the server. The server stores the received information in a database and creates a user profile. When the user makes a request for cooking suggestions, the server refers to the user profile and selects appropriate cooking steps. These are then provided to the user via voice or text.
[1472] Recipe posting function to the community
[1473] We provide a mechanism to automatically generate posting pages using a generative AI model so that users can easily post cooking instructions. When a user enters new cooking instructions, the device sends the information to a server. The server then uses the generative AI model to refine the received instructions and automatically generate an attractive title and description. The generated posting page is saved in a database and made publicly available for other users to view.
[1474] Data refinement based on user feedback
[1475] We implement a mechanism to collect feedback from users, update the ratings of cooking steps, and exclude low-rated steps from the system's learning data. When a user enters feedback for a cooking step they have tried, the device sends the information to a server. The server stores the feedback data in a database and periodically aggregates and analyzes it. By excluding low-rated steps from the learning data in this way, the accuracy of the algorithm is improved.
[1476] Recognizing and responding to user emotions using an emotion engine
[1477] The system recognizes emotions from the user's voice and suggests optimal cooking steps based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy. When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server uses a speech recognition API to convert it into text data and then uses an emotion engine to analyze the user's emotions. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking steps that best suit the user's emotions are selected. For example, if a user requests, in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine recognizes the user's fatigue and suggests a simple omelet that requires little effort to make."
[1478] Through the above process, the system of the present invention can provide cooking instructions through voice interaction, always providing the latest, high-quality cooking instructions while reflecting the user's preferences and emotions, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[1479] Prompt Sentence Examples
[1480] When a user asks, "What's your recommended dish for tonight?", the system responds, "How about spaghetti arrabiata? It's easy and delicious." In this way, the system smoothly suggests the appropriate cooking procedure for a specific situation.
[1481] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1482] Implementing a voice-interactive UI
[1483] (Processing flow)
[1484] Step 1: User makes a voice request
[1485] Step 2: The device captures the audio data and sends it to the server
[1486] Step 3: The server uses a speech recognition API to convert the data into text.
[1487] Step 4: The server analyzes the request
[1488] Step 5: The server selects the cooking instructions based on the request.
[1489] Step 6: The server converts the selected recipe into text data.
[1490] Step 7: The server uses the speech synthesis API to convert the data into audio.
[1491] Step 8: Respond to the user by sending voice data to the device
[1492] (Specific explanation)
[1493] Step 1:
[1494] The user makes a voice request. For example, the user says, "What would you like for dinner tonight?"
[1495] Step 2:
[1496] The device captures the voice data and sends it to the server. Specifically, it uses the device's microphone to capture the user's voice as digital data. The device then sends the captured voice data to the server via the Internet. (Input) Voice data (Output) Data sent to the server
[1497] Step 3:
[1498] The server converts the received voice data into text data using a speech recognition API. For example, the server analyzes the received voice data using a speech recognition API such as Google Cloud Speech-to-Text API. This converts the voice into text data. (Input) Voice data (Output) Text data
[1499] Step 4:
[1500] The server analyzes the request. The server passes the text data to a natural language processing engine for analysis and understands the user's request. For example, it extracts keywords such as "dinner," "recommendations," and "tell me." (Input) Text data (Output) Analysis results
[1501] Step 5:
[1502] The server selects the cooking procedure according to the request. Based on the analysis results, the server refers to the user profile and database to select the appropriate cooking procedure. For example, "Spaghetti Arrabbiata" is selected. (Input) Analysis results (Output) Selected cooking procedure
[1503] Step 6:
[1504] The server converts the selected cooking instructions into text data. The selected cooking instructions are formatted for presentation to the user. For example, a sentence such as "Today's recommended dish is spaghetti arrabbiata" is generated. (Input) Selected cooking instructions (Output) Text data
[1505] Step 7:
[1506] The server converts the text data into audio data using a speech synthesis API. The server converts the text data into audio data using a speech synthesis API such as Google Cloud Text-to-Speech API. (Input) Text data (Output) Audio data
[1507] Step 8:
[1508] The server responds to the user by transmitting the generated voice data to the terminal, and the terminal plays back the received voice data to provide a response to the user.
[1509] (Input) Voice data (Output) Voice response provided to the user
[1510] Register user preferences and introduce the most suitable menu
[1511] (Processing flow)
[1512] Step 1: User registers preferences and allergy information
[1513] Step 2: The device sends the information to the server
[1514] Step 3: The server saves the information to a database
[1515] Step 4: User requests dish suggestions
[1516] Step 5: The server looks up the user profile and selects the best recipe
[1517] Step 6: The server provides the recipe via voice or text
[1518] (Specific explanation)
[1519] Step 1:
[1520] The user registers their preferences and allergy information, for example, by typing "I'm a vegetarian and have a nut allergy" into the terminal.
[1521] Step 2:
[1522] The device sends information to the server. Specifically, the device converts the input information into data packets and sends them to the server. (Input) Preference and allergy information (Output) Data sent to the server
[1523] Step 3:
[1524] The server saves the information in a database. The server analyzes the received user information and saves it in a database as a user profile. For example, it classifies the information by adding tags such as "vegetarian" or "nut allergy." (Input) Data sent (Output) User profile saved in the database
[1525] Step 4:
[1526] The user requests food suggestions, for example, "Tell me what's good for dinner today."
[1527] Step 5:
[1528] The server refers to the user profile and selects the most suitable cooking procedure. For example, "Select vegetarian curry without nuts." (Input) User profile (Output) Selected cooking procedure
[1529] Step 6:
[1530] The server provides cooking instructions by voice or text. The server generates the selected cooking instructions as text data and converts it into voice data using a speech synthesis API. It is played on the device. (Input) Selected cooking instructions (Output) Voice response provided to the user
[1531] Recipe posting function to the community
[1532] (Processing flow)
[1533] Step 1: User enters new cooking instructions
[1534] Step 2: The device sends the recipe information to the server
[1535] Step 3: The server refines the recipe information using the generated AI model
[1536] Step 4: The server automatically generates the submission page
[1537] Step 5: Save the post to the database and publish it
[1538] (Specific explanation)
[1539] Step 1:
[1540] A user inputs new cooking instructions, for example, "new pancake recipe" into the terminal and clicks the submit button.
[1541] Step 2:
[1542] The terminal sends recipe information to the server. The terminal sends the input recipe information to the server as a data packet. (Input) Cooking procedure information (Output) Data sent to the server
[1543] Step 3:
[1544] The server refines the recipe information using a generative AI model. The server analyzes the received step-by-step information using the generative AI model and generates an appealing title and description. For example, it generates a title such as "Easy recipe for fluffy pancakes." (Input) Received step-by-step information (Output) Generated title and description
[1545] Step 4:
[1546] The server automatically generates a posting page. Based on the generated title and description, the server automatically generates a recipe posting page. (Input) Generated title and description (Output) Posting page
[1547] Step 5:
[1548] Save the submitted page in the database and make it public. Save the generated submitted page in the database and make it public. Display it in the web interface so that other users can view it. (Input) Submitted page (Output) Saved in the database and published
[1549] Data refinement based on user feedback
[1550] (Processing flow)
[1551] Step 1: User provides feedback on cooking instructions
[1552] Step 2: The device sends the feedback information to the server
[1553] Step 3: The server stores the feedback data in a database
[1554] Step 4: The server periodically aggregates and analyzes the feedback.
[1555] Step 5: The server removes cooking instructions with low ratings from the training data
[1556] (Specific explanation)
[1557] Step 1:
[1558] The user inputs feedback on the cooking procedure, for example, "This recipe was bland" into the terminal.
[1559] Step 2:
[1560] The terminal sends feedback information to the server. The terminal converts the input feedback information into data packets and sends them to the server. (Input) Feedback information (Output) Data sent to the server
[1561] Step 3:
[1562] The server saves the feedback data in a database. For example, it saves data such as "Restaurant Egg Curry", "Bland taste", and "Rating: 2". (Input) Feedback information sent. (Output) Feedback saved in the database.
[1563] Step 4:
[1564] The server periodically collects and analyzes feedback. Feedback data is periodically collected, and the ratings from many users are analyzed to determine which recipes are highly rated and which are poorly rated. (Input) Feedback stored in the database (Output) Analysis results
[1565] Step 5:
[1566] The server removes cooking steps with low ratings from the training data. By removing recipes with low ratings based on the aggregation results from the training data, the system improves its suggestion accuracy. (Input) Analysis results (Output) Updated training data
[1567] Recognizing and responding to user emotions using an emotion engine
[1568] (Processing flow)
[1569] Step 1: User makes a voice request
[1570] Step 2: The device captures the audio data and sends it to the server
[1571] Step 3: The server uses a speech recognition API to convert the data into text.
[1572] Step 4: The server analyzes the emotion information using the emotion engine
[1573] Step 5: The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal recipe.
[1574] Step 6: The server adds emotion information to the feedback data
[1575] (Specific explanation)
[1576] Step 1:
[1577] The user makes a request by voice. For example, the user may request in a tired voice, "I want a quick dinner."
[1578] Step 2:
[1579] The device captures the audio data and sends it to the server. The device captures the audio data and sends it to the server. (Input) Audio data (Output) Data sent to the server
[1580] Step 3:
[1581] The server uses a speech recognition API to convert the voice data into text data. The server uses a speech recognition API to convert the voice data into text data. (Input) Voice data (Output) Text data
[1582] Step 4:
[1583] The server uses an emotion engine to analyze emotional information. The server uses an emotion engine to analyze emotional information from the tone and pace of the user's voice. For example, emotional information indicating "fatigue" is extracted. (Input) Text data (Output) Emotional information
[1584] Step 5:
[1585] The server reflects the emotional information in the user profile and recipe selection algorithm to select the optimal cooking procedure. The server selects the optimal recipe for the user based on the emotional information. For example, it selects a "simple omelet" that can be made with minimal effort. (Input) Emotional information (Output) Selected cooking procedure
[1586] Step 6:
[1587] The server adds the emotional information to the feedback data. The emotional information is saved as feedback data along with the suggested cooking steps, and will be used for future suggestions. (Input) Emotional information (Output) Feedback data
[1588] (Application example 2)
[1589] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1590] Conventional recipe provision systems suggest recipes by analyzing users' voice requests, but they have the problem of being unable to provide recipes that reflect the user's emotions or the situation of the day. This means that users cannot receive suggestions for dishes that suit their physical condition or mood, resulting in a lack of convenience. In addition, there is a lack of a mechanism for improving the accuracy of the system by reflecting user feedback.
[1591] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1592] In this invention, the server includes means for capturing user voice, means for converting voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for analyzing the user's emotions, means for recommending recipes based on the emotion information, and means for collecting user feedback and improving the accuracy of the system. This makes it possible to provide individualized recipes according to each user's emotions, thereby realizing more convenient recipe suggestions.
[1593] The "means for capturing the user's voice" is a mechanism by which the electronic device receives the voice uttered by the user and captures it as digital data.
[1594] The "means for converting voice data into text data" is a mechanism that analyzes the captured voice data and performs a process to express the contents as a string of characters.
[1595] The "means for analyzing text data and selecting cooking recipes that meet the user's requests" refers to an algorithm and system that understands the user's requests based on the converted text data and selects appropriate cooking recipes.
[1596] The "means for converting the selected recipe into audio data" is a mechanism that performs the process of converting the selected recipe from a character string into an audio signal.
[1597] The "means for providing audio data to the user" refers to a device or program for playing back the converted audio data to the user and conveying information.
[1598] "Means for analyzing user emotions" refers to algorithms or systems that determine and analyze the user's emotional state from the tone and content of the voice.
[1599] "Means for recommending cooking recipes based on emotional information" is a system that selects and recommends the most appropriate cooking recipe based on the user's current mood and state based on the results of emotional analysis.
[1600] "Means for collecting user feedback and improving the accuracy of the system" refers to the process and mechanisms for collecting user evaluations and opinions and using them to improve the system's algorithms and database.
[1601] This invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. In particular, it aims to provide more personalized services by combining it with an emotion engine that recognizes the user's emotions.
[1602] System Configuration
[1603] Hardware
[1604] Smartphone: A device that captures the user's voice and receives the analysis results.
[1605] Server: Analyzes voice data, analyzes emotions, selects recipes, and aggregates feedback data
[1606] software
[1607] SpeechRecognition API: Converting voice data into text data
[1608] Emotion Analysis API: Analyzing user emotions from text data
[1609] Text to Speech API: Convert text data into audio data
[1610] Database: Stores user profile information, cooking recipes, and feedback data
[1611] Operation overview
[1612] 1. Audio Capture
[1613] A user speaks to their smartphone, asking, "What would you like for dinner tonight?" The smartphone captures the voice through its built-in microphone and sends this voice data to the server.
[1614] 2. Voice Recognition
[1615] When the server receives the voice data, it uses the SpeechRecognition API to convert the voice data into text data.
[1616] 3. Emotion analysis
[1617] The server then sends the converted text data to the Emotion Analysis API, which analyzes the user's emotions. The analysis result can be, for example, "I'm tired."
[1618] 4. Recipe Selection
[1619] The server selects the most suitable recipe for the user's request based on the results of sentiment analysis and the user profile database. For tired users, it selects easy-to-make dishes.
[1620] 5. Speech Synthesis
[1621] The selected cooking recipe is again treated as text data and converted into audio data using the Text to Speech API.
[1622] 6. Provide answers
[1623] The smartphone receives the converted voice data and provides the user with a spoken response such as, "You seem tired today, how about a quick sandwich?"
[1624] 7. Feedback Collection
[1625] After users try the suggested recipes, they provide feedback. The server collects this feedback data and stores it in a database. In the future, the feedback will be used to improve the accuracy of the algorithm.
[1626] Specific examples
[1627] For example, when a user makes a voice request such as "Tell me something easy to eat," the system captures the voice and the server analyzes the voice data. If the emotion engine recognizes the emotion "tired" from the user's voice, the system will suggest a quick and easy-to-prepare dish such as a "sandwich." The following is an example of a prompt sentence that can be input to the generative AI model:
[1628] Convert the user's voice into text, feed that text into a sentiment analysis engine, select the sentiment, and output a food recommendation accordingly.
[1629] Input: I'm tired today and want a quick meal.
[1630] Emotions: Tired
[1631] Output: You seem tired, would you like some lunch?
[1632] These procedures provide personalized information in response to user requests, improving user convenience.
[1633] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1634] Step 1:
[1635] The user speaks a question into the smartphone. The user's voice data is captured by the smartphone's microphone. The input is the user's voice, and the output is the captured voice data. The voice data is converted into a digital format and sent to the server.
[1636] Step 2:
[1637] The server receives the audio data. The input is digital audio data, and the output is text data. The server converts the audio data into text data using the SpeechRecognition API. In this process, the audio content is converted into a string of characters.
[1638] Step 3:
[1639] The server sends the text data to the Emotion Analysis API and analyzes the user's emotions. The input is text data and the output is emotional information. The server obtains the analysis result, for example, an emotion such as "tired."
[1640] Step 4:
[1641] The server selects the optimal recipe for the user's request based on the emotional information and the user profile database. The input is the emotional analysis results and the user profile, and the output is the selected recipe. In particular, if the user is tired, an easy-to-make dish is selected.
[1642] Step 5:
[1643] The server converts the selected recipe into audio data using the Text to Speech API. The input is the text data of the selected recipe, and the output is audio data. In this process, the recipe in text format is converted back into audio format.
[1644] Step 6:
[1645] The server sends the voice data to the smartphone. The smartphone receives this data and suggests recipes to the user by voice. The input is the converted voice data, and the output is the voice output to the user.
[1646] Step 7:
[1647] After the user has tried the suggested recipe, they input their feedback into their smartphone. The input is the user's feedback, and the output is the collected feedback data. The feedback data is sent to the server and stored in a database.
[1648] Step 8:
[1649] The server analyzes the feedback data and updates the algorithm to improve the system's accuracy. The input is the accumulated feedback data, and the output is an improved recommendation algorithm. This will enable more personalized and accurate recipe suggestions in the future.
[1650] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1651] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1652] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1653] [Fourth embodiment]
[1654] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1655] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1656] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1657] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1658] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1659] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1660] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1661] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1662] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1663] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1664] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1665] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1666] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1667] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[1668] 1. Implementing a voice-activated UI
[1669] Program Overview:
[1670] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1671] Processing Description:
[1672] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[1673] Examples:
[1674] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[1675] 2. Register user preferences and introduce the most suitable menu
[1676] Program Overview:
[1677] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1678] Processing Description:
[1679] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[1680] Examples:
[1681] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[1682] 3. Recipe posting function to the community
[1683] Program Overview:
[1684] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1685] Processing Description:
[1686] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[1687] Examples:
[1688] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[1689] 4. Data refinement based on user feedback
[1690] Program Overview:
[1691] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1692] Processing Description:
[1693] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[1694] Examples:
[1695] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[1696] In this way, the system of the present invention provides cooking recipes based on voice interaction and reflects user preferences and community participation, allowing users to enjoy a more comfortable and convenient cooking experience.
[1697] The processing flow will be explained below.
[1698] 1. Implementing a voice-activated UI
[1699] Program Overview:
[1700] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1701] Processing Steps:
[1702] Step 1:
[1703] Users speak into the device to ask questions or make requests about food, such as "What would you like for dinner tonight?"
[1704] Step 2:
[1705] The device captures the user's voice. The device uses a built-in microphone to obtain the voice data.
[1706] Step 3:
[1707] The device transmits the captured audio data to the server, and the device transmits the audio data via the Internet.
[1708] Step 4:
[1709] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[1710] Step 5:
[1711] The server analyzes the text data to understand the user's intent, and then searches for appropriate recipe information based on the analyzed data.
[1712] Step 6:
[1713] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[1714] Step 7:
[1715] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[1716] Step 8:
[1717] The server transmits the generated voice data to the terminal, and the voice data is transmitted via the Internet.
[1718] Step 9:
[1719] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[1720] 2. Register user preferences and introduce the most suitable menu
[1721] Program Overview:
[1722] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1723] Processing Steps:
[1724] Step 1:
[1725] The user inputs their food preferences and allergy information through the terminal, such as "vegetarian" or "nut allergy."
[1726] Step 2:
[1727] The device sends the input preference information to the server, and the input data is sent to the server via the Internet.
[1728] Step 3:
[1729] The server stores the received information in a database as a user profile. Profile data is created and stored for each user.
[1730] Step 4:
[1731] The user inputs a specific food request into the device, for example, asking "What do you recommend for lunch today?"
[1732] Step 5:
[1733] The device captures the audio and sends it to a server, which then sends the audio data over the internet.
[1734] Step 6:
[1735] The server passes the received voice data to the voice recognition API and converts it into text data. The text data obtained from the voice recognition API is analyzed.
[1736] Step 7:
[1737] The server references the user profile and uses a filtering algorithm to select the best recipes, based on the user's preferences and allergies.
[1738] Step 8:
[1739] The selected recipe information is passed to the speech synthesis API and converted into voice data, which is then returned to the server.
[1740] Step 9:
[1741] The server sends the audio data to the terminal. The audio data is sent via the Internet.
[1742] Step 10:
[1743] The terminal plays the audio data to provide information to the user. The audio data is played through the terminal's speaker.
[1744] 3. Recipe posting function to the community
[1745] Program Overview:
[1746] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1747] Processing Steps:
[1748] Step 1:
[1749] The user inputs new cooking recipe information into the terminal, for example, "original cookie recipe."
[1750] Step 2:
[1751] The terminal sends the input recipe information to the server, and the input data is sent to the server via the Internet.
[1752] Step 3:
[1753] The server passes the received recipe information to the generation AI model for refinement. The generation AI analyzes the recipe information and generates a more attractive title and description.
[1754] Step 4:
[1755] The server saves the generated submission page data, including the title and description, in a database. The generated submission page is then prepared for public viewing.
[1756] Step 5:
[1757] The server generates a URL for publishing and configures the publishing settings. The submission page will then be viewable by the general public.
[1758] Step 6:
[1759] General users can view published recipe pages through their devices. They can also view the posting page using the device's browser.
[1760] 4. Data refinement based on user feedback
[1761] Program Overview:
[1762] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1763] Processing Steps:
[1764] Step 1:
[1765] The user enters feedback about the recipe they tried into the device, for example, "This recipe was bland."
[1766] Step 2:
[1767] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[1768] Step 3:
[1769] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[1770] Step 4:
[1771] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[1772] Step 5:
[1773] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[1774] Step 6:
[1775] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[1776] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is capable of always providing the latest, high-quality recipe information while reflecting the user's preferences and community participation.
[1777] Example 1
[1778] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1779] Conventional cooking assistant systems are limited to analyzing users' voice input and providing cooking recipes, but lack the ability to reflect users' preferences and allergy information or filter recipes based on user feedback. Furthermore, they lack the ability to automatically refine recipe posts when users post new recipes. This has made improving the user experience a challenge.
[1780] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1781] In this invention, the server includes means for capturing a user's voice, means for converting the voice data into text data, means for analyzing the text data and selecting a recipe according to the user's request, means for converting the selected recipe into voice data, means for providing the voice data to the user, means for collecting user feedback and storing it in a database, means for excluding low-rated recipes from the training data, means for polishing the recipe entered by the user using a generative AI model, and means for inputting prompt sentences into the generative AI model. This makes it possible to provide optimal recipes that reflect the user's preferences and feedback, and to automatically polish new recipes.
[1782] A "means for capturing voice" is a device or method for capturing a user's spoken voice as digital data.
[1783] The "means for converting voice data into text data" is a process for converting captured voice data into text information using voice recognition technology.
[1784] "Means for analyzing text data and selecting cooking recipes that meet user requests" refers to a method that uses natural language processing technology to understand user requests from text data and select the optimal cooking recipe to meet those requests.
[1785] The "means for converting the selected recipe into voice data" is a method for converting text-format recipe information into a format that can be output as voice using voice synthesis technology.
[1786] A "means for providing audio data to a user" is a method for transmitting audio data to a user's device and playing it back.
[1787] The "means for collecting user feedback and storing it in a database" refers to a method for collecting, organizing, and storing the ratings and comments provided by users in a database.
[1788] "Means for excluding low-rated recipes from learning data" refers to the process of analyzing user rating data and removing recipes that have received a certain level of low rating from the system's recommendation candidates.
[1789] "Means for brushing up recipes entered by users using generative AI models" refers to a method of improving recipe information provided by users using generative AI technology to make the content more appealing.
[1790] A "means for inputting prompt text into a generative AI model" is a method for inputting text containing specific instructions or information into a generative AI model and obtaining output based on that text.
[1791] MODE FOR CARRYING OUT THE INVENTION
[1792] The present invention relates to a voice interactive cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes. Specific embodiments of the present invention will be described below.
[1793] 1. System Configuration
[1794] This system consists of a device used by the user and a central server. The device has a microphone to capture the user's voice and a speaker to play back the voice data. The server uses a speech recognition API, a speech synthesis API, a generative AI model, and a database to select and respond to the user's request with the optimal recipe.
[1795] 2. Implementing a Voice-Based Interactive UI
[1796] When a user asks, "What would you like for dinner today?", the device captures the voice and sends the voice data to the server. The server uses a speech recognition API such as Amazon Transcribe to convert the voice data into text data. The converted text data is then analyzed to select a cooking recipe that meets the user's request. The results of this selection are then prepared as text data and converted into voice data using a speech synthesis API such as Amazon Polly. The voice data is then sent to the device and provided to the user through the speaker.
[1797] Examples:
[1798] When a user asks, "What's your recommended dish for tonight?" the system will respond aloud with, "How about spaghetti arrabiata? It's easy and delicious."
[1799] 3. Recipe selection that reflects user preferences
[1800] Users enter their food preferences and allergy information into their device and send it to the server. The server stores this information in a database and manages it as a user profile. When a user requests recipe suggestions, the server references this user profile and selects the most suitable recipe.
[1801] Examples:
[1802] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a nut-free vegetarian curry."
[1803] 4. Automatically improve recipe posts
[1804] When a user enters a new recipe into their device and posts it, the server passes the information to a generative AI model to generate an appealing title and description, and the resulting post page is stored in a database and made available for other users to view.
[1805] Examples:
[1806] When a user enters and posts a "new pancake recipe," the system automatically generates a title such as "Easy recipe for fluffy pancakes" and an introductory text, creating a posting page. For example, the prompt for the generative AI model could be "Generate an attractive cooking recipe posting page based on the following content:"
[1807] 5. Data refinement based on user feedback
[1808] The feedback provided by the user is sent from the device to the server and stored in a database. The server periodically aggregates the feedback data and removes low-rated recipes from the learning data. This allows the system to provide only highly rated recipes to users in the future.
[1809] Examples:
[1810] If a user provides feedback such as "this recipe was bland," the server will filter out that recipe if similar feedback continues, improving the accuracy of the system's recipe selection algorithm.
[1811] As described above, the voice-activated cooking assistant system of the present invention provides optimal recipes for users by combining voice input analysis, user profile utilization, recipe refinement using generative AI models, and feedback collection and analysis, thereby enabling users to enjoy a more comfortable and convenient cooking experience.
[1812] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1813] Step 1:
[1814] The user provides voice input. The user says, "What would you like for dinner tonight?"
[1815] Step 2:
[1816] The device captures voice data and sends it to the server. The voice data is obtained in digital form through a microphone and sent to the server using an HTTP request. The input is the user's voice and the output is digital voice data sent to the server.
[1817] Step 3:
[1818] The server receives the voice data and passes it to the voice recognition API. The voice data is stored in cloud storage, and its URL is used as input for the voice recognition API. The input is the URL of the voice data, and the output is text data.
[1819] Step 4:
[1820] The server converts the audio data into text data using a speech recognition API. The text is extracted from the JSON format data returned as an API response. The input is the URL of the audio data and the API response, and the output is the extracted text data.
[1821] Step 5:
[1822] The server analyzes the text data, understands the user's request, and selects the most suitable recipe. It uses natural language processing technology to extract the "dish name" and "conditions" from the text data, and queries the database to search for matching recipes. The input is text data, and the output is information about the selected recipe.
[1823] Step 6:
[1824] The server prepares the selected recipe as text data, structuring it as "Spaghetti Arrabbiata is recommended." The input is the selected recipe information, and the output is the response text data.
[1825] Step 7:
[1826] The server uses a speech synthesis API such as Amazon Polly to convert text data into speech data. The request to the API includes text data and speech settings. The input is text data, and the output is the generated speech data.
[1827] Step 8:
[1828] The server sends the generated voice data to the device and provides it to the user through the device's speaker. The voice data is sent to the device as an HTTP response, and the device plays the data. The input is the voice data, and the output is the voice response provided to the user.
[1829] Step 9:
[1830] The user inputs their food preferences and allergy information into the device, which then sends it to the server. The input information is sent in JSON format. The input is the user's preferences and allergy information, and the output is the user profile data sent to the server.
[1831] Step 10:
[1832] The server stores the received information in a database. It associates the information with the user ID and adds it to the corresponding record in the database. The input is the user profile data, and the output is the user information stored in the database.
[1833] Step 11:
[1834] The user inputs a new cooking recipe and the device sends the information to the server. The input information is sent in JSON format. The input is the cooking recipe entered by the user, and the output is the recipe data sent to the server.
[1835] Step 12:
[1836] The server inputs the received recipe information into the generative AI model. The generative AI model then inputs text data containing specific prompts. The input is the user's recipe information and prompt, and the output is the generated title and description.
[1837] Step 13:
[1838] The server saves the generated post, including the title and description, in a database where it can be viewed by other users. The input is the generated post data, and the output is the post saved in the database.
[1839] Step 14:
[1840] The user inputs feedback for the recipes they have tried, and the terminal sends this information to the server. The input is the user's feedback, and the output is the feedback data sent to the server.
[1841] Step 15:
[1842] The server saves the feedback data in a database. It saves the rating scores and comments in the appropriate fields. The input is the feedback data, and the output is the rating information saved in the database.
[1843] Step 16:
[1844] The server periodically aggregates the feedback data and removes poorly rated recipes from the training data. It then runs a data analysis process to identify and remove poorly rated recipes. The input is the feedback data, and the output is updated training data.
[1845] The above are the specific processing steps of the program of this system.
[1846] (Application example 1)
[1847] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1848] When users use food delivery services in their daily lives, they face challenges such as the time it takes to select the optimal menu and the need to individually confirm preferences and allergy information. Furthermore, when users select a delivery menu, they are unable to make the optimal selection based on their past order history and preferences. Recipe posting and rating functions for sharing with other users are also time-consuming and labor-intensive, as they must be done manually.
[1849] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1850] In this invention, the server includes a means for capturing the user's voice, a means for converting the voice data into text data, and a means for analyzing the text data and selecting a recipe or delivery menu item according to the user's request. This allows the user to quickly request a delivery menu item by voice, streamlining the ordering process. The server also includes a function for storing user preferences and allergy information in a database and recommending delivery menu items based on that information, enabling optimal suggestions for each individual user. Furthermore, by adding a function for refining and automatically generating recipes entered by the user, users can efficiently post and rate recipes for sharing with other users.
[1851] "User" refers to a person who uses the system to make a voice request and receive a cooking recipe or delivery menu.
[1852] "Means for capturing audio" refers to a mechanism for collecting the user's audio using a device such as a microphone.
[1853] "Voice data" refers to data that expresses the user's voice as digital information.
[1854] "Means for converting into text data" refers to a mechanism for converting voice data into text information using voice recognition technology.
[1855] "Text data" refers to digital information that has been converted from audio data into text form.
[1856] "User request" refers to a request or question uttered by a user.
[1857] "Cooking recipes or delivery menus" refers to the food preparation methods and delivery options suggested by the system.
[1858] "Means for selection" refers to a mechanism for selecting the most suitable cooking recipe or delivery menu based on the user's request.
[1859] "Means for converting into voice data" refers to a mechanism that uses voice synthesis technology to convey the selected recipe or delivery menu to the user in voice form.
[1860] "Means of providing" refers to devices such as speakers and headphones used to deliver audio data to users.
[1861] "User's food preferences and allergy information" refers to information about dietary preferences and ingredients to avoid that the user has registered in the system.
[1862] A "database" refers to a system for centrally managing user information and data processed by the system.
[1863] "Recommendation means" refers to a mechanism for suggesting optimal recipes or delivery menus by taking into consideration the user's food preferences and allergy information.
[1864] "Inputted recipe" refers to information on food cooking methods and ingredients that a user has registered or provided to the system.
[1865] "Means of polishing and automatically generating a post page" refers to a mechanism that uses a generative AI model to format information entered by a user into an attractive form and automatically convert it into a publishable format.
[1866] To implement this invention, it is necessary to build a system that proposes optimal cooking recipes or delivery menus based on a user's voice request. This system is configured around a voice-interactive user interface.
[1867] System configuration
[1868] 1. How to capture the user's voice:
[1869] Hardware: A smartphone or tablet with a built-in microphone.
[1870] Role: Collects user voice in real time.
[1871] 2. How to convert audio data to text data:
[1872] Software: speech_recognition library.
[1873] Role: Converts voice data into text.
[1874] 3. A method for analyzing text data and selecting cooking recipes or delivery menus according to user requests:
[1875] Technologies used: Natural Language Processing (NLP) and databases.
[1876] Role: Analyzes user requests and selects the most suitable cooking recipe or delivery menu.
[1877] 4. Means for converting selected information into audio data:
[1878] Software: pyttsx3 library.
[1879] Role: Converts selected recipes or delivery menus into audio data.
[1880] 5. Means of providing audio data to the user:
[1881] Hardware: Smartphone speakers and headphones.
[1882] Role: Communicates information to the user audibly.
[1883] 6. A means to store user's food preferences and allergy information in a database and make recommendations:
[1884] Database: A specialized user profile database.
[1885] Role: Stores user preferences and allergy information to improve future recommendations.
[1886] 7. How to polish a recipe entered by a user and automatically generate a posting page:
[1887] Technology used: Content generation using generative AI models.
[1888] Role: To polish recipe information provided by users to make it more appealing and publish it as an automatically generated posting page.
[1889] Program operation explanation
[1890] The server captures voice data input by the user via a smartphone or tablet and converts the voice data into text using the speech_recognition library. It then analyzes the text data using natural language processing technology to select the optimal recipe or delivery menu item for the user's request. This selection process takes into account the user's preferences and allergies. The selected information is then converted back into voice data using the pyttsx3 library and delivered to the user via the smartphone's speakers or headphones.
[1891] Additionally, when a user enters a new recipe, the recipe information is refined using a generative AI model and automatically generated into an attractive posting page, which is then saved in a database and made publicly available for other users to view.
[1892] Specific examples
[1893] When a user speaks to their smartphone, "What delivery do you recommend for dinner tonight?", the system responds by voice, "Would you like pizza or sushi? Which would you like to order?" If the user selects pizza, the system will suggest the best pizza delivery service.
[1894] Prompt Sentence Examples
[1895] "The following input should contain a list of restaurants that offer pizza and sushi delivery. If the user selects pizza, output the names of the top 5 restaurants listed.
[1896] Input: I want pizza.
[1897] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1898] Step 1:
[1899] The user issues a voice request to the smartphone. For example, say, "What delivery do you recommend for dinner tonight?" The voice is captured by the smartphone's microphone.
[1900] Step 2:
[1901] The device sends the captured audio data to the server. The input is audio data, which is then sent to the server.
[1902] Step 3:
[1903] The server converts the received voice data into text data using the speech_recognition library. The input is voice data, and the output is the converted text data. This conversion uses speech recognition technology.
[1904] Step 4:
[1905] The server analyzes the text data using natural language processing (NLP) technology to understand the user's request. The input is text data, and the output is the analyzed request data. As a concrete example, it identifies that the user's request is for "dinner delivery."
[1906] Step 5:
[1907] The server retrieves the user's preferences and allergy information from the database, compares it with the analysis results, and recommends the most suitable recipe or delivery menu. The input is the analyzed request content and user profile data, and the output is the selected recipe or delivery menu.
[1908] Step 6:
[1909] The server converts the selected recipe or delivery menu into voice data using the pyttsx3 library. The input is text-based recipe or menu information, and the output is voice data.
[1910] Step 7:
[1911] The server sends the generated voice data to the terminal. The input is the voice data, and the output is the transfer of the voice data to the terminal.
[1912] Step 8:
[1913] The device then provides the received voice data to the user through a speaker or headphones. The input is voice data, and the output is audible audio for the user. Specifically, the smartphone responds with a voice message asking, "Would you like pizza or sushi? Which would you like to order?"
[1914] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1915] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking recipes, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[1916] 1. Implementing a voice-activated UI
[1917] Program Overview:
[1918] Using speech recognition and speech synthesis APIs, we will build a system that analyzes voice input from users in real time and generates appropriate responses.
[1919] Processing Description:
[1920] First, the user asks verbally, "What would you like for dinner today?" The device captures this voice and sends it to the server as voice data. The server passes the received voice data to a voice recognition API and converts it into text data. After obtaining the text data, the server analyzes the user's request and selects an appropriate recipe. The selected recipe is converted back into text data, and then converted into voice data using a voice synthesis API. This voice data is sent to the device, and the answer is provided to the user via voice.
[1921] Examples:
[1922] When a user asks aloud, "What's your recommended dish for tonight?", the system responds aloud, "How about spaghetti arrabiata? It's easy and delicious."
[1923] 2. Register user preferences and introduce the most suitable menu
[1924] Program Overview:
[1925] The system implements an algorithm that stores the user's pre-set food preferences and allergy information in a database and recommends recipes based on that information.
[1926] Processing Description:
[1927] Users enter their preferences and allergy information into their device and send it to the server. The server stores the received information in a database as a user profile. Each time a user makes a request for cooking suggestions, the server refers to this user profile and selects the most suitable recipe. The selected recipe is then provided to the user via voice or text.
[1928] Examples:
[1929] If a user sets the answer as "I'm a vegetarian and have a nut allergy," the system will suggest "I recommend a vegetarian curry without nuts."
[1930] 3. Recipe posting function to the community
[1931] Program Overview:
[1932] We provide a system that uses generative AI to automatically generate posting pages so that users can easily post cooking recipes.
[1933] Processing Description:
[1934] When a user inputs a new recipe, the device sends the information to the server. The server then refines the received recipe information using a generative AI model and automatically generates an appealing title and description. The generated post page is saved in a database and made publicly available for other users to view.
[1935] Examples:
[1936] When a user enters and posts a "new pancake recipe," the system generates a title and description, such as "Easy recipe for fluffy pancakes," and automatically creates a recipe page.
[1937] 4. Data refinement based on user feedback
[1938] Program Overview:
[1939] We will implement a mechanism to collect feedback from users, update recipe ratings, and exclude low-rated recipes from the system's learning data.
[1940] Processing Description:
[1941] When a user enters feedback on a recipe they have tried, the device sends that information to the server. The server stores the feedback data in a database and periodically aggregates and analyzes it. Based on the results of this analysis, recipes with low ratings are removed from the training data, improving the accuracy of the algorithm.
[1942] Examples:
[1943] Users provide feedback such as "This recipe was bland." The system aggregates this information and, if similar low ratings continue, removes the recipe, allowing it to provide only more highly rated recipes in the future.
[1944] 5. Recognizing and responding to user emotions using an emotion engine
[1945] Program Overview:
[1946] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[1947] Processing Description:
[1948] When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server converts the voice data into text using a speech recognition API and then analyzes the user's emotions using an emotion engine. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking recipe that best suits the user's emotions is selected. In addition, emotional information is added to the feedback data to help with future recipe suggestions.
[1949] Examples:
[1950] If a user requests in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine will recognize the user's tired emotions and suggest a simple omelet that requires little effort to make."
[1951] Through the above process, the system of the present invention can provide cooking recipes through voice dialogue, always providing the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[1952] The processing flow will be explained below.
[1953] Recognizing and responding to user emotions using an emotion engine
[1954] Program Overview:
[1955] The system recognizes emotions from the user's voice and suggests optimal recipes based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy.
[1956] Processing Steps:
[1957] Step 1:
[1958] The user speaks into the device to make a cooking request, such as, "I'm tired. Do you have a quick dinner?"
[1959] Step 2:
[1960] The device captures the user's voice. The device's microphone is used to obtain the voice data.
[1961] Step 3:
[1962] The device sends the captured audio data to a server, which then transmits the audio data over the Internet.
[1963] Step 4:
[1964] The server passes the received voice data to the voice recognition API and converts it into text data. The voice recognition API analyzes the voice data and returns it to the server as text data.
[1965] Step 5:
[1966] The server passes the text data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the text data and the intonation and tone of the voice to recognize the user's emotions.
[1967] Step 6:
[1968] The server reflects the analyzed emotional information in the user profile, which is then added to the user profile and used to select the next recipe.
[1969] Step 7:
[1970] The server selects the optimal recipe based on the user's request and emotional information. The selection algorithm takes the user's emotional information into account.
[1971] Step 8:
[1972] The server stores the selected recipe information as text data, and retrieves related recipe information from the database.
[1973] Step 9:
[1974] The server passes the text data to the speech synthesis API, which converts it into speech data. The speech synthesis API returns the text data to the server as speech data.
[1975] Step 10:
[1976] The server transmits the generated voice data to the terminal, which then transmits the voice data via the Internet.
[1977] Step 11:
[1978] The device uses the audio playback function to provide the received audio data to the user, and the audio is played through the device's speakers or headphones.
[1979] Step 12:
[1980] The user tries the recipe and enters feedback into the device, for example, "This recipe was very easy and delicious."
[1981] Step 13:
[1982] The terminal transmits the input feedback data to the server, which then transmits the feedback data via the Internet.
[1983] Step 14:
[1984] The server stores the received feedback data in a database and aggregates the feedback data for each recipe.
[1985] Step 15:
[1986] The server periodically analyzes the collected feedback data, identifies recipes with low ratings, and calculates the ratings using a data analysis algorithm.
[1987] Step 16:
[1988] The server performs an update process to remove low-rated recipes from the training data, so that they are not used in the next recommendation or generation task.
[1989] Step 17:
[1990] The server uses the updated learning data to recommend new recipes, improving the system's algorithm and providing more appropriate recipes.
[1991] Through the above processing steps, the system of the present invention provides cooking recipes through voice dialogue, and is able to provide the latest, high-quality recipe information while reflecting the user's preferences and feelings, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[1992] Example 2
[1993] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1994] Conventional cooking assistant systems have difficulty reflecting users' preferences and allergy information when providing appropriate cooking instructions in response to a user's voice request. They also lack the ability to post the cooking instructions entered by the user in an appealing format, or the ability to improve the system's accuracy based on user feedback. Furthermore, they are unable to suggest recipes that take the user's emotions into account, making it difficult to provide personalized services.
[1995] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's voice, means for converting voice data into text data, means for analyzing the text data and selecting cooking procedures according to the user's request, means for converting the selected cooking procedures into voice data, means for providing the voice data to the user, means for storing the user's cooking preferences and allergy information in a database and recommending cooking procedures based on the stored information, means for improving the cooking procedures entered by the user using a generative AI model and automatically generating a posting page, means for collecting user feedback, updating the ratings of the cooking procedures, and excluding low-rated cooking procedures from the system's learning data, and means for recognizing emotions from the user's voice and suggesting optimal cooking procedures based on the emotions, and means for adding user emotion information to the feedback data. This makes it possible to always provide the latest, high-quality cooking procedures taking into account the user's preferences and emotions.
[1996] "User" refers to any individual or group of people who operate the system and use the cooking recipe suggestions and voice interface.
[1997] "Voice capture means" refers to a device or software mechanism for capturing a user's speech as digital data.
[1998] "Voice data" refers to data that is a digital representation of a user's voice.
[1999] "Text data" refers to data in the form of a string of characters obtained by analyzing voice data.
[2000] "Means for analyzing text data" refers to software or algorithms that analyze text data converted from speech and understand the user's intent and request.
[2001] A "cooking recipe" refers to a document that lists the steps or methods for preparing a dish.
[2002] "Means for converting into voice data" refers to software or devices that convert text data back into voice format using voice synthesis technology.
[2003] A "database" refers to a system for efficiently storing and managing various data such as user preferences and allergy information.
[2004] "Recommendation tools" refers to algorithms and software functions that select optimal cooking instructions based on a user's profile information and past data.
[2005] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generates text, images, etc. based on user input.
[2006] "Posting Page" refers to a web page or part of an application where a user can share and publish cooking instructions.
[2007] "Feedback" refers to opinions such as ratings and comments provided by users to the system.
[2008] An "emotion engine" refers to a technology or software component that analyzes emotions from a user's voice or text and understands the user's emotional state.
[2009] "Information processing device" refers to the entire system including a computer and peripheral devices for inputting, processing, and outputting data.
[2010] The present invention relates to a voice-activated cooking assistant system that analyzes a user's voice requests and provides optimal cooking procedures, and aims to provide a more personalized service by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention are described below.
[2011] Implementing a voice-interactive UI
[2012] The program for this system primarily uses speech recognition and speech synthesis APIs to build a mechanism for analyzing voice input from users in real time and generating appropriate responses. Specifically, it uses the Google Cloud Speech-to-Text API as the speech recognition API and the Google Cloud Text-to-Speech API as the speech synthesis API. When a user asks a question out loud, such as "What would you like for dinner tonight?", the device captures this speech and sends it to the server as audio data. The server passes the received audio data to the speech recognition API, converts it into text data, and analyzes the user's request. After selecting the appropriate cooking instructions, it converts it into audio data using the speech synthesis API and sends it to the device, where it provides the user with a spoken response.
[2013] Register user preferences and introduce the most suitable menu
[2014] The system implements an algorithm that recommends cooking steps based on a database of cooking preferences and allergy information preset by the user. The user enters their preferences and allergy information into the device and sends it to the server. The server stores the received information in a database and creates a user profile. When the user makes a request for cooking suggestions, the server refers to the user profile and selects appropriate cooking steps. These are then provided to the user via voice or text.
[2015] Recipe posting function to the community
[2016] We provide a mechanism to automatically generate posting pages using a generative AI model so that users can easily post cooking instructions. When a user enters new cooking instructions, the device sends the information to a server. The server then uses the generative AI model to refine the received instructions and automatically generate an attractive title and description. The generated posting page is saved in a database and made publicly available for other users to view.
[2017] Data refinement based on user feedback
[2018] We implement a mechanism to collect feedback from users, update the ratings of cooking steps, and exclude low-rated steps from the system's learning data. When a user enters feedback for a cooking step they have tried, the device sends the information to a server. The server stores the feedback data in a database and periodically aggregates and analyzes it. By excluding low-rated steps from the learning data in this way, the accuracy of the algorithm is improved.
[2019] Recognizing and responding to user emotions using an emotion engine
[2020] The system recognizes emotions from the user's voice and suggests optimal cooking steps based on those emotions. It also adds the user's emotional information to the feedback data to improve the system's accuracy. When a user makes a request by voice, the voice data is captured by the device and sent to the server. The server uses a speech recognition API to convert it into text data and then uses an emotion engine to analyze the user's emotions. The analyzed emotional information is reflected in the user profile and recipe selection algorithm, and the cooking steps that best suit the user's emotions are selected. For example, if a user requests, in a tired voice, "I want an easy-to-make dinner," the system will respond by saying, "The emotion engine recognizes the user's fatigue and suggests a simple omelet that requires little effort to make."
[2021] Through the above process, the system of the present invention can provide cooking instructions through voice interaction, always providing the latest, high-quality cooking instructions while reflecting the user's preferences and emotions, thereby allowing the user to enjoy a more comfortable and personalized cooking experience.
[2022] Prompt Sentence Examples
[2023] When a user asks, "What's your recommended dish for tonight?", the system responds, "How about spaghetti arrabiata? It's easy and delicious." In this way, the system smoothly suggests the appropriate cooking procedure for a specific situation.
[2024] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2025] Implementing a voice-interactive UI
[2026] (Processing flow)
[2027] Step 1: User makes a voice request
[2028] Step 2: The device captures the audio data and sends it to the server
[2029] Step 3: The server uses a speech recognition API to convert the data into text.
[2030] Step 4: The server analyzes the request
[2031] Step 5: The server selects the cooking instructions based on the request.
[2032] Step 6: The server converts the selected recipe into text data.
[2033] Step 7: The server uses the speech synthesis API to convert the data into audio.
[2034] Step 8: Respond to the user by sending voice data to the device
[2035] (Specific explanation)
[2036] Step 1:
[2037] The user makes a voice request. For example, the user says, "What would you like for dinner tonight?"
[2038] Step 2:
[2039] The device captures the voice data and sends it to the server. Specifically, it uses the device's microphone to capture the user's voice as digital data. The device then sends the captured voice data to the server via the Internet. (Input) Voice data (Output) Data sent to the server
[2040] Step 3:
[2041] The server converts the received voice data into text data using a speech recognition API. For example, the server analyzes the received voice data using a speech recognition API such as Google Cloud Speech-to-Text API. This converts the voice into text data. (Input) Voice data (Output) Text data
[2042] Step 4:
[2043] The server analyzes the request. The server passes the text data to a natural language processing engine for analysis and understands the user's request. For example, it extracts keywords such as "dinner," "recommendations," and "tell me." (Input) Text data (Output) Analysis results
[2044] Step 5:
[2045] The server selects the cooking procedure according to the request. Based on the analysis results, the server refers to the user profile and database to select the appropriate cooking procedure. For example, "Spaghetti Arrabbiata" is selected. (Input) Analysis results (Output) Selected cooking procedure
[2046] Step 6:
[2047] The server converts the selected cooking instructions into text data. The selected cooking instructions are formatted for presentation to the user. For example, a sentence such as "Today's recommended dish is spaghetti arrabbiata" is generated. (Input) Selected cooking instructions (Output) Text data
[2048] Step 7:
[2049] The server converts the text data into audio data using a speech synthesis API. The server converts the text data into audio data using a speech synthesis API such as Google Cloud Text-to-Speech API. (Input) Text data (Output) Audio data
[2050] Step 8:
[2051] The server responds to the user by transmitting the generated voice data to the terminal, and the terminal plays back the received voice data to provide a response to the user.
[2052] (Input) Voice data (Output) Voice response provided to the user
[2053] Register user preferences and introduce the most suitable menu
[2054] (Processing flow)
[2055] Step 1: User registers preferences and allergy information
[2056] Step 2: The device sends the information to the server
[2057] Step 3: The server saves the information to a database
[2058] Step 4: User requests dish suggestions
[2059] Step 5: The server looks up the user profile and selects the best recipe
[2060] Step 6: The server provides the recipe via voice or text
[2061] (Specific explanation)
[2062] Step 1:
[2063] The user registers their preferences and allergy information, for example, by typing "I'm a vegetarian and have a nut allergy" into the terminal.
[2064] Step 2:
[2065] The device sends information to the server. Specifically, the device converts the input information into data packets and sends them to the server. (Input) Preference and allergy information (Output) Data sent to the server
[2066] Step 3:
[2067] The server saves the information in a database. The server analyzes the received user information and saves it in a database as a user profile. For example, it classifies the information by adding tags such as "vegetarian" or "nut allergy." (Input) Data sent (Output) User profile saved in the database
[2068] Step 4:
[2069] The user requests food suggestions, for example, "Tell me what's good for dinner today."
[2070] Step 5:
[2071] The server refers to the user profile and selects the most suitable cooking procedure. For example, "Select vegetarian curry without nuts." (Input) User ...
Claims
1. means for capturing the user's voice; means for converting voice data into text data; A means for analyzing the text data and selecting a cooking recipe according to a user's request; A means for converting the selected cooking recipe into audio data; A system including means for providing audio data to a user.
2. 2. The system according to claim 1, further comprising means for storing information on the user's cooking preferences and allergies in a database and recommending cooking recipes based thereon.
3. The system according to claim 1 , further comprising means for improving a cooking recipe input by a user and automatically generating a posting page.
4. The system according to claim 1, further comprising means for collecting feedback on cooking recipes input by users, aggregating the feedback data at regular intervals, and updating the recipe ratings.
5. 10. The system of claim 1, further comprising means for removing poorly rated recipes from the training data based on the feedback data to improve the accuracy of the system's algorithms.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A