system

The system addresses the lack of personalized coffee customization by using generative models and voice recognition to tailor coffee brewing to individual preferences and health conditions, enhancing the coffee experience.

JP2026074884APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing coffee extraction methods do not fully utilize the characteristics of individual coffee beans, and consumers lack easy means to customize coffee according to their preferences and health conditions.

Method used

A system that utilizes characteristic data of coffee beans to generate optimal recipes using a generative model, provides customization options through voice recognition, and integrates with household appliances for personalized brewing.

Benefits of technology

Enables consumers to easily customize coffee based on their preferences and health conditions, maximizing the characteristics of coffee beans and providing a richer coffee experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074884000001_ABST
    Figure 2026074884000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving characteristic data about coffee beans, A means of using a generative model that generates an optimal coffee recipe based on the characteristic data, Means for adjusting the coffee extraction based on the generated recipe, A means of presenting customization options based on user information, A means for receiving and processing user commands using speech recognition technology, A means of operating in conjunction with other household electrical appliances, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Regarding coffee that many consumers use daily, there is a problem that an extraction method that fully utilizes the characteristics of individual coffee beans has not been provided. In addition, means for consumers to easily customize coffee according to their preferences and health conditions are limited. Therefore, there is a demand to maximize the characteristics of coffee beans and provide coffee according to the needs of each user, but there is a problem that the conventional methods for achieving this are burdensome for consumers.

Means for Solving the Problems

[0005] This invention provides a system that receives characteristic data about coffee beans and uses it to generate an optimal coffee recipe using a generative model. Based on the generated recipe, it is possible to automatically adjust the coffee brewing conditions. Furthermore, it has the function to present customization options in an interactive format according to the user's preferences and health condition, and to execute them based on user instructions using voice recognition technology. In addition, this system can be linked with other household electrical appliances to provide users with a richer coffee experience.

[0006] "Coffee beans" are the seeds of a plant used to extract coffee beverages, and their characteristics vary depending on the origin and type of coffee bean.

[0007] "Characteristic data" refers to specific information about coffee beans, such as type, origin, and roasting level.

[0008] A "generative model" is a form of algorithm or artificial intelligence used to generate the optimal coffee recipe based on input data.

[0009] A "recipe" is information that describes the extraction method and procedure for bringing out the best flavor when using a specific type of coffee bean.

[0010] "Customization options" refer to choices presented to adjust the taste, strength, and other aspects of coffee based on the user's preferences and health condition.

[0011] "Voice recognition technology" is a technology that analyzes the voice spoken by a user, interprets their intent, and executes instructions accordingly.

[0012] "Household electrical appliances" is a general term referring to various devices and appliances that use electricity within the home. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, a numbered communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system aimed at providing the optimal coffee by utilizing information about coffee beans. The system mainly consists of three components: a server, a terminal, and a user, each with its own role.

[0035] Server roles and operations

[0036] The server maintains a database that manages and updates characteristic data such as coffee bean type, origin, and roast level. It also uses a generative model based on the collected data to generate the optimal coffee recipe. When a request is sent from a user via the terminal, the server processes the relevant data and performs the necessary calculations to provide a recipe tailored to the user's preferences and health condition.

[0037] Terminal role and operation

[0038] The terminal receives coffee bean information and personal preferences entered by the user and provides an interface for sending this information to the server. Furthermore, it receives recipe information sent from the server and generates an operation screen for the user to customize based on that information. It also has voice recognition technology, which allows it to analyze the voice spoken by the user and execute the commands necessary for the brewing process.

[0039] User roles and actions

[0040] Users input information about the coffee beans they have on hand into the device, and also record their taste preferences and health status. Furthermore, they can adjust the recipe using the customization options presented on the device. Finally, by giving a voice command to start brewing and operating the system, they can obtain coffee of their desired quality.

[0041] Specific example

[0042] For example, consider a user who wants to use "Ethiopian Yirgacheffe beans" and brew coffee to their liking. First, the user enters the details of the beans on the device, and then specifies their preferred taste. The device sends this information to the server, which generates the optimal brewing method and sends it back to the device. The device makes fine adjustments according to the user's wishes and can finally initiate brewing via voice recognition. Through this process, the user can easily enjoy a customized cup of coffee tailored to their needs.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user inputs information about their coffee beans (e.g., origin and type), as well as their own taste preferences and health status, into the terminal. The terminal formats this information and stores it in a way that allows for subsequent data processing.

[0046] Step 2:

[0047] The terminal generates an API request to send the entered data to the server. This request contains the user's input data, which the server uses as the basis for processing.

[0048] Step 3:

[0049] The server receives requests sent from terminals and analyzes the data. Based on the analysis results, it uses a generative model to generate recipes suitable for the characteristics of the coffee beans. The generated recipes include recommended brewing conditions.

[0050] Step 4:

[0051] The server sends the generated recipe to the terminal. The terminal prepares a user interface to display the received recipe in a user-friendly format.

[0052] Step 5:

[0053] Users can view the generated recipe along with the provided customization options through their device's display screen. They can further adjust the flavor and concentration according to their preferences and then see the results.

[0054] Step 6:

[0055] Once the adjustments are complete, the user uses voice commands to give instructions to the device, such as "Please prepare some coffee." The voice is then converted to text by the device's speech recognition technology.

[0056] Step 7:

[0057] The terminal uses voice recognition to confirm the user's instructions and then controls devices such as a coffee grinder to begin the actual coffee brewing process. Once brewing is complete, the user is notified, and the finished coffee is served.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] Conventional beverage preparation systems have faced challenges in generating optimal recipes tailored to individual users' preferences and health conditions, and in easily brewing and customizing them. Furthermore, they lacked sufficient voice control and integration capabilities with other household electrical appliances, resulting in poor usability.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes a device for receiving characteristic data about coffee beans, a device for utilizing a generation model that generates optimal beverage preparation instructions based on the characteristic data, and a device for adjusting the extraction of the beverage based on the generated instructions. This enables the generation of recipes tailored to the individual needs of the user and facilitates easy extraction operations.

[0063] "Characteristic data about coffee beans" refers to data that includes information such as the type of coffee bean, its origin, roast level, and quality.

[0064] A "generative model" is a machine learning model that generates optimal beverage preparation instructions for the user based on the data received.

[0065] "Beverage preparation instructions" are instructions that specify the extraction conditions and procedures necessary for preparing a beverage.

[0066] "Personalized settings options" is a feature that provides a customizable menu of options tailored to the user's preferences and health condition.

[0067] "Voice recognition technology" is a technology that analyzes the user's voice input and interprets it as appropriate instructions.

[0068] "Household electrical appliances" refers to various electrical devices used in the home, and this system can be operated in conjunction with them.

[0069] This invention is a system that utilizes the characteristic information of coffee beans possessed by the user and provides optimal beverage preparation instructions based on that information. The system mainly consists of three components: a server device, a terminal device, and the user, each playing a different role.

[0070] Server device operation

[0071] The server device maintains a database that manages and updates the received coffee bean characteristic data. Specifically, the server uses open-source data management software to collect and organize this information. Next, based on the data provided by the user, it generates optimal beverage preparation instructions using a generative AI model. In this process, the server uses a programming language and leverages machine learning libraries to implement the model and process the data.

[0072] Terminal device operation

[0073] The terminal device provides an interface for receiving coffee bean information and personal settings entered by the user. It also receives cooking instructions sent from the server and generates an operation screen for the user to customize their coffee based on those instructions. Specifically, the terminal is designed as an application that runs on a mobile communication device, and is easily operated by the user using touch controls and voice recognition technology.

[0074] User actions

[0075] Users input information about the coffee beans they have on hand via a terminal device and record their taste preferences and health status. They can further refine the brewing instructions using customization options. Finally, they give voice commands to the terminal to begin brewing, and the machine brews the beverage according to their requests.

[0076] Specific example

[0077] For example, consider a scenario where a user wants to brew their preferred coffee using specific Ethiopian beans. The user first inputs details of the beans into a terminal device, specifying their desired flavor profile, such as "fruity and acidic." This information is sent to a server device, which then inputs a prompt message into its AI model: "Ethiopian Yirgacheffe beans, fruity, acidic flavor, suggest optimal brewing method." This model generates optimal brewing instructions and sends the result back to the terminal device. The user can further customize the information using the provided options and finally initiate brewing using voice recognition to obtain their desired quality beverage.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The user inputs information about the characteristics of the coffee beans on the terminal device. Using a touchscreen or voice input function, the user enters information such as "Bean type: Ethiopian," "Flavor preference: Fruity," and "Roast level: Medium." This information is structured in JSON format and prepared for transmission to the server.

[0081] Step 2:

[0082] The terminal sends the information entered by the user to the server. This input is information about coffee beans entered by the user. The terminal uses the HTTPS protocol to securely send this data to the server and internally receives a confirmation message that the transmission is complete.

[0083] Step 3:

[0084] The server operates a generative AI model based on the received coffee bean characteristic data. First, the server analyzes the received data and forms a prompt sentence such as "Ethiopian beans, fruity, medium roast." This prompt sentence is input into the generative AI model, which then generates optimal beverage preparation instructions based on the collected data. These output instructions are temporarily stored in the server's database.

[0085] Step 4:

[0086] The server sends back the generated beverage preparation instructions to the terminal. These beverage preparation instructions are prepared on the terminal to facilitate user operation. Based on the received instructions, the terminal displays them in a visually easy-to-understand format on its user interface.

[0087] Step 5:

[0088] Users can review the customization options on their device and make adjustments as needed. Here, they can use sliders and dropdown menus to modify the instructions. The results of the customizations are immediately reflected on the device.

[0089] Step 6:

[0090] The user ultimately gives a voice command to the terminal device to start the extraction process. The final output is the commencement of the actual extraction process. The terminal uses voice recognition technology to process this command and sends a signal to the extraction machine to execute the instructions. This allows the user to obtain a beverage that meets their preferences.

[0091] (Application Example 1)

[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0093] Traditional systems are insufficient to provide optimal beverages tailored to the individual preferences and health conditions of each customer. Furthermore, there is a need for appropriate means to streamline the in-store ordering process and provide personalized services. Additionally, improving the efficiency of store services and enhancing customer convenience through integration with mobile communication terminals is a key challenge.

[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0095] In this invention, the server includes means for receiving characteristic data about coffee beans, means for using a generation model that generates an optimal beverage recipe based on the characteristic data, and means for performing operations in cooperation with a mobile communication terminal. This speeds up order processing at the point of sale and enables the provision of beverages optimized for the user.

[0096] "Characteristic data regarding coffee beans" includes data such as the type of coffee bean, origin, roast level, and other physical characteristics.

[0097] An "optimal beverage recipe" is a guideline regarding the characteristics and brewing methods of a beverage, generated based on the user's preferences and health condition.

[0098] "Methods using generative models" refer to mechanisms for creating optimal beverage recipes using algorithms based on data.

[0099] "Operation via collaboration with mobile communication terminals" refers to a method of operating a system through communication devices such as smartphones, and performing data transmission, reception, and execution of instructions.

[0100] "Speech recognition technology" is a technology that analyzes speech data, converts it into text, and then understands and executes commands based on that text.

[0101] This system consists of three components: servers, terminals, and users.

[0102] The server first manages characteristic data about coffee beans in a database. This characteristic data includes information such as the place of origin, type of bean, and roast level. Based on this data, it utilizes a generative AI model to generate optimal beverage recipes tailored to each user's preferences and health condition. The generated recipes are sent to the terminal and include customization options according to the user's requests. The server uses cloud computing platforms (e.g., AWS® or Google® Cloud) to perform high-speed data processing.

[0103] The terminal provides an interface for receiving information about coffee beans and taste preferences from users. Mobile communication devices such as smartphones and tablets are used, and the collected data is transmitted to a server via the internet. The terminal also utilizes speech recognition technology; for example, by using the Azure Cognitive Services speech recognition API, it can analyze voice commands uttered by users and execute appropriate instructions.

[0104] Users can input their coffee preferences into the system, review dynamically generated recipes, and customize them as needed. For example, a user can voice-instruct, "Please tell me the best way to brew a fruity coffee using Ethiopian Yirgacheffe beans," and receive a coffee tailored to their taste.

[0105] This system configuration allows for the provision of optimal beverages tailored to each individual user, improving the in-store experience.

[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0107] Step 1:

[0108] Users use their smartphones to input information such as the type of coffee beans they like, their preferred taste, and their health status. This data includes the type of beans (e.g., Ethiopian Yirgacheffe), taste profile, and desired caffeine level. This data is collected via the device's interface and transmitted to a server over the internet.

[0109] Step 2:

[0110] The server stores the received coffee bean characteristic data and user preference data, and passes this to a generative AI model. The generative AI model processes this data and calculates the optimal beverage recipe customized for each user. Based on the input data, it searches for existing recipes in the database, and the AI ​​analyzes each data combination to generate a recipe that includes the optimal brewing method. As a result of this process, the generated recipe is stored in a temporary resource on the server.

[0111] Step 3:

[0112] Recipes generated on the server are sent to the terminal in real time. The terminal receives this information and displays it visually to the user. The user can then customize the recipe if necessary. For example, the user can use an interface to adjust the level of sweetness and richness to their liking and send their selections back to the server from the terminal.

[0113] Step 4:

[0114] Based on the final customized recipe displayed on the terminal, the user issues a voice command to start brewing. The terminal recognizes this voice and converts it to text using the Azure Cognitive Services speech recognition API. The recognized command is then processed as a signal to control the brewing process. For example, if the voice input is "Please start brewing," the system will analyze it and send a command to the coffee maker accordingly.

[0115] Step 5:

[0116] The server logs the completed process based on the finalized customization information that has been re-received. This allows the system to retain the user's satisfied coffee profile and data, and continuously collect information for future reconstruction. The collected data also serves as the basis for statistical analysis, which can be used to further improve the model.

[0117] This process provides customized beverages that match the user's preferences, resulting in a personalized beverage experience.

[0118] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0119] This invention is a system that recognizes the user's emotions and provides the optimal coffee based on them. In particular, it can combine characteristic data about coffee beans with user information to customize the coffee according to the user's emotional state. This makes it possible to more individualize the coffee brewing process and provide a coffee experience that is tailored to the user's mood and state of mind.

[0120] Server roles and operations

[0121] The server not only manages a database containing various information about coffee beans, but also analyzes emotional data sent by users. Using an emotion engine, it analyzes the user's emotional state as data in real time and generates the most suitable recipe for each user. This includes processing voice data and other related data.

[0122] Terminal role and operation

[0123] The device receives voice input from the user and analyzes it using speech recognition technology. Furthermore, an emotion engine estimates the user's emotions based on the tone and word choice read from the voice. The estimated emotions are sent to a server for processing. The device then presents the user with a recipe generated considering the emotions and supports customization as needed.

[0124] User roles and actions

[0125] Users not only input information about coffee beans into the device, but also enjoy personalized adjustments based on their emotions. For example, they can easily prepare a coffee tailored to their mood, such as when they want to relax or concentrate, by giving voice commands.

[0126] Specific example

[0127] Let's take an example where a user is feeling stressed and wants to calm down to alleviate that stress. The user speaks to the device, saying, "I'd like a coffee that will help me relax a bit." The device analyzes the voice, inferring the user's emotions from their tone of voice and converting it into data. The server receives this data and uses an emotion engine to generate a recipe that is expected to have the most relaxing effect. This recipe is presented to the user via the device, and after final confirmation, the coffee is brewed and served. Through this process, the user can obtain a special cup of coffee that matches their emotions.

[0128] The following describes the processing flow.

[0129] Step 1:

[0130] The user communicates their current emotional state to the device via voice. For example, they might use a voice command such as, "I want to calm down a bit today." The device then uses speech recognition technology to convert this into text.

[0131] Step 2:

[0132] The device analyzes the text data converted from the speech and the resulting tone of voice to estimate the user's emotions. An emotion engine is used for this purpose. The estimated results are then generated as data and sent to the server.

[0133] Step 3:

[0134] The server analyzes the emotional data received from the terminal and uses an emotion engine to interpret the user's current mood and desires. Based on this, it applies a coffee recipe generation model to construct the optimal coffee recipe for the user.

[0135] Step 4:

[0136] The server sends the created recipe to the terminal. The terminal receives this information and prepares to display it to the user in an easy-to-understand visual format.

[0137] Step 5:

[0138] The user looks at the device screen and reviews the suggested recipe. They can then customize it further as needed or accept the suggestion as is.

[0139] Step 6:

[0140] Once the user is satisfied with the recipe, they can instruct the device to start brewing the coffee using voice or touch controls.

[0141] Step 7:

[0142] The device controls the coffee grinder and brewing equipment according to the user's instructions, and brews coffee according to the specified recipe. When the finished coffee is ready, the device notifies the user.

[0143] (Example 2)

[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0145] In modern society, users demand customized beverage experiences that cater to diverse emotions and preferences. However, conventional beverage delivery systems have struggled to provide individualized beverages that accurately reflect users' emotional states. This problem hinders improvements in user satisfaction, creating a need for technology that more faithfully reflects user emotions and delivers ideal beverages.

[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0147] In this invention, the server includes means for acquiring characteristic information about coffee, means for receiving and analyzing voice input from the user, and means for estimating the user's emotions based on the acquired voice data. This makes it possible to provide optimized beverages that meet the user's emotions and individual needs.

[0148] "Coffee characteristic information" refers to data related to the characteristics of coffee, such as its taste, aroma, origin, and production method.

[0149] "Means for receiving and analyzing voice input from a user" refers to the process and technology for recognizing a user's voice instructions, analyzing their content, and converting them into digital data.

[0150] "Means for estimating user emotions" refers to algorithms or systems that analyze and infer a user's emotional state based on voice data and other relevant information.

[0151] A "generative model for generating optimal beverage recipes" is an algorithm or model that calculates and suggests the most suitable preparation method and ingredients for a beverage based on the user's emotional state and preferences.

[0152] "Means for adjusting extraction conditions and providing beverages" refers to the technology of setting the parameters necessary for beverage extraction (temperature, pressure, time, etc.) based on a generated recipe, and preparing the actual beverage accordingly.

[0153] "Means for coordinating and operating other home appliances" refers to technical processes or devices that communicate with coffee machines and other related home appliances to automate or coordinate a series of processes.

[0154] This invention is a system for providing customized beverages that respond to a user's emotions. The system includes acquiring characteristic information about coffee, voice analysis, emotion estimation, recipe generation using a generative AI model, and providing the beverage under optimal brewing conditions.

[0155] server

[0156] The server has a database that manages extensive characteristic information about coffee and analyzes emotional data from users. This includes processing data obtained from voice input in real time and identifying the user's emotional state using an emotion analysis engine. Based on the emotional data, the server uses a generative AI model to calculate and generate the optimal beverage recipe for the user.

[0157] terminal

[0158] The terminal's role is to receive the user's voice input. This process is achieved through a microphone and speech recognition technology. The speech recognition technology converts the user's utterance into text data and sends it to the emotion engine. The terminal also presents the generated recipe to the user, and after receiving the user's final confirmation, operates related equipment such as a coffee machine to dispense the beverage.

[0159] User

[0160] Users provide their emotional state and beverage preferences to the device via voice input. This information is sent to a server via the device, where the most suitable beverage recipe is generated. For example, if a user says, "I want to relax," the system will suggest a beverage recipe that enhances relaxation based on that emotion.

[0161] Example of a prompt

[0162] One example of a prompt is: "Please suggest ingredients and extraction methods to provide the optimal beverage recipe for when the user wants to relax." This prompt is an instruction to use a generative AI model to derive a recipe that meets the user's specific needs.

[0163] In this way, the present invention makes it possible to provide a high-quality and rapid beverage experience that responds to the user's emotions.

[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0165] Step 1:

[0166] The device acquires the user's voice input. When the user speaks, voice data is captured by the microphone. The input is voice data that expresses the user's requests and emotions. The device performs noise cancellation and processes the input into pure voice data. The output is clear voice data.

[0167] Step 2:

[0168] The device converts acquired audio data into text data using speech recognition technology. The input is audio data, and natural language processing algorithms are used in the analysis process. In this stage of data processing, the audio waveform is converted into text, and the output is text data representing the sentences spoken by the user.

[0169] Step 3:

[0170] The device estimates the user's emotions from text data. The emotion engine analyzes keywords and context within the text to estimate the emotional state. The input is text data, the emotion analysis results are output, and the emotional state is digitized.

[0171] Step 4:

[0172] The server receives the emotion analysis results and retrieves characteristic information about coffee from the database. The input consists of emotional state data and user preference information. A generative AI model integrates this data to generate the optimal beverage recipe. The output is a beverage recipe tailored to the user's needs.

[0173] Step 5:

[0174] The terminal presents the generated beverage recipe to the user and requests final confirmation. The input is beverage recipe data, and the terminal uses this to send information to the user visually or audibly. Once the user confirms, the final request is approved as output.

[0175] Step 6:

[0176] The server sets the extraction conditions and controls other appliances such as coffee machines based on approved beverage recipes. The input is the final approved recipe data, which is used to physically extract the beverage. The output is the extracted beverage.

[0177] This series of processes enables the system to provide personalized beverages tailored to the user's emotions and preferences.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0180] In modern restaurants, providing beverages tailored to each customer's emotional state and preferences is not easy. Traditional service methods lack the flexibility to meet diverse customer needs, and suggesting beverages based on a customer's emotional state is particularly challenging. This necessitates providing more personalized service to customers.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes means for analyzing the user's emotional state in real time using emotion analysis technology, means for using a generation model that generates an optimal beverage recipe based on the analyzed emotional state, and means for adjusting the beverage delivery based on the generated recipe. This makes it possible to provide personalized beverages that correspond to the user's emotional state.

[0183] "Emotion analysis technology" is a technology that estimates a user's emotional state in real time from voice and other data.

[0184] A "generative model" refers to an algorithm or method that generates optimal results or suggestions based on input data.

[0185] "Means for coordinating beverage provision" refers to a system for optimizing the process of preparing beverages according to generated recipes and providing them to users.

[0186] "Customization options" refer to settings and choices that allow users to select the most suitable services or products according to their preferences and needs.

[0187] "Speech recognition technology" is a technology that converts voice input into text data and then analyzes it.

[0188] An "output device" is a device or equipment used to present processed information to the user.

[0189] To realize this invention, a system combining multiple technological elements is required. This system consists of a cloud server, a user-operated terminal, and an output device.

[0190] The server receives the user's voice data and converts it into text data using speech recognition technology. This text data is then used to estimate the user's emotional state using sentiment analysis technology. Specifically, speech recognition is performed using the Python `speech_recognition` library, and the emotional state is analyzed using "EmotionEngine." Based on the analyzed emotional information, the server generates the optimal beverage recipe using a generative AI model. This generative model combines user information and emotional state to suggest personalized beverages.

[0191] The terminal is responsible for receiving voice input from the user and sending it to the server. Voice input is mainly performed on devices such as smartphones and tablets, and a voice recognition process is executed. The user communicates their emotional needs for the desired beverage through a simple interface. This makes it easy for the user to receive a beverage that meets their needs.

[0192] The output device presents the generated recipe to the user visually or audibly. For example, it might display details of the suggested beverage on a display in the store or on the user's smartphone screen. The user can then choose their preferred option from the presented choices and receive their beverage after final confirmation.

[0193] For example, if a cafe customer says they want to relax, the system analyzes their voice tone and input to suggest a relaxing beverage. In this example, the generative AI model might recommend a decaf latte. The user agrees to the suggested beverage and confirms their order.

[0194] Example of a prompt:

[0195] "Please suggest beverage recipes that should be offered to users who want to relax."

[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0197] Step 1:

[0198] The device receives voice input from the user. This input includes the user's current emotional state and a description of the beverage they desire. The microphone on the device captures the voice and records it as audio data.

[0199] Step 2:

[0200] The server receives audio data sent from the terminal. It uses a speech recognition engine to convert this audio data into text data. This conversion process utilizes the Python library `speech_recognition` to format the text from the audio. The converted text clearly reflects what the user said.

[0201] Step 3:

[0202] The server processes text data using sentiment analysis technology. This process utilizes "EmotionEngine" to estimate the user's emotional state from the text. This estimation is based on the tone and keywords extracted from the text, resulting in the generation of data indicating the user's emotional state.

[0203] Step 4:

[0204] The server uses a generative AI model to generate optimal beverage recipes based on estimated emotional states. The model takes emotional data and information about the user's past preferences as input and outputs the optimal beverage corresponding to the user's current emotional state. This generation is performed using specific prompts.

[0205] Step 5:

[0206] The server sends the generated recipe to the terminal. The terminal presents this beverage information to the user visually or audibly. The user can confirm the suggested beverage through on-screen displays or audio guidance.

[0207] Step 6:

[0208] The user reviews the presented recipe and makes a selection. The user's selection is entered into the terminal and notified to the server, confirming the final order. The system then completes the provision of a beverage tailored to the user's emotional state.

[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0212] [Second Embodiment]

[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0225] This invention is a system aimed at providing the optimal coffee by utilizing information about coffee beans. The system mainly consists of three components: a server, a terminal, and a user, each with its own role.

[0226] Server roles and operations

[0227] The server maintains a database that manages and updates characteristic data such as coffee bean type, origin, and roast level. It also uses a generative model based on the collected data to generate the optimal coffee recipe. When a request is sent from a user via the terminal, the server processes the relevant data and performs the necessary calculations to provide a recipe tailored to the user's preferences and health condition.

[0228] Terminal role and operation

[0229] The terminal receives coffee bean information and personal preferences entered by the user and provides an interface for sending this information to the server. Furthermore, it receives recipe information sent from the server and generates an operation screen for the user to customize based on that information. It also has voice recognition technology, which allows it to analyze the voice spoken by the user and execute the commands necessary for the brewing process.

[0230] User roles and actions

[0231] Users input information about the coffee beans they have on hand into the device, and also record their taste preferences and health status. Furthermore, they can adjust the recipe using the customization options presented on the device. Finally, by giving a voice command to start brewing and operating the system, they can obtain coffee of their desired quality.

[0232] Specific example

[0233] For example, consider a user who wants to use "Ethiopian Yirgacheffe beans" and brew coffee to their liking. First, the user enters the details of the beans on the device, and then specifies their preferred taste. The device sends this information to the server, which generates the optimal brewing method and sends it back to the device. The device makes fine adjustments according to the user's wishes and can finally initiate brewing via voice recognition. Through this process, the user can easily enjoy a customized cup of coffee tailored to their needs.

[0234] The following describes the processing flow.

[0235] Step 1:

[0236] The user inputs information about their coffee beans (e.g., origin and type), as well as their own taste preferences and health status, into the terminal. The terminal formats this information and stores it in a way that allows for subsequent data processing.

[0237] Step 2:

[0238] The terminal generates an API request to send the entered data to the server. This request contains the user's input data, which the server uses as the basis for processing.

[0239] Step 3:

[0240] The server receives requests sent from terminals and analyzes the data. Based on the analysis results, it uses a generative model to generate recipes suitable for the characteristics of the coffee beans. The generated recipes include recommended brewing conditions.

[0241] Step 4:

[0242] The server sends the generated recipe to the terminal. The terminal prepares a user interface to display the received recipe in a user-friendly format.

[0243] Step 5:

[0244] Users can view the generated recipe along with the provided customization options through their device's display screen. They can further adjust the flavor and concentration according to their preferences and then see the results.

[0245] Step 6:

[0246] Once the adjustments are complete, the user uses voice commands to give instructions to the device, such as "Please prepare some coffee." The voice is then converted to text by the device's speech recognition technology.

[0247] Step 7:

[0248] The terminal uses voice recognition to confirm the user's instructions and then controls devices such as a coffee grinder to begin the actual coffee brewing process. Once brewing is complete, the user is notified, and the finished coffee is served.

[0249] (Example 1)

[0250] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0251] Conventional beverage preparation systems have faced challenges in generating optimal recipes tailored to individual users' preferences and health conditions, and in easily brewing and customizing them. Furthermore, they lacked sufficient voice control and integration capabilities with other household electrical appliances, resulting in poor usability.

[0252] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0253] In this invention, the server includes a device for receiving characteristic data about coffee beans, a device for utilizing a generation model that generates optimal beverage preparation instructions based on the characteristic data, and a device for adjusting the extraction of the beverage based on the generated instructions. This enables the generation of recipes tailored to the individual needs of the user and facilitates easy extraction operations.

[0254] "Characteristic data about coffee beans" refers to data that includes information such as the type of coffee bean, its origin, roast level, and quality.

[0255] A "generative model" is a machine learning model that generates optimal beverage preparation instructions for the user based on the data received.

[0256] "Beverage preparation instructions" are instructions that specify the extraction conditions and procedures necessary for preparing a beverage.

[0257] "Personalized settings options" is a feature that provides a customizable menu of options tailored to the user's preferences and health condition.

[0258] "Voice recognition technology" is a technology that analyzes the user's voice input and interprets it as appropriate instructions.

[0259] "Household electrical appliances" refers to various electrical devices used in the home, and this system can be operated in conjunction with them.

[0260] This invention is a system that utilizes the characteristic information of coffee beans possessed by the user and provides optimal beverage preparation instructions based on that information. The system mainly consists of three components: a server device, a terminal device, and the user, each playing a different role.

[0261] Server device operation

[0262] The server device maintains a database that manages and updates the received coffee bean characteristic data. Specifically, the server uses open-source data management software to collect and organize this information. Next, based on the data provided by the user, it generates optimal beverage preparation instructions using a generative AI model. In this process, the server uses a programming language and leverages machine learning libraries to implement the model and process the data.

[0263] Terminal device operation

[0264] The terminal device provides an interface for receiving coffee bean information and personal settings entered by the user. It also receives cooking instructions sent from the server and generates an operation screen for the user to customize their coffee based on those instructions. Specifically, the terminal is designed as an application that runs on a mobile communication device, and is easily operated by the user using touch controls and voice recognition technology.

[0265] User actions

[0266] Users input information about the coffee beans they have on hand via a terminal device and record their taste preferences and health status. They can further refine the brewing instructions using customization options. Finally, they give voice commands to the terminal to begin brewing, and the machine brews the beverage according to their requests.

[0267] Specific example

[0268] For example, consider a scenario where a user wants to brew their preferred coffee using specific Ethiopian beans. The user first inputs details of the beans into a terminal device, specifying their desired flavor profile, such as "fruity and acidic." This information is sent to a server device, which then inputs a prompt message into its AI model: "Ethiopian Yirgacheffe beans, fruity, acidic flavor, suggest optimal brewing method." This model generates optimal brewing instructions and sends the result back to the terminal device. The user can further customize the information using the provided options and finally initiate brewing using voice recognition to obtain their desired quality beverage.

[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0270] Step 1:

[0271] The user inputs information about the characteristics of the coffee beans on the terminal device. Using a touchscreen or voice input function, the user enters information such as "Bean type: Ethiopian," "Flavor preference: Fruity," and "Roast level: Medium." This information is structured in JSON format and prepared for transmission to the server.

[0272] Step 2:

[0273] The terminal sends the information entered by the user to the server. This input is information about coffee beans entered by the user. The terminal uses the HTTPS protocol to securely send this data to the server and internally receives a confirmation message that the transmission is complete.

[0274] Step 3:

[0275] The server operates a generative AI model based on the received coffee bean characteristic data. First, the server analyzes the received data and forms a prompt sentence such as "Ethiopian beans, fruity, medium roast." This prompt sentence is input into the generative AI model, which then generates optimal beverage preparation instructions based on the collected data. These output instructions are temporarily stored in the server's database.

[0276] Step 4:

[0277] The server sends back the generated beverage preparation instructions to the terminal. These beverage preparation instructions are prepared on the terminal to facilitate user operation. Based on the received instructions, the terminal displays them in a visually easy-to-understand format on its user interface.

[0278] Step 5:

[0279] Users can review the customization options on their device and make adjustments as needed. Here, they can use sliders and dropdown menus to modify the instructions. The results of the customizations are immediately reflected on the device.

[0280] Step 6:

[0281] Finally, the user gives a voice instruction to the terminal device to start extraction. The final output is that the actual extraction process starts. The terminal processes this command using voice recognition technology and sends a signal to the extraction device to execute the instruction. As a result, the user can obtain the beverage that meets their own needs.

[0282] (Application Example 1)

[0283] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0284] In order to provide an optimal beverage according to the preferences and health conditions of individual users, the conventional system is not sufficient. Also, an appropriate means for smoothing the in-store ordering process and providing individualized services is required. Furthermore, it is an issue to improve the convenience of users by streamlining the store service through cooperation with a mobile communication terminal.

[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0286] In this invention, the server includes means for receiving characteristic data regarding coffee beans, means for using a generation model that generates an optimal beverage recipe based on the characteristic data, and means for performing an operation through cooperation with a mobile communication terminal. As a result, the in-store order processing is speeded up, and it becomes possible to provide an optimized beverage to the user.

[0287] The "characteristic data regarding coffee beans" is data including the type of coffee beans, the place of origin, the roasting degree, and other physical characteristics.

[0288] The "optimal beverage recipe" is a guideline regarding the characteristics and extraction method of a beverage, generated based on the preferences and health conditions of the user.

[0289] "Methods using generative models" refer to mechanisms for creating optimal beverage recipes using algorithms based on data.

[0290] "Operation via collaboration with mobile communication terminals" refers to a method of operating a system through communication devices such as smartphones, and performing data transmission, reception, and execution of instructions.

[0291] "Speech recognition technology" is a technology that analyzes speech data, converts it into text, and then understands and executes commands based on that text.

[0292] This system consists of three components: servers, terminals, and users.

[0293] The server first manages characteristic data about coffee beans in a database. This characteristic data includes information such as the place of origin, type of bean, and roast level. Based on this data, it utilizes a generative AI model to generate optimal beverage recipes tailored to each user's preferences and health condition. The generated recipes are sent to the terminal and include customization options according to the user's requests. The server uses cloud computing platforms (e.g., AWS and Google Cloud) to perform high-speed data processing.

[0294] The terminal provides an interface for receiving information about coffee beans and taste preferences from users. Mobile communication devices such as smartphones and tablets are used, and the collected data is transmitted to a server via the internet. The terminal also utilizes speech recognition technology; for example, by using the Azure Cognitive Services speech recognition API, it can analyze voice commands uttered by users and execute appropriate instructions.

[0295] Users can input their coffee preferences into the system, review dynamically generated recipes, and customize them as needed. For example, a user can voice-instruct, "Please tell me the best way to brew a fruity coffee using Ethiopian Yirgacheffe beans," and receive a coffee tailored to their taste.

[0296] This system configuration allows for the provision of optimal beverages tailored to each individual user, improving the in-store experience.

[0297] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0298] Step 1:

[0299] Users use their smartphones to input information such as the type of coffee beans they like, their preferred taste, and their health status. This data includes the type of beans (e.g., Ethiopian Yirgacheffe), taste profile, and desired caffeine level. This data is collected via the device's interface and transmitted to a server over the internet.

[0300] Step 2:

[0301] The server stores the received coffee bean characteristic data and user preference data, and passes this to a generative AI model. The generative AI model processes this data and calculates the optimal beverage recipe customized for each user. Based on the input data, it searches for existing recipes in the database, and the AI ​​analyzes each data combination to generate a recipe that includes the optimal brewing method. As a result of this process, the generated recipe is stored in a temporary resource on the server.

[0302] Step 3:

[0303] The recipe generated by the server is sent to the terminal in real time. The terminal receives this information and visually displays it to the user. Here, the user can customize the recipe if necessary. For example, the user uses an interface to adjust the degree of sweetness and concentration according to their condition, and resends the selection from the terminal to the server.

[0304] Step 4:

[0305] Based on the final customized recipe displayed on the terminal, the user issues a voice command to start extraction. The terminal recognizes this voice and converts the voice into text using the speech recognition API of Azure Cognitive Services. The recognized command is processed as a signal to control the extraction process as it is. For example, a voice input such as "Please start extraction" is analyzed, and the system sends an instruction to the coffee maker according to the instruction.

[0306] Step 5:

[0307] Based on the re-received final customization information, the server records the completed process in the log. As a result, it is possible to retain the profile and data of the coffee that the user is satisfied with and continue to collect information for later reconstruction. The collected data also serves as a basis for statistical analysis and can be used to improve further models.

[0308] Through this process, a customized beverage that matches the user's preference is provided, and a personalized beverage experience is realized.

[0309] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.

[0310] This invention is a system that recognizes the user's emotions and provides the optimal coffee based on them. In particular, it can combine characteristic data about coffee beans with user information to customize the coffee according to the user's emotional state. This makes it possible to more individualize the coffee brewing process and provide a coffee experience that is tailored to the user's mood and state of mind.

[0311] Server roles and operations

[0312] The server not only manages a database containing various information about coffee beans, but also analyzes emotional data sent by users. Using an emotion engine, it analyzes the user's emotional state as data in real time and generates the most suitable recipe for each user. This includes processing voice data and other related data.

[0313] Terminal role and operation

[0314] The device receives voice input from the user and analyzes it using speech recognition technology. Furthermore, an emotion engine estimates the user's emotions based on the tone and word choice read from the voice. The estimated emotions are sent to a server for processing. The device then presents the user with a recipe generated considering the emotions and supports customization as needed.

[0315] User roles and actions

[0316] Users not only input information about coffee beans into the device, but also enjoy personalized adjustments based on their emotions. For example, they can easily prepare a coffee tailored to their mood, such as when they want to relax or concentrate, by giving voice commands.

[0317] Specific example

[0318] Let's take an example where a user is feeling stressed and wants to calm down to alleviate that stress. The user speaks to the device, saying, "I'd like a coffee that will help me relax a bit." The device analyzes the voice, inferring the user's emotions from their tone of voice and converting it into data. The server receives this data and uses an emotion engine to generate a recipe that is expected to have the most relaxing effect. This recipe is presented to the user via the device, and after final confirmation, the coffee is brewed and served. Through this process, the user can obtain a special cup of coffee that matches their emotions.

[0319] The following describes the processing flow.

[0320] Step 1:

[0321] The user communicates their current emotional state to the device via voice. For example, they might use a voice command such as, "I want to calm down a bit today." The device then uses speech recognition technology to convert this into text.

[0322] Step 2:

[0323] The device analyzes the text data converted from the speech and the resulting tone of voice to estimate the user's emotions. An emotion engine is used for this purpose. The estimated results are then generated as data and sent to the server.

[0324] Step 3:

[0325] The server analyzes the emotional data received from the terminal and uses an emotion engine to interpret the user's current mood and desires. Based on this, it applies a coffee recipe generation model to construct the optimal coffee recipe for the user.

[0326] Step 4:

[0327] The server sends the created recipe to the terminal. The terminal receives this information and prepares to display it to the user in an easy-to-understand visual format.

[0328] Step 5:

[0329] The user looks at the device screen and reviews the suggested recipe. They can then customize it further as needed or accept the suggestion as is.

[0330] Step 6:

[0331] Once the user is satisfied with the recipe, they can instruct the device to start brewing the coffee using voice or touch controls.

[0332] Step 7:

[0333] The device controls the coffee grinder and brewing equipment according to the user's instructions, and brews coffee according to the specified recipe. When the finished coffee is ready, the device notifies the user.

[0334] (Example 2)

[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0336] In modern society, users demand customized beverage experiences that cater to diverse emotions and preferences. However, conventional beverage delivery systems have struggled to provide individualized beverages that accurately reflect users' emotional states. This problem hinders improvements in user satisfaction, creating a need for technology that more faithfully reflects user emotions and delivers ideal beverages.

[0337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0338] In this invention, the server includes means for acquiring characteristic information about coffee, means for receiving and analyzing voice input from the user, and means for estimating the user's emotions based on the acquired voice data. This makes it possible to provide optimized beverages that meet the user's emotions and individual needs.

[0339] "Coffee characteristic information" refers to data related to the characteristics of coffee, such as its taste, aroma, origin, and production method.

[0340] "Means for receiving and analyzing voice input from a user" refers to the process and technology for recognizing a user's voice instructions, analyzing their content, and converting them into digital data.

[0341] "Means for estimating user emotions" refers to algorithms or systems that analyze and infer a user's emotional state based on voice data and other relevant information.

[0342] A "generative model for generating optimal beverage recipes" is an algorithm or model that calculates and suggests the most suitable preparation method and ingredients for a beverage based on the user's emotional state and preferences.

[0343] "Means for adjusting extraction conditions and providing beverages" refers to the technology of setting the parameters necessary for beverage extraction (temperature, pressure, time, etc.) based on a generated recipe, and preparing the actual beverage accordingly.

[0344] "Means for coordinating and operating other home appliances" refers to technical processes or devices that communicate with coffee machines and other related home appliances to automate or coordinate a series of processes.

[0345] This invention is a system for providing customized beverages that respond to a user's emotions. The system includes acquiring characteristic information about coffee, voice analysis, emotion estimation, recipe generation using a generative AI model, and providing the beverage under optimal brewing conditions.

[0346] server

[0347] The server has a database that manages extensive characteristic information about coffee and analyzes emotional data from users. This includes processing data obtained from voice input in real time and identifying the user's emotional state using an emotion analysis engine. Based on the emotional data, the server uses a generative AI model to calculate and generate the optimal beverage recipe for the user.

[0348] terminal

[0349] The terminal's role is to receive the user's voice input. This process is achieved through a microphone and speech recognition technology. The speech recognition technology converts the user's utterance into text data and sends it to the emotion engine. The terminal also presents the generated recipe to the user, and after receiving the user's final confirmation, operates related equipment such as a coffee machine to dispense the beverage.

[0350] User

[0351] Users provide their emotional state and beverage preferences to the device via voice input. This information is sent to a server via the device, where the most suitable beverage recipe is generated. For example, if a user says, "I want to relax," the system will suggest a beverage recipe that enhances relaxation based on that emotion.

[0352] Example of a prompt

[0353] One example of a prompt is: "Please suggest ingredients and extraction methods to provide the optimal beverage recipe for when the user wants to relax." This prompt is an instruction to use a generative AI model to derive a recipe that meets the user's specific needs.

[0354] In this way, the present invention makes it possible to provide a high-quality and rapid beverage experience that responds to the user's emotions.

[0355] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0356] Step 1:

[0357] The device acquires the user's voice input. When the user speaks, voice data is captured by the microphone. The input is voice data that expresses the user's requests and emotions. The device performs noise cancellation and processes the input into pure voice data. The output is clear voice data.

[0358] Step 2:

[0359] The device converts acquired audio data into text data using speech recognition technology. The input is audio data, and natural language processing algorithms are used in the analysis process. In this stage of data processing, the audio waveform is converted into text, and the output is text data representing the sentences spoken by the user.

[0360] Step 3:

[0361] The device estimates the user's emotions from text data. The emotion engine analyzes keywords and context within the text to estimate the emotional state. The input is text data, the emotion analysis results are output, and the emotional state is digitized.

[0362] Step 4:

[0363] The server receives the emotion analysis results and retrieves characteristic information about coffee from the database. The input consists of emotional state data and user preference information. A generative AI model integrates this data to generate the optimal beverage recipe. The output is a beverage recipe tailored to the user's needs.

[0364] Step 5:

[0365] The terminal presents the generated beverage recipe to the user and requests final confirmation. The input is beverage recipe data, and the terminal uses this to send information to the user visually or audibly. Once the user confirms, the final request is approved as output.

[0366] Step 6:

[0367] The server sets the extraction conditions and controls other appliances such as coffee machines based on approved beverage recipes. The input is the final approved recipe data, which is used to physically extract the beverage. The output is the extracted beverage.

[0368] This series of processes enables the system to provide personalized beverages tailored to the user's emotions and preferences.

[0369] (Application Example 2)

[0370] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0371] In modern restaurants, providing beverages tailored to each customer's emotional state and preferences is not easy. Traditional service methods lack the flexibility to meet diverse customer needs, and suggesting beverages based on a customer's emotional state is particularly challenging. This necessitates providing more personalized service to customers.

[0372] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0373] In this invention, the server includes means for analyzing the user's emotional state in real time using emotion analysis technology, means for using a generation model that generates an optimal beverage recipe based on the analyzed emotional state, and means for adjusting the beverage delivery based on the generated recipe. This makes it possible to provide personalized beverages that correspond to the user's emotional state.

[0374] "Emotion analysis technology" is a technology that estimates a user's emotional state in real time from voice and other data.

[0375] A "generative model" refers to an algorithm or method that generates optimal results or suggestions based on input data.

[0376] "Means for coordinating beverage provision" refers to a system for optimizing the process of preparing beverages according to generated recipes and providing them to users.

[0377] "Customization options" refer to settings and choices that allow users to select the most suitable services or products according to their preferences and needs.

[0378] "Speech recognition technology" is a technology that converts voice input into text data and then analyzes it.

[0379] An "output device" is a device or equipment used to present processed information to the user.

[0380] To realize this invention, a system combining multiple technological elements is required. This system consists of a cloud server, a user-operated terminal, and an output device.

[0381] The server receives the user's voice data and converts it into text data using speech recognition technology. This text data is then used to estimate the user's emotional state using sentiment analysis technology. Specifically, speech recognition is performed using the Python `speech_recognition` library, and the emotional state is analyzed using "EmotionEngine." Based on the analyzed emotional information, the server generates the optimal beverage recipe using a generative AI model. This generative model combines user information and emotional state to suggest personalized beverages.

[0382] The terminal is responsible for receiving voice input from the user and sending it to the server. Voice input is mainly performed on devices such as smartphones and tablets, and a voice recognition process is executed. The user communicates their emotional needs for the desired beverage through a simple interface. This makes it easy for the user to receive a beverage that meets their needs.

[0383] The output device presents the generated recipe to the user visually or audibly. For example, it might display details of the suggested beverage on a display in the store or on the user's smartphone screen. The user can then choose their preferred option from the presented choices and receive their beverage after final confirmation.

[0384] For example, if a cafe customer says they want to relax, the system analyzes their voice tone and input to suggest a relaxing beverage. In this example, the generative AI model might recommend a decaf latte. The user agrees to the suggested beverage and confirms their order.

[0385] Example of a prompt:

[0386] "Please suggest beverage recipes that should be offered to users who want to relax."

[0387] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0388] Step 1:

[0389] The device receives voice input from the user. This input includes the user's current emotional state and a description of the beverage they desire. The microphone on the device captures the voice and records it as audio data.

[0390] Step 2:

[0391] The server receives audio data sent from the terminal. It uses a speech recognition engine to convert this audio data into text data. This conversion process utilizes the Python library `speech_recognition` to format the text from the audio. The converted text clearly reflects what the user said.

[0392] Step 3:

[0393] The server processes text data using sentiment analysis technology. This process utilizes "EmotionEngine" to estimate the user's emotional state from the text. This estimation is based on the tone and keywords extracted from the text, resulting in the generation of data indicating the user's emotional state.

[0394] Step 4:

[0395] The server uses a generative AI model to generate optimal beverage recipes based on estimated emotional states. The model takes emotional data and information about the user's past preferences as input and outputs the optimal beverage corresponding to the user's current emotional state. This generation is performed using specific prompts.

[0396] Step 5:

[0397] The server sends the generated recipe to the terminal. The terminal presents this beverage information to the user visually or audibly. The user can confirm the suggested beverage through on-screen displays or audio guidance.

[0398] Step 6:

[0399] The user reviews the presented recipe and makes a selection. The user's selection is entered into the terminal and notified to the server, confirming the final order. The system then completes the provision of a beverage tailored to the user's emotional state.

[0400] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0401] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0402] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0403] [Third Embodiment]

[0404] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0405] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0406] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0407] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0408] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0410] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0411] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0412] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0413] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0414] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0415] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0416] This invention is a system aimed at providing the optimal coffee by utilizing information about coffee beans. The system mainly consists of three components: a server, a terminal, and a user, each with its own role.

[0417] Server roles and operations

[0418] The server maintains a database that manages and updates characteristic data such as coffee bean type, origin, and roast level. It also uses a generative model based on the collected data to generate the optimal coffee recipe. When a request is sent from a user via the terminal, the server processes the relevant data and performs the necessary calculations to provide a recipe tailored to the user's preferences and health condition.

[0419] Terminal role and operation

[0420] The terminal receives coffee bean information and personal preferences entered by the user and provides an interface for sending this information to the server. Furthermore, it receives recipe information sent from the server and generates an operation screen for the user to customize based on that information. It also has voice recognition technology, which allows it to analyze the voice spoken by the user and execute the commands necessary for the brewing process.

[0421] User roles and actions

[0422] Users input information about the coffee beans they have on hand into the device, and also record their taste preferences and health status. Furthermore, they can adjust the recipe using the customization options presented on the device. Finally, by giving a voice command to start brewing and operating the system, they can obtain coffee of their desired quality.

[0423] Specific example

[0424] For example, consider a user who wants to use "Ethiopian Yirgacheffe beans" and brew coffee to their liking. First, the user enters the details of the beans on the device, and then specifies their preferred taste. The device sends this information to the server, which generates the optimal brewing method and sends it back to the device. The device makes fine adjustments according to the user's wishes and can finally initiate brewing via voice recognition. Through this process, the user can easily enjoy a customized cup of coffee tailored to their needs.

[0425] The following describes the processing flow.

[0426] Step 1:

[0427] The user inputs information about their coffee beans (e.g., origin and type), as well as their own taste preferences and health status, into the terminal. The terminal formats this information and stores it in a way that allows for subsequent data processing.

[0428] Step 2:

[0429] The terminal generates an API request to send the entered data to the server. This request contains the user's input data, which the server uses as the basis for processing.

[0430] Step 3:

[0431] The server receives requests sent from terminals and analyzes the data. Based on the analysis results, it uses a generative model to generate recipes suitable for the characteristics of the coffee beans. The generated recipes include recommended brewing conditions.

[0432] Step 4:

[0433] The server sends the generated recipe to the terminal. The terminal prepares a user interface to display the received recipe in a user-friendly format.

[0434] Step 5:

[0435] Users can view the generated recipe along with the provided customization options through their device's display screen. They can further adjust the flavor and concentration according to their preferences and then see the results.

[0436] Step 6:

[0437] Once the adjustments are complete, the user uses voice commands to give instructions to the device, such as "Please prepare some coffee." The voice is then converted to text by the device's speech recognition technology.

[0438] Step 7:

[0439] The terminal uses voice recognition to confirm the user's instructions and then controls devices such as a coffee grinder to begin the actual coffee brewing process. Once brewing is complete, the user is notified, and the finished coffee is served.

[0440] (Example 1)

[0441] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0442] Conventional beverage preparation systems have faced challenges in generating optimal recipes tailored to individual users' preferences and health conditions, and in easily brewing and customizing them. Furthermore, they lacked sufficient voice control and integration capabilities with other household electrical appliances, resulting in poor usability.

[0443] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0444] In this invention, the server includes a device for receiving characteristic data about coffee beans, a device for utilizing a generation model that generates optimal beverage preparation instructions based on the characteristic data, and a device for adjusting the extraction of the beverage based on the generated instructions. This enables the generation of recipes tailored to the individual needs of the user and facilitates easy extraction operations.

[0445] "Characteristic data about coffee beans" refers to data that includes information such as the type of coffee bean, its origin, roast level, and quality.

[0446] A "generative model" is a machine learning model that generates optimal beverage preparation instructions for the user based on the data received.

[0447] "Beverage preparation instructions" are instructions that specify the extraction conditions and procedures necessary for preparing a beverage.

[0448] "Personalized settings options" is a feature that provides a customizable menu of options tailored to the user's preferences and health condition.

[0449] "Voice recognition technology" is a technology that analyzes the user's voice input and interprets it as appropriate instructions.

[0450] "Household electrical appliances" refers to various electrical devices used in the home, and this system can be operated in conjunction with them.

[0451] This invention is a system that utilizes the characteristic information of coffee beans possessed by the user and provides optimal beverage preparation instructions based on that information. The system mainly consists of three components: a server device, a terminal device, and the user, each playing a different role.

[0452] Server device operation

[0453] The server device maintains a database that manages and updates the received coffee bean characteristic data. Specifically, the server uses open-source data management software to collect and organize this information. Next, based on the data provided by the user, it generates optimal beverage preparation instructions using a generative AI model. In this process, the server uses a programming language and leverages machine learning libraries to implement the model and process the data.

[0454] Terminal device operation

[0455] The terminal device provides an interface for receiving coffee bean information and personal settings entered by the user. It also receives cooking instructions sent from the server and generates an operation screen for the user to customize their coffee based on those instructions. Specifically, the terminal is designed as an application that runs on a mobile communication device, and is easily operated by the user using touch controls and voice recognition technology.

[0456] User actions

[0457] Users input information about the coffee beans they have on hand via a terminal device and record their taste preferences and health status. They can further refine the brewing instructions using customization options. Finally, they give voice commands to the terminal to begin brewing, and the machine brews the beverage according to their requests.

[0458] Specific example

[0459] For example, consider a scenario where a user wants to brew their preferred coffee using specific Ethiopian beans. The user first inputs details of the beans into a terminal device, specifying their desired flavor profile, such as "fruity and acidic." This information is sent to a server device, which then inputs a prompt message into its AI model: "Ethiopian Yirgacheffe beans, fruity, acidic flavor, suggest optimal brewing method." This model generates optimal brewing instructions and sends the result back to the terminal device. The user can further customize the information using the provided options and finally initiate brewing using voice recognition to obtain their desired quality beverage.

[0460] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0461] Step 1:

[0462] The user inputs information about the characteristics of the coffee beans on the terminal device. Using a touchscreen or voice input function, the user enters information such as "Bean type: Ethiopian," "Flavor preference: Fruity," and "Roast level: Medium." This information is structured in JSON format and prepared for transmission to the server.

[0463] Step 2:

[0464] The terminal sends the information entered by the user to the server. This input is information about coffee beans entered by the user. The terminal uses the HTTPS protocol to securely send this data to the server and internally receives a confirmation message that the transmission is complete.

[0465] Step 3:

[0466] The server operates a generative AI model based on the received coffee bean characteristic data. First, the server analyzes the received data and forms a prompt sentence such as "Ethiopian beans, fruity, medium roast." This prompt sentence is input into the generative AI model, which then generates optimal beverage preparation instructions based on the collected data. These output instructions are temporarily stored in the server's database.

[0467] Step 4:

[0468] The server sends back the generated beverage preparation instructions to the terminal. These beverage preparation instructions are prepared on the terminal to facilitate user operation. Based on the received instructions, the terminal displays them in a visually easy-to-understand format on its user interface.

[0469] Step 5:

[0470] Users can review the customization options on their device and make adjustments as needed. Here, they can use sliders and dropdown menus to modify the instructions. The results of the customizations are immediately reflected on the device.

[0471] Step 6:

[0472] The user ultimately gives a voice command to the terminal device to start the extraction process. The final output is the commencement of the actual extraction process. The terminal uses voice recognition technology to process this command and sends a signal to the extraction machine to execute the instructions. This allows the user to obtain a beverage that meets their preferences.

[0473] (Application Example 1)

[0474] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0475] Traditional systems are insufficient to provide optimal beverages tailored to the individual preferences and health conditions of each customer. Furthermore, there is a need for appropriate means to streamline the in-store ordering process and provide personalized services. Additionally, improving the efficiency of store services and enhancing customer convenience through integration with mobile communication terminals is a key challenge.

[0476] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0477] In this invention, the server includes means for receiving characteristic data about coffee beans, means for using a generation model that generates an optimal beverage recipe based on the characteristic data, and means for performing operations in cooperation with a mobile communication terminal. This speeds up order processing at the point of sale and enables the provision of beverages optimized for the user.

[0478] "Characteristic data regarding coffee beans" includes data such as the type of coffee bean, origin, roast level, and other physical characteristics.

[0479] An "optimal beverage recipe" is a guideline regarding the characteristics and brewing methods of a beverage, generated based on the user's preferences and health condition.

[0480] "Methods using generative models" refer to mechanisms for creating optimal beverage recipes using algorithms based on data.

[0481] "Operation via collaboration with mobile communication terminals" refers to a method of operating a system through communication devices such as smartphones, and performing data transmission, reception, and execution of instructions.

[0482] "Speech recognition technology" is a technology that analyzes speech data, converts it into text, and then understands and executes commands based on that text.

[0483] This system consists of three components: servers, terminals, and users.

[0484] The server first manages characteristic data about coffee beans in a database. This characteristic data includes information such as the place of origin, type of bean, and roast level. Based on this data, it utilizes a generative AI model to generate optimal beverage recipes tailored to each user's preferences and health condition. The generated recipes are sent to the terminal and include customization options according to the user's requests. The server uses cloud computing platforms (e.g., AWS and Google Cloud) to perform high-speed data processing.

[0485] The terminal provides an interface for receiving information about coffee beans and taste preferences from users. Mobile communication devices such as smartphones and tablets are used, and the collected data is transmitted to a server via the internet. The terminal also utilizes speech recognition technology; for example, by using the Azure Cognitive Services speech recognition API, it can analyze voice commands uttered by users and execute appropriate instructions.

[0486] Users can input their coffee preferences into the system, review dynamically generated recipes, and customize them as needed. For example, a user can voice-instruct, "Please tell me the best way to brew a fruity coffee using Ethiopian Yirgacheffe beans," and receive a coffee tailored to their taste.

[0487] This system configuration allows for the provision of optimal beverages tailored to each individual user, improving the in-store experience.

[0488] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0489] Step 1:

[0490] Users use their smartphones to input information such as the type of coffee beans they like, their preferred taste, and their health status. This data includes the type of beans (e.g., Ethiopian Yirgacheffe), taste profile, and desired caffeine level. This data is collected via the device's interface and transmitted to a server over the internet.

[0491] Step 2:

[0492] The server stores the received coffee bean characteristic data and user preference data, and passes this to a generative AI model. The generative AI model processes this data and calculates the optimal beverage recipe customized for each user. Based on the input data, it searches for existing recipes in the database, and the AI ​​analyzes each data combination to generate a recipe that includes the optimal brewing method. As a result of this process, the generated recipe is stored in a temporary resource on the server.

[0493] Step 3:

[0494] Recipes generated on the server are sent to the terminal in real time. The terminal receives this information and displays it visually to the user. The user can then customize the recipe if necessary. For example, the user can use an interface to adjust the level of sweetness and richness to their liking and send their selections back to the server from the terminal.

[0495] Step 4:

[0496] Based on the final customized recipe displayed on the terminal, the user issues a voice command to start brewing. The terminal recognizes this voice and converts it to text using the Azure Cognitive Services speech recognition API. The recognized command is then processed as a signal to control the brewing process. For example, if the voice input is "Please start brewing," the system will analyze it and send a command to the coffee maker accordingly.

[0497] Step 5:

[0498] The server logs the completed process based on the finalized customization information that has been re-received. This allows the system to retain the user's satisfied coffee profile and data, and continuously collect information for future reconstruction. The collected data also serves as the basis for statistical analysis, which can be used to further improve the model.

[0499] This process provides customized beverages that match the user's preferences, resulting in a personalized beverage experience.

[0500] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0501] This invention is a system that recognizes the user's emotions and provides the optimal coffee based on them. In particular, it can combine characteristic data about coffee beans with user information to customize the coffee according to the user's emotional state. This makes it possible to more individualize the coffee brewing process and provide a coffee experience that is tailored to the user's mood and state of mind.

[0502] Server roles and operations

[0503] The server not only manages a database containing various information about coffee beans, but also analyzes emotional data sent by users. Using an emotion engine, it analyzes the user's emotional state as data in real time and generates the most suitable recipe for each user. This includes processing voice data and other related data.

[0504] Terminal role and operation

[0505] The device receives voice input from the user and analyzes it using speech recognition technology. Furthermore, an emotion engine estimates the user's emotions based on the tone and word choice read from the voice. The estimated emotions are sent to a server for processing. The device then presents the user with a recipe generated considering the emotions and supports customization as needed.

[0506] User roles and actions

[0507] Users not only input information about coffee beans into the device, but also enjoy personalized adjustments based on their emotions. For example, they can easily prepare a coffee tailored to their mood, such as when they want to relax or concentrate, by giving voice commands.

[0508] Specific example

[0509] Let's take an example where a user is feeling stressed and wants to calm down to alleviate that stress. The user speaks to the device, saying, "I'd like a coffee that will help me relax a bit." The device analyzes the voice, inferring the user's emotions from their tone of voice and converting it into data. The server receives this data and uses an emotion engine to generate a recipe that is expected to have the most relaxing effect. This recipe is presented to the user via the device, and after final confirmation, the coffee is brewed and served. Through this process, the user can obtain a special cup of coffee that matches their emotions.

[0510] The following describes the processing flow.

[0511] Step 1:

[0512] The user communicates their current emotional state to the device via voice. For example, they might use a voice command such as, "I want to calm down a bit today." The device then uses speech recognition technology to convert this into text.

[0513] Step 2:

[0514] The device analyzes the text data converted from the speech and the resulting tone of voice to estimate the user's emotions. An emotion engine is used for this purpose. The estimated results are then generated as data and sent to the server.

[0515] Step 3:

[0516] The server analyzes the emotional data received from the terminal and uses an emotion engine to interpret the user's current mood and desires. Based on this, it applies a coffee recipe generation model to construct the optimal coffee recipe for the user.

[0517] Step 4:

[0518] The server sends the created recipe to the terminal. The terminal receives this information and prepares to display it to the user in an easy-to-understand visual format.

[0519] Step 5:

[0520] The user looks at the device screen and reviews the suggested recipe. They can then customize it further as needed or accept the suggestion as is.

[0521] Step 6:

[0522] Once the user is satisfied with the recipe, they can instruct the device to start brewing the coffee using voice or touch controls.

[0523] Step 7:

[0524] The device controls the coffee grinder and brewing equipment according to the user's instructions, and brews coffee according to the specified recipe. When the finished coffee is ready, the device notifies the user.

[0525] (Example 2)

[0526] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0527] In modern society, users demand customized beverage experiences that cater to diverse emotions and preferences. However, conventional beverage delivery systems have struggled to provide individualized beverages that accurately reflect users' emotional states. This problem hinders improvements in user satisfaction, creating a need for technology that more faithfully reflects user emotions and delivers ideal beverages.

[0528] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0529] In this invention, the server includes means for acquiring characteristic information about coffee, means for receiving and analyzing voice input from the user, and means for estimating the user's emotions based on the acquired voice data. This makes it possible to provide optimized beverages that meet the user's emotions and individual needs.

[0530] "Coffee characteristic information" refers to data related to the characteristics of coffee, such as its taste, aroma, origin, and production method.

[0531] "Means for receiving and analyzing voice input from a user" refers to the process and technology for recognizing a user's voice instructions, analyzing their content, and converting them into digital data.

[0532] "Means for estimating user emotions" refers to algorithms or systems that analyze and infer a user's emotional state based on voice data and other relevant information.

[0533] A "generative model for generating optimal beverage recipes" is an algorithm or model that calculates and suggests the most suitable preparation method and ingredients for a beverage based on the user's emotional state and preferences.

[0534] "Means for adjusting extraction conditions and providing beverages" refers to the technology of setting the parameters necessary for beverage extraction (temperature, pressure, time, etc.) based on a generated recipe, and preparing the actual beverage accordingly.

[0535] "Means for coordinating and operating other home appliances" refers to technical processes or devices that communicate with coffee machines and other related home appliances to automate or coordinate a series of processes.

[0536] This invention is a system for providing customized beverages that respond to a user's emotions. The system includes acquiring characteristic information about coffee, voice analysis, emotion estimation, recipe generation using a generative AI model, and providing the beverage under optimal brewing conditions.

[0537] server

[0538] The server has a database that manages extensive characteristic information about coffee and analyzes emotional data from users. This includes processing data obtained from voice input in real time and identifying the user's emotional state using an emotion analysis engine. Based on the emotional data, the server uses a generative AI model to calculate and generate the optimal beverage recipe for the user.

[0539] terminal

[0540] The terminal's role is to receive the user's voice input. This process is achieved through a microphone and speech recognition technology. The speech recognition technology converts the user's utterance into text data and sends it to the emotion engine. The terminal also presents the generated recipe to the user, and after receiving the user's final confirmation, operates related equipment such as a coffee machine to dispense the beverage.

[0541] User

[0542] Users provide their emotional state and beverage preferences to the device via voice input. This information is sent to a server via the device, where the most suitable beverage recipe is generated. For example, if a user says, "I want to relax," the system will suggest a beverage recipe that enhances relaxation based on that emotion.

[0543] Example of a prompt

[0544] One example of a prompt is: "Please suggest ingredients and extraction methods to provide the optimal beverage recipe for when the user wants to relax." This prompt is an instruction to use a generative AI model to derive a recipe that meets the user's specific needs.

[0545] In this way, the present invention makes it possible to provide a high-quality and rapid beverage experience that responds to the user's emotions.

[0546] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0547] Step 1:

[0548] The device acquires the user's voice input. When the user speaks, voice data is captured by the microphone. The input is voice data that expresses the user's requests and emotions. The device performs noise cancellation and processes the input into pure voice data. The output is clear voice data.

[0549] Step 2:

[0550] The device converts acquired audio data into text data using speech recognition technology. The input is audio data, and natural language processing algorithms are used in the analysis process. In this stage of data processing, the audio waveform is converted into text, and the output is text data representing the sentences spoken by the user.

[0551] Step 3:

[0552] The device estimates the user's emotions from text data. The emotion engine analyzes keywords and context within the text to estimate the emotional state. The input is text data, the emotion analysis results are output, and the emotional state is digitized.

[0553] Step 4:

[0554] The server receives the emotion analysis results and retrieves characteristic information about coffee from the database. The input consists of emotional state data and user preference information. A generative AI model integrates this data to generate the optimal beverage recipe. The output is a beverage recipe tailored to the user's needs.

[0555] Step 5:

[0556] The terminal presents the generated beverage recipe to the user and requests final confirmation. The input is beverage recipe data, and the terminal uses this to send information to the user visually or audibly. Once the user confirms, the final request is approved as output.

[0557] Step 6:

[0558] The server sets the extraction conditions and controls other appliances such as coffee machines based on approved beverage recipes. The input is the final approved recipe data, which is used to physically extract the beverage. The output is the extracted beverage.

[0559] This series of processes enables the system to provide personalized beverages tailored to the user's emotions and preferences.

[0560] (Application Example 2)

[0561] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0562] In modern restaurants, providing beverages tailored to each customer's emotional state and preferences is not easy. Traditional service methods lack the flexibility to meet diverse customer needs, and suggesting beverages based on a customer's emotional state is particularly challenging. This necessitates providing more personalized service to customers.

[0563] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0564] In this invention, the server includes means for analyzing the user's emotional state in real time using emotion analysis technology, means for using a generation model that generates an optimal beverage recipe based on the analyzed emotional state, and means for adjusting the beverage delivery based on the generated recipe. This makes it possible to provide personalized beverages that correspond to the user's emotional state.

[0565] "Emotion analysis technology" is a technology that estimates a user's emotional state in real time from voice and other data.

[0566] A "generative model" refers to an algorithm or method that generates optimal results or suggestions based on input data.

[0567] "Means for coordinating beverage provision" refers to a system for optimizing the process of preparing beverages according to generated recipes and providing them to users.

[0568] "Customization options" refer to settings and choices that allow users to select the most suitable services or products according to their preferences and needs.

[0569] "Speech recognition technology" is a technology that converts voice input into text data and then analyzes it.

[0570] An "output device" is a device or equipment used to present processed information to the user.

[0571] To realize this invention, a system combining multiple technological elements is required. This system consists of a cloud server, a user-operated terminal, and an output device.

[0572] The server receives the user's voice data and converts it into text data using speech recognition technology. This text data is then used to estimate the user's emotional state using sentiment analysis technology. Specifically, speech recognition is performed using the Python `speech_recognition` library, and the emotional state is analyzed using "EmotionEngine." Based on the analyzed emotional information, the server generates the optimal beverage recipe using a generative AI model. This generative model combines user information and emotional state to suggest personalized beverages.

[0573] The terminal is responsible for receiving voice input from the user and sending it to the server. Voice input is mainly performed on devices such as smartphones and tablets, and a voice recognition process is executed. The user communicates their emotional needs for the desired beverage through a simple interface. This makes it easy for the user to receive a beverage that meets their needs.

[0574] The output device presents the generated recipe to the user visually or audibly. For example, it might display details of the suggested beverage on a display in the store or on the user's smartphone screen. The user can then choose their preferred option from the presented choices and receive their beverage after final confirmation.

[0575] For example, if a cafe customer says they want to relax, the system analyzes their voice tone and input to suggest a relaxing beverage. In this example, the generative AI model might recommend a decaf latte. The user agrees to the suggested beverage and confirms their order.

[0576] Example of a prompt:

[0577] "Please suggest beverage recipes that should be offered to users who want to relax."

[0578] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0579] Step 1:

[0580] The device receives voice input from the user. This input includes the user's current emotional state and a description of the beverage they desire. The microphone on the device captures the voice and records it as audio data.

[0581] Step 2:

[0582] The server receives audio data sent from the terminal. It uses a speech recognition engine to convert this audio data into text data. This conversion process utilizes the Python library `speech_recognition` to format the text from the audio. The converted text clearly reflects what the user said.

[0583] Step 3:

[0584] The server processes text data using sentiment analysis technology. This process utilizes "EmotionEngine" to estimate the user's emotional state from the text. This estimation is based on the tone and keywords extracted from the text, resulting in the generation of data indicating the user's emotional state.

[0585] Step 4:

[0586] The server uses a generative AI model to generate optimal beverage recipes based on estimated emotional states. The model takes emotional data and information about the user's past preferences as input and outputs the optimal beverage corresponding to the user's current emotional state. This generation is performed using specific prompts.

[0587] Step 5:

[0588] The server sends the generated recipe to the terminal. The terminal presents this beverage information to the user visually or audibly. The user can confirm the suggested beverage through on-screen displays or audio guidance.

[0589] Step 6:

[0590] The user reviews the presented recipe and makes a selection. The user's selection is entered into the terminal and notified to the server, confirming the final order. The system then completes the provision of a beverage tailored to the user's emotional state.

[0591] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0592] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0593] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0594] [Fourth Embodiment]

[0595] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0596] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0597] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0598] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0599] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0600] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0601] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0602] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0603] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0604] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0605] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0606] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0607] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0608] This invention is a system aimed at providing the optimal coffee by utilizing information about coffee beans. The system mainly consists of three components: a server, a terminal, and a user, each with its own role.

[0609] Server roles and operations

[0610] The server maintains a database that manages and updates characteristic data such as coffee bean type, origin, and roast level. It also uses a generative model based on the collected data to generate the optimal coffee recipe. When a request is sent from a user via the terminal, the server processes the relevant data and performs the necessary calculations to provide a recipe tailored to the user's preferences and health condition.

[0611] Terminal role and operation

[0612] The terminal receives coffee bean information and personal preferences entered by the user and provides an interface for sending this information to the server. Furthermore, it receives recipe information sent from the server and generates an operation screen for the user to customize based on that information. It also has voice recognition technology, which allows it to analyze the voice spoken by the user and execute the commands necessary for the brewing process.

[0613] User roles and actions

[0614] Users input information about the coffee beans they have on hand into the device, and also record their taste preferences and health status. Furthermore, they can adjust the recipe using the customization options presented on the device. Finally, by giving a voice command to start brewing and operating the system, they can obtain coffee of their desired quality.

[0615] Specific example

[0616] For example, consider a user who wants to use "Ethiopian Yirgacheffe beans" and brew coffee to their liking. First, the user enters the details of the beans on the device, and then specifies their preferred taste. The device sends this information to the server, which generates the optimal brewing method and sends it back to the device. The device makes fine adjustments according to the user's wishes and can finally initiate brewing via voice recognition. Through this process, the user can easily enjoy a customized cup of coffee tailored to their needs.

[0617] The following describes the processing flow.

[0618] Step 1:

[0619] The user inputs information about their coffee beans (e.g., origin and type), as well as their own taste preferences and health status, into the terminal. The terminal formats this information and stores it in a way that allows for subsequent data processing.

[0620] Step 2:

[0621] The terminal generates an API request to send the entered data to the server. This request contains the user's input data, which the server uses as the basis for processing.

[0622] Step 3:

[0623] The server receives requests sent from terminals and analyzes the data. Based on the analysis results, it uses a generative model to generate recipes suitable for the characteristics of the coffee beans. The generated recipes include recommended brewing conditions.

[0624] Step 4:

[0625] The server sends the generated recipe to the terminal. The terminal prepares a user interface to display the received recipe in a user-friendly format.

[0626] Step 5:

[0627] Users can view the generated recipe along with the provided customization options through their device's display screen. They can further adjust the flavor and concentration according to their preferences and then see the results.

[0628] Step 6:

[0629] Once the adjustments are complete, the user uses voice commands to give instructions to the device, such as "Please prepare some coffee." The voice is then converted to text by the device's speech recognition technology.

[0630] Step 7:

[0631] The terminal uses voice recognition to confirm the user's instructions and then controls devices such as a coffee grinder to begin the actual coffee brewing process. Once brewing is complete, the user is notified, and the finished coffee is served.

[0632] (Example 1)

[0633] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0634] Conventional beverage preparation systems have faced challenges in generating optimal recipes tailored to individual users' preferences and health conditions, and in easily brewing and customizing them. Furthermore, they lacked sufficient voice control and integration capabilities with other household electrical appliances, resulting in poor usability.

[0635] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0636] In this invention, the server includes a device for receiving characteristic data about coffee beans, a device for utilizing a generation model that generates optimal beverage preparation instructions based on the characteristic data, and a device for adjusting the extraction of the beverage based on the generated instructions. This enables the generation of recipes tailored to the individual needs of the user and facilitates easy extraction operations.

[0637] "Characteristic data about coffee beans" refers to data that includes information such as the type of coffee bean, its origin, roast level, and quality.

[0638] A "generative model" is a machine learning model that generates optimal beverage preparation instructions for the user based on the data received.

[0639] "Beverage preparation instructions" are instructions that specify the extraction conditions and procedures necessary for preparing a beverage.

[0640] "Personalized settings options" is a feature that provides a customizable menu of options tailored to the user's preferences and health condition.

[0641] "Voice recognition technology" is a technology that analyzes the user's voice input and interprets it as appropriate instructions.

[0642] "Household electrical appliances" refers to various electrical devices used in the home, and this system can be operated in conjunction with them.

[0643] This invention is a system that utilizes the characteristic information of coffee beans possessed by the user and provides optimal beverage preparation instructions based on that information. The system mainly consists of three components: a server device, a terminal device, and the user, each playing a different role.

[0644] Server device operation

[0645] The server device maintains a database that manages and updates the received coffee bean characteristic data. Specifically, the server uses open-source data management software to collect and organize this information. Next, based on the data provided by the user, it generates optimal beverage preparation instructions using a generative AI model. In this process, the server uses a programming language and leverages machine learning libraries to implement the model and process the data.

[0646] Terminal device operation

[0647] The terminal device provides an interface for receiving coffee bean information and personal settings entered by the user. It also receives cooking instructions sent from the server and generates an operation screen for the user to customize their coffee based on those instructions. Specifically, the terminal is designed as an application that runs on a mobile communication device, and is easily operated by the user using touch controls and voice recognition technology.

[0648] User actions

[0649] Users input information about the coffee beans they have on hand via a terminal device and record their taste preferences and health status. They can further refine the brewing instructions using customization options. Finally, they give voice commands to the terminal to begin brewing, and the machine brews the beverage according to their requests.

[0650] Specific example

[0651] For example, consider a scenario where a user wants to brew their preferred coffee using specific Ethiopian beans. The user first inputs details of the beans into a terminal device, specifying their desired flavor profile, such as "fruity and acidic." This information is sent to a server device, which then inputs a prompt message into its AI model: "Ethiopian Yirgacheffe beans, fruity, acidic flavor, suggest optimal brewing method." This model generates optimal brewing instructions and sends the result back to the terminal device. The user can further customize the information using the provided options and finally initiate brewing using voice recognition to obtain their desired quality beverage.

[0652] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0653] Step 1:

[0654] The user inputs information about the characteristics of the coffee beans on the terminal device. Using a touchscreen or voice input function, the user enters information such as "Bean type: Ethiopian," "Flavor preference: Fruity," and "Roast level: Medium." This information is structured in JSON format and prepared for transmission to the server.

[0655] Step 2:

[0656] The terminal sends the information entered by the user to the server. This input is information about coffee beans entered by the user. The terminal uses the HTTPS protocol to securely send this data to the server and internally receives a confirmation message that the transmission is complete.

[0657] Step 3:

[0658] The server operates a generative AI model based on the received coffee bean characteristic data. First, the server analyzes the received data and forms a prompt sentence such as "Ethiopian beans, fruity, medium roast." This prompt sentence is input into the generative AI model, which then generates optimal beverage preparation instructions based on the collected data. These output instructions are temporarily stored in the server's database.

[0659] Step 4:

[0660] The server sends back the generated beverage preparation instructions to the terminal. These beverage preparation instructions are prepared on the terminal to facilitate user operation. Based on the received instructions, the terminal displays them in a visually easy-to-understand format on its user interface.

[0661] Step 5:

[0662] Users can review the customization options on their device and make adjustments as needed. Here, they can use sliders and dropdown menus to modify the instructions. The results of the customizations are immediately reflected on the device.

[0663] Step 6:

[0664] The user ultimately gives a voice command to the terminal device to start the extraction process. The final output is the commencement of the actual extraction process. The terminal uses voice recognition technology to process this command and sends a signal to the extraction machine to execute the instructions. This allows the user to obtain a beverage that meets their preferences.

[0665] (Application Example 1)

[0666] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0667] Traditional systems are insufficient to provide optimal beverages tailored to the individual preferences and health conditions of each customer. Furthermore, there is a need for appropriate means to streamline the in-store ordering process and provide personalized services. Additionally, improving the efficiency of store services and enhancing customer convenience through integration with mobile communication terminals is a key challenge.

[0668] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0669] In this invention, the server includes means for receiving characteristic data about coffee beans, means for using a generation model that generates an optimal beverage recipe based on the characteristic data, and means for performing operations in cooperation with a mobile communication terminal. This speeds up order processing at the point of sale and enables the provision of beverages optimized for the user.

[0670] "Characteristic data regarding coffee beans" includes data such as the type of coffee bean, origin, roast level, and other physical characteristics.

[0671] An "optimal beverage recipe" is a guideline regarding the characteristics and brewing methods of a beverage, generated based on the user's preferences and health condition.

[0672] "Methods using generative models" refer to mechanisms for creating optimal beverage recipes using algorithms based on data.

[0673] "Operation via collaboration with mobile communication terminals" refers to a method of operating a system through communication devices such as smartphones, and performing data transmission, reception, and execution of instructions.

[0674] "Speech recognition technology" is a technology that analyzes speech data, converts it into text, and then understands and executes commands based on that text.

[0675] This system consists of three components: servers, terminals, and users.

[0676] The server first manages characteristic data about coffee beans in a database. This characteristic data includes information such as the place of origin, type of bean, and roast level. Based on this data, it utilizes a generative AI model to generate optimal beverage recipes tailored to each user's preferences and health condition. The generated recipes are sent to the terminal and include customization options according to the user's requests. The server uses cloud computing platforms (e.g., AWS and Google Cloud) to perform high-speed data processing.

[0677] The terminal provides an interface for receiving information about coffee beans and taste preferences from users. Mobile communication devices such as smartphones and tablets are used, and the collected data is transmitted to a server via the internet. The terminal also utilizes speech recognition technology; for example, by using the Azure Cognitive Services speech recognition API, it can analyze voice commands uttered by users and execute appropriate instructions.

[0678] Users can input their coffee preferences into the system, review dynamically generated recipes, and customize them as needed. For example, a user can voice-instruct, "Please tell me the best way to brew a fruity coffee using Ethiopian Yirgacheffe beans," and receive a coffee tailored to their taste.

[0679] This system configuration allows for the provision of optimal beverages tailored to each individual user, improving the in-store experience.

[0680] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0681] Step 1:

[0682] Users use their smartphones to input information such as the type of coffee beans they like, their preferred taste, and their health status. This data includes the type of beans (e.g., Ethiopian Yirgacheffe), taste profile, and desired caffeine level. This data is collected via the device's interface and transmitted to a server over the internet.

[0683] Step 2:

[0684] The server stores the received coffee bean characteristic data and user preference data, and passes this to a generative AI model. The generative AI model processes this data and calculates the optimal beverage recipe customized for each user. Based on the input data, it searches for existing recipes in the database, and the AI ​​analyzes each data combination to generate a recipe that includes the optimal brewing method. As a result of this process, the generated recipe is stored in a temporary resource on the server.

[0685] Step 3:

[0686] Recipes generated on the server are sent to the terminal in real time. The terminal receives this information and displays it visually to the user. The user can then customize the recipe if necessary. For example, the user can use an interface to adjust the level of sweetness and richness to their liking and send their selections back to the server from the terminal.

[0687] Step 4:

[0688] Based on the final customized recipe displayed on the terminal, the user issues a voice command to start brewing. The terminal recognizes this voice and converts it to text using the Azure Cognitive Services speech recognition API. The recognized command is then processed as a signal to control the brewing process. For example, if the voice input is "Please start brewing," the system will analyze it and send a command to the coffee maker accordingly.

[0689] Step 5:

[0690] The server logs the completed process based on the finalized customization information that has been re-received. This allows the system to retain the user's satisfied coffee profile and data, and continuously collect information for future reconstruction. The collected data also serves as the basis for statistical analysis, which can be used to further improve the model.

[0691] This process provides customized beverages that match the user's preferences, resulting in a personalized beverage experience.

[0692] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0693] This invention is a system that recognizes the user's emotions and provides the optimal coffee based on them. In particular, it can combine characteristic data about coffee beans with user information to customize the coffee according to the user's emotional state. This makes it possible to more individualize the coffee brewing process and provide a coffee experience that is tailored to the user's mood and state of mind.

[0694] Server roles and operations

[0695] The server not only manages a database containing various information about coffee beans, but also analyzes emotional data sent by users. Using an emotion engine, it analyzes the user's emotional state as data in real time and generates the most suitable recipe for each user. This includes processing voice data and other related data.

[0696] Terminal role and operation

[0697] The device receives voice input from the user and analyzes it using speech recognition technology. Furthermore, an emotion engine estimates the user's emotions based on the tone and word choice read from the voice. The estimated emotions are sent to a server for processing. The device then presents the user with a recipe generated considering the emotions and supports customization as needed.

[0698] User roles and actions

[0699] Users not only input information about coffee beans into the device, but also enjoy personalized adjustments based on their emotions. For example, they can easily prepare a coffee tailored to their mood, such as when they want to relax or concentrate, by giving voice commands.

[0700] Specific example

[0701] Let's take an example where a user is feeling stressed and wants to calm down to alleviate that stress. The user speaks to the device, saying, "I'd like a coffee that will help me relax a bit." The device analyzes the voice, inferring the user's emotions from their tone of voice and converting it into data. The server receives this data and uses an emotion engine to generate a recipe that is expected to have the most relaxing effect. This recipe is presented to the user via the device, and after final confirmation, the coffee is brewed and served. Through this process, the user can obtain a special cup of coffee that matches their emotions.

[0702] The following describes the processing flow.

[0703] Step 1:

[0704] The user communicates their current emotional state to the device via voice. For example, they might use a voice command such as, "I want to calm down a bit today." The device then uses speech recognition technology to convert this into text.

[0705] Step 2:

[0706] The device analyzes the text data converted from the speech and the resulting tone of voice to estimate the user's emotions. An emotion engine is used for this purpose. The estimated results are then generated as data and sent to the server.

[0707] Step 3:

[0708] The server analyzes the emotional data received from the terminal and uses an emotion engine to interpret the user's current mood and desires. Based on this, it applies a coffee recipe generation model to construct the optimal coffee recipe for the user.

[0709] Step 4:

[0710] The server sends the created recipe to the terminal. The terminal receives this information and prepares to display it to the user in an easy-to-understand visual format.

[0711] Step 5:

[0712] The user looks at the device screen and reviews the suggested recipe. They can then customize it further as needed or accept the suggestion as is.

[0713] Step 6:

[0714] Once the user is satisfied with the recipe, they can instruct the device to start brewing the coffee using voice or touch controls.

[0715] Step 7:

[0716] The device controls the coffee grinder and brewing equipment according to the user's instructions, and brews coffee according to the specified recipe. When the finished coffee is ready, the device notifies the user.

[0717] (Example 2)

[0718] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0719] In modern society, users demand customized beverage experiences that cater to diverse emotions and preferences. However, conventional beverage delivery systems have struggled to provide individualized beverages that accurately reflect users' emotional states. This problem hinders improvements in user satisfaction, creating a need for technology that more faithfully reflects user emotions and delivers ideal beverages.

[0720] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0721] In this invention, the server includes means for acquiring characteristic information about coffee, means for receiving and analyzing voice input from the user, and means for estimating the user's emotions based on the acquired voice data. This makes it possible to provide optimized beverages that meet the user's emotions and individual needs.

[0722] "Coffee characteristic information" refers to data related to the characteristics of coffee, such as its taste, aroma, origin, and production method.

[0723] "Means for receiving and analyzing voice input from a user" refers to the process and technology for recognizing a user's voice instructions, analyzing their content, and converting them into digital data.

[0724] "Means for estimating user emotions" refers to algorithms or systems that analyze and infer a user's emotional state based on voice data and other relevant information.

[0725] A "generative model for generating optimal beverage recipes" is an algorithm or model that calculates and suggests the most suitable preparation method and ingredients for a beverage based on the user's emotional state and preferences.

[0726] "Means for adjusting extraction conditions and providing beverages" refers to the technology of setting the parameters necessary for beverage extraction (temperature, pressure, time, etc.) based on a generated recipe, and preparing the actual beverage accordingly.

[0727] "Means for coordinating and operating other home appliances" refers to technical processes or devices that communicate with coffee machines and other related home appliances to automate or coordinate a series of processes.

[0728] This invention is a system for providing customized beverages that respond to a user's emotions. The system includes acquiring characteristic information about coffee, voice analysis, emotion estimation, recipe generation using a generative AI model, and providing the beverage under optimal brewing conditions.

[0729] server

[0730] The server has a database that manages extensive characteristic information about coffee and analyzes emotional data from users. This includes processing data obtained from voice input in real time and identifying the user's emotional state using an emotion analysis engine. Based on the emotional data, the server uses a generative AI model to calculate and generate the optimal beverage recipe for the user.

[0731] terminal

[0732] The terminal's role is to receive the user's voice input. This process is achieved through a microphone and speech recognition technology. The speech recognition technology converts the user's utterance into text data and sends it to the emotion engine. The terminal also presents the generated recipe to the user, and after receiving the user's final confirmation, operates related equipment such as a coffee machine to dispense the beverage.

[0733] User

[0734] Users provide their emotional state and beverage preferences to the device via voice input. This information is sent to a server via the device, where the most suitable beverage recipe is generated. For example, if a user says, "I want to relax," the system will suggest a beverage recipe that enhances relaxation based on that emotion.

[0735] Example of a prompt

[0736] One example of a prompt is: "Please suggest ingredients and extraction methods to provide the optimal beverage recipe for when the user wants to relax." This prompt is an instruction to use a generative AI model to derive a recipe that meets the user's specific needs.

[0737] In this way, the present invention makes it possible to provide a high-quality and rapid beverage experience that responds to the user's emotions.

[0738] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0739] Step 1:

[0740] The device acquires the user's voice input. When the user speaks, voice data is captured by the microphone. The input is voice data that expresses the user's requests and emotions. The device performs noise cancellation and processes the input into pure voice data. The output is clear voice data.

[0741] Step 2:

[0742] The device converts acquired audio data into text data using speech recognition technology. The input is audio data, and natural language processing algorithms are used in the analysis process. In this stage of data processing, the audio waveform is converted into text, and the output is text data representing the sentences spoken by the user.

[0743] Step 3:

[0744] The device estimates the user's emotions from text data. The emotion engine analyzes keywords and context within the text to estimate the emotional state. The input is text data, the emotion analysis results are output, and the emotional state is digitized.

[0745] Step 4:

[0746] The server receives the emotion analysis results and retrieves characteristic information about coffee from the database. The input consists of emotional state data and user preference information. A generative AI model integrates this data to generate the optimal beverage recipe. The output is a beverage recipe tailored to the user's needs.

[0747] Step 5:

[0748] The terminal presents the generated beverage recipe to the user and requests final confirmation. The input is beverage recipe data, and the terminal uses this to send information to the user visually or audibly. Once the user confirms, the final request is approved as output.

[0749] Step 6:

[0750] The server sets the extraction conditions and controls other appliances such as coffee machines based on approved beverage recipes. The input is the final approved recipe data, which is used to physically extract the beverage. The output is the extracted beverage.

[0751] This series of processes enables the system to provide personalized beverages tailored to the user's emotions and preferences.

[0752] (Application Example 2)

[0753] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0754] In modern restaurants, providing beverages tailored to each customer's emotional state and preferences is not easy. Traditional service methods lack the flexibility to meet diverse customer needs, and suggesting beverages based on a customer's emotional state is particularly challenging. This necessitates providing more personalized service to customers.

[0755] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0756] In this invention, the server includes means for analyzing the user's emotional state in real time using emotion analysis technology, means for using a generation model that generates an optimal beverage recipe based on the analyzed emotional state, and means for adjusting the beverage delivery based on the generated recipe. This makes it possible to provide personalized beverages that correspond to the user's emotional state.

[0757] "Emotion analysis technology" is a technology that estimates a user's emotional state in real time from voice and other data.

[0758] A "generative model" refers to an algorithm or method that generates optimal results or suggestions based on input data.

[0759] "Means for coordinating beverage provision" refers to a system for optimizing the process of preparing beverages according to generated recipes and providing them to users.

[0760] "Customization options" refer to settings and choices that allow users to select the most suitable services or products according to their preferences and needs.

[0761] "Speech recognition technology" is a technology that converts voice input into text data and then analyzes it.

[0762] An "output device" is a device or equipment used to present processed information to the user.

[0763] To realize this invention, a system combining multiple technological elements is required. This system consists of a cloud server, a user-operated terminal, and an output device.

[0764] The server receives the user's voice data and converts it into text data using speech recognition technology. This text data is then used to estimate the user's emotional state using sentiment analysis technology. Specifically, speech recognition is performed using the Python `speech_recognition` library, and the emotional state is analyzed using "EmotionEngine." Based on the analyzed emotional information, the server generates the optimal beverage recipe using a generative AI model. This generative model combines user information and emotional state to suggest personalized beverages.

[0765] The terminal is responsible for receiving voice input from the user and sending it to the server. Voice input is mainly performed on devices such as smartphones and tablets, and a voice recognition process is executed. The user communicates their emotional needs for the desired beverage through a simple interface. This makes it easy for the user to receive a beverage that meets their needs.

[0766] The output device presents the generated recipe to the user visually or audibly. For example, it might display details of the suggested beverage on a display in the store or on the user's smartphone screen. The user can then choose their preferred option from the presented choices and receive their beverage after final confirmation.

[0767] For example, if a cafe customer says they want to relax, the system analyzes their voice tone and input to suggest a relaxing beverage. In this example, the generative AI model might recommend a decaf latte. The user agrees to the suggested beverage and confirms their order.

[0768] Example of a prompt:

[0769] "Please suggest beverage recipes that should be offered to users who want to relax."

[0770] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0771] Step 1:

[0772] The device receives voice input from the user. This input includes the user's current emotional state and a description of the beverage they desire. The microphone on the device captures the voice and records it as audio data.

[0773] Step 2:

[0774] The server receives audio data sent from the terminal. It uses a speech recognition engine to convert this audio data into text data. This conversion process utilizes the Python library `speech_recognition` to format the text from the audio. The converted text clearly reflects what the user said.

[0775] Step 3:

[0776] The server processes text data using sentiment analysis technology. This process utilizes "EmotionEngine" to estimate the user's emotional state from the text. This estimation is based on the tone and keywords extracted from the text, resulting in the generation of data indicating the user's emotional state.

[0777] Step 4:

[0778] The server uses a generative AI model to generate optimal beverage recipes based on estimated emotional states. The model takes emotional data and information about the user's past preferences as input and outputs the optimal beverage corresponding to the user's current emotional state. This generation is performed using specific prompts.

[0779] Step 5:

[0780] The server sends the generated recipe to the terminal. The terminal presents this beverage information to the user visually or audibly. The user can confirm the suggested beverage through on-screen displays or audio guidance.

[0781] Step 6:

[0782] The user reviews the presented recipe and makes a selection. The user's selection is entered into the terminal and notified to the server, confirming the final order. The system then completes the provision of a beverage tailored to the user's emotional state.

[0783] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0784] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0785] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0786] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0787] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0788] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0789] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0790] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0791] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0792] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0793] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0794] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0795] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0796] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0797] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0798] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0799] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0800] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0801] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0802] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0803] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0804] The following is further disclosed regarding the embodiments described above.

[0805] (Claim 1)

[0806] A means for receiving characteristic data about coffee beans,

[0807] A means of using a generative model that generates an optimal coffee recipe based on the characteristic data,

[0808] Means for adjusting the coffee extraction based on the generated recipe,

[0809] A means of presenting customization options based on user information,

[0810] A means for receiving and processing user commands using speech recognition technology,

[0811] A means of operating in conjunction with other household electrical appliances,

[0812] A system that includes this.

[0813] (Claim 2)

[0814] The system according to claim 1, which utilizes a database containing information on the origin and producer of coffee beans.

[0815] (Claim 3)

[0816] The system according to claim 1, which dynamically generates recipes based on the user's health status and preferences.

[0817] "Example 1"

[0818] (Claim 1)

[0819] A device that receives characteristic data about coffee beans,

[0820] A device that utilizes a generation model that generates optimal beverage preparation instructions based on the characteristic data,

[0821] A device that adjusts the extraction of a beverage based on the generated instructions,

[0822] A device that presents personalized setting options based on user information,

[0823] A device that uses speech recognition technology to receive and process user instructions,

[0824] A device that operates in conjunction with other household electrical appliances,

[0825] A system that includes this.

[0826] (Claim 2)

[0827] The system according to claim 1, which uses a set of information including information about the origin and manufacturer of coffee beans.

[0828] (Claim 3)

[0829] The system according to claim 1, which dynamically generates beverage preparation instructions based on the user's health condition and personal preferences.

[0830] "Application Example 1"

[0831] (Claim 1)

[0832] A means for receiving characteristic data about coffee beans,

[0833] A means of using a generative model that generates an optimal beverage recipe based on the characteristic data,

[0834] Means for adjusting the beverage based on the generated recipe,

[0835] A means of presenting customized options based on user information,

[0836] A means for receiving and processing user commands using speech recognition technology,

[0837] A means of performing operations in cooperation with a mobile communication terminal,

[0838] A system that includes this.

[0839] (Claim 2)

[0840] The system according to claim 1, utilizing an information management device that includes information on the origin and manufacturer of coffee beans.

[0841] (Claim 3)

[0842] The system according to claim 1, which dynamically generates recipes based on the user's health condition and preferences.

[0843] "Example 2 of combining an emotion engine"

[0844] (Claim 1)

[0845] A means of obtaining characteristic information about coffee,

[0846] A means of receiving and analyzing voice input from the user,

[0847] A method for estimating the user's emotions based on acquired audio data,

[0848] A means of using a generative model that generates an optimal beverage recipe based on estimated emotional information,

[0849] A means of providing a beverage by adjusting the extraction conditions based on the generated recipe,

[0850] A means for coordinating and operating the aforementioned system with other home appliances,

[0851] A system that includes this.

[0852] (Claim 2)

[0853] The system according to claim 1, which utilizes a database containing geographical and producer information related to coffee.

[0854] (Claim 3)

[0855] The system according to claim 1, which dynamically generates a beverage preparation method based on the user's emotional state and personal preferences.

[0856] "Application example 2 of combining emotional engines"

[0857] (Claim 1)

[0858] A means of analyzing a user's emotional state in real time using emotion analysis technology,

[0859] A method using a generative model that generates the optimal beverage recipe based on the analyzed emotional state,

[0860] Means for adjusting the provision of beverages based on the generated recipe,

[0861] A means of presenting customization options based on user information,

[0862] A means for receiving and processing user voice commands using speech recognition technology,

[0863] A means for presenting customized beverage information to an output device,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] The system according to claim 1, which provides a beverage that corresponds to the emotional state via voice input in a store environment.

[0867] (Claim 3)

[0868] The system according to claim 1, which automatically recommends a beverage recipe based on an estimated emotional state and supports the preparation of the beverage. [Explanation of symbols]

[0869] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving characteristic data about coffee beans, A means of using a generative model that generates an optimal coffee recipe based on the characteristic data, Means for adjusting the coffee extraction based on the generated recipe, A means of presenting customization options based on user information, A means for receiving and processing user commands using speech recognition technology, A means of operating in conjunction with other household electrical appliances, A system that includes this.

2. The system according to claim 1, which utilizes a database containing information on the origin and producer of coffee beans.

3. The system according to claim 1, which dynamically generates recipes based on the user's health status and preferences.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A