System
The system addresses the lack of nutritional and exercise guidance in traditional health management by using an emotional engine and a generation AI model to provide personalized health recommendations based on emotional state and nutritional content.
Patent Information
- Application Number
- JP2024181992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-23
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-08
AI Technical Summary
Traditional health management systems for children lack sufficient nutritional information and exercise guidance, making it difficult for users to make healthy choices.
A system that includes an emotional engine to recognize user emotions, a nutritional component acquisition means to retrieve food information from a database, and a generation AI model to propose foods and exercises based on emotional state and nutritional content.
The system enables users to easily access nutritional information and exercise suggestions tailored to their emotional state, facilitating healthier choices and more effective health management.
Smart Images

Figure 2025071787000001_ABST
Abstract
Description
[Technical field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional health management for children involves insufficient provision of nutritional information (including information on nutritional components) and exercise guidance, making it difficult for children to make healthy choices. [Means for solving the problem]
[0005] A system according to one embodiment of the present disclosure includes an emotion engine that recognizes a user's emotions, an acquisition means that acquires information about foods the user is viewing on a terminal, a nutritional component acquisition means that acquires information about the nutritional components of the foods the user is viewing from a food database, a suggestion means that suggests foods suitable for the user's emotions based on the user's emotions and the information about the nutritional components of the foods by inputting a specific prompt sentence into a generative AI model, a generation means that uses the generative AI model to generate a guide for a question about food from the user, a customization means that customizes the guide for the question based on the user's emotions, and a provision means that provides foods suitable for the user's emotions and the customized guide to the user via the user's terminal.
[0006] A system according to one aspect of the present disclosure includes a nutritional component acquisition means for acquiring information regarding the nutritional components of the food a user is viewing from a food database, an exercise information means for acquiring information regarding exercise from an exercise database, an answer generation means for generating a guide to a user's question regarding exercise using the data generation model, and a display means for displaying a guide including an exercise form corresponding to the question on the specific terminal.
[0007] In a system according to one aspect of the present disclosure, the information processing means includes a voice recognition means for converting voice commands from a user into text.
[0008] A system according to one aspect of the present disclosure makes it easier to make healthy choices.
[0009] A "food database" may be interpreted as a database in which nutritional information, food characteristics, etc. are recorded.
[0010] The "exercise database" may be interpreted as a database in which types of exercise, their effects, points to note, etc. are recorded.
[0011] The "data generation model" may be interpreted as an interactive artificial intelligence system that uses natural language processing technology. The data generation model may be interpreted as a generative AI that generates specific answers and advice in response to user questions. [Brief description of the drawings]
[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0016] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by the processor.
[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0018] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0020] [First embodiment]
[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0025] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.
[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.
[0028] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.
[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0032] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0033] The embodiment for implementing the present disclosure includes the following elements.
[0034] 1. Server: The server has a function to access a food database and an exercise database. The server has a function to receive information from a user and generate specific questions for the generative AI. The server has a function to receive answers generated by the generative AI and send them to a terminal. The terminal may be interpreted as a device such as a mobile phone (e.g., a smartphone or tablet) or a computer, and can display and operate information.
[0035] 2. Terminal: The terminal may be interpreted as smart glasses and has a function of displaying information through the smart glasses or a screen. The smart glasses may be interpreted as a glasses-type device and has a function of displaying information and playing audio and video. The terminal has a function of displaying the answer received from the server.
[0036] 3. User: A user can use the system via a server or a terminal when viewing food or performing exercise.
[0037] In this way, users can understand the nutritional information of foods and the appropriate intake amount, and learn about the effects of exercise and the correct form. Specifically, when a user looks at a food, the nutritional information is displayed on the screen of the smart glasses or device. In addition, when the user exercises, exercise suggestions and form tips are displayed. Furthermore, when the user asks the generative AI a question, specific advice and information are obtained.
[0038] In the above manner, users can utilize a system that supports healthy choices.
[0039] The process flow will be explained below.
[0040] Step 1: The user sends information about whether to view food or exercise to the server.
[0041] Step 2: When the user views a food, the server obtains the nutritional information of the corresponding food from the food database.
[0042] Step 3: When the user exercises, the server obtains the corresponding exercise information from the exercise database.
[0043] Step 4: Based on the acquired information, the server generates specific questions for the generative AI.
[0044] Step 5: The generative AI analyzes the received question and generates an appropriate answer.
[0045] Step 6: The server receives the generated answer from the generative AI.
[0046] Step 7: The server sends the received response to the terminal.
[0047] Step 8: The device displays the received answer on the screen of the smart glasses or the device.
[0048] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0049] 1: Information on whether the user is viewing food or exercising
[0050] 2: Nutritional information of food obtained from food database
[0051] 3: Exercise information obtained from the exercise database
[0052] 4: Specific questions for generative AI
[0053] 5: Answers generated by generative AI
[0054] 6: Answer sent from the server to the device
[0055] 7: Answer displayed on the terminal
[0056] Example 1
[0057] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0058] Conventional health management systems make it difficult for users to easily obtain nutritional information about foods and appropriate forms for exercise. They also make it difficult to provide quick and accurate answers to specific questions from users. For this reason, there is a demand for a system that allows users to quickly and easily obtain information to make healthy choices.
[0059] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0060] In this invention, the server includes an information processing means for generating a question for the generative AI model based on information provided by the user, an answer receiving means for receiving an answer generated from the generative AI model and transmitting it to the user's terminal, and a display means for displaying information corresponding to the question on a specific terminal. This allows the user to easily obtain nutritional information on foods and an appropriate form for exercise, and to obtain a quick and accurate answer to a specific question. The server may include an acquisition means for acquiring information on the food viewed by the user on the terminal, a nutritional information acquisition means for acquiring information on the nutritional information of foods from a food database, a suggestion means for suggesting foods suitable for the user's emotions based on the user's emotions and the information on the nutritional information of foods by inputting a specific prompt sentence into the generative AI model, a generation means for generating a guide for a question about food from the user using the generative AI model, a customization means for customizing the guide for the question based on the user's emotions, and a provision means for providing the user with foods suitable for the user's emotions and the customized guide via the user's terminal.
[0061] The nutritional component acquisition means is a means for acquiring information on the nutritional components of a food viewed by a user. Specifically, when a user is viewing a specific food on a terminal, an acquisition means for acquiring information on the food viewed by the user on the terminal acquires the food information, and then the nutritional component acquisition means acquires information on the nutritional components of the food viewed by the user from a food database.
[0062] An "information processing means" is a means for generating questions for a generative AI model based on information provided by a user.
[0063] The "answer receiving means" is a means for receiving an answer generated from the generative AI model and transmitting it to the user's terminal.
[0064] The "display means" is a means for displaying information corresponding to a question on a specific terminal.
[0065] The "nutritional component acquisition means" is a means for acquiring information about the nutritional components of the food the user is viewing from the food database.
[0066] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0067] The "answer generation means" is a means for generating guides for the user's questions about exercise and food using a data generation model.
[0068] A "voice recognition means" is a means for converting voice commands from a user into text.
[0069] A "terminal" is a device used by a user, such as smart glasses, that has the ability to display information and play audio and video.
[0070] A "generative AI model" is a model that uses pre-trained data to generate appropriate answers to user questions.
[0071] This invention is a system for allowing a user to easily obtain nutritional information on foods and appropriate forms of exercise, and to obtain quick and accurate answers to specific questions. This system is mainly composed of three elements: a server, a terminal, and a user. The server may include: a nutritional information acquisition means for acquiring information on nutritional information on foods from a food database; a suggestion means for suggesting foods suitable for the user's emotions based on the user's emotions and information on nutritional information on foods by inputting a specific prompt sentence into a generative AI model; a generation means for generating a guide for a question about food from the user using the generative AI model; a customization means for customizing the guide for the question based on the user's emotions; and a provision means for providing the user with foods suitable for the user's emotions and the customized guide via the user's terminal. The server may include an exercise information means for acquiring information on exercise from an exercise database; and a generation means for generating an appropriate form for the exercise asked by the user as a guide for the question by inputting a specific prompt sentence into a generative AI model, and the provision means may provide the appropriate form for the exercise to the user via the user's terminal.
[0072] Server Roles
[0073] The server uses a high-performance cloud server (for example, a cloud computing service). The server has the following functions:
[0074] Information processing means: The server creates appropriate questions for the generative AI model based on information provided by the user, such as voice commands. It uses voice recognition software (e.g., a voice recognition service) to convert voice to text.
[0075] Answer receiving means: Receives the answer generated by the generative AI model and sends it to the user's device. During this process, it checks the accuracy of the answer and adjusts the format if necessary.
[0076] Database access function: Accesses food database and exercise database and obtains necessary information from each database.
[0077] Specifically, the server uses voice recognition software to convert the user's voice command into text, and then sends the text, such as "What are the nutrients in this apple?", as a question to the generative AI model.
[0078] Terminal Roles
[0079] The terminal is a device used by a user, and in this invention we assume it is a smart glass. This terminal has the following functions:
[0080] Display means: Displays the information sent from the server. Has the function of visually presenting information on the smart glasses display.
[0081] User interface: Provides an interface for voice command input and touch operation.
[0082] Specifically, when a user picks up an apple and says to the smart glasses, "Tell me the nutritional value of this apple," the apple's calories and nutritional components will be displayed on the smart glasses' display.
[0083] User Roles
[0084] When users view foods or exercise, they receive information via the server or device, which helps them make healthy choices.
[0085] The specific action involves a user picking up an apple and inputting a voice command to the smart glasses, such as, "Tell me the nutritional value of this apple." The answer generated by the generative AI model is, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C," which is then displayed on the smart glasses' display.
[0086] Examples of prompt statements
[0087] Examples of prompts include:
[0088] "How much vitamin C does this apple contain?"
[0089] "How many calories can you burn in 30 minutes of running?"
[0090] As a result, by using the system of the present invention, a user can easily obtain nutritional information about foods and appropriate exercise forms, and make healthy choices based on that information.
[0091] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0092] Divide the program flow into processing steps
[0093] Step 1: User Input
[0094] Step 2: Server processes information and generates a query
[0095] Step 3: Generate answers using generative AI models
[0096] Step 4: Server receives and sends response
[0097] Step 5: Displaying information via terminal
[0098] Detailed explanation of each processing step
[0099] Step 1: User Input
[0100] When a user uses the system, they input the necessary information into the system through voice commands, touch operations, etc. This allows the system to collect information about foods and exercise that interest the user.
[0101] Specific behavior:
[0102] The user picks up an apple and asks the smart glasses verbally, "What are the nutrients in this apple?"
[0103] Input: Voice command "What are the nutritional values of this apple?"
[0104] Output: Audio data
[0105] Step 2: Server processes information and generates a query
[0106] The server uses speech recognition software to convert the voice data into text, which is then used to generate appropriate questions for the generative AI model.
[0107] Specific behavior:
[0108] The server uses a speech recognition service to convert the user's voice commands into text, which is then formatted into an appropriate question and sent to the generative AI model.
[0109] Input: Audio data
[0110] Output: Text data "What are the nutrients in this apple?"
[0111] Step 3: Generate answers using generative AI models
[0112] The generative AI model generates an answer based on the question sent by the server, which contains the most relevant information for the user's question.
[0113] Specific behavior:
[0114] The generative AI model receives text data such as "Tell me the nutrients in this apple," and based on that generates an answer containing information about the apple's nutritional composition.
[0115] Input: Question text "What are the nutrients in this apple?"
[0116] Output: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0117] Step 4: Server receives and sends response
[0118] The server receives the answer sent by the generative AI model, checks the accuracy of the answer, and then sends the answer to the user's device.
[0119] Specific behavior:
[0120] The server receives the answer from the generative AI model, checks its content, adjusts the format, and sends it to the user's smart glasses.
[0121] Input: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0122] Output: Formatted response data
[0123] Step 5: Displaying information via terminal
[0124] The terminal (smart glasses) visually presents the answers sent from the server to the user, who can then check the information directly on the display.
[0125] Specific behavior:
[0126] The answer appears on the smart glasses' display, and the user can review the information, such as, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0127] Input: Formatted response data
[0128] Output: Information displayed to the user
[0129] The above is the specific processing flow of the program of this system, which allows users to easily obtain information on the nutritional content of foods and the appropriate form for exercise.
[0130] (Application example 1)
[0131] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0132] In the past, guidance for children to make healthy choices was limited, making it difficult to provide comprehensive nutritional information and exercise suggestions for food. In addition, there was a lack of easy access to specific health advice on food and exercise in real time. This made it difficult for children to make healthy choices on a daily basis.
[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0134] In this invention, the server includes a nutritional component acquisition means for acquiring information on the nutritional components of the food the child is looking at from a food database, a first answer generation means for generating a guide for a question about food using a data generation model, a first display means for displaying a guide including the nutritional components of the food corresponding to the question on a specific screen, an exercise information acquisition means for acquiring exercise information based on the calories of the food from an exercise database, a second answer generation means for suggesting an exercise based on the calorie information and displaying the guide, a second display means for displaying a guide including the nutritional component information of the food and the exercise suggestion on the specific screen, and a health advice generation means for providing specific health advice based on a generative AI model. This enables children to make healthy choices on a daily basis by being provided with nutritional component information of foods and appropriate exercise suggestions in real time.
[0135] A "food database" is a database that stores information about the nutritional components, effects, and precautions of foods.
[0136] The "nutritional component acquisition means" is a mechanism for acquiring nutritional component information on a specific food from a food database.
[0137] A "data generative model" is an algorithm or system that generates guides or advice based on given data.
[0138] The "first answer generation means" is a mechanism that uses a data generation model to generate guides for questions about food.
[0139] A "specific screen" is a screen displayed on a mobile device such as smart glasses or a smartphone.
[0140] The "first display means" is a mechanism for displaying a guide including the nutritional information of the food corresponding to the question on a specific screen.
[0141] The "exercise database" is a database that stores information on various types of exercise and calorie consumption data.
[0142] The "exercise information acquisition means" is a mechanism for acquiring information about a specific exercise from the exercise database.
[0143] The "second answer generating means" is a mechanism for suggesting exercises based on calorie information and generating a guide for the exercises.
[0144] The "second display means" is a mechanism for displaying a guide including nutritional information on foods and exercise suggestions on a specific screen.
[0145] A "generative AI model" is an artificial intelligence model that provides specific health advice in response to user questions.
[0146] A "health advice generation means" is a mechanism that generates specific health advice using a generative AI model.
[0147] The system for implementing this invention consists of a food database, an exercise database, a generative AI model, a server, terminals such as smart glasses and smartphones, and a user.
[0148] First, when a user browses a food menu using smart glasses or a smartphone, the camera or QR code (registered trademark) reader built into the smart glasses or smartphone acquires food information. This information is sent to a server, which then acquires nutritional information for the target food from a food database. This information is acquired by a nutritional information acquisition means and stored in the server.
[0149] The server then uses the data generation model to generate a specific guide based on the acquired nutritional information, including the nutritional information of the foods displayed on a particular screen (user's smart glasses or smartphone). The first answer generation means is responsible for this process.
[0150] Furthermore, the server accesses the exercise database to acquire appropriate exercise information based on the calorie consumption of the target food. Based on the information acquired by the exercise information acquisition means, the server generates an exercise suggestion using the second answer generation means. The generated exercise suggestion is displayed on a specific screen together with nutritional information of the food. The exercise suggestion also includes an explanation of a specific form and points to note.
[0151] The information is finally provided to the user through the second display means. This system allows the user to check the nutritional information of food and exercise suggestions in real time.
[0152] Additionally, the device also includes a means to generate specific health advice for users using generative AI models, enabling users to receive specific, actionable advice in response to nutrition and exercise questions, helping them make healthier choices.
[0153] As a concrete example, when a user scans "chicken salad" with their smartphone, the server displays its nutritional content (e.g., protein, vitamins, calories). The server also generates exercise suggestions based on the "chicken salad" (e.g., 20 minutes of walking, 10 minutes of stretching) and displays them to the user. Then, when the user asks the generative AI model, "How healthy is this meal?", the server returns specific advice.
[0154] Examples of prompts include:
[0155] "User is trying to order 'Chicken Salad'. Please advise nutritional information and appropriate exercise."
[0156] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0157] Step 1:
[0158] A user scans a food menu with smart glasses or a smartphone.
[0159] Input: Food information obtained by smart glasses or smartphone (e.g., barcode or QR code (registered trademark)).
[0160] Output: Food information data sent to the server.
[0161] Step 2:
[0162] Based on the food information received by the server, the nutritional components of the corresponding food are retrieved from a food database.
[0163] Input: Food information data submitted by the user.
[0164] Output: Nutritional information of foods retrieved from a food database.
[0165] Specific operation: The server compares the food information data with the food database and extracts the corresponding nutritional information.
[0166] Step 3:
[0167] The server uses the data generation model to generate a guide based on the acquired nutritional information.
[0168] Input: Nutritional information for food.
[0169] Output:Nutrition guide information.
[0170] Specific operation: The server inputs nutritional information into a data generation model to generate food guides.
[0171] Step 4:
[0172] The nutritional guide information generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0173] Enter: nutrition guide information.
[0174] Output: Nutrition facts guide displayed on your device.
[0175] Specific operation: The server sends nutrition guide information to the terminal and displays it on a specific screen.
[0176] Step 5:
[0177] The server obtains appropriate exercise suggestion information from an exercise database based on the calorie information of the food.
[0178] Input: Food calorie information.
[0179] Output: Exercise suggestion information obtained from the exercise database.
[0180] Specific operation: The server matches the calorie information with the exercise database and extracts appropriate exercise suggestions.
[0181] Step 6:
[0182] The server uses the second answer generating means to specifically generate the exercise suggestion.
[0183] Input: Exercise suggestion information obtained from the exercise database.
[0184] Output: A specific exercise suggestion guide.
[0185] Specific operation: The server generates an exercise suggestion guide based on the exercise suggestion information.
[0186] Step 7:
[0187] The exercise suggestion guide generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0188] Enter: a concrete exercise suggestion guide.
[0189] Output: Exercise suggestion guide displayed on the device.
[0190] Specific operation: The server sends the exercise suggestion guide to the terminal and displays it on a specific screen.
[0191] Step 8:
[0192] Users can feed additional questions to the generative AI model, for example specific questions about diet or exercise.
[0193] Input: The user's question.
[0194] Output: The question data to be sent to the server.
[0195] Specific operation: A user uses a terminal to input a question, which is then sent to the server.
[0196] Step 9:
[0197] A generative AI model generates specific health advice based on user questions.
[0198] Input: Question data from the user.
[0199] Output: The generated health advice.
[0200] Specific operation: The server inputs the question data into a generative AI model to generate specific health advice.
[0201] Step 10:
[0202] The health advice generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0203] Input: The generated health advice.
[0204] Output: Health advice displayed on the terminal.
[0205] Specific operation: The server sends health advice to the terminal and displays it on a specific screen.
[0206] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0207] The embodiment for implementing the present disclosure includes the following elements.
[0208] 1. Server: The server has the ability to access the food database and exercise database. The server has the ability to receive information from the user and generate specific questions for the generative AI. The server combines an emotion engine to recognize the user's emotional state and customize appropriate advice and information.
[0209] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying answers received from the server and customized information.
[0210] 3. User: The user uses the system via the server or a terminal when viewing food or exercising. The user provides information to the emotion engine through facial expressions and voice that indicate the emotional state.
[0211] In the above-mentioned form, the user can receive individually customized advice and information through a system that combines an emotion engine. Specifically, the emotion engine that recognizes the user's emotional state analyzes information such as facial expressions and voice to estimate the user's emotional state. Then, advice and information regarding food and exercise are customized according to the user's emotional state, and are sent from the server to the terminal. The terminal then displays the received answers and customized information on the smart glasses or screen.
[0212] In this manner, the user can receive individual support tailored to their emotional state, enabling more effective health management.
[0213] The process flow will be explained below.
[0214] Step 1: The user sends information about whether to view food or exercise to the server.
[0215] Step 2: The server utilizes an emotion engine that recognizes the user's emotions and estimates the user's emotional state by analyzing information such as facial expressions and voice.
[0216] Step 3: The server obtains the nutritional information of the corresponding food from the food database based on the user's emotional state.
[0217] Step 4: The server customizes the appropriate advice and information to match the user's emotional state. For example, if the user is feeling stressed, it will suggest foods and exercises that will help them relax.
[0218] Step 5: The server uses generative AI to generate answers to the user's specific questions, including advice and information tailored to the user's emotional state.
[0219] Step 6: The server sends the generated response to the terminal.
[0220] Step 7: The device displays the received answers on the smart glasses or screen, providing advice and information customized to the user's emotional state.
[0221] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0222] 1: Information on whether the user is viewing food or exercising
[0223] 2: Estimation results of user’s emotional state
[0224] 3: Nutritional information of food obtained from food database
[0225] 4: Tailor advice and information to your emotional state
[0226] 5: Answers generated by generative AI
[0227] 6: Answer sent from the server to the device
[0228] 7: Answer displayed on the terminal
[0229] Example 2
[0230] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0231] Conventional health management systems provide uniform advice without considering the user's emotional state, making it difficult to provide appropriate support according to the situation and feelings of each individual user. In particular, when the user's emotional state affects the effectiveness of health management, there is a problem in that the effect cannot be maximized.
[0232] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0233] In this invention, the server includes a nutritional component acquisition means for acquiring information from a food database, an exercise information acquisition means for acquiring information from an exercise database, an emotion analysis means for analyzing the user's emotion state, an answer generation means for generating a guide using a generative AI model, and a customization means for customizing the guide based on the user's emotion state. This makes it possible to provide individual health advice according to the user's emotion state.
[0234] The "nutritional component acquisition means" is a means for acquiring information on the nutritional components of the food the user is viewing from a food database.
[0235] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0236] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data and estimating the user's emotional state.
[0237] An "answer generation means" is a means for generating guides for food and exercise questions posed by a user using a generative AI model.
[0238] A "customization method" is a method for individually adjusting and modifying the generated guide based on the user's emotional state.
[0239] A "display means" is a means for visually presenting a customized guide to a user.
[0240] The system for implementing this invention includes three main elements: a server, a terminal, and a user. Each element works in conjunction with each other to provide individual health advice and support to the user.
[0241] The server has access to the food database and exercise database. Specifically, the server uses the following software and hardware: SQL server and NoSQL database are used to access the food database and exercise database. OpenCV (registered trademark) and Google (registered trademark) Speech-to-Text API for voice recognition are used for user sentiment analysis. OpenAI (registered trademark) GPT (Generative Pre-trained Transformer) is applied as the generative AI model.
[0242] As a concrete example, consider the case where a user voice-inputs a request such as "I would like to know what meals are recommended after today's exercise." The user makes the request into the microphone of the smart glasses. This voice data is sent to the server and converted into text using the Google Speech-to-Text API. Furthermore, the server uses OpenCV to analyze the user's facial expressions and determine their emotional state, such as "fatigue" or "energy."
[0243] For generative AI models, the prompts generated are:
[0244] "What's the best thing to eat when you're tired after a workout?"
[0245] The answers returned by the generative AI model include specific advice, such as "A protein shake, banana, and oatmeal would be good."
[0246] The server then generates a customized guide based on the results of the sentiment analysis, adding additional information such as "We also recommend a cold sports drink" if the user is particularly tired. This information is sent to the device via HTTP or WebSocket.
[0247] The device displays information to the user using the smart glasses or smartphone display. For example, the smart glasses display might show "recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink." Based on this information, the user can take actual health management actions. Specifically, the user can prepare a protein shake while looking at the smart glasses and consume it with a banana.
[0248] Through the above process, users can receive personalized health advice based on their emotional state, enabling more effective health management.
[0249] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0250] Step 1:
[0251] The user enters a request
[0252] The user voices a request into the microphone of the smart glasses, such as "I want to know what meals you recommend after today's exercise." This voice data is sent from the smart glasses to the server.
[0253] Input: Audio data
[0254] Output: Audio data sent to the server
[0255] Step 2:
[0256] The server converts the voice data into text.
[0257] The server converts the received voice data into text using the Google Speech-to-Text API. At this stage, the voice data is converted into text data such as "I would like to know what meals I should eat after my exercise today."
[0258] Input: Audio data
[0259] Output: Text data
[0260] Step 3:
[0261] The server analyzes the user's emotional state.
[0262] The server uses OpenCV to analyze the facial expression data sent from the smart glasses to determine the user's emotional state. For example, it analyzes the user's facial muscle movements and voice tone to determine the user's emotional state, such as "fatigue" or "vigor."
[0263] Input: facial expression data, voice data
[0264] Output: Sentiment analysis result (e.g. user is tired)
[0265] Step 4:
[0266] The server generates and sends prompts to the generative AI model.
[0267] The server generates a prompt sentence to send to the generative AI model based on the text data and the results of sentiment analysis. For example, it generates a prompt sentence such as "What is the best meal to eat when you are tired after exercise?"
[0268] Input: Text data, emotion analysis results
[0269] Output: Prompt statement
[0270] Step 5:
[0271] Generative AI models return answers
[0272] A generative AI model (e.g., OpenAI's GPT) generates an answer to a prompt, such as "Protein shakes, bananas, and oatmeal would be good."
[0273] Input: Prompt statement
[0274] Output: The generated answer
[0275] Step 6:
[0276] Server customizes answer
[0277] The server customizes the generated answer based on the user's emotional state: for example, if the user is particularly tired, the server may add additional information to the generated answer, such as "I also recommend a cold sports drink."
[0278] Input: Generated answers, sentiment analysis results
[0279] Output: Customized answer
[0280] Step 7:
[0281] The server sends customized information to the device.
[0282] The server sends the customized response to the terminal using HTTP communication or WebSocket.
[0283] Input: Customized Answer
[0284] Output: Customized answer sent to the terminal
[0285] Step 8:
[0286] The device displays the information
[0287] The device (smart glasses) displays the customized information sent from the server on the display. For example, the smart glasses display shows "Recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink."
[0288] Input: Customized Answer
[0289] Output: Visually displayed information
[0290] Step 9:
[0291] Users check and use the information
[0292] The user checks the information displayed on the device and takes actual action based on that information. For example, the user prepares a protein shake while looking at the display on the smart glasses and consumes it with a banana.
[0293] Input: Visually displayed information
[0294] Output: User action (e.g. consume a protein shake and a banana)
[0295] (Application example 2)
[0296] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0297] Conventional food delivery services and health management applications have difficulty responding to individual users' emotional states, and are therefore unable to suggest meals or exercises that are optimal for the user's emotions and conditions. The present invention aims to provide more effective health management and a more satisfying user experience by utilizing emotion recognition technology to enable customized suggestions of foods and exercises according to the user's emotional state.
[0298] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes an emotion recognition means for recognizing the emotional state of the user, a nutritional component acquisition means for acquiring information on the nutritional components of foods corresponding to the user's emotion from a food database, an answer generation means for generating a guide for a question about food using a data generation model, and a display means for displaying a guide including the nutritional components of foods corresponding to the question on a specific screen. This makes it possible to provide individual support tailored to the emotional state of the user.
[0299] An "emotion recognition means" is a device or software for analyzing and estimating a user's emotional state from facial expressions and voice.
[0300] The term "nutritional component acquisition means" refers to a device or software for acquiring nutritional information contained in a specific food from a food database.
[0301] A "data generation model" is a module that uses specific algorithms or generative AI to generate guides and suggestions for various user questions.
[0302] An "answer generation means" is a device or software that uses a data generation model to generate answers to a user's questions about food and exercise.
[0303] "Display means" refers to a screen or device for displaying the generated answers and guidance to the user.
[0304] The "exercise information means" refers to a device or software for acquiring information about exercise from an exercise database.
[0305] "Movement form" refers to the specific movements and postures of a particular movement or exercise.
[0306] "Food efficacy" refers to the specific effect or benefit that a particular food has on health.
[0307] "Food precautions" refer to the points or risks that should be taken into consideration when consuming a particular food.
[0308] A system for implementing the present invention consists of three main components: a server, a terminal, and a user.
[0309] server
[0310] The server has the following functions:
[0311] 1. Emotion recognition means
[0312] Hardware and software: Using a smartphone camera and microphone, OpenCV (registered trademark), Google Cloud Vision API, IBM Watson (registered trademark), etc., the system analyzes and estimates the user's emotional state in real time from their facial expressions and voice.
[0313] Example: When a user speaks into the camera on their smartphone, the server receives the video and audio data and analyzes the emotions using emotion recognition means.
[0314] 2. Means of obtaining nutritional information
[0315] Hardware / Software: Uses SQL to access the food database and retrieve nutritional information for specific foods.
[0316] Example: In response to a request for "avocado," the server retrieves the nutritional information for avocado from a database.
[0317] 3. Data Generation Model
[0318] Hardware / Software: Uses OpenAI ChatGPT(registered trademark) to generate guides and answers to user questions about food and exercise.
[0319] Example: If you ask, "What foods are good to eat when you're feeling stressed?" the generative AI will make suggestions such as "avocado and salmon."
[0320] 4. Answer generation means
[0321] Hardware / Software: Using Python (registered trademark), Flask (registered trademark), etc., the answers obtained from the data generation model are optimized for the user and sent to the terminal. The server may include a generating means for including optimized nutritional information of food based on the emotional state of the user recognized by the emotion engine in a guide for food-related questions from the user.
[0322] Example: The generated proposal content is converted into JSON format and sent to the terminal.
[0323] Terminal
[0324] The terminal has the following functions:
[0325] 1. Display means
[0326] Hardware / Software: Information is displayed on users' smartphones and smart glasses using React Native (registered trademark), ANDROID (registered trademark) SDK, and iOS SDK.
[0327] Example: When a user opens the app, suggestions for foods and exercises that are effective in reducing stress appear on the screen.
[0328] User
[0329] A user uses the system in the following way:
[0330] 1. Providing an emotional state
[0331] Users use a smartphone or smart glasses to provide their emotional state to the system through facial expressions and voice.
[0332] Example: When a user opens the app and says, "I'm very tired today," the emotion recognition means analyzes it and sends it to the server.
[0333] 2. Receiving proposal information
[0334] The user selects ingredients and exercises based on the information received.
[0335] Example: A user orders a recommended "avocado and salmon salad" for delivery and does the suggested stretching exercises while waiting.
[0336] Examples of prompt statements
[0337] "User's emotional state: Stress"
[0338] "Contents of information provided: Menu suggestions using ingredients that are effective in reducing stress, and simple relaxation exercises."
[0339] "Databases used: Food database, exercise database"
[0340] "Method: Analyze the user's emotional state and generate optimized suggestions."
[0341] With the above method, users can receive dietary suggestions and exercise support that are individually customized according to their emotional state, enabling them to manage their health more effectively.
[0342] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0343] Processing flow
[0344] Step 1:
[0345] input:
[0346] The user uses a smartphone or smart glasses to provide the system with their emotional state through facial expressions and voice.
[0347] Specific behavior:
[0348] The user says, "I'm tired today."
[0349] Input data: User's facial expression and voice data.
[0350] Step 2:
[0351] input:
[0352] The server receives the user's facial expression and voice data.
[0353] Specific behavior:
[0354] The server analyzes this data using emotion recognition means (smartphone camera, microphone, OpenCV, Google Cloud Vision API, IBM Watson) to estimate the emotional state.
[0355] Data processing: Analyze facial expressions and voice data to infer emotional states (e.g. "stress" or "fatigue").
[0356] Output data: Emotional state data as the analysis result.
[0357] Step 3:
[0358] input:
[0359] Based on the emotional state data, the server sends a request to a food database.
[0360] Specific behavior:
[0361] The server uses RDS to retrieve the relevant nutritional information from a food database.
[0362] Data calculation: Filtering optimized food information based on emotional state (e.g. "stress").
[0363] Output data: Nutritional information of filtered foods.
[0364] Step 4:
[0365] input:
[0366] The server uses the emotional state data and the obtained nutritional information of the food to initiate a guide generation process based on a data generation model.
[0367] Specific behavior:
[0368] Use OpenAI ChatGPT to generate custom guides and suggestions for user questions.
[0369] Example prompt: "User's emotional state: Stressed"
[0370] Data Computation: Using generative AI models to provide optimal meal suggestions based on your emotional state, e.g., “avocado and salmon salad.”
[0371] Output data: Generated diet and exercise guide content.
[0372] Step 5:
[0373] input:
[0374] The server converts the generated proposal content into JSON format and sends it to the terminal.
[0375] Specific behavior:
[0376] The server uses Python or Flask to convert the suggestions obtained from the data generation model into JSON.
[0377] Data processing: Converting guide content into a sendable data format.
[0378] Output data: Guide content in JSON format.
[0379] Step 6:
[0380] input:
[0381] The terminal receives the guide contents sent from the server and displays them to the user.
[0382] Specific behavior:
[0383] The device uses React Native, Android SDK, or iOS SDK to display the guide content on the user's smartphone or smart glasses.
[0384] Display: Suggestions for diet and exercise that fit the user's emotional state.
[0385] Output: Customized guide content displayed on screen.
[0386] Step 7:
[0387] input:
[0388] The user selects ingredients and exercises based on the information displayed on the terminal.
[0389] Specific behavior:
[0390] The user selects the suggested ingredients and orders delivery.
[0391] The user performs a displayed video of a relaxing stretch.
[0392] User action: Select and execute.
[0393] Output: The food choices and exercise the user actually makes.
[0394] Through the above processing steps, the user can receive customized diet and exercise suggestions according to his / her emotional state, enabling the user to carry out more effective health management.
[0395] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0396] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0397] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0398] [Second embodiment]
[0399] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0400] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0401] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0402] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0403] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0404] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0405] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0406] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0407] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0408] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0409] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0410] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0411] The embodiment for implementing the present disclosure includes the following elements.
[0412] 1. Server: The server has the function of accessing the food database and the exercise database. The server has the function of receiving information from the user and generating specific questions for the generative AI. The server has the function of receiving the answers generated by the generative AI and sending them to the terminal.
[0413] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying the answer received from the server.
[0414] 3. User: The user uses the system via the server or terminal when viewing food or performing exercise.
[0415] In this way, users can understand the nutritional information of foods and the appropriate intake amount, and learn about the effects of exercise and the correct form. Specifically, when a user looks at a food, the nutritional information is displayed on the screen of the smart glasses or device. In addition, when the user exercises, exercise suggestions and form tips are displayed. Furthermore, when the user asks the generative AI a question, specific advice and information are obtained.
[0416] In the above manner, users can utilize a system that supports healthy choices.
[0417] The process flow will be explained below.
[0418] Step 1: The user sends information about whether to view food or exercise to the server.
[0419] Step 2: When the user views a food, the server obtains the nutritional information of the corresponding food from the food database.
[0420] Step 3: When the user exercises, the server obtains the corresponding exercise information from the exercise database.
[0421] Step 4: Based on the acquired information, the server generates specific questions for the generative AI.
[0422] Step 5: The generative AI analyzes the received question and generates an appropriate answer.
[0423] Step 6: The server receives the generated answer from the generative AI.
[0424] Step 7: The server sends the received response to the terminal.
[0425] Step 8: The device displays the received answer on the smart glasses or screen.
[0426] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0427] 1: Information on whether the user is viewing food or exercising
[0428] 2: Nutritional information of food obtained from food database
[0429] 3: Exercise information obtained from the exercise database
[0430] 4: Specific questions for generative AI
[0431] 5: Answers generated by generative AI
[0432] 6: Answer sent from the server to the device
[0433] 7: Answer displayed on the terminal
[0434] Example 1
[0435] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0436] Conventional health management systems make it difficult for users to easily obtain nutritional information about foods and appropriate forms for exercise. They also make it difficult to provide quick and accurate answers to specific questions from users. For this reason, there is a demand for a system that allows users to quickly and easily obtain information to make healthy choices.
[0437] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0438] In this invention, the server includes an information processing means for generating questions for the generative AI model based on information provided by the user, an answer receiving means for receiving the answer generated from the generative AI model and transmitting it to the user's terminal, and a display means for displaying information corresponding to the question on a specific terminal, thereby enabling the user to easily obtain nutritional information on foods and appropriate exercise form, and to obtain quick and accurate answers to specific questions.
[0439] An "information processing means" is a means for generating questions for a generative AI model based on information provided by a user.
[0440] The "answer receiving means" is a means for receiving an answer generated from the generative AI model and transmitting it to the user's terminal.
[0441] The "display means" is a means for displaying information corresponding to a question on a specific terminal.
[0442] The "nutritional component acquisition means" is a means for acquiring information about the nutritional components of the food the user is viewing from the food database.
[0443] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0444] The "answer generation means" is a means for generating guides for the user's questions about exercise and food using a data generation model.
[0445] A "voice recognition means" is a means for converting voice commands from a user into text.
[0446] A "terminal" is a device used by a user, such as smart glasses, that has the ability to display information and play audio and video.
[0447] A "generative AI model" is a model that uses pre-trained data to generate appropriate answers to user questions.
[0448] This invention is a system that allows users to easily obtain nutritional information about foods and appropriate exercise forms, and to receive quick and accurate answers to specific questions. This system is mainly composed of three elements: a server, a terminal, and a user.
[0449] Server Roles
[0450] The server uses a high-performance cloud server (for example, a cloud computing service). The server has the following functions:
[0451] Information processing means: The server creates appropriate questions for the generative AI model based on information provided by the user, such as voice commands. It uses voice recognition software (e.g., a voice recognition service) to convert voice to text.
[0452] Answer receiving means: Receives the answer generated by the generative AI model and sends it to the user's device. During this process, the accuracy of the answer is checked and the format is adjusted if necessary.
[0453] Database access function: Accesses food database and exercise database and obtains necessary information from each database.
[0454] Specifically, the server uses voice recognition software to convert the user's voice command into text, and then sends the text, such as "What are the nutrients in this apple?", as a question to the generative AI model.
[0455] Terminal Roles
[0456] The terminal is a device used by a user, and in this invention we assume it is a smart glass. This terminal has the following functions:
[0457] Display means: Displays the information sent from the server. Has the function of visually presenting information on the smart glasses display.
[0458] User interface: Provides an interface for voice command input and touch operation.
[0459] Specifically, when a user picks up an apple and says to the smart glasses, "Tell me the nutritional value of this apple," the apple's calories and nutritional components will be displayed on the smart glasses' display.
[0460] User Roles
[0461] When users view foods or exercise, information is obtained via the server or device, which supports the user in making healthy choices.
[0462] The specific action involves a user picking up an apple and inputting a voice command to the smart glasses, such as, "Tell me the nutritional value of this apple." The answer generated by the generative AI model is, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C," which is then displayed on the smart glasses' display.
[0463] Examples of prompt statements
[0464] Examples of prompts include:
[0465] "How much vitamin C does this apple contain?"
[0466] "How many calories can you burn in 30 minutes of running?"
[0467] As a result, by using the system of the present invention, a user can easily obtain nutritional information about foods and appropriate exercise forms, and make healthy choices based on that information.
[0468] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0469] Divide the program flow into processing steps
[0470] Step 1: User Input
[0471] Step 2: Server processes information and generates a query
[0472] Step 3: Generate answers using a generative AI model
[0473] Step 4: Server receives and sends response
[0474] Step 5: Displaying information via terminal
[0475] Detailed explanation of each processing step
[0476] Step 1: User Input
[0477] When a user uses the system, they input the necessary information into the system through voice commands, touch operations, etc. This allows the system to collect information about foods and exercise that interest the user.
[0478] Specific behavior:
[0479] The user picks up an apple and asks the smart glasses verbally, "What are the nutrients in this apple?"
[0480] Input: Voice command "What are the nutritional values of this apple?"
[0481] Output: Audio data
[0482] Step 2: Server processes information and generates a query
[0483] The server uses speech recognition software to convert the voice data into text, which is then used to generate appropriate questions for the generative AI model.
[0484] Specific behavior:
[0485] The server uses a speech recognition service to convert the user's voice commands into text, which is then formatted into an appropriate question and sent to the generative AI model.
[0486] Input: Audio data
[0487] Output: Text data "What are the nutrients in this apple?"
[0488] Step 3: Generate answers using a generative AI model
[0489] The generative AI model generates an answer based on the question sent by the server, which contains the most relevant information for the user's question.
[0490] Specific behavior:
[0491] The generative AI model receives text data such as "Tell me the nutritional value of this apple," and generates an answer containing information about the apple's nutritional composition based on that data.
[0492] Input: Question text "What are the nutrients in this apple?"
[0493] Output: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0494] Step 4: Server receives and sends response
[0495] The server receives the answer sent by the generative AI model, checks the accuracy of the answer, and then sends the answer to the user's device.
[0496] Specific behavior:
[0497] The server receives the answer from the generative AI model, checks its content, adjusts the format, and sends it to the user's smart glasses.
[0498] Input: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0499] Output: Formatted response data
[0500] Step 5: Displaying information via terminal
[0501] The terminal (smart glasses) visually presents the answers sent from the server to the user, who can then check the information directly on the display.
[0502] Specific behavior:
[0503] The answer appears on the smart glasses' display, and the user can review the information, such as, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0504] Input: Formatted response data
[0505] Output: Information displayed to the user
[0506] The above is the specific processing flow of the program of this system, which allows users to easily obtain information on the nutritional content of foods and the appropriate form for exercise.
[0507] (Application example 1)
[0508] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0509] In the past, guidance for children to make healthy choices was limited, making it difficult to provide comprehensive nutritional information and exercise suggestions for food. In addition, there was a lack of easy access to specific health advice on food and exercise in real time. This made it difficult for children to make healthy choices on a daily basis.
[0510] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0511] In this invention, the server includes a nutritional component acquisition means for acquiring information on the nutritional components of the food the child is looking at from a food database, a first answer generation means for generating a guide for a question about food using a data generation model, a first display means for displaying a guide including the nutritional components of the food corresponding to the question on a specific screen, an exercise information acquisition means for acquiring exercise information based on the calories of the food from an exercise database, a second answer generation means for suggesting an exercise based on the calorie information and displaying the guide, a second display means for displaying a guide including the nutritional component information of the food and the exercise suggestion on the specific screen, and a health advice generation means for providing specific health advice based on a generative AI model. This enables children to make healthy choices on a daily basis by being provided with nutritional component information of foods and appropriate exercise suggestions in real time.
[0512] A "food database" is a database that stores information about the nutritional components, effects, and precautions of foods.
[0513] The "nutritional component acquisition means" is a mechanism for acquiring nutritional component information on a specific food from a food database.
[0514] A "data generative model" is an algorithm or system that generates guides or advice based on given data.
[0515] The "first answer generation means" is a mechanism that uses a data generation model to generate guides for questions about food.
[0516] A "specific screen" is a screen displayed on a mobile device such as smart glasses or a smartphone.
[0517] The "first display means" is a mechanism for displaying a guide including the nutritional information of the food corresponding to the question on a specific screen.
[0518] The "exercise database" is a database that stores information on various types of exercise and calorie consumption data.
[0519] The "exercise information acquisition means" is a mechanism for acquiring information about a specific exercise from the exercise database.
[0520] The "second answer generating means" is a mechanism for suggesting exercises based on calorie information and generating a guide for the exercises.
[0521] The "second display means" is a mechanism for displaying a guide including nutritional information on foods and exercise suggestions on a specific screen.
[0522] A "generative AI model" is an artificial intelligence model that provides specific health advice in response to user questions.
[0523] A "health advice generation means" is a mechanism that generates specific health advice using a generative AI model.
[0524] The system for implementing this invention consists of a food database, an exercise database, a generative AI model, a server, terminals such as smart glasses and smartphones, and a user.
[0525] First, when a user browses a food menu using smart glasses or a smartphone, the camera or QR code reader built into the smart glasses or smartphone acquires food information. This information is sent to a server, which then acquires nutritional information for the target food from a food database. This information is acquired by a nutritional information acquisition means and stored on the server.
[0526] The server then uses the data generation model to generate a specific guide based on the acquired nutritional information, including the nutritional information of the foods displayed on a particular screen (user's smart glasses or smartphone). The first answer generation means is responsible for this process.
[0527] Furthermore, the server accesses the exercise database to acquire appropriate exercise information based on the calorie consumption of the target food. Based on the information acquired by the exercise information acquisition means, the server generates an exercise suggestion using the second answer generation means. The generated exercise suggestion is displayed on a specific screen together with nutritional information of the food. The exercise suggestion also includes an explanation of a specific form and points to note.
[0528] The information is finally provided to the user through the second display means. This system allows the user to check the nutritional information of food and exercise suggestions in real time.
[0529] Additionally, the device also includes a means to generate specific health advice for users using generative AI models, enabling users to receive specific, actionable advice in response to nutrition and exercise questions, helping them make healthier choices.
[0530] As a concrete example, when a user scans "chicken salad" with their smartphone, the server displays its nutritional content (e.g., protein, vitamins, calories). The server also generates exercise suggestions based on the "chicken salad" (e.g., 20 minutes of walking, 10 minutes of stretching) and displays them to the user. Then, when the user asks the generative AI model, "How healthy is this meal?", the server returns specific advice.
[0531] Examples of prompts include:
[0532] "User is trying to order 'Chicken Salad'. Please advise nutritional information and appropriate exercise."
[0533] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0534] Step 1:
[0535] A user scans a food menu with smart glasses or a smartphone.
[0536] Input: Food information obtained via smart glasses or smartphone (e.g., barcode or QR code).
[0537] Output: Food information data sent to the server.
[0538] Step 2:
[0539] Based on the food information received by the server, the nutritional components of the corresponding food are retrieved from a food database.
[0540] Input: Food information data submitted by the user.
[0541] Output: Nutritional information of foods retrieved from a food database.
[0542] Specific operation: The server compares the food information data with the food database and extracts the corresponding nutritional information.
[0543] Step 3:
[0544] The server uses the data generation model to generate a guide based on the acquired nutritional information.
[0545] Input: Nutritional information for food.
[0546] Output:Nutrition guide information.
[0547] Specific operation: The server inputs nutritional information into a data generation model to generate food guides.
[0548] Step 4:
[0549] The nutritional guide information generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0550] Enter: nutrition guide information.
[0551] Output: Nutrition facts guide displayed on your device.
[0552] Specific operation: The server sends nutrition guide information to the terminal and displays it on a specific screen.
[0553] Step 5:
[0554] The server obtains appropriate exercise suggestion information from an exercise database based on the calorie information of the food.
[0555] Input: Food calorie information.
[0556] Output: Exercise suggestion information obtained from the exercise database.
[0557] Specific operation: The server matches the calorie information with the exercise database and extracts appropriate exercise suggestions.
[0558] Step 6:
[0559] The server uses the second answer generating means to specifically generate the exercise suggestion.
[0560] Input: Exercise suggestion information obtained from the exercise database.
[0561] Output: A specific exercise suggestion guide.
[0562] Specific operation: The server generates an exercise suggestion guide based on the exercise suggestion information.
[0563] Step 7:
[0564] The exercise suggestion guide generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0565] Enter: a concrete exercise suggestion guide.
[0566] Output: Exercise suggestion guide displayed on the device.
[0567] Specific operation: The server sends the exercise suggestion guide to the terminal and displays it on a specific screen.
[0568] Step 8:
[0569] Users can feed additional questions to the generative AI model, for example specific questions about diet or exercise.
[0570] Input: The user's question.
[0571] Output: The question data to be sent to the server.
[0572] Specific operation: A user uses a terminal to input a question, which is then sent to the server.
[0573] Step 9:
[0574] A generative AI model generates specific health advice based on user questions.
[0575] Input: Question data from the user.
[0576] Output: The generated health advice.
[0577] Specific operation: The server inputs the question data into a generative AI model to generate specific health advice.
[0578] Step 10:
[0579] The health advice generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0580] Input: The generated health advice.
[0581] Output: Health advice displayed on the terminal.
[0582] Specific operation: The server sends health advice to the terminal and displays it on a specific screen.
[0583] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0584] The embodiment for implementing the present disclosure includes the following elements.
[0585] 1. Server: The server has the ability to access the food database and the exercise database. The server has the ability to receive information from the user and generate specific questions for the generative AI. The server combines an emotion engine to recognize the user's emotional state and customize appropriate advice and information.
[0586] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying answers received from the server and customized information.
[0587] 3. User: The user uses the system via the server or a terminal when viewing food or exercising. The user provides information to the emotion engine through facial expressions and voice that indicate the emotional state.
[0588] In this way, users can receive advice and information that is individually customized through a system that combines an emotion engine. Specifically, the emotion engine that recognizes the user's emotional state analyzes information such as facial expressions and voice, and predicts the user's emotional state.
[0589] Then, advice and information about food and exercise are customized according to the user's emotional state and sent from the server to the device. The device displays the received answers and customized information on the smart glasses or screen.
[0590] In this manner, the user can receive individual support tailored to their emotional state, enabling more effective health management.
[0591] The process flow will be explained below.
[0592] Step 1: The user sends information about whether to view food or exercise to the server.
[0593] Step 2: The server utilizes an emotion engine that recognizes the user's emotions and estimates the user's emotional state by analyzing information such as facial expressions and voice.
[0594] Step 3: The server obtains the nutritional information of the corresponding food from the food database based on the user's emotional state.
[0595] Step 4: The server customizes the appropriate advice and information to match the user's emotional state. For example, if the user is feeling stressed, it will suggest foods and exercises that will help them relax.
[0596] Step 5: The server uses generative AI to generate answers to the user's specific questions, including advice and information tailored to the user's emotional state.
[0597] Step 6: The server sends the generated response to the terminal.
[0598] Step 7: The device displays the received answers on the smart glasses or screen, providing advice and information customized to the user's emotional state.
[0599] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0600] 1: Information on whether the user is viewing food or exercising
[0601] 2: Estimation results of user’s emotional state
[0602] 3: Nutritional information of food obtained from food database
[0603] 4: Tailor advice and information to your emotional state
[0604] 5: Answers generated by generative AI
[0605] 6: Answer sent from the server to the device
[0606] 7: Answer displayed on the terminal
[0607] Example 2
[0608] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0609] Conventional health management systems provide uniform advice without considering the user's emotional state, making it difficult to provide appropriate support according to the situation and feelings of each individual user. In particular, when the user's emotional state affects the effectiveness of health management, there is a problem in that the effect cannot be maximized.
[0610] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0611] In this invention, the server includes a nutritional component acquisition means for acquiring information from a food database, an exercise information acquisition means for acquiring information from an exercise database, an emotion analysis means for analyzing the user's emotion state, an answer generation means for generating a guide using a generative AI model, and a customization means for customizing the guide based on the user's emotion state. This makes it possible to provide individual health advice according to the user's emotion state.
[0612] The "nutritional component acquisition means" is a means for acquiring information on the nutritional components of the food the user is viewing from a food database.
[0613] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0614] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data and estimating the user's emotional state.
[0615] An "answer generation means" is a means for generating guides for food and exercise questions posed by a user using a generative AI model.
[0616] A "customization method" is a method for individually adjusting and modifying the generated guide based on the user's emotional state.
[0617] A "display means" is a means for visually presenting a customized guide to a user.
[0618] The system for implementing this invention includes three main elements: a server, a terminal, and a user. Each element works in conjunction with each other to provide individual health advice and support to the user.
[0619] The server has access to the food database and exercise database. Specifically, the server uses the following software and hardware: A SQL server and NoSQL database are used to access the food database and exercise database. OpenCV and Google Speech-to-Text API for voice recognition are used for user sentiment analysis. OpenAI's GPT (Generative Pre-trained Transformer) is applied as the generative AI model.
[0620] As a concrete example, consider the case where a user voice-inputs a request such as "I would like to know what meals are recommended after today's exercise." The user makes the request into the microphone of the smart glasses. This voice data is sent to the server and converted into text using the Google Speech-to-Text API. Furthermore, the server uses OpenCV to analyze the user's facial expressions and determine their emotional state, such as "fatigue" or "energy."
[0621] For generative AI models, the prompts generated are:
[0622] "What's the best thing to eat when you're tired after a workout?"
[0623] The answers returned by the generative AI model include specific advice, such as "A protein shake, banana, and oatmeal would be good."
[0624] The server then generates a customized guide based on the results of the sentiment analysis, adding additional information such as "We also recommend a cold sports drink" if the user is particularly tired. This information is sent to the device via HTTP or WebSocket.
[0625] The device displays information to the user using the smart glasses or smartphone display. For example, the smart glasses display might show "recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink." Based on this information, the user can take actual health management actions. Specifically, the user can prepare a protein shake while looking at the smart glasses and consume it with a banana.
[0626] Through the above process, users can receive personalized health advice based on their emotional state, enabling more effective health management.
[0627] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0628] Step 1:
[0629] The user enters a request
[0630] The user voices a request into the microphone of the smart glasses, such as "I want to know what meals you recommend after today's exercise." This voice data is sent from the smart glasses to the server.
[0631] Input: Audio data
[0632] Output: Audio data sent to the server
[0633] Step 2:
[0634] The server converts the voice data into text.
[0635] The server converts the received voice data into text using the Google Speech-to-Text API. At this stage, the voice data is converted into text data such as "I would like to know what meals I should eat after my exercise today."
[0636] Input: Audio data
[0637] Output: Text data
[0638] Step 3:
[0639] The server analyzes the user's emotional state.
[0640] The server uses OpenCV to analyze the facial expression data sent from the smart glasses to determine the user's emotional state. For example, it analyzes the user's facial muscle movements and voice tone to determine the user's emotional state, such as "fatigue" or "vigor."
[0641] Input: facial expression data, voice data
[0642] Output: Sentiment analysis result (e.g. user is tired)
[0643] Step 4:
[0644] The server generates and sends prompts to the generative AI model.
[0645] The server generates a prompt sentence to send to the generative AI model based on the text data and the results of sentiment analysis. For example, it generates a prompt sentence such as "What is the best meal to eat when you are tired after exercise?"
[0646] Input: Text data, emotion analysis results
[0647] Output: Prompt statement
[0648] Step 5:
[0649] Generative AI models return answers
[0650] A generative AI model (e.g., OpenAI's GPT) generates an answer to a prompt, such as "Protein shakes, bananas, and oatmeal would be good."
[0651] Input: Prompt statement
[0652] Output: The generated answer
[0653] Step 6:
[0654] Server customizes answer
[0655] The server customizes the generated answer based on the user's emotional state: for example, if the user is particularly tired, the server may add additional information to the generated answer, such as "I also recommend a cold sports drink."
[0656] Input: Generated answers, sentiment analysis results
[0657] Output: Customized answer
[0658] Step 7:
[0659] The server sends customized information to the device.
[0660] The server sends the customized response to the terminal using HTTP communication or WebSocket.
[0661] Input: Customized Answer
[0662] Output: Customized answer sent to the terminal
[0663] Step 8:
[0664] The device displays the information
[0665] The device (smart glasses) displays the customized information sent from the server on the display. For example, the smart glasses display shows "Recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink."
[0666] Input: Customized Answer
[0667] Output: Visually displayed information
[0668] Step 9:
[0669] Users check and use the information
[0670] The user checks the information displayed on the device and takes actual action based on that information. For example, the user prepares a protein shake while looking at the display on the smart glasses and consumes it with a banana.
[0671] Input: Visually displayed information
[0672] Output: User action (e.g. consume a protein shake and a banana)
[0673] (Application example 2)
[0674] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0675] Conventional food delivery services and health management applications have difficulty responding to individual users' emotional states, and are therefore unable to suggest meals or exercises that are optimal for the user's emotions and conditions. The present invention aims to provide more effective health management and a more satisfying user experience by utilizing emotion recognition technology to enable customized suggestions of foods and exercises according to the user's emotional state.
[0676] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes an emotion recognition means for recognizing the emotional state of the user, a nutritional component acquisition means for acquiring information on the nutritional components of foods corresponding to the user's emotion from a food database, an answer generation means for generating a guide for a question about food using a data generation model, and a display means for displaying a guide including the nutritional components of foods corresponding to the question on a specific screen. This makes it possible to provide individual support tailored to the emotional state of the user.
[0677] An "emotion recognition means" is a device or software for analyzing and estimating a user's emotional state from facial expressions and voice.
[0678] The term "nutritional component acquisition means" refers to a device or software for acquiring nutritional information contained in a specific food from a food database.
[0679] A "data generation model" is a module that uses specific algorithms or generative AI to generate guides and suggestions for various user questions.
[0680] An "answer generation means" is a device or software that uses a data generation model to generate answers to a user's questions about food and exercise.
[0681] "Display means" refers to a screen or device for displaying the generated answers and guidance to the user.
[0682] The "exercise information means" refers to a device or software for acquiring information about exercise from an exercise database.
[0683] "Movement form" refers to the specific movements and postures of a particular movement or exercise.
[0684] "Food efficacy" refers to the specific effect or benefit that a particular food has on health.
[0685] "Food precautions" refer to the points or risks that should be taken into consideration when consuming a particular food.
[0686] A system for implementing the present invention consists of three main components: a server, a terminal, and a user.
[0687] server
[0688] The server has the following functions:
[0689] 1. Emotion recognition means
[0690] Hardware and software: Using a smartphone camera and microphone, OpenCV, Google Cloud Vision API, IBM Watson, etc., the system analyzes and estimates the user's emotional state in real time from their facial expressions and voice.
[0691] Example: When a user speaks into the camera on their smartphone, the server receives the video and audio data and analyzes the emotions using emotion recognition means.
[0692] 2. Means of obtaining nutritional information
[0693] Hardware / Software: Uses RDS and SQL to access the food database and retrieve nutritional information for specific foods.
[0694] Example: In response to a request for "avocado," the server retrieves the nutritional information for avocado from a database.
[0695] 3. Data Generation Model
[0696] Hardware / Software: Uses OpenAI ChatGPT to generate guides and answers to user questions about food and exercise.
[0697] Example: If you ask, "What foods are good to eat when you're feeling stressed?" the generative AI will make suggestions such as "avocado and salmon."
[0698] 4. Answer generation means
[0699] Hardware / Software: Using Python, Flask, etc., the answers obtained from the data generation model are optimized for the user and sent to the terminal. The server may include a generating means for including in the guide for the user's food questions the nutritional content of food optimized based on the user's emotional state recognized by the emotion engine.
[0700] Example: The generated proposal content is converted into JSON format and sent to the terminal.
[0701] Terminal
[0702] The terminal has the following functions:
[0703] 1. Display means
[0704] Hardware / Software: Information is displayed on users' smartphones and smart glasses using React Native, Android SDK, and iOS SDK.
[0705] Example: When a user opens the app, suggestions for foods and exercises that are effective in reducing stress appear on the screen.
[0706] User
[0707] A user uses the system in the following way:
[0708] 1. Providing an emotional state
[0709] Users use a smartphone or smart glasses to provide their emotional state to the system through facial expressions and voice.
[0710] Example: When a user opens the app and says, "I'm very tired today," the emotion recognition means analyzes it and sends it to the server.
[0711] 2. Receiving proposal information
[0712] The user selects ingredients and exercises based on the information received.
[0713] Example: A user orders a recommended "avocado and salmon salad" for delivery and does the suggested stretching exercises while waiting.
[0714] Examples of prompt statements
[0715] "User's emotional state: Stress"
[0716] "Contents of information provided: Menu suggestions using ingredients that are effective in reducing stress, and simple relaxation exercises."
[0717] "Databases used: Food database, exercise database"
[0718] "Method: Analyze the user's emotional state and generate optimized suggestions."
[0719] With the above method, users can receive dietary suggestions and exercise support that are individually customized according to their emotional state, enabling them to manage their health more effectively.
[0720] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0721] Processing flow
[0722] Step 1:
[0723] input:
[0724] The user uses a smartphone or smart glasses to provide the system with their emotional state through facial expressions and voice.
[0725] Specific behavior:
[0726] The user says, "I'm tired today."
[0727] Input data: User's facial expression and voice data.
[0728] Step 2:
[0729] input:
[0730] The server receives the user's facial expression and voice data.
[0731] Specific behavior:
[0732] The server analyzes this data using emotion recognition means (smartphone camera, microphone, OpenCV, Google Cloud Vision API, IBM Watson) to estimate the emotional state.
[0733] Data processing: Analyze facial expressions and voice data to infer emotional states (e.g. "stress" or "fatigue").
[0734] Output data: Emotional state data as the analysis result.
[0735] Step 3:
[0736] input:
[0737] Based on the emotional state data, the server sends a request to a food database.
[0738] Specific behavior:
[0739] The server uses RDS to retrieve the relevant nutritional information from a food database.
[0740] Data calculation: Filtering optimized food information based on emotional state (e.g. "stress").
[0741] Output data: Nutritional information of filtered foods.
[0742] Step 4:
[0743] input:
[0744] The server uses the emotional state data and the obtained nutritional information of the food to initiate a guide generation process based on a data generation model.
[0745] Specific behavior:
[0746] Use OpenAI ChatGPT to generate custom guides and suggestions for user questions.
[0747] Example prompt: "User's emotional state: Stressed"
[0748] Data Computation: Using generative AI models to provide optimal meal suggestions based on your emotional state, e.g., “avocado and salmon salad.”
[0749] Output data: Generated diet and exercise guide content.
[0750] Step 5:
[0751] input:
[0752] The server converts the generated proposal content into JSON format and sends it to the terminal.
[0753] Specific behavior:
[0754] The server uses Python or Flask to convert the suggestions obtained from the data generation model into JSON.
[0755] Data processing: Converting guide content into a sendable data format.
[0756] Output data: Guide content in JSON format.
[0757] Step 6:
[0758] input:
[0759] The terminal receives the guide contents sent from the server and displays them to the user.
[0760] Specific behavior:
[0761] The device uses React Native, Android SDK, or iOS SDK to display the guide content on the user's smartphone or smart glasses.
[0762] Display: Suggestions for diet and exercise that fit the user's emotional state.
[0763] Output: Customized guide content displayed on screen.
[0764] Step 7:
[0765] input:
[0766] The user selects ingredients and exercises based on the information displayed on the terminal.
[0767] Specific behavior:
[0768] The user selects the suggested ingredients and orders delivery.
[0769] The user performs a displayed video of a relaxing stretch.
[0770] User action: Select and execute.
[0771] Output: The food choices and exercise the user actually makes.
[0772] Through the above processing steps, the user can receive customized diet and exercise suggestions according to his / her emotional state, enabling the user to carry out more effective health management.
[0773] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0774] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0775] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0776] [Third embodiment]
[0777] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0778] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0779] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0780] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0781] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0782] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0783] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0784] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0785] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0786] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0787] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0788] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".
[0789] The embodiment for implementing the present disclosure includes the following elements.
[0790] 1. Server: The server has the function of accessing the food database and the exercise database. The server has the function of receiving information from the user and generating specific questions for the generative AI. The server has the function of receiving the answers generated by the generative AI and sending them to the terminal.
[0791] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying the answer received from the server.
[0792] 3. User: The user uses the system via the server or terminal when viewing food or performing exercise.
[0793] In this way, users can understand the nutritional information of foods and the appropriate intake amount, and learn about the effects of exercise and the correct form. Specifically, when a user looks at a food, the nutritional information is displayed on the screen of the smart glasses or device. In addition, when the user exercises, exercise suggestions and form tips are displayed. Furthermore, when the user asks the generative AI a question, specific advice and information are obtained.
[0794] In the above manner, users can utilize a system that supports healthy choices.
[0795] The process flow will be explained below.
[0796] Step 1: The user sends information about whether to view food or exercise to the server.
[0797] Step 2: When the user views a food, the server obtains the nutritional information of the corresponding food from the food database.
[0798] Step 3: When the user exercises, the server obtains the corresponding exercise information from the exercise database.
[0799] Step 4: Based on the acquired information, the server generates specific questions for the generative AI.
[0800] Step 5: The generative AI analyzes the received question and generates an appropriate answer.
[0801] Step 6: The server receives the generated answer from the generative AI.
[0802] Step 7: The server sends the received response to the terminal.
[0803] Step 8: The device displays the received answer on the smart glasses or screen.
[0804] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0805] 1: Information on whether the user is viewing food or exercising
[0806] 2: Nutritional information of food obtained from food database
[0807] 3: Exercise information obtained from the exercise database
[0808] 4: Specific questions for generative AI
[0809] 5: Answers generated by generative AI
[0810] 6: Answer sent from the server to the device
[0811] 7: Answer displayed on the terminal
[0812] Example 1
[0813] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0814] Conventional health management systems make it difficult for users to easily obtain nutritional information about foods and appropriate forms for exercise. They also make it difficult to provide quick and accurate answers to specific questions from users. For this reason, there is a demand for a system that allows users to quickly and easily obtain information to make healthy choices.
[0815] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0816] In this invention, the server includes an information processing means for generating questions for the generative AI model based on information provided by the user, an answer receiving means for receiving the answer generated from the generative AI model and transmitting it to the user's terminal, and a display means for displaying information corresponding to the question on a specific terminal, thereby enabling the user to easily obtain nutritional information on foods and appropriate exercise form, and to obtain quick and accurate answers to specific questions.
[0817] An "information processing means" is a means for generating questions for a generative AI model based on information provided by a user.
[0818] The "answer receiving means" is a means for receiving an answer generated from the generative AI model and transmitting it to the user's terminal.
[0819] The "display means" is a means for displaying information corresponding to a question on a specific terminal.
[0820] The "nutritional component acquisition means" is a means for acquiring information about the nutritional components of the food the user is viewing from the food database.
[0821] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0822] The "answer generation means" is a means for generating guides for the user's questions about exercise and food using a data generation model.
[0823] A "voice recognition means" is a means for converting voice commands from a user into text.
[0824] A "terminal" is a device used by a user, such as smart glasses, that has the ability to display information and play audio and video.
[0825] A "generative AI model" is a model that uses pre-trained data to generate appropriate answers to user questions.
[0826] This invention is a system that allows users to easily obtain nutritional information about foods and appropriate exercise forms, and to receive quick and accurate answers to specific questions. This system is mainly composed of three elements: a server, a terminal, and a user.
[0827] Server Roles
[0828] The server uses a high-performance cloud server (for example, a cloud computing service). The server has the following functions:
[0829] Information processing means: The server creates appropriate questions for the generative AI model based on information provided by the user, such as voice commands. It uses voice recognition software (e.g., a voice recognition service) to convert voice to text.
[0830] Answer receiving means: Receives the answer generated by the generative AI model and sends it to the user's device. During this process, the accuracy of the answer is checked and the format is adjusted if necessary.
[0831] Database access function: Accesses food database and exercise database and obtains necessary information from each database.
[0832] Specifically, the server uses voice recognition software to convert the user's voice command into text, and then sends the text, such as "What are the nutrients in this apple?", as a question to the generative AI model.
[0833] Terminal Roles
[0834] The terminal is a device used by a user, and in this invention we assume it is a smart glass. This terminal has the following functions:
[0835] Display means: Displays the information sent from the server. Has the function of visually presenting information on the smart glasses display.
[0836] User interface: Provides an interface for voice command input and touch operation.
[0837] Specifically, when a user picks up an apple and says to the smart glasses, "Tell me the nutritional value of this apple," the apple's calories and nutritional components will be displayed on the smart glasses' display.
[0838] User Roles
[0839] When users view foods or exercise, they receive information via the server or device, which helps them make healthy choices.
[0840] The specific action involves a user picking up an apple and inputting a voice command to the smart glasses, such as, "Tell me the nutritional value of this apple." The answer generated by the generative AI model is, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C," which is then displayed on the smart glasses' display.
[0841] Examples of prompt statements
[0842] Examples of prompts include:
[0843] "How much vitamin C does this apple contain?"
[0844] "How many calories can you burn in 30 minutes of running?"
[0845] As a result, by using the system of the present invention, a user can easily obtain nutritional information about foods and appropriate exercise forms, and make healthy choices based on that information.
[0846] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0847] Divide the program flow into processing steps
[0848] Step 1: User Input
[0849] Step 2: Server processes information and generates a query
[0850] Step 3: Generate answers using generative AI models
[0851] Step 4: Server receives and sends response
[0852] Step 5: Displaying information via terminal
[0853] Detailed explanation of each processing step
[0854] Step 1: User Input
[0855] When a user uses the system, they input the necessary information into the system through voice commands, touch operations, etc. This allows the system to collect information about foods and exercise that interest the user.
[0856] Specific behavior:
[0857] The user picks up an apple and asks the smart glasses verbally, "What are the nutrients in this apple?"
[0858] Input: Voice command "What are the nutritional values of this apple?"
[0859] Output: Audio data
[0860] Step 2: Server processes information and generates a query
[0861] The server uses speech recognition software to convert the voice data into text, which is then used to generate appropriate questions for the generative AI model.
[0862] Specific behavior:
[0863] The server uses a speech recognition service to convert the user's voice commands into text, which is then formatted into an appropriate question and sent to the generative AI model.
[0864] Input: Audio data
[0865] Output: Text data "What are the nutrients in this apple?"
[0866] Step 3: Generate answers using generative AI models
[0867] The generative AI model generates an answer based on the question sent by the server, which contains the most relevant information for the user's question.
[0868] Specific behavior:
[0869] The generative AI model receives text data such as "Tell me the nutritional value of this apple," and generates an answer containing information about the apple's nutritional composition based on that data.
[0870] Input: Question text "What are the nutrients in this apple?"
[0871] Output: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0872] Step 4: Server receives and sends response
[0873] The server receives the answer sent by the generative AI model, checks the accuracy of the answer, and then sends the answer to the user's device.
[0874] Specific behavior:
[0875] The server receives the answer from the generative AI model, checks its content, adjusts the format, and sends it to the user's smart glasses.
[0876] Input: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0877] Output: Formatted response data
[0878] Step 5: Displaying information via terminal
[0879] The terminal (smart glasses) visually presents the answers sent from the server to the user, who can then check the information directly on the display.
[0880] Specific behavior:
[0881] The answer appears on the smart glasses' display, and the user can review the information, such as, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[0882] Input: Formatted response data
[0883] Output: Information displayed to the user
[0884] The above is the specific processing flow of the program of this system, which allows users to easily obtain information on the nutritional content of foods and the appropriate form for exercise.
[0885] (Application example 1)
[0886] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0887] In the past, guidance for children to make healthy choices was limited, making it difficult to provide comprehensive nutritional information and exercise suggestions for food. In addition, there was a lack of easy access to specific health advice on food and exercise in real time. This made it difficult for children to make healthy choices on a daily basis.
[0888] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0889] In this invention, the server includes a nutritional component acquisition means for acquiring information on the nutritional components of the food the child is looking at from a food database, a first answer generation means for generating a guide for a question about food using a data generation model, a first display means for displaying a guide including the nutritional components of the food corresponding to the question on a specific screen, an exercise information acquisition means for acquiring exercise information based on the calories of the food from an exercise database, a second answer generation means for suggesting an exercise based on the calorie information and displaying the guide, a second display means for displaying a guide including the nutritional component information of the food and the exercise suggestion on the specific screen, and a health advice generation means for providing specific health advice based on a generative AI model. This enables children to make healthy choices on a daily basis by being provided with nutritional component information of foods and appropriate exercise suggestions in real time.
[0890] A "food database" is a database that stores information about the nutritional components, effects, and precautions of foods.
[0891] The "nutritional component acquisition means" is a mechanism for acquiring nutritional component information on a specific food from a food database.
[0892] A "data generative model" is an algorithm or system that generates guides or advice based on given data.
[0893] The "first answer generation means" is a mechanism that uses a data generation model to generate guides for questions about food.
[0894] A "specific screen" is a screen displayed on a mobile device such as smart glasses or a smartphone.
[0895] The "first display means" is a mechanism for displaying a guide including the nutritional information of the food corresponding to the question on a specific screen.
[0896] The "exercise database" is a database that stores information on various types of exercise and calorie consumption data.
[0897] The "exercise information acquisition means" is a mechanism for acquiring information about a specific exercise from the exercise database.
[0898] The "second answer generating means" is a mechanism for suggesting exercises based on calorie information and generating a guide for the exercises.
[0899] The "second display means" is a mechanism for displaying a guide including nutritional information on foods and exercise suggestions on a specific screen.
[0900] A "generative AI model" is an artificial intelligence model that provides specific health advice in response to user questions.
[0901] A "health advice generation means" is a mechanism that generates specific health advice using a generative AI model.
[0902] The system for implementing this invention consists of a food database, an exercise database, a generative AI model, a server, terminals such as smart glasses and smartphones, and a user.
[0903] First, when a user browses a food menu using smart glasses or a smartphone, the camera or QR code reader built into the smart glasses or smartphone acquires food information. This information is sent to a server, which then acquires nutritional information for the target food from a food database. This information is acquired by a nutritional information acquisition means and stored on the server.
[0904] The server then uses the data generation model to generate a specific guide based on the acquired nutritional information, including the nutritional information of the foods displayed on a particular screen (user's smart glasses or smartphone). The first answer generation means is responsible for this process.
[0905] Furthermore, the server accesses the exercise database to acquire appropriate exercise information based on the calorie consumption of the target food. Based on the information acquired by the exercise information acquisition means, the server generates an exercise suggestion using the second answer generation means. The generated exercise suggestion is displayed on a specific screen together with nutritional information of the food. The exercise suggestion also includes an explanation of a specific form and points to note.
[0906] The information is finally provided to the user through the second display means. This system allows the user to check the nutritional information of food and exercise suggestions in real time.
[0907] Additionally, the device also includes a means to generate specific health advice for users using generative AI models, enabling users to receive specific, actionable advice in response to nutrition and exercise questions, helping them make healthier choices.
[0908] As a concrete example, when a user scans "chicken salad" with their smartphone, the server displays its nutritional content (e.g., protein, vitamins, calories). The server also generates exercise suggestions based on the "chicken salad" (e.g., 20 minutes of walking, 10 minutes of stretching) and displays them to the user. Then, when the user asks the generative AI model, "How healthy is this meal?", the server returns specific advice.
[0909] Examples of prompts include:
[0910] "User is trying to order 'Chicken Salad'. Please advise nutritional information and appropriate exercise."
[0911] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0912] Step 1:
[0913] A user scans a food menu with smart glasses or a smartphone.
[0914] Input: Food information obtained via smart glasses or smartphone (e.g., barcode or QR code).
[0915] Output: Food information data sent to the server.
[0916] Step 2:
[0917] Based on the food information received by the server, the nutritional components of the corresponding food are retrieved from a food database.
[0918] Input: Food information data submitted by the user.
[0919] Output: Nutritional information of foods retrieved from a food database.
[0920] Specific operation: The server compares the food information data with the food database and extracts the corresponding nutritional information.
[0921] Step 3:
[0922] The server uses the data generation model to generate a guide based on the acquired nutritional information.
[0923] Input: Nutritional information for food.
[0924] Output:Nutrition guide information.
[0925] Specific operation: The server inputs nutritional information into a data generation model to generate food guides.
[0926] Step 4:
[0927] The nutritional guide information generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0928] Enter: nutrition guide information.
[0929] Output: Nutrition facts guide displayed on your device.
[0930] Specific operation: The server sends nutrition guide information to the terminal and displays it on a specific screen.
[0931] Step 5:
[0932] The server obtains appropriate exercise suggestion information from an exercise database based on the calorie information of the food.
[0933] Input: Food calorie information.
[0934] Output: Exercise suggestion information obtained from the exercise database.
[0935] Specific operation: The server matches the calorie information with the exercise database and extracts appropriate exercise suggestions.
[0936] Step 6:
[0937] The server uses the second answer generating means to specifically generate the exercise suggestion.
[0938] Input: Exercise suggestion information obtained from the exercise database.
[0939] Output: A specific exercise suggestion guide.
[0940] Specific operation: The server generates an exercise suggestion guide based on the exercise suggestion information.
[0941] Step 7:
[0942] The exercise suggestion guide generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0943] Enter: a concrete exercise suggestion guide.
[0944] Output: Exercise suggestion guide displayed on the device.
[0945] Specific operation: The server sends the exercise suggestion guide to the terminal and displays it on a specific screen.
[0946] Step 8:
[0947] Users can feed additional questions to the generative AI model, for example specific questions about diet or exercise.
[0948] Input: The user's question.
[0949] Output: The question data to be sent to the server.
[0950] Specific operation: A user uses a terminal to input a question, which is then sent to the server.
[0951] Step 9:
[0952] A generative AI model generates specific health advice based on user questions.
[0953] Input: Question data from the user.
[0954] Output: The generated health advice.
[0955] Specific operation: The server inputs the question data into a generative AI model to generate specific health advice.
[0956] Step 10:
[0957] The health advice generated by the server is displayed on a specific screen (smart glasses or smartphone).
[0958] Input: The generated health advice.
[0959] Output: Health advice displayed on the terminal.
[0960] Specific operation: The server sends health advice to the terminal and displays it on a specific screen.
[0961] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0962] The embodiment for implementing the present disclosure includes the following elements.
[0963] 1. Server: The server has the ability to access the food database and the exercise database. The server has the ability to receive information from the user and generate specific questions for the generative AI. The server combines an emotion engine to recognize the user's emotional state and customize appropriate advice and information.
[0964] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying answers received from the server and customized information.
[0965] 3. User: The user uses the system via the server or a terminal when viewing food or exercising. The user provides information to the emotion engine through facial expressions and voice that indicate the emotional state.
[0966] In the above-mentioned form, the user can receive individually customized advice and information through a system that combines an emotion engine. Specifically, the emotion engine that recognizes the user's emotional state analyzes information such as facial expressions and voice to estimate the user's emotional state. Then, advice and information regarding food and exercise are customized according to the user's emotional state, and are sent from the server to the terminal. The terminal displays the received answers and customized information on the smart glasses or screen.
[0967] In this manner, the user can receive individual support tailored to their emotional state, enabling more effective health management.
[0968] The process flow will be explained below.
[0969] Step 1: The user sends information about whether to view food or exercise to the server.
[0970] Step 2: The server utilizes an emotion engine that recognizes the user's emotions and estimates the user's emotional state by analyzing information such as facial expressions and voice.
[0971] Step 3: The server obtains the nutritional information of the corresponding food from the food database based on the user's emotional state.
[0972] Step 4: The server customizes the appropriate advice and information to match the user's emotional state. For example, if the user is feeling stressed, it will suggest foods and exercises that will help them relax.
[0973] Step 5: The server uses generative AI to generate answers to the user's specific questions, including advice and information tailored to the user's emotional state.
[0974] Step 6: The server sends the generated response to the terminal.
[0975] Step 7: The device displays the received answers on the smart glasses or screen, providing advice and information customized to the user's emotional state.
[0976] Example answers displayed on the screen of the smart glasses or device may include the following information:
[0977] 1: Information on whether the user is viewing food or exercising
[0978] 2: Estimation results of user’s emotional state
[0979] 3: Nutritional information of food obtained from food database
[0980] 4: Tailor advice and information to your emotional state
[0981] 5: Answers generated by generative AI
[0982] 6: Answer sent from the server to the device
[0983] 7: Answer displayed on the terminal
[0984] Example 2
[0985] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0986] Conventional health management systems provide uniform advice without considering the user's emotional state, making it difficult to provide appropriate support according to the situation and feelings of each individual user. In particular, when the user's emotional state affects the effectiveness of health management, there is a problem in that the effect cannot be maximized.
[0987] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0988] In this invention, the server includes a nutritional component acquisition means for acquiring information from a food database, an exercise information acquisition means for acquiring information from an exercise database, an emotion analysis means for analyzing the user's emotion state, an answer generation means for generating a guide using a generative AI model, and a customization means for customizing the guide based on the user's emotion state. This makes it possible to provide individual health advice according to the user's emotion state.
[0989] The "nutritional component acquisition means" is a means for acquiring information on the nutritional components of the food the user is viewing from a food database.
[0990] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[0991] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data and estimating the user's emotional state.
[0992] An "answer generation means" is a means for generating guides for food and exercise questions posed by a user using a generative AI model.
[0993] A "customization method" is a method for individually adjusting and modifying the generated guide based on the user's emotional state.
[0994] A "display means" is a means for visually presenting a customized guide to a user.
[0995] The system for implementing this invention includes three main elements: a server, a terminal, and a user. Each element works in conjunction with each other to provide individual health advice and support to the user.
[0996] The server has access to the food database and exercise database. Specifically, the server uses the following software and hardware: A SQL server and NoSQL database are used to access the food database and exercise database. OpenCV and Google Speech-to-Text API for voice recognition are used for user sentiment analysis. OpenAI's GPT (Generative Pre-trained Transformer) is applied as the generative AI model.
[0997] As a concrete example, consider the case where a user voice-inputs a request such as "I would like to know what meals are recommended after today's exercise." The user makes the request into the microphone of the smart glasses. This voice data is sent to the server and converted into text using the Google Speech-to-Text API. Furthermore, the server uses OpenCV to analyze the user's facial expressions and determine their emotional state, such as "fatigue" or "energy."
[0998] For generative AI models, the prompts generated are:
[0999] "What's the best thing to eat when you're tired after a workout?"
[1000] The answers returned by the generative AI model include specific advice, such as "A protein shake, banana, and oatmeal would be good."
[1001] The server then generates a customized guide based on the results of the sentiment analysis, adding additional information such as "We also recommend a cold sports drink" if the user is particularly tired. This information is sent to the device via HTTP or WebSocket.
[1002] The device displays information to the user using the smart glasses or smartphone display. For example, the smart glasses display might show "recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink." Based on this information, the user can take actual health management actions. Specifically, the user can prepare a protein shake while looking at the smart glasses and consume it with a banana.
[1003] Through the above process, users can receive personalized health advice based on their emotional state, enabling more effective health management.
[1004] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1005] Step 1:
[1006] The user enters a request
[1007] The user voices a request into the microphone of the smart glasses, such as "I want to know what meals you recommend after today's exercise." This voice data is sent from the smart glasses to the server.
[1008] Input: Audio data
[1009] Output: Audio data sent to the server
[1010] Step 2:
[1011] The server converts the voice data into text.
[1012] The server converts the received voice data into text using the Google Speech-to-Text API. At this stage, the voice data is converted into text data such as "I would like to know what meals I should eat after my exercise today."
[1013] Input: Audio data
[1014] Output: Text data
[1015] Step 3:
[1016] The server analyzes the user's emotional state.
[1017] The server uses OpenCV to analyze the facial expression data sent from the smart glasses to determine the user's emotional state. For example, it analyzes the user's facial muscle movements and voice tone to determine the user's emotional state, such as "fatigue" or "vigor."
[1018] Input: facial expression data, voice data
[1019] Output: Sentiment analysis result (e.g. user is tired)
[1020] Step 4:
[1021] The server generates and sends prompts to the generative AI model.
[1022] The server generates a prompt sentence to send to the generative AI model based on the text data and the sentiment analysis results, for example, "What is the best meal to eat when you're tired after exercise?"
[1023] Input: Text data, emotion analysis results
[1024] Output: Prompt statement
[1025] Step 5:
[1026] Generative AI models return answers
[1027] A generative AI model (e.g., OpenAI's GPT) generates an answer to a prompt, such as "Protein shakes, bananas, and oatmeal would be good."
[1028] Input: Prompt statement
[1029] Output: The generated answer
[1030] Step 6:
[1031] Server customizes answer
[1032] The server customizes the generated answer based on the user's emotional state: for example, if the user is particularly tired, the server may add additional information to the generated answer, such as "I also recommend a cold sports drink."
[1033] Input: Generated answers, sentiment analysis results
[1034] Output: Customized answer
[1035] Step 7:
[1036] The server sends customized information to the device.
[1037] The server sends the customized response to the terminal using HTTP communication or WebSocket.
[1038] Input: Customized Answer
[1039] Output: Customized answer sent to the terminal
[1040] Step 8:
[1041] The device displays the information
[1042] The device (smart glasses) displays the customized information sent from the server on the display. For example, the smart glasses display shows "Recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink."
[1043] Input: Customized Answer
[1044] Output: Visually displayed information
[1045] Step 9:
[1046] Users check and use the information
[1047] The user checks the information displayed on the device and takes actual action based on that information. For example, the user prepares a protein shake while looking at the display on the smart glasses and consumes it with a banana.
[1048] Input: Visually displayed information
[1049] Output: User action (e.g. consume a protein shake and a banana)
[1050] (Application example 2)
[1051] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".
[1052] Conventional food delivery services and health management applications have difficulty responding to individual users' emotional states, and are therefore unable to suggest meals or exercises that are optimal for the user's emotions and conditions. The present invention aims to provide more effective health management and a more satisfying user experience by utilizing emotion recognition technology to enable customized suggestions of foods and exercises according to the user's emotional state.
[1053] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes an emotion recognition means for recognizing the emotional state of the user, a nutritional component acquisition means for acquiring information on the nutritional components of foods corresponding to the user's emotion from a food database, an answer generation means for generating a guide for a question about food using a data generation model, and a display means for displaying a guide including the nutritional components of foods corresponding to the question on a specific screen. This makes it possible to provide individual support tailored to the emotional state of the user.
[1054] An "emotion recognition means" is a device or software for analyzing and estimating a user's emotional state from facial expressions and voice.
[1055] The term "nutritional component acquisition means" refers to a device or software for acquiring nutritional information contained in a specific food from a food database.
[1056] A "data generation model" is a module that uses specific algorithms or generative AI to generate guides and suggestions for various user questions.
[1057] An "answer generation means" is a device or software that uses a data generation model to generate answers to a user's questions about food and exercise.
[1058] "Display means" refers to a screen or device for displaying the generated answers and guidance to the user.
[1059] The "exercise information means" refers to a device or software for acquiring information about exercise from an exercise database.
[1060] "Movement form" refers to the specific movements and postures of a particular movement or exercise.
[1061] "Food efficacy" refers to the specific effect or benefit that a particular food has on health.
[1062] "Food precautions" refer to the points or risks that should be taken into consideration when consuming a particular food.
[1063] A system for implementing the present invention consists of three main components: a server, a terminal, and a user.
[1064] server
[1065] The server has the following functions:
[1066] 1. Emotion recognition means
[1067] Hardware and software: Using a smartphone camera and microphone, OpenCV, Google Cloud Vision API, IBM Watson, etc., the system analyzes and estimates the user's emotional state in real time from their facial expressions and voice.
[1068] Example: When a user speaks into the camera on their smartphone, the server receives the video and audio data and analyzes the emotions using emotion recognition means.
[1069] 2. Means of obtaining nutritional information
[1070] Hardware / Software: Uses RDS and SQL to access food database and get nutritional information of specific foods.
[1071] Example: In response to a request for "avocado," the server retrieves the nutritional information for avocado from a database.
[1072] 3. Data Generation Model
[1073] Hardware / Software: Uses OpenAI ChatGPT to generate guides and answers to user questions about food and exercise.
[1074] Example: If you ask, "What foods are good to eat when you're feeling stressed?" the generative AI will make suggestions such as "avocado and salmon."
[1075] 4. Answer generation means
[1076] Hardware / Software: Using Python, Flask, etc., answers obtained from data-generating models are optimized for the user and sent to the device.
[1077] Example: The generated proposal content is converted into JSON format and sent to the terminal.
[1078] Terminal
[1079] The terminal has the following functions:
[1080] 1. Display means
[1081] Hardware / Software: Information is displayed on users' smartphones and smart glasses using React Native, Android SDK, and iOS SDK.
[1082] Example: When a user opens the app, suggestions for foods and exercises that are effective in reducing stress appear on the screen.
[1083] User
[1084] A user uses the system in the following way:
[1085] 1. Providing an emotional state
[1086] Users use a smartphone or smart glasses to provide their emotional state to the system through facial expressions and voice.
[1087] Example: When a user opens the app and says, "I'm very tired today," the emotion recognition means analyzes it and sends it to the server.
[1088] 2. Receiving proposal information
[1089] The user selects ingredients and exercises based on the information received.
[1090] Example: A user orders a recommended "avocado and salmon salad" for delivery and does the suggested stretching exercises while waiting.
[1091] Examples of prompt statements
[1092] "User's emotional state: Stress"
[1093] "Contents of information provided: Menu suggestions using ingredients that are effective in reducing stress, and simple relaxation exercises."
[1094] "Databases used: Food database, exercise database"
[1095] "Method: Analyze the user's emotional state and generate optimized suggestions."
[1096] With the above method, users can receive dietary suggestions and exercise support that are individually customized according to their emotional state, enabling them to manage their health more effectively.
[1097] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1098] Processing flow
[1099] Step 1:
[1100] input:
[1101] The user uses a smartphone or smart glasses to provide the system with their emotional state through facial expressions and voice.
[1102] Specific behavior:
[1103] The user says, "I'm tired today."
[1104] Input data: User's facial expression and voice data.
[1105] Step 2:
[1106] input:
[1107] The server receives the user's facial expression and voice data.
[1108] Specific behavior:
[1109] The server analyzes this data using emotion recognition means (smartphone camera, microphone, OpenCV, Google Cloud Vision API, IBM Watson) to estimate the emotional state.
[1110] Data processing: Analyze facial expressions and voice data to infer emotional states (e.g. "stress" or "fatigue").
[1111] Output data: Emotional state data as the analysis result.
[1112] Step 3:
[1113] input:
[1114] Based on the emotional state data, the server sends a request to a food database.
[1115] Specific behavior:
[1116] The server uses RDS to retrieve the relevant nutritional information from a food database.
[1117] Data calculation: Filtering optimized food information based on emotional state (e.g. "stress").
[1118] Output data: Nutritional information of filtered foods.
[1119] Step 4:
[1120] input:
[1121] The server uses the emotional state data and the obtained nutritional information of the food to initiate a guide generation process based on a data generation model.
[1122] Specific behavior:
[1123] Use OpenAI ChatGPT to generate custom guides and suggestions for user questions.
[1124] Example prompt: "User's emotional state: Stressed"
[1125] Data Computation: Using generative AI models to provide optimal meal suggestions based on your emotional state, e.g., “avocado and salmon salad.”
[1126] Output data: Generated diet and exercise guide content.
[1127] Step 5:
[1128] input:
[1129] The server converts the generated proposal content into JSON format and sends it to the terminal.
[1130] Specific behavior:
[1131] The server uses Python or Flask to convert the suggestions obtained from the data generation model into JSON.
[1132] Data processing: Converting guide content into a sendable data format.
[1133] Output data: Guide content in JSON format.
[1134] Step 6:
[1135] input:
[1136] The terminal receives the guide contents sent from the server and displays them to the user.
[1137] Specific behavior:
[1138] The device uses React Native, Android SDK, or iOS SDK to display the guide content on the user's smartphone or smart glasses.
[1139] Display: Suggestions for diet and exercise that fit the user's emotional state.
[1140] Output: Customized guide content displayed on screen.
[1141] Step 7:
[1142] input:
[1143] The user selects ingredients and exercises based on the information displayed on the terminal.
[1144] Specific behavior:
[1145] The user selects the suggested ingredients and orders delivery.
[1146] The user performs a displayed video of a relaxing stretch.
[1147] User action: Select and execute.
[1148] Output: The food choices and exercise the user actually makes.
[1149] Through the above processing steps, the user can receive customized diet and exercise suggestions according to his / her emotional state, enabling the user to carry out more effective health management.
[1150] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1151] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1152] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1153] [Fourth embodiment]
[1154] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1155] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1156] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[1157] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1158] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[1159] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[1160] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[1161] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1162] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1163] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1164] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1165] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1166] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1167] The embodiment for implementing the present disclosure includes the following elements.
[1168] 1. Server: The server has the function of accessing the food database and the exercise database. The server has the function of receiving information from the user and generating specific questions for the generative AI. The server has the function of receiving the answers generated by the generative AI and sending them to the terminal.
[1169] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying the answer received from the server.
[1170] 3. User: The user uses the system via the server or terminal when viewing food or performing exercise.
[1171] In this way, users can understand the nutritional information of foods and the appropriate intake amount, and learn about the effects of exercise and the correct form. Specifically, when a user looks at a food, the nutritional information is displayed on the screen of the smart glasses or device. In addition, when the user exercises, exercise suggestions and form tips are displayed. Furthermore, when the user asks the generative AI a question, specific advice and information are obtained.
[1172] In the above manner, users can utilize a system that supports healthy choices.
[1173] The process flow will be explained below.
[1174] Step 1: The user sends information about whether to view food or exercise to the server.
[1175] Step 2: When the user views a food, the server obtains the nutritional information of the corresponding food from the food database.
[1176] Step 3: When the user exercises, the server obtains the corresponding exercise information from the exercise database.
[1177] Step 4: Based on the acquired information, the server generates specific questions for the generative AI.
[1178] Step 5: The generative AI analyzes the received question and generates an appropriate answer.
[1179] Step 6: The server receives the generated answer from the generative AI.
[1180] Step 7: The server sends the received response to the terminal.
[1181] Step 8: The device displays the received answer on the smart glasses or screen.
[1182] Example answers displayed on the screen of the smart glasses or device may include the following information:
[1183] 1: Information on whether the user is viewing food or exercising
[1184] 2: Nutritional information of food obtained from food database
[1185] 3: Exercise information obtained from the exercise database
[1186] 4: Specific questions for generative AI
[1187] 5: Answers generated by generative AI
[1188] 6: Answer sent from the server to the device
[1189] 7: Answer displayed on the terminal
[1190] Example 1
[1191] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1192] Conventional health management systems make it difficult for users to easily obtain nutritional information about foods and appropriate forms for exercise. They also make it difficult to provide quick and accurate answers to specific questions from users. For this reason, there is a demand for a system that allows users to quickly and easily obtain information to make healthy choices.
[1193] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1194] In this invention, the server includes an information processing means for generating questions for the generative AI model based on information provided by the user, an answer receiving means for receiving the answer generated from the generative AI model and transmitting it to the user's terminal, and a display means for displaying information corresponding to the question on a specific terminal, thereby enabling the user to easily obtain nutritional information on foods and appropriate exercise form, and to obtain quick and accurate answers to specific questions.
[1195] An "information processing means" is a means for generating questions for a generative AI model based on information provided by a user.
[1196] The "answer receiving means" is a means for receiving an answer generated from the generative AI model and transmitting it to the user's terminal.
[1197] The "display means" is a means for displaying information corresponding to a question on a specific terminal.
[1198] The "nutritional component acquisition means" is a means for acquiring information about the nutritional components of the food the user is viewing from the food database.
[1199] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[1200] The "answer generation means" is a means for generating guides for the user's questions about exercise and food using a data generation model.
[1201] A "voice recognition means" is a means for converting voice commands from a user into text.
[1202] A "terminal" is a device used by a user, such as smart glasses, that has the ability to display information and play audio and video.
[1203] A "generative AI model" is a model that uses pre-trained data to generate appropriate answers to user questions.
[1204] This invention is a system that allows users to easily obtain nutritional information about foods and appropriate exercise forms, and to receive quick and accurate answers to specific questions. This system is mainly composed of three elements: a server, a terminal, and a user.
[1205] Server Roles
[1206] The server uses a high-performance cloud server (for example, a cloud computing service). The server has the following functions:
[1207] Information processing means: The server creates appropriate questions for the generative AI model based on information provided by the user, such as voice commands. It uses voice recognition software (e.g., a voice recognition service) to convert voice to text.
[1208] Answer receiving means: Receives the answer generated by the generative AI model and sends it to the user's device. During this process, the accuracy of the answer is checked and the format is adjusted if necessary.
[1209] Database access function: Accesses food database and exercise database and obtains necessary information from each database.
[1210] Specifically, the server uses voice recognition software to convert the user's voice command into text, and then sends the text, such as "What are the nutrients in this apple?", as a question to the generative AI model.
[1211] Terminal Roles
[1212] The terminal is a device used by a user, and in this invention we assume it is a smart glass. This terminal has the following functions:
[1213] Display means: Displays the information sent from the server. Has the function of visually presenting information on the smart glasses display.
[1214] User interface: Provides an interface for voice command input and touch operation.
[1215] Specifically, when a user picks up an apple and says to the smart glasses, "Tell me the nutritional value of this apple," the apple's calories and nutritional components will be displayed on the smart glasses' display.
[1216] User Roles
[1217] When users view foods or exercise, information is obtained via the server or device, which supports the user in making healthy choices.
[1218] The specific action involves a user picking up an apple and inputting a voice command to the smart glasses, such as, "Tell me the nutritional value of this apple." The answer generated by the generative AI model is, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C," which is then displayed on the smart glasses' display.
[1219] Examples of prompt statements
[1220] Examples of prompts include:
[1221] "How much vitamin C does this apple contain?"
[1222] "How many calories can you burn in 30 minutes of running?"
[1223] As a result, by using the system of the present invention, a user can easily obtain nutritional information about foods and appropriate exercise forms, and make healthy choices based on that information.
[1224] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1225] Divide the program flow into processing steps
[1226] Step 1: User Input
[1227] Step 2: Server processes information and generates a query
[1228] Step 3: Generate answers using generative AI models
[1229] Step 4: Server receives and sends response
[1230] Step 5: Displaying information via terminal
[1231] Detailed explanation of each processing step
[1232] Step 1: User Input
[1233] When a user uses the system, they input the necessary information into the system through voice commands, touch operations, etc. This allows the system to collect information about foods and exercise that interest the user.
[1234] Specific behavior:
[1235] The user picks up an apple and asks the smart glasses verbally, "What are the nutrients in this apple?"
[1236] Input: Voice command "What are the nutritional values of this apple?"
[1237] Output: Audio data
[1238] Step 2: Server processes information and generates a query
[1239] The server uses speech recognition software to convert the voice data into text, which is then used to generate appropriate questions for the generative AI model.
[1240] Specific behavior:
[1241] The server uses a speech recognition service to convert the user's voice commands into text, which is then formatted into an appropriate question and sent to the generative AI model.
[1242] Input: Audio data
[1243] Output: Text data "What are the nutrients in this apple?"
[1244] Step 3: Generate answers using a generative AI model
[1245] The generative AI model generates an answer based on the question sent by the server, which contains the most relevant information for the user's question.
[1246] Specific behavior:
[1247] The generative AI model receives text data such as "Tell me the nutritional value of this apple," and generates an answer containing information about the apple's nutritional composition based on that data.
[1248] Input: Question text "What are the nutrients in this apple?"
[1249] Output: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[1250] Step 4: Server receives and sends response
[1251] The server receives the answer sent by the generative AI model, checks the accuracy of the answer, and then sends the answer to the user's device.
[1252] Specific behavior:
[1253] The server receives the answer from the generative AI model, checks its content, adjusts the format, and sends it to the user's smart glasses.
[1254] Input: Answer text "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[1255] Output: Formatted response data
[1256] Step 5: Displaying information via terminal
[1257] The terminal (smart glasses) visually presents the answers sent from the server to the user, who can then check the information directly on the display.
[1258] Specific behavior:
[1259] The answer appears on the smart glasses' display, and the user can review the information, such as, "This apple contains approximately 52 calories, 0.3g fat, 14g carbohydrates, and 5.4g vitamin C."
[1260] Input: Formatted response data
[1261] Output: Information displayed to the user
[1262] The above is the specific processing flow of the program of this system, which allows users to easily obtain information on the nutritional content of foods and the appropriate form for exercise.
[1263] (Application example 1)
[1264] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1265] In the past, guidance for children to make healthy choices was limited, making it difficult to provide comprehensive nutritional information and exercise suggestions for food. In addition, there was a lack of easy access to specific health advice on food and exercise in real time. This made it difficult for children to make healthy choices on a daily basis.
[1266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1267] In this invention, the server includes a nutritional component acquisition means for acquiring information on the nutritional components of the food the child is looking at from a food database, a first answer generation means for generating a guide for a question about food using a data generation model, a first display means for displaying a guide including the nutritional components of the food corresponding to the question on a specific screen, an exercise information acquisition means for acquiring exercise information based on the calories of the food from an exercise database, a second answer generation means for suggesting an exercise based on the calorie information and displaying the guide, a second display means for displaying a guide including the nutritional component information of the food and the exercise suggestion on the specific screen, and a health advice generation means for providing specific health advice based on a generative AI model. This enables children to make healthy choices on a daily basis by being provided with nutritional component information of foods and appropriate exercise suggestions in real time.
[1268] A "food database" is a database that stores information about the nutritional components, effects, and precautions of foods.
[1269] The "nutritional component acquisition means" is a mechanism for acquiring nutritional component information on a specific food from a food database.
[1270] A "data generative model" is an algorithm or system that generates guides or advice based on given data.
[1271] The "first answer generation means" is a mechanism that uses a data generation model to generate guides for questions about food.
[1272] A "specific screen" is a screen displayed on a mobile device such as smart glasses or a smartphone.
[1273] The "first display means" is a mechanism for displaying a guide including the nutritional information of the food corresponding to the question on a specific screen.
[1274] The "exercise database" is a database that stores information on various types of exercise and calorie consumption data.
[1275] The "exercise information acquisition means" is a mechanism for acquiring information about a specific exercise from the exercise database.
[1276] The "second answer generating means" is a mechanism for suggesting exercises based on calorie information and generating a guide for the exercises.
[1277] The "second display means" is a mechanism for displaying a guide including nutritional information on foods and exercise suggestions on a specific screen.
[1278] A "generative AI model" is an artificial intelligence model that provides specific health advice in response to user questions.
[1279] A "health advice generation means" is a mechanism that generates specific health advice using a generative AI model.
[1280] The system for implementing this invention consists of a food database, an exercise database, a generative AI model, a server, terminals such as smart glasses and smartphones, and a user.
[1281] First, when a user browses a food menu using smart glasses or a smartphone, the camera or QR code reader built into the smart glasses or smartphone acquires food information. This information is sent to a server, which then acquires nutritional information for the target food from a food database. This information is acquired by a nutritional information acquisition means and stored on the server.
[1282] The server then uses the data generation model to generate a specific guide based on the acquired nutritional information, including the nutritional information of the foods displayed on a particular screen (user's smart glasses or smartphone). The first answer generation means is responsible for this process.
[1283] Furthermore, the server accesses the exercise database to acquire appropriate exercise information based on the calorie consumption of the target food. Based on the information acquired by the exercise information acquisition means, the server generates an exercise suggestion using the second answer generation means. The generated exercise suggestion is displayed on a specific screen together with nutritional information of the food. The exercise suggestion also includes an explanation of a specific form and points to note.
[1284] The information is finally provided to the user through the second display means. This system allows the user to check the nutritional information of food and exercise suggestions in real time.
[1285] Additionally, the device also includes a means to generate specific health advice for users using generative AI models, enabling users to receive specific, actionable advice in response to nutrition and exercise questions, helping them make healthier choices.
[1286] As a concrete example, when a user scans "chicken salad" with their smartphone, the server displays its nutritional content (e.g., protein, vitamins, calories). The server also generates exercise suggestions based on the "chicken salad" (e.g., 20 minutes of walking, 10 minutes of stretching) and displays them to the user. Then, when the user asks the generative AI model, "How healthy is this meal?", the server returns specific advice.
[1287] Examples of prompts include:
[1288] "User is trying to order 'Chicken Salad'. Please advise nutritional information and appropriate exercise."
[1289] The flow of the specific process in the application example 1 will be described with reference to FIG.
[1290] Step 1:
[1291] A user scans a food menu with smart glasses or a smartphone.
[1292] Input: Food information obtained via smart glasses or smartphone (e.g., barcode or QR code).
[1293] Output: Food information data sent to the server.
[1294] Step 2:
[1295] Based on the food information received by the server, the nutritional components of the corresponding food are retrieved from a food database.
[1296] Input: Food information data submitted by the user.
[1297] Output: Nutritional information of foods retrieved from a food database.
[1298] Specific operation: The server compares the food information data with the food database and extracts the corresponding nutritional information.
[1299] Step 3:
[1300] The server uses the data generation model to generate a guide based on the acquired nutritional information.
[1301] Input: Nutritional information for food.
[1302] Output:Nutrition guide information.
[1303] Specific operation: The server inputs nutritional information into a data generation model to generate food guides.
[1304] Step 4:
[1305] The nutritional guide information generated by the server is displayed on a specific screen (smart glasses or smartphone).
[1306] Enter: nutrition guide information.
[1307] Output: Nutrition facts guide displayed on your device.
[1308] Specific operation: The server sends nutrition guide information to the terminal and displays it on a specific screen.
[1309] Step 5:
[1310] The server obtains appropriate exercise suggestion information from an exercise database based on the calorie information of the food.
[1311] Input: Food calorie information.
[1312] Output: Exercise suggestion information obtained from the exercise database.
[1313] Specific operation: The server matches the calorie information with the exercise database and extracts appropriate exercise suggestions.
[1314] Step 6:
[1315] The server uses the second answer generating means to specifically generate the exercise suggestion.
[1316] Input: Exercise suggestion information obtained from the exercise database.
[1317] Output: A specific exercise suggestion guide.
[1318] Specific operation: The server generates an exercise suggestion guide based on the exercise suggestion information.
[1319] Step 7:
[1320] The exercise suggestion guide generated by the server is displayed on a specific screen (smart glasses or smartphone).
[1321] Enter: a concrete exercise suggestion guide.
[1322] Output: Exercise suggestion guide displayed on the device.
[1323] Specific operation: The server sends the exercise suggestion guide to the terminal and displays it on a specific screen.
[1324] Step 8:
[1325] Users can feed additional questions to the generative AI model, for example specific questions about diet or exercise.
[1326] Input: The user's question.
[1327] Output: The question data to be sent to the server.
[1328] Specific operation: A user uses a terminal to input a question, which is then sent to the server.
[1329] Step 9:
[1330] A generative AI model generates specific health advice based on user questions.
[1331] Input: Question data from the user.
[1332] Output: The generated health advice.
[1333] Specific operation: The server inputs the question data into a generative AI model to generate specific health advice.
[1334] Step 10:
[1335] The health advice generated by the server is displayed on a specific screen (smart glasses or smartphone).
[1336] Input: The generated health advice.
[1337] Output: Health advice displayed on the terminal.
[1338] Specific operation: The server sends health advice to the terminal and displays it on a specific screen.
[1339] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1340] The embodiment for implementing the present disclosure includes the following elements.
[1341] 1. Server: The server has the ability to access the food database and the exercise database. The server has the ability to receive information from the user and generate specific questions for the generative AI. The server combines an emotion engine to recognize the user's emotional state and customize appropriate advice and information.
[1342] 2. Terminal: The terminal has the function of displaying information through smart glasses or a screen. The terminal has the function of displaying answers received from the server and customized information.
[1343] 3. User: The user uses the system via the server or a terminal when viewing food or exercising. The user provides information to the emotion engine through facial expressions and voice that indicate the emotional state.
[1344] In the above-mentioned form, the user can receive individually customized advice and information through a system that combines an emotion engine. Specifically, the emotion engine that recognizes the user's emotional state analyzes information such as facial expressions and voice to estimate the user's emotional state. Then, advice and information regarding food and exercise are customized according to the user's emotional state, and are sent from the server to the terminal. The terminal displays the received answers and customized information on the smart glasses or screen.
[1345] In this manner, the user can receive individual support tailored to their emotional state, enabling more effective health management.
[1346] The process flow will be explained below.
[1347] Step 1: The user sends information about whether to view food or exercise to the server.
[1348] Step 2: The server utilizes an emotion engine that recognizes the user's emotions and estimates the user's emotional state by analyzing information such as facial expressions and voice.
[1349] Step 3: The server obtains the nutritional information of the corresponding food from the food database based on the user's emotional state.
[1350] Step 4: The server customizes the appropriate advice and information to match the user's emotional state. For example, if the user is feeling stressed, it will suggest foods and exercises that will help them relax.
[1351] Step 5: The server uses generative AI to generate answers to the user's specific questions, including advice and information tailored to the user's emotional state.
[1352] Step 6: The server sends the generated response to the terminal.
[1353] Step 7: The device displays the received answers on the smart glasses or screen, providing advice and information customized to the user's emotional state.
[1354] Example answers displayed on the screen of the smart glasses or device may include the following information:
[1355] 1: Information on whether the user is viewing food or exercising
[1356] 2: Estimation results of user’s emotional state
[1357] 3: Nutritional information of food obtained from food database
[1358] 4: Tailor advice and information to your emotional state
[1359] 5: Answers generated by generative AI
[1360] 6: Answer sent from the server to the device
[1361] 7: Answer displayed on the terminal
[1362] Example 2
[1363] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1364] Conventional health management systems provide uniform advice without considering the user's emotional state, making it difficult to provide appropriate support according to the situation and feelings of each individual user. In particular, when the user's emotional state affects the effectiveness of health management, there is a problem in that the effect cannot be maximized.
[1365] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1366] In this invention, the server includes a nutritional component acquisition means for acquiring information from a food database, an exercise information acquisition means for acquiring information from an exercise database, an emotion analysis means for analyzing the user's emotion state, an answer generation means for generating a guide using a generative AI model, and a customization means for customizing the guide based on the user's emotion state. This makes it possible to provide individual health advice according to the user's emotion state.
[1367] The "nutritional component acquisition means" is a means for acquiring information on the nutritional components of the food the user is viewing from a food database.
[1368] The "exercise information means" is a means for acquiring information about exercise from an exercise database.
[1369] The "emotion analysis means" is a means for analyzing the user's facial expressions and voice data and estimating the user's emotional state.
[1370] An "answer generation means" is a means for generating guides for food and exercise questions posed by a user using a generative AI model.
[1371] A "customization method" is a method for individually adjusting and modifying the generated guide based on the user's emotional state.
[1372] A "display means" is a means for visually presenting a customized guide to a user.
[1373] The system for implementing this invention includes three main elements: a server, a terminal, and a user. Each element works in conjunction with each other to provide individual health advice and support to the user.
[1374] The server has access to the food database and exercise database. Specifically, the server uses the following software and hardware: A SQL server and NoSQL database are used to access the food database and exercise database. OpenCV and Google Speech-to-Text API for voice recognition are used for user sentiment analysis. OpenAI's GPT (Generative Pre-trained Transformer) is applied as the generative AI model.
[1375] As a concrete example, consider the case where a user voice-inputs a request such as "I would like to know what meals are recommended after today's exercise." The user makes the request into the microphone of the smart glasses. This voice data is sent to the server and converted into text using the Google Speech-to-Text API. Furthermore, the server uses OpenCV to analyze the user's facial expressions and determine their emotional state, such as "fatigue" or "energy."
[1376] For generative AI models, the prompts generated look like this:
[1377] "What's the best thing to eat when you're tired after a workout?"
[1378] The answers returned by the generative AI model include specific advice, such as "A protein shake, banana, and oatmeal would be good."
[1379] The server then generates a customized guide based on the results of the sentiment analysis, adding additional information such as "We also recommend a cold sports drink" if the user is particularly tired. This information is sent to the device via HTTP or WebSocket.
[1380] The device displays information to the user using the smart glasses or smartphone display. For example, the smart glasses display might show "recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink." Based on this information, the user can take actual health management actions. Specifically, the user can prepare a protein shake while looking at the smart glasses and consume it with a banana.
[1381] Through the above process, users can receive personalized health advice based on their emotional state, enabling more effective health management.
[1382] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1383] Step 1:
[1384] The user enters a request
[1385] The user voices a request into the microphone of the smart glasses, such as "I want to know what meals you recommend after today's exercise." This voice data is sent from the smart glasses to the server.
[1386] Input: Audio data
[1387] Output: Audio data sent to the server
[1388] Step 2:
[1389] The server converts the voice data into text.
[1390] The server converts the received voice data into text using the Google Speech-to-Text API. At this stage, the voice data is converted into text data such as "I would like to know what meals I should eat after my exercise today."
[1391] Input: Audio data
[1392] Output: Text data
[1393] Step 3:
[1394] The server analyzes the user's emotional state.
[1395] The server uses OpenCV to analyze the facial expression data sent from the smart glasses to determine the user's emotional state. For example, it analyzes the user's facial muscle movements and voice tone to determine the user's emotional state, such as "fatigue" or "vigor."
[1396] Input: facial expression data, voice data
[1397] Output: Sentiment analysis result (e.g. user is tired)
[1398] Step 4:
[1399] The server generates and sends prompts to the generative AI model.
[1400] The server generates a prompt sentence to send to the generative AI model based on the text data and the sentiment analysis results, for example, "What is the best meal to eat when you're tired after exercise?"
[1401] Input: Text data, emotion analysis results
[1402] Output: Prompt statement
[1403] Step 5:
[1404] Generative AI models return answers
[1405] A generative AI model (e.g., OpenAI's GPT) generates an answer to a prompt, such as "Protein shakes, bananas, and oatmeal would be good."
[1406] Input: Prompt statement
[1407] Output: The generated answer
[1408] Step 6:
[1409] Server customizes answer
[1410] The server customizes the generated answer based on the user's emotional state: for example, if the user is particularly tired, the server may add additional information to the generated answer, such as "I also recommend a cold sports drink."
[1411] Input: Generated answers, sentiment analysis results
[1412] Output: Customized answer
[1413] Step 7:
[1414] The server sends customized information to the device.
[1415] The server sends the customized response to the terminal using HTTP communication or WebSocket.
[1416] Input: Customized Answer
[1417] Output: Customized answer sent to the terminal
[1418] Step 8:
[1419] The device displays the information
[1420] The device (smart glasses) displays the customized information sent from the server on the display. For example, the smart glasses display shows "Recommended meals after exercise: protein shake, banana, oatmeal, cold sports drink."
[1421] Input: Customized Answer
[1422] Output: Visually displayed information
[1423] Step 9:
[1424] Users check and use the information
[1425] The user checks the information displayed on the device and takes actual action based on that information. For example, the user prepares a protein shake while looking at the display on the smart glasses and consumes it with a banana.
[1426] Input: Visually displayed information
[1427] Output: User action (e.g. consume a protein shake and a banana)
[1428] (Application example 2)
[1429] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1430] Conventional food delivery services and health management applications have difficulty responding to individual users' emotional states, and are therefore unable to suggest meals or exercises that are optimal for the user's emotions and conditions. The present invention aims to provide more effective health management and a more satisfying user experience by utilizing emotion recognition technology to enable customized suggestions of foods and exercises according to the user's emotional state.
[1431] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes an emotion recognition means for recognizing the emotional state of the user, a nutritional component acquisition means for acquiring information on the nutritional components of foods corresponding to the user's emotion from a food database, an answer generation means for generating a guide for a question about food using a data generation model, and a display means for displaying a guide including the nutritional components of foods corresponding to the question on a specific screen. This makes it possible to provide individual support tailored to the emotional state of the user.
[1432] An "emotion recognition means" is a device or software for analyzing and estimating a user's emotional state from facial expressions and voice.
[1433] The term "nutritional component acquisition means" refers to a device or software for acquiring nutritional information contained in a specific food from a food database.
[1434] A "data generation model" is a module that uses specific algorithms or generative AI to generate guides and suggestions for various user questions.
[1435] An "answer generation means" is a device or software that uses a data generation model to generate answers to a user's questions about food and exercise.
[1436] "Display means" refers to a screen or device for displaying the generated answers and guidance to the user.
[1437] The "exercise information means" refers to a device or software for acquiring information about exercise from an exercise database.
[1438] "Movement form" refers to the specific movements and postures of a particular movement or exercise.
[1439] "Food efficacy" refers to the specific effect or benefit that a particular food has on health.
[1440] "Food precautions" refer to the points or risks that should be taken into consideration when consuming a particular food.
[1441] A system for implementing the present invention consists of three main components: a server, a terminal, and a user.
[1442] server
[1443] The server has the following functions:
[1444] 1. Emotion recognition means
[1445] Hardware and software: Using a smartphone camera and microphone, OpenCV, Google Cloud Vision API, IBM Watson, etc., the system analyzes and estimates the user's emotional state in real time from their facial expressions and voice.
[1446] Example: When a user speaks into the camera on their smartphone, the server receives the video and audio data and analyzes the emotions using emotion recognition means.
[1447] 2. Means of obtaining nutritional information
[1448] Hardware / Software: Uses RDS and SQL to access the food database and retrieve nutritional information for specific foods.
[1449] Example: In response to a request for "avocado," the server retrieves the nutritional information for avocado from a database.
[1450] 3. Data Generation Model
[1451] Hardware / Software: Uses OpenAI ChatGPT to generate guides and answers to user questions about food and exercise.
[1452] Example: If you ask, "What foods are good to eat when you're feeling stressed?" the generative AI will make suggestions such as "avocado and salmon."
[1453] 4. Answer generation means
[1454] Hardware / Software: Using Python, Flask, etc., answers obtained from data-generating models are optimized for the user and sent to the device.
[1455] Example: The generated proposal content is converted into JSON format and sent to the terminal.
[1456] Terminal
[1457] The terminal has the following functions:
[1458] 1. Display means
[1459] Hardware / Software: Information is displayed on users' smartphones and smart glasses using React Native, Android SDK, and iOS SDK.
[1460] Example: When a user opens the app, suggestions for foods and exercises that are effective in reducing stress appear on the screen.
[1461] User
[1462] A user uses the system in the following way:
[1463] 1. Providing an emotional state
[1464] Users use a smartphone or smart glasses to provide their emotional state to the system through facial expressions and voice.
[1465] Example: When a user opens the app and says, "I'm very tired today," the emotion recognition means analyzes it and sends it to the server.
[1466] 2. Receiving proposal information
[1467] The user selects ingredients and exercises based on the information received.
[1468] Example: A user orders a recommended "avocado and salmon salad" for delivery and does the suggested stretching exercises while waiting.
[1469] Examples of prompt statements
[1470] "User's emotional state: Stress"
[1471] "Contents of information provided: Menu suggestions using ingredients that are effective in reducing stress, and simple relaxation exercises."
[1472] "Databases used: Food database, exercise database"
[1473] "Method: Analyze the user's emotional state and generate optimized suggestions."
[1474] With the above method, users can receive dietary suggestions and exercise support that are individually customized according to their emotional state, enabling them to manage their health more effectively.
[1475] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1476] Processing flow
[1477] Step 1:
[1478] input:
[1479] The user uses a smartphone or smart glasses to provide the system with their emotional state through facial expressions and voice.
[1480] Specific behavior:
[1481] The user says, "I'm tired today."
[1482] Input data: User's facial expression and voice data.
[1483] Step 2:
[1484] input:
[1485] The server receives the user's facial expression and voice data.
[1486] Specific behavior:
[1487] The server analyzes this data using emotion recognition means (smartphone camera, microphone, OpenCV, Google Cloud Vision API, IBM Watson) to estimate the emotional state.
[1488] Data processing: Analyze facial expressions and voice data to infer emotional states (e.g. "stress" or "fatigue").
[1489] Output data: Emotional state data as the analysis result.
[1490] Step 3:
[1491] input:
[1492] Based on the emotional state data, the server sends a request to a food database.
[1493] Specific behavior:
[1494] The server uses RDS to retrieve the relevant nutritional information from a food database.
[1495] Data calculation: Filtering optimized food information based on emotional state (e.g. "stress").
[1496] Output data: Nutritional information of filtered foods.
[1497] Step 4:
[1498] input:
[1499] The server uses the emotional state data and the obtained nutritional information of the food to initiate a guide generation process based on a data generation model.
[1500] Specific behavior:
[1501] Use OpenAI ChatGPT to generate custom guides and suggestions for user questions.
[1502] Example prompt: "User's emotional state: Stressed"
[1503] Data Computation: Using generative AI models to provide optimal meal suggestions based on your emotional state, e.g., “avocado and salmon salad.”
[1504] Output data: Generated diet and exercise guide content.
[1505] Step 5:
[1506] input:
[1507] The server converts the generated proposal content into JSON format and sends it to the terminal.
[1508] Specific behavior:
[1509] The server uses Python or Flask to convert the suggestions obtained from the data generation model into JSON.
[1510] Data processing: Converting guide content into a sendable data format.
[1511] Output data: Guide content in JSON format.
[1512] Step 6:
[1513] input:
[1514] The terminal receives the guide contents sent from the server and displays them to the user.
[1515] Specific behavior:
[1516] The device uses React Native, Android SDK, or iOS SDK to display the guide content on the user's smartphone or smart glasses.
[1517] Display: Suggestions for diet and exercise that fit the user's emotional state.
[1518] Output: Customized guide content displayed on screen.
[1519] Step 7:
[1520] input:
[1521] The user selects ingredients and exercises based on the information displayed on the terminal.
[1522] Specific behavior:
[1523] The user selects the suggested ingredients and orders delivery.
[1524] The user performs a displayed video of a relaxing stretch.
[1525] User action: Select and execute.
[1526] Output: The food choices and exercise the user actually makes.
[1527] Through the above processing steps, the user can receive customized diet and exercise suggestions according to his / her emotional state, enabling the user to carry out more effective health management.
[1528] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1529] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1530] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.
[1531] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.
[1532] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1533] These emotions are distributed in the 3 o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1534] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1535] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.
[1536] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."
[1537] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.
[1538] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).
[1539] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.
[1540] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1541] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.
[1542] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1543] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.
[1544] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[1545] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1546] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.
[1547] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.
[1548] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.
[1549] The following is further disclosed regarding the above embodiment.
[1550] (Appendix 1) A system that acts as a guide to help children make healthy choices, comprising: a nutritional information obtaining means for obtaining information on the nutritional information of the food viewed by the child from a food database; A first answer generation means for generating a guide for a question about food using a data generation model; a first display means for displaying a guide including nutritional components of the food corresponding to the question on a specific screen; A system including:
[1551] (Appendix 2) exercise information means for acquiring information regarding exercise from an exercise database; A second answer generation means for generating a guide for a question about a child's exercise using the data generation model; A second display means for displaying a guide including an exercise form corresponding to the question on the specific screen; 2. The system of claim 1, comprising:
[1552] (Appendix 3) The system of claim 1, wherein the first answer generation means further provides guidance on at least one of the efficacy and precautions of the food. (Appendix 4) The system of any one of claims 1 to 3, further comprising an emotion engine for recognizing the child's emotions, the emotion engine recognizing the child's emotions, and a customization means for customizing the guide based on the recognized emotions.
[1553] "Example 1"
[1554] (Claim 1) An information processing means for generating questions for a generative AI model based on information provided by a user; An answer receiving means for receiving an answer generated from the generative AI model and transmitting the answer to a user's terminal; a display means for displaying information corresponding to the question on a specific terminal; A system including:
[1555] (Claim 2) A nutritional component acquiring means for acquiring information on nutritional components of the food viewed by the user from a food database; exercise information means for acquiring information regarding exercise from an exercise database; An answer generation means for generating a guide for a user's exercise-related question using the data generation model; a display means for displaying a guide including an exercise form corresponding to the question on the specific terminal; 2. The system of claim 1, comprising:
[1556] (Claim 3) 2. The system of claim 1, wherein the information processing means includes speech recognition means for converting voice commands from a user into text.
[1557] "Application example 1" (Claim 1) a nutritional information obtaining means for obtaining information on the nutritional information of the food viewed by the child from a food database; A first answer generation means for generating a guide for a question about food using a data generation model; a first display means for displaying a guide including nutritional components of the food corresponding to the question on a specific screen; an exercise information acquiring means for acquiring exercise information based on food calories from an exercise database; A second answer generating means for suggesting an exercise based on the calorie information and displaying a guide therefor; A second display means for displaying a guide including nutritional information of food and exercise suggestions on the specific screen; A health advice generating means for providing specific health advice based on a generative AI model; A system including:
[1558] (Claim 2) The system according to claim 1 , wherein the first answer generating means further provides guidance on at least one of efficacy and precautions of food.
[1559] (Claim 3) The system of claim 1, wherein the specific screen displays real-time health advice using a mobile device such as smart glasses or a smartphone.
[1560] "Example 2 of combining emotion engines"
[1561] (Claim 1) A nutritional component acquiring means for acquiring information on nutritional components of the food viewed by the user from a food database; exercise information means for acquiring information regarding exercise from an exercise database; An emotion analysis means for analyzing an emotional state of a user; an answer generation means for generating guides for food and exercise questions using a generative AI model; customization means for customizing the guide based on the emotional state of the user; a display means for displaying the guide customized by the customization means; A system including:
[1562] (Claim 2) The system according to claim 1, which provides guidance on at least one of the efficacy and precautions of food.
[1563] (Claim 3) 10. The system of claim 1, wherein the information is customized based on an emotional state of the user.
[1564] "Application example 2 when combining emotion engines"
[1565] (Claim 1) an emotion recognition means for recognizing an emotional state of a user; A nutritional component acquiring means for acquiring information on nutritional components of food corresponding to the user's emotion from a food database; A first answer generation means for generating a guide for a question about food using a data generation model; a first display means for displaying a guide including nutritional components of the food corresponding to the question on a specific screen; A system including:
[1566] (Claim 2) exercise information means for acquiring information regarding exercise from an exercise database; A second answer generation means for generating a guide for a question about exercise according to a user's emotional state using the data generation model; A second display means for displaying a guide including an exercise form corresponding to the question on a specific screen; 2. The system of claim 1, comprising:
[1567] (Claim 3) The system according to claim 1 , wherein the first answer generating means further provides guidance on at least one of efficacy and precautions of food.
[1568] (Claim 1) an emotion engine that recognizes the user's emotions; A nutritional component acquisition means for acquiring information on nutritional components of food from a food database; A suggestion means for suggesting foods suitable for the user's emotion based on the user's emotion and information on nutritional components of the foods by inputting a specific prompt sentence into a generative AI model; A generating means for generating a guide for a food question from the user using a generative AI model; A customization means for customizing a guide for the question based on the emotion of the user; A providing means for providing the user with food suitable for the emotion of the user and the customized guide via a terminal of the user; A system including:
[1569] (Claim 2) The system of claim 1 , wherein the generating means includes in the guide for the question nutritional content of foods optimized based on the emotional state of the user recognized by the emotion engine.
[1570] (Claim 3) exercise information means for acquiring information regarding exercise from an exercise database; A generating means for generating an appropriate form for the exercise requested by the user as a guide for the question by inputting a specific prompt sentence into a generating AI model; Including, The system of claim 1 , wherein the providing means provides the appropriate form for the exercise to the user via a terminal of the user. [Explanation of symbols]
[1571] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. an emotion engine that recognizes the user's emotions; An acquisition means for acquiring information about food viewed by the user on the terminal; A nutritional component acquiring means for acquiring information on nutritional components of the food viewed by the user from a food database; A suggestion means for suggesting foods suitable for the user's emotion based on the user's emotion and information on nutritional components of the foods by inputting a specific prompt sentence into a generative AI model; A generating means for generating a guide for a food question from the user using a generative AI model; A customization means for customizing a guide for the question based on the emotion of the user; A providing means for providing the user with food suitable for the emotion of the user and the customized guide via a terminal of the user; A system including:
2. The system of claim 1 , wherein the generating means includes in the guide for the question nutritional content of foods optimized based on the emotional state of the user recognized by the emotion engine.
3. exercise information means for acquiring information regarding exercise from an exercise database; A generating means for generating an appropriate form for the exercise requested by the user as a guide for the question by inputting a specific prompt sentence into a generating AI model; Including, The system according to claim 1 , wherein the providing means provides the appropriate form for the exercise to the user via a terminal of the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A