System

A voice-based system for generating personalized training menus addresses the challenge of accessing tailored health training for elderly and disabled individuals by using advanced speech recognition and AI to provide effective home-based exercise solutions.

JP2026034169APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137290
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Elderly people and individuals with physical disabilities face challenges in accessing tailored health training information online and often struggle to maintain consistent exercise due to unfamiliarity with the internet and the burden of hospital visits.

Method used

A system that acquires user health conditions through voice input, converts it to text, analyzes the condition using natural language processing, generates a personalized training menu, and provides it with text and video instructions, utilizing advanced speech recognition and generative AI to facilitate easy and effective home-based training.

Benefits of technology

Enables elderly and physically challenged individuals to maintain their health effectively at home by providing customized training menus and instructions, improving their quality of life through easy access and correct exercise execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034169000001_ABST
    Figure 2026034169000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining a voice input; means for converting the obtained voice input into text; means for analyzing a health condition of a user based on the voice input; means for generating an appropriate training menu based on the health condition; and means for providing the generated training menu to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to provide a training menu that elderly people and people with physical disabilities can comfortably complete at home in their spare time. However, many elderly people are unfamiliar with the Internet and find it difficult to obtain appropriate training information online. Furthermore, there is currently no easy way to obtain individual training menus tailored to their health condition. Furthermore, the burden of visiting a hospital itself is significant, making it difficult for many people to exercise consistently to maintain their health. There is a need to solve these problems and improve the health and quality of life (QOL) of elderly people and people with physical disabilities. [Means for solving the problem]

[0005] The present invention is a system that acquires a user's health condition and concerns using voice input and proposes an appropriate training menu based on that information. Specifically, the system includes a means for acquiring voice input, a means for converting the acquired voice input into text, a means for analyzing the user's health condition based on the voice input, a means for generating an appropriate training menu based on the health condition, and a means for providing the generated training menu to the user. This allows the user to easily acquire and implement the optimal training menu by simply inputting voice. The generated training menu also includes text and reference videos, making it visually easy to understand. Furthermore, by including a software program that performs a series of processes, it is possible to efficiently and flexibly provide training tailored to various health conditions. In this way, the present invention is a system that enables elderly people and people with physical challenges to maintain their health easily and effectively at home.

[0006] "Voice input" is the process of capturing a user's spoken voice as digital data.

[0007] "Text" is character information converted from audio data.

[0008] "Health status" refers to the user's physical and mental state, including information about specific symptoms and physical conditions.

[0009] "Analysis" is the process of interpreting information and deriving specific results or conclusions based on acquired data.

[0010] A "training menu" refers to a series of instructions or guides for exercises or exercises suggested based on the user's health condition.

[0011] "Generation" is the process of creating new data or content based on given information.

[0012] "Providing" refers to presenting the generated training menu to the user and making it available for use.

[0013] A "system" is a collection of devices and programs that consist of multiple means and functions that perform tasks ranging from obtaining voice input to providing training menus. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0036] Acquiring voice input

[0037] The user talks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0038] Audio data conversion

[0039] The device inputs the received voice data into a voice recognition engine and converts it into text data, which generates the text "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0040] Health status analysis

[0041] The server receives the text data sent from the device and analyzes the content using natural language processing (NLP) technology. For example, it identifies the user's health condition based on the keywords "stiff shoulders" and "severe."

[0042] Training menu generation

[0043] The server generates an appropriate training menu based on the results of the health analysis. Using the generation AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness. Specific exercise instructions include "stretch both arms up and slowly tilt them to the left and right," along with a link to a reference video demonstrating this exercise.

[0044] Providing training instructions

[0045] The server transmits the generated training menu to the terminal, which includes specific instructions including text and video links.

[0046] Viewing and Running Training

[0047] The device displays the training menu received from the server on the user interface. The user performs the exercises while looking at the displayed training menu. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" displayed on the screen. The user can perform the exercises in the correct way while playing a reference video.

[0048] Follow-up

[0049] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes the data again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0050] In this way, the present invention is a system that allows elderly people and users with physical disabilities to easily and effectively train at home to maintain their health. By providing specific training instructions and reference videos, users can exercise in the correct way, which is expected to contribute to maintaining health and improving quality of life.

[0051] The processing flow will be explained below.

[0052] Step 1:

[0053] The user talks to the system about their physical condition and concerns. For example, "I have severe shoulder stiffness. Please tell me what kind of exercise I should do."

[0054] Step 2:

[0055] The terminal receives the user's voice through a microphone.

[0056] Step 3:

[0057] The device inputs the received voice data into a voice recognition engine and converts it into text data. For example, the generated text is "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0058] Step 4:

[0059] The terminal transmits the converted text data to the server.

[0060] Step 5:

[0061] The server receives the text data sent from the terminal.

[0062] Step 6:

[0063] The server analyzes the text data using natural language processing (NLP) technology, specifically extracting the keywords "stiff shoulders" and "severe" to determine the user's health condition.

[0064] Step 7:

[0065] The server generates an appropriate training menu based on the analysis of the user's health status. Using the AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness.

[0066] Step 8:

[0067] The server then creates a training menu by documenting specific exercises and adding links to reference videos. For example, "Extend both arms overhead and slowly tilt them back and forth three times."

[0068] Step 9:

[0069] The server transmits the generated training menu to the terminal.

[0070] Step 10:

[0071] The terminal displays the training menu received from the server on the user interface.

[0072] Step 11:

[0073] The user performs exercises while viewing the displayed training menu. They can also play reference videos to check the exercise methods.

[0074] Step 12:

[0075] If the user has any questions or concerns during the training, they can input their voice again, and the device will pick up the voice.

[0076] Step 13:

[0077] The terminal converts the newly received voice data back into text data and transmits it to the server.

[0078] Step 14:

[0079] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[0080] Step 15:

[0081] The terminal displays any additional instructions or explanations received to the user.

[0082] Example 1

[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0084] There is a need for a system that allows elderly people and users with physical challenges to easily obtain and implement appropriate training menus at home through voice input. However, existing systems have issues with the accuracy of speech recognition and text analysis, as well as the degree of customization of training menus, making it difficult for users to achieve the results they expect. For this reason, there is a need for a system that utilizes more advanced speech recognition, natural language processing, and generative AI technologies to provide appropriate training menus tailored to the user's health condition.

[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0086] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text data, means for analyzing the user's health condition based on the text data using natural language processing technology, means for generating an appropriate training menu using a generative AI model based on the analysis results of the health condition, means for transmitting the generated training menu to a terminal and providing it to the user, and means for displaying and executing training based on the training menu, thereby enabling the user to easily and effectively perform health maintenance training at home.

[0087] "Means for obtaining voice input" refers to a device or a function of a device that has the function of capturing the voice spoken by a user to the system and recording it as digital voice data.

[0088] A "means for converting speech input into text data" is software or a system capable of analyzing recorded speech data and converting it into a corresponding text string.

[0089] "Natural language processing technology" refers to a series of algorithms and techniques for analyzing text data and understanding its content and intent.

[0090] "Means for analyzing a user's health condition" refers to software or a system that has the function of extracting and analyzing information about a user's physical condition and health from input text data using natural language processing technology.

[0091] A "generative AI model" is a model that uses artificial intelligence technology to generate appropriate responses and suggestions from input data.

[0092] "Means for generating a training menu" refers to software or a system that has the function of automatically creating an appropriate training menu based on the user's health condition analyzed using a generative AI model.

[0093] The "means for transmitting a training menu to a terminal" refers to software or a system that has the function of transferring the generated training menu as data to a terminal and providing it to the user.

[0094] "Means for displaying and executing training" refers to a device or software that has the function of displaying the training menu provided to the terminal on a user interface and allowing the user to perform the training according to the instructions.

[0095] "Link to reference video" refers to a URL or data that provides access to a video that shows specific methods for performing a training menu.

[0096] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0097] Acquiring voice input

[0098] The user talks to the system about their physical condition and concerns. For example, they might say something like, "I have really bad shoulder stiffness. Please tell me what kind of exercise I should do." In order to obtain clear voice data, it is important to speak in an environment that eliminates as much background noise as possible.

[0099] Audio data conversion

[0100] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specific voice recognition engines that are used include Google (registered trademark) Cloud Speech-to-Text API. For example, voice input data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do" is converted into text.

[0101] Health status analysis

[0102] The server receives the text data sent from the device and analyzes the content using natural language processing technology. A natural language processing engine such as AWS (registered trademark) Comprehend is used here. Specifically, keywords such as "stiff shoulders" and "severe" are identified from the text data, and the user's health condition is analyzed.

[0103] Training menu generation

[0104] The server uses a generative AI model based on the analysis results to generate an appropriate training menu. Specifically, it uses a generative AI model such as OpenAI's GPT-3. For example, a prompt such as "Please generate a training menu recommended for elderly people with severe shoulder stiffness" can be input, and the server will suggest "shoulder stretching exercises" based on that. The generated content includes instructions such as "Extend both arms up and slowly tilt them to the left and right" and links to reference videos.

[0105] Providing training instructions

[0106] The server sends the generated training menu to the device, which contains text instructions and links to reference videos.

[0107] Viewing and Running Training

[0108] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of the web or mobile application. The user performs the training by following the "shoulder stretch exercise" instructions displayed on the screen. The user can perform the exercises with the correct form while playing the reference video.

[0109] Follow-up

[0110] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine again and converts it into text data. The server then analyzes the text data and generates additional instructions or explanations. For example, it might generate additional instructions such as, "Please refer to the following video for shoulder stretching exercises," and send the video link to the device. The device then displays the newly received instructions or explanations to the user, who can then continue their training based on them.

[0111] Examples of specific examples and prompts

[0112] Example: A 70-year-old elderly person uses this system for stiff shoulders. He / she speaks to the system, saying, "I have severe shoulder stiffness, so please tell me what kind of exercises I should do." The system then suggests shoulder stretching exercises and provides a link to a video.

[0113] Example prompt sentence:

[0114] "Generate a recommended training menu for elderly people with severe shoulder stiffness."

[0115] "Please provide exercise instructions and a link to a helpful video to help relieve shoulder stiffness."

[0116] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0117] Step 1:

[0118] The user speaks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." The input is the user's voice, and the output is the voice data received by the device.

[0119] Step 2:

[0120] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specifically, it uses the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." Here, the voice recognition engine analyzes the phonemes of the voice and generates a corresponding string of characters.

[0121] Step 3:

[0122] The server receives the text data sent from the device and analyzes it using natural language processing technology. Specifically, AWS Comprehend is used to extract keywords from the text data and identify the health condition. The input is text data, and the output is analysis results such as "stiff shoulders" or "severe." Natural language processing technology extracts keywords and performs sentiment analysis.

[0123] Step 4:

[0124] The server uses a generative AI model to generate an appropriate training menu based on the health status analysis results. Specifically, it uses OpenAI's GPT-3 and inputs the prompt "Please generate a recommended training menu for elderly people with severe shoulder stiffness." The inputs are the analysis results and the prompt, and the output is a training menu such as "Extend both arms up and slowly tilt them left and right" and a link to a reference video. The generative AI model suggests optimal training content based on the user's health status.

[0125] Step 5:

[0126] The server sends the generated training menu to the terminal. The input is the training menu and video link, and the output is the data to be sent to the terminal. The server transfers the text data of the training menu and the video link to the terminal.

[0127] Step 6:

[0128] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of a web or mobile application. The input is the data received from the server, and the output is the display on the interface that the user can visually confirm.

[0129] Step 7:

[0130] The user performs exercises according to the displayed training menu. Specifically, the user performs the exercises by following the instructions displayed on the screen: "Extend both arms upwards and slowly tilt them to the left and right." A reference video is played and the user performs the exercises while checking the movements. The input is the displayed training menu, and the output is the user performing the exercises with the correct form.

[0131] Step 8:

[0132] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine, which converts it back into text data. The input is voice data, and the output is text data.

[0133] Step 9:

[0134] The server analyzes the newly received text data and generates additional instructions or explanations. The input is new text data, and the output is additional instructions or explanations. For example, the server generates an additional instruction such as "Please refer to the following video" and sends it to the device.

[0135] Step 10:

[0136] The terminal displays newly received instructions and explanations to the user. The input is the additional instructions and explanations sent from the server, and the output is what is displayed on the user interface. The user continues training based on that.

[0137] (Application example 1)

[0138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0139] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, or elsewhere. However, existing systems have problems in that users cannot acquire effective training menus in real time or cannot properly execute acquired menus. For example, the accuracy of voice input or voice recognition may be low, resulting in an inappropriate menu not being generated. Furthermore, the generated menu may be difficult for users to understand, resulting in the inability to perform exercises in the correct manner. Furthermore, specialized equipment or devices are required, which may be difficult for elderly people and users with physical challenges to implement. There is a need for a system that solves these problems and allows users to easily and effectively acquire and implement training menus.

[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0141] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition based on the voice input, means for generating an appropriate training menu based on the health condition, means for providing the generated training menu to the user, and means for displaying the generated training menu on a smart device, thereby enabling the user to obtain an effective training menu in real time at home, at a fitness facility, or the like, and to exercise in the correct way via the smart device.

[0142] The "means for acquiring voice input" is a device or method for acquiring the voice uttered by the user as digital data.

[0143] A "means for converting captured speech input to text" is a device or method for converting speech data into natural language text data.

[0144] The "means for analyzing a user's health condition based on a voice input" refers to a device or method for analyzing text data obtained from a voice input and determining a user's health condition.

[0145] The "means for generating an appropriate training menu based on health condition" refers to a device or method that automatically creates an appropriate exercise and stretching menu based on the analyzed health condition.

[0146] The "means for providing the generated training menu to the user" refers to a device or method that determines how to provide the generated training menu to the user and implements that method.

[0147] The "means for displaying the generated training menu on a smart device" refers to a device or method for visually displaying the generated training menu on a device such as a smartphone or smart glasses used by the user.

[0148] A "generative AI model" is a model that uses artificial intelligence techniques to generate new information from data.

[0149] A "prompt" is an input text given to a generative AI model, used to cause the model to generate specific instructions or information.

[0150] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, etc. The main components of this system are a voice input acquisition means, a means for converting the acquired voice input into text, a means for analyzing the user's health condition, a means for generating a training menu, a means for providing the generated training menu to the user, and a means for displaying the menu on a smart device.

[0151] Program processing

[0152] The whole system includes the following steps:

[0153] 1. Voice input acquisition: The user speaks to a smart device (e.g., smart glasses or smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is acquired using a microphone.

[0154] 2. Speech Recognition: The captured voice data is converted into text data using the Google Cloud Speech-to-Text API, which transcribes the user's speech into text.

[0155] 3. Text analysis: The server analyzes the text data using natural language processing (NLP) techniques (e.g., spaCy) to identify the user's health condition. For example, the text "My left knee hurts" is analyzed and it is recognized that the user has a knee problem.

[0156] 4. Training menu generation: The server generates an appropriate training menu using a generative AI model (e.g., OpenAI's GPT) based on the analysis results. An example of this prompt might be, "Please suggest an effective exercise menu for my left knee pain." The generated training menu includes specific exercise content and links to reference videos.

[0157] 5. Providing training menu: The generated training menu is sent to the smart device, where the user can check the instructions on the display of their smart glasses or smartphone.

[0158] 6. Display and execution of training: The generated exercise instructions are displayed on the display of the user device. For example, specific exercise instructions such as "Perform knee support stretches" are displayed, and a reference video can also be viewed.

[0159] Hardware and software used

[0160] Smart devices (e.g. smart glasses, smartphones): Used for voice input and display.

[0161] Cloud server: Used for voice recognition, text analysis, and training menu generation.

[0162] Speech recognition API (e.g., Google Cloud Speech-to-Text API): Used to convert voice data into text.

[0163] Natural language processing libraries (e.g. spaCy): Used to analyze text data.

[0164] Generative AI models (e.g., OpenAI's GPT): Used to generate training menus.

[0165] Specific examples

[0166] For example, if an elderly person says to a trainer at a fitness gym, "My left knee hurts. Please tell me what exercises I should do," the smart glasses will recognize the speech and convert it into text. The server then analyzes the text and determines that the person has pain in their left knee. An appropriate training menu is generated using a generative AI model, and the trainer is shown instructions such as, "Perform stretches to support your knee." The smart glasses also display a video showing how to perform the exercises correctly.

[0167] Prompt Sentence Examples

[0168] "Please suggest an effective exercise routine for my left knee pain."

[0169] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0170] Step 1:

[0171] Acquiring voice input

[0172] The user speaks into the microphone of their smart device (e.g., smart glasses, smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is picked up by the microphone on the smart device.

[0173] Input: User's voice

[0174] Output: Audio data

[0175] Step 2:

[0176] Converting audio data to text

[0177] The device sends the captured voice data to the Google Cloud Speech-to-Text API and converts it into text data. For example, a speech like "My left knee hurts, so please tell me what exercises I should do" is converted into the corresponding text.

[0178] Input: Audio data

[0179] Output: Text data

[0180] Step 3:

[0181] Text data analysis

[0182] The server receives the text data sent from the device and analyzes it using natural language processing (NLP) technology. For example, the keywords "knee" and "pain" may be extracted as analysis results, and the user's health condition may be identified as "pain in the left knee."

[0183] Input: Text data

[0184] Output: Health status analysis results

[0185] Step 4:

[0186] Training menu generation

[0187] Based on the analysis results, the server inputs prompts into a generative AI model (e.g., OpenAI's GPT) to generate an appropriate training menu. For example, a prompt such as "Please suggest an effective exercise menu for left knee pain" generates specific exercise content and links to reference videos.

[0188] Input: Analysis results regarding health status, prompt text

[0189] Output: Training menu

[0190] Step 5:

[0191] Sending and providing training menus

[0192] The server sends the generated training menu to the user's smart device, which then notifies the user of the menu. Specifically, the generated training menu is displayed on the display of smart glasses or a smartphone.

[0193] Input: Training Menu

[0194] Output: Training menu displayed on a smart device

[0195] Step 6:

[0196] Viewing and Running Training

[0197] The user device supports the user in exercising accurately based on the displayed training menu. For example, exercise instructions such as "Perform stretches to support your knees" are displayed on the smart glasses along with a reference video. The user exercises while watching the video.

[0198] Input: Training menu displayed on smart device

[0199] Output: User executes training

[0200] Each processing step allows the user to obtain an appropriate training menu in real time and perform exercises in the correct way through their smart device.

[0201] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0202] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[0203] Acquiring and converting voice input

[0204] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0205] The device receives the user's voice through a microphone, inputs this voice data into a speech recognition engine, and converts it into text data. The converted text data will contain the following content: "I've been suffering from severe shoulder stiffness lately and I'm feeling depressed. What kind of exercise is good?"

[0206] Health and Emotion Analysis

[0207] The terminal transmits the converted text data to the server.

[0208] The server analyzes the received text data using natural language processing (NLP) technology to identify the user's health condition, for example by extracting keywords such as "stiff shoulders" and "feeling depressed."

[0209] Furthermore, the server uses an emotion engine to analyze the user's emotions, for example, to identify the emotion "feeling depressed."

[0210] Creation and provision of training menus

[0211] The server generates an appropriate training menu based on the analysis results. Using the generation AI, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. This includes specific exercise instructions and links to reference videos. For example, instructions might include "extend both arms overhead and slowly tilt them back and forth three times."

[0212] The server selects a training menu based on the user's emotions and transmits it to the terminal. The selected menu is most effective for the user, for example, a menu that enhances relaxation.

[0213] Running the training

[0214] The terminal displays the training menu received from the server on the user interface.

[0215] The user looks at the displayed training menu and performs the exercises. For example, they perform the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. They can perform the exercises in the correct way by playing the reference video.

[0216] Follow-up

[0217] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they can ask, "Can you tell me more about the shoulder stretching exercise?"

[0218] The terminal converts the new voice data into text data again and transmits it to the server.

[0219] The server analyzes again and generates additional instructions or explanations, which are sent to the terminal.

[0220] The terminal displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0221] In this way, the present invention is a system that enables elderly people and users with physical disabilities to easily and effectively train at home to maintain their health, and provides personalized and optimized instruction through emotion recognition. By providing specific training instructions and reference videos, it is expected that users will exercise in the correct way, contributing to maintaining their health and improving their quality of life.

[0222] The processing flow will be explained below.

[0223] Step 1:

[0224] The user talks to the system about their physical condition, worries, and feelings. For example, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0225] Step 2:

[0226] The terminal receives the user's voice through a microphone.

[0227] Step 3:

[0228] The device inputs the received voice data into a voice recognition engine and converts it into text data. The converted text data becomes, "I've been having really bad shoulder pain lately and it's making me feel depressed. What kind of exercise would be good?"

[0229] Step 4:

[0230] The terminal transmits the converted text data to the server.

[0231] Step 5:

[0232] The server receives the text data sent from the terminal.

[0233] Step 6:

[0234] The server uses natural language processing (NLP) technology to analyze the text data and identify the user's health condition. Specifically, it extracts and analyzes the keywords "stiff shoulders" and "feeling depressed."

[0235] Step 7:

[0236] The server then uses an emotion engine to analyze the emotions in the text data. From the part "feeling depressed," the server identifies the user's emotion as "depressed."

[0237] Step 8:

[0238] The server generates an appropriate training menu based on the analysis results. Using the AI ​​generation, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. The exercises include specific instructions such as "extend both arms overhead and slowly tilt them back and forth three times" and reference videos.

[0239] Step 9:

[0240] The server adjusts the training menu and delivery method based on the emotions recognized by the emotion engine. For example, for a user who is feeling depressed, it selects an exercise menu with a high relaxation effect.

[0241] Step 10:

[0242] The server transmits the generated training menu to the terminal. The adjusted training menu includes content that has a high relaxation effect.

[0243] Step 11:

[0244] The terminal displays the training menu received from the server on the user interface.

[0245] Step 12:

[0246] The user performs the exercises while viewing the displayed training menu. For example, the user performs the exercises according to the instructions for "shoulder stretching exercises." The user performs the exercises in the correct way while playing the reference video.

[0247] Step 13:

[0248] If the user has any questions or concerns during the training, they can input their voice again and the device will pick up the voice. For example, they might ask, "Can you tell me more about the shoulder stretching exercise?"

[0249] Step 14:

[0250] The terminal converts the newly received voice data back into text data and transmits it to the server.

[0251] Step 15:

[0252] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[0253] Step 16:

[0254] The terminal displays any additional instructions or explanations received to the user.

[0255] Example 2

[0256] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0257] It is difficult for elderly people and users with physical challenges to acquire and implement appropriate and effective training programs at home. Providing optimal training programs based on the user's current health condition and emotions is also insufficient. Conventional systems have been unable to recognize the user's emotions and provide individually optimized instruction.

[0258] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0259] In this invention, the server includes a means for acquiring voice input, a means for converting the acquired voice input into text, a means for analyzing the text data using natural language processing technology and emotion analysis technology to determine the health condition and emotions, a means for generating a training menu using a generative AI model, and a means for providing the generated menu to the user, thereby making it possible to provide an individually optimized training menu based on the user's health condition and emotions.

[0260] "Means for obtaining voice input" refers to devices or techniques for receiving voices spoken by a user to the system.

[0261] The "means for converting to text" refers to a speech recognition device or technology for converting acquired voice data into character data.

[0262] "Natural language processing technology" is a technology for analyzing text data and understanding and processing human language, and is used to identify a user's health condition and emotions.

[0263] "Emotion analysis technology" is a technology for analyzing and identifying emotions contained in text data.

[0264] A "generative AI model" is an artificial intelligence model that automatically generates outputs such as training menus based on input data.

[0265] The "means for generating a training menu" refers to a technique or device for creating an exercise program suited to a user based on the analyzed health condition and emotions.

[0266] "Means of providing" refers to the devices and technologies used to display and communicate the generated training menu to the user.

[0267] "Text data" is character data converted from voice input using voice recognition technology.

[0268] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[0269] Hardware and software used

[0270] This system uses the following hardware and software:

[0271] Device: A device used by a user, such as a smartphone or tablet, that includes a microphone, display, and internet connectivity.

[0272] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[0273] Natural language processing technology: Using the Google Cloud Natural Language API, health status and emotions are analyzed from text data.

[0274] Sentiment analysis technology: IBM Watson (registered trademark) Tone Analyzer is used to identify emotions from text data.

[0275] Generative AI model: OpenAI GPT-4 (registered trademark) is used to generate a training menu based on the analysis results.

[0276] System Operation

[0277] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0278] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API. The converted text data will contain the following content: "I've been having a lot of stiff shoulders lately and I'm feeling depressed. What kind of exercise is good?"

[0279] The device then sends the converted text data to a server, which analyzes it using the Google Cloud Natural Language API to identify the user's health condition. For example, it extracts keywords such as "stiff shoulders" and "feeling depressed." It also uses IBM Watson Tone Analyzer to identify the emotion of "feeling depressed."

[0280] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu tailored to the user's health condition and emotions. Specific training menus include "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. For example, instructions such as "extend both arms overhead and slowly tilt them back and forth three times" are included, along with links to reference videos.

[0281] The generated training menu is sent from the server to the terminal and displayed on the user interface. The user looks at the displayed training menu and performs the exercises. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. The user can perform the exercises in the correct way by playing the reference video.

[0282] If the user has any questions or concerns during the training, they can request more detailed explanations by voice input again. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes it again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0283] Specific examples

[0284] For example, if a user uses the prompt "Tell me some exercises for stiff shoulders that I can do while sitting in a chair," the system will do the following:

[0285] User: Can you tell me some exercises I can do to relieve shoulder stiffness while sitting in a chair?

[0286] Device: Receives audio and converts it to text using the Google Cloud Speech-to-Text API.

[0287] Server: Receives text data and analyzes it using Google Cloud Natural Language API and IBM Watson Tone Analyzer.

[0288] Server: Sends prompts to the generative AI model (OpenAI GPT-4) to generate an appropriate training menu.

[0289] Device: The generated training menu is displayed. Specifically, the instructions are displayed such as "Perform three sets of stretching exercises that involve slowly rotating your shoulders while sitting in a chair."

[0290] This system allows users to easily and effectively train at home to maintain their health, and receive personalized guidance that is optimized by recognizing their emotions.

[0291] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0292] Step 1:

[0293] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0294] Input: User's voice data.

[0295] Output: Audio data received through the microphone.

[0296] Step 2:

[0297] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API.

[0298] Specific operation: The device analyzes the speech waveform and generates corresponding text using a language model.

[0299] Input: Received audio data.

[0300] Output: Text data: "I've been having really bad shoulder pain lately and it's been making me feel depressed. What kind of exercise would be good?"

[0301] Step 3:

[0302] The terminal transmits the converted text data to the server using the HTTPS protocol.

[0303] Specific operation: Sends text data to the server as an HTTP request.

[0304] Input: Text data.

[0305] Output: The text data sent to the server.

[0306] Step 4:

[0307] The server analyzes the received text data using the Google Cloud Natural Language API to determine the user's health status.

[0308] Specific operation: The server performs morphological analysis on the text data and extracts keywords such as "stiff shoulders" and "feeling depressed."

[0309] Input: The text data sent.

[0310] Output: Analysis results on health status and emotions (e.g., stiff shoulders, depression).

[0311] Step 5:

[0312] The server also uses IBM Watson Tone Analyzer to identify emotions from the text data.

[0313] Specific operation: The server performs sentiment analysis on the text data to identify emotions such as "depressed."

[0314] Input: Text data.

[0315] Output: Sentiment analysis result: "I feel depressed."

[0316] Step 6:

[0317] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu that corresponds to the user's health condition and emotions.

[0318] Specific operation: Input prompts into the generative AI model, and generate an appropriate training menu (e.g., shoulder stretching exercises, relaxation exercises) based on them.

[0319] Input: Analysis results (health status and emotions).

[0320] Output: The generated training menu.

[0321] Step 7:

[0322] The server sends the generated training menu to the terminal using the HTTPS protocol.

[0323] Specific operation: The generated training menu is sent to the terminal as an HTTP response.

[0324] Input: The generated training menu.

[0325] Output: The training menu sent to the device.

[0326] Step 8:

[0327] The terminal displays the training menu received from the server to the user.

[0328] Specific operation: Display a training menu (e.g., shoulder stretching exercises, relaxation exercise instructions, and reference video links) on the user interface.

[0329] Input: Training menu received from the server.

[0330] Output: The training menu displayed in the user interface.

[0331] Step 9:

[0332] The user performs exercises according to the displayed training menu, for example, following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen.

[0333] Specific movements: Watch the reference video and perform the training using the correct form.

[0334] Input: The displayed training menu and reference video.

[0335] Output: The training actions performed by the user.

[0336] Step 10:

[0337] If a user has a question during the exercise, they can ask for more detailed explanation by speaking again, for example, "Can you tell me more about the shoulder stretch exercise?"

[0338] Input: Voice input for questions or to ask for further clarification.

[0339] Output: Audio data received through the microphone.

[0340] Step 11:

[0341] The device then converts the new voice data back into text using the Google Cloud Speech-to-Text API and sends it to the server.

[0342] Specific operation: Analyzes the audio waveform, generates the corresponding text, and sends the text to the server.

[0343] Input: New audio data.

[0344] Output: Text data and data sent to the server.

[0345] Step 12:

[0346] The server again analyzes using the Google Cloud Natural Language API and IBM Watson Tone Analyzer to generate additional instructions and explanations.

[0347] Specific Actions: Health and emotional analysis is performed on the text data, and generative AI models are used to generate additional instructions and explanations (e.g., detailed instructions for shoulder stretching exercises).

[0348] Input: New text data.

[0349] Output: Additional instructions or explanations.

[0350] Step 13:

[0351] The terminal displays the newly received instructions and explanations to the user, who then continues training based on them.

[0352] What it does: Display new instructions or explanations in the user interface.

[0353] Input: Any additional instructions or explanations received from the server.

[0354] Output: Instructions or explanations displayed in the user interface.

[0355] (Application example 2)

[0356] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0357] The problem that this invention aims to solve is to provide a system that enables elderly people and users with physical disabilities to acquire and implement appropriate training menus at home or in a physical store.Furthermore, by recognizing the user's emotions and adjusting the training menu and delivery method, the system provides optimal instruction for each user, contributing to maintaining health and improving QOL.

[0358] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0359] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition and emotions based on the voice input, means for generating an appropriate training menu based on the health condition and emotions, means for selecting and providing the generated training menu based on the user's emotions, and means for receiving feedback on the provided training menu and reanalyzing it, thereby enabling the user to acquire and appropriately perform optimal training in real time according to their current physical condition and emotions.

[0360] "Voice input" is voice data uttered by a user.

[0361] "Text data" refers to data obtained by converting voice input into text information.

[0362] "Health status" is information that indicates the physical and mental status of the user.

[0363] "Emotion" is information that indicates the user's mood or emotions.

[0364] A "training menu" is specific instructions for exercise and movement that are generated based on the user's health and mood.

[0365] "Smart terminals" are devices with advanced information processing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[0366] "Feedback" refers to comments or status reports provided by users during or after training.

[0367] "Real-time" means that there is no delay between data acquisition, processing, and providing results.

[0368] The "server" is an information processing device that analyzes voice data and text data and generates and provides training menus.

[0369] "Analysis" is a process of extracting information from voice and text data to understand the user's health condition and emotions.

[0370] A "generative AI model" is an algorithm or software that uses artificial intelligence to create new data such as training menus.

[0371] A "prompt sentence" is an input sentence given to a generative AI model, and serves as a hint for the AI ​​to generate new data based on that sentence.

[0372] This invention is a system that analyzes a user's health condition and emotions based on voice input, and generates and provides a training menu. The specific configuration of the system is as follows.

[0373] First, the device receives voice input from the user and converts it into text data. For voice recognition, smart devices such as smartphones with microphones, smart glasses, and head-mounted displays are used. For conversion, Google Speech Recognition API or other voice recognition software is used.

[0374] The acquired text data is then sent to a server, where it uses natural language processing (NLP) technology to analyze the user's health status and emotions. The NLP technology uses the Sentiment Analysis pipeline from the Transformers library. Based on the analyzed information, the server uses a generative AI model to generate an appropriate training menu. The generative AI model uses OpenAI's API.

[0375] The generated training menu is provided in real time, including text and reference videos. It is displayed on a smart device, and the user can follow it to carry out the training. Furthermore, the user can provide feedback during the training, which is also captured as voice input and analyzed again by the server. This allows the system to provide even more optimized training instructions.

[0376] Specific examples

[0377] As a concrete example, suppose a user wearing smart glasses at a fitness gym says, "My shoulder hurts today and I'm feeling a little unwell." This voice input is picked up through the smart glasses' microphone and converted into text by a speech recognition engine. The text data, "My shoulder hurts today and I'm feeling a little unwell," is sent to a server, where NLP technology and an emotion analysis engine analyze the user's health condition (shoulder pain) and emotion (unwell).

[0378] Based on the analysis results, OpenAI's generative AI model is used to create prompts and generate appropriate training menus. Examples of prompts include:

[0379] Example prompt sentence:

[0380] User condition: My shoulder hurts today and I'm feeling a bit unwell.

[0381] Emotion: Negative

[0382] So what to do?

[0383] The generated training menu might be, for example, "Start with some light stretching and then relax with some deep breathing exercises." This menu is displayed on the smart glasses' display, allowing the user to carry out the training.

[0384] This system allows users to obtain the optimal training in real time based on their physical condition and emotions at that moment, and then carry it out appropriately.

[0385] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0386] Step 1:

[0387] Acquiring voice input

[0388] The device picks up the user's voice through a microphone, and the user makes specific utterances such as "My shoulder hurts today, and I'm feeling a bit unwell."

[0389] Input: User's voice data

[0390] Output: Audio data file in the device

[0391] Step 2:

[0392] Converting audio data to text

[0393] The device converts the acquired voice data into text using voice recognition software such as Google Speech Recognition API. For example, the voice data "My shoulder hurts today and I'm feeling a bit unwell" is acquired as text data.

[0394] Input: Audio data file

[0395] Output: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[0396] Step 3:

[0397] Text data analysis

[0398] The server receives the text data sent from the device and analyzes the user's health status and emotions using natural language processing (NLP) techniques. Specifically, it uses the Sentiment Analysis pipeline in the Transformers library to analyze emotions and extract information about the health status from the text.

[0399] Input: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[0400] Output: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[0401] Step 4:

[0402] Training menu generation

[0403] The server generates an appropriate training menu based on the analysis results. It uses OpenAI's generative AI model to create prompts, and then generates a training menu based on those prompts. For example, it uses the prompt, "My shoulder hurts today, and I'm feeling a bit unwell. Emotion: Negative. So, what should I do?"

[0404] Input: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[0405] Output: Training menu (suggestions for light stretching and relaxation)

[0406] Step 5:

[0407] Providing training menus

[0408] The device displays the training menu received from the server on the user interface, and specific training content is displayed as text and reference video links on the display of the smart glasses or smartphone.

[0409] Input: Training Menu

[0410] Output: Training instructions displayed on the smart device

[0411] Step 6:

[0412] Get feedback

[0413] The user can provide additional instructions or give feedback by voice during the training, such as "Please tell me more about the shoulder stretching exercise."

[0414] Input: Audio data as feedback

[0415] Output: Audio data file in the device

[0416] Step 7:

[0417] Reanalyzing Feedback

[0418] The device then converts the newly acquired voice data back into text and sends it back to the server, which analyzes the feedback data, generates additional instructions and explanations, and sends them back to the device.

[0419] Input: Text data as feedback

[0420] Output: Additional training instructions or explanations

[0421] Step 8:

[0422] Providing additional instructions

[0423] The terminal displays the additional instructions and explanations received from the server on the user interface, and the user continues training based on them.

[0424] Input: Additional training instructions or explanations

[0425] Output: Additional instructions or explanations displayed on the smart device

[0426] This allows the entire process to proceed effectively in real time, allowing users to receive the training that best suits their physical and emotional state at that time.

[0427] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0428] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0429] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0430] [Second embodiment]

[0431] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0432] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0433] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0434] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0435] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0436] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0437] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0438] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0439] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0440] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0441] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0442] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0443] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0444] Acquiring voice input

[0445] The user talks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0446] Audio data conversion

[0447] The device inputs the received voice data into a voice recognition engine and converts it into text data, which generates the text "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0448] Health status analysis

[0449] The server receives the text data sent from the device and analyzes the content using natural language processing (NLP) technology. For example, it identifies the user's health condition based on the keywords "stiff shoulders" and "severe."

[0450] Training menu generation

[0451] The server generates an appropriate training menu based on the results of the health analysis. Using the generation AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness. Specific exercise instructions include "stretch both arms up and slowly tilt them to the left and right," along with a link to a reference video demonstrating this exercise.

[0452] Providing training instructions

[0453] The server transmits the generated training menu to the terminal, which includes specific instructions including text and video links.

[0454] Viewing and Running Training

[0455] The device displays the training menu received from the server on the user interface. The user performs the exercises while looking at the displayed training menu. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" displayed on the screen. The user can perform the exercises in the correct way while playing a reference video.

[0456] Follow-up

[0457] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes the data again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0458] In this way, the present invention is a system that allows elderly people and users with physical disabilities to easily and effectively train at home to maintain their health. By providing specific training instructions and reference videos, users can exercise in the correct way, which is expected to contribute to maintaining health and improving quality of life.

[0459] The processing flow will be explained below.

[0460] Step 1:

[0461] The user talks to the system about their physical condition and concerns. For example, "I have severe shoulder stiffness. Please tell me what kind of exercise I should do."

[0462] Step 2:

[0463] The terminal receives the user's voice through a microphone.

[0464] Step 3:

[0465] The device inputs the received voice data into a voice recognition engine and converts it into text data. For example, the generated text is "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0466] Step 4:

[0467] The terminal transmits the converted text data to the server.

[0468] Step 5:

[0469] The server receives the text data sent from the terminal.

[0470] Step 6:

[0471] The server analyzes the text data using natural language processing (NLP) technology, specifically extracting the keywords "stiff shoulders" and "severe" to determine the user's health condition.

[0472] Step 7:

[0473] The server generates an appropriate training menu based on the analysis of the user's health status. Using the AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness.

[0474] Step 8:

[0475] The server then creates a training menu by documenting specific exercises and adding links to reference videos. For example, "Extend both arms overhead and slowly tilt them back and forth three times."

[0476] Step 9:

[0477] The server transmits the generated training menu to the terminal.

[0478] Step 10:

[0479] The terminal displays the training menu received from the server on the user interface.

[0480] Step 11:

[0481] The user performs exercises while viewing the displayed training menu. They can also play reference videos to check the exercise methods.

[0482] Step 12:

[0483] If the user has any questions or concerns during the training, they can input their voice again, and the device will pick up the voice.

[0484] Step 13:

[0485] The terminal converts the newly received voice data back into text data and transmits it to the server.

[0486] Step 14:

[0487] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[0488] Step 15:

[0489] The terminal displays any additional instructions or explanations received to the user.

[0490] Example 1

[0491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0492] There is a need for a system that allows elderly people and users with physical challenges to easily obtain and implement appropriate training menus at home through voice input. However, existing systems have issues with the accuracy of speech recognition and text analysis, as well as the degree of customization of training menus, making it difficult for users to achieve the results they expect. For this reason, there is a need for a system that utilizes more advanced speech recognition, natural language processing, and generative AI technologies to provide appropriate training menus tailored to the user's health condition.

[0493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0494] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text data, means for analyzing the user's health condition based on the text data using natural language processing technology, means for generating an appropriate training menu using a generative AI model based on the analysis results of the health condition, means for transmitting the generated training menu to a terminal and providing it to the user, and means for displaying and executing training based on the training menu, thereby enabling the user to easily and effectively perform health maintenance training at home.

[0495] "Means for obtaining voice input" refers to a device or a function of a device that has the function of capturing the voice spoken by a user to the system and recording it as digital voice data.

[0496] A "means for converting speech input into text data" is software or a system capable of analyzing recorded speech data and converting it into a corresponding text string.

[0497] "Natural language processing technology" refers to a series of algorithms and techniques for analyzing text data and understanding its content and intent.

[0498] "Means for analyzing a user's health condition" refers to software or a system that has the function of extracting and analyzing information about a user's physical condition and health from input text data using natural language processing technology.

[0499] A "generative AI model" is a model that uses artificial intelligence technology to generate appropriate responses and suggestions from input data.

[0500] "Means for generating a training menu" refers to software or a system that has the function of automatically creating an appropriate training menu based on the user's health condition analyzed using a generative AI model.

[0501] The "means for transmitting a training menu to a terminal" refers to software or a system that has the function of transferring the generated training menu as data to a terminal and providing it to the user.

[0502] "Means for displaying and executing training" refers to a device or software that has the function of displaying the training menu provided to the terminal on a user interface and allowing the user to perform the training according to the instructions.

[0503] "Link to reference video" refers to a URL or data that provides access to a video that shows specific methods for performing a training menu.

[0504] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0505] Acquiring voice input

[0506] The user talks to the system about their physical condition and concerns. For example, they might say something like, "I have really bad shoulder stiffness. Please tell me what kind of exercise I should do." In order to obtain clear voice data, it is important to speak in an environment that eliminates as much background noise as possible.

[0507] Audio data conversion

[0508] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specific voice recognition engines such as Google Cloud Speech-to-Text API are used. For example, voice input data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do" is converted into text.

[0509] Health status analysis

[0510] The server receives the text data sent from the device and analyzes the content using natural language processing technology. A natural language processing engine such as AWS Comprehend is used here. Specifically, keywords such as "stiff shoulders" and "severe" are identified from the text data, and the user's health condition is analyzed.

[0511] Training menu generation

[0512] The server uses a generative AI model based on the analysis results to generate an appropriate training menu. Specifically, it uses a generative AI model such as OpenAI's GPT-3. For example, a prompt such as "Please generate a recommended training menu for an elderly person with severe shoulder stiffness" can be input, and the server will suggest "shoulder stretching exercises" based on that. The generated content includes instructions such as "Extend both arms up and slowly tilt them to the left and right" as well as links to reference videos.

[0513] Providing training instructions

[0514] The server sends the generated training menu to the device, which contains text instructions and links to reference videos.

[0515] Viewing and Running Training

[0516] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of the web or mobile application. The user performs the training by following the "shoulder stretch exercise" instructions displayed on the screen. The user can perform the exercises with the correct form while playing the reference video.

[0517] Follow-up

[0518] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine again and converts it into text data. The server then analyzes the text data and generates additional instructions or explanations. For example, it might generate additional instructions such as, "Please refer to the following video for shoulder stretching exercises," and send the video link to the device. The device then displays the newly received instructions or explanations to the user, who can then continue their training based on them.

[0519] Examples of specific examples and prompts

[0520] Example: A 70-year-old elderly person uses this system for stiff shoulders. He / she speaks to the system, saying, "I have severe shoulder stiffness, so please tell me what kind of exercises I should do." The system then suggests shoulder stretching exercises and provides a link to a video.

[0521] Example prompt sentence:

[0522] "Generate a recommended training menu for elderly people with severe shoulder stiffness."

[0523] "Please provide exercise instructions and a link to a helpful video to help relieve shoulder stiffness."

[0524] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0525] Step 1:

[0526] The user speaks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." The input is the user's voice, and the output is the voice data received by the device.

[0527] Step 2:

[0528] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specifically, it uses the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." Here, the voice recognition engine analyzes the phonemes of the voice and generates a corresponding string of characters.

[0529] Step 3:

[0530] The server receives the text data sent from the device and analyzes it using natural language processing technology. Specifically, AWS Comprehend is used to extract keywords from the text data and identify the health condition. The input is text data, and the output is analysis results such as "stiff shoulders" or "severe." Natural language processing technology extracts keywords and performs sentiment analysis.

[0531] Step 4:

[0532] The server uses a generative AI model to generate an appropriate training menu based on the health status analysis results. Specifically, it uses OpenAI's GPT-3 and inputs the prompt "Please generate a recommended training menu for elderly people with severe shoulder stiffness." The inputs are the analysis results and the prompt, and the output is a training menu such as "Extend both arms up and slowly tilt them left and right" and a link to a reference video. The generative AI model suggests optimal training content based on the user's health status.

[0533] Step 5:

[0534] The server sends the generated training menu to the terminal. The input is the training menu and video link, and the output is the data to be sent to the terminal. The server transfers the text data of the training menu and the video link to the terminal.

[0535] Step 6:

[0536] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of a web or mobile application. The input is the data received from the server, and the output is the display on the interface that the user can visually confirm.

[0537] Step 7:

[0538] The user performs exercises according to the displayed training menu. Specifically, the user performs the exercises by following the instructions displayed on the screen: "Extend both arms upwards and slowly tilt them to the left and right." A reference video is played and the user performs the exercises while checking the movements. The input is the displayed training menu, and the output is the user performing the exercises with the correct form.

[0539] Step 8:

[0540] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine, which converts it back into text data. The input is voice data, and the output is text data.

[0541] Step 9:

[0542] The server analyzes the newly received text data and generates additional instructions or explanations. The input is new text data, and the output is additional instructions or explanations. For example, the server generates an additional instruction such as "Please refer to the following video" and sends it to the device.

[0543] Step 10:

[0544] The terminal displays newly received instructions and explanations to the user. The input is the additional instructions and explanations sent from the server, and the output is what is displayed on the user interface. The user continues training based on that.

[0545] (Application example 1)

[0546] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0547] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, or elsewhere. However, existing systems have problems in that users cannot acquire effective training menus in real time or cannot properly execute acquired menus. For example, the accuracy of voice input or voice recognition may be low, resulting in an inappropriate menu not being generated. Furthermore, the generated menu may be difficult for users to understand, resulting in the inability to perform exercises in the correct manner. Furthermore, specialized equipment or devices are required, which may be difficult for elderly people and users with physical challenges to implement. There is a need for a system that solves these problems and allows users to easily and effectively acquire and implement training menus.

[0548] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0549] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition based on the voice input, means for generating an appropriate training menu based on the health condition, means for providing the generated training menu to the user, and means for displaying the generated training menu on a smart device, thereby enabling the user to obtain an effective training menu in real time at home, at a fitness facility, or the like, and to exercise in the correct way via the smart device.

[0550] The "means for acquiring voice input" is a device or method for acquiring the voice uttered by the user as digital data.

[0551] A "means for converting captured speech input to text" is a device or method for converting speech data into natural language text data.

[0552] The "means for analyzing a user's health condition based on a voice input" refers to a device or method for analyzing text data obtained from a voice input and determining a user's health condition.

[0553] The "means for generating an appropriate training menu based on health condition" refers to a device or method that automatically creates an appropriate exercise and stretching menu based on the analyzed health condition.

[0554] The "means for providing the generated training menu to the user" refers to a device or method that determines how to provide the generated training menu to the user and implements that method.

[0555] The "means for displaying the generated training menu on a smart device" refers to a device or method for visually displaying the generated training menu on a device such as a smartphone or smart glasses used by the user.

[0556] A "generative AI model" is a model that uses artificial intelligence techniques to generate new information from data.

[0557] A "prompt" is an input text given to a generative AI model, used to cause the model to generate specific instructions or information.

[0558] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, etc. The main components of this system are a voice input acquisition means, a means for converting the acquired voice input into text, a means for analyzing the user's health condition, a means for generating a training menu, a means for providing the generated training menu to the user, and a means for displaying the menu on a smart device.

[0559] Program processing

[0560] The whole system includes the following steps:

[0561] 1. Voice input acquisition: The user speaks to a smart device (e.g., smart glasses or smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is acquired using a microphone.

[0562] 2. Speech Recognition: The captured voice data is converted into text data using the Google Cloud Speech-to-Text API, which transcribes the user's speech into text.

[0563] 3. Text analysis: The server analyzes the text data using natural language processing (NLP) techniques (e.g., spaCy) to identify the user's health condition. For example, the text "My left knee hurts" is analyzed and it is recognized that the user has a knee problem.

[0564] 4. Training menu generation: The server generates an appropriate training menu using a generative AI model (e.g., OpenAI's GPT) based on the analysis results. An example of this prompt might be, "Please suggest an effective exercise menu for my left knee pain." The generated training menu includes specific exercise content and links to reference videos.

[0565] 5. Providing training menu: The generated training menu is sent to the smart device, where the user can check the instructions on the display of their smart glasses or smartphone.

[0566] 6. Display and execution of training: The generated exercise instructions are displayed on the display of the user device. For example, specific exercise instructions such as "Perform knee support stretches" are displayed, and a reference video can also be viewed.

[0567] Hardware and software used

[0568] Smart devices (e.g. smart glasses, smartphones): Used for voice input and display.

[0569] Cloud server: Used for voice recognition, text analysis, and training menu generation.

[0570] Speech recognition API (e.g., Google Cloud Speech-to-Text API): Used to convert voice data into text.

[0571] Natural language processing libraries (e.g. spaCy): Used to analyze text data.

[0572] Generative AI models (e.g., OpenAI's GPT): Used to generate training menus.

[0573] Specific examples

[0574] For example, if an elderly person says to a trainer at a fitness gym, "My left knee hurts. Please tell me what exercises I should do," the smart glasses will recognize the speech and convert it into text. The server then analyzes the text and determines that the person has pain in their left knee. An appropriate training menu is generated using a generative AI model, and the trainer is shown instructions such as, "Perform stretches to support your knee." The smart glasses also display a video showing how to perform the exercises correctly.

[0575] Prompt Sentence Examples

[0576] "Please suggest an effective exercise routine for my left knee pain."

[0577] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0578] Step 1:

[0579] Acquiring voice input

[0580] The user speaks into the microphone of their smart device (e.g., smart glasses, smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is picked up by the microphone on the smart device.

[0581] Input: User's voice

[0582] Output: Audio data

[0583] Step 2:

[0584] Converting audio data to text

[0585] The device sends the captured voice data to the Google Cloud Speech-to-Text API and converts it into text data. For example, a speech like "My left knee hurts, so please tell me what exercises I should do" is converted into the corresponding text.

[0586] Input: Audio data

[0587] Output: Text data

[0588] Step 3:

[0589] Text data analysis

[0590] The server receives the text data sent from the device and analyzes it using natural language processing (NLP) technology. For example, the keywords "knee" and "pain" may be extracted as analysis results, and the user's health condition may be identified as "pain in the left knee."

[0591] Input: Text data

[0592] Output: Health status analysis results

[0593] Step 4:

[0594] Training menu generation

[0595] Based on the analysis results, the server inputs prompts into a generative AI model (e.g., OpenAI's GPT) to generate an appropriate training menu. For example, a prompt such as "Please suggest an effective exercise menu for left knee pain" generates specific exercise content and links to reference videos.

[0596] Input: Analysis results regarding health status, prompt text

[0597] Output: Training menu

[0598] Step 5:

[0599] Sending and providing training menus

[0600] The server sends the generated training menu to the user's smart device, which then notifies the user of the menu. Specifically, the generated training menu is displayed on the display of smart glasses or a smartphone.

[0601] Input: Training Menu

[0602] Output: Training menu displayed on a smart device

[0603] Step 6:

[0604] Viewing and Running Training

[0605] The user device supports the user in exercising accurately based on the displayed training menu. For example, exercise instructions such as "Perform stretches to support your knees" are displayed on the smart glasses along with a reference video. The user exercises while watching the video.

[0606] Input: Training menu displayed on smart device

[0607] Output: User executes training

[0608] Each processing step allows the user to obtain an appropriate training menu in real time and perform exercises in the correct way through their smart device.

[0609] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0610] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[0611] Acquiring and converting voice input

[0612] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0613] The device receives the user's voice through a microphone, inputs this voice data into a speech recognition engine, and converts it into text data. The converted text data will contain the following content: "I've been suffering from severe shoulder stiffness lately and I'm feeling depressed. What kind of exercise is good?"

[0614] Health and Emotion Analysis

[0615] The terminal transmits the converted text data to the server.

[0616] The server analyzes the received text data using natural language processing (NLP) technology to identify the user's health condition, for example by extracting keywords such as "stiff shoulders" and "feeling depressed."

[0617] Furthermore, the server uses an emotion engine to analyze the user's emotions, for example, to identify the emotion "feeling depressed."

[0618] Creation and provision of training menus

[0619] The server generates an appropriate training menu based on the analysis results. Using the generation AI, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. This includes specific exercise instructions and links to reference videos. For example, instructions might include "extend both arms overhead and slowly tilt them back and forth three times."

[0620] The server selects a training menu based on the user's emotions and transmits it to the terminal. The selected menu is most effective for the user, for example, a menu that enhances relaxation.

[0621] Running the training

[0622] The terminal displays the training menu received from the server on the user interface.

[0623] The user looks at the displayed training menu and performs the exercises. For example, they perform the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. They can perform the exercises in the correct way by playing the reference video.

[0624] Follow-up

[0625] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they can ask, "Can you tell me more about the shoulder stretching exercise?"

[0626] The terminal converts the new voice data into text data again and transmits it to the server.

[0627] The server analyzes again and generates additional instructions or explanations, which are sent to the terminal.

[0628] The terminal displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0629] In this way, the present invention is a system that enables elderly people and users with physical disabilities to easily and effectively train at home to maintain their health, and provides personalized and optimized instruction through emotion recognition. By providing specific training instructions and reference videos, it is expected that users will exercise in the correct way, contributing to maintaining their health and improving their quality of life.

[0630] The processing flow will be explained below.

[0631] Step 1:

[0632] The user talks to the system about their physical condition, worries, and feelings. For example, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0633] Step 2:

[0634] The terminal receives the user's voice through a microphone.

[0635] Step 3:

[0636] The device inputs the received voice data into a voice recognition engine and converts it into text data. The converted text data becomes, "I've been having really bad shoulder pain lately and it's making me feel depressed. What kind of exercise would be good?"

[0637] Step 4:

[0638] The terminal transmits the converted text data to the server.

[0639] Step 5:

[0640] The server receives the text data sent from the terminal.

[0641] Step 6:

[0642] The server uses natural language processing (NLP) technology to analyze the text data and identify the user's health condition. Specifically, it extracts and analyzes the keywords "stiff shoulders" and "feeling depressed."

[0643] Step 7:

[0644] The server then uses an emotion engine to analyze the emotions in the text data. From the part "feeling depressed," the server identifies the user's emotion as "depressed."

[0645] Step 8:

[0646] The server generates an appropriate training menu based on the analysis results. Using the AI ​​generation, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. The exercises include specific instructions such as "extend both arms overhead and slowly tilt them back and forth three times" and reference videos.

[0647] Step 9:

[0648] The server adjusts the training menu and delivery method based on the emotions recognized by the emotion engine. For example, for a user who is feeling depressed, it selects an exercise menu with a high relaxation effect.

[0649] Step 10:

[0650] The server transmits the generated training menu to the terminal. The adjusted training menu includes content that has a high relaxation effect.

[0651] Step 11:

[0652] The terminal displays the training menu received from the server on the user interface.

[0653] Step 12:

[0654] The user performs the exercises while viewing the displayed training menu. For example, the user performs the exercises according to the instructions for "shoulder stretching exercises." The user performs the exercises in the correct way while playing the reference video.

[0655] Step 13:

[0656] If the user has any questions or concerns during the training, they can input their voice again and the device will pick up the voice. For example, they might ask, "Can you tell me more about the shoulder stretching exercise?"

[0657] Step 14:

[0658] The terminal converts the newly received voice data back into text data and transmits it to the server.

[0659] Step 15:

[0660] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[0661] Step 16:

[0662] The terminal displays any additional instructions or explanations received to the user.

[0663] Example 2

[0664] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0665] It is difficult for elderly people and users with physical challenges to acquire and implement appropriate and effective training programs at home. Providing optimal training programs based on the user's current health condition and emotions is also insufficient. Conventional systems have been unable to recognize the user's emotions and provide individually optimized instruction.

[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0667] In this invention, the server includes a means for acquiring voice input, a means for converting the acquired voice input into text, a means for analyzing the text data using natural language processing technology and emotion analysis technology to determine the health condition and emotions, a means for generating a training menu using a generative AI model, and a means for providing the generated menu to the user, thereby making it possible to provide an individually optimized training menu based on the user's health condition and emotions.

[0668] "Means for obtaining voice input" refers to devices or techniques for receiving voices spoken by a user to the system.

[0669] The "means for converting to text" refers to a speech recognition device or technology for converting acquired voice data into character data.

[0670] "Natural language processing technology" is a technology for analyzing text data and understanding and processing human language, and is used to identify a user's health condition and emotions.

[0671] "Emotion analysis technology" is a technology for analyzing and identifying emotions contained in text data.

[0672] A "generative AI model" is an artificial intelligence model that automatically generates outputs such as training menus based on input data.

[0673] The "means for generating a training menu" refers to a technique or device for creating an exercise program suited to a user based on the analyzed health condition and emotions.

[0674] "Means of providing" refers to the devices and technologies used to display and communicate the generated training menu to the user.

[0675] "Text data" is character data converted from voice input using voice recognition technology.

[0676] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[0677] Hardware and software used

[0678] This system uses the following hardware and software:

[0679] Device: A device used by a user, such as a smartphone or tablet, that includes a microphone, display, and internet connectivity.

[0680] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[0681] Natural language processing technology: Using the Google Cloud Natural Language API, health status and emotions are analyzed from text data.

[0682] Sentiment analysis technology: Uses IBM Watson Tone Analyzer to identify emotions from text data.

[0683] Generative AI model: Uses OpenAI GPT-4 to generate a training menu based on the analysis results.

[0684] System Operation

[0685] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0686] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API. The converted text data will contain the following content: "I've been having a lot of stiff shoulders lately and I'm feeling depressed. What kind of exercise is good?"

[0687] The device then sends the converted text data to a server, which analyzes it using the Google Cloud Natural Language API to identify the user's health condition. For example, it extracts keywords such as "stiff shoulders" and "feeling depressed." It also uses IBM Watson Tone Analyzer to identify the emotion of "feeling depressed."

[0688] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu tailored to the user's health condition and emotions. Specific training menus include "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. For example, instructions such as "extend both arms overhead and slowly tilt them back and forth three times" are included, along with links to reference videos.

[0689] The generated training menu is sent from the server to the terminal and displayed on the user interface. The user looks at the displayed training menu and performs the exercises. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. The user can perform the exercises in the correct way by playing the reference video.

[0690] If the user has any questions or concerns during the training, they can request more detailed explanations by voice input again. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes it again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0691] Specific examples

[0692] For example, if a user uses the prompt "Tell me some exercises for stiff shoulders that I can do while sitting in a chair," the system will do the following:

[0693] User: Can you tell me some exercises I can do to relieve shoulder stiffness while sitting in a chair?

[0694] Device: Receives audio and converts it to text using the Google Cloud Speech-to-Text API.

[0695] Server: Receives text data and analyzes it using Google Cloud Natural Language API and IBM Watson Tone Analyzer.

[0696] Server: Sends prompts to the generative AI model (OpenAI GPT-4) to generate an appropriate training menu.

[0697] Device: The generated training menu is displayed. Specifically, the instructions are displayed such as "Perform three sets of stretching exercises that involve slowly rotating your shoulders while sitting in a chair."

[0698] This system allows users to easily and effectively train at home to maintain their health, and receive personalized guidance that is optimized by recognizing their emotions.

[0699] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0700] Step 1:

[0701] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[0702] Input: User's voice data.

[0703] Output: Audio data received through the microphone.

[0704] Step 2:

[0705] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API.

[0706] Specific operation: The device analyzes the speech waveform and generates corresponding text using a language model.

[0707] Input: Received audio data.

[0708] Output: Text data: "I've been having really bad shoulder pain lately and it's been making me feel depressed. What kind of exercise would be good?"

[0709] Step 3:

[0710] The terminal transmits the converted text data to the server using the HTTPS protocol.

[0711] Specific operation: Sends text data to the server as an HTTP request.

[0712] Input: Text data.

[0713] Output: The text data sent to the server.

[0714] Step 4:

[0715] The server analyzes the received text data using the Google Cloud Natural Language API to determine the user's health status.

[0716] Specific operation: The server performs morphological analysis on the text data and extracts keywords such as "stiff shoulders" and "feeling depressed."

[0717] Input: The text data sent.

[0718] Output: Analysis results on health status and emotions (e.g., stiff shoulders, depression).

[0719] Step 5:

[0720] The server also uses IBM Watson Tone Analyzer to identify emotions from the text data.

[0721] Specific operation: The server performs sentiment analysis on the text data to identify emotions such as "depressed."

[0722] Input: Text data.

[0723] Output: Sentiment analysis result: "I feel depressed."

[0724] Step 6:

[0725] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu that corresponds to the user's health condition and emotions.

[0726] Specific operation: Input prompts into the generative AI model, and generate an appropriate training menu (e.g., shoulder stretching exercises, relaxation exercises) based on them.

[0727] Input: Analysis results (health status and emotions).

[0728] Output: The generated training menu.

[0729] Step 7:

[0730] The server sends the generated training menu to the terminal using the HTTPS protocol.

[0731] Specific operation: The generated training menu is sent to the terminal as an HTTP response.

[0732] Input: The generated training menu.

[0733] Output: The training menu sent to the device.

[0734] Step 8:

[0735] The terminal displays the training menu received from the server to the user.

[0736] Specific operation: Display a training menu (e.g., shoulder stretching exercises, relaxation exercise instructions, and reference video links) on the user interface.

[0737] Input: Training menu received from the server.

[0738] Output: The training menu displayed in the user interface.

[0739] Step 9:

[0740] The user performs exercises according to the displayed training menu, for example, following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen.

[0741] Specific movements: Watch the reference video and perform the training using the correct form.

[0742] Input: The displayed training menu and reference video.

[0743] Output: The training actions performed by the user.

[0744] Step 10:

[0745] If a user has a question during the exercise, they can ask for more detailed explanation by speaking again, for example, "Can you tell me more about the shoulder stretch exercise?"

[0746] Input: Voice input for questions or to ask for further clarification.

[0747] Output: Audio data received through the microphone.

[0748] Step 11:

[0749] The device then converts the new voice data back into text using the Google Cloud Speech-to-Text API and sends it to the server.

[0750] Specific operation: Analyzes the audio waveform, generates the corresponding text, and sends the text to the server.

[0751] Input: New audio data.

[0752] Output: Text data and data sent to the server.

[0753] Step 12:

[0754] The server again analyzes using the Google Cloud Natural Language API and IBM Watson Tone Analyzer to generate additional instructions and explanations.

[0755] Specific Actions: Health and emotional analysis is performed on the text data, and generative AI models are used to generate additional instructions and explanations (e.g., detailed instructions for shoulder stretching exercises).

[0756] Input: New text data.

[0757] Output: Additional instructions or explanations.

[0758] Step 13:

[0759] The terminal displays the newly received instructions and explanations to the user, who then continues training based on them.

[0760] What it does: Display new instructions or explanations in the user interface.

[0761] Input: Any additional instructions or explanations received from the server.

[0762] Output: Instructions or explanations displayed in the user interface.

[0763] (Application example 2)

[0764] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0765] The problem that this invention aims to solve is to provide a system that enables elderly people and users with physical disabilities to acquire and implement appropriate training menus at home or in a physical store.Furthermore, by recognizing the user's emotions and adjusting the training menu and delivery method, the system provides optimal instruction for each user, contributing to maintaining health and improving QOL.

[0766] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0767] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition and emotions based on the voice input, means for generating an appropriate training menu based on the health condition and emotions, means for selecting and providing the generated training menu based on the user's emotions, and means for receiving feedback on the provided training menu and reanalyzing it, thereby enabling the user to acquire and appropriately perform optimal training in real time according to their current physical condition and emotions.

[0768] "Voice input" is voice data uttered by a user.

[0769] "Text data" refers to data obtained by converting voice input into text information.

[0770] "Health status" is information that indicates the physical and mental status of the user.

[0771] "Emotion" is information that indicates the user's mood or emotions.

[0772] A "training menu" is specific instructions for exercise and movement that are generated based on the user's health and mood.

[0773] "Smart terminals" are devices with advanced information processing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[0774] "Feedback" refers to comments or status reports provided by users during or after training.

[0775] "Real-time" means that there is no delay between data acquisition, processing, and providing results.

[0776] The "server" is an information processing device that analyzes voice data and text data and generates and provides training menus.

[0777] "Analysis" is a process of extracting information from voice and text data to understand the user's health condition and emotions.

[0778] A "generative AI model" is an algorithm or software that uses artificial intelligence to create new data such as training menus.

[0779] A "prompt sentence" is an input sentence given to a generative AI model, and serves as a hint for the AI ​​to generate new data based on that sentence.

[0780] This invention is a system that analyzes a user's health condition and emotions based on voice input, and generates and provides a training menu. The specific configuration of the system is as follows.

[0781] First, the device receives voice input from the user and converts it into text data. For voice recognition, smart devices such as smartphones with microphones, smart glasses, and head-mounted displays are used. For conversion, Google Speech Recognition API or other voice recognition software is used.

[0782] The acquired text data is then sent to a server, where it uses natural language processing (NLP) technology to analyze the user's health status and emotions. The NLP technology uses the Sentiment Analysis pipeline from the Transformers library. Based on the analyzed information, the server uses a generative AI model to generate an appropriate training menu. The generative AI model uses OpenAI's API.

[0783] The generated training menu is provided in real time, including text and reference videos. It is displayed on a smart device, and the user can follow it to carry out the training. Furthermore, the user can provide feedback during the training, which is also captured as voice input and analyzed again by the server. This allows the system to provide even more optimized training instructions.

[0784] Specific examples

[0785] As a concrete example, suppose a user wearing smart glasses at a fitness gym says, "My shoulder hurts today and I'm feeling a little unwell." This voice input is picked up through the smart glasses' microphone and converted into text by a speech recognition engine. The text data, "My shoulder hurts today and I'm feeling a little unwell," is sent to a server, where NLP technology and an emotion analysis engine analyze the user's health condition (shoulder pain) and emotion (unwell).

[0786] Based on the analysis results, OpenAI's generative AI model is used to create prompts and generate appropriate training menus. Examples of prompts include:

[0787] Example prompt sentence:

[0788] User condition: My shoulder hurts today and I'm feeling a bit unwell.

[0789] Emotion: Negative

[0790] So what to do?

[0791] The generated training menu might be, for example, "Start with some light stretching and then relax with some deep breathing exercises." This menu is displayed on the smart glasses' display, allowing the user to carry out the training.

[0792] This system allows users to obtain the optimal training in real time based on their physical condition and emotions at that moment, and then carry it out appropriately.

[0793] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0794] Step 1:

[0795] Acquiring voice input

[0796] The device picks up the user's voice through a microphone, and the user makes specific utterances such as "My shoulder hurts today, and I'm feeling a bit unwell."

[0797] Input: User's voice data

[0798] Output: Audio data file in the device

[0799] Step 2:

[0800] Converting audio data to text

[0801] The device converts the acquired voice data into text using voice recognition software such as Google Speech Recognition API. For example, the voice data "My shoulder hurts today and I'm feeling a bit unwell" is acquired as text data.

[0802] Input: Audio data file

[0803] Output: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[0804] Step 3:

[0805] Text data analysis

[0806] The server receives the text data sent from the device and analyzes the user's health status and emotions using natural language processing (NLP) techniques. Specifically, it uses the Sentiment Analysis pipeline in the Transformers library to analyze emotions and extract information about the health status from the text.

[0807] Input: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[0808] Output: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[0809] Step 4:

[0810] Training menu generation

[0811] The server generates an appropriate training menu based on the analysis results. It uses OpenAI's generative AI model to create prompts, and then generates a training menu based on those prompts. For example, it uses the prompt, "My shoulder hurts today, and I'm feeling a bit unwell. Emotion: Negative. So, what should I do?"

[0812] Input: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[0813] Output: Training menu (suggestions for light stretching and relaxation)

[0814] Step 5:

[0815] Providing training menus

[0816] The device displays the training menu received from the server on the user interface, and specific training content is displayed as text and reference video links on the display of the smart glasses or smartphone.

[0817] Input: Training Menu

[0818] Output: Training instructions displayed on the smart device

[0819] Step 6:

[0820] Get feedback

[0821] The user can provide additional instructions or give feedback by voice during the training, such as "Please tell me more about the shoulder stretching exercise."

[0822] Input: Audio data as feedback

[0823] Output: Audio data file in the device

[0824] Step 7:

[0825] Reanalyzing Feedback

[0826] The device then converts the newly acquired voice data back into text and sends it back to the server, which analyzes the feedback data, generates additional instructions and explanations, and sends them back to the device.

[0827] Input: Text data as feedback

[0828] Output: Additional training instructions or explanations

[0829] Step 8:

[0830] Providing additional instructions

[0831] The terminal displays the additional instructions and explanations received from the server on the user interface, and the user continues training based on them.

[0832] Input: Additional training instructions or explanations

[0833] Output: Additional instructions or explanations displayed on the smart device

[0834] This allows the entire process to proceed effectively in real time, allowing users to receive the training that best suits their physical and emotional state at that time.

[0835] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0836] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0837] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0838] [Third embodiment]

[0839] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0840] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0841] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0842] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0843] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0844] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0845] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0846] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0847] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0848] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0849] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0850] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0851] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0852] Acquiring voice input

[0853] The user talks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0854] Audio data conversion

[0855] The device inputs the received voice data into a voice recognition engine and converts it into text data, which generates the text "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0856] Health status analysis

[0857] The server receives the text data sent from the device and analyzes the content using natural language processing (NLP) technology. For example, it identifies the user's health condition based on the keywords "stiff shoulders" and "severe."

[0858] Training menu generation

[0859] The server generates an appropriate training menu based on the results of the health analysis. Using the generation AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness. Specific exercise instructions include "stretch both arms up and slowly tilt them to the left and right," along with a link to a reference video demonstrating this exercise.

[0860] Providing training instructions

[0861] The server transmits the generated training menu to the terminal, which includes specific instructions including text and video links.

[0862] Viewing and Running Training

[0863] The device displays the training menu received from the server on the user interface. The user performs the exercises while looking at the displayed training menu. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" displayed on the screen. The user can perform the exercises in the correct way while playing a reference video.

[0864] Follow-up

[0865] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes the data again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[0866] In this way, the present invention is a system that allows elderly people and users with physical disabilities to easily and effectively train at home to maintain their health. By providing specific training instructions and reference videos, users can exercise in the correct way, which is expected to contribute to maintaining health and improving quality of life.

[0867] The processing flow will be explained below.

[0868] Step 1:

[0869] The user talks to the system about their physical condition and concerns. For example, "I have severe shoulder stiffness. Please tell me what kind of exercise I should do."

[0870] Step 2:

[0871] The terminal receives the user's voice through a microphone.

[0872] Step 3:

[0873] The device inputs the received voice data into a voice recognition engine and converts it into text data. For example, the generated text is "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[0874] Step 4:

[0875] The terminal transmits the converted text data to the server.

[0876] Step 5:

[0877] The server receives the text data sent from the terminal.

[0878] Step 6:

[0879] The server analyzes the text data using natural language processing (NLP) technology, specifically extracting the keywords "stiff shoulders" and "severe" to determine the user's health condition.

[0880] Step 7:

[0881] The server generates an appropriate training menu based on the analysis of the user's health status. Using the AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness.

[0882] Step 8:

[0883] The server then creates a training menu by documenting specific exercises and adding links to reference videos. For example, "Extend both arms overhead and slowly tilt them back and forth three times."

[0884] Step 9:

[0885] The server transmits the generated training menu to the terminal.

[0886] Step 10:

[0887] The terminal displays the training menu received from the server on the user interface.

[0888] Step 11:

[0889] The user performs exercises while viewing the displayed training menu. They can also play reference videos to check the exercise methods.

[0890] Step 12:

[0891] If the user has any questions or concerns during the training, they can input their voice again, and the device will pick up the voice.

[0892] Step 13:

[0893] The terminal converts the newly received voice data back into text data and transmits it to the server.

[0894] Step 14:

[0895] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[0896] Step 15:

[0897] The terminal displays any additional instructions or explanations received to the user.

[0898] Example 1

[0899] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0900] There is a need for a system that allows elderly people and users with physical challenges to easily obtain and implement appropriate training menus at home through voice input. However, existing systems have issues with the accuracy of speech recognition and text analysis, as well as the degree of customization of training menus, making it difficult for users to achieve the results they expect. For this reason, there is a need for a system that utilizes more advanced speech recognition, natural language processing, and generative AI technologies to provide appropriate training menus tailored to the user's health condition.

[0901] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0902] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text data, means for analyzing the user's health condition based on the text data using natural language processing technology, means for generating an appropriate training menu using a generative AI model based on the analysis results of the health condition, means for transmitting the generated training menu to a terminal and providing it to the user, and means for displaying and executing training based on the training menu, thereby enabling the user to easily and effectively perform health maintenance training at home.

[0903] "Means for obtaining voice input" refers to a device or a function of a device that has the function of capturing the voice spoken by a user to the system and recording it as digital voice data.

[0904] A "means for converting speech input into text data" is software or a system capable of analyzing recorded speech data and converting it into a corresponding text string.

[0905] "Natural language processing technology" refers to a series of algorithms and techniques for analyzing text data and understanding its content and intent.

[0906] "Means for analyzing a user's health condition" refers to software or a system that has the function of extracting and analyzing information about a user's physical condition and health from input text data using natural language processing technology.

[0907] A "generative AI model" is a model that uses artificial intelligence technology to generate appropriate responses and suggestions from input data.

[0908] "Means for generating a training menu" refers to software or a system that has the function of automatically creating an appropriate training menu based on the user's health condition analyzed using a generative AI model.

[0909] The "means for transmitting a training menu to a terminal" refers to software or a system that has the function of transferring the generated training menu as data to a terminal and providing it to the user.

[0910] "Means for displaying and executing training" refers to a device or software that has the function of displaying the training menu provided to the terminal on a user interface and allowing the user to perform the training according to the instructions.

[0911] "Link to reference video" refers to a URL or data that provides access to a video that shows specific methods for performing a training menu.

[0912] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[0913] Acquiring voice input

[0914] The user talks to the system about their physical condition and concerns. For example, they might say something like, "I have really bad shoulder stiffness. Please tell me what kind of exercise I should do." In order to obtain clear voice data, it is important to speak in an environment that eliminates as much background noise as possible.

[0915] Audio data conversion

[0916] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specific voice recognition engines such as Google Cloud Speech-to-Text API are used. For example, voice input data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do" is converted into text.

[0917] Health status analysis

[0918] The server receives the text data sent from the device and analyzes the content using natural language processing technology. A natural language processing engine such as AWS Comprehend is used here. Specifically, keywords such as "stiff shoulders" and "severe" are identified from the text data, and the user's health condition is analyzed.

[0919] Training menu generation

[0920] The server uses a generative AI model based on the analysis results to generate an appropriate training menu. Specifically, it uses a generative AI model such as OpenAI's GPT-3. For example, a prompt such as "Please generate a recommended training menu for an elderly person with severe shoulder stiffness" can be input, and the server will suggest "shoulder stretching exercises" based on that. The generated content includes instructions such as "Extend both arms up and slowly tilt them to the left and right" as well as links to reference videos.

[0921] Providing training instructions

[0922] The server sends the generated training menu to the device, which contains text instructions and links to reference videos.

[0923] Viewing and Running Training

[0924] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of the web or mobile application. The user performs the training by following the "shoulder stretch exercise" instructions displayed on the screen. The user can perform the exercises with the correct form while playing the reference video.

[0925] Follow-up

[0926] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine again and converts it into text data. The server then analyzes the text data and generates additional instructions or explanations. For example, it might generate additional instructions such as, "Please refer to the following video for shoulder stretching exercises," and send the video link to the device. The device then displays the newly received instructions or explanations to the user, who can then continue their training based on them.

[0927] Examples of specific examples and prompts

[0928] Example: A 70-year-old elderly person uses this system for stiff shoulders. He / she speaks to the system, saying, "I have severe shoulder stiffness, so please tell me what kind of exercises I should do." The system then suggests shoulder stretching exercises and provides a link to a video.

[0929] Example prompt sentence:

[0930] "Generate a recommended training menu for elderly people with severe shoulder stiffness."

[0931] "Please provide exercise instructions and a link to a helpful video to help relieve shoulder stiffness."

[0932] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0933] Step 1:

[0934] The user speaks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." The input is the user's voice, and the output is the voice data received by the device.

[0935] Step 2:

[0936] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specifically, it uses the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." Here, the voice recognition engine analyzes the phonemes of the voice and generates a corresponding string of characters.

[0937] Step 3:

[0938] The server receives the text data sent from the device and analyzes it using natural language processing technology. Specifically, AWS Comprehend is used to extract keywords from the text data and identify the health condition. The input is text data, and the output is analysis results such as "stiff shoulders" or "severe." Natural language processing technology extracts keywords and performs sentiment analysis.

[0939] Step 4:

[0940] The server uses a generative AI model to generate an appropriate training menu based on the health status analysis results. Specifically, it uses OpenAI's GPT-3 and inputs the prompt "Please generate a recommended training menu for elderly people with severe shoulder stiffness." The inputs are the analysis results and the prompt, and the output is a training menu such as "Extend both arms up and slowly tilt them left and right" and a link to a reference video. The generative AI model suggests optimal training content based on the user's health status.

[0941] Step 5:

[0942] The server sends the generated training menu to the terminal. The input is the training menu and video link, and the output is the data to be sent to the terminal. The server transfers the text data of the training menu and the video link to the terminal.

[0943] Step 6:

[0944] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of a web or mobile application. The input is the data received from the server, and the output is the display on the interface that the user can visually confirm.

[0945] Step 7:

[0946] The user performs exercises according to the displayed training menu. Specifically, the user performs the exercises by following the instructions displayed on the screen: "Extend both arms upwards and slowly tilt them to the left and right." A reference video is played and the user performs the exercises while checking the movements. The input is the displayed training menu, and the output is the user performing the exercises with the correct form.

[0947] Step 8:

[0948] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine, which converts it back into text data. The input is voice data, and the output is text data.

[0949] Step 9:

[0950] The server analyzes the newly received text data and generates additional instructions or explanations. The input is new text data, and the output is additional instructions or explanations. For example, the server generates an additional instruction such as "Please refer to the following video" and sends it to the device.

[0951] Step 10:

[0952] The terminal displays newly received instructions and explanations to the user. The input is the additional instructions and explanations sent from the server, and the output is what is displayed on the user interface. The user continues training based on that.

[0953] (Application example 1)

[0954] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0955] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, or elsewhere. However, existing systems have problems in that users cannot acquire effective training menus in real time or cannot properly execute acquired menus. For example, the accuracy of voice input or voice recognition may be low, resulting in an inappropriate menu not being generated. Furthermore, the generated menu may be difficult for users to understand, resulting in the inability to perform exercises in the correct manner. Furthermore, specialized equipment or devices are required, which may be difficult for elderly people and users with physical challenges to implement. There is a need for a system that solves these problems and allows users to easily and effectively acquire and implement training menus.

[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0957] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition based on the voice input, means for generating an appropriate training menu based on the health condition, means for providing the generated training menu to the user, and means for displaying the generated training menu on a smart device, thereby enabling the user to obtain an effective training menu in real time at home, at a fitness facility, or the like, and to exercise in the correct way via the smart device.

[0958] The "means for acquiring voice input" is a device or method for acquiring the voice uttered by the user as digital data.

[0959] A "means for converting captured speech input to text" is a device or method for converting speech data into natural language text data.

[0960] The "means for analyzing a user's health condition based on a voice input" refers to a device or method for analyzing text data obtained from a voice input and determining a user's health condition.

[0961] The "means for generating an appropriate training menu based on health condition" refers to a device or method that automatically creates an appropriate exercise and stretching menu based on the analyzed health condition.

[0962] The "means for providing the generated training menu to the user" refers to a device or method that determines how to provide the generated training menu to the user and implements that method.

[0963] The "means for displaying the generated training menu on a smart device" refers to a device or method for visually displaying the generated training menu on a device such as a smartphone or smart glasses used by the user.

[0964] A "generative AI model" is a model that uses artificial intelligence techniques to generate new information from data.

[0965] A "prompt" is an input text given to a generative AI model, used to cause the model to generate specific instructions or information.

[0966] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, etc. The main components of this system are a voice input acquisition means, a means for converting the acquired voice input into text, a means for analyzing the user's health condition, a means for generating a training menu, a means for providing the generated training menu to the user, and a means for displaying the menu on a smart device.

[0967] Program processing

[0968] The whole system includes the following steps:

[0969] 1. Voice input acquisition: The user speaks to a smart device (e.g., smart glasses or smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is acquired using a microphone.

[0970] 2. Speech Recognition: The captured voice data is converted into text data using the Google Cloud Speech-to-Text API, which transcribes the user's speech into text.

[0971] 3. Text analysis: The server analyzes the text data using natural language processing (NLP) techniques (e.g., spaCy) to identify the user's health condition. For example, the text "My left knee hurts" is analyzed and it is recognized that the user has a knee problem.

[0972] 4. Training menu generation: The server generates an appropriate training menu using a generative AI model (e.g., OpenAI's GPT) based on the analysis results. An example of this prompt might be, "Please suggest an effective exercise menu for my left knee pain." The generated training menu includes specific exercise content and links to reference videos.

[0973] 5. Providing training menu: The generated training menu is sent to the smart device, where the user can check the instructions on the display of their smart glasses or smartphone.

[0974] 6. Display and execution of training: The generated exercise instructions are displayed on the display of the user device. For example, specific exercise instructions such as "Perform knee support stretches" are displayed, and a reference video can also be viewed.

[0975] Hardware and software used

[0976] Smart devices (e.g. smart glasses, smartphones): Used for voice input and display.

[0977] Cloud server: Used for voice recognition, text analysis, and training menu generation.

[0978] Speech recognition API (e.g., Google Cloud Speech-to-Text API): Used to convert voice data into text.

[0979] Natural language processing libraries (e.g. spaCy): Used to analyze text data.

[0980] Generative AI models (e.g., OpenAI's GPT): Used to generate training menus.

[0981] Specific examples

[0982] For example, if an elderly person says to a trainer at a fitness gym, "My left knee hurts. Please tell me what exercises I should do," the smart glasses will recognize the speech and convert it into text. The server then analyzes the text and determines that the person has pain in their left knee. An appropriate training menu is generated using a generative AI model, and the trainer is shown instructions such as, "Perform stretches to support your knee." The smart glasses also display a video showing how to perform the exercises correctly.

[0983] Prompt Sentence Examples

[0984] "Please suggest an effective exercise routine for my left knee pain."

[0985] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0986] Step 1:

[0987] Acquiring voice input

[0988] The user speaks into the microphone of their smart device (e.g., smart glasses, smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is picked up by the microphone on the smart device.

[0989] Input: User's voice

[0990] Output: Audio data

[0991] Step 2:

[0992] Converting audio data to text

[0993] The device sends the captured voice data to the Google Cloud Speech-to-Text API and converts it into text data. For example, a speech like "My left knee hurts, so please tell me what exercises I should do" is converted into the corresponding text.

[0994] Input: Audio data

[0995] Output: Text data

[0996] Step 3:

[0997] Text data analysis

[0998] The server receives the text data sent from the device and analyzes it using natural language processing (NLP) technology. For example, the keywords "knee" and "pain" may be extracted as analysis results, and the user's health condition may be identified as "pain in the left knee."

[0999] Input: Text data

[1000] Output: Health status analysis results

[1001] Step 4:

[1002] Training menu generation

[1003] Based on the analysis results, the server inputs prompts into a generative AI model (e.g., OpenAI's GPT) to generate an appropriate training menu. For example, a prompt such as "Please suggest an effective exercise menu for left knee pain" generates specific exercise content and links to reference videos.

[1004] Input: Analysis results regarding health status, prompt text

[1005] Output: Training menu

[1006] Step 5:

[1007] Sending and providing training menus

[1008] The server sends the generated training menu to the user's smart device, which then notifies the user of the menu. Specifically, the generated training menu is displayed on the display of smart glasses or a smartphone.

[1009] Input: Training Menu

[1010] Output: Training menu displayed on a smart device

[1011] Step 6:

[1012] Viewing and Running Training

[1013] The user device supports the user in exercising accurately based on the displayed training menu. For example, exercise instructions such as "Perform stretches to support your knees" are displayed on the smart glasses along with a reference video. The user exercises while watching the video.

[1014] Input: Training menu displayed on smart device

[1015] Output: User executes training

[1016] Each processing step allows the user to obtain an appropriate training menu in real time and perform exercises in the correct way through their smart device.

[1017] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1018] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[1019] Acquiring and converting voice input

[1020] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1021] The device receives the user's voice through a microphone, inputs this voice data into a speech recognition engine, and converts it into text data. The converted text data will contain the following content: "I've been suffering from severe shoulder stiffness lately and I'm feeling depressed. What kind of exercise is good?"

[1022] Health and Emotion Analysis

[1023] The terminal transmits the converted text data to the server.

[1024] The server analyzes the received text data using natural language processing (NLP) technology to identify the user's health condition, for example by extracting keywords such as "stiff shoulders" and "feeling depressed."

[1025] Furthermore, the server uses an emotion engine to analyze the user's emotions, for example, to identify the emotion "feeling depressed."

[1026] Creation and provision of training menus

[1027] The server generates an appropriate training menu based on the analysis results. Using the generation AI, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. This includes specific exercise instructions and links to reference videos. For example, instructions might include "extend both arms overhead and slowly tilt them back and forth three times."

[1028] The server selects a training menu based on the user's emotions and transmits it to the terminal. The selected menu is most effective for the user, for example, a menu that enhances relaxation.

[1029] Running the training

[1030] The terminal displays the training menu received from the server on the user interface.

[1031] The user looks at the displayed training menu and performs the exercises. For example, they perform the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. They can perform the exercises in the correct way by playing the reference video.

[1032] Follow-up

[1033] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they can ask, "Can you tell me more about the shoulder stretching exercise?"

[1034] The terminal converts the new voice data into text data again and transmits it to the server.

[1035] The server analyzes again and generates additional instructions or explanations, which are sent to the terminal.

[1036] The terminal displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[1037] In this way, the present invention is a system that enables elderly people and users with physical disabilities to easily and effectively train at home to maintain their health, and provides personalized and optimized instruction through emotion recognition. By providing specific training instructions and reference videos, it is expected that users will exercise in the correct way, contributing to maintaining their health and improving their quality of life.

[1038] The processing flow will be explained below.

[1039] Step 1:

[1040] The user talks to the system about their physical condition, worries, and feelings. For example, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1041] Step 2:

[1042] The terminal receives the user's voice through a microphone.

[1043] Step 3:

[1044] The device inputs the received voice data into a voice recognition engine and converts it into text data. The converted text data becomes, "I've been having really bad shoulder pain lately and it's making me feel depressed. What kind of exercise would be good?"

[1045] Step 4:

[1046] The terminal transmits the converted text data to the server.

[1047] Step 5:

[1048] The server receives the text data sent from the terminal.

[1049] Step 6:

[1050] The server uses natural language processing (NLP) technology to analyze the text data and identify the user's health condition. Specifically, it extracts and analyzes the keywords "stiff shoulders" and "feeling depressed."

[1051] Step 7:

[1052] The server then uses an emotion engine to analyze the emotions in the text data. From the part "feeling depressed," the server identifies the user's emotion as "depressed."

[1053] Step 8:

[1054] The server generates an appropriate training menu based on the analysis results. Using the AI ​​generation, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. The exercises include specific instructions such as "extend both arms overhead and slowly tilt them back and forth three times" and reference videos.

[1055] Step 9:

[1056] The server adjusts the training menu and delivery method based on the emotions recognized by the emotion engine. For example, for a user who is feeling depressed, it selects an exercise menu with a high relaxation effect.

[1057] Step 10:

[1058] The server transmits the generated training menu to the terminal. The adjusted training menu includes content that has a high relaxation effect.

[1059] Step 11:

[1060] The terminal displays the training menu received from the server on the user interface.

[1061] Step 12:

[1062] The user performs the exercises while viewing the displayed training menu. For example, the user performs the exercises according to the instructions for "shoulder stretching exercises." The user performs the exercises in the correct way while playing the reference video.

[1063] Step 13:

[1064] If the user has any questions or concerns during the training, they can input their voice again and the device will pick up the voice. For example, they might ask, "Can you tell me more about the shoulder stretching exercise?"

[1065] Step 14:

[1066] The terminal converts the newly received voice data back into text data and transmits it to the server.

[1067] Step 15:

[1068] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[1069] Step 16:

[1070] The terminal displays any additional instructions or explanations received to the user.

[1071] Example 2

[1072] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1073] It is difficult for elderly people and users with physical challenges to acquire and implement appropriate and effective training programs at home. Providing optimal training programs based on the user's current health condition and emotions is also insufficient. Conventional systems have been unable to recognize the user's emotions and provide individually optimized instruction.

[1074] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1075] In this invention, the server includes a means for acquiring voice input, a means for converting the acquired voice input into text, a means for analyzing the text data using natural language processing technology and emotion analysis technology to determine the health condition and emotions, a means for generating a training menu using a generative AI model, and a means for providing the generated menu to the user, thereby making it possible to provide an individually optimized training menu based on the user's health condition and emotions.

[1076] "Means for obtaining voice input" refers to devices or techniques for receiving voices spoken by a user to the system.

[1077] The "means for converting to text" refers to a speech recognition device or technology for converting acquired voice data into character data.

[1078] "Natural language processing technology" is a technology for analyzing text data and understanding and processing human language, and is used to identify a user's health condition and emotions.

[1079] "Emotion analysis technology" is a technology for analyzing and identifying emotions contained in text data.

[1080] A "generative AI model" is an artificial intelligence model that automatically generates outputs such as training menus based on input data.

[1081] The "means for generating a training menu" refers to a technique or device for creating an exercise program suited to a user based on the analyzed health condition and emotions.

[1082] "Means of providing" refers to the devices and technologies used to display and communicate the generated training menu to the user.

[1083] "Text data" is character data converted from voice input using voice recognition technology.

[1084] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[1085] Hardware and software used

[1086] This system uses the following hardware and software:

[1087] Device: A device used by a user, such as a smartphone or tablet, that includes a microphone, display, and internet connectivity.

[1088] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[1089] Natural language processing technology: Using the Google Cloud Natural Language API, health status and emotions are analyzed from text data.

[1090] Sentiment analysis technology: Uses IBM Watson Tone Analyzer to identify emotions from text data.

[1091] Generative AI model: Uses OpenAI GPT-4 to generate a training menu based on the analysis results.

[1092] System Operation

[1093] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1094] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API. The converted text data will contain the following content: "I've been having a lot of stiff shoulders lately and I'm feeling depressed. What kind of exercise is good?"

[1095] The device then sends the converted text data to a server, which analyzes it using the Google Cloud Natural Language API to identify the user's health condition. For example, it extracts keywords such as "stiff shoulders" and "feeling depressed." It also uses IBM Watson Tone Analyzer to identify the emotion of "feeling depressed."

[1096] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu tailored to the user's health condition and emotions. Specific training menus include "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. For example, instructions such as "extend both arms overhead and slowly tilt them back and forth three times" are included, along with links to reference videos.

[1097] The generated training menu is sent from the server to the terminal and displayed on the user interface. The user looks at the displayed training menu and performs the exercises. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. The user can perform the exercises in the correct way by playing the reference video.

[1098] If the user has any questions or concerns during the training, they can request more detailed explanations by voice input again. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes it again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[1099] Specific examples

[1100] For example, if a user uses the prompt "Tell me some exercises for stiff shoulders that I can do while sitting in a chair," the system will do the following:

[1101] User: Can you tell me some exercises I can do to relieve shoulder stiffness while sitting in a chair?

[1102] Device: Receives audio and converts it to text using the Google Cloud Speech-to-Text API.

[1103] Server: Receives text data and analyzes it using Google Cloud Natural Language API and IBM Watson Tone Analyzer.

[1104] Server: Sends prompts to the generative AI model (OpenAI GPT-4) to generate an appropriate training menu.

[1105] Device: The generated training menu is displayed. Specifically, the instructions are displayed such as "Perform three sets of stretching exercises that involve slowly rotating your shoulders while sitting in a chair."

[1106] This system allows users to easily and effectively train at home to maintain their health, and receive personalized guidance that is optimized by recognizing their emotions.

[1107] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1108] Step 1:

[1109] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1110] Input: User's voice data.

[1111] Output: Audio data received through the microphone.

[1112] Step 2:

[1113] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API.

[1114] Specific operation: The device analyzes the speech waveform and generates corresponding text using a language model.

[1115] Input: Received audio data.

[1116] Output: Text data: "I've been having really bad shoulder pain lately and it's been making me feel depressed. What kind of exercise would be good?"

[1117] Step 3:

[1118] The terminal transmits the converted text data to the server using the HTTPS protocol.

[1119] Specific operation: Sends text data to the server as an HTTP request.

[1120] Input: Text data.

[1121] Output: The text data sent to the server.

[1122] Step 4:

[1123] The server analyzes the received text data using the Google Cloud Natural Language API to determine the user's health status.

[1124] Specific operation: The server performs morphological analysis on the text data and extracts keywords such as "stiff shoulders" and "feeling depressed."

[1125] Input: The text data sent.

[1126] Output: Analysis results on health status and emotions (e.g., stiff shoulders, depression).

[1127] Step 5:

[1128] The server also uses IBM Watson Tone Analyzer to identify emotions from the text data.

[1129] Specific operation: The server performs sentiment analysis on the text data to identify emotions such as "depressed."

[1130] Input: Text data.

[1131] Output: Sentiment analysis result: "I feel depressed."

[1132] Step 6:

[1133] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu that corresponds to the user's health condition and emotions.

[1134] Specific operation: Input prompts into the generative AI model, and generate an appropriate training menu (e.g., shoulder stretching exercises, relaxation exercises) based on them.

[1135] Input: Analysis results (health status and emotions).

[1136] Output: The generated training menu.

[1137] Step 7:

[1138] The server sends the generated training menu to the terminal using the HTTPS protocol.

[1139] Specific operation: The generated training menu is sent to the terminal as an HTTP response.

[1140] Input: The generated training menu.

[1141] Output: The training menu sent to the device.

[1142] Step 8:

[1143] The terminal displays the training menu received from the server to the user.

[1144] Specific operation: Display a training menu (e.g., shoulder stretching exercises, relaxation exercise instructions, and reference video links) on the user interface.

[1145] Input: Training menu received from the server.

[1146] Output: The training menu displayed in the user interface.

[1147] Step 9:

[1148] The user performs exercises according to the displayed training menu, for example, following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen.

[1149] Specific movements: Watch the reference video and perform the training using the correct form.

[1150] Input: The displayed training menu and reference video.

[1151] Output: The training actions performed by the user.

[1152] Step 10:

[1153] If a user has a question during the exercise, they can ask for more detailed explanation by speaking again, for example, "Can you tell me more about the shoulder stretch exercise?"

[1154] Input: Voice input for questions or to ask for further clarification.

[1155] Output: Audio data received through the microphone.

[1156] Step 11:

[1157] The device then converts the new voice data back into text using the Google Cloud Speech-to-Text API and sends it to the server.

[1158] Specific operation: Analyzes the audio waveform, generates the corresponding text, and sends the text to the server.

[1159] Input: New audio data.

[1160] Output: Text data and data sent to the server.

[1161] Step 12:

[1162] The server again analyzes using the Google Cloud Natural Language API and IBM Watson Tone Analyzer to generate additional instructions and explanations.

[1163] Specific Actions: Health and emotional analysis is performed on the text data, and generative AI models are used to generate additional instructions and explanations (e.g., detailed instructions for shoulder stretching exercises).

[1164] Input: New text data.

[1165] Output: Additional instructions or explanations.

[1166] Step 13:

[1167] The terminal displays the newly received instructions and explanations to the user, who then continues training based on them.

[1168] What it does: Display new instructions or explanations in the user interface.

[1169] Input: Any additional instructions or explanations received from the server.

[1170] Output: Instructions or explanations displayed in the user interface.

[1171] (Application example 2)

[1172] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1173] The problem that this invention aims to solve is to provide a system that enables elderly people and users with physical disabilities to acquire and implement appropriate training menus at home or in a physical store.Furthermore, by recognizing the user's emotions and adjusting the training menu and delivery method, the system provides optimal instruction for each user, contributing to maintaining health and improving QOL.

[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1175] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition and emotions based on the voice input, means for generating an appropriate training menu based on the health condition and emotions, means for selecting and providing the generated training menu based on the user's emotions, and means for receiving feedback on the provided training menu and reanalyzing it, thereby enabling the user to acquire and appropriately perform optimal training in real time according to their current physical condition and emotions.

[1176] "Voice input" is voice data uttered by a user.

[1177] "Text data" refers to data obtained by converting voice input into text information.

[1178] "Health status" is information that indicates the physical and mental status of the user.

[1179] "Emotion" is information that indicates the user's mood or emotions.

[1180] A "training menu" is specific instructions for exercise and movement that are generated based on the user's health and mood.

[1181] "Smart terminals" are devices with advanced information processing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[1182] "Feedback" refers to comments or status reports provided by users during or after training.

[1183] "Real-time" means that there is no delay between data acquisition, processing, and providing results.

[1184] The "server" is an information processing device that analyzes voice data and text data and generates and provides training menus.

[1185] "Analysis" is a process of extracting information from voice and text data to understand the user's health condition and emotions.

[1186] A "generative AI model" is an algorithm or software that uses artificial intelligence to create new data such as training menus.

[1187] A "prompt sentence" is an input sentence given to a generative AI model, and serves as a hint for the AI ​​to generate new data based on that sentence.

[1188] This invention is a system that analyzes a user's health condition and emotions based on voice input, and generates and provides a training menu. The specific configuration of the system is as follows.

[1189] First, the device receives voice input from the user and converts it into text data. For voice recognition, smart devices such as smartphones with microphones, smart glasses, and head-mounted displays are used. For conversion, Google Speech Recognition API or other voice recognition software is used.

[1190] The acquired text data is then sent to a server, where it uses natural language processing (NLP) technology to analyze the user's health status and emotions. The NLP technology uses the Sentiment Analysis pipeline from the Transformers library. Based on the analyzed information, the server uses a generative AI model to generate an appropriate training menu. The generative AI model uses OpenAI's API.

[1191] The generated training menu is provided in real time, including text and reference videos. It is displayed on a smart device, and the user can follow it to carry out the training. Furthermore, the user can provide feedback during the training, which is also captured as voice input and analyzed again by the server. This allows the system to provide even more optimized training instructions.

[1192] Specific examples

[1193] As a concrete example, suppose a user wearing smart glasses at a fitness gym says, "My shoulder hurts today and I'm feeling a little unwell." This voice input is picked up through the smart glasses' microphone and converted into text by a speech recognition engine. The text data, "My shoulder hurts today and I'm feeling a little unwell," is sent to a server, where NLP technology and an emotion analysis engine analyze the user's health condition (shoulder pain) and emotion (unwell).

[1194] Based on the analysis results, OpenAI's generative AI model is used to create prompts and generate appropriate training menus. Examples of prompts include:

[1195] Example prompt sentence:

[1196] User condition: My shoulder hurts today and I'm feeling a bit unwell.

[1197] Emotion: Negative

[1198] So what to do?

[1199] The generated training menu might be, for example, "Start with some light stretching and then relax with some deep breathing exercises." This menu is displayed on the smart glasses' display, allowing the user to carry out the training.

[1200] This system allows users to obtain the optimal training in real time based on their physical condition and emotions at that moment, and then carry it out appropriately.

[1201] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1202] Step 1:

[1203] Acquiring voice input

[1204] The device picks up the user's voice through a microphone, and the user makes specific utterances such as "My shoulder hurts today, and I'm feeling a bit unwell."

[1205] Input: User's voice data

[1206] Output: Audio data file in the device

[1207] Step 2:

[1208] Converting audio data to text

[1209] The device converts the acquired voice data into text using voice recognition software such as Google Speech Recognition API. For example, the voice data "My shoulder hurts today and I'm feeling a bit unwell" is acquired as text data.

[1210] Input: Audio data file

[1211] Output: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[1212] Step 3:

[1213] Text data analysis

[1214] The server receives the text data sent from the device and analyzes the user's health status and emotions using natural language processing (NLP) techniques. Specifically, it uses the Sentiment Analysis pipeline in the Transformers library to analyze emotions and extract information about the health status from the text.

[1215] Input: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[1216] Output: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[1217] Step 4:

[1218] Training menu generation

[1219] The server generates an appropriate training menu based on the analysis results. It uses OpenAI's generative AI model to create prompts, and then generates a training menu based on those prompts. For example, it uses the prompt, "My shoulder hurts today, and I'm feeling a bit unwell. Emotion: Negative. So, what should I do?"

[1220] Input: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[1221] Output: Training menu (suggestions for light stretching and relaxation)

[1222] Step 5:

[1223] Providing training menus

[1224] The device displays the training menu received from the server on the user interface, and specific training content is displayed as text and reference video links on the display of the smart glasses or smartphone.

[1225] Input: Training Menu

[1226] Output: Training instructions displayed on the smart device

[1227] Step 6:

[1228] Get feedback

[1229] The user can provide additional instructions or give feedback by voice during the training, such as "Please tell me more about the shoulder stretching exercise."

[1230] Input: Audio data as feedback

[1231] Output: Audio data file in the device

[1232] Step 7:

[1233] Reanalyzing Feedback

[1234] The device then converts the newly acquired voice data back into text and sends it back to the server, which analyzes the feedback data, generates additional instructions and explanations, and sends them back to the device.

[1235] Input: Text data as feedback

[1236] Output: Additional training instructions or explanations

[1237] Step 8:

[1238] Providing additional instructions

[1239] The terminal displays the additional instructions and explanations received from the server on the user interface, and the user continues training based on them.

[1240] Input: Additional training instructions or explanations

[1241] Output: Additional instructions or explanations displayed on the smart device

[1242] This allows the entire process to proceed effectively in real time, allowing users to receive the training that best suits their physical and emotional state at that time.

[1243] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1244] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1245] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1246] [Fourth embodiment]

[1247] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1248] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1249] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1250] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1251] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1252] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1253] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1254] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1255] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1256] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1257] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1258] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1259] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1260] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[1261] Acquiring voice input

[1262] The user talks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[1263] Audio data conversion

[1264] The device inputs the received voice data into a voice recognition engine and converts it into text data, which generates the text "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[1265] Health status analysis

[1266] The server receives the text data sent from the device and analyzes the content using natural language processing (NLP) technology. For example, it identifies the user's health condition based on the keywords "stiff shoulders" and "severe."

[1267] Training menu generation

[1268] The server generates an appropriate training menu based on the results of the health analysis. Using the generation AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness. Specific exercise instructions include "stretch both arms up and slowly tilt them to the left and right," along with a link to a reference video demonstrating this exercise.

[1269] Providing training instructions

[1270] The server transmits the generated training menu to the terminal, which includes specific instructions including text and video links.

[1271] Viewing and Running Training

[1272] The device displays the training menu received from the server on the user interface. The user performs the exercises while looking at the displayed training menu. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" displayed on the screen. The user can perform the exercises in the correct way while playing a reference video.

[1273] Follow-up

[1274] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes the data again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[1275] In this way, the present invention is a system that allows elderly people and users with physical disabilities to easily and effectively train at home to maintain their health. By providing specific training instructions and reference videos, users can exercise in the correct way, which is expected to contribute to maintaining health and improving quality of life.

[1276] The processing flow will be explained below.

[1277] Step 1:

[1278] The user talks to the system about their physical condition and concerns. For example, "I have severe shoulder stiffness. Please tell me what kind of exercise I should do."

[1279] Step 2:

[1280] The terminal receives the user's voice through a microphone.

[1281] Step 3:

[1282] The device inputs the received voice data into a voice recognition engine and converts it into text data. For example, the generated text is "I have severe shoulder stiffness, so please tell me what kind of exercise I should do."

[1283] Step 4:

[1284] The terminal transmits the converted text data to the server.

[1285] Step 5:

[1286] The server receives the text data sent from the terminal.

[1287] Step 6:

[1288] The server analyzes the text data using natural language processing (NLP) technology, specifically extracting the keywords "stiff shoulders" and "severe" to determine the user's health condition.

[1289] Step 7:

[1290] The server generates an appropriate training menu based on the analysis of the user's health status. Using the AI, it suggests "shoulder stretching exercises" to alleviate shoulder stiffness.

[1291] Step 8:

[1292] The server then creates a training menu by documenting specific exercises and adding links to reference videos. For example, "Extend both arms overhead and slowly tilt them back and forth three times."

[1293] Step 9:

[1294] The server transmits the generated training menu to the terminal.

[1295] Step 10:

[1296] The terminal displays the training menu received from the server on the user interface.

[1297] Step 11:

[1298] The user performs exercises while viewing the displayed training menu. They can also play reference videos to check the exercise methods.

[1299] Step 12:

[1300] If the user has any questions or concerns during the training, they can input their voice again, and the device will pick up the voice.

[1301] Step 13:

[1302] The terminal converts the newly received voice data back into text data and transmits it to the server.

[1303] Step 14:

[1304] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[1305] Step 15:

[1306] The terminal displays any additional instructions or explanations received to the user.

[1307] Example 1

[1308] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1309] There is a need for a system that allows elderly people and users with physical challenges to easily obtain and implement appropriate training menus at home through voice input. However, existing systems have issues with the accuracy of speech recognition and text analysis, as well as the degree of customization of training menus, making it difficult for users to achieve the results they expect. For this reason, there is a need for a system that utilizes more advanced speech recognition, natural language processing, and generative AI technologies to provide appropriate training menus tailored to the user's health condition.

[1310] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1311] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text data, means for analyzing the user's health condition based on the text data using natural language processing technology, means for generating an appropriate training menu using a generative AI model based on the analysis results of the health condition, means for transmitting the generated training menu to a terminal and providing it to the user, and means for displaying and executing training based on the training menu, thereby enabling the user to easily and effectively perform health maintenance training at home.

[1312] "Means for obtaining voice input" refers to a device or a function of a device that has the function of capturing the voice spoken by a user to the system and recording it as digital voice data.

[1313] A "means for converting speech input into text data" is software or a system capable of analyzing recorded speech data and converting it into a corresponding text string.

[1314] "Natural language processing technology" refers to a series of algorithms and techniques for analyzing text data and understanding its content and intent.

[1315] "Means for analyzing a user's health condition" refers to software or a system that has the function of extracting and analyzing information about a user's physical condition and health from input text data using natural language processing technology.

[1316] A "generative AI model" is a model that uses artificial intelligence technology to generate appropriate responses and suggestions from input data.

[1317] "Means for generating a training menu" refers to software or a system that has the function of automatically creating an appropriate training menu based on the user's health condition analyzed using a generative AI model.

[1318] The "means for transmitting a training menu to a terminal" refers to software or a system that has the function of transferring the generated training menu as data to a terminal and providing it to the user.

[1319] "Means for displaying and executing training" refers to a device or software that has the function of displaying the training menu provided to the terminal on a user interface and allowing the user to perform the training according to the instructions.

[1320] "Link to reference video" refers to a URL or data that provides access to a video that shows specific methods for performing a training menu.

[1321] This invention is a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home. The system based on this invention is mainly composed of the following stages: voice input, voice recognition, text analysis, training menu generation, and training implementation. Specific embodiments of each stage are described below.

[1322] Acquiring voice input

[1323] The user talks to the system about their physical condition and concerns. For example, they might say something like, "I have really bad shoulder stiffness. Please tell me what kind of exercise I should do." In order to obtain clear voice data, it is important to speak in an environment that eliminates as much background noise as possible.

[1324] Audio data conversion

[1325] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specific voice recognition engines such as Google Cloud Speech-to-Text API are used. For example, voice input data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do" is converted into text.

[1326] Health status analysis

[1327] The server receives the text data sent from the device and analyzes the content using natural language processing technology. A natural language processing engine such as AWS Comprehend is used here. Specifically, keywords such as "stiff shoulders" and "severe" are identified from the text data, and the user's health condition is analyzed.

[1328] Training menu generation

[1329] The server uses a generative AI model based on the analysis results to generate an appropriate training menu. Specifically, it uses a generative AI model such as OpenAI's GPT-3. For example, a prompt such as "Please generate a recommended training menu for an elderly person with severe shoulder stiffness" can be input, and the server will suggest "shoulder stretching exercises" based on that. The generated content includes instructions such as "Extend both arms up and slowly tilt them to the left and right" as well as links to reference videos.

[1330] Providing training instructions

[1331] The server sends the generated training menu to the device, which contains text instructions and links to reference videos.

[1332] Viewing and Running Training

[1333] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of the web or mobile application. The user performs the training by following the "shoulder stretch exercise" instructions displayed on the screen. The user can perform the exercises with the correct form while playing the reference video.

[1334] Follow-up

[1335] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine again and converts it into text data. The server then analyzes the text data and generates additional instructions or explanations. For example, it might generate additional instructions such as, "Please refer to the following video for shoulder stretching exercises," and send the video link to the device. The device then displays the newly received instructions or explanations to the user, who can then continue their training based on them.

[1336] Examples of specific examples and prompts

[1337] Example: A 70-year-old elderly person uses this system for stiff shoulders. He / she speaks to the system, saying, "I have severe shoulder stiffness, so please tell me what kind of exercises I should do." The system then suggests shoulder stretching exercises and provides a link to a video.

[1338] Example prompt sentence:

[1339] "Generate a recommended training menu for elderly people with severe shoulder stiffness."

[1340] "Please provide exercise instructions and a link to a helpful video to help relieve shoulder stiffness."

[1341] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1342] Step 1:

[1343] The user speaks to the system about their physical condition and concerns. For example, they might say, "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." The input is the user's voice, and the output is the voice data received by the device.

[1344] Step 2:

[1345] The device inputs the received voice data into a voice recognition engine and converts it into text data. Specifically, it uses the Google Cloud Speech-to-Text API. The input is voice data, and the output is text data such as "I have severe shoulder stiffness, so please tell me what kind of exercise I should do." Here, the voice recognition engine analyzes the phonemes of the voice and generates a corresponding string of characters.

[1346] Step 3:

[1347] The server receives the text data sent from the device and analyzes it using natural language processing technology. Specifically, AWS Comprehend is used to extract keywords from the text data and identify the health condition. The input is text data, and the output is analysis results such as "stiff shoulders" or "severe." Natural language processing technology extracts keywords and performs sentiment analysis.

[1348] Step 4:

[1349] The server uses a generative AI model to generate an appropriate training menu based on the health status analysis results. Specifically, it uses OpenAI's GPT-3 and inputs the prompt "Please generate a recommended training menu for elderly people with severe shoulder stiffness." The inputs are the analysis results and the prompt, and the output is a training menu such as "Extend both arms up and slowly tilt them left and right" and a link to a reference video. The generative AI model suggests optimal training content based on the user's health status.

[1350] Step 5:

[1351] The server sends the generated training menu to the terminal. The input is the training menu and video link, and the output is the data to be sent to the terminal. The server transfers the text data of the training menu and the video link to the terminal.

[1352] Step 6:

[1353] The device displays the received training menu on the user interface. Specifically, it displays text instructions and video links on the screen of a web or mobile application. The input is the data received from the server, and the output is the display on the interface that the user can visually confirm.

[1354] Step 7:

[1355] The user performs exercises according to the displayed training menu. Specifically, the user performs the exercises by following the instructions displayed on the screen: "Extend both arms upwards and slowly tilt them to the left and right." A reference video is played and the user performs the exercises while checking the movements. The input is the displayed training menu, and the output is the user performing the exercises with the correct form.

[1356] Step 8:

[1357] If the user has any questions or concerns during the training, they can input the voice again. For example, they might ask, "Please tell me more about shoulder stretching exercises." The device then passes the new voice data to the voice recognition engine, which converts it back into text data. The input is voice data, and the output is text data.

[1358] Step 9:

[1359] The server analyzes the newly received text data and generates additional instructions or explanations. The input is new text data, and the output is additional instructions or explanations. For example, the server generates an additional instruction such as "Please refer to the following video" and sends it to the device.

[1360] Step 10:

[1361] The terminal displays newly received instructions and explanations to the user. The input is the additional instructions and explanations sent from the server, and the output is what is displayed on the user interface. The user continues training based on that.

[1362] (Application example 1)

[1363] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1364] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, or elsewhere. However, existing systems have problems in that users cannot acquire effective training menus in real time or cannot properly execute acquired menus. For example, the accuracy of voice input or voice recognition may be low, resulting in an inappropriate menu not being generated. Furthermore, the generated menu may be difficult for users to understand, resulting in the inability to perform exercises in the correct manner. Furthermore, specialized equipment or devices are required, which may be difficult for elderly people and users with physical challenges to implement. There is a need for a system that solves these problems and allows users to easily and effectively acquire and implement training menus.

[1365] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1366] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition based on the voice input, means for generating an appropriate training menu based on the health condition, means for providing the generated training menu to the user, and means for displaying the generated training menu on a smart device, thereby enabling the user to obtain an effective training menu in real time at home, at a fitness facility, or the like, and to exercise in the correct way via the smart device.

[1367] The "means for acquiring voice input" is a device or method for acquiring the voice uttered by the user as digital data.

[1368] A "means for converting captured speech input to text" is a device or method for converting speech data into natural language text data.

[1369] The "means for analyzing a user's health condition based on a voice input" refers to a device or method for analyzing text data obtained from a voice input and determining a user's health condition.

[1370] The "means for generating an appropriate training menu based on health condition" refers to a device or method that automatically creates an appropriate exercise and stretching menu based on the analyzed health condition.

[1371] The "means for providing the generated training menu to the user" refers to a device or method that determines how to provide the generated training menu to the user and implements that method.

[1372] The "means for displaying the generated training menu on a smart device" refers to a device or method for visually displaying the generated training menu on a device such as a smartphone or smart glasses used by the user.

[1373] A "generative AI model" is a model that uses artificial intelligence techniques to generate new information from data.

[1374] A "prompt" is an input text given to a generative AI model, used to cause the model to generate specific instructions or information.

[1375] The present invention relates to a system that enables elderly people and users with physical challenges to acquire and implement appropriate training menus at home, at fitness facilities, etc. The main components of this system are a voice input acquisition means, a means for converting the acquired voice input into text, a means for analyzing the user's health condition, a means for generating a training menu, a means for providing the generated training menu to the user, and a means for displaying the menu on a smart device.

[1376] Program processing

[1377] The whole system includes the following steps:

[1378] 1. Voice input acquisition: The user speaks to a smart device (e.g., smart glasses or smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is acquired using a microphone.

[1379] 2. Speech Recognition: The captured voice data is converted into text data using the Google Cloud Speech-to-Text API, which transcribes the user's speech into text.

[1380] 3. Text analysis: The server analyzes the text data using natural language processing (NLP) techniques (e.g., spaCy) to identify the user's health condition. For example, the text "My left knee hurts" is analyzed and it is recognized that the user has a knee problem.

[1381] 4. Training menu generation: The server generates an appropriate training menu using a generative AI model (e.g., OpenAI's GPT) based on the analysis results. An example of this prompt might be, "Please suggest an effective exercise menu for my left knee pain." The generated training menu includes specific exercise content and links to reference videos.

[1382] 5. Providing training menu: The generated training menu is sent to the smart device, where the user can check the instructions on the display of their smart glasses or smartphone.

[1383] 6. Display and execution of training: The generated exercise instructions are displayed on the display of the user device. For example, specific exercise instructions such as "Perform knee support stretches" are displayed, and a reference video can also be viewed.

[1384] Hardware and software used

[1385] Smart devices (e.g. smart glasses, smartphones): Used for voice input and display.

[1386] Cloud server: Used for voice recognition, text analysis, and training menu generation.

[1387] Speech recognition API (e.g., Google Cloud Speech-to-Text API): Used to convert voice data into text.

[1388] Natural language processing libraries (e.g. spaCy): Used to analyze text data.

[1389] Generative AI models (e.g., OpenAI's GPT): Used to generate training menus.

[1390] Specific examples

[1391] For example, if an elderly person says to a trainer at a fitness gym, "My left knee hurts. Please tell me what exercises I should do," the smart glasses will recognize the speech and convert it into text. The server then analyzes the text and determines that the person has pain in their left knee. An appropriate training menu is generated using a generative AI model, and the trainer is shown instructions such as, "Perform stretches to support your knee." The smart glasses also display a video showing how to perform the exercises correctly.

[1392] Prompt Sentence Examples

[1393] "Please suggest an effective exercise routine for my left knee pain."

[1394] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1395] Step 1:

[1396] Acquiring voice input

[1397] The user speaks into the microphone of their smart device (e.g., smart glasses, smartphone) about their physical condition and areas they would like to improve. For example, they might say, "My left knee hurts, so please tell me what exercises I should do." This voice data is picked up by the microphone on the smart device.

[1398] Input: User's voice

[1399] Output: Audio data

[1400] Step 2:

[1401] Converting audio data to text

[1402] The device sends the captured voice data to the Google Cloud Speech-to-Text API and converts it into text data. For example, a speech like "My left knee hurts, so please tell me what exercises I should do" is converted into the corresponding text.

[1403] Input: Audio data

[1404] Output: Text data

[1405] Step 3:

[1406] Text data analysis

[1407] The server receives the text data sent from the device and analyzes it using natural language processing (NLP) technology. For example, the keywords "knee" and "pain" may be extracted as analysis results, and the user's health condition may be identified as "pain in the left knee."

[1408] Input: Text data

[1409] Output: Health status analysis results

[1410] Step 4:

[1411] Training menu generation

[1412] Based on the analysis results, the server inputs prompts into a generative AI model (e.g., OpenAI's GPT) to generate an appropriate training menu. For example, a prompt such as "Please suggest an effective exercise menu for left knee pain" generates specific exercise content and links to reference videos.

[1413] Input: Analysis results regarding health status, prompt text

[1414] Output: Training menu

[1415] Step 5:

[1416] Sending and providing training menus

[1417] The server sends the generated training menu to the user's smart device, which then notifies the user of the menu. Specifically, the generated training menu is displayed on the display of smart glasses or a smartphone.

[1418] Input: Training Menu

[1419] Output: Training menu displayed on a smart device

[1420] Step 6:

[1421] Viewing and Running Training

[1422] The user device supports the user in exercising accurately based on the displayed training menu. For example, exercise instructions such as "Perform stretches to support your knees" are displayed on the smart glasses along with a reference video. The user exercises while watching the video.

[1423] Input: Training menu displayed on smart device

[1424] Output: User executes training

[1425] Each processing step allows the user to obtain an appropriate training menu in real time and perform exercises in the correct way through their smart device.

[1426] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1427] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[1428] Acquiring and converting voice input

[1429] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1430] The device receives the user's voice through a microphone, inputs this voice data into a speech recognition engine, and converts it into text data. The converted text data will contain the following content: "I've been suffering from severe shoulder stiffness lately and I'm feeling depressed. What kind of exercise is good?"

[1431] Health and Emotion Analysis

[1432] The terminal transmits the converted text data to the server.

[1433] The server analyzes the received text data using natural language processing (NLP) technology to identify the user's health condition, for example by extracting keywords such as "stiff shoulders" and "feeling depressed."

[1434] Furthermore, the server uses an emotion engine to analyze the user's emotions, for example, to identify the emotion "feeling depressed."

[1435] Creation and provision of training menus

[1436] The server generates an appropriate training menu based on the analysis results. Using the generation AI, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. This includes specific exercise instructions and links to reference videos. For example, instructions might include "extend both arms overhead and slowly tilt them back and forth three times."

[1437] The server selects a training menu based on the user's emotions and transmits it to the terminal. The selected menu is most effective for the user, for example, a menu that enhances relaxation.

[1438] Running the training

[1439] The terminal displays the training menu received from the server on the user interface.

[1440] The user looks at the displayed training menu and performs the exercises. For example, they perform the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. They can perform the exercises in the correct way by playing the reference video.

[1441] Follow-up

[1442] If the user has any questions or concerns during the training, they can input the information again by voice. For example, they can ask, "Can you tell me more about the shoulder stretching exercise?"

[1443] The terminal converts the new voice data into text data again and transmits it to the server.

[1444] The server analyzes again and generates additional instructions or explanations, which are sent to the terminal.

[1445] The terminal displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[1446] In this way, the present invention is a system that enables elderly people and users with physical disabilities to easily and effectively train at home to maintain their health, and provides personalized and optimized instruction through emotion recognition. By providing specific training instructions and reference videos, it is expected that users will exercise in the correct way, contributing to maintaining their health and improving their quality of life.

[1447] The processing flow will be explained below.

[1448] Step 1:

[1449] The user talks to the system about their physical condition, worries, and feelings. For example, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1450] Step 2:

[1451] The terminal receives the user's voice through a microphone.

[1452] Step 3:

[1453] The device inputs the received voice data into a voice recognition engine and converts it into text data. The converted text data becomes, "I've been having really bad shoulder pain lately and it's making me feel depressed. What kind of exercise would be good?"

[1454] Step 4:

[1455] The terminal transmits the converted text data to the server.

[1456] Step 5:

[1457] The server receives the text data sent from the terminal.

[1458] Step 6:

[1459] The server uses natural language processing (NLP) technology to analyze the text data and identify the user's health condition. Specifically, it extracts and analyzes the keywords "stiff shoulders" and "feeling depressed."

[1460] Step 7:

[1461] The server then uses an emotion engine to analyze the emotions in the text data. From the part "feeling depressed," the server identifies the user's emotion as "depressed."

[1462] Step 8:

[1463] The server generates an appropriate training menu based on the analysis results. Using the AI ​​generation, it suggests "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. The exercises include specific instructions such as "extend both arms overhead and slowly tilt them back and forth three times" and reference videos.

[1464] Step 9:

[1465] The server adjusts the training menu and delivery method based on the emotions recognized by the emotion engine. For example, for a user who is feeling depressed, it selects an exercise menu with a high relaxation effect.

[1466] Step 10:

[1467] The server transmits the generated training menu to the terminal. The adjusted training menu includes content that has a high relaxation effect.

[1468] Step 11:

[1469] The terminal displays the training menu received from the server on the user interface.

[1470] Step 12:

[1471] The user performs the exercises while viewing the displayed training menu. For example, the user performs the exercises according to the instructions for "shoulder stretching exercises." The user performs the exercises in the correct way while playing the reference video.

[1472] Step 13:

[1473] If the user has any questions or concerns during the training, they can input their voice again and the device will pick up the voice. For example, they might ask, "Can you tell me more about the shoulder stretching exercise?"

[1474] Step 14:

[1475] The terminal converts the newly received voice data back into text data and transmits it to the server.

[1476] Step 15:

[1477] The server performs the analysis again, generates answers to any unclear points and additional instructions, and sends them to the terminal.

[1478] Step 16:

[1479] The terminal displays any additional instructions or explanations received to the user.

[1480] Example 2

[1481] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1482] It is difficult for elderly people and users with physical challenges to acquire and implement appropriate and effective training programs at home. Providing optimal training programs based on the user's current health condition and emotions is also insufficient. Conventional systems have been unable to recognize the user's emotions and provide individually optimized instruction.

[1483] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1484] In this invention, the server includes a means for acquiring voice input, a means for converting the acquired voice input into text, a means for analyzing the text data using natural language processing technology and emotion analysis technology to determine the health condition and emotions, a means for generating a training menu using a generative AI model, and a means for providing the generated menu to the user, thereby making it possible to provide an individually optimized training menu based on the user's health condition and emotions.

[1485] "Means for obtaining voice input" refers to devices or techniques for receiving voices spoken by a user to the system.

[1486] The "means for converting to text" refers to a speech recognition device or technology for converting acquired voice data into character data.

[1487] "Natural language processing technology" is a technology for analyzing text data and understanding and processing human language, and is used to identify a user's health condition and emotions.

[1488] "Emotion analysis technology" is a technology for analyzing and identifying emotions contained in text data.

[1489] A "generative AI model" is an artificial intelligence model that automatically generates outputs such as training menus based on input data.

[1490] The "means for generating a training menu" refers to a technique or device for creating an exercise program suited to a user based on the analyzed health condition and emotions.

[1491] "Means of providing" refers to the devices and technologies used to display and communicate the generated training menu to the user.

[1492] "Text data" is character data converted from voice input using voice recognition technology.

[1493] This invention relates to a system that allows elderly people and users with physical disabilities to acquire and implement appropriate training menus at home. Furthermore, it features an emotion engine that recognizes the user's emotions and adjusts the training menu and delivery method accordingly.

[1494] Hardware and software used

[1495] This system uses the following hardware and software:

[1496] Device: A device used by a user, such as a smartphone or tablet, that includes a microphone, display, and internet connectivity.

[1497] Speech recognition engine: Uses the Google Cloud Speech-to-Text API to convert voice data into text data.

[1498] Natural language processing technology: Using the Google Cloud Natural Language API, health status and emotions are analyzed from text data.

[1499] Sentiment analysis technology: Uses IBM Watson Tone Analyzer to identify emotions from text data.

[1500] Generative AI model: Uses OpenAI GPT-4 to generate a training menu based on the analysis results.

[1501] System Operation

[1502] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1503] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API. The converted text data will contain the following content: "I've been having a lot of stiff shoulders lately and I'm feeling depressed. What kind of exercise is good?"

[1504] The device then sends the converted text data to a server, which analyzes it using the Google Cloud Natural Language API to identify the user's health condition. For example, it extracts keywords such as "stiff shoulders" and "feeling depressed." It also uses IBM Watson Tone Analyzer to identify the emotion of "feeling depressed."

[1505] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu tailored to the user's health condition and emotions. Specific training menus include "shoulder stretching exercises" to relieve shoulder stiffness and "relaxation exercises" to improve mood. For example, instructions such as "extend both arms overhead and slowly tilt them back and forth three times" are included, along with links to reference videos.

[1506] The generated training menu is sent from the server to the terminal and displayed on the user interface. The user looks at the displayed training menu and performs the exercises. For example, the user performs the exercises by following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen. The user can perform the exercises in the correct way by playing the reference video.

[1507] If the user has any questions or concerns during the training, they can request more detailed explanations by voice input again. For example, they might ask, "Could you please explain in more detail the shoulder stretching exercise?" The device then converts the new voice data into text data again and sends it to the server. The server then analyzes it again, generates additional instructions and explanations, and sends them to the device. The device then displays the newly received instructions and explanations to the user, who can then continue their training based on them.

[1508] Specific examples

[1509] For example, if a user uses the prompt "Tell me some exercises for stiff shoulders that I can do while sitting in a chair," the system will do the following:

[1510] User: Can you tell me some exercises I can do to relieve shoulder stiffness while sitting in a chair?

[1511] Device: Receives audio and converts it to text using the Google Cloud Speech-to-Text API.

[1512] Server: Receives text data and analyzes it using Google Cloud Natural Language API and IBM Watson Tone Analyzer.

[1513] Server: Sends prompts to the generative AI model (OpenAI GPT-4) to generate an appropriate training menu.

[1514] Device: The generated training menu is displayed. Specifically, the instructions are displayed such as "Perform three sets of stretching exercises that involve slowly rotating your shoulders while sitting in a chair."

[1515] This system allows users to easily and effectively train at home to maintain their health, and receive personalized guidance that is optimized by recognizing their emotions.

[1516] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1517] Step 1:

[1518] The user talks to the system about their physical condition, worries, and emotions. For example, they might say, "I've been having really bad shoulder pain lately and I'm feeling depressed. What kind of exercise would be good?"

[1519] Input: User's voice data.

[1520] Output: Audio data received through the microphone.

[1521] Step 2:

[1522] The device receives the user's voice through a microphone and converts this voice data into text data using the Google Cloud Speech-to-Text API.

[1523] Specific operation: The device analyzes the speech waveform and generates corresponding text using a language model.

[1524] Input: Received audio data.

[1525] Output: Text data: "I've been having really bad shoulder pain lately and it's been making me feel depressed. What kind of exercise would be good?"

[1526] Step 3:

[1527] The terminal transmits the converted text data to the server using the HTTPS protocol.

[1528] Specific operation: Sends text data to the server as an HTTP request.

[1529] Input: Text data.

[1530] Output: The text data sent to the server.

[1531] Step 4:

[1532] The server analyzes the received text data using the Google Cloud Natural Language API to determine the user's health status.

[1533] Specific operation: The server performs morphological analysis on the text data and extracts keywords such as "stiff shoulders" and "feeling depressed."

[1534] Input: The text data sent.

[1535] Output: Analysis results on health status and emotions (e.g., stiff shoulders, depression).

[1536] Step 5:

[1537] The server also uses IBM Watson Tone Analyzer to identify emotions from the text data.

[1538] Specific operation: The server performs sentiment analysis on the text data to identify emotions such as "depressed."

[1539] Input: Text data.

[1540] Output: Sentiment analysis result: "I feel depressed."

[1541] Step 6:

[1542] Based on the analysis results, the server uses a generative AI model (OpenAI GPT-4) to generate a training menu that corresponds to the user's health condition and emotions.

[1543] Specific operation: Input prompts into the generative AI model, and generate an appropriate training menu (e.g., shoulder stretching exercises, relaxation exercises) based on them.

[1544] Input: Analysis results (health status and emotions).

[1545] Output: The generated training menu.

[1546] Step 7:

[1547] The server sends the generated training menu to the terminal using the HTTPS protocol.

[1548] Specific operation: The generated training menu is sent to the terminal as an HTTP response.

[1549] Input: The generated training menu.

[1550] Output: The training menu sent to the device.

[1551] Step 8:

[1552] The terminal displays the training menu received from the server to the user.

[1553] Specific operation: Display a training menu (e.g., shoulder stretching exercises, relaxation exercise instructions, and reference video links) on the user interface.

[1554] Input: Training menu received from the server.

[1555] Output: The training menu displayed in the user interface.

[1556] Step 9:

[1557] The user performs exercises according to the displayed training menu, for example, following the instructions for "shoulder stretching exercises" and "relaxation exercises" displayed on the screen.

[1558] Specific movements: Watch the reference video and perform the training using the correct form.

[1559] Input: The displayed training menu and reference video.

[1560] Output: The training actions performed by the user.

[1561] Step 10:

[1562] If a user has a question during the exercise, they can ask for more detailed explanation by speaking again, for example, "Can you tell me more about the shoulder stretch exercise?"

[1563] Input: Voice input for questions or to ask for further clarification.

[1564] Output: Audio data received through the microphone.

[1565] Step 11:

[1566] The device then converts the new voice data back into text using the Google Cloud Speech-to-Text API and sends it to the server.

[1567] Specific operation: Analyzes the audio waveform, generates the corresponding text, and sends the text to the server.

[1568] Input: New audio data.

[1569] Output: Text data and data sent to the server.

[1570] Step 12:

[1571] The server again analyzes using the Google Cloud Natural Language API and IBM Watson Tone Analyzer to generate additional instructions and explanations.

[1572] Specific Actions: Health and emotional analysis is performed on the text data, and generative AI models are used to generate additional instructions and explanations (e.g., detailed instructions for shoulder stretching exercises).

[1573] Input: New text data.

[1574] Output: Additional instructions or explanations.

[1575] Step 13:

[1576] The terminal displays the newly received instructions and explanations to the user, who then continues training based on them.

[1577] What it does: Display new instructions or explanations in the user interface.

[1578] Input: Any additional instructions or explanations received from the server.

[1579] Output: Instructions or explanations displayed in the user interface.

[1580] (Application example 2)

[1581] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1582] The problem that this invention aims to solve is to provide a system that enables elderly people and users with physical disabilities to acquire and implement appropriate training menus at home or in a physical store.Furthermore, by recognizing the user's emotions and adjusting the training menu and delivery method, the system provides optimal instruction for each user, contributing to maintaining health and improving QOL.

[1583] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1584] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice input into text, means for analyzing the user's health condition and emotions based on the voice input, means for generating an appropriate training menu based on the health condition and emotions, means for selecting and providing the generated training menu based on the user's emotions, and means for receiving feedback on the provided training menu and reanalyzing it, thereby enabling the user to acquire and appropriately perform optimal training in real time according to their current physical condition and emotions.

[1585] "Voice input" is voice data uttered by a user.

[1586] "Text data" refers to data obtained by converting voice input into text information.

[1587] "Health status" is information that indicates the physical and mental status of the user.

[1588] "Emotion" is information that indicates the user's mood or emotions.

[1589] A "training menu" is specific instructions for exercise and movement that are generated based on the user's health and mood.

[1590] "Smart terminals" are devices with advanced information processing capabilities, such as smartphones, smart glasses, and head-mounted displays.

[1591] "Feedback" refers to comments or status reports provided by users during or after training.

[1592] "Real-time" means that there is no delay between data acquisition, processing, and providing results.

[1593] The "server" is an information processing device that analyzes voice data and text data and generates and provides training menus.

[1594] "Analysis" is a process of extracting information from voice and text data to understand the user's health condition and emotions.

[1595] A "generative AI model" is an algorithm or software that uses artificial intelligence to create new data such as training menus.

[1596] A "prompt sentence" is an input sentence given to a generative AI model, and serves as a hint for the AI ​​to generate new data based on that sentence.

[1597] This invention is a system that analyzes a user's health condition and emotions based on voice input, and generates and provides a training menu. The specific configuration of the system is as follows.

[1598] First, the device receives voice input from the user and converts it into text data. For voice recognition, smart devices such as smartphones with microphones, smart glasses, and head-mounted displays are used. For conversion, Google Speech Recognition API or other voice recognition software is used.

[1599] The acquired text data is then sent to a server, where it uses natural language processing (NLP) technology to analyze the user's health status and emotions. The NLP technology uses the Sentiment Analysis pipeline from the Transformers library. Based on the analyzed information, the server uses a generative AI model to generate an appropriate training menu. The generative AI model uses OpenAI's API.

[1600] The generated training menu is provided in real time, including text and reference videos. It is displayed on a smart device, and the user can follow it to carry out the training. Furthermore, the user can provide feedback during the training, which is also captured as voice input and analyzed again by the server. This allows the system to provide even more optimized training instructions.

[1601] Specific examples

[1602] As a concrete example, suppose a user wearing smart glasses at a fitness gym says, "My shoulder hurts today and I'm feeling a little unwell." This voice input is picked up through the smart glasses' microphone and converted into text by a speech recognition engine. The text data, "My shoulder hurts today and I'm feeling a little unwell," is sent to a server, where NLP technology and an emotion analysis engine analyze the user's health condition (shoulder pain) and emotion (unwell).

[1603] Based on the analysis results, OpenAI's generative AI model is used to create prompts and generate appropriate training menus. Examples of prompts include:

[1604] Example prompt sentence:

[1605] User condition: My shoulder hurts today and I'm feeling a bit unwell.

[1606] Emotion: Negative

[1607] So what to do?

[1608] The generated training menu might be, for example, "Start with some light stretching and then relax with some deep breathing exercises." This menu is displayed on the smart glasses' display, allowing the user to carry out the training.

[1609] This system allows users to obtain the optimal training in real time based on their physical condition and emotions at that moment, and then carry it out appropriately.

[1610] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1611] Step 1:

[1612] Acquiring voice input

[1613] The device picks up the user's voice through a microphone, and the user makes specific utterances such as "My shoulder hurts today, and I'm feeling a bit unwell."

[1614] Input: User's voice data

[1615] Output: Audio data file in the device

[1616] Step 2:

[1617] Converting audio data to text

[1618] The device converts the acquired voice data into text using voice recognition software such as Google Speech Recognition API. For example, the voice data "My shoulder hurts today and I'm feeling a bit unwell" is acquired as text data.

[1619] Input: Audio data file

[1620] Output: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[1621] Step 3:

[1622] Text data analysis

[1623] The server receives the text data sent from the device and analyzes the user's health status and emotions using natural language processing (NLP) techniques. Specifically, it uses the Sentiment Analysis pipeline in the Transformers library to analyze emotions and extract information about the health status from the text.

[1624] Input: Text data "My shoulder hurts today and I'm feeling a bit unwell."

[1625] Output: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[1626] Step 4:

[1627] Training menu generation

[1628] The server generates an appropriate training menu based on the analysis results. It uses OpenAI's generative AI model to create prompts, and then generates a training menu based on those prompts. For example, it uses the prompt, "My shoulder hurts today, and I'm feeling a bit unwell. Emotion: Negative. So, what should I do?"

[1629] Input: Analysis results of health condition (shoulder pain) and emotion (feeling unwell)

[1630] Output: Training menu (suggestions for light stretching and relaxation)

[1631] Step 5:

[1632] Providing training menus

[1633] The device displays the training menu received from the server on the user interface, and specific training content is displayed as text and reference video links on the display of the smart glasses or smartphone.

[1634] Input: Training Menu

[1635] Output: Training instructions displayed on the smart device

[1636] Step 6:

[1637] Get feedback

[1638] The user can provide additional instructions or give feedback by voice during the training, such as "Please tell me more about the shoulder stretching exercise."

[1639] Input: Audio data as feedback

[1640] Output: Audio data file in the device

[1641] Step 7:

[1642] Reanalyzing Feedback

[1643] The device then converts the newly acquired voice data back into text and sends it back to the server, which analyzes the feedback data, generates additional instructions and explanations, and sends them back to the device.

[1644] Input: Text data as feedback

[1645] Output: Additional training instructions or explanations

[1646] Step 8:

[1647] Providing additional instructions

[1648] The terminal displays the additional instructions and explanations received from the server on the user interface, and the user continues training based on them.

[1649] Input: Additional training instructions or explanations

[1650] Output: Additional instructions or explanations displayed on the smart device

[1651] This allows the entire process to proceed effectively in real time, allowing users to receive the training that best suits their physical and emotional state at that time.

[1652] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1653] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1654] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1655] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1656] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1657] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1658] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1659] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1660] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1661] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1662] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1663] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1664] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1665] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1666] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1667] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1668] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1669] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1670] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1671] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1672] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1673] The following is further disclosed regarding the above embodiment.

[1674] (Claim 1)

[1675] a means for obtaining a voice input;

[1676] means for converting the captured speech input into text;

[1677] means for analyzing the user's health condition based on the voice input;

[1678] a means for generating an appropriate training menu based on the health condition;

[1679] means for providing the generated training menu to a user;

[1680] A system including:

[1681] (Claim 2)

[1682] The system of claim 1 , wherein the generated training menu includes text and reference videos.

[1683] (Claim 3)

[1684] The system of claim 1 further includes a software program that performs a series of processes: acquiring the voice input, converting the acquired voice data into text data, analyzing the text data to determine health status, generating a training menu, and providing the generated menu to the user.

[1685] "Example 1"

[1686] (Claim 1)

[1687] a means for obtaining a voice input;

[1688] means for converting the acquired voice input into text data;

[1689] means for analyzing the user's health condition based on the text data using natural language processing technology;

[1690] A means for generating an appropriate training menu using a generation AI model based on the analysis results of the health condition;

[1691] means for transmitting the generated training menu to a terminal and providing it to a user;

[1692] means for displaying and executing training based on the training menu;

[1693] A system including:

[1694] (Claim 2)

[1695] The system of claim 1, wherein the generated training menu includes text and links to reference videos.

[1696] (Claim 3)

[1697] The system of claim 1 further includes a software program that performs a series of processes: acquiring the voice input, converting the acquired voice data into text data, analyzing the text data to determine the health condition, generating a training menu using a generative AI model, and transmitting the generated menu to a terminal to provide it to the user.

[1698] "Application Example 1"

[1699] (Claim 1)

[1700] a means for obtaining a voice input;

[1701] means for converting the captured speech input into text;

[1702] means for analyzing the user's health condition based on the voice input;

[1703] a means for generating an appropriate training menu based on the health condition;

[1704] means for providing the generated training menu to a user;

[1705] a means for displaying the generated training menu on a smart device;

[1706] A system including:

[1707] (Claim 2)

[1708] The system of claim 1 , wherein the generated training menu includes text and reference videos.

[1709] (Claim 3)

[1710] The system of claim 1 further includes a software program that performs a series of processes: acquiring the voice input, converting the acquired voice data into text data, analyzing the text data to determine the health condition, generating a training menu using a generative AI model, providing the generated menu to the user, and displaying it on a smart device.

[1711] "Example 2: Combining Emotion Engines"

[1712] (Claim 1)

[1713] a means for obtaining a voice input;

[1714] means for converting the captured speech input into text;

[1715] A means for analyzing the text data using natural language processing technology to analyze the user's health condition and emotions;

[1716] A means for generating an appropriate training menu using a generation AI model based on the health condition and emotion;

[1717] means for providing the generated training menu to a user;

[1718] a means for a user to perform training based on the provided training menu;

[1719] A system including:

[1720] (Claim 2)

[1721] The system of claim 1 , wherein the generated training menu includes text and reference videos.

[1722] (Claim 3)

[1723] The system of claim 1 further includes a software program that performs a series of processes: acquiring the voice input, converting the acquired voice data into text data, analyzing the text data using natural language processing technology and emotion analysis technology to determine the health condition and emotions, generating a training menu using a generative AI model, and providing the generated menu to the user.

[1724] "Application example 2 when combining emotion engines"

[1725] (Claim 1)

[1726] a means for obtaining a voice input;

[1727] means for converting the captured speech input into text;

[1728] means for analyzing the user's health condition and emotions based on the voice input;

[1729] A means for generating an appropriate training menu based on the health condition and emotion;

[1730] means for selecting and providing the generated training menu based on the user's feelings;

[1731] The system further includes means for receiving and reanalyzing feedback on the provided training menu.

[1732] (Claim 2)

[1733] The system of claim 1, wherein the generated training menu includes text and reference videos, and further provides real-time training support using a smart device.

[1734] (Claim 3)

[1735] 2. The system of claim 1, further comprising a software program that performs a series of processes of acquiring the voice input, converting the acquired voice data into text data, analyzing the text data to determine the user's health condition and emotions, generating a training menu, providing the generated menu based on the user's emotions, and further reanalyzing the feedback. [Explanation of symbols]

[1736] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining a voice input; means for converting the captured speech input into text; means for analyzing the user's health condition based on the voice input; a means for generating an appropriate training menu based on the health condition; means for providing the generated training menu to a user; A system including:

2. The system of claim 1 , wherein the generated training menu includes text and reference videos.

3. The system of claim 1 further includes a software program that performs a series of processes: acquiring the voice input, converting the acquired voice data into text data, analyzing the text data to determine health status, generating a training menu, and providing the generated menu to the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A