system
The system addresses the challenge of real-time fitness form guidance by using multimodal generation AI and AR to track and provide immediate feedback, ensuring proper exercise technique and enhancing training efficacy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing systems face challenges in providing real-time fitness form guidance and training with the correct form, leading to difficulties in maintaining proper exercise technique and reducing the risk of injury.
A system utilizing multimodal generation AI and AR technology to track, analyze, and provide real-time feedback on user movements, using a trainer avatar to demonstrate the correct form through augmented reality.
Enables users to train with the correct form, reducing the risk of injury and maximizing fitness effectiveness by providing immediate guidance and personalized feedback.
Smart Images

Figure 2026073214000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0006] , , , , , ,
[0005] , , , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that it is difficult to perform fitness form guidance in real time and it is difficult to train with the correct form.
[0005] The system according to the embodiment aims to enable a user to train with the correct form. <The system according to this embodiment comprises a tracking unit, an analysis unit, and a feedback unit. The tracking unit tracks the user's movements. The analysis unit analyzes the movement data tracked by the tracking unit. The feedback unit provides feedback based on the data analyzed by the analysis unit. [Effects of the Invention]
[0007] The system according to this embodiment allows users to perform training with the correct form. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The fitness guide system according to an embodiment of the present invention is a system that provides real-time fitness form guidance utilizing multimodal generation AI and AR technology. When a user trains at home or in the gym, the fitness guide system tracks their movements with a camera, and the generation AI analyzes and provides guidance on form in real time. An avatar of a trainer that appears through AR provides visual feedback on the correct form, enabling an experience close to personalized instruction. The fitness guide system provides support that meets the needs of general consumers, fitness gyms, and personal trainers. Training with the correct form reduces the risk of injury and maximizes fitness effectiveness. For example, performing squats with the correct form reduces the burden on the knees and lower back and allows for effective muscle training. Thus, the fitness guide system utilizing multimodal generation AI and AR technology is a very beneficial tool for users. As a result, the fitness guide system tracks, analyzes, and provides feedback on the user's movements in real time, enabling them to train with the correct form.
[0029] The fitness guide system according to this embodiment comprises a tracking unit, an analysis unit, and a feedback unit. The tracking unit tracks the user's movements. The tracking unit captures the user's movements in real time, for example, using a camera. The camera may include, for example, a high-resolution camera or a depth sensor. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. The analysis unit analyzes the movement data tracked by the tracking unit. The analysis unit analyzes the movement data using, for example, a generative AI to analyze the user's form. The generative AI may be, for example, a text generation AI or a multimodal generation AI. Based on the movement data, the analysis unit compares it with the correct form and identifies which parts are incorrect. The feedback unit provides feedback based on the data analyzed by the analysis unit. The feedback unit may, for example, use an avatar of a trainer appearing through AR to visually provide feedback on the correct form. The feedback unit can provide guidance to the user in real time. As a result, the fitness guide system according to this embodiment can train with the correct form by tracking, analyzing, and providing feedback on the user's movements in real time.
[0030] The tracking unit tracks the user's movements. For example, the tracking unit captures the user's movements in real time using a camera. The camera can include, for example, a high-resolution camera or a depth sensor. Specifically, a high-resolution camera can capture the user's subtle movements and facial expressions in detail, while a depth sensor can accurately measure the user's position and distance. This allows the tracking unit to capture the user's movements in three dimensions and acquire more accurate data. Furthermore, by using multiple cameras in combination, the tracking unit can capture the user's movements from multiple angles, eliminating blind spots. For example, by placing cameras in front, on the sides, and behind, the user's entire body movements can be tracked in detail. The tracking unit also processes the data acquired from the cameras in real time and can immediately transmit the user's movements to the analysis unit. This enables real-time tracking of the user's movements and provides immediate feedback. Additionally, the tracking unit can save the user's movement data and compare it with past data to understand the user's progress. This allows the user to confirm the effectiveness of their training and maintain motivation.
[0031] The analysis unit analyzes the motion data tracked by the tracking unit. For example, the analysis unit uses a generative AI to analyze the motion data and analyze the user's form. The generative AI can be, for example, a text-generating AI or a multimodal-generating AI. Specifically, the generative AI receives the user's motion data as input and compares it against a reference database for comparison with correct form. The generative AI analyzes the user's motion frame by frame and evaluates the accuracy of the motion in each frame. For example, if the user is performing squats, the generative AI analyzes the knee angle, back position, foot placement, etc., and compares them to the correct form. The generative AI analyzes each element of the motion in detail and identifies which parts are incorrect. For example, it can detect problems such as knees turning inward or a rounded back. Furthermore, the analysis unit can evaluate the improvement in the user's form by comparing it with past motion data. This allows the user to see changes in their form and feel the effects of their training. The analysis unit can also propose an individualized training plan based on the user's motion data. For example, if there are many problems with a particular movement, it can suggest exercises to focus on improving that movement. This allows the analysis unit to analyze user behavior in detail and provide feedback tailored to individual needs.
[0032] The feedback unit provides feedback based on data analyzed by the analysis unit. For example, the feedback unit uses a trainer avatar that appears through AR to visually demonstrate the correct form. Specifically, the trainer avatar appears in front of the user through an AR device worn by the user and demonstrates the correct form. The user can correct their own movements while watching the trainer avatar's actions. Furthermore, the feedback unit can also provide specific instructions through voice guidance and text messages. For example, it provides specific instructions in real time, such as "Open your knees a little more outward" or "Keep your back straight." This allows the user to combine visual feedback and voice guidance to train with more accurate form. The feedback unit can also evaluate the user's training progress based on their movement data and provide feedback to boost motivation. For example, it provides positive feedback such as "Your form has improved since last time" or "Great progress." This allows the user to check their progress and maintain their motivation for training. In addition, the feedback unit can collect user feedback and use it to improve the system. For example, it can adjust the content and method of feedback based on the feedback provided by the user to provide more effective guidance. This allows the feedback unit to provide real-time guidance to the user and help them train using the correct form.
[0033] The feedback unit can provide visual feedback on the correct form through a trainer avatar that appears via AR. For example, the feedback unit can display the trainer avatar in front of the user and demonstrate the correct form. The trainer avatar can move in real time in accordance with the user's movements. The feedback unit enables the user to intuitively understand the correct form. By providing visual feedback through AR, the user can intuitively understand the correct form. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the movements of the trainer avatar into a generating AI and have the generating AI execute the avatar's movements.
[0034] The tracking unit can capture the user's movements in real time using a camera. The tracking unit can capture the user's movements using, for example, a high-resolution camera. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. The tracking unit can also capture the user's movements using, for example, a depth sensor. This allows for accurate, real-time tracking of the user's movements using a camera. Some or all of the above-described processes in the tracking unit may be performed using, for example, AI, or without AI. For example, the tracking unit can input the movement data acquired from the camera into a generating AI and have the generating AI perform analysis of the movement data.
[0035] The analysis unit can compare the user's action data with a correct form and identify which parts are incorrect. The analysis unit can analyze the user's form by, for example, using a generation AI to analyze the action data. The generation AI can be, for example, a text generation AI or a multimodal generation AI. The analysis unit compares the action data with a correct form and identifies which parts are incorrect. This makes it easier to identify errors in the user's actions by comparing them with a correct form. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input action data into a generation AI and have the generation AI perform a comparison with a correct form.
[0036] The feedback unit can provide real-time guidance to the user based on the analysis results. The feedback unit provides, for example, audio or visual feedback to the user. The feedback unit helps the user intuitively understand the correct form. The feedback unit can, for example, display a trainer avatar in front of the user and demonstrate the correct form. This allows the user to immediately correct their form through real-time guidance. Some or all of the above processing in the feedback unit may be performed using, for example, AI, or not using AI. For example, the feedback unit can input the analysis results into a generating AI and have the generating AI perform real-time guidance.
[0037] The feedback unit can provide support tailored to the needs of general consumers, fitness gyms, and personal trainers. For example, the feedback unit can provide simple feedback to general consumers and detailed feedback to fitness gyms. For personal trainers, the feedback unit can propose training plans and manage progress. This allows the unit to serve a wide range of users by providing support that meets the diverse needs of users. Some or all of the above-described processes in the feedback unit may be performed using AI, for example, or not. For example, the feedback unit can input feedback tailored to the user's needs into a generating AI and have the generating AI execute the optimal feedback.
[0038] The tracking unit can optimize its tracking algorithm by referring to the user's past behavior data during tracking. For example, the tracking unit adjusts its tracking algorithm based on data from training the user has performed in the past. The tracking unit learns specific behavior patterns from the user's past behavior data to improve tracking accuracy. The tracking unit customizes its tracking algorithm by referring to the user's past training history. This allows for improved tracking accuracy by referring to past behavior data. Some or all of the above processes in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input the user's past behavior data into a generating AI and have the generating AI perform the optimization of the tracking algorithm.
[0039] The tracking unit can dynamically change tracking parameters in accordance with the user's movement speed and rhythm during tracking. For example, if the user is performing fast movements, the tracking unit increases the tracking frame rate to improve accuracy. If the user is performing slow movements, the tracking unit decreases the tracking frame rate to conserve resources. The tracking unit adjusts the tracking parameters in real time according to the user's movement rhythm. This allows for more accurate tracking by adjusting the tracking parameters according to the movement speed and rhythm. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input data on the user's movement speed and rhythm into a generating AI and have the generating AI perform the dynamic changes to the tracking parameters.
[0040] The tracking unit can analyze the user's ambient sounds during tracking and incorporate them as background information for the movements. For example, the tracking unit can analyze the ambient sounds around the user to improve tracking accuracy. If the user is training while listening to music, the tracking unit adjusts the tracking parameters to match the rhythm. The tracking unit analyzes the noise level around the user to optimize tracking accuracy. In this way, tracking accuracy can be improved by analyzing ambient sounds. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input the user's ambient sound data into a generating AI and have the generating AI perform the analysis of the ambient sounds.
[0041] The tracking unit can improve tracking accuracy by taking into account the influence of the user's clothing and accessories during tracking. For example, the tracking unit analyzes the color and pattern of the clothing the user is wearing to improve tracking accuracy. If the user is wearing accessories, the tracking unit adjusts the tracking parameters to take their influence into account. The tracking unit analyzes the movement of the user's clothing and accessories to optimize tracking accuracy. In this way, tracking accuracy can be improved by taking into account the influence of clothing and accessories. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input data on the user's clothing and accessories into a generating AI and have the generating AI perform the improvement of tracking accuracy.
[0042] The analysis unit can improve the accuracy of its analysis by referring to the user's past training data during the analysis process. For example, the analysis unit adjusts the analysis algorithm based on the user's past training data. The analysis unit learns specific behavioral patterns from the user's past training data to improve the accuracy of the analysis. The analysis unit customizes the analysis algorithm by referring to the user's past training history. This allows the accuracy of the analysis to be improved by referring to past training data. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's past training data into a generating AI and have the generating AI perform the task of improving the accuracy of the analysis.
[0043] The analysis unit can customize the analysis criteria according to the user's body type and muscle mass during analysis. For example, the analysis unit adjusts the analysis criteria based on the user's body type data. The analysis unit customizes the analysis criteria taking into account the user's muscle mass. The analysis unit optimizes the analysis algorithm according to the user's body type and muscle mass. This allows for more accurate analysis by customizing the analysis criteria according to body type and muscle mass. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's body type and muscle mass data into a generating AI and have the generating AI perform the customization of the analysis criteria.
[0044] The analysis unit can determine the priority of analysis based on the user's training goals during the analysis process. For example, if the user's training goal is to improve muscle strength, the analysis unit will prioritize analyzing data related to muscle strength. If the user's training goal is to lose weight, the analysis unit will prioritize analyzing data related to calorie consumption. If the user's training goal is to improve flexibility, the analysis unit will prioritize analyzing data related to flexibility. By determining the priority of analysis based on training goals, more effective feedback can be provided. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's training goal data into a generating AI and have the generating AI determine the priority of analysis.
[0045] The analysis unit can improve the accuracy of its analysis by referring to the user's diet and sleep data during the analysis process. For example, the analysis unit adjusts the analysis algorithm based on the user's diet data. The analysis unit improves the accuracy of its analysis by considering the user's sleep data. The analysis unit optimizes the analysis algorithm by referring to the user's diet and sleep data. In this way, the accuracy of the analysis can be improved by referring to the diet and sleep data. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without using AI. For example, the analysis unit can input the user's diet and sleep data into a generating AI and have the generating AI perform the improvement of the analysis accuracy.
[0046] The feedback unit can select the optimal feedback method by referring to the user's past feedback history when providing feedback. For example, the feedback unit selects the optimal feedback method based on feedback the user has received in the past. The feedback unit learns specific feedback patterns from the user's past feedback history and provides the optimal method. The feedback unit customizes the content of the feedback by referring to the user's past feedback history. In this way, the optimal feedback method can be provided by referring to past feedback history. Some or all of the above processes in the feedback unit may be performed using AI, for example, or without using AI. For example, the feedback unit can input the user's past feedback history into a generating AI and have the generating AI select the optimal feedback method.
[0047] The feedback unit can adjust the level of detail of the feedback according to the user's training progress. For example, if the user's training progress is good, the feedback unit will provide detailed feedback. If the user's training progress is behind, the feedback unit will provide concise feedback. The feedback unit customizes the content of the feedback according to the user's training progress. This allows for more appropriate feedback to be provided by adjusting the level of detail of the feedback according to the training progress. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's training progress data into a generating AI and have the generating AI perform the adjustment of the level of detail of the feedback.
[0048] The feedback unit can change the format of the feedback based on the user's training environment. For example, if the user is training at home, the feedback unit provides audio feedback. If the user is training at a gym, the feedback unit provides visual feedback. If the user is training outdoors, the feedback unit provides vibration feedback. By changing the format of the feedback based on the training environment, more appropriate feedback can be provided. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's training environment data into a generating AI and have the generating AI perform the change in the format of the feedback.
[0049] The feedback unit can customize the content of feedback by referring to the user's social media activity. For example, the feedback unit can customize the content of feedback based on training content shared by the user on social media. The feedback unit can provide feedback by referring to advice from fitness influencers that the user follows on social media. The feedback unit can provide feedback that increases the user's motivation for training based on the user's social media activity. In this way, by referring to social media activity, the feedback unit can provide the user with the most suitable feedback. Some or all of the above processes in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input the user's social media activity data into a generating AI and have the generating AI perform the customization of the feedback content.
[0050] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0051] The fitness guidance system can also include a heart rate monitoring unit that monitors the user's heart rate. The heart rate monitoring unit acquires the user's heart rate data in real time and transmits it to the analysis unit. Based on the heart rate data, the analysis unit can evaluate the user's exercise intensity and recommend an appropriate training intensity. For example, if the user's heart rate is too high, the analysis unit will instruct the user to lower the training intensity, and conversely, if the heart rate is too low, it will instruct the user to increase the intensity. It can also estimate the user's fatigue level based on the heart rate data and recommend rest as needed. This allows the user to perform optimal training tailored to their physical condition.
[0052] The fitness guide system may also include a meal data acquisition unit that obtains the user's meal data. The meal data acquisition unit records the contents of the meals consumed by the user and transmits them to the analysis unit. Based on the meal data, the analysis unit can evaluate the user's nutritional status and provide dietary advice to maximize the effectiveness of training. For example, if the user is deficient in protein, the analysis unit can recommend meals high in protein, and conversely, if the user is consuming too many calories, it can recommend calorie restriction. This allows the user to manage their health from both a training and dietary perspective.
[0053] The fitness guidance system can also include a sleep data acquisition unit that collects the user's sleep data. The sleep data acquisition unit records the user's sleep patterns and transmits them to the analysis unit. Based on the sleep data, the analysis unit can evaluate the user's fatigue level and adjust the intensity and content of the training. For example, if the user has not gotten enough sleep, the analysis unit can recommend lighter training, and conversely, if the user has gotten enough sleep, it can recommend harder training. This allows the user to perform training that is optimal for their sleep condition.
[0054] The fitness guide system can also include an environmental monitoring unit that monitors the user's training environment. The environmental monitoring unit acquires temperature, humidity, noise level, and other data of the user's training environment in real time and transmits it to the analysis unit. Based on the environmental data, the analysis unit can suggest optimal environmental conditions for training. For example, if the temperature is too high, the analysis unit can instruct the user to stop training, and conversely, if the temperature is appropriate, it can instruct the user to continue training. This allows the user to train in an optimal environment.
[0055] The fitness guide system may also include a training plan optimization unit that optimizes the training plan by referencing the user's training history. The training plan optimization unit proposes the optimal training plan based on data from the user's past training. For example, it can evaluate the effectiveness of past training and propose the most effective training again. It can also adjust the plan to avoid training that the user has struggled with in the past. This allows the user to execute an optimal training plan based on their own training history.
[0056] The following briefly describes the processing flow for example form 1.
[0057] Step 1: The tracking unit tracks the user's movements. The tracking unit captures the user's movements in real time, for example, using a camera. The camera may include, for example, a high-resolution camera or a depth sensor. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. Step 2: The analysis unit analyzes the motion data tracked by the tracking unit. The analysis unit analyzes the motion data using, for example, a generation AI to analyze the user's form. The generation AI can be, for example, a text generation AI or a multimodal generation AI. Based on the motion data, the analysis unit compares it with a correct form to identify which parts are incorrect. Step 3: The feedback unit provides feedback based on the data analyzed by the analysis unit. For example, the feedback unit uses an AR-enabled trainer avatar to visually provide feedback on the correct form. The feedback unit can provide guidance to the user in real time.
[0058] (Example of form 2) The fitness guide system according to an embodiment of the present invention is a system that provides real-time fitness form guidance utilizing multimodal generation AI and AR technology. When a user trains at home or in the gym, the fitness guide system tracks their movements with a camera, and the generation AI analyzes and provides guidance on form in real time. An avatar of a trainer that appears through AR provides visual feedback on the correct form, enabling an experience close to personalized instruction. The fitness guide system provides support that meets the needs of general consumers, fitness gyms, and personal trainers. Training with the correct form reduces the risk of injury and maximizes fitness effectiveness. For example, performing squats with the correct form reduces the burden on the knees and lower back and allows for effective muscle training. Thus, the fitness guide system utilizing multimodal generation AI and AR technology is a very beneficial tool for users. As a result, the fitness guide system tracks, analyzes, and provides feedback on the user's movements in real time, enabling them to train with the correct form.
[0059] The fitness guide system according to this embodiment comprises a tracking unit, an analysis unit, and a feedback unit. The tracking unit tracks the user's movements. The tracking unit captures the user's movements in real time, for example, using a camera. The camera may include, for example, a high-resolution camera or a depth sensor. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. The analysis unit analyzes the movement data tracked by the tracking unit. The analysis unit analyzes the movement data using, for example, a generative AI to analyze the user's form. The generative AI may be, for example, a text generation AI or a multimodal generation AI. Based on the movement data, the analysis unit compares it with the correct form and identifies which parts are incorrect. The feedback unit provides feedback based on the data analyzed by the analysis unit. The feedback unit may, for example, use an avatar of a trainer appearing through AR to visually provide feedback on the correct form. The feedback unit can provide guidance to the user in real time. As a result, the fitness guide system according to this embodiment can train with the correct form by tracking, analyzing, and providing feedback on the user's movements in real time.
[0060] The tracking unit tracks the user's movements. For example, the tracking unit captures the user's movements in real time using a camera. The camera can include, for example, a high-resolution camera or a depth sensor. Specifically, a high-resolution camera can capture the user's subtle movements and facial expressions in detail, while a depth sensor can accurately measure the user's position and distance. This allows the tracking unit to capture the user's movements in three dimensions and acquire more accurate data. Furthermore, by using multiple cameras in combination, the tracking unit can capture the user's movements from multiple angles, eliminating blind spots. For example, by placing cameras in front, on the sides, and behind, the user's entire body movements can be tracked in detail. The tracking unit also processes the data acquired from the cameras in real time and can immediately transmit the user's movements to the analysis unit. This enables real-time tracking of the user's movements and provides immediate feedback. Additionally, the tracking unit can save the user's movement data and compare it with past data to understand the user's progress. This allows the user to confirm the effectiveness of their training and maintain motivation.
[0061] The analysis unit analyzes the motion data tracked by the tracking unit. For example, the analysis unit uses a generative AI to analyze the motion data and analyze the user's form. The generative AI can be, for example, a text-generating AI or a multimodal-generating AI. Specifically, the generative AI receives the user's motion data as input and compares it against a reference database for comparison with correct form. The generative AI analyzes the user's motion frame by frame and evaluates the accuracy of the motion in each frame. For example, if the user is performing squats, the generative AI analyzes the knee angle, back position, foot placement, etc., and compares them to the correct form. The generative AI analyzes each element of the motion in detail and identifies which parts are incorrect. For example, it can detect problems such as knees turning inward or a rounded back. Furthermore, the analysis unit can evaluate the improvement in the user's form by comparing it with past motion data. This allows the user to see changes in their form and feel the effects of their training. The analysis unit can also propose an individualized training plan based on the user's motion data. For example, if there are many problems with a particular movement, it can suggest exercises to focus on improving that movement. This allows the analysis unit to analyze user behavior in detail and provide feedback tailored to individual needs.
[0062] The feedback unit provides feedback based on data analyzed by the analysis unit. For example, the feedback unit uses a trainer avatar that appears through AR to visually demonstrate the correct form. Specifically, the trainer avatar appears in front of the user through an AR device worn by the user and demonstrates the correct form. The user can correct their own movements while watching the trainer avatar's actions. Furthermore, the feedback unit can also provide specific instructions through voice guidance and text messages. For example, it provides specific instructions in real time, such as "Open your knees a little more outward" or "Keep your back straight." This allows the user to combine visual feedback and voice guidance to train with more accurate form. The feedback unit can also evaluate the user's training progress based on their movement data and provide feedback to boost motivation. For example, it provides positive feedback such as "Your form has improved since last time" or "Great progress." This allows the user to check their progress and maintain their motivation for training. In addition, the feedback unit can collect user feedback and use it to improve the system. For example, it can adjust the content and method of feedback based on the feedback provided by the user to provide more effective guidance. This allows the feedback unit to provide real-time guidance to the user and help them train using the correct form.
[0063] The feedback unit can provide visual feedback on the correct form through a trainer avatar that appears via AR. For example, the feedback unit can display the trainer avatar in front of the user and demonstrate the correct form. The trainer avatar can move in real time in accordance with the user's movements. The feedback unit enables the user to intuitively understand the correct form. By providing visual feedback through AR, the user can intuitively understand the correct form. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the movements of the trainer avatar into a generating AI and have the generating AI execute the avatar's movements.
[0064] The tracking unit can capture the user's movements in real time using a camera. The tracking unit can capture the user's movements using, for example, a high-resolution camera. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. The tracking unit can also capture the user's movements using, for example, a depth sensor. This allows for accurate, real-time tracking of the user's movements using a camera. Some or all of the above-described processes in the tracking unit may be performed using, for example, AI, or without AI. For example, the tracking unit can input the movement data acquired from the camera into a generating AI and have the generating AI perform analysis of the movement data.
[0065] The analysis unit can compare the user's action data with a correct form and identify which parts are incorrect. The analysis unit can analyze the user's form by, for example, using a generation AI to analyze the action data. The generation AI can be, for example, a text generation AI or a multimodal generation AI. The analysis unit compares the action data with a correct form and identifies which parts are incorrect. This makes it easier to identify errors in the user's actions by comparing them with a correct form. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input action data into a generation AI and have the generation AI perform a comparison with a correct form.
[0066] The feedback unit can provide real-time guidance to the user based on the analysis results. The feedback unit provides, for example, audio or visual feedback to the user. The feedback unit helps the user intuitively understand the correct form. The feedback unit can, for example, display a trainer avatar in front of the user and demonstrate the correct form. This allows the user to immediately correct their form through real-time guidance. Some or all of the above processing in the feedback unit may be performed using, for example, AI, or not using AI. For example, the feedback unit can input the analysis results into a generating AI and have the generating AI perform real-time guidance.
[0067] The feedback unit can provide support tailored to the needs of general consumers, fitness gyms, and personal trainers. For example, the feedback unit can provide simple feedback to general consumers and detailed feedback to fitness gyms. For personal trainers, the feedback unit can propose training plans and manage progress. This allows the unit to serve a wide range of users by providing support that meets the diverse needs of users. Some or all of the above-described processes in the feedback unit may be performed using AI, for example, or not. For example, the feedback unit can input feedback tailored to the user's needs into a generating AI and have the generating AI execute the optimal feedback.
[0068] The tracking unit can estimate the user's emotions and adjust the tracking accuracy based on the estimated emotions. For example, if the user is stressed, the tracking unit increases the tracking accuracy to provide more detailed feedback. If the user is relaxed, the tracking unit slightly loosens the tracking accuracy to prioritize natural movements. If the user is focused, the tracking unit optimizes the tracking accuracy to provide the most effective feedback. This allows for more appropriate feedback by adjusting the tracking accuracy according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the tracking unit may be performed using AI or not. For example, the tracking unit can input user emotion data into a generative AI and have the generative AI adjust the tracking accuracy.
[0069] The tracking unit can optimize its tracking algorithm by referring to the user's past behavior data during tracking. For example, the tracking unit adjusts its tracking algorithm based on data from training the user has performed in the past. The tracking unit learns specific behavior patterns from the user's past behavior data to improve tracking accuracy. The tracking unit customizes its tracking algorithm by referring to the user's past training history. This allows for improved tracking accuracy by referring to past behavior data. Some or all of the above processes in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input the user's past behavior data into a generating AI and have the generating AI perform the optimization of the tracking algorithm.
[0070] The tracking unit can dynamically change tracking parameters in accordance with the user's movement speed and rhythm during tracking. For example, if the user is performing fast movements, the tracking unit increases the tracking frame rate to improve accuracy. If the user is performing slow movements, the tracking unit decreases the tracking frame rate to conserve resources. The tracking unit adjusts the tracking parameters in real time according to the user's movement rhythm. This allows for more accurate tracking by adjusting the tracking parameters according to the movement speed and rhythm. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input data on the user's movement speed and rhythm into a generating AI and have the generating AI perform the dynamic changes to the tracking parameters.
[0071] The tracking unit can estimate the user's emotions and adjust the tracking frequency based on the estimated emotions. For example, if the user is stressed, the tracking unit increases the tracking frequency to provide more detailed feedback. If the user is relaxed, the tracking unit slightly reduces the tracking frequency to emphasize natural movements. If the user is focused, the tracking unit optimizes the tracking frequency to provide the most effective feedback. This allows for more appropriate feedback by adjusting the tracking frequency according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the tracking unit may be performed using AI or not. For example, the tracking unit can input user emotion data into a generative AI and have the generative AI adjust the tracking frequency.
[0072] The tracking unit can analyze the user's ambient sounds during tracking and incorporate them as background information for the movements. For example, the tracking unit can analyze the ambient sounds around the user to improve tracking accuracy. If the user is training while listening to music, the tracking unit adjusts the tracking parameters to match the rhythm. The tracking unit analyzes the noise level around the user to optimize tracking accuracy. In this way, tracking accuracy can be improved by analyzing ambient sounds. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input the user's ambient sound data into a generating AI and have the generating AI perform the analysis of the ambient sounds.
[0073] The tracking unit can improve tracking accuracy by taking into account the influence of the user's clothing and accessories during tracking. For example, the tracking unit analyzes the color and pattern of the clothing the user is wearing to improve tracking accuracy. If the user is wearing accessories, the tracking unit adjusts the tracking parameters to take their influence into account. The tracking unit analyzes the movement of the user's clothing and accessories to optimize tracking accuracy. In this way, tracking accuracy can be improved by taking into account the influence of clothing and accessories. Some or all of the above processing in the tracking unit may be performed using AI, for example, or without AI. For example, the tracking unit can input data on the user's clothing and accessories into a generating AI and have the generating AI perform the improvement of tracking accuracy.
[0074] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. For example, if the user is stressed, the analysis unit adjusts the analysis algorithm with high accuracy and provides detailed feedback. If the user is relaxed, the analysis unit slightly loosens the analysis algorithm and prioritizes natural behavior. If the user is focused, the analysis unit optimizes the analysis algorithm and provides the most effective feedback. In this way, by adjusting the analysis algorithm according to the user's emotions, more appropriate feedback can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI perform the adjustment of the analysis algorithm.
[0075] The analysis unit can improve the accuracy of its analysis by referring to the user's past training data during the analysis process. For example, the analysis unit adjusts the analysis algorithm based on the user's past training data. The analysis unit learns specific behavioral patterns from the user's past training data to improve the accuracy of the analysis. The analysis unit customizes the analysis algorithm by referring to the user's past training history. This allows the accuracy of the analysis to be improved by referring to past training data. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's past training data into a generating AI and have the generating AI perform the task of improving the accuracy of the analysis.
[0076] The analysis unit can customize the analysis criteria according to the user's body type and muscle mass during analysis. For example, the analysis unit adjusts the analysis criteria based on the user's body type data. The analysis unit customizes the analysis criteria taking into account the user's muscle mass. The analysis unit optimizes the analysis algorithm according to the user's body type and muscle mass. This allows for more accurate analysis by customizing the analysis criteria according to body type and muscle mass. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's body type and muscle mass data into a generating AI and have the generating AI perform the customization of the analysis criteria.
[0077] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. For example, if the user is stressed, the analysis unit provides a simple and highly visible display method. If the user is relaxed, the analysis unit provides a display method that includes detailed information. If the user is focused, the analysis unit provides a display method that gets straight to the point. By adjusting the display method according to the user's emotions, more appropriate feedback can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or not using AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI adjust the display method of the analysis results.
[0078] The analysis unit can determine the priority of analysis based on the user's training goals during the analysis process. For example, if the user's training goal is to improve muscle strength, the analysis unit will prioritize analyzing data related to muscle strength. If the user's training goal is to lose weight, the analysis unit will prioritize analyzing data related to calorie consumption. If the user's training goal is to improve flexibility, the analysis unit will prioritize analyzing data related to flexibility. By determining the priority of analysis based on training goals, more effective feedback can be provided. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's training goal data into a generating AI and have the generating AI determine the priority of analysis.
[0079] The analysis unit can improve the accuracy of its analysis by referring to the user's diet and sleep data during the analysis process. For example, the analysis unit adjusts the analysis algorithm based on the user's diet data. The analysis unit improves the accuracy of its analysis by considering the user's sleep data. The analysis unit optimizes the analysis algorithm by referring to the user's diet and sleep data. In this way, the accuracy of the analysis can be improved by referring to the diet and sleep data. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without using AI. For example, the analysis unit can input the user's diet and sleep data into a generating AI and have the generating AI perform the improvement of the analysis accuracy.
[0080] The feedback unit can estimate the user's emotions and adjust the content of the feedback based on the estimated emotions. For example, if the user is stressed, the feedback unit will provide gentle feedback. If the user is relaxed, the feedback unit will provide detailed feedback. If the user is focused, the feedback unit will provide concise feedback. By adjusting the content of the feedback according to the user's emotions, more appropriate feedback can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input user emotion data into a generative AI and have the generative AI adjust the content of the feedback.
[0081] The feedback unit can select the optimal feedback method by referring to the user's past feedback history when providing feedback. For example, the feedback unit selects the optimal feedback method based on feedback the user has received in the past. The feedback unit learns specific feedback patterns from the user's past feedback history and provides the optimal method. The feedback unit customizes the content of the feedback by referring to the user's past feedback history. In this way, the optimal feedback method can be provided by referring to past feedback history. Some or all of the above processes in the feedback unit may be performed using AI, for example, or without using AI. For example, the feedback unit can input the user's past feedback history into a generating AI and have the generating AI select the optimal feedback method.
[0082] The feedback unit can adjust the level of detail of the feedback according to the user's training progress. For example, if the user's training progress is good, the feedback unit will provide detailed feedback. If the user's training progress is behind, the feedback unit will provide concise feedback. The feedback unit customizes the content of the feedback according to the user's training progress. This allows for more appropriate feedback to be provided by adjusting the level of detail of the feedback according to the training progress. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's training progress data into a generating AI and have the generating AI perform the adjustment of the level of detail of the feedback.
[0083] The feedback unit can estimate the user's emotions and adjust the timing of feedback based on the estimated emotions. For example, if the user is stressed, the feedback unit can delay the timing of feedback to provide it in a relaxed state. If the user is relaxed, the feedback unit can advance the timing of feedback to provide it immediately. If the user is focused, the feedback unit can optimize the timing of feedback to provide it at the most effective time. In this way, by adjusting the timing of feedback according to the user's emotions, more effective feedback can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input user emotion data into the generative AI and have the generative AI perform the adjustment of the feedback timing.
[0084] The feedback unit can change the format of the feedback based on the user's training environment. For example, if the user is training at home, the feedback unit provides audio feedback. If the user is training at a gym, the feedback unit provides visual feedback. If the user is training outdoors, the feedback unit provides vibration feedback. By changing the format of the feedback based on the training environment, more appropriate feedback can be provided. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's training environment data into a generating AI and have the generating AI perform the change in the format of the feedback.
[0085] The feedback unit can customize the content of feedback by referring to the user's social media activity. For example, the feedback unit can customize the content of feedback based on training content shared by the user on social media. The feedback unit can provide feedback by referring to advice from fitness influencers that the user follows on social media. The feedback unit can provide feedback that increases the user's motivation for training based on the user's social media activity. In this way, by referring to social media activity, the feedback unit can provide the user with the most suitable feedback. Some or all of the above processes in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input the user's social media activity data into a generating AI and have the generating AI perform the customization of the feedback content.
[0086] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0087] The fitness guidance system can also include a heart rate monitoring unit that monitors the user's heart rate. The heart rate monitoring unit acquires the user's heart rate data in real time and transmits it to the analysis unit. Based on the heart rate data, the analysis unit can evaluate the user's exercise intensity and recommend an appropriate training intensity. For example, if the user's heart rate is too high, the analysis unit will instruct the user to lower the training intensity, and conversely, if the heart rate is too low, it will instruct the user to increase the intensity. It can also estimate the user's fatigue level based on the heart rate data and recommend rest as needed. This allows the user to perform optimal training tailored to their physical condition.
[0088] The fitness guidance system can also include a training plan customization unit that estimates the user's emotions and customizes the training plan based on those emotions. The training plan customization unit can recommend relaxing exercises if the user is stressed, and more challenging exercises if the user is relaxed. It can also recommend exercises to maintain concentration if the user is focused. This allows the system to provide an optimal training plan tailored to the user's emotional state.
[0089] The fitness guide system may also include a meal data acquisition unit that obtains the user's meal data. The meal data acquisition unit records the contents of the meals consumed by the user and transmits them to the analysis unit. Based on the meal data, the analysis unit can evaluate the user's nutritional status and provide dietary advice to maximize the effectiveness of training. For example, if the user is deficient in protein, the analysis unit can recommend meals high in protein, and conversely, if the user is consuming too many calories, it can recommend calorie restriction. This allows the user to manage their health from both a training and dietary perspective.
[0090] The fitness guidance system can also include a motivation enhancement unit that estimates the user's emotions and increases training motivation based on those emotions. This unit can provide encouraging messages if the user is stressed, and provide feedback that fosters a sense of accomplishment if the user is relaxed. It can also provide advice to maintain focus if the user is concentrating. This allows for the provision of appropriate motivation enhancement measures tailored to the user's emotional state.
[0091] The fitness guidance system can also include a sleep data acquisition unit that collects the user's sleep data. The sleep data acquisition unit records the user's sleep patterns and transmits them to the analysis unit. Based on the sleep data, the analysis unit can evaluate the user's fatigue level and adjust the intensity and content of the training. For example, if the user has not gotten enough sleep, the analysis unit can recommend lighter training, and conversely, if the user has gotten enough sleep, it can recommend harder training. This allows the user to perform training that is optimal for their sleep condition.
[0092] The fitness guidance system can also include a progress evaluation unit that estimates the user's emotions and evaluates training progress based on those emotions. The progress evaluation unit can perform a gentler progress evaluation if the user is stressed, and a more detailed evaluation if the user is relaxed. It can also optimize the progress evaluation to provide the most effective feedback when the user is focused. This allows for appropriate progress evaluation tailored to the user's emotional state.
[0093] The fitness guide system can also include an environmental monitoring unit that monitors the user's training environment. The environmental monitoring unit acquires temperature, humidity, noise level, and other data of the user's training environment in real time and transmits it to the analysis unit. Based on the environmental data, the analysis unit can suggest optimal environmental conditions for training. For example, if the temperature is too high, the analysis unit can instruct the user to stop training, and conversely, if the temperature is appropriate, it can instruct the user to continue training. This allows the user to train in an optimal environment.
[0094] The fitness guidance system may also include a rest timing adjustment unit that estimates the user's emotions and adjusts the timing of rest periods during training based on those emotions. The rest timing adjustment unit can instruct the user to take an earlier rest if they are feeling stressed, and to delay rests if they are relaxed. It can also instruct the user to take a rest at the optimal time if they are focused. This allows for the provision of appropriate rest timings tailored to the user's emotional state.
[0095] The fitness guide system may also include a training plan optimization unit that optimizes the training plan by referencing the user's training history. The training plan optimization unit proposes the optimal training plan based on data from the user's past training. For example, it can evaluate the effectiveness of past training and propose the most effective training again. It can also adjust the plan to avoid training that the user has struggled with in the past. This allows the user to execute an optimal training plan based on their own training history.
[0096] The fitness guidance system may also include a feedback content adjustment unit that estimates the user's emotions and adjusts the training feedback content based on those emotions. The feedback content adjustment unit can provide gentle feedback when the user is stressed, and detailed feedback when the user is relaxed. It can also provide concise feedback when the user is focused. This allows for the provision of appropriate feedback content tailored to the user's emotional state.
[0097] The following briefly describes the processing flow for example form 2.
[0098] Step 1: The tracking unit tracks the user's movements. The tracking unit captures the user's movements in real time, for example, using a camera. The camera may include, for example, a high-resolution camera or a depth sensor. Based on the data acquired from the camera, the tracking unit can accurately track the user's movements. Step 2: The analysis unit analyzes the motion data tracked by the tracking unit. The analysis unit analyzes the motion data using, for example, a generation AI to analyze the user's form. The generation AI can be, for example, a text generation AI or a multimodal generation AI. Based on the motion data, the analysis unit compares it with a correct form to identify which parts are incorrect. Step 3: The feedback unit provides feedback based on the data analyzed by the analysis unit. For example, the feedback unit uses an AR-enabled trainer avatar to visually provide feedback on the correct form. The feedback unit can provide guidance to the user in real time.
[0099] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0100] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0101] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0102] Each of the multiple elements described above, including the tracking unit, analysis unit, and feedback unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the tracking unit captures the user's movements in real time using the camera 42 of the smart device 14. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the movement data using generated AI and analyzes the user's form. The feedback unit uses, for example, the output device 40 of the smart device 14 to provide visual feedback on the correct form through a trainer avatar appearing via AR. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0103] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0104] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0105] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0106] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0107] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0108] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0109] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0110] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0111] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0112] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0113] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0114] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0115] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0116] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0117] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0118] Each of the multiple elements described above, including the tracking unit, analysis unit, and feedback unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the tracking unit captures the user's movements in real time using the camera 42 of the smart glasses 214. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the movement data using generated AI and analyzes the user's form. The feedback unit uses, for example, the speaker 240 of the smart glasses 214 to provide visual feedback on the correct form through a trainer avatar appearing via AR. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0119] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0120] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0121] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0122] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0123] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0124] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0125] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0126] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0127] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0128] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0129] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0130] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0131] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0132] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0133] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0134] Each of the multiple elements described above, including the tracking unit, analysis unit, and feedback unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the tracking unit captures the user's movements in real time using the camera 42 of the headset terminal 314. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the movement data using generated AI and analyzes the user's form. The feedback unit uses the display 343 of the headset terminal 314 to provide visual feedback on the correct form through a trainer avatar appearing via AR. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0135] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0136] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0137] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0138] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0139] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0140] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0141] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0142] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0143] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0144] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0145] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0146] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0147] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0148] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0149] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0150] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0151] Each of the multiple elements described above, including the tracking unit, analysis unit, and feedback unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the tracking unit captures the user's movements in real time using the camera 42 of the robot 414. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the movement data using generated AI and analyzes the user's form. The feedback unit uses, for example, the speaker 240 of the robot 414 to provide visual feedback on the correct form through a trainer avatar appearing via AR. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0152] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0153] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0154] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0155] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0156] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0157] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0158] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0159] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0160] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0161] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0162] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0163] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0164] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0165] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0166] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0167] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0168] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0169] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0170] (Note 1) A tracking unit that tracks user actions, An analysis unit analyzes the motion data tracked by the aforementioned tracking unit, The system includes a feedback unit that provides feedback based on the data analyzed by the analysis unit. A system characterized by the following features. (Note 2) The aforementioned feedback unit is An AR-enabled trainer avatar provides visual feedback on correct form. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned tracking unit is Capture user actions in real time using a camera. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned analysis unit, Based on user behavior data, the system compares it to a correct form to identify which parts are incorrect. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned feedback unit is Based on the analysis results, provide real-time guidance to users. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned feedback unit is We provide support tailored to the needs of general consumers, fitness gyms, and personal trainers. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned tracking unit is It estimates the user's emotions and adjusts the tracking accuracy based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned tracking unit is During tracking, the tracking algorithm is optimized by referencing the user's past behavior data. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned tracking unit is During tracking, the tracking parameters are dynamically changed according to the user's movement speed and rhythm. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned tracking unit is It estimates the user's emotions and adjusts the tracking frequency based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned tracking unit is During tracking, the system analyzes the user's ambient sounds and incorporates them as background information for their actions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned tracking unit is During tracking, we improve tracking accuracy by taking into account the influence of the user's clothing and accessories. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, It estimates the user's emotions and adjusts the analysis algorithm based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the system references the user's past training data to improve the accuracy of the analysis. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, the analysis criteria are customized according to the user's body type and muscle mass. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the analysis priorities are determined based on the user's training objectives. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the accuracy of the analysis is improved by referencing the user's diet and sleep data. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned feedback unit is It estimates the user's emotions and adjusts the content of the feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned feedback unit is When providing feedback, the system selects the most suitable feedback method by referring to the user's past feedback history. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned feedback unit is When providing feedback, adjust the level of detail in the feedback based on the user's training progress. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned feedback unit is It estimates the user's emotions and adjusts the timing of feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned feedback unit is When providing feedback, the format of the feedback will be changed based on the user's training environment. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned feedback unit is When providing feedback, the content of the feedback is customized by referencing the user's social media activity. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0171] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A tracking unit that tracks user actions, An analysis unit analyzes the motion data tracked by the aforementioned tracking unit, The system includes a feedback unit that provides feedback based on the data analyzed by the analysis unit. A system characterized by the following features.
2. The aforementioned feedback unit is An AR-enabled trainer avatar provides visual feedback on correct form. The system according to feature 1.
3. The aforementioned tracking unit is Capture user actions in real time using a camera. The system according to feature 1.
4. The aforementioned analysis unit, Based on user behavior data, the system compares it to a correct form to identify which parts are incorrect. The system according to feature 1.
5. The aforementioned feedback unit is Based on the analysis results, provide real-time guidance to users. The system according to feature 1.
6. The aforementioned feedback unit is We provide support tailored to the needs of general consumers, fitness gyms, and personal trainers. The system according to feature 1.
7. The aforementioned tracking unit is It estimates the user's emotions and adjusts the tracking accuracy based on the estimated user emotions. The system according to feature 1.
8. The aforementioned tracking unit is During tracking, the tracking algorithm is optimized by referencing the user's past behavior data. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A