system
The system addresses the challenge of beginners finding suitable practice methods by using AI to analyze physical characteristics and provide personalized playing techniques and feedback, ensuring continuous instrument learning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Beginners of musical instruments face difficulty in finding a practice method suitable for their physique and finger length, and often lack individually optimized guidance.
A system comprising an acquisition unit, analysis unit, generation unit, and evaluation unit that uses AI to analyze physical characteristics, generate optimal playing forms and fingering techniques, provide feedback, and evaluate progress, enabling individually optimized instruction.
Provides optimal playing form and fingering techniques tailored to the user's physique and finger length, preventing discouragement and supporting continuous instrument playing by offering personalized and available guidance.
Smart Images

Figure 2026073275000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0006] , , , , , , ,
[0005] , , ,
[0003] , , , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that it is difficult for beginners of musical instruments to find a practice method suitable for their physique and finger length, and it is difficult to receive individually optimized guidance.
[0005] The system according to the embodiment aims to provide an optimal performance form and finger technique based on the user's physique and finger length.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an acquisition unit, an analysis unit, a generation unit, a provision unit, and an evaluation unit. The acquisition unit acquires physical size data. The analysis unit analyzes the physical size data acquired by the acquisition unit. The generation unit generates optimal playing forms and fingering techniques based on the data analyzed by the analysis unit. The provision unit provides feedback generated by the generation unit. The evaluation unit evaluates the user's progress based on the feedback provided by the provision unit and proposes new practice methods and tasks. [Effects of the Invention]
[0007] The system according to this embodiment can provide the optimal playing form and fingering technique based on the user's physique and finger length. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The Personalized Play Coach, according to an embodiment of the present invention, is an AI system designed to solve the problems faced by beginner and intermediate musicians. Many beginner musicians often become discouraged because they cannot find practice methods that suit their personal characteristics, such as their physique and finger length. For example, it is common to hear stories of people giving up on guitar because they couldn't play the F chord. While it is possible to attend music lessons and receive instruction, whether or not one receives individually optimized instruction depends on the instructor's skills and experience. The Personalized Play Coach uses multimodal generative AI to provide individualized instruction on the optimal playing style and fingering techniques for the player, based on information such as the user's physique and finger length. This solves problems such as "holding the instrument in a way that puts a strain on the body, such as causing tendinitis" and "not being able to reach or move fingers properly," thus rescuing beginners who tend to give up at the beginning of learning to play an instrument. Furthermore, it is possible to provide comprehensive corrective instruction to improve the skills of intermediate musicians who tend to develop bad habits, taking into account their skeletal structure and range of motion of their joints. For example, a user takes a photo of their physique using their smartphone camera, and a generative AI automatically analyzes physical characteristics such as height, limb length, and joint positions to generate a personalized profile for the user. Next, when the user films a video of themselves playing an instrument with their smartphone, the generative AI analyzes the video and compares it to ideal playing form and fingering techniques. Based on the analysis results from the generative AI, specific feedback is generated. This feedback is overlaid on the user's performance using augmented reality (AR), making it intuitively understandable. The generative AI also regularly evaluates the user's progress and suggests new practice methods and challenges tailored to their skill level. This ensures that even after beginners have mastered basic skills, the system continues to provide advanced feedback and performance challenges for intermediate players. Personalized Play Coach prevents beginners from giving up and supports continued instrument playing by providing individually optimized instruction, always-available support, and efficient skill-building opportunities. This allows users to practice at their own pace and in a way that suits their challenges, enabling them to experience the joy of playing an instrument.This allows the personalized play coach to provide optimal playing form and fingering techniques based on the user's physical data, and to evaluate progress, enabling individually optimized instruction.
[0029] The personalized play coach according to this embodiment comprises an acquisition unit, an analysis unit, a generation unit, a provision unit, and an evaluation unit. The acquisition unit acquires the user's physical size data. The acquisition unit can acquire the user's physical size data, for example, using a smartphone camera. The acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions. The analysis unit analyzes the physical size data acquired by the acquisition unit. The analysis unit can generate a user-specific profile based on the acquired physical size data, for example. The analysis unit analyzes the user's physical size data in detail and generates basic data for providing individually optimized instruction. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. The generation unit can generate the ideal playing form and fingering technique, for example, by analyzing a user's performance video. The generation unit uses a generation AI to analyze a user's performance video and generate the optimal playing form and fingering technique. The provision unit provides the feedback generated by the generation unit. The providing unit can, for example, overlay feedback generated by the generation AI onto the user's performance using augmented reality (AR). The providing unit displays feedback using AR technology to make it easier for the user to intuitively understand the feedback. The evaluation unit evaluates the user's progress based on the feedback provided by the providing unit and proposes new practice methods and assignments. The evaluation unit can, for example, periodically evaluate the user's progress and propose new practice methods and assignments according to their skill level. The evaluation unit evaluates progress in detail in order to provide appropriate practice methods and assignments according to the user's skill level. As a result, the personalized play coach according to the embodiment can provide optimal playing form and fingering techniques based on the user's physical data and evaluate progress, enabling individually optimized instruction.
[0030] The data acquisition unit acquires the user's physical characteristics. For example, the data acquisition unit can acquire the user's physical characteristics using a smartphone camera. Specifically, it uses the smartphone camera to capture the user's entire body and uses image processing technology to automatically analyze the user's physical characteristics such as height, limb length, and joint positions. This utilizes an image recognition algorithm based on deep learning, enabling the acquisition of the user's physical characteristics with high accuracy. For example, by having the user stand in front of the smartphone camera and strike a specific pose, the data acquisition unit captures images from multiple angles and generates a 3D model. Based on this 3D model, the user's physical characteristics can be analyzed in detail. The data acquisition unit can also acquire data from wearable devices that the user uses daily. For example, by collecting data such as heart rate, steps taken, and calories burned from smartwatches and fitness trackers and integrating it with the user's physical characteristics data, a more comprehensive profile can be created. As a result, the data acquisition unit can collect user physical characteristics data from multiple angles and provide detailed and accurate data.
[0031] The analysis unit analyzes the physique data acquired by the acquisition unit. For example, the analysis unit can generate a user-specific profile based on the acquired physique data. Specifically, it analyzes the data sent from the acquisition unit to understand the user's physical characteristics in detail. This involves a process of classifying the user's physique data and extracting features using machine learning algorithms. For example, based on data such as the user's height, limb length, and joint positions, it generates basic data to identify the optimal playing form and fingering techniques for the user's physique. Using this data, the analysis unit creates a profile to provide individually optimized instruction tailored to the user's physique. This profile includes not only the user's physical characteristics but also past playing data and practice history. This allows the analysis unit to analyze the user's physique data in detail and generate basic data for providing individually optimized instruction. Furthermore, the analysis unit can continuously monitor the user's physique data and update the profile as needed. This allows the analysis unit to always provide optimal instruction based on the latest data.
[0032] The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. For example, the generation unit can analyze a user's performance video and generate the ideal playing form and fingering technique. Specifically, it uses a generation AI to analyze the user's performance video and generate the optimal playing form and fingering technique. The generation AI utilizes video analysis technology using deep learning to analyze the user's performance movements in detail. For example, it analyzes the user's hand movements, finger positions, and body posture when playing the instrument to identify the ideal playing form and fingering technique. Based on this data, the generation unit can generate the optimal playing form and fingering technique for the user's physique and playing style. Furthermore, the generation unit generates feedback to provide the user with the generated playing form and fingering technique. This includes specific performance instruction and practice method suggestions. For example, it can generate feedback that specifically shows practice methods for improving a particular performance technique or areas for improvement in playing form. As a result, the generation unit can generate the optimal playing form and fingering technique for the user's physique and playing style, and provide individually optimized instruction.
[0033] The service provider provides feedback generated by the generation unit. For example, the service provider can overlay feedback generated by the generation AI onto the user's performance using augmented reality (AR). Specifically, it uses AR technology to overlay feedback in real time onto the user's performance video. This allows the user to intuitively understand which parts need improvement while watching their own performance. For example, while the user is playing an instrument, the service provider can display real-time feedback on hand position and finger movements, showing correct playing form and fingering techniques. The service provider can use visual guides and animations to make the feedback easier for the user to understand intuitively. This makes it easier for the user to grasp specific areas for improvement while watching their own performance. The service provider can also record the user's response to the feedback and incorporate it into the next feedback. This allows the service provider to provide appropriate feedback according to the user's progress, achieving individually optimized instruction. Furthermore, the service provider can flexibly adjust the way feedback is displayed depending on the environment and device the user is using to receive the feedback. For example, it can provide feedback displays compatible with various devices such as smartphones, tablets, and AR glasses. This ensures that the service provider can receive optimal feedback in any environment.
[0034] The evaluation unit assesses the user's progress based on feedback provided by the service provider and proposes new practice methods and assignments. Specifically, the evaluation unit analyzes the user's performance data and feedback history to evaluate the user's skill level and progress in detail. This includes a process that uses machine learning algorithms to analyze the user's performance data and evaluate the degree of skill improvement and assignment completion. For example, it can evaluate how much time the user spent to acquire a particular performance technique and how accurately they were able to perform it. Based on this data, the evaluation unit proposes new practice methods and assignments tailored to the user's skill level. For example, it can specifically indicate practice methods for improving a particular performance technique and the next tasks the user should master. Furthermore, the evaluation unit regularly evaluates the user's progress and provides feedback to support the user in continuously improving their skills. This allows the evaluation unit to provide appropriate practice methods and assignments according to the user's skill level, achieving individually optimized instruction. In addition, the evaluation unit can perform more accurate evaluations by recording the user's responses to feedback and reflecting them in the next evaluation. This allows the evaluation unit to conduct a detailed assessment of the user's progress and generate foundational data for providing individually optimized instruction.
[0035] The service provider can overlay feedback generated by a generative AI onto the user's performance using augmented reality (AR). For example, the service provider can use a generative AI to analyze the user's performance video, generate ideal playing form and fingering techniques, and display that feedback in AR. The service provider uses AR technology to display the feedback in a way that makes it easier for the user to intuitively understand it. For example, the service provider can visually display ideal playing form and fingering techniques overlaid on the user's performance video. This allows the user to practice while comparing their own performance with the ideal playing form. Some or all of the above processing in the service provider may be performed using a generative AI, or it may be performed without a generative AI. For example, the service provider can display the feedback using a system that takes feedback generated by a generative AI as input and outputs an AR display. This makes it easier for the user to intuitively understand the feedback.
[0036] The evaluation unit can periodically assess the user's progress and suggest new practice methods and assignments tailored to their skill level. For example, the evaluation unit can periodically analyze the user's performance videos and assess their progress. The evaluation unit can suggest appropriate practice methods and assignments according to the user's skill level. For example, the evaluation unit can suggest basic practice methods for beginners. It can also suggest advanced practice methods and assignments for intermediate users. The evaluation unit provides detailed evaluations of the user's progress and appropriate feedback tailored to their skill level. This allows users to receive practice methods and assignments appropriate to their skill level, enabling them to improve their skills efficiently. Some or all of the above processes in the evaluation unit may be performed using AI, or not. For example, the evaluation unit can input the user's performance videos into an AI and have the AI perform the progress evaluation. This allows for the provision of appropriate practice methods and assignments tailored to the user's skill level.
[0037] The data acquisition unit can acquire user body size data using a smartphone camera. For example, the data acquisition unit can acquire data when a user takes a picture of their body with their smartphone camera. The data acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions, using the smartphone camera. The data acquisition unit provides a simple method using a smartphone camera so that users can easily acquire body size data. For example, the data acquisition unit can analyze image data taken with a smartphone camera and extract body size data. This allows users to easily acquire body size data without needing special equipment. Some or all of the above processing in the data acquisition unit may be performed using AI, or not. For example, the data acquisition unit can input image data taken with a smartphone camera into an AI and have the AI perform the analysis of body size data. This allows users to easily acquire body size data.
[0038] The analysis unit can generate a user-specific profile based on the acquired physique data. For example, the analysis unit can analyze the acquired physique data in detail and generate a user-specific profile. Based on the user's physique data, the analysis unit generates basic data for providing individually optimized instruction. For example, the analysis unit analyzes the user's physical characteristics such as height, limb length, and joint positions and generates a user-specific profile. This allows the user to receive playing form and fingering techniques that are optimal for their physique. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the acquired physique data into AI and have the AI generate a user-specific profile. This enables individually optimized instruction by generating a user-specific profile.
[0039] The generation unit can analyze the user's performance video and generate ideal playing form and fingering techniques. For example, the generation unit inputs the user's performance video into a generation AI and has the generation AI execute the ideal playing form and fingering techniques. The generation unit uses the generation AI to analyze the user's performance video and generate the optimal playing form and fingering techniques. For example, the generation unit can analyze the user's performance video in detail and generate ideal playing form and fingering techniques. This allows the user to practice while comparing their own performance with the ideal playing form. Some or all of the above processing in the generation unit may be performed using a generation AI, or without using a generation AI. For example, the generation unit can input the user's performance video into a generation AI and have the generation AI execute the generation of ideal playing form and fingering techniques. This allows the ideal playing form and fingering techniques to be provided by analyzing the user's performance video.
[0040] The acquisition unit can analyze the user's past physical size data and select the optimal acquisition method. For example, the acquisition unit can select the method that yielded the most accurate data based on the user's past physical size data. The acquisition unit can also analyze fluctuations in the user's past physical size data and select a stable data acquisition method. The acquisition unit can also consider the frequency of acquiring the user's past physical size data and select the optimal acquisition timing. This allows for the acquisition of physical size data in the most optimal way based on past data. Some or all of the above-described processes in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's past physical size data into AI and have the AI select the optimal acquisition method. This allows for the acquisition of physical size data in the most optimal way based on past data.
[0041] The acquisition unit can filter body size data based on the user's current health status and lifestyle. For example, the acquisition unit can consider the user's current health status and acquire body size data when the user is feeling well. The acquisition unit can also analyze the user's lifestyle and acquire body size data at the most appropriate time. The acquisition unit can also select the type of data to acquire based on the user's health status and lifestyle. This allows for the acquisition of data tailored to the user's health status and lifestyle. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's health data into AI and have the AI perform the filtering. This allows for the acquisition of data tailored to the user's health status and lifestyle.
[0042] The data acquisition unit can prioritize acquiring highly relevant data when acquiring body size data, taking into account the user's geographical location. For example, if the user is at high altitude, the data acquisition unit can prioritize acquiring data related to oxygen concentration and atmospheric pressure. If the user is in an urban area, the data acquisition unit can also prioritize acquiring data related to ambient noise and vibration. If the user is indoors, the data acquisition unit can also prioritize acquiring data related to indoor temperature and humidity. This allows for the acquisition of highly relevant data based on the user's geographical location. Some or all of the above processing in the data acquisition unit may be performed using AI, for example, or without AI. For example, the data acquisition unit can input the user's geographical location data into AI and have AI acquire highly relevant data. This allows for the acquisition of highly relevant data based on the user's geographical location.
[0043] The acquisition unit can analyze the user's social media activity and acquire relevant data when acquiring body size data. For example, if the user posts about exercise on social media, the acquisition unit can acquire body size data related to that activity. The acquisition unit can also acquire body size data based on information shared by the user about health on social media. The acquisition unit can also acquire body size data related to an event if the user participates in a specific event on social media. This allows for the acquisition of relevant data based on the user's social media activity. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's social media data into AI and have the AI acquire the relevant data. This allows for the acquisition of relevant data based on the user's social media activity.
[0044] The analysis unit can adjust the level of detail of the analysis based on the importance of the body size data during the analysis. For example, the analysis unit performs a detailed analysis on important body size data. The analysis unit can also perform a simplified analysis on basic body size data. The analysis unit can also perform a detailed analysis tailored to a specific purpose for body size data relevant to that purpose. This allows for detailed analysis of important data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI adjust the level of detail of the analysis based on importance. This allows for detailed analysis of important data.
[0045] The analysis unit can apply different analysis algorithms depending on the category of body size data during analysis. For example, the analysis unit can apply a skeletal analysis algorithm to skeletal data. The analysis unit can also apply a muscle analysis algorithm to muscle data. The analysis unit can also apply a joint analysis algorithm to joint data. This allows for appropriate analysis according to the data category. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI apply the appropriate analysis algorithm according to the category. This allows for appropriate analysis according to the data category.
[0046] The analysis unit can determine the priority of analysis based on when the body size data was acquired. For example, the analysis unit may prioritize the analysis of the most recent body size data. The analysis unit can also analyze the most recent data while referring to past body size data. The analysis unit can also prioritize the analysis of body size data acquired during a specific period. This allows for the prioritization of the analysis of the most recent data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI determine the priority of analysis based on the acquisition date. This allows for the prioritization of the analysis of the most recent data.
[0047] The analysis unit can adjust the order of analysis based on the relevance of the body size data during analysis. For example, the analysis unit may prioritize the analysis of important body size data. The analysis unit may also prioritize the analysis of highly relevant body size data. The analysis unit may also prioritize the analysis of body size data related to a specific purpose. This allows for the prioritization of highly relevant data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI adjust the order of analysis based on relevance. This allows for the prioritization of highly relevant data.
[0048] The generation unit can adjust the level of detail of the generated data based on the importance of the physical size data during generation. For example, the generation unit can generate detailed playing forms and fingering techniques based on important physical size data. The generation unit can also generate simplified playing forms and fingering techniques based on basic physical size data. The generation unit can also generate detailed playing forms and fingering techniques tailored to a specific purpose based on physical size data relevant to that purpose. This allows for the generation of detailed playing forms and fingering techniques based on important data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physical size data into AI and have the AI adjust the level of detail of the generated data based on its importance. This allows for the generation of detailed playing forms and fingering techniques based on important data.
[0049] The generation unit can apply different generation algorithms depending on the category of the physique data during generation. For example, the generation unit can generate playing forms and fingering techniques suitable for the skeleton based on skeletal data. The generation unit can also generate playing forms and fingering techniques suitable for the muscles based on muscle data. The generation unit can also generate playing forms and fingering techniques suitable for the joints based on joint data. This makes it possible to generate appropriate playing forms and fingering techniques according to the data category. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI execute the application of a generation algorithm according to the category. This makes it possible to generate appropriate playing forms and fingering techniques according to the data category.
[0050] The generation unit can determine the generation priority based on when the physique data was acquired during generation. For example, the generation unit can generate playing forms and fingering techniques based on the latest physique data. The generation unit can also generate playing forms and fingering techniques based on the latest data while referring to past physique data. The generation unit can also generate playing forms and fingering techniques based on physique data acquired during a specific period. This allows for the generation of playing forms and fingering techniques based on the latest data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI determine the generation priority based on the acquisition date. This allows for the generation of playing forms and fingering techniques based on the latest data.
[0051] The generation unit can adjust the generation order based on the relevance of the physique data during generation. For example, the generation unit can prioritize the generation of playing forms and fingering techniques based on important physique data. The generation unit can also prioritize the generation of playing forms and fingering techniques based on highly relevant physique data. The generation unit can also prioritize the generation of playing forms and fingering techniques based on physique data relevant to a specific purpose. This allows for the priority generation of playing forms and fingering techniques based on highly relevant data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI adjust the generation order based on relevance. This allows for the priority generation of playing forms and fingering techniques based on highly relevant data.
[0052] The service provider can select the optimal display method by referring to the user's past feedback history when providing feedback. For example, the service provider can select the most effective display method based on the user's past feedback history. The service provider can also analyze the user's past feedback history and select a display method that is easy to understand. The service provider can also adjust the display order of feedback by referring to the user's past feedback history. This allows the service provider to select the optimal display method based on past feedback history. Some or all of the above processes in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the user's feedback history into AI and have the AI select the optimal display method. This allows the service provider to select the optimal display method based on past feedback history.
[0053] The feedback system can customize the content of the feedback based on the user's current playing status when providing feedback. For example, the system can analyze the user's current playing status and provide appropriate feedback. The system can also adjust the level of detail of the feedback according to the user's playing status. The system can also determine the priority of the feedback based on the user's playing status. This allows for the provision of appropriate feedback according to the current playing status. Some or all of the above processes in the feedback system may be performed using AI, for example, or not using AI. For example, the system can input the user's playing data into AI and have the AI customize the feedback. This allows for the provision of appropriate feedback according to the current playing status.
[0054] The service provider can select the optimal feedback method when providing feedback, taking into account the user's geographical location. For example, if the user is outdoors, the service provider may prioritize providing visual feedback. If the user is indoors, the service provider may also provide detailed feedback. If the user is on the move, the service provider may also provide concise and to-the-point feedback. This allows the service provider to select the optimal feedback method based on geographical location information. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the user's geographical location data into AI and have the AI select the optimal feedback method. This allows the service provider to select the optimal feedback method based on geographical location information.
[0055] The service provider can analyze a user's social media activity and suggest methods for providing feedback when providing feedback. For example, if a user posts about exercise on social media, the service provider can provide feedback related to that activity. If a user shares health-related information on social media, the service provider can also provide feedback based on that information. If a user participates in a specific event on social media, the service provider can also provide feedback related to that event. This allows the service provider to suggest appropriate feedback methods based on social media activity. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the user's social media data into AI and have the AI suggest feedback methods. This allows the service provider to suggest appropriate feedback methods based on social media activity.
[0056] The evaluation unit can analyze the user's past performance data during evaluation to select the optimal evaluation method. For example, the evaluation unit can select the most effective evaluation method based on the user's past performance data. The evaluation unit can also analyze the user's past performance data to select an evaluation method that is easy to understand. The evaluation unit can also adjust the display order of evaluations based on the user's past performance data. This allows the evaluation unit to select the optimal evaluation method based on past performance data. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's performance data into AI and have the AI select the optimal evaluation method. This allows the evaluation unit to select the optimal evaluation method based on past performance data.
[0057] The evaluation unit can customize the evaluation criteria based on the user's current skill level during the evaluation process. For example, the evaluation unit can analyze the user's current skill level and set appropriate evaluation criteria. The evaluation unit can also adjust the level of detail in the evaluation according to the user's skill level. The evaluation unit can also determine the priority of the evaluation based on the user's skill level. This allows for the setting of appropriate evaluation criteria according to the current skill level. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or not using AI. For example, the evaluation unit can input the user's skill data into AI and have the AI perform the customization of the evaluation criteria. This allows for the setting of appropriate evaluation criteria according to the current skill level.
[0058] The evaluation unit can select the optimal evaluation method during evaluation, taking into account the user's geographical location information. For example, if the user is outdoors, the evaluation unit may prioritize providing a visual evaluation. If the user is indoors, the evaluation unit may also provide a detailed evaluation. If the user is on the move, the evaluation unit may also provide a concise and to-the-point evaluation. This allows the evaluation unit to select the optimal evaluation method based on geographical location information. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's geographical location data into AI and have the AI select the optimal evaluation method. This allows the evaluation unit to select the optimal evaluation method based on geographical location information.
[0059] The evaluation unit can analyze a user's social media activity during the evaluation process and propose evaluation methods. For example, if a user posts about exercise on social media, the evaluation unit can provide an evaluation related to that activity. If a user shares health-related information on social media, the evaluation unit can also provide an evaluation based on that information. If a user participates in a specific event on social media, the evaluation unit can also provide an evaluation related to that event. This allows the evaluation unit to propose appropriate evaluation methods based on social media activity. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or not using AI. For example, the evaluation unit can input the user's social media data into AI and have the AI propose evaluation methods. This allows the evaluation unit to propose appropriate evaluation methods based on social media activity.
[0060] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0061] The acquisition unit can acquire not only the user's physical size data but also their lifestyle data. For example, it can acquire data such as the user's sleep patterns, diet, and exercise habits, and use this data to estimate the user's physical condition and energy level. The analysis unit can generate an optimal training schedule that takes into account the user's physical condition and energy level based on the acquired lifestyle data. The generation unit generates training content that is appropriate for the user's physical condition and energy level based on the analysis results. The provision unit provides the generated training content to the user, supporting them so that they can continue training without overexerting themselves. This enables individually optimized guidance tailored to the user's lifestyle.
[0062] The evaluation unit can analyze the user's past practice data and select the optimal evaluation method. For example, it can select the most effective evaluation method based on the user's past practice data. The evaluation unit can also analyze the user's past practice data and select an evaluation method that is easy to understand. The evaluation unit can also adjust the display order of evaluations based on the user's past practice data. This allows for the selection of the optimal evaluation method based on past practice data. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's practice data into AI and have the AI select the optimal evaluation method. This allows for the selection of the optimal evaluation method based on past practice data.
[0063] The acquisition unit can acquire not only the user's physical size data but also their health data. For example, it can acquire data such as the user's heart rate, blood pressure, and body temperature, and estimate the user's health status based on this data. The analysis unit can generate an optimal training schedule that takes the user's health status into account based on the acquired health data. The generation unit generates training content that is appropriate for the user's health status based on the analysis results. The provision unit provides the generated training content to the user, supporting them so that they can continue training without overexerting themselves. This enables individually optimized instruction tailored to the user's health status.
[0064] The generation unit can analyze the user's performance video as well as their audio data to generate the optimal playing form and fingering technique. For example, it can acquire audio data from the user's performance and analyze the dynamics and rhythmic accuracy. Based on the analysis results of the audio data, the generation unit can adjust the user's playing form and fingering technique. The provision unit provides the generated playing form and fingering technique to the user, supporting them in improving their musical expressiveness. This enables individually optimized instruction utilizing audio data.
[0065] The evaluation unit can perform evaluations by considering the user's musical preferences and goals in addition to the user's performance data. For example, it can acquire the user's preferred music genres and desired performance style, and set evaluation criteria based on this information. The evaluation unit can adjust the level of detail of the evaluation and the content of the feedback according to the user's musical preferences and goals. This allows for the provision of appropriate evaluations tailored to the user's individual goals. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's musical preference and goal data into AI and have the AI set the evaluation criteria. This allows for the provision of appropriate evaluations tailored to the user's musical preferences and goals.
[0066] The data acquisition unit can acquire the user's geographical location information in addition to the user's physical size data. For example, if the user is at high altitude, it can acquire data related to oxygen concentration and atmospheric pressure. If the user is in an urban area, it can also acquire data related to ambient noise and vibration. If the user is indoors, it can also acquire data related to indoor temperature and humidity. This allows for the acquisition of highly relevant data based on the user's geographical location information. Some or all of the above-described processing in the data acquisition unit may be performed using AI, for example, or without AI. For example, the data acquisition unit can input the user's geographical location data into AI and have AI acquire highly relevant data. This allows for the acquisition of highly relevant data based on the user's geographical location information.
[0067] The following briefly describes the processing flow for example form 1.
[0068] Step 1: The acquisition unit acquires the user's physical size data. The acquisition unit can acquire the user's physical size data, for example, using a smartphone camera. The acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions. Step 2: The analysis unit analyzes the body size data acquired by the acquisition unit. The analysis unit can, for example, generate a user-specific profile based on the acquired body size data. The analysis unit analyzes the user's body size data in detail and generates basic data for providing individually optimized instruction. Step 3: The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. For example, the generation unit can analyze a user's performance video and generate the ideal playing form and fingering technique. The generation unit uses a generation AI to analyze the user's performance video and generate the optimal playing form and fingering technique. Step 4: The provider unit provides the feedback generated by the generator unit. The provider unit can, for example, overlay the feedback generated by the generation AI onto the user's performance using augmented reality (AR). The provider unit uses AR technology to display the feedback in a way that makes it easier for the user to intuitively understand it. Step 5: The evaluation unit assesses the user's progress based on the feedback provided by the delivery unit and proposes new practice methods and tasks. For example, the evaluation unit can periodically assess the user's progress and propose new practice methods and tasks appropriate to their skill level. The evaluation unit conducts a detailed assessment of progress in order to provide appropriate practice methods and tasks according to the user's skill level.
[0069] (Example of form 2) The Personalized Play Coach, according to an embodiment of the present invention, is an AI system designed to solve the problems faced by beginner and intermediate musicians. Many beginner musicians often become discouraged because they cannot find practice methods that suit their personal characteristics, such as their physique and finger length. For example, it is common to hear stories of people giving up on guitar because they couldn't play the F chord. While it is possible to attend music lessons and receive instruction, whether or not one receives individually optimized instruction depends on the instructor's skills and experience. The Personalized Play Coach uses multimodal generative AI to provide individualized instruction on the optimal playing style and fingering techniques for the player, based on information such as the user's physique and finger length. This solves problems such as "holding the instrument in a way that puts a strain on the body, such as causing tendinitis" and "not being able to reach or move fingers properly," thus rescuing beginners who tend to give up at the beginning of learning to play an instrument. Furthermore, it is possible to provide comprehensive corrective instruction to improve the skills of intermediate musicians who tend to develop bad habits, taking into account their skeletal structure and range of motion of their joints. For example, a user takes a photo of their physique using their smartphone camera, and a generative AI automatically analyzes physical characteristics such as height, limb length, and joint positions to generate a personalized profile for the user. Next, when the user films a video of themselves playing an instrument with their smartphone, the generative AI analyzes the video and compares it to ideal playing form and fingering techniques. Based on the analysis results from the generative AI, specific feedback is generated. This feedback is overlaid on the user's performance using augmented reality (AR), making it intuitively understandable. The generative AI also regularly evaluates the user's progress and suggests new practice methods and challenges tailored to their skill level. This ensures that even after beginners have mastered basic skills, the system continues to provide advanced feedback and performance challenges for intermediate players. Personalized Play Coach prevents beginners from giving up and supports continued instrument playing by providing individually optimized instruction, always-available support, and efficient skill-building opportunities. This allows users to practice at their own pace and in a way that suits their challenges, enabling them to experience the joy of playing an instrument.This allows the personalized play coach to provide optimal playing form and fingering techniques based on the user's physical data, and to evaluate progress, enabling individually optimized instruction.
[0070] The personalized play coach according to this embodiment comprises an acquisition unit, an analysis unit, a generation unit, a provision unit, and an evaluation unit. The acquisition unit acquires the user's physical size data. The acquisition unit can acquire the user's physical size data, for example, using a smartphone camera. The acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions. The analysis unit analyzes the physical size data acquired by the acquisition unit. The analysis unit can generate a user-specific profile based on the acquired physical size data, for example. The analysis unit analyzes the user's physical size data in detail and generates basic data for providing individually optimized instruction. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. The generation unit can generate the ideal playing form and fingering technique, for example, by analyzing a user's performance video. The generation unit uses a generation AI to analyze a user's performance video and generate the optimal playing form and fingering technique. The provision unit provides the feedback generated by the generation unit. The providing unit can, for example, overlay feedback generated by the generation AI onto the user's performance using augmented reality (AR). The providing unit displays feedback using AR technology to make it easier for the user to intuitively understand the feedback. The evaluation unit evaluates the user's progress based on the feedback provided by the providing unit and proposes new practice methods and assignments. The evaluation unit can, for example, periodically evaluate the user's progress and propose new practice methods and assignments according to their skill level. The evaluation unit evaluates progress in detail in order to provide appropriate practice methods and assignments according to the user's skill level. As a result, the personalized play coach according to the embodiment can provide optimal playing form and fingering techniques based on the user's physical data and evaluate progress, enabling individually optimized instruction.
[0071] The data acquisition unit acquires the user's physical characteristics. For example, the data acquisition unit can acquire the user's physical characteristics using a smartphone camera. Specifically, it uses the smartphone camera to capture the user's entire body and uses image processing technology to automatically analyze the user's physical characteristics such as height, limb length, and joint positions. This utilizes an image recognition algorithm based on deep learning, enabling the acquisition of the user's physical characteristics with high accuracy. For example, by having the user stand in front of the smartphone camera and strike a specific pose, the data acquisition unit captures images from multiple angles and generates a 3D model. Based on this 3D model, the user's physical characteristics can be analyzed in detail. The data acquisition unit can also acquire data from wearable devices that the user uses daily. For example, by collecting data such as heart rate, steps taken, and calories burned from smartwatches and fitness trackers and integrating it with the user's physical characteristics data, a more comprehensive profile can be created. As a result, the data acquisition unit can collect user physical characteristics data from multiple angles and provide detailed and accurate data.
[0072] The analysis unit analyzes the physique data acquired by the acquisition unit. For example, the analysis unit can generate a user-specific profile based on the acquired physique data. Specifically, it analyzes the data sent from the acquisition unit to understand the user's physical characteristics in detail. This involves a process of classifying the user's physique data and extracting features using machine learning algorithms. For example, based on data such as the user's height, limb length, and joint positions, it generates basic data to identify the optimal playing form and fingering techniques for the user's physique. Using this data, the analysis unit creates a profile to provide individually optimized instruction tailored to the user's physique. This profile includes not only the user's physical characteristics but also past playing data and practice history. This allows the analysis unit to analyze the user's physique data in detail and generate basic data for providing individually optimized instruction. Furthermore, the analysis unit can continuously monitor the user's physique data and update the profile as needed. This allows the analysis unit to always provide optimal instruction based on the latest data.
[0073] The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. For example, the generation unit can analyze a user's performance video and generate the ideal playing form and fingering technique. Specifically, it uses a generation AI to analyze the user's performance video and generate the optimal playing form and fingering technique. The generation AI utilizes video analysis technology using deep learning to analyze the user's performance movements in detail. For example, it analyzes the user's hand movements, finger positions, and body posture when playing the instrument to identify the ideal playing form and fingering technique. Based on this data, the generation unit can generate the optimal playing form and fingering technique for the user's physique and playing style. Furthermore, the generation unit generates feedback to provide the user with the generated playing form and fingering technique. This includes specific performance instruction and practice method suggestions. For example, it can generate feedback that specifically shows practice methods for improving a particular performance technique or areas for improvement in playing form. As a result, the generation unit can generate the optimal playing form and fingering technique for the user's physique and playing style, and provide individually optimized instruction.
[0074] The service provider provides feedback generated by the generation unit. For example, the service provider can overlay feedback generated by the generation AI onto the user's performance using augmented reality (AR). Specifically, it uses AR technology to overlay feedback in real time onto the user's performance video. This allows the user to intuitively understand which parts need improvement while watching their own performance. For example, while the user is playing an instrument, the service provider can display real-time feedback on hand position and finger movements, showing correct playing form and fingering techniques. The service provider can use visual guides and animations to make the feedback easier for the user to understand intuitively. This makes it easier for the user to grasp specific areas for improvement while watching their own performance. The service provider can also record the user's response to the feedback and incorporate it into the next feedback. This allows the service provider to provide appropriate feedback according to the user's progress, achieving individually optimized instruction. Furthermore, the service provider can flexibly adjust the way feedback is displayed depending on the environment and device the user is using to receive the feedback. For example, it can provide feedback displays compatible with various devices such as smartphones, tablets, and AR glasses. This ensures that the service provider can receive optimal feedback in any environment.
[0075] The evaluation unit assesses the user's progress based on feedback provided by the service provider and proposes new practice methods and assignments. Specifically, the evaluation unit analyzes the user's performance data and feedback history to evaluate the user's skill level and progress in detail. This includes a process that uses machine learning algorithms to analyze the user's performance data and evaluate the degree of skill improvement and assignment completion. For example, it can evaluate how much time the user spent to acquire a particular performance technique and how accurately they were able to perform it. Based on this data, the evaluation unit proposes new practice methods and assignments tailored to the user's skill level. For example, it can specifically indicate practice methods for improving a particular performance technique and the next tasks the user should master. Furthermore, the evaluation unit regularly evaluates the user's progress and provides feedback to support the user in continuously improving their skills. This allows the evaluation unit to provide appropriate practice methods and assignments according to the user's skill level, achieving individually optimized instruction. In addition, the evaluation unit can perform more accurate evaluations by recording the user's responses to feedback and reflecting them in the next evaluation. This allows the evaluation unit to conduct a detailed assessment of the user's progress and generate foundational data for providing individually optimized instruction.
[0076] The service provider can overlay feedback generated by a generative AI onto the user's performance using augmented reality (AR). For example, the service provider can use a generative AI to analyze the user's performance video, generate ideal playing form and fingering techniques, and display that feedback in AR. The service provider uses AR technology to display the feedback in a way that makes it easier for the user to intuitively understand it. For example, the service provider can visually display ideal playing form and fingering techniques overlaid on the user's performance video. This allows the user to practice while comparing their own performance with the ideal playing form. Some or all of the above processing in the service provider may be performed using a generative AI, or it may be performed without a generative AI. For example, the service provider can display the feedback using a system that takes feedback generated by a generative AI as input and outputs an AR display. This makes it easier for the user to intuitively understand the feedback.
[0077] The evaluation unit can periodically assess the user's progress and suggest new practice methods and assignments tailored to their skill level. For example, the evaluation unit can periodically analyze the user's performance videos and assess their progress. The evaluation unit can suggest appropriate practice methods and assignments according to the user's skill level. For example, the evaluation unit can suggest basic practice methods for beginners. It can also suggest advanced practice methods and assignments for intermediate users. The evaluation unit provides detailed evaluations of the user's progress and appropriate feedback tailored to their skill level. This allows users to receive practice methods and assignments appropriate to their skill level, enabling them to improve their skills efficiently. Some or all of the above processes in the evaluation unit may be performed using AI, or not. For example, the evaluation unit can input the user's performance videos into an AI and have the AI perform the progress evaluation. This allows for the provision of appropriate practice methods and assignments tailored to the user's skill level.
[0078] The data acquisition unit can acquire user body size data using a smartphone camera. For example, the data acquisition unit can acquire data when a user takes a picture of their body with their smartphone camera. The data acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions, using the smartphone camera. The data acquisition unit provides a simple method using a smartphone camera so that users can easily acquire body size data. For example, the data acquisition unit can analyze image data taken with a smartphone camera and extract body size data. This allows users to easily acquire body size data without needing special equipment. Some or all of the above processing in the data acquisition unit may be performed using AI, or not. For example, the data acquisition unit can input image data taken with a smartphone camera into an AI and have the AI perform the analysis of body size data. This allows users to easily acquire body size data.
[0079] The analysis unit can generate a user-specific profile based on the acquired physique data. For example, the analysis unit can analyze the acquired physique data in detail and generate a user-specific profile. Based on the user's physique data, the analysis unit generates basic data for providing individually optimized instruction. For example, the analysis unit analyzes the user's physical characteristics such as height, limb length, and joint positions and generates a user-specific profile. This allows the user to receive playing form and fingering techniques that are optimal for their physique. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the acquired physique data into AI and have the AI generate a user-specific profile. This enables individually optimized instruction by generating a user-specific profile.
[0080] The generation unit can analyze the user's performance video and generate ideal playing form and fingering techniques. For example, the generation unit inputs the user's performance video into a generation AI and has the generation AI execute the ideal playing form and fingering techniques. The generation unit uses the generation AI to analyze the user's performance video and generate the optimal playing form and fingering techniques. For example, the generation unit can analyze the user's performance video in detail and generate ideal playing form and fingering techniques. This allows the user to practice while comparing their own performance with the ideal playing form. Some or all of the above processing in the generation unit may be performed using a generation AI, or without using a generation AI. For example, the generation unit can input the user's performance video into a generation AI and have the generation AI execute the generation of ideal playing form and fingering techniques. This allows the ideal playing form and fingering techniques to be provided by analyzing the user's performance video.
[0081] The acquisition unit can estimate the user's emotions and adjust the timing of acquiring body size data based on the estimated user emotions. For example, if the user is relaxed, the acquisition unit can encourage relaxed shooting to acquire body size data in a natural posture. If the user is tense, the acquisition unit can also provide guidance to relieve tension and acquire body size data in a relaxed state. If the user is in a hurry, the acquisition unit can also provide a simplified procedure to quickly acquire body size data. This allows for the acquisition of body size data at the optimal timing according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows for the acquisition of body size data at the optimal timing according to the user's emotions.
[0082] The acquisition unit can analyze the user's past physical size data and select the optimal acquisition method. For example, the acquisition unit can select the method that yielded the most accurate data based on the user's past physical size data. The acquisition unit can also analyze fluctuations in the user's past physical size data and select a stable data acquisition method. The acquisition unit can also consider the frequency of acquiring the user's past physical size data and select the optimal acquisition timing. This allows for the acquisition of physical size data in the most optimal way based on past data. Some or all of the above-described processes in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's past physical size data into AI and have the AI select the optimal acquisition method. This allows for the acquisition of physical size data in the most optimal way based on past data.
[0083] The acquisition unit can filter body size data based on the user's current health status and lifestyle. For example, the acquisition unit can consider the user's current health status and acquire body size data when the user is feeling well. The acquisition unit can also analyze the user's lifestyle and acquire body size data at the most appropriate time. The acquisition unit can also select the type of data to acquire based on the user's health status and lifestyle. This allows for the acquisition of data tailored to the user's health status and lifestyle. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's health data into AI and have the AI perform the filtering. This allows for the acquisition of data tailored to the user's health status and lifestyle.
[0084] The acquisition unit can estimate the user's emotions and determine the priority of physique data to acquire based on the estimated user emotions. For example, if the user is relaxed, the acquisition unit may prioritize acquiring detailed physique data. If the user is tense, the acquisition unit may also prioritize acquiring basic physique data. If the user is in a hurry, the acquisition unit may also prioritize acquiring the most important physique data. This allows for the priority acquisition of important data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the acquisition unit may be performed using AI, or not using AI. For example, the acquisition unit can input the user's facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows for the priority acquisition of important data according to the user's emotions.
[0085] The data acquisition unit can prioritize acquiring highly relevant data when acquiring body size data, taking into account the user's geographical location. For example, if the user is at high altitude, the data acquisition unit can prioritize acquiring data related to oxygen concentration and atmospheric pressure. If the user is in an urban area, the data acquisition unit can also prioritize acquiring data related to ambient noise and vibration. If the user is indoors, the data acquisition unit can also prioritize acquiring data related to indoor temperature and humidity. This allows for the acquisition of highly relevant data based on the user's geographical location. Some or all of the above processing in the data acquisition unit may be performed using AI, for example, or without AI. For example, the data acquisition unit can input the user's geographical location data into AI and have AI acquire highly relevant data. This allows for the acquisition of highly relevant data based on the user's geographical location.
[0086] The acquisition unit can analyze the user's social media activity and acquire relevant data when acquiring body size data. For example, if the user posts about exercise on social media, the acquisition unit can acquire body size data related to that activity. The acquisition unit can also acquire body size data based on information shared by the user about health on social media. The acquisition unit can also acquire body size data related to an event if the user participates in a specific event on social media. This allows for the acquisition of relevant data based on the user's social media activity. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's social media data into AI and have the AI acquire the relevant data. This allows for the acquisition of relevant data based on the user's social media activity.
[0087] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the user is relaxed, the analysis unit can provide detailed analysis results. If the user is tense, the analysis unit can also provide concise and to-the-point analysis results. If the user is in a hurry, the analysis unit can also provide visual analysis results for quick understanding. This allows the presentation of the analysis results to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, or not using AI. For example, the analysis unit can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the presentation of the analysis results to be adjusted according to the user's emotions.
[0088] The analysis unit can adjust the level of detail of the analysis based on the importance of the body size data during the analysis. For example, the analysis unit performs a detailed analysis on important body size data. The analysis unit can also perform a simplified analysis on basic body size data. The analysis unit can also perform a detailed analysis tailored to a specific purpose for body size data relevant to that purpose. This allows for detailed analysis of important data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI adjust the level of detail of the analysis based on importance. This allows for detailed analysis of important data.
[0089] The analysis unit can apply different analysis algorithms depending on the category of body size data during analysis. For example, the analysis unit can apply a skeletal analysis algorithm to skeletal data. The analysis unit can also apply a muscle analysis algorithm to muscle data. The analysis unit can also apply a joint analysis algorithm to joint data. This allows for appropriate analysis according to the data category. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI apply the appropriate analysis algorithm according to the category. This allows for appropriate analysis according to the data category.
[0090] The analysis unit can estimate the user's emotions and adjust the length of the analysis based on the estimated emotions. For example, if the user is relaxed, the analysis unit can perform a detailed analysis and provide a longer report. If the user is tense, the analysis unit can also perform a concise analysis and provide a shorter report. If the user is in a hurry, the analysis unit can also perform a short, to-the-point analysis. This allows the length of the analysis to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, or not using AI. For example, the analysis unit can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the length of the analysis to be adjusted according to the user's emotions.
[0091] The analysis unit can determine the priority of analysis based on when the body size data was acquired. For example, the analysis unit may prioritize the analysis of the most recent body size data. The analysis unit can also analyze the most recent data while referring to past body size data. The analysis unit can also prioritize the analysis of body size data acquired during a specific period. This allows for the prioritization of the analysis of the most recent data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI determine the priority of analysis based on the acquisition date. This allows for the prioritization of the analysis of the most recent data.
[0092] The analysis unit can adjust the order of analysis based on the relevance of the body size data during analysis. For example, the analysis unit may prioritize the analysis of important body size data. The analysis unit may also prioritize the analysis of highly relevant body size data. The analysis unit may also prioritize the analysis of body size data related to a specific purpose. This allows for the prioritization of highly relevant data. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input body size data into AI and have the AI adjust the order of analysis based on relevance. This allows for the prioritization of highly relevant data.
[0093] The generation unit can estimate the user's emotions and adjust the way it expresses the performance form and fingering techniques based on the estimated emotions. For example, if the user is relaxed, the generation unit can provide detailed performance forms and fingering techniques. If the user is tense, the generation unit can also provide concise and to-the-point performance forms and fingering techniques. If the user is in a hurry, the generation unit can also provide visual performance forms and fingering techniques that can be quickly understood. This allows the expression of performance forms and fingering techniques to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation. This allows the expression of performance forms and fingering techniques to be adjusted according to the user's emotions.
[0094] The generation unit can adjust the level of detail of the generated data based on the importance of the physical size data during generation. For example, the generation unit can generate detailed playing forms and fingering techniques based on important physical size data. The generation unit can also generate simplified playing forms and fingering techniques based on basic physical size data. The generation unit can also generate detailed playing forms and fingering techniques tailored to a specific purpose based on physical size data relevant to that purpose. This allows for the generation of detailed playing forms and fingering techniques based on important data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physical size data into AI and have the AI adjust the level of detail of the generated data based on its importance. This allows for the generation of detailed playing forms and fingering techniques based on important data.
[0095] The generation unit can apply different generation algorithms depending on the category of the physique data during generation. For example, the generation unit can generate playing forms and fingering techniques suitable for the skeleton based on skeletal data. The generation unit can also generate playing forms and fingering techniques suitable for the muscles based on muscle data. The generation unit can also generate playing forms and fingering techniques suitable for the joints based on joint data. This makes it possible to generate appropriate playing forms and fingering techniques according to the data category. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI execute the application of a generation algorithm according to the category. This makes it possible to generate appropriate playing forms and fingering techniques according to the data category.
[0096] The generation unit can estimate the user's emotions and adjust the length of the generated playing forms and fingering techniques based on the estimated user emotions. For example, if the user is relaxed, the generation unit can provide detailed playing forms and fingering techniques. If the user is tense, the generation unit can also provide concise and to-the-point playing forms and fingering techniques. If the user is in a hurry, the generation unit can also provide visual playing forms and fingering techniques for quick understanding. This allows the length of the playing forms and fingering techniques to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation. This allows the length of the playing forms and fingering techniques to be adjusted according to the user's emotions.
[0097] The generation unit can determine the generation priority based on when the physique data was acquired during generation. For example, the generation unit can generate playing forms and fingering techniques based on the latest physique data. The generation unit can also generate playing forms and fingering techniques based on the latest data while referring to past physique data. The generation unit can also generate playing forms and fingering techniques based on physique data acquired during a specific period. This allows for the generation of playing forms and fingering techniques based on the latest data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI determine the generation priority based on the acquisition date. This allows for the generation of playing forms and fingering techniques based on the latest data.
[0098] The generation unit can adjust the generation order based on the relevance of the physique data during generation. For example, the generation unit can prioritize the generation of playing forms and fingering techniques based on important physique data. The generation unit can also prioritize the generation of playing forms and fingering techniques based on highly relevant physique data. The generation unit can also prioritize the generation of playing forms and fingering techniques based on physique data relevant to a specific purpose. This allows for the priority generation of playing forms and fingering techniques based on highly relevant data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input physique data into AI and have the AI adjust the generation order based on relevance. This allows for the priority generation of playing forms and fingering techniques based on highly relevant data.
[0099] The service provider can estimate the user's emotions and adjust how feedback is displayed based on the estimated emotions. For example, if the user is relaxed, the service provider can provide detailed feedback. If the user is stressed, the service provider can also provide concise and to-the-point feedback. If the user is in a hurry, the service provider can also provide visual feedback for quick understanding. This allows the service provider to adjust how feedback is displayed according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI or not using AI. For example, the service provider can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the service provider to adjust how feedback is displayed according to the user's emotions.
[0100] The service provider can select the optimal display method by referring to the user's past feedback history when providing feedback. For example, the service provider can select the most effective display method based on the user's past feedback history. The service provider can also analyze the user's past feedback history and select a display method that is easy to understand. The service provider can also adjust the display order of feedback by referring to the user's past feedback history. This allows the service provider to select the optimal display method based on past feedback history. Some or all of the above processes in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the user's feedback history into AI and have the AI select the optimal display method. This allows the service provider to select the optimal display method based on past feedback history.
[0101] The feedback system can customize the content of the feedback based on the user's current playing status when providing feedback. For example, the system can analyze the user's current playing status and provide appropriate feedback. The system can also adjust the level of detail of the feedback according to the user's playing status. The system can also determine the priority of the feedback based on the user's playing status. This allows for the provision of appropriate feedback according to the current playing status. Some or all of the above processes in the feedback system may be performed using AI, for example, or not using AI. For example, the system can input the user's playing data into AI and have the AI customize the feedback. This allows for the provision of appropriate feedback according to the current playing status.
[0102] The service provider can estimate the user's emotions and prioritize feedback based on those emotions. For example, if the user is relaxed, the service provider may prioritize detailed feedback. If the user is stressed, the service provider may prioritize concise and to-the-point feedback. If the user is in a hurry, the service provider may prioritize visual feedback for quick understanding. This allows for the priority of important feedback according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI or not. For example, the service provider can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows for the priority of important feedback according to the user's emotions.
[0103] The service provider can select the optimal feedback method when providing feedback, taking into account the user's geographical location. For example, if the user is outdoors, the service provider may prioritize providing visual feedback. If the user is indoors, the service provider may also provide detailed feedback. If the user is on the move, the service provider may also provide concise and to-the-point feedback. This allows the service provider to select the optimal feedback method based on geographical location information. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the user's geographical location data into AI and have the AI select the optimal feedback method. This allows the service provider to select the optimal feedback method based on geographical location information.
[0104] The service provider can analyze a user's social media activity and suggest methods for providing feedback when providing feedback. For example, if a user posts about exercise on social media, the service provider can provide feedback related to that activity. If a user shares health-related information on social media, the service provider can also provide feedback based on that information. If a user participates in a specific event on social media, the service provider can also provide feedback related to that event. This allows the service provider to suggest appropriate feedback methods based on social media activity. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the user's social media data into AI and have the AI suggest feedback methods. This allows the service provider to suggest appropriate feedback methods based on social media activity.
[0105] The evaluation unit can estimate the user's emotions and adjust the evaluation method based on the estimated emotions. For example, if the user is relaxed, the evaluation unit can provide a detailed evaluation. If the user is tense, the evaluation unit can also provide a concise and to-the-point evaluation. If the user is in a hurry, the evaluation unit can also provide a visual evaluation that can be quickly understood. This allows the evaluation method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or not using AI. For example, the evaluation unit can input the user's facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the evaluation method to be adjusted according to the user's emotions.
[0106] The evaluation unit can analyze the user's past performance data during evaluation to select the optimal evaluation method. For example, the evaluation unit can select the most effective evaluation method based on the user's past performance data. The evaluation unit can also analyze the user's past performance data to select an evaluation method that is easy to understand. The evaluation unit can also adjust the display order of evaluations based on the user's past performance data. This allows the evaluation unit to select the optimal evaluation method based on past performance data. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's performance data into AI and have the AI select the optimal evaluation method. This allows the evaluation unit to select the optimal evaluation method based on past performance data.
[0107] The evaluation unit can customize the evaluation criteria based on the user's current skill level during the evaluation process. For example, the evaluation unit can analyze the user's current skill level and set appropriate evaluation criteria. The evaluation unit can also adjust the level of detail in the evaluation according to the user's skill level. The evaluation unit can also determine the priority of the evaluation based on the user's skill level. This allows for the setting of appropriate evaluation criteria according to the current skill level. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or not using AI. For example, the evaluation unit can input the user's skill data into AI and have the AI perform the customization of the evaluation criteria. This allows for the setting of appropriate evaluation criteria according to the current skill level.
[0108] The evaluation unit can estimate the user's emotions and determine the priority of evaluations based on the estimated emotions. For example, if the user is relaxed, the evaluation unit may prioritize providing detailed evaluations. If the user is tense, the evaluation unit may also prioritize providing concise and to-the-point evaluations. If the user is in a hurry, the evaluation unit may also prioritize providing visual evaluations that can be quickly understood. This allows for the prioritization of important evaluations according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the evaluation unit may be performed using AI or not using AI. For example, the evaluation unit can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows for the prioritization of important evaluations according to the user's emotions.
[0109] The evaluation unit can select the optimal evaluation method during evaluation, taking into account the user's geographical location information. For example, if the user is outdoors, the evaluation unit may prioritize providing a visual evaluation. If the user is indoors, the evaluation unit may also provide a detailed evaluation. If the user is on the move, the evaluation unit may also provide a concise and to-the-point evaluation. This allows the evaluation unit to select the optimal evaluation method based on geographical location information. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's geographical location data into AI and have the AI select the optimal evaluation method. This allows the evaluation unit to select the optimal evaluation method based on geographical location information.
[0110] The evaluation unit can analyze a user's social media activity during the evaluation process and propose evaluation methods. For example, if a user posts about exercise on social media, the evaluation unit can provide an evaluation related to that activity. If a user shares health-related information on social media, the evaluation unit can also provide an evaluation based on that information. If a user participates in a specific event on social media, the evaluation unit can also provide an evaluation related to that event. This allows the evaluation unit to propose appropriate evaluation methods based on social media activity. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or not using AI. For example, the evaluation unit can input the user's social media data into AI and have the AI propose evaluation methods. This allows the evaluation unit to propose appropriate evaluation methods based on social media activity.
[0111] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0112] The acquisition unit can acquire not only the user's physical size data but also their lifestyle data. For example, it can acquire data such as the user's sleep patterns, diet, and exercise habits, and use this data to estimate the user's physical condition and energy level. The analysis unit can generate an optimal training schedule that takes into account the user's physical condition and energy level based on the acquired lifestyle data. The generation unit generates training content that is appropriate for the user's physical condition and energy level based on the analysis results. The provision unit provides the generated training content to the user, supporting them so that they can continue training without overexerting themselves. This enables individually optimized guidance tailored to the user's lifestyle.
[0113] The service provider can estimate the user's emotions and adjust the content of the feedback based on the estimated emotions. For example, if the user is relaxed, detailed feedback can be provided. If the user is tense, concise and to-the-point feedback can be provided. If the user is in a hurry, visual feedback can be provided for quick understanding. This allows for the provision of appropriate feedback according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, or not using AI. For example, the service provider can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows for the provision of appropriate feedback according to the user's emotions.
[0114] The evaluation unit can analyze the user's past practice data and select the optimal evaluation method. For example, it can select the most effective evaluation method based on the user's past practice data. The evaluation unit can also analyze the user's past practice data and select an evaluation method that is easy to understand. The evaluation unit can also adjust the display order of evaluations based on the user's past practice data. This allows for the selection of the optimal evaluation method based on past practice data. Some or all of the above processes in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's practice data into AI and have the AI select the optimal evaluation method. This allows for the selection of the optimal evaluation method based on past practice data.
[0115] The acquisition unit can acquire not only the user's physical size data but also their health data. For example, it can acquire data such as the user's heart rate, blood pressure, and body temperature, and estimate the user's health status based on this data. The analysis unit can generate an optimal training schedule that takes the user's health status into account based on the acquired health data. The generation unit generates training content that is appropriate for the user's health status based on the analysis results. The provision unit provides the generated training content to the user, supporting them so that they can continue training without overexerting themselves. This enables individually optimized instruction tailored to the user's health status.
[0116] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the user is relaxed, it can provide detailed analysis results. If the user is tense, it can provide concise and to-the-point analysis results. If the user is in a hurry, it can provide visual analysis results for quick understanding. This allows the presentation of the analysis results to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the presentation of the analysis results to be adjusted according to the user's emotions.
[0117] The generation unit can analyze the user's performance video as well as their audio data to generate the optimal playing form and fingering technique. For example, it can acquire audio data from the user's performance and analyze the dynamics and rhythmic accuracy. Based on the analysis results of the audio data, the generation unit can adjust the user's playing form and fingering technique. The provision unit provides the generated playing form and fingering technique to the user, supporting them in improving their musical expressiveness. This enables individually optimized instruction utilizing audio data.
[0118] The acquisition unit can estimate the user's emotions and adjust the timing of acquiring body size data based on the estimated user emotions. For example, if the user is relaxed, it can encourage them to take photos in a relaxed state to acquire body size data in a natural posture. If the user is tense, it can provide guidance to relieve tension and acquire body size data in a relaxed state. If the user is in a hurry, it can provide a simplified procedure to quickly acquire body size data. This allows for the acquisition of body size data at the optimal timing according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the acquisition unit may be performed using AI, or not using AI. For example, the acquisition unit can input the user's facial expression data into the generative AI and have the generative AI perform emotion estimation. This allows for the acquisition of body size data at the optimal timing according to the user's emotions.
[0119] The evaluation unit can perform evaluations by considering the user's musical preferences and goals in addition to the user's performance data. For example, it can acquire the user's preferred music genres and desired performance style, and set evaluation criteria based on this information. The evaluation unit can adjust the level of detail of the evaluation and the content of the feedback according to the user's musical preferences and goals. This allows for the provision of appropriate evaluations tailored to the user's individual goals. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's musical preference and goal data into AI and have the AI set the evaluation criteria. This allows for the provision of appropriate evaluations tailored to the user's musical preferences and goals.
[0120] The service provider can estimate the user's emotions and adjust how feedback is displayed based on the estimated emotions. For example, if the user is relaxed, detailed feedback can be provided. If the user is tense, concise and to-the-point feedback can be provided. If the user is in a hurry, visual feedback can be provided for quick understanding. This allows the feedback display method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI or not using AI. For example, the service provider can input user facial expression data into a generative AI and have the generative AI perform emotion estimation. This allows the feedback display method to be adjusted according to the user's emotions.
[0121] The data acquisition unit can acquire the user's geographical location information in addition to the user's physical size data. For example, if the user is at high altitude, it can acquire data related to oxygen concentration and atmospheric pressure. If the user is in an urban area, it can also acquire data related to ambient noise and vibration. If the user is indoors, it can also acquire data related to indoor temperature and humidity. This allows for the acquisition of highly relevant data based on the user's geographical location information. Some or all of the above-described processing in the data acquisition unit may be performed using AI, for example, or without AI. For example, the data acquisition unit can input the user's geographical location data into AI and have AI acquire highly relevant data. This allows for the acquisition of highly relevant data based on the user's geographical location information.
[0122] The following briefly describes the processing flow for example form 2.
[0123] Step 1: The acquisition unit acquires the user's physical size data. The acquisition unit can acquire the user's physical size data, for example, using a smartphone camera. The acquisition unit automatically analyzes the user's physical characteristics, such as height, limb length, and joint positions. Step 2: The analysis unit analyzes the body size data acquired by the acquisition unit. The analysis unit can, for example, generate a user-specific profile based on the acquired body size data. The analysis unit analyzes the user's body size data in detail and generates basic data for providing individually optimized instruction. Step 3: The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the analysis unit. For example, the generation unit can analyze a user's performance video and generate the ideal playing form and fingering technique. The generation unit uses a generation AI to analyze the user's performance video and generate the optimal playing form and fingering technique. Step 4: The provider unit provides the feedback generated by the generator unit. The provider unit can, for example, overlay the feedback generated by the generation AI onto the user's performance using augmented reality (AR). The provider unit uses AR technology to display the feedback in a way that makes it easier for the user to intuitively understand it. Step 5: The evaluation unit assesses the user's progress based on the feedback provided by the delivery unit and proposes new practice methods and tasks. For example, the evaluation unit can periodically assess the user's progress and propose new practice methods and tasks appropriate to their skill level. The evaluation unit conducts a detailed assessment of progress in order to provide appropriate practice methods and tasks according to the user's skill level.
[0124] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0125] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0126] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0127] Each of the multiple elements described above, including the acquisition unit, analysis unit, generation unit, provision unit, and evaluation unit, is implemented, for example, in at least one of the smart device 14 and the data processing unit 12. For example, the acquisition unit acquires the user's physique data using the camera 42 of the smart device 14. The analysis unit analyzes the acquired physique data by the specific processing unit 290 of the data processing unit 12 and generates a user-specific profile. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the specific processing unit 290 of the data processing unit 12. The provision unit displays the feedback generated by the control unit 46A of the smart device 14 overlaid on the user's performance using augmented reality. The evaluation unit evaluates the user's progress using the specific processing unit 290 of the data processing unit 12 and proposes new practice methods and tasks. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0128] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0129] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0130] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0131] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0132] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0133] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0134] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0135] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0136] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0137] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0138] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0139] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0140] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0141] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0142] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0143] Each of the multiple elements described above, including the acquisition unit, analysis unit, generation unit, provision unit, and evaluation unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the acquisition unit acquires the user's physique data using the camera 42 of the smart glasses 214. The analysis unit analyzes the physique data acquired by the specific processing unit 290 of the data processing unit 12 and generates a user-specific profile. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the specific processing unit 290 of the data processing unit 12. The provision unit displays the feedback generated by the control unit 46A of the smart glasses 214 overlaid on the user's performance using augmented reality. The evaluation unit evaluates the user's progress using the specific processing unit 290 of the data processing unit 12 and proposes new practice methods and tasks. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0144] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0145] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0146] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0147] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0148] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0149] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0150] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0151] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0152] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0153] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0154] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0155] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0156] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0157] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0158] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0159] Each of the multiple elements described above, including the acquisition unit, analysis unit, generation unit, provision unit, and evaluation unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the acquisition unit acquires the user's physique data using the camera 42 of the headset terminal 314. The analysis unit analyzes the physique data acquired by the specific processing unit 290 of the data processing unit 12 and generates a user-specific profile. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the specific processing unit 290 of the data processing unit 12. The provision unit displays the feedback generated by the control unit 46A of the headset terminal 314 overlaid on the user's performance using augmented reality (AR). The evaluation unit evaluates the user's progress using the specific processing unit 290 of the data processing unit 12 and proposes new practice methods and tasks. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0160] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0161] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0162] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0163] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0164] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0165] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0166] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0167] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0168] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0169] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0170] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0171] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0172] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0173] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0174] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0175] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0176] Each of the multiple elements described above, including the acquisition unit, analysis unit, generation unit, provision unit, and evaluation unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the acquisition unit acquires the user's physique data using the camera 42 of the robot 414. The analysis unit analyzes the physique data acquired by the specific processing unit 290 of the data processing unit 12 and generates a user-specific profile. The generation unit generates the optimal playing form and fingering technique based on the data analyzed by the specific processing unit 290 of the data processing unit 12. The provision unit displays the feedback generated by the control unit 46A of the robot 414 overlaid on the user's performance using augmented reality. The evaluation unit evaluates the user's progress using the specific processing unit 290 of the data processing unit 12 and proposes new practice methods and tasks. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0177] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0178] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0179] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0180] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0181] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0182] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0183] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0184] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0185] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0186] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0187] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0188] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0189] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0190] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0191] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0192] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0193] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0194] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0195] (Note 1) An acquisition unit that acquires body size data, An analysis unit analyzes the body size data acquired by the acquisition unit, A generation unit generates the optimal playing form and fingering technique based on the data analyzed by the aforementioned analysis unit, A providing unit that provides the feedback generated by the generation unit, The system includes an evaluation unit that evaluates the user's progress based on the feedback provided by the aforementioned provision unit and proposes new practice methods and tasks. A system characterized by the following features. (Note 2) The aforementioned supply unit is, The AI-generated feedback is overlaid onto the user's performance using augmented reality (AR). The system described in Appendix 1, characterized by the features described herein. (Note 3) The evaluation unit, Regularly evaluate user progress and suggest new practice methods and challenges tailored to their skill level. The system described in Appendix 1, characterized by the features described herein. (Note 4) The acquisition unit is, The system uses the smartphone's camera to acquire the user's body size data. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned analysis unit, A user-specific profile is generated based on the acquired body size data. The system described in Appendix 1, characterized by the features described herein. (Note 6) The generating unit is It analyzes user performance videos and generates ideal playing form and fingering techniques. The system described in Appendix 1, characterized by the features described herein. (Note 7) The acquisition unit is, The system estimates the user's emotions and adjusts the timing of acquiring body size data based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The acquisition unit is, Analyze the user's past body size data and select the optimal acquisition method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The acquisition unit is, When acquiring body size data, filtering is performed based on the user's current health status and lifestyle. The system described in Appendix 1, characterized by the features described herein. (Note 10) The acquisition unit is, The system estimates the user's emotions and determines the priority of body size data to acquire based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The acquisition unit is, When acquiring body size data, the system prioritizes acquiring highly relevant data by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The acquisition unit is, When acquiring body size data, we analyze the user's social media activity and obtain related data. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the representation of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the level of detail of the analysis is adjusted based on the importance of the body size data. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of body size data. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, It estimates the user's emotions and adjusts the length of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when the body size data was acquired. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of body size data. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is It estimates the user's emotions and adjusts the way performance forms and fingering techniques are expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is During generation, adjust the level of detail based on the importance of the body size data. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is During generation, different generation algorithms are applied depending on the category of body size data. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is It estimates the user's emotions and adjusts the length of the playing form and fingering techniques generated based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is During generation, the generation priority is determined based on when the body size data was acquired. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is During generation, the generation order is adjusted based on the relevance of body size data. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned supply unit is, It estimates the user's emotions and adjusts how feedback is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned supply unit is, When providing feedback, the system will refer to the user's past feedback history to select the most suitable display method. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned supply unit is, When providing feedback, the content of the feedback will be customized based on the user's current playing status. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned supply unit is, It estimates the user's emotions and prioritizes feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned supply unit is, When providing feedback, the most suitable feedback method will be selected, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned supply unit is, When providing feedback, we analyze the user's social media activity and suggest methods for providing feedback. The system described in Appendix 1, characterized by the features described herein. (Note 31) The evaluation unit, It estimates the user's emotions and adjusts the evaluation method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The evaluation unit, During evaluation, the system analyzes the user's past performance data to select the optimal evaluation method. The system described in Appendix 1, characterized by the features described herein. (Note 33) The evaluation unit, During the evaluation process, the evaluation criteria are customized based on the user's current skill level. The system described in Appendix 1, characterized by the features described herein. (Note 34) The evaluation unit, It estimates the user's emotions and determines the priority of evaluations based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The evaluation unit, During evaluation, the optimal evaluation method is selected by considering the user's geographical location information. The system described in Appendix 1, characterized by the features described herein. (Note 36) The evaluation unit, During the evaluation process, we will analyze users' social media activity and propose evaluation methods. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0196] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. An acquisition unit that acquires body size data, An analysis unit analyzes the body size data acquired by the acquisition unit, A generation unit generates the optimal playing form and fingering technique based on the data analyzed by the aforementioned analysis unit, A providing unit that provides the feedback generated by the generation unit, The system includes an evaluation unit that evaluates the user's progress based on the feedback provided by the aforementioned provision unit and proposes new practice methods and tasks. A system characterized by the following features.
2. The aforementioned supply unit is, The AI-generated feedback is overlaid onto the user's performance using augmented reality (AR). The system according to feature 1.
3. The evaluation unit, Regularly evaluate user progress and suggest new practice methods and challenges tailored to their skill level. The system according to feature 1.
4. The acquisition unit is, The system uses the smartphone's camera to acquire the user's body size data. The system according to feature 1.
5. The aforementioned analysis unit, A user-specific profile is generated based on the acquired body size data. The system according to feature 1.
6. The generating unit is It analyzes user performance videos and generates ideal playing form and fingering techniques. The system according to feature 1.
7. The acquisition unit is, The system estimates the user's emotions and adjusts the timing of acquiring body size data based on the estimated emotions. The system according to feature 1.
8. The acquisition unit is, Analyze the user's past body size data and select the optimal acquisition method. The system according to feature 1.
9. The acquisition unit is, When acquiring body size data, filtering is performed based on the user's current health status and lifestyle. The system according to feature 1.
10. The acquisition unit is, The system estimates the user's emotions and determines the priority of body size data to acquire based on the estimated user emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A