system

The system addresses the challenge of individualized music learning by using AI to analyze performance, generate tailored practice menus, and offer online instruction, enhancing skill improvement and inclusivity for all learners.

JP2026068398APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional music learning methods require direct teacher guidance, making individualized instruction difficult, and lack effective evaluation of performance, particularly for learners with hearing impairments, leading to inefficient skill improvement and potential continuation of incorrect techniques.

Method used

A system that uses AI to analyze user performance via camera, evaluates form, generates personalized practice menus, incorporates game elements, and provides online instruction and sign language recognition for inclusive support.

Benefits of technology

Provides an efficient and engaging learning environment that improves musical skills through personalized practice menus and real-time feedback, accommodating diverse learning needs including those with hearing impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068398000001_ABST
    Figure 2026068398000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A terminal means equipped with a camera for capturing the user's performance movements, A server means equipped with an analysis device that analyzes acquired video data of performance movements and extracts the user's skeletal information, A server means equipped with a control device that evaluates the user's playing form based on extracted skeletal information and generates a practice menu according to the evaluation results, A terminal device that presents the generated practice menu as a user interface and provides a learning environment that incorporates game elements, A server means equipped with a communication device that provides online instruction and feedback to users as they progress through their musical activities, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] [In the conventional music learning method, direct guidance by a teacher is required, individual guidance is difficult, and it is a problem that the learner's performance ability and technical improvement cannot be efficiently carried out. In addition, since the learner has limited opportunities to receive a detailed analysis of their own performance form, there is a risk of continuing to practice with an incorrect form. In particular, there is a lack of a learning environment that takes into account the accessibility of people with hearing impairments, etc., and it is required to meet various learning needs including these.]

Means for Solving the Problems

[0005] [This invention provides a system that acquires a user's performance movements using a camera and extracts the user's skeletal information by analyzing the video data using AI technology. Based on this analysis, the system evaluates the performance form and generates individually optimized practice menus, thereby promoting effective learning. Furthermore, the generated practice menus are presented on a user interface incorporating game elements to motivate learning. In addition, the system is equipped with communication functions that enable online instruction and feedback, and solves the above-mentioned problems by providing inclusive support functions using sign language recognition technology, particularly for people with hearing impairments.]

[0006] A "recording device" is a device used to capture the user's performance actions as video.

[0007] A "terminal" is a device used by the user to control the camera and to display practice menus.

[0008] An "analysis device" is a device installed on the server side that analyzes acquired video data and extracts the user's skeletal information.

[0009] A "server" is a primary computing device that performs data analysis and control, and communicates with multiple terminals to manage and operate the entire performance learning system.

[0010] A "control device" is a device that evaluates the user's playing form based on analyzed skeletal information and generates and provides practice menus.

[0011] A "user interface" is a function that provides screens and control systems for users to visually check and operate practice menus and instructional content.

[0012] A "communication device" is a device that sends and receives data between the user and the server, enabling online lessons and feedback.

[0013] "Sign language recognition technology" is a technology that recognizes sign language as visual information for people with hearing impairments, and then analyzes and reflects its content.

[0014] "Game elements" refer to competitive or reward-based features introduced to make user learning enjoyable and sustainable. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Embodiment for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a music learning system designed to improve the user's performance skills. The user records their performance via a recording device connected to a terminal. The terminal has the function of transmitting this video data to a server.

[0037] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This skeletal information allows for an accurate evaluation of the user's playing form. The analysis results are compared on the server with exemplary form and used to identify areas for improvement and strengths of the user.

[0038] Based on the evaluation of the user's playing form, the server generates an individually optimized practice menu. This menu includes a variety of exercises aimed at improving playing technique, and is composed of appropriate exercises selected from a database managed by the server.

[0039] The generated practice menu is presented to the user via the device. The device uses a user interface incorporating game elements to visualize learning progress and enhance the user's motivation to learn. Furthermore, users can receive online instruction at any time, and the server supports real-time communication for this purpose.

[0040] Furthermore, as a consideration for people with hearing impairments, the terminal utilizes sign language recognition technology to provide an environment where users can confirm necessary information in sign language. The server generates appropriate feedback corresponding to this and presents it through the terminal.

[0041] For example, when a user learns to play the guitar, the performance data captured by the camera is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server then suggests specific practice menus regarding finger movements and strokes. These are displayed on the device, allowing the user to improve their playing skills in an enjoyable way while receiving instruction.

[0042] In this way, the present invention provides users with a personalized and effective music learning experience, supporting the improvement of their performance skills. This system creates a more efficient and engaging learning environment than conventional methods.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user activates the device's recording function and prepares to record their musical performance. Once ready, the device starts recording, capturing the performance in real time.

[0046] Step 2:

[0047] The device sends recorded video data to the server. This transmission is primarily done over the network, enabling real-time data reception.

[0048] Step 3:

[0049] The server receives video data transmitted from the terminal and performs skeletal detection using an analysis device. This extracts skeletal information such as the user's hands and arms.

[0050] Step 4:

[0051] The server compares the user's playing form to a model form based on the extracted skeletal information. It then identifies areas that need improvement based on the form's evaluation criteria.

[0052] Step 5:

[0053] The server automatically generates individually optimized practice menus based on the form evaluation results. These menus will include specific technical exercises and formal drills.

[0054] Step 6:

[0055] The device presents the generated practice menu to the user. A gamified user interface is used to enhance learning motivation during the presentation.

[0056] Step 7:

[0057] The user practices based on the provided menu, records their performance, and sends it from their device to the server. They then receive feedback from the server.

[0058] Step 8:

[0059] The server evaluates the video data sent back by the user and provides online direct instruction and feedback as needed. For users with hearing impairments, it also uses sign language recognition technology to provide appropriate feedback.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] [To improve musical performance skills, there is a need for a system that provides efficient practice menus tailored to individual user characteristics, and enables evaluation and improvement of performance. Furthermore, an environment is needed where multiple learners, each in different circumstances, can receive appropriate feedback online.]

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for [equipped with an information processing device for extracting target operation information], means for [equipped with a control device for generating a practice plan], and means for [equipped with a communication device for providing instruction and feedback over a communication network]. This enables [the provision of personalized practice menus to users and efficient performance evaluation and instruction via online means].

[0065] An "image acquisition device" refers to a device used to acquire video footage of a subject's performance.

[0066] An "information processing device" refers to a computer device used to analyze acquired data and extract and process specific information.

[0067] "Motion information" refers to information extracted as data related to musical performance and body movements.

[0068] A "control device" refers to a device that has the function of performing specific instructions or actions based on data.

[0069] A "practice plan" refers to a set of exercises and training guidelines developed to improve the user's playing technique.

[0070] "User display device" refers to a terminal or device used to visually present information.

[0071] "Communication equipment" refers to hardware or software used to send and receive data and instructions over a network.

[0072] A "reference pattern" refers to information set as model data that represents ideal behavior or performance.

[0073] This invention is a learning system aimed at improving the user's musical performance skills. The user can acquire video footage of their performance using an image acquisition device connected to an information processing device. The acquired video footage is transmitted from the information processing device to a server via a communication device.

[0074] The server is equipped with an information processing device capable of analyzing the received video data, for example, by using a skeletal analysis algorithm to extract user movement information. This makes it possible to accurately recognize movements during performance and compare them with reference patterns stored in a database.

[0075] The server uses a control device to evaluate the user's performance based on these comparison results. The practice plan generated from the evaluation results is optimized for the user and presented to the user display device via an information processing device. This allows the user to progress through their learning while visually checking their progress through an interface that includes game elements.

[0076] Furthermore, the server enables online instruction and feedback using communication equipment. This functionality allows users to receive real-time instruction even from remote locations.

[0077] Furthermore, by utilizing motion recognition technology, it is possible to provide support tailored to users with specific needs. The goal is to create a learning environment suitable for people with hearing impairments.

[0078] For example, when a user learns to play the guitar, data captured by an image acquisition device is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server proposes a practice plan for finger movements and strokes, which is displayed on the user's display device. This allows the user to improve their playing skills while having fun.

[0079] A concrete example of a prompt message to be input into the generating AI model is, "Analyze the user's playing form and suggest the optimal practice program." In this way, the present invention provides users with a personalized and efficient music learning experience and supports the improvement of their playing skills.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The user uses an image acquisition device connected to an information processing device to acquire video data of their performance. At this stage, the input is the user's performance, and the output is image data. The user plays the instrument, and their actions are recorded as video.

[0083] Step 2:

[0084] The terminal sends this acquired video data to the server. In this step, the input is image data, and the output is data transfer to the server. The terminal uploads the video data over the network.

[0085] Step 3:

[0086] The server receives the transmitted video data and extracts motion information using an analysis device. The input here is video data, and the output is motion information data. Specifically, the server applies an analysis model to calculate the performer's skeletal structure and movement patterns.

[0087] Step 4:

[0088] The server compares the extracted motion information with a reference pattern to evaluate the performance form. The input for this step is motion information data and a reference pattern, and the output is the evaluation result. The server uses an evaluation algorithm to identify areas that can be improved.

[0089] Step 5:

[0090] The server generates a user-specific practice plan based on the evaluation results. The input for this step is the evaluation results, and the output is the practice plan. The server selects appropriate practice items from the database and constructs a menu tailored to the user.

[0091] Step 6:

[0092] The terminal presents the generated practice plan to the user via a user display device. Here, the input is the practice plan, and the output is a visual learning interface. Based on the information displayed on the screen, the user checks their learning progress.

[0093] Step 7:

[0094] Users receive online instruction and feedback as needed. Input consists of the user's performance status and advice from the instructor, while output is feedback to the user. The server communicates in real time, delivering information to the user.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] To efficiently improve the skills of operators in the workplace, it is necessary to provide individual motion analysis and appropriate feedback promptly. However, with current technology, it is difficult to accurately grasp the differences in operation among operators and point out areas for improvement in real time. Furthermore, there is a lack of support for operators with specific needs, such as those with hearing impairments.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes: a portable terminal means equipped with a camera for acquiring the operator's work movements; an information processing means equipped with an analysis device for analyzing the acquired video data of the work movements and extracting the operator's morphological information; and an information processing means equipped with a control device for evaluating the operator's work posture based on the extracted morphological information and generating an educational menu according to the evaluation results. This makes it possible to provide individual motion analysis of the operator and optimal feedback in real time.

[0100] An "operator" refers to a person who operates or manages machinery or equipment at a work site.

[0101] "Operational actions" refer to the specific actions and behaviors that an operator performs when operating machinery or equipment.

[0102] A "recording device" is a device installed to record images or videos, and mainly refers to a camera.

[0103] A "mobile device" refers to a device that is easy to carry and has communication and computer functions, such as a smartphone or tablet.

[0104] "Video data" refers to a file that records visual information acquired by a camera in digital format.

[0105] An "analysis device" refers to a device that has computing resources to process acquired data and extract specific information.

[0106] An "information processing device" refers to a computer system that receives and processes data and generates or transforms necessary information.

[0107] "Morphological information" refers to information about the operator's body position and posture.

[0108] "Evaluation" refers to the process of judging the operator's actions based on extracted data and in accordance with specific criteria.

[0109] "Educational menu" refers to training programs and instructional content designed to improve the skills of operators.

[0110] A "user interface" refers to the screens and operating methods that provide interaction for exchanging information between a system and an operator.

[0111] "Playful elements" refer to game-like features incorporated to make work or learning more enjoyable.

[0112] A "communication device" refers to a device equipped with hardware and software that enables the transmission and reception of data.

[0113] "Standard posture" refers to the ideal or standard posture in an action.

[0114] The system for implementing this invention includes a series of processes for efficiently analyzing the operator's work movements and promoting skill improvement. First, the operator records their work movements as video data using a camera mounted on a mobile terminal. The terminal transmits this video data to an information processing device. The information processing device processes the received video data with an analysis device and extracts the operator's morphological information. Based on this morphological information, it evaluates the operator's work posture and generates a specific training menu using a generated AI model.

[0115] The server presents the educational menu generated above to the operator via a user interface on a mobile device. This educational menu incorporates playful elements, providing an environment where skills can be improved in an enjoyable way. For example, using welding work in an automobile manufacturing plant as an example, the operator films basic movements, and the information processing device analyzes the differences from ideal welding movements. Practice menus to approach ideal movements are provided in a game format, which can increase the operator's motivation to learn.

[0116] Furthermore, the server provides real-time feedback using communication devices and supports remote instruction as needed. In this process, appropriate feedback can be provided using sign language recognition technology for the hearing impaired, creating an environment where all operators can acquire skills equally. An example of a prompt from the generated AI model is, "Based on welding motion analysis data, please suggest a step-by-step practice menu for beginner workers."

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] The terminal records the operator's actions as video data using a camera. The input is real-time work actions, and the output is a video data file. Specifically, when the operator starts an action, the terminal's camera tracks it and records continuous video.

[0120] Step 2:

[0121] The terminal transmits the acquired video data to the information processing device. The input is the video data generated in step 1, and the output is the data stream sent to the server. Specifically, the terminal performs the process of uploading the data to the specified address on the server via the internet connection.

[0122] Step 3:

[0123] The server processes the received video data with an analysis device and extracts morphological information of the operator. The input is the video data transmitted in step 2, and the output is a dataset showing the position and movement of the operator's body. Specifically, the analysis software (e.g., OpenPose) is used to execute a skeletal tracking algorithm and extract the operator's movement information in real time.

[0124] Step 4:

[0125] The server evaluates the operator's working posture based on the extracted morphological information and generates an educational menu using a generating AI model. The input is the morphological information obtained in step 3, and the output is an educational menu suitable for the operator. Specifically, the AI ​​analysis uses this data to detect shortcomings in the operator's movements and creates a training plan for improvement. An example of a prompt used in the generating AI model is, "Based on the welding motion analysis data, please suggest a step-by-step practice menu for a novice worker."

[0126] Step 5:

[0127] The server sends the generated educational menu to the terminal, which then displays it in a user interface. The input is the educational menu generated in step 4, and the output is the learning program displayed on the terminal screen. Specifically, the terminal analyzes the educational menu and generates a GUI to display it in a way that is easy for the user to understand.

[0128] Step 6:

[0129] The user improves their work based on the provided training menu and re-records the results for feedback. The input is the work action based on the training menu, and the output is new video data for feedback. Specifically, the process involves the user independently training and re-recording the improved action.

[0130] Step 7:

[0131] The server uses communication equipment to provide real-time feedback and, if necessary, supports remote instruction. The input is the video data for feedback obtained in step 6, and the output is specific improvement instructions for the operator. Specifically, the server performs real-time analysis and immediately delivers instructional messages to the operator indicating areas that need improvement.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention is a music learning system that incorporates an emotion engine to improve the user's performance experience. The user records their instrument performance using a recording device connected to a terminal. The terminal sends this recorded data to a server, which analyzes the user's performance actions in detail.

[0134] The server extracts the user's skeletal information based on the received video data and evaluates their playing form. It also compares the analyzed skeletal information with the model form to analyze performance. Based on this analysis, it identifies areas that need improvement and generates a personalized practice menu.

[0135] Furthermore, this invention utilizes an emotion engine to recognize emotions from the user's facial expressions and movements. This emotion data is used to evaluate the user's motivation and stress levels, and to adjust the generated training menus and interfaces.

[0136] The device integrates this information and presents it to the user through a user interface that incorporates game elements. It detects the user's mood during practice in real time and flexibly adjusts the practice content and teaching methods according to the situation, thereby maximizing learning effectiveness.

[0137] For example, when a user practices the piano, the system analyzes the user's playing and provides feedback on finger position and speed. At the same time, the emotion engine evaluates the user's facial expressions, and if it determines that the user is feeling tense or anxious, it immediately provides relaxation exercises and positive feedback to create a more comfortable learning environment.

[0138] Furthermore, sign language recognition technology is incorporated to support users with hearing impairments, allowing information to be communicated using sign language. This feature creates an environment where all users can learn at their own pace.

[0139] In this way, by incorporating an emotion engine, the user's performance experience becomes an intelligent learning system that responds sensitively not only to individuality but also to emotions, supporting significant improvement in musical technique and sustained motivation to learn.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] The user activates the recording device connected to the terminal and prepares to begin playing their instrument. The terminal begins recording the performance in real time.

[0143] Step 2:

[0144] The device sends recorded video data to the server. This data contains detailed information about the user's performance and is delivered to the server via the network.

[0145] Step 3:

[0146] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This information serves as basic data for evaluating the performance form.

[0147] Step 4:

[0148] The server compares the extracted skeletal information with the model form to evaluate the accuracy and areas for improvement of the performance. Based on this evaluation, it generates a practice menu tailored to the user.

[0149] Step 5:

[0150] The server uses an emotion engine to analyze the user's facial expressions from video data and evaluate their emotions. Understanding their emotional state provides information to adjust the user's practice environment.

[0151] Step 6:

[0152] The device presents the user with practice menus received from the server and emotionally-based adjustment information. The presentation utilizes an interface that includes game elements to enhance the user's motivation to learn.

[0153] Step 7:

[0154] The user practices according to the provided training menu. During this time, the device continuously monitors the user's facial expressions and sends any changes to the server.

[0155] Step 8:

[0156] The server adjusts the content of feedback and online guidance as needed based on the emotional data received in real time. It also uses an emotional engine to provide guidance that helps users relax.

[0157] Step 9:

[0158] The emotion engine activates features such as adding relaxation exercises to the training menu or sending encouraging messages when the user is feeling tense or anxious.

[0159] (Example 2)

[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0161] Providing an effective learning environment that fully considers each user's individual skills and psychological state in music education is challenging. Furthermore, providing appropriate instruction and feedback to all users, including those with hearing impairments, is also a challenge. In addition, it is necessary to provide an autonomous and individualized learning experience by recognizing emotions in real time and utilizing that information in the learning process.

[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0163] In this invention, the server includes: an information processing device equipped with an analysis device that analyzes acquired motion data and extracts user structural information; an information processing device equipped with a control device that evaluates the user's movements based on the extracted structural information and generates a training plan according to the evaluation results; and a means further equipped with a function that recognizes the user's psychological state using emotion analysis technology and makes adjustments to the training plan and operation screen. This makes it possible to provide users with an individualized learning environment while adjusting learning based on emotions in real time.

[0164] A "recording device" is a device used to photograph a user's actions and record them as data.

[0165] "Device" is a broad term referring to hardware or software used to perform a specific function.

[0166] An "information processing device" is a computer or system used to process data and generate the necessary output.

[0167] "Structural information" refers to data that represents the position and movement of the user's body, obtained through analysis.

[0168] A "control device" is a device that includes a mechanism for managing a specific process or device and controlling its operation in accordance with its purpose.

[0169] A "training plan" refers to a series of practice menus and tasks aimed at improving the user's skills.

[0170] "Entertainment elements" are elements designed to attract users' interest and promote learning while they enjoy themselves.

[0171] A "communication network" refers to the network infrastructure used for sending and receiving information.

[0172] "Emotional analysis technology" is a technology that evaluates a user's psychological state based on their facial expressions and actions.

[0173] A "standard movement" refers to a typical or ideal movement that serves as the basis for evaluation.

[0174] This invention is a system for improving music learning. The system is composed of a combination of a recording device, an information processing device, a control device, emotion analysis technology, and the like.

[0175] Users record their musical performances using a recording device connected to their terminal. This recording device is capable of high-resolution data collection, capturing the user's movements and facial expressions in detail. The recorded data is transmitted digitally to an information processing device.

[0176] The terminal, as an information processing device, analyzes data and extracts the user's structural information. This process utilizes known libraries and algorithms (e.g., computer vision technology). Structural information refers to data that systematically represents the user's body position and movement.

[0177] The server evaluates the user's playing form based on the obtained structural information. This evaluation involves comparison with a baseline performance, and a training plan is generated based on the results. By using a generative AI model, a learning program optimized for the user is provided.

[0178] Furthermore, the server utilizes emotion analysis technology to evaluate the user's psychological state in real time. Emotional data is extracted from the user's facial expressions and actions, and the training plan and user interface are adjusted accordingly.

[0179] The device presents the generated training plan to the user as an interface. The interface incorporates entertainment elements to make learning enjoyable for the user. For example, if the user is practicing the piano, the system provides feedback on finger position and speed, and suggests relaxation exercises as needed.

[0180] An example of a prompt message would be, "Send this data to the server and have it analyzed using the sentiment engine." This example helps users understand exactly how to use the system.

[0181] In this way, this invention can maximize the effectiveness of music learning by providing an individualized learning environment.

[0182] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0183] Step 1:

[0184] The user begins playing and records the performance using a recording device connected to the terminal. The input for recording is the user's movements and facial expressions while playing. This data is recorded in high resolution and includes the user's subtle movements and expressions. Digital video data is generated as output.

[0185] Step 2:

[0186] The terminal transmits the recorded video data to the information processing device. The input is the video data generated in step 1. The information processing device receives this data and converts it into a data format suitable for analysis. The output is the video data formatted into an analyzable format.

[0187] Step 3:

[0188] The server extracts the user's structural information based on the video data transmitted from the terminal. The input for this step is analyzable video data. The server uses computer vision algorithms to identify the user's body position and the movement of each part, and generates a digital skeletal model. The output is structural information data that provides a detailed representation of the user's movements.

[0189] Step 4:

[0190] The server evaluates the user's playing form based on the structural information obtained. The input is structural information data. The control unit within the server compares the form to a standard operation and determines whether the form is good or bad. Based on the evaluation results, specific areas for improvement are identified. The output is an evaluation report that includes the areas for improvement.

[0191] Step 5:

[0192] The server generates a training plan optimized for the user based on the evaluation report. The input is the evaluation report obtained in step 4. The generated AI model is used to create individual training content and tasks. The output is a training plan customized for the user.

[0193] Step 6:

[0194] The server uses emotion analysis technology to analyze the user's psychological state from video footage. The input is video data, including the user's facial expressions. The server uses AI to evaluate the user's emotional state. The output is emotion evaluation data.

[0195] Step 7:

[0196] The terminal presents the user with an operation screen based on the training plan and emotion assessment data received from the server. The input consists of the user's training plan and emotion assessment data. The operation screen reflects real-time feedback tailored to the user's emotions. The output is an interactive user interface.

[0197] Step 8:

[0198] The device optimizes the user's learning experience based on information received during performance. Input consists of real-time user movement and emotion data. This allows for immediate suggestions of the most effective practice and instruction methods based on the user's current state. Output consists of suggestions and instructions to facilitate the user's skill improvement.

[0199] (Application Example 2)

[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0201] In modern society, a problem exists in the limited opportunities and instruction available for improving musical instrument playing skills. Furthermore, providing feedback that takes into account the user's unique emotional state is difficult, resulting in challenges in maintaining user motivation. Additionally, there is a lack of efficient mechanisms for providing real-time, individualized instruction and feedback in physical stores where direct experiences such as instrument trials take place. These circumstances make it difficult to achieve sustained and efficient improvement in playing skills.

[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0203] In this invention, the server includes terminal means equipped with a camera for capturing user movements, information processing means equipped with an analysis device for analyzing acquired video data of movements and extracting user physical information, and information processing means equipped with an emotion analysis device for analyzing facial emotions and providing feedback according to the emotional state. This makes it possible to provide real-time feedback tailored to each individual user and appropriate guidance based on their emotional state.

[0204] A "camera for capturing user movements" is a device used to visually capture the physical movements of a user while they are performing an instrument or engaging in an activity.

[0205] An "information processing device" is a computer system used to analyze, store, or transmit digital data.

[0206] "Physical information" refers to data about the user's body movements and posture, specifically including information such as the position and angle of the skeleton.

[0207] A "control device" is a device that performs processing according to a specific purpose based on the data it receives.

[0208] A "display surface" refers to a visual device or screen used to present information to a user.

[0209] A "communication device" is a device used to send and receive data and instructions, and has the function of connecting to servers and terminals via a network.

[0210] An "emotion analysis device" is a device that analyzes a user's emotions from their facial expressions and actions.

[0211] "Feedback" refers to information that provides guidance and evaluation to the user based on the data and analysis results obtained.

[0212] To implement this invention, a system is constructed that integrates a series of devices and software. The server receives video data from a terminal connected to a camera and smart glasses for capturing the user's movements. The terminal's camera captures the user's performance and movements in high resolution and transmits the video to the server. At this time, image processing software such as OpenCV is used to capture the video data in real time.

[0213] The server uses an information processing device to analyze the received video data. Specifically, it uses libraries such as PoseNet and OpenPose to extract skeletal information. This extracts the user's physical information as digital data. This data is compared with a pre-configured model form to evaluate the performance form. The emotion analysis device uses emotion recognition software such as facial_emotion_recognition to analyze the user's facial expressions from the video data and understand their emotional state. Based on this information, the server generates appropriate feedback according to the user's emotional state.

[0214] User feedback is presented in real time on the display screen. Through the interface of the display device or smart glasses, users receive visual and audible feedback. This allows users to correct their technique on the spot and improve their playing skills. For example, if the server analyzes the user's finger movements during a performance and detects incorrect form, a guide to correct it will appear on the display. At the same time, if the user is feeling tense or anxious, positive feedback to encourage relaxation is also provided.

[0215] An example of a prompt used to implement this system is: "Analyze the user's performance video and evaluate areas for improvement in their playing form and emotional state. If the user is nervous, provide advice to help them relax and generate positive feedback." Based on this prompt, the generating AI model provides instruction optimized for the user.

[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0217] Step 1:

[0218] The user begins playing music using smart glasses or a camera-equipped device, recording their performance in real time. The input data is the user's performance video, which the device continuously captures. This data is saved digitally as it is necessary for subsequent processing.

[0219] Step 2:

[0220] The terminal sends the recorded performance video data to the server. The server receives this data and stores it in a database. The input is the user's performance video data, and the output is the raw video data sent to the server. The transmitted data is necessary to prepare the video for analysis.

[0221] Step 3:

[0222] The server uses PoseNet and OpenPose to analyze the user's body information from received video data and extract skeletal information. The input is video data of the user performing, and by analyzing this data, the user's skeletal information is extracted. The output is digital data related to the user's body movements. This skeletal information is used to evaluate form.

[0223] Step 4:

[0224] The server uses emotion recognition software, such as facial_emotion_recognition, to analyze the user's facial expressions from the video and obtain emotional information. The input is the user's video data, and the output is data indicating the user's emotional state. Through this process, the user's motivation and tension level are evaluated.

[0225] Step 5:

[0226] The server compares the extracted skeletal information with a model form to evaluate the performance form and generates feedback based on emotional information. The inputs are skeletal information and emotional information, and the output is feedback data provided to the user. This feedback includes areas for improvement in performance technique and advice tailored to the emotional state.

[0227] Step 6:

[0228] The generated feedback is displayed on the device's screen in real time. Users receive this feedback through smart glasses or a display. The input is feedback data, and the output is instruction and advice conveyed to the user. This allows users to improve their performance and maintain their motivation.

[0229] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0230] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0231] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0232] [Second Embodiment]

[0233] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0234] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0235] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0236] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0237] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0238] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0239] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0240] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0241] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0242] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0243] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0244] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0245] This invention is a music learning system designed to improve the user's performance skills. The user records their performance via a recording device connected to a terminal. The terminal has the function of transmitting this video data to a server.

[0246] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This skeletal information allows for an accurate evaluation of the user's playing form. The analysis results are compared on the server with exemplary form and used to identify areas for improvement and strengths of the user.

[0247] Based on the evaluation of the user's playing form, the server generates an individually optimized practice menu. This menu includes a variety of exercises aimed at improving playing technique, and is composed of appropriate exercises selected from a database managed by the server.

[0248] The generated practice menu is presented to the user via the device. The device uses a user interface incorporating game elements to visualize learning progress and enhance the user's motivation to learn. Furthermore, users can receive online instruction at any time, and the server supports real-time communication for this purpose.

[0249] Furthermore, as a consideration for people with hearing impairments, the terminal utilizes sign language recognition technology to provide an environment where users can confirm necessary information in sign language. The server generates appropriate feedback corresponding to this and presents it through the terminal.

[0250] For example, when a user learns to play the guitar, the performance data captured by the camera is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server then suggests specific practice menus regarding finger movements and strokes. These are displayed on the device, allowing the user to improve their playing skills in an enjoyable way while receiving instruction.

[0251] In this way, the present invention provides users with a personalized and effective music learning experience, supporting the improvement of their performance skills. This system creates a more efficient and engaging learning environment than conventional methods.

[0252] The following describes the processing flow.

[0253] Step 1:

[0254] The user activates the device's recording function and prepares to record their musical performance. Once ready, the device starts recording, capturing the performance in real time.

[0255] Step 2:

[0256] The device sends recorded video data to the server. This transmission is primarily done over the network, enabling real-time data reception.

[0257] Step 3:

[0258] The server receives video data transmitted from the terminal and performs skeletal detection using an analysis device. This extracts skeletal information such as the user's hands and arms.

[0259] Step 4:

[0260] The server compares the user's playing form to a model form based on the extracted skeletal information. It then identifies areas that need improvement based on the form's evaluation criteria.

[0261] Step 5:

[0262] The server automatically generates individually optimized practice menus based on the form evaluation results. These menus will include specific technical exercises and formal drills.

[0263] Step 6:

[0264] The device presents the generated practice menu to the user. A gamified user interface is used to enhance learning motivation during the presentation.

[0265] Step 7:

[0266] The user practices based on the provided menu, records their performance, and sends it from their device to the server. They then receive feedback from the server.

[0267] Step 8:

[0268] The server evaluates the video data sent back by the user and provides online direct instruction and feedback as needed. For users with hearing impairments, it also uses sign language recognition technology to provide appropriate feedback.

[0269] (Example 1)

[0270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0271] [To improve musical performance skills, there is a need for a system that provides efficient practice menus tailored to individual user characteristics, and enables evaluation and improvement of performance. Furthermore, an environment is needed where multiple learners, each in different circumstances, can receive appropriate feedback online.]

[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0273] In this invention, the server includes means for [equipped with an information processing device for extracting target operation information], means for [equipped with a control device for generating a practice plan], and means for [equipped with a communication device for providing instruction and feedback over a communication network]. This enables [the provision of personalized practice menus to users and efficient performance evaluation and instruction via online means].

[0274] An "image acquisition device" refers to a device used to acquire video footage of a subject's performance.

[0275] An "information processing device" refers to a computer device used to analyze acquired data and extract and process specific information.

[0276] "Motion information" refers to information extracted as data related to musical performance and body movements.

[0277] The "control device" refers to a device having a function of performing specific instructions or operations based on data.

[0278] The "practice plan" refers to a series of exercise and training guidelines developed to improve the user's performance technique.

[0279] The "user display device" refers to a terminal or device for visually presenting information.

[0280] The "communication device" refers to hardware or software for transmitting and receiving data and instructions via a network.

[0281] The "reference pattern" refers to information set as model data indicating ideal operations or performances.

[0282] This invention is a learning system aimed at improving the user's music performance technique. The user can use an image acquisition device connected to an information processing device to acquire their own performance as a video. The acquired video is transmitted to a server via a communication device by the information processing device.

[0283] The server is equipped with an information processing device capable of analyzing the received video data, and extracts the user's motion information using, for example, a skeleton analysis algorithm. Thereby, it is possible to accurately recognize the motion during performance and compare it with the reference pattern stored in the database.

[0284] The server uses a control device to evaluate the user's performance based on this comparison result. The practice plan generated from the evaluation result is constructed in a form optimized for the user and presented to the user display device via the information processing device. Thereby, the user can proceed with learning while visually confirming their progress through an interface including game elements.

[0285] In addition, the server enables online guidance and feedback using a communication device. With this function, users can receive real-time guidance even from a remote location.

[0286] Furthermore, by utilizing motion recognition technology, support for users with specific needs can also be provided. The aim is to construct a learning environment suitable for people with hearing impairments.

[0287] As a specific example, when a user is learning to play the guitar, the data captured by the image acquisition device is analyzed by the server to evaluate the position and form of the user's hand. Based on the analysis results, the server proposes a practice plan for finger movement and stroke, and displays it on the user display device. In this way, the user can improve their playing skills while enjoying the process.

[0288] As a specific example of the prompt sentence input into the generative AI model, there is "Analyze the user's playing form and present an optimal practice program." Thus, the present invention provides an efficient music learning experience customized for users and supports the improvement of playing techniques.

[0289] The flow of the specific process in Example 1 will be described using FIG. 11.

[0290] Step 1:

[0291] The user uses an image acquisition device connected to the information processing device to obtain video data of the performance. The input at this stage is the user's performance act, and the output is image data. The user plays the instrument, and the movement is recorded as video.

[0292] Step 2:

[0293] The terminal transmits the acquired video data to the server. The input at this step is image data, and the output is data transfer to the server. The terminal uploads the video data through the network.

[0294] Step 3:

[0295] The server receives the transmitted video data and extracts motion information using an analysis device. The input here is video data, and the output is motion information data. Specifically, the server applies an analysis model to calculate the performer's skeletal structure and movement patterns.

[0296] Step 4:

[0297] The server compares the extracted motion information with a reference pattern to evaluate the performance form. The input for this step is motion information data and a reference pattern, and the output is the evaluation result. The server uses an evaluation algorithm to identify areas that can be improved.

[0298] Step 5:

[0299] The server generates a user-specific practice plan based on the evaluation results. The input for this step is the evaluation results, and the output is the practice plan. The server selects appropriate practice items from the database and constructs a menu tailored to the user.

[0300] Step 6:

[0301] The terminal presents the generated practice plan to the user via a user display device. Here, the input is the practice plan, and the output is a visual learning interface. Based on the information displayed on the screen, the user checks their learning progress.

[0302] Step 7:

[0303] Users receive online instruction and feedback as needed. Input consists of the user's performance status and advice from the instructor, while output is feedback to the user. The server communicates in real time, delivering information to the user.

[0304] (Application Example 1)

[0305] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0306] In order to efficiently improve the skills of operators at the work site, it is necessary to provide individual motion analysis and appropriate feedback promptly. However, with the current technology, it is difficult to accurately grasp the differences in the motions of each operator and point out areas for improvement in real time. In addition, there is also a lack of support for operators with specific needs such as those with hearing impairments.

[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0308] [[ID=I3]] In this invention, the server includes: portable terminal means equipped with a photographing device for acquiring the work motions of an operator; information processing device means equipped with an analysis device for analyzing the video data of the acquired work motions and extracting the morphological information of the operator; and information processing device means equipped with a control device for evaluating the work posture of the operator based on the extracted morphological information and generating an educational menu according to the evaluation result. Thereby, it becomes possible to provide individual motion analysis and optimal feedback for the operator in real time.

[0309] The "operator" refers to a person who operates or manages machines and equipment at the work site.

[0310] The "work motion" refers to the specific motions and behaviors performed by the operator when operating machines and equipment.

[0311] The "photographing device" is a device installed for recording images and videos, mainly referring to a camera.

[0312] The "portable terminal" refers to a device that is easy to carry and has communication and computer functions, such as a smartphone or a tablet.

[0313] "Video data" refers to a file that records visual information acquired by a camera in digital format.

[0314] An "analysis device" refers to a device that has computing resources to process acquired data and extract specific information.

[0315] An "information processing device" refers to a computer system that receives and processes data and generates or transforms necessary information.

[0316] "Morphological information" refers to information about the operator's body position and posture.

[0317] "Evaluation" refers to the process of judging the operator's actions based on extracted data and in accordance with specific criteria.

[0318] "Educational menu" refers to training programs and instructional content designed to improve the skills of operators.

[0319] A "user interface" refers to the screens and operating methods that provide interaction for exchanging information between a system and an operator.

[0320] "Playful elements" refer to game-like features incorporated to make work or learning more enjoyable.

[0321] A "communication device" refers to a device equipped with hardware and software that enables the transmission and reception of data.

[0322] "Standard posture" refers to the ideal or standard posture in an action.

[0323] The system for implementing this invention includes a series of processes for efficiently analyzing the operator's work movements and promoting skill improvement. First, the operator records their work movements as video data using a camera mounted on a mobile terminal. The terminal transmits this video data to an information processing device. The information processing device processes the received video data with an analysis device and extracts the operator's morphological information. Based on this morphological information, it evaluates the operator's work posture and generates a specific training menu using a generated AI model.

[0324] The server presents the educational menu generated above to the operator via a user interface on a mobile device. This educational menu incorporates playful elements, providing an environment where skills can be improved in an enjoyable way. For example, using welding work in an automobile manufacturing plant as an example, the operator films basic movements, and the information processing device analyzes the differences from ideal welding movements. Practice menus to approach ideal movements are provided in a game format, which can increase the operator's motivation to learn.

[0325] Furthermore, the server provides real-time feedback using communication devices and supports remote instruction as needed. In this process, appropriate feedback can be provided using sign language recognition technology for the hearing impaired, creating an environment where all operators can acquire skills equally. An example of a prompt from the generated AI model is, "Based on welding motion analysis data, please suggest a step-by-step practice menu for beginner workers."

[0326] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0327] Step 1:

[0328] The terminal records the operator's actions as video data using a camera. The input is real-time work actions, and the output is a video data file. Specifically, when the operator starts an action, the terminal's camera tracks it and records continuous video.

[0329] Step 2:

[0330] The terminal transmits the acquired video data to the information processing device. The input is the video data generated in step 1, and the output is the data stream sent to the server. Specifically, the terminal performs the process of uploading the data to the specified address on the server via the internet connection.

[0331] Step 3:

[0332] The server processes the received video data with an analysis device and extracts morphological information of the operator. The input is the video data transmitted in step 2, and the output is a dataset showing the position and movement of the operator's body. Specifically, the analysis software (e.g., OpenPose) is used to execute a skeletal tracking algorithm and extract the operator's movement information in real time.

[0333] Step 4:

[0334] The server evaluates the operator's working posture based on the extracted morphological information and generates an educational menu using a generating AI model. The input is the morphological information obtained in step 3, and the output is an educational menu suitable for the operator. Specifically, the AI ​​analysis uses this data to detect shortcomings in the operator's movements and creates a training plan for improvement. An example of a prompt used in the generating AI model is, "Based on the welding motion analysis data, please suggest a step-by-step practice menu for a novice worker."

[0335] Step 5:

[0336] The server sends the generated educational menu to the terminal, which then displays it in a user interface. The input is the educational menu generated in step 4, and the output is the learning program displayed on the terminal screen. Specifically, the terminal analyzes the educational menu and generates a GUI to display it in a way that is easy for the user to understand.

[0337] Step 6:

[0338] The user improves their work based on the provided training menu and re-records the results for feedback. The input is the work action based on the training menu, and the output is new video data for feedback. Specifically, the process involves the user independently training and re-recording the improved action.

[0339] Step 7:

[0340] The server uses communication equipment to provide real-time feedback and, if necessary, supports remote instruction. The input is the video data for feedback obtained in step 6, and the output is specific improvement instructions for the operator. Specifically, the server performs real-time analysis and immediately delivers instructional messages to the operator indicating areas that need improvement.

[0341] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0342] This invention is a music learning system that incorporates an emotion engine to improve the user's performance experience. The user records their instrument performance using a recording device connected to a terminal. The terminal sends this recorded data to a server, which analyzes the user's performance actions in detail.

[0343] The server extracts the user's skeletal information based on the received video data and evaluates their playing form. It also compares the analyzed skeletal information with the model form to analyze performance. Based on this analysis, it identifies areas that need improvement and generates a personalized practice menu.

[0344] Furthermore, this invention utilizes an emotion engine to recognize emotions from the user's facial expressions and movements. This emotion data is used to evaluate the user's motivation and stress levels, and to adjust the generated training menus and interfaces.

[0345] The device integrates this information and presents it to the user through a user interface that incorporates game elements. It detects the user's mood during practice in real time and flexibly adjusts the practice content and teaching methods according to the situation, thereby maximizing learning effectiveness.

[0346] For example, when a user practices the piano, the system analyzes the user's playing and provides feedback on finger position and speed. At the same time, the emotion engine evaluates the user's facial expressions, and if it determines that the user is feeling tense or anxious, it immediately provides relaxation exercises and positive feedback to create a more comfortable learning environment.

[0347] Furthermore, sign language recognition technology is incorporated to support users with hearing impairments, allowing information to be communicated using sign language. This feature creates an environment where all users can learn at their own pace.

[0348] In this way, by incorporating an emotion engine, the user's performance experience becomes an intelligent learning system that responds sensitively not only to individuality but also to emotions, supporting significant improvement in musical technique and sustained motivation to learn.

[0349] The following describes the processing flow.

[0350] Step 1:

[0351] The user activates the recording device connected to the terminal and prepares to begin playing their instrument. The terminal begins recording the performance in real time.

[0352] Step 2:

[0353] The device sends recorded video data to the server. This data contains detailed information about the user's performance and is delivered to the server via the network.

[0354] Step 3:

[0355] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This information serves as the basic data for evaluating the performance form.

[0356] Step 4:

[0357] The server compares the extracted skeletal information with the model form and evaluates the accuracy and areas for improvement of the performance. Based on this evaluation, it generates a practice menu tailored to the user.

[0358] Step 5:

[0359] The server uses an emotion engine to analyze the user's facial expressions from video data and evaluate their emotions. Understanding their emotional state provides information to adjust the user's practice environment.

[0360] Step 6:

[0361] The device presents the user with practice menus received from the server and emotionally-based adjustment information. The presentation utilizes an interface that includes game elements to enhance the user's motivation to learn.

[0362] Step 7:

[0363] The user practices according to the provided training menu. During this time, the device continuously monitors the user's facial expressions and sends any changes to the server.

[0364] Step 8:

[0365] The server adjusts the content of feedback and online guidance as needed based on the emotional data received in real time. It also uses an emotional engine to provide guidance that helps users relax.

[0366] Step 9:

[0367] The emotion engine activates features such as adding relaxation exercises to the training menu or sending encouraging messages when the user is feeling tense or anxious.

[0368] (Example 2)

[0369] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0370] Providing an effective learning environment that fully considers each user's individual skills and psychological state in music education is challenging. Furthermore, providing appropriate instruction and feedback to all users, including those with hearing impairments, is also a challenge. In addition, it is necessary to provide an autonomous and individualized learning experience by recognizing emotions in real time and utilizing that information in the learning process.

[0371] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0372] In this invention, the server includes: an information processing device equipped with an analysis device that analyzes acquired motion data and extracts user structural information; an information processing device equipped with a control device that evaluates the user's movements based on the extracted structural information and generates a training plan according to the evaluation results; and a means further equipped with a function that recognizes the user's psychological state using emotion analysis technology and makes adjustments to the training plan and operation screen. This makes it possible to provide users with an individualized learning environment while adjusting learning based on emotions in real time.

[0373] A "recording device" is a device used to photograph a user's actions and record them as data.

[0374] "Device" is a broad term referring to hardware or software used to perform a specific function.

[0375] An "information processing device" is a computer or system used to process data and generate the necessary output.

[0376] "Structural information" refers to data that represents the position and movement of the user's body, obtained through analysis.

[0377] A "control device" is a device that includes a mechanism for managing a specific process or device and controlling its operation in accordance with its purpose.

[0378] A "training plan" refers to a series of practice menus and tasks aimed at improving the user's skills.

[0379] "Entertainment elements" are elements designed to attract users' interest and promote learning while they enjoy themselves.

[0380] A "communication network" refers to the network infrastructure used for sending and receiving information.

[0381] "Emotional analysis technology" is a technology that evaluates a user's psychological state based on their facial expressions and actions.

[0382] A "standard movement" refers to a typical or ideal movement that serves as the basis for evaluation.

[0383] This invention is a system for improving music learning. The system is composed of a combination of a recording device, an information processing device, a control device, emotion analysis technology, and the like.

[0384] Users record their musical performances using a recording device connected to their terminal. This recording device is capable of high-resolution data collection, capturing the user's movements and facial expressions in detail. The recorded data is transmitted digitally to an information processing device.

[0385] The terminal, as an information processing device, analyzes data and extracts the user's structural information. This process utilizes known libraries and algorithms (e.g., computer vision technology). Structural information refers to data that systematically represents the user's body position and movement.

[0386] The server evaluates the user's playing form based on the obtained structural information. This evaluation involves comparison with a baseline performance, and a training plan is generated based on the results. By using a generative AI model, a learning program optimized for the user is provided.

[0387] Furthermore, the server utilizes emotion analysis technology to evaluate the user's psychological state in real time. Emotional data is extracted from the user's facial expressions and actions, and the training plan and user interface are adjusted accordingly.

[0388] The device presents the generated training plan to the user as an interface. The interface incorporates entertainment elements to make learning enjoyable for the user. For example, if the user is practicing the piano, the system provides feedback on finger position and speed, and suggests relaxation exercises as needed.

[0389] An example of a prompt message would be, "Send this data to the server and have it analyzed using the sentiment engine." This example helps users understand exactly how to use the system.

[0390] In this way, this invention can maximize the effectiveness of music learning by providing an individualized learning environment.

[0391] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0392] Step 1:

[0393] The user begins playing and records the performance using a recording device connected to the terminal. The input for recording is the user's movements and facial expressions while playing. This data is recorded in high resolution and includes the user's subtle movements and expressions. Digital video data is generated as output.

[0394] Step 2:

[0395] The terminal transmits the recorded video data to the information processing device. The input is the video data generated in step 1. The information processing device receives this data and converts it into a data format suitable for analysis. The output is the video data formatted into an analyzable format.

[0396] Step 3:

[0397] The server extracts the user's structural information based on the video data transmitted from the terminal. The input for this step is analyzable video data. The server uses computer vision algorithms to identify the user's body position and the movement of each part, and generates a digital skeletal model. The output is structural information data that provides a detailed representation of the user's movements.

[0398] Step 4:

[0399] The server evaluates the user's playing form based on the structural information obtained. The input is structural information data. The control unit within the server compares the form to a standard operation and determines whether the form is good or bad. Based on the evaluation results, specific areas for improvement are identified. The output is an evaluation report that includes the areas for improvement.

[0400] Step 5:

[0401] The server generates a training plan optimized for the user based on the evaluation report. The input is the evaluation report obtained in step 4. The generated AI model is used to create individual training content and tasks. The output is a training plan customized for the user.

[0402] Step 6:

[0403] The server uses emotion analysis technology to analyze the user's psychological state from video footage. The input is video data, including the user's facial expressions. The server uses AI to evaluate the user's emotional state. The output is emotion evaluation data.

[0404] Step 7:

[0405] The terminal presents the user with an operation screen based on the training plan and emotion assessment data received from the server. The input consists of the user's training plan and emotion assessment data. The operation screen reflects real-time feedback tailored to the user's emotions. The output is an interactive user interface.

[0406] Step 8:

[0407] The device optimizes the user's learning experience based on information received during performance. Input consists of real-time user movement and emotion data. This allows for immediate suggestions of the most effective practice and instruction methods based on the user's current state. Output consists of suggestions and instructions to facilitate the user's skill improvement.

[0408] (Application Example 2)

[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0410] In modern society, a problem exists in the limited opportunities and instruction available for improving musical instrument playing skills. Furthermore, providing feedback that takes into account the user's unique emotional state is difficult, resulting in challenges in maintaining user motivation. Additionally, there is a lack of efficient mechanisms for providing real-time, individualized instruction and feedback in physical stores where direct experiences such as instrument trials take place. These circumstances make it difficult to achieve sustained and efficient improvement in playing skills.

[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0412] In this invention, the server includes terminal means equipped with a camera for capturing user movements, information processing means equipped with an analysis device for analyzing acquired video data of movements and extracting user physical information, and information processing means equipped with an emotion analysis device for analyzing facial emotions and providing feedback according to the emotional state. This makes it possible to provide real-time feedback tailored to each individual user and appropriate guidance based on their emotional state.

[0413] A "camera for capturing user movements" is a device used to visually capture the physical movements of a user while they are performing an instrument or engaging in an activity.

[0414] An "information processing device" is a computer system used to analyze, store, or transmit digital data.

[0415] "Physical information" refers to data about the user's body movements and posture, specifically including information such as the position and angle of the skeleton.

[0416] A "control device" is a device that performs processing according to a specific purpose based on the data it receives.

[0417] A "display surface" refers to a visual device or screen used to present information to a user.

[0418] A "communication device" is a device used to send and receive data and instructions, and has the function of connecting to servers and terminals via a network.

[0419] An "emotion analysis device" is a device that analyzes a user's emotions from their facial expressions and actions.

[0420] "Feedback" refers to information that provides guidance and evaluation to the user based on the data and analysis results obtained.

[0421] To implement this invention, a system is constructed that integrates a series of devices and software. The server receives video data from a terminal connected to a camera and smart glasses for capturing the user's movements. The terminal's camera captures the user's performance and movements in high resolution and transmits the video to the server. At this time, image processing software such as OpenCV is used to capture the video data in real time.

[0422] The server uses an information processing device to analyze the received video data. Specifically, it uses libraries such as PoseNet and OpenPose to extract skeletal information. This extracts the user's physical information as digital data. This data is compared with a pre-configured model form to evaluate the performance form. The emotion analysis device uses emotion recognition software such as facial_emotion_recognition to analyze the user's facial expressions from the video data and understand their emotional state. Based on this information, the server generates appropriate feedback according to the user's emotional state.

[0423] User feedback is presented in real time on the display screen. Through the interface of the display device or smart glasses, users receive visual and audible feedback. This allows users to correct their technique on the spot and improve their playing skills. For example, if the server analyzes the user's finger movements during a performance and detects incorrect form, a guide to correct it will appear on the display. At the same time, if the user is feeling tense or anxious, positive feedback to encourage relaxation is also provided.

[0424] An example of a prompt used to implement this system is: "Analyze the user's performance video and evaluate areas for improvement in their playing form and emotional state. If the user is nervous, provide advice to help them relax and generate positive feedback." Based on this prompt, the generating AI model provides instruction optimized for the user.

[0425] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0426] Step 1:

[0427] The user begins playing music using smart glasses or a camera-equipped device, recording their performance in real time. The input data is the user's performance video, which the device continuously captures. This data is saved digitally as it is necessary for subsequent processing.

[0428] Step 2:

[0429] The terminal sends the recorded performance video data to the server. The server receives this data and stores it in a database. The input is the user's performance video data, and the output is the raw video data sent to the server. The transmitted data is necessary to prepare the video for analysis.

[0430] Step 3:

[0431] The server uses PoseNet and OpenPose to analyze the user's body information from received video data and extract skeletal information. The input is video data of the user performing, and by analyzing this data, the user's skeletal information is extracted. The output is digital data related to the user's body movements. This skeletal information is used to evaluate form.

[0432] Step 4:

[0433] The server uses emotion recognition software, such as facial_emotion_recognition, to analyze the user's facial expressions from the video and obtain emotional information. The input is the user's video data, and the output is data indicating the user's emotional state. Through this process, the user's motivation and tension level are evaluated.

[0434] Step 5:

[0435] The server compares the extracted skeletal information with a model form to evaluate the performance form and generates feedback based on emotional information. The inputs are skeletal information and emotional information, and the output is feedback data provided to the user. This feedback includes areas for improvement in performance technique and advice tailored to the emotional state.

[0436] Step 6:

[0437] The generated feedback is displayed on the device's screen in real time. Users receive this feedback through smart glasses or a display. The input is feedback data, and the output is instruction and advice conveyed to the user. This allows users to improve their performance and maintain their motivation.

[0438] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0439] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0440] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0441] [Third Embodiment]

[0442] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0443] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0444] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0445] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0446] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0447] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0448] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0449] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0450] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0451] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0452] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0453] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0454] This invention is a music learning system designed to improve the user's performance skills. The user records their performance via a recording device connected to a terminal. The terminal has the function of transmitting this video data to a server.

[0455] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This skeletal information allows for an accurate evaluation of the user's playing form. The analysis results are compared on the server with exemplary form and used to identify areas for improvement and strengths of the user.

[0456] Based on the evaluation of the user's playing form, the server generates an individually optimized practice menu. This menu includes a variety of exercises aimed at improving playing technique, and is composed of appropriate exercises selected from a database managed by the server.

[0457] The generated practice menu is presented to the user via the device. The device uses a user interface incorporating game elements to visualize learning progress and enhance the user's motivation to learn. Furthermore, users can receive online instruction at any time, and the server supports real-time communication for this purpose.

[0458] Furthermore, as a consideration for people with hearing impairments, the terminal utilizes sign language recognition technology to provide an environment where users can confirm necessary information in sign language. The server generates appropriate feedback corresponding to this and presents it through the terminal.

[0459] For example, when a user learns to play the guitar, the performance data captured by the camera is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server then suggests specific practice menus regarding finger movements and strokes. These are displayed on the device, allowing the user to improve their playing skills in an enjoyable way while receiving instruction.

[0460] In this way, the present invention provides users with a personalized and effective music learning experience, supporting the improvement of their performance skills. This system creates a more efficient and engaging learning environment than conventional methods.

[0461] The following describes the processing flow.

[0462] Step 1:

[0463] The user activates the device's recording function and prepares to record their musical performance. Once ready, the device starts recording, capturing the performance in real time.

[0464] Step 2:

[0465] The device sends recorded video data to the server. This transmission is primarily done over the network, enabling real-time data reception.

[0466] Step 3:

[0467] The server receives video data transmitted from the terminal and performs skeletal detection using an analysis device. This extracts skeletal information such as the user's hands and arms.

[0468] Step 4:

[0469] The server compares the user's playing form to a model form based on the extracted skeletal information. It then identifies areas that need improvement based on the form's evaluation criteria.

[0470] Step 5:

[0471] The server automatically generates individually optimized practice menus based on the form evaluation results. These menus will include specific technical exercises and formal drills.

[0472] Step 6:

[0473] The device presents the generated practice menu to the user. A gamified user interface is used to enhance learning motivation during the presentation.

[0474] Step 7:

[0475] The user practices based on the provided menu, records their performance, and sends it from their device to the server. They then receive feedback from the server.

[0476] Step 8:

[0477] The server evaluates the video data sent back by the user and provides online direct instruction and feedback as needed. For users with hearing impairments, it also uses sign language recognition technology to provide appropriate feedback.

[0478] (Example 1)

[0479] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0480] [To improve musical performance skills, there is a need for a system that provides efficient practice menus tailored to individual user characteristics, and enables evaluation and improvement of performance. Furthermore, an environment is needed where multiple learners, each in different circumstances, can receive appropriate feedback online.]

[0481] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0482] In this invention, the server includes means for [equipped with an information processing device for extracting target operation information], means for [equipped with a control device for generating a practice plan], and means for [equipped with a communication device for providing instruction and feedback over a communication network]. This enables [the provision of personalized practice menus to users and efficient performance evaluation and instruction via online means].

[0483] An "image acquisition device" refers to a device used to acquire video footage of a subject's performance.

[0484] An "information processing device" refers to a computer device used to analyze acquired data and extract and process specific information.

[0485] "Motion information" refers to information extracted as data related to musical performance and body movements.

[0486] A "control device" refers to a device that has the function of performing specific instructions or actions based on data.

[0487] A "practice plan" refers to a set of exercises and training guidelines developed to improve the user's playing technique.

[0488] "User display device" refers to a terminal or device used to visually present information.

[0489] "Communication equipment" refers to hardware or software used to send and receive data and instructions over a network.

[0490] A "reference pattern" refers to information set as model data that represents ideal behavior or performance.

[0491] This invention is a learning system aimed at improving the user's musical performance skills. The user can acquire video footage of their performance using an image acquisition device connected to an information processing device. The acquired video footage is transmitted from the information processing device to a server via a communication device.

[0492] The server is equipped with an information processing device capable of analyzing the received video data, for example, by using a skeletal analysis algorithm to extract user movement information. This makes it possible to accurately recognize movements during performance and compare them with reference patterns stored in a database.

[0493] The server uses a control device to evaluate the user's performance based on these comparison results. The practice plan generated from the evaluation results is optimized for the user and presented to the user display device via an information processing device. This allows the user to progress through their learning while visually checking their progress through an interface that includes game elements.

[0494] Furthermore, the server enables online instruction and feedback using communication equipment. This functionality allows users to receive real-time instruction even from remote locations.

[0495] Furthermore, by utilizing motion recognition technology, it is possible to provide support tailored to users with specific needs. The goal is to create a learning environment suitable for people with hearing impairments.

[0496] For example, when a user learns to play the guitar, data captured by an image acquisition device is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server proposes a practice plan for finger movements and strokes, which is displayed on the user's display device. This allows the user to improve their playing skills while having fun.

[0497] A concrete example of a prompt message to be input into the generating AI model is, "Analyze the user's playing form and suggest the optimal practice program." In this way, the present invention provides users with a personalized and efficient music learning experience and supports the improvement of their playing skills.

[0498] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0499] Step 1:

[0500] The user uses an image acquisition device connected to an information processing device to acquire video data of their performance. At this stage, the input is the user's performance, and the output is image data. The user plays the instrument, and their actions are recorded as video.

[0501] Step 2:

[0502] The terminal sends this acquired video data to the server. In this step, the input is image data, and the output is data transfer to the server. The terminal uploads the video data over the network.

[0503] Step 3:

[0504] The server receives the transmitted video data and extracts motion information using an analysis device. The input here is video data, and the output is motion information data. Specifically, the server applies an analysis model to calculate the performer's skeletal structure and movement patterns.

[0505] Step 4:

[0506] The server compares the extracted motion information with a reference pattern to evaluate the performance form. The input for this step is motion information data and a reference pattern, and the output is the evaluation result. The server uses an evaluation algorithm to identify areas that can be improved.

[0507] Step 5:

[0508] The server generates a user-specific practice plan based on the evaluation results. The input for this step is the evaluation results, and the output is the practice plan. The server selects appropriate practice items from the database and constructs a menu tailored to the user.

[0509] Step 6:

[0510] The terminal presents the generated practice plan to the user via a user display device. Here, the input is the practice plan, and the output is a visual learning interface. Based on the information displayed on the screen, the user checks their learning progress.

[0511] Step 7:

[0512] Users receive online instruction and feedback as needed. Input consists of the user's performance status and advice from the instructor, while output is feedback to the user. The server communicates in real time, delivering information to the user.

[0513] (Application Example 1)

[0514] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0515] To efficiently improve the skills of operators in the workplace, it is necessary to provide individual motion analysis and appropriate feedback promptly. However, with current technology, it is difficult to accurately grasp the differences in operation among operators and point out areas for improvement in real time. Furthermore, there is a lack of support for operators with specific needs, such as those with hearing impairments.

[0516] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0517] In this invention, the server includes: a portable terminal means equipped with a camera for acquiring the operator's work movements; an information processing means equipped with an analysis device for analyzing the acquired video data of the work movements and extracting the operator's morphological information; and an information processing means equipped with a control device for evaluating the operator's work posture based on the extracted morphological information and generating an educational menu according to the evaluation results. This makes it possible to provide individual motion analysis of the operator and optimal feedback in real time.

[0518] An "operator" refers to a person who operates or manages machinery or equipment at a work site.

[0519] "Operational actions" refer to the specific actions and behaviors that an operator performs when operating machinery or equipment.

[0520] A "recording device" is a device installed to record images or videos, and mainly refers to a camera.

[0521] A "mobile device" refers to a device that is easy to carry and has communication and computer functions, such as a smartphone or tablet.

[0522] "Video data" refers to a file that records visual information acquired by a camera in digital format.

[0523] An "analysis device" refers to a device that has computing resources to process acquired data and extract specific information.

[0524] An "information processing device" refers to a computer system that receives and processes data and generates or transforms necessary information.

[0525] "Morphological information" refers to information about the operator's body position and posture.

[0526] "Evaluation" refers to the process of judging the operator's actions based on extracted data and in accordance with specific criteria.

[0527] "Educational menu" refers to training programs and instructional content designed to improve the skills of operators.

[0528] A "user interface" refers to the screens and operating methods that provide interaction for exchanging information between a system and an operator.

[0529] "Playful elements" refer to game-like features incorporated to make work or learning more enjoyable.

[0530] A "communication device" refers to a device equipped with hardware and software that enables the transmission and reception of data.

[0531] "Standard posture" refers to the ideal or standard posture in an action.

[0532] The system for implementing this invention includes a series of processes for efficiently analyzing the operator's work movements and promoting skill improvement. First, the operator records their work movements as video data using a camera mounted on a mobile terminal. The terminal transmits this video data to an information processing device. The information processing device processes the received video data with an analysis device and extracts the operator's morphological information. Based on this morphological information, it evaluates the operator's work posture and generates a specific training menu using a generated AI model.

[0533] The server presents the educational menu generated above to the operator via a user interface on a mobile device. This educational menu incorporates playful elements, providing an environment where skills can be improved in an enjoyable way. For example, using welding work in an automobile manufacturing plant as an example, the operator films basic movements, and the information processing device analyzes the differences from ideal welding movements. Practice menus to approach ideal movements are provided in a game format, which can increase the operator's motivation to learn.

[0534] Furthermore, the server provides real-time feedback using communication devices and supports remote instruction as needed. In this process, appropriate feedback can be provided using sign language recognition technology for the hearing impaired, creating an environment where all operators can acquire skills equally. An example of a prompt from the generated AI model is, "Based on welding motion analysis data, please suggest a step-by-step practice menu for beginner workers."

[0535] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0536] Step 1:

[0537] The terminal records the operator's actions as video data using a camera. The input is real-time work actions, and the output is a video data file. Specifically, when the operator starts an action, the terminal's camera tracks it and records continuous video.

[0538] Step 2:

[0539] The terminal transmits the acquired video data to the information processing device. The input is the video data generated in step 1, and the output is the data stream sent to the server. Specifically, the terminal performs the process of uploading the data to the specified address on the server via the internet connection.

[0540] Step 3:

[0541] The server processes the received video data with an analysis device and extracts morphological information of the operator. The input is the video data transmitted in step 2, and the output is a dataset showing the position and movement of the operator's body. Specifically, the analysis software (e.g., OpenPose) is used to execute a skeletal tracking algorithm and extract the operator's movement information in real time.

[0542] Step 4:

[0543] The server evaluates the operator's working posture based on the extracted morphological information and generates an educational menu using a generating AI model. The input is the morphological information obtained in step 3, and the output is an educational menu suitable for the operator. Specifically, the AI ​​analysis uses this data to detect shortcomings in the operator's movements and creates a training plan for improvement. An example of a prompt used in the generating AI model is, "Based on the welding motion analysis data, please suggest a step-by-step practice menu for a novice worker."

[0544] Step 5:

[0545] The server sends the generated educational menu to the terminal, which then displays it in a user interface. The input is the educational menu generated in step 4, and the output is the learning program displayed on the terminal screen. Specifically, the terminal analyzes the educational menu and generates a GUI to display it in a way that is easy for the user to understand.

[0546] Step 6:

[0547] The user improves their work based on the provided training menu and re-records the results for feedback. The input is the work action based on the training menu, and the output is new video data for feedback. Specifically, the process involves the user independently training and re-recording the improved action.

[0548] Step 7:

[0549] The server uses communication equipment to provide real-time feedback and, if necessary, supports remote instruction. The input is the video data for feedback obtained in step 6, and the output is specific improvement instructions for the operator. Specifically, the server performs real-time analysis and immediately delivers instructional messages to the operator indicating areas that need improvement.

[0550] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0551] This invention is a music learning system that incorporates an emotion engine to improve the user's performance experience. The user records their instrument performance using a recording device connected to a terminal. The terminal sends this recorded data to a server, which analyzes the user's performance actions in detail.

[0552] The server extracts the user's skeletal information based on the received video data and evaluates their playing form. It also compares the analyzed skeletal information with the model form to analyze performance. Based on this analysis, it identifies areas that need improvement and generates a personalized practice menu.

[0553] Furthermore, this invention utilizes an emotion engine to recognize emotions from the user's facial expressions and movements. This emotion data is used to evaluate the user's motivation and stress levels, and to adjust the generated training menus and interfaces.

[0554] The device integrates this information and presents it to the user through a user interface that incorporates game elements. It detects the user's mood during practice in real time and flexibly adjusts the practice content and teaching methods according to the situation, thereby maximizing learning effectiveness.

[0555] For example, when a user practices the piano, the system analyzes the user's playing and provides feedback on finger position and speed. At the same time, the emotion engine evaluates the user's facial expressions, and if it determines that the user is feeling tense or anxious, it immediately provides relaxation exercises and positive feedback to create a more comfortable learning environment.

[0556] Furthermore, sign language recognition technology is incorporated to support users with hearing impairments, allowing information to be communicated using sign language. This feature creates an environment where all users can learn at their own pace.

[0557] In this way, by incorporating an emotion engine, the user's performance experience becomes an intelligent learning system that responds sensitively not only to individuality but also to emotions, supporting significant improvement in musical technique and sustained motivation to learn.

[0558] The following describes the processing flow.

[0559] Step 1:

[0560] The user activates the recording device connected to the terminal and prepares to begin playing their instrument. The terminal begins recording the performance in real time.

[0561] Step 2:

[0562] The device sends recorded video data to the server. This data contains detailed information about the user's performance and is delivered to the server via the network.

[0563] Step 3:

[0564] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This information serves as the basic data for evaluating the performance form.

[0565] Step 4:

[0566] The server compares the extracted skeletal information with the model form and evaluates the accuracy and areas for improvement of the performance. Based on this evaluation, it generates a practice menu tailored to the user.

[0567] Step 5:

[0568] The server uses an emotion engine to analyze the user's facial expressions from video data and evaluate their emotions. Understanding their emotional state provides information to adjust the user's practice environment.

[0569] Step 6:

[0570] The device presents the user with practice menus received from the server and emotionally-based adjustment information. The presentation utilizes an interface that includes game elements to enhance the user's motivation to learn.

[0571] Step 7:

[0572] The user practices according to the provided training menu. During this time, the device continuously monitors the user's facial expressions and sends any changes to the server.

[0573] Step 8:

[0574] The server adjusts the content of feedback and online guidance as needed based on the emotional data received in real time. It also uses an emotional engine to provide guidance that helps users relax.

[0575] Step 9:

[0576] The emotion engine activates features such as adding relaxation exercises to the training menu or sending encouraging messages when the user is feeling tense or anxious.

[0577] (Example 2)

[0578] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0579] Providing an effective learning environment that fully considers each user's individual skills and psychological state in music education is challenging. Furthermore, providing appropriate instruction and feedback to all users, including those with hearing impairments, is also a challenge. In addition, it is necessary to provide an autonomous and individualized learning experience by recognizing emotions in real time and utilizing that information in the learning process.

[0580] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0581] In this invention, the server includes: an information processing device equipped with an analysis device that analyzes acquired motion data and extracts user structural information; an information processing device equipped with a control device that evaluates the user's movements based on the extracted structural information and generates a training plan according to the evaluation results; and a means further equipped with a function that recognizes the user's psychological state using emotion analysis technology and makes adjustments to the training plan and operation screen. This makes it possible to provide users with an individualized learning environment while adjusting learning based on emotions in real time.

[0582] A "recording device" is a device used to photograph a user's actions and record them as data.

[0583] "Device" is a broad term referring to hardware or software used to perform a specific function.

[0584] An "information processing device" is a computer or system used to process data and generate the necessary output.

[0585] "Structural information" refers to data that represents the position and movement of the user's body, obtained through analysis.

[0586] A "control device" is a device that includes a mechanism for managing a specific process or device and controlling its operation in accordance with its purpose.

[0587] A "training plan" refers to a series of practice menus and tasks aimed at improving the user's skills.

[0588] "Entertainment elements" are elements designed to attract users' interest and promote learning while they enjoy themselves.

[0589] A "communication network" refers to the network infrastructure used for sending and receiving information.

[0590] "Emotional analysis technology" is a technology that evaluates a user's psychological state based on their facial expressions and actions.

[0591] A "standard movement" refers to a typical or ideal movement that serves as the basis for evaluation.

[0592] This invention is a system for improving music learning. The system is composed of a combination of a recording device, an information processing device, a control device, emotion analysis technology, and the like.

[0593] Users record their musical performances using a recording device connected to their terminal. This recording device is capable of high-resolution data collection, capturing the user's movements and facial expressions in detail. The recorded data is transmitted digitally to an information processing device.

[0594] The terminal, as an information processing device, analyzes data and extracts the user's structural information. This process utilizes known libraries and algorithms (e.g., computer vision technology). Structural information refers to data that systematically represents the user's body position and movement.

[0595] The server evaluates the user's playing form based on the obtained structural information. This evaluation involves comparison with a baseline performance, and a training plan is generated based on the results. By using a generative AI model, a learning program optimized for the user is provided.

[0596] Furthermore, the server utilizes emotion analysis technology to evaluate the user's psychological state in real time. Emotional data is extracted from the user's facial expressions and actions, and the training plan and user interface are adjusted accordingly.

[0597] The device presents the generated training plan to the user as an interface. The interface incorporates entertainment elements to make learning enjoyable for the user. For example, if the user is practicing the piano, the system provides feedback on finger position and speed, and suggests relaxation exercises as needed.

[0598] An example of a prompt message would be, "Send this data to the server and have it analyzed using the sentiment engine." This example helps users understand exactly how to use the system.

[0599] In this way, this invention can maximize the effectiveness of music learning by providing an individualized learning environment.

[0600] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0601] Step 1:

[0602] The user begins playing and records the performance using a recording device connected to the terminal. The input for recording is the user's movements and facial expressions while playing. This data is recorded in high resolution and includes the user's subtle movements and expressions. Digital video data is generated as output.

[0603] Step 2:

[0604] The terminal transmits the recorded video data to the information processing device. The input is the video data generated in step 1. The information processing device receives this data and converts it into a data format suitable for analysis. The output is the video data formatted into an analyzable format.

[0605] Step 3:

[0606] The server extracts the user's structural information based on the video data transmitted from the terminal. The input for this step is analyzable video data. The server uses computer vision algorithms to identify the user's body position and the movement of each part, and generates a digital skeletal model. The output is structural information data that provides a detailed representation of the user's movements.

[0607] Step 4:

[0608] The server evaluates the user's playing form based on the structural information obtained. The input is structural information data. The control unit within the server compares the form to a standard operation and determines whether the form is good or bad. Based on the evaluation results, specific areas for improvement are identified. The output is an evaluation report that includes the areas for improvement.

[0609] Step 5:

[0610] The server generates a training plan optimized for the user based on the evaluation report. The input is the evaluation report obtained in step 4. The generated AI model is used to create individual training content and tasks. The output is a training plan customized for the user.

[0611] Step 6:

[0612] The server uses emotion analysis technology to analyze the user's psychological state from video footage. The input is video data, including the user's facial expressions. The server uses AI to evaluate the user's emotional state. The output is emotion evaluation data.

[0613] Step 7:

[0614] The terminal presents the user with an operation screen based on the training plan and emotion assessment data received from the server. The input consists of the user's training plan and emotion assessment data. The operation screen reflects real-time feedback tailored to the user's emotions. The output is an interactive user interface.

[0615] Step 8:

[0616] The device optimizes the user's learning experience based on information received during performance. Input consists of real-time user movement and emotion data. This allows for immediate suggestions of the most effective practice and instruction methods based on the user's current state. Output consists of suggestions and instructions to facilitate the user's skill improvement.

[0617] (Application Example 2)

[0618] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0619] In modern society, a problem exists in the limited opportunities and instruction available for improving musical instrument playing skills. Furthermore, providing feedback that takes into account the user's unique emotional state is difficult, resulting in challenges in maintaining user motivation. Additionally, there is a lack of efficient mechanisms for providing real-time, individualized instruction and feedback in physical stores where direct experiences such as instrument trials take place. These circumstances make it difficult to achieve sustained and efficient improvement in playing skills.

[0620] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0621] In this invention, the server includes terminal means equipped with a camera for capturing user movements, information processing means equipped with an analysis device for analyzing acquired video data of movements and extracting user physical information, and information processing means equipped with an emotion analysis device for analyzing facial emotions and providing feedback according to the emotional state. This makes it possible to provide real-time feedback tailored to each individual user and appropriate guidance based on their emotional state.

[0622] A "camera for capturing user movements" is a device used to visually capture the physical movements of a user while they are performing an instrument or engaging in an activity.

[0623] An "information processing device" is a computer system used to analyze, store, or transmit digital data.

[0624] "Physical information" refers to data about the user's body movements and posture, specifically including information such as the position and angle of the skeleton.

[0625] A "control device" is a device that performs processing according to a specific purpose based on the data it receives.

[0626] A "display surface" refers to a visual device or screen used to present information to a user.

[0627] A "communication device" is a device used to send and receive data and instructions, and has the function of connecting to servers and terminals via a network.

[0628] An "emotion analysis device" is a device that analyzes a user's emotions from their facial expressions and actions.

[0629] "Feedback" refers to information that provides guidance and evaluation to the user based on the data and analysis results obtained.

[0630] To implement this invention, a system is constructed that integrates a series of devices and software. The server receives video data from a terminal connected to a camera and smart glasses for capturing the user's movements. The terminal's camera captures the user's performance and movements in high resolution and transmits the video to the server. At this time, image processing software such as OpenCV is used to capture the video data in real time.

[0631] The server uses an information processing device to analyze the received video data. Specifically, it uses libraries such as PoseNet and OpenPose to extract skeletal information. This extracts the user's physical information as digital data. This data is compared with a pre-configured model form to evaluate the performance form. The emotion analysis device uses emotion recognition software such as facial_emotion_recognition to analyze the user's facial expressions from the video data and understand their emotional state. Based on this information, the server generates appropriate feedback according to the user's emotional state.

[0632] User feedback is presented in real time on the display screen. Through the interface of the display device or smart glasses, users receive visual and audible feedback. This allows users to correct their technique on the spot and improve their playing skills. For example, if the server analyzes the user's finger movements during a performance and detects incorrect form, a guide to correct it will appear on the display. At the same time, if the user is feeling tense or anxious, positive feedback to encourage relaxation is also provided.

[0633] An example of a prompt used to implement this system is: "Analyze the user's performance video and evaluate areas for improvement in their playing form and emotional state. If the user is nervous, provide advice to help them relax and generate positive feedback." Based on this prompt, the generating AI model provides instruction optimized for the user.

[0634] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0635] Step 1:

[0636] The user begins playing music using smart glasses or a camera-equipped device, recording their performance in real time. The input data is the user's performance video, which the device continuously captures. This data is saved digitally as it is necessary for subsequent processing.

[0637] Step 2:

[0638] The terminal sends the recorded performance video data to the server. The server receives this data and stores it in a database. The input is the user's performance video data, and the output is the raw video data sent to the server. The transmitted data is necessary to prepare the video for analysis.

[0639] Step 3:

[0640] The server uses PoseNet and OpenPose to analyze the user's body information from received video data and extract skeletal information. The input is video data of the user performing, and by analyzing this data, the user's skeletal information is extracted. The output is digital data related to the user's body movements. This skeletal information is used to evaluate form.

[0641] Step 4:

[0642] The server uses emotion recognition software, such as facial_emotion_recognition, to analyze the user's facial expressions from the video and obtain emotional information. The input is the user's video data, and the output is data indicating the user's emotional state. Through this process, the user's motivation and tension level are evaluated.

[0643] Step 5:

[0644] The server compares the extracted skeletal information with a model form to evaluate the performance form and generates feedback based on emotional information. The inputs are skeletal information and emotional information, and the output is feedback data provided to the user. This feedback includes areas for improvement in performance technique and advice tailored to the emotional state.

[0645] Step 6:

[0646] The generated feedback is displayed on the device's screen in real time. Users receive this feedback through smart glasses or a display. The input is feedback data, and the output is instruction and advice conveyed to the user. This allows users to improve their performance and maintain their motivation.

[0647] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0648] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0649] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0650] [Fourth Embodiment]

[0651] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0652] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0653] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0654] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0655] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0656] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0657] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0658] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0659] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0660] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0661] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0662] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0663] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0664] This invention is a music learning system designed to improve the user's performance skills. The user records their performance via a recording device connected to a terminal. The terminal has the function of transmitting this video data to a server.

[0665] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This skeletal information allows for an accurate evaluation of the user's playing form. The analysis results are compared on the server with exemplary form and used to identify areas for improvement and strengths of the user.

[0666] Based on the evaluation of the user's playing form, the server generates an individually optimized practice menu. This menu includes a variety of exercises aimed at improving playing technique, and is composed of appropriate exercises selected from a database managed by the server.

[0667] The generated practice menu is presented to the user via the device. The device uses a user interface incorporating game elements to visualize learning progress and enhance the user's motivation to learn. Furthermore, users can receive online instruction at any time, and the server supports real-time communication for this purpose.

[0668] Furthermore, as a consideration for people with hearing impairments, the terminal utilizes sign language recognition technology to provide an environment where users can confirm necessary information in sign language. The server generates appropriate feedback corresponding to this and presents it through the terminal.

[0669] For example, when a user learns to play the guitar, the performance data captured by the camera is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server then suggests specific practice menus regarding finger movements and strokes. These are displayed on the device, allowing the user to improve their playing skills in an enjoyable way while receiving instruction.

[0670] In this way, the present invention provides users with a personalized and effective music learning experience, supporting the improvement of their performance skills. This system creates a more efficient and engaging learning environment than conventional methods.

[0671] The following describes the processing flow.

[0672] Step 1:

[0673] The user activates the device's recording function and prepares to record their musical performance. Once ready, the device starts recording, capturing the performance in real time.

[0674] Step 2:

[0675] The device sends recorded video data to the server. This transmission is primarily done over the network, enabling real-time data reception.

[0676] Step 3:

[0677] The server receives video data transmitted from the terminal and performs skeletal detection using an analysis device. This extracts skeletal information such as the user's hands and arms.

[0678] Step 4:

[0679] The server compares the user's playing form to a model form based on the extracted skeletal information. It then identifies areas that need improvement based on the form's evaluation criteria.

[0680] Step 5:

[0681] The server automatically generates individually optimized practice menus based on the form evaluation results. These menus will include specific technical exercises and formal drills.

[0682] Step 6:

[0683] The device presents the generated practice menu to the user. A gamified user interface is used to enhance learning motivation during the presentation.

[0684] Step 7:

[0685] The user practices based on the provided menu, records their performance, and sends it from their device to the server. They then receive feedback from the server.

[0686] Step 8:

[0687] The server evaluates the video data sent back by the user and provides online direct instruction and feedback as needed. For users with hearing impairments, it also uses sign language recognition technology to provide appropriate feedback.

[0688] (Example 1)

[0689] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0690] [To improve musical performance skills, there is a need for a system that provides efficient practice menus tailored to individual user characteristics, and enables evaluation and improvement of performance. Furthermore, an environment is needed where multiple learners, each in different circumstances, can receive appropriate feedback online.]

[0691] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0692] In this invention, the server includes means for [equipped with an information processing device for extracting target operation information], means for [equipped with a control device for generating a practice plan], and means for [equipped with a communication device for providing instruction and feedback over a communication network]. This enables [the provision of personalized practice menus to users and efficient performance evaluation and instruction via online means].

[0693] An "image acquisition device" refers to a device used to acquire video footage of a subject's performance.

[0694] An "information processing device" refers to a computer device used to analyze acquired data and extract and process specific information.

[0695] "Motion information" refers to information extracted as data related to musical performance and body movements.

[0696] A "control device" refers to a device that has the function of performing specific instructions or actions based on data.

[0697] A "practice plan" refers to a set of exercises and training guidelines developed to improve the user's playing technique.

[0698] "User display device" refers to a terminal or device used to visually present information.

[0699] "Communication equipment" refers to hardware or software used to send and receive data and instructions over a network.

[0700] A "reference pattern" refers to information set as model data that represents ideal behavior or performance.

[0701] This invention is a learning system aimed at improving the user's musical performance skills. The user can acquire video footage of their performance using an image acquisition device connected to an information processing device. The acquired video footage is transmitted from the information processing device to a server via a communication device.

[0702] The server is equipped with an information processing device capable of analyzing the received video data, for example, by using a skeletal analysis algorithm to extract user movement information. This makes it possible to accurately recognize movements during performance and compare them with reference patterns stored in a database.

[0703] The server uses a control device to evaluate the user's performance based on these comparison results. The practice plan generated from the evaluation results is optimized for the user and presented to the user display device via an information processing device. This allows the user to progress through their learning while visually checking their progress through an interface that includes game elements.

[0704] Furthermore, the server enables online instruction and feedback using communication equipment. This functionality allows users to receive real-time instruction even from remote locations.

[0705] Furthermore, by utilizing motion recognition technology, it is possible to provide support tailored to users with specific needs. The goal is to create a learning environment suitable for people with hearing impairments.

[0706] For example, when a user learns to play the guitar, data captured by an image acquisition device is analyzed on a server to evaluate the user's hand position and form. Based on the analysis results, the server proposes a practice plan for finger movements and strokes, which is displayed on the user's display device. This allows the user to improve their playing skills while having fun.

[0707] A concrete example of a prompt message to be input into the generating AI model is, "Analyze the user's playing form and suggest the optimal practice program." In this way, the present invention provides users with a personalized and efficient music learning experience and supports the improvement of their playing skills.

[0708] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0709] Step 1:

[0710] The user uses an image acquisition device connected to an information processing device to acquire video data of their performance. At this stage, the input is the user's performance, and the output is image data. The user plays the instrument, and their actions are recorded as video.

[0711] Step 2:

[0712] The terminal sends this acquired video data to the server. In this step, the input is image data, and the output is data transfer to the server. The terminal uploads the video data over the network.

[0713] Step 3:

[0714] The server receives the transmitted video data and extracts motion information using an analysis device. The input here is video data, and the output is motion information data. Specifically, the server applies an analysis model to calculate the performer's skeletal structure and movement patterns.

[0715] Step 4:

[0716] The server compares the extracted motion information with a reference pattern to evaluate the performance form. The input for this step is motion information data and a reference pattern, and the output is the evaluation result. The server uses an evaluation algorithm to identify areas that can be improved.

[0717] Step 5:

[0718] The server generates a user-specific practice plan based on the evaluation results. The input for this step is the evaluation results, and the output is the practice plan. The server selects appropriate practice items from the database and constructs a menu tailored to the user.

[0719] Step 6:

[0720] The terminal presents the generated practice plan to the user via a user display device. Here, the input is the practice plan, and the output is a visual learning interface. Based on the information displayed on the screen, the user checks their learning progress.

[0721] Step 7:

[0722] Users receive online instruction and feedback as needed. Input consists of the user's performance status and advice from the instructor, while output is feedback to the user. The server communicates in real time, delivering information to the user.

[0723] (Application Example 1)

[0724] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0725] To efficiently improve the skills of operators in the workplace, it is necessary to provide individual motion analysis and appropriate feedback promptly. However, with current technology, it is difficult to accurately grasp the differences in operation among operators and point out areas for improvement in real time. Furthermore, there is a lack of support for operators with specific needs, such as those with hearing impairments.

[0726] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0727] In this invention, the server includes: a portable terminal means equipped with a camera for acquiring the operator's work movements; an information processing means equipped with an analysis device for analyzing the acquired video data of the work movements and extracting the operator's morphological information; and an information processing means equipped with a control device for evaluating the operator's work posture based on the extracted morphological information and generating an educational menu according to the evaluation results. This makes it possible to provide individual motion analysis of the operator and optimal feedback in real time.

[0728] An "operator" refers to a person who operates or manages machinery or equipment at a work site.

[0729] "Operational actions" refer to the specific actions and behaviors that an operator performs when operating machinery or equipment.

[0730] A "recording device" is a device installed to record images or videos, and mainly refers to a camera.

[0731] A "mobile device" refers to a device that is easy to carry and has communication and computer functions, such as a smartphone or tablet.

[0732] "Video data" refers to a file that records visual information acquired by a camera in digital format.

[0733] An "analysis device" refers to a device that has computing resources to process acquired data and extract specific information.

[0734] An "information processing device" refers to a computer system that receives and processes data and generates or transforms necessary information.

[0735] "Morphological information" refers to information about the operator's body position and posture.

[0736] "Evaluation" refers to the process of judging the operator's actions based on extracted data and in accordance with specific criteria.

[0737] "Educational menu" refers to training programs and instructional content designed to improve the skills of operators.

[0738] A "user interface" refers to the screens and operating methods that provide interaction for exchanging information between a system and an operator.

[0739] "Playful elements" refer to game-like features incorporated to make work or learning more enjoyable.

[0740] A "communication device" refers to a device equipped with hardware and software that enables the transmission and reception of data.

[0741] "Standard posture" refers to the ideal or standard posture in an action.

[0742] The system for implementing this invention includes a series of processes for efficiently analyzing the operator's work movements and promoting skill improvement. First, the operator records their work movements as video data using a camera mounted on a mobile terminal. The terminal transmits this video data to an information processing device. The information processing device processes the received video data with an analysis device and extracts the operator's morphological information. Based on this morphological information, it evaluates the operator's work posture and generates a specific training menu using a generated AI model.

[0743] The server presents the educational menu generated above to the operator via a user interface on a mobile device. This educational menu incorporates playful elements, providing an environment where skills can be improved in an enjoyable way. For example, using welding work in an automobile manufacturing plant as an example, the operator films basic movements, and the information processing device analyzes the differences from ideal welding movements. Practice menus to approach ideal movements are provided in a game format, which can increase the operator's motivation to learn.

[0744] Furthermore, the server provides real-time feedback using communication devices and supports remote instruction as needed. In this process, appropriate feedback can be provided using sign language recognition technology for the hearing impaired, creating an environment where all operators can acquire skills equally. An example of a prompt from the generated AI model is, "Based on welding motion analysis data, please suggest a step-by-step practice menu for beginner workers."

[0745] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0746] Step 1:

[0747] The terminal records the operator's actions as video data using a camera. The input is real-time work actions, and the output is a video data file. Specifically, when the operator starts an action, the terminal's camera tracks it and records continuous video.

[0748] Step 2:

[0749] The terminal transmits the acquired video data to the information processing device. The input is the video data generated in step 1, and the output is the data stream sent to the server. Specifically, the terminal performs the process of uploading the data to the specified address on the server via the internet connection.

[0750] Step 3:

[0751] The server processes the received video data with an analysis device and extracts morphological information of the operator. The input is the video data transmitted in step 2, and the output is a dataset showing the position and movement of the operator's body. Specifically, the analysis software (e.g., OpenPose) is used to execute a skeletal tracking algorithm and extract the operator's movement information in real time.

[0752] Step 4:

[0753] The server evaluates the operator's working posture based on the extracted morphological information and generates an educational menu using a generating AI model. The input is the morphological information obtained in step 3, and the output is an educational menu suitable for the operator. Specifically, the AI ​​analysis uses this data to detect shortcomings in the operator's movements and creates a training plan for improvement. An example of a prompt used in the generating AI model is, "Based on the welding motion analysis data, please suggest a step-by-step practice menu for a novice worker."

[0754] Step 5:

[0755] The server sends the generated educational menu to the terminal, which then displays it in a user interface. The input is the educational menu generated in step 4, and the output is the learning program displayed on the terminal screen. Specifically, the terminal analyzes the educational menu and generates a GUI to display it in a way that is easy for the user to understand.

[0756] Step 6:

[0757] The user improves their work based on the provided training menu and re-records the results for feedback. The input is the work action based on the training menu, and the output is new video data for feedback. Specifically, the process involves the user independently training and re-recording the improved action.

[0758] Step 7:

[0759] The server uses communication equipment to provide real-time feedback and, if necessary, supports remote instruction. The input is the video data for feedback obtained in step 6, and the output is specific improvement instructions for the operator. Specifically, the server performs real-time analysis and immediately delivers instructional messages to the operator indicating areas that need improvement.

[0760] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0761] This invention is a music learning system that incorporates an emotion engine to improve the user's performance experience. The user records their instrument performance using a recording device connected to a terminal. The terminal sends this recorded data to a server, which analyzes the user's performance actions in detail.

[0762] The server extracts the user's skeletal information based on the received video data and evaluates their playing form. It also compares the analyzed skeletal information with the model form to analyze performance. Based on this analysis, it identifies areas that need improvement and generates a personalized practice menu.

[0763] Furthermore, this invention utilizes an emotion engine to recognize emotions from the user's facial expressions and movements. This emotion data is used to evaluate the user's motivation and stress levels, and to adjust the generated training menus and interfaces.

[0764] The device integrates this information and presents it to the user through a user interface that incorporates game elements. It detects the user's mood during practice in real time and flexibly adjusts the practice content and teaching methods according to the situation, thereby maximizing learning effectiveness.

[0765] For example, when a user practices the piano, the system analyzes the user's playing and provides feedback on finger position and speed. At the same time, the emotion engine evaluates the user's facial expressions, and if it determines that the user is feeling tense or anxious, it immediately provides relaxation exercises and positive feedback to create a more comfortable learning environment.

[0766] Furthermore, sign language recognition technology is incorporated to support users with hearing impairments, allowing information to be communicated using sign language. This feature creates an environment where all users can learn at their own pace.

[0767] In this way, by incorporating an emotion engine, the user's performance experience becomes an intelligent learning system that responds sensitively not only to individuality but also to emotions, supporting significant improvement in musical technique and sustained motivation to learn.

[0768] The following describes the processing flow.

[0769] Step 1:

[0770] The user activates the recording device connected to the terminal and prepares to begin playing their instrument. The terminal begins recording the performance in real time.

[0771] Step 2:

[0772] The device sends recorded video data to the server. This data contains detailed information about the user's performance and is delivered to the server via the network.

[0773] Step 3:

[0774] The server receives the transmitted video data and uses an analysis device to extract the user's skeletal information. This information serves as the basic data for evaluating the performance form.

[0775] Step 4:

[0776] The server compares the extracted skeletal information with the model form and evaluates the accuracy and areas for improvement of the performance. Based on this evaluation, it generates a practice menu tailored to the user.

[0777] Step 5:

[0778] The server uses an emotion engine to analyze the user's facial expressions from video data and evaluate their emotions. Understanding their emotional state provides information to adjust the user's practice environment.

[0779] Step 6:

[0780] The device presents the user with practice menus received from the server and emotionally-based adjustment information. The presentation utilizes an interface that includes game elements to enhance the user's motivation to learn.

[0781] Step 7:

[0782] The user practices according to the provided training menu. During this time, the device continuously monitors the user's facial expressions and sends any changes to the server.

[0783] Step 8:

[0784] The server adjusts the content of feedback and online guidance as needed based on the emotional data received in real time. It also uses an emotional engine to provide guidance that helps users relax.

[0785] Step 9:

[0786] The emotion engine activates features such as adding relaxation exercises to the training menu or sending encouraging messages when the user is feeling tense or anxious.

[0787] (Example 2)

[0788] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0789] Providing an effective learning environment that fully considers each user's individual skills and psychological state in music education is challenging. Furthermore, providing appropriate instruction and feedback to all users, including those with hearing impairments, is also a challenge. In addition, it is necessary to provide an autonomous and individualized learning experience by recognizing emotions in real time and utilizing that information in the learning process.

[0790] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0791] In this invention, the server includes: an information processing device equipped with an analysis device that analyzes acquired motion data and extracts user structural information; an information processing device equipped with a control device that evaluates the user's movements based on the extracted structural information and generates a training plan according to the evaluation results; and a means further equipped with a function that recognizes the user's psychological state using emotion analysis technology and makes adjustments to the training plan and operation screen. This makes it possible to provide users with an individualized learning environment while adjusting learning based on emotions in real time.

[0792] A "recording device" is a device used to photograph a user's actions and record them as data.

[0793] "Device" is a broad term referring to hardware or software used to perform a specific function.

[0794] An "information processing device" is a computer or system used to process data and generate the necessary output.

[0795] "Structural information" refers to data that represents the position and movement of the user's body, obtained through analysis.

[0796] A "control device" is a device that includes a mechanism for managing a specific process or device and controlling its operation in accordance with its purpose.

[0797] A "training plan" refers to a series of practice menus and tasks aimed at improving the user's skills.

[0798] "Entertainment elements" are elements designed to attract users' interest and promote learning while they enjoy themselves.

[0799] A "communication network" refers to the network infrastructure used for sending and receiving information.

[0800] "Emotional analysis technology" is a technology that evaluates a user's psychological state based on their facial expressions and actions.

[0801] A "standard movement" refers to a typical or ideal movement that serves as the basis for evaluation.

[0802] This invention is a system for improving music learning. The system is composed of a combination of a recording device, an information processing device, a control device, emotion analysis technology, and the like.

[0803] Users record their musical performances using a recording device connected to their terminal. This recording device is capable of high-resolution data collection, capturing the user's movements and facial expressions in detail. The recorded data is transmitted digitally to an information processing device.

[0804] The terminal, as an information processing device, analyzes data and extracts the user's structural information. This process utilizes known libraries and algorithms (e.g., computer vision technology). Structural information refers to data that systematically represents the user's body position and movement.

[0805] The server evaluates the user's playing form based on the obtained structural information. This evaluation involves comparison with a baseline performance, and a training plan is generated based on the results. By using a generative AI model, a learning program optimized for the user is provided.

[0806] Furthermore, the server utilizes emotion analysis technology to evaluate the user's psychological state in real time. Emotional data is extracted from the user's facial expressions and actions, and the training plan and user interface are adjusted accordingly.

[0807] The device presents the generated training plan to the user as an interface. The interface incorporates entertainment elements to make learning enjoyable for the user. For example, if the user is practicing the piano, the system provides feedback on finger position and speed, and suggests relaxation exercises as needed.

[0808] An example of a prompt message would be, "Send this data to the server and have it analyzed using the sentiment engine." This example helps users understand exactly how to use the system.

[0809] In this way, this invention can maximize the effectiveness of music learning by providing an individualized learning environment.

[0810] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0811] Step 1:

[0812] The user begins playing and records the performance using a recording device connected to the terminal. The input for recording is the user's movements and facial expressions while playing. This data is recorded in high resolution and includes the user's subtle movements and expressions. Digital video data is generated as output.

[0813] Step 2:

[0814] The terminal transmits the recorded video data to the information processing device. The input is the video data generated in step 1. The information processing device receives this data and converts it into a data format suitable for analysis. The output is the video data formatted into an analyzable format.

[0815] Step 3:

[0816] The server extracts the user's structural information based on the video data transmitted from the terminal. The input for this step is analyzable video data. The server uses computer vision algorithms to identify the user's body position and the movement of each part, and generates a digital skeletal model. The output is structural information data that provides a detailed representation of the user's movements.

[0817] Step 4:

[0818] The server evaluates the user's playing form based on the structural information obtained. The input is structural information data. The control unit within the server compares the form to a standard operation and determines whether the form is good or bad. Based on the evaluation results, specific areas for improvement are identified. The output is an evaluation report that includes the areas for improvement.

[0819] Step 5:

[0820] The server generates a training plan optimized for the user based on the evaluation report. The input is the evaluation report obtained in step 4. The generated AI model is used to create individual training content and tasks. The output is a training plan customized for the user.

[0821] Step 6:

[0822] The server uses emotion analysis technology to analyze the user's psychological state from video footage. The input is video data, including the user's facial expressions. The server uses AI to evaluate the user's emotional state. The output is emotion evaluation data.

[0823] Step 7:

[0824] The terminal presents the user with an operation screen based on the training plan and emotion assessment data received from the server. The input consists of the user's training plan and emotion assessment data. The operation screen reflects real-time feedback tailored to the user's emotions. The output is an interactive user interface.

[0825] Step 8:

[0826] The device optimizes the user's learning experience based on information received during performance. Input consists of real-time user movement and emotion data. This allows for immediate suggestions of the most effective practice and instruction methods based on the user's current state. Output consists of suggestions and instructions to facilitate the user's skill improvement.

[0827] (Application Example 2)

[0828] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0829] In modern society, a problem exists in the limited opportunities and instruction available for improving musical instrument playing skills. Furthermore, providing feedback that takes into account the user's unique emotional state is difficult, resulting in challenges in maintaining user motivation. Additionally, there is a lack of efficient mechanisms for providing real-time, individualized instruction and feedback in physical stores where direct experiences such as instrument trials take place. These circumstances make it difficult to achieve sustained and efficient improvement in playing skills.

[0830] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0831] In this invention, the server includes terminal means equipped with a camera for capturing user movements, information processing means equipped with an analysis device for analyzing acquired video data of movements and extracting user physical information, and information processing means equipped with an emotion analysis device for analyzing facial emotions and providing feedback according to the emotional state. This makes it possible to provide real-time feedback tailored to each individual user and appropriate guidance based on their emotional state.

[0832] A "camera for capturing user movements" is a device used to visually capture the physical movements of a user while they are performing an instrument or engaging in an activity.

[0833] An "information processing device" is a computer system used to analyze, store, or transmit digital data.

[0834] "Physical information" refers to data about the user's body movements and posture, specifically including information such as the position and angle of the skeleton.

[0835] A "control device" is a device that performs processing according to a specific purpose based on the data it receives.

[0836] A "display surface" refers to a visual device or screen used to present information to a user.

[0837] A "communication device" is a device used to send and receive data and instructions, and has the function of connecting to servers and terminals via a network.

[0838] An "emotion analysis device" is a device that analyzes a user's emotions from their facial expressions and actions.

[0839] "Feedback" refers to information that provides guidance and evaluation to the user based on the data and analysis results obtained.

[0840] To implement this invention, a system is constructed that integrates a series of devices and software. The server receives video data from a terminal connected to a camera and smart glasses for capturing the user's movements. The terminal's camera captures the user's performance and movements in high resolution and transmits the video to the server. At this time, image processing software such as OpenCV is used to capture the video data in real time.

[0841] The server uses an information processing device to analyze the received video data. Specifically, it uses libraries such as PoseNet and OpenPose to extract skeletal information. This extracts the user's physical information as digital data. This data is compared with a pre-configured model form to evaluate the performance form. The emotion analysis device uses emotion recognition software such as facial_emotion_recognition to analyze the user's facial expressions from the video data and understand their emotional state. Based on this information, the server generates appropriate feedback according to the user's emotional state.

[0842] User feedback is presented in real time on the display screen. Through the interface of the display device or smart glasses, users receive visual and audible feedback. This allows users to correct their technique on the spot and improve their playing skills. For example, if the server analyzes the user's finger movements during a performance and detects incorrect form, a guide to correct it will appear on the display. At the same time, if the user is feeling tense or anxious, positive feedback to encourage relaxation is also provided.

[0843] An example of a prompt used to implement this system is: "Analyze the user's performance video and evaluate areas for improvement in their playing form and emotional state. If the user is nervous, provide advice to help them relax and generate positive feedback." Based on this prompt, the generating AI model provides instruction optimized for the user.

[0844] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0845] Step 1:

[0846] The user begins playing music using smart glasses or a camera-equipped device, recording their performance in real time. The input data is the user's performance video, which the device continuously captures. This data is saved digitally as it is necessary for subsequent processing.

[0847] Step 2:

[0848] The terminal sends the recorded performance video data to the server. The server receives this data and stores it in a database. The input is the user's performance video data, and the output is the raw video data sent to the server. The transmitted data is necessary to prepare the video for analysis.

[0849] Step 3:

[0850] The server uses PoseNet and OpenPose to analyze the user's body information from received video data and extract skeletal information. The input is video data of the user performing, and by analyzing this data, the user's skeletal information is extracted. The output is digital data related to the user's body movements. This skeletal information is used to evaluate form.

[0851] Step 4:

[0852] The server uses emotion recognition software, such as facial_emotion_recognition, to analyze the user's facial expressions from the video and obtain emotional information. The input is the user's video data, and the output is data indicating the user's emotional state. Through this process, the user's motivation and tension level are evaluated.

[0853] Step 5:

[0854] The server compares the extracted skeletal information with a model form to evaluate the performance form and generates feedback based on emotional information. The inputs are skeletal information and emotional information, and the output is feedback data provided to the user. This feedback includes areas for improvement in performance technique and advice tailored to the emotional state.

[0855] Step 6:

[0856] The generated feedback is displayed on the device's screen in real time. Users receive this feedback through smart glasses or a display. The input is feedback data, and the output is instruction and advice conveyed to the user. This allows users to improve their performance and maintain their motivation.

[0857] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0858] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0859] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0860] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0861] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0862] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0863] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0864] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0865] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0866] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0867] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0868] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0869] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0870] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0871] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0872] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0873] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0874] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0875] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0876] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0877] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0878] The following is further disclosed regarding the embodiments described above.

[0879] (Claim 1)

[0880] [A terminal means equipped with a camera for capturing the user's performance actions,

[0881] [A server means equipped with an analysis device that analyzes acquired video data of performance movements and extracts the user's skeletal information,

[0882] [A server means equipped with a control device that evaluates the user's playing form based on extracted skeletal information and generates a practice menu according to the evaluation results,

[0883] [A terminal device that presents the generated practice menu as a user interface and provides a learning environment that incorporates game elements,

[0884] [A server means equipped with a communication device that provides online instruction and feedback to users as they progress through their musical activities,

[0885] A system that includes this.

[0886] (Claim 2)

[0887] The system according to claim 1, further comprising a function to provide performance support for hearing-impaired persons using sign language recognition technology.

[0888] (Claim 3)

[0889] The system according to claim 1, further comprising a function to compare the user's playing form with a model form and a function to provide guidance for improving playing technique.

[0890] "Example 1"

[0891] (Claim 1)

[0892] [An information processing device equipped with an image acquisition device for acquiring the performance actions of the target,

[0893] [An information processing device that analyzes video information of acquired performance movements and extracts information about the target movements,

[0894] [An information processing device equipped with a control device that evaluates the target performance pattern based on extracted motion information and generates a practice plan according to the evaluation result,

[0895] [An information processing device that presents the generated practice plan as a user display device and provides a learning environment that incorporates competitive elements,

[0896] [An information processing device equipped with a communication device that provides guidance and feedback via a communication network when the subject is performing a musical act,

[0897] A system that includes this.

[0898] (Claim 2)

[0899] The system according to claim 1, further comprising a function that provides performance support to individuals with specific needs using motion recognition technology.

[0900] (Claim 3)

[0901] The system according to claim 1, further comprising a function to compare the target performance pattern with a reference pattern and a function to provide instruction for improving performance technique.

[0902] "Application Example 1"

[0903] (Claim 1)

[0904] [A portable terminal equipped with a camera for capturing the operator's work movements,

[0905] [An information processing device equipped with an analysis device that analyzes acquired video data of work movements and extracts morphological information of the operator,

[0906] [An information processing device equipped with a control device that evaluates the operator's working posture based on extracted morphological information and generates an educational menu according to the evaluation results,

[0907] [A mobile device that presents the generated educational menu as a user interface and provides a learning environment that incorporates playful elements,

[0908] [An information processing device equipped with a communication device that provides remote guidance and feedback to the operator as they carry out work activities,

[0909] A system that includes this.

[0910] (Claim 2)

[0911] The system according to claim 1, further comprising a function to provide work support for hearing-impaired persons using sign language recognition technology.

[0912] (Claim 3)

[0913] The system according to claim 1, further comprising a function to compare the operator's working posture with a standard posture and a function to provide guidance for improving skills.

[0914] "Example 2 of combining an emotion engine"

[0915] (Claim 1)

[0916] [Device means equipped with a recording device for acquiring user actions,

[0917] [An information processing device equipped with an analysis device that analyzes acquired motion data and extracts user structural information,

[0918] [An information processing device equipped with a control device that evaluates the user's movements based on extracted structural information and generates a training plan according to the evaluation results,

[0919] [A device that presents the generated training plan as a screen for the user to operate, and provides a learning environment that incorporates entertainment elements,

[0920] [Information processing device equipped with a communication device that provides guidance and responses via a communication network when a user is carrying out an activity,

[0921] [A means to further incorporate the function of recognizing the user's psychological state using emotion analysis technology and making adjustments to the training plan and operation screen,

[0922] A system that includes this.

[0923] (Claim 2)

[0924] The system according to claim 1, further comprising a function to provide activity support for hearing-impaired persons using manual communication recognition technology.

[0925] (Claim 3)

[0926] The system according to claim 1, further comprising a function to compare the user's movements with reference movements and a function to provide instruction for skill improvement.

[0927] "Application example 2 when combining with an emotional engine"

[0928] (Claim 1)

[0929] [A terminal means equipped with a camera for capturing user actions,

[0930] [An information processing device equipped with an analysis device that analyzes acquired video data of movements and extracts the user's physical information,

[0931] [An information processing device equipped with a control device that evaluates the user's form based on extracted physical information and generates an instruction program according to the evaluation results,

[0932] [A terminal device that displays the generated instructional program on a screen and provides a learning environment that incorporates game elements,

[0933] [An information processing device equipped with a communication device that provides online guidance and feedback to users as they progress through their activities,

[0934] [An information processing device equipped with an emotion analysis device that analyzes facial emotions and provides feedback according to the emotional state,

[0935] A system that includes this.

[0936] (Claim 2)

[0937] The system according to claim 1, further comprising a function to provide real-time feedback to various presentation devices based on information obtained from physical analysis and emotional analysis.

[0938] (Claim 3)

[0939] The system according to claim 1, further comprising a function to compare a user's form against comparison criteria and a function to provide support for technical improvement. [Explanation of Symbols]

[0940] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal means equipped with a camera for capturing the user's performance movements, A server means equipped with an analysis device that analyzes acquired video data of performance movements and extracts the user's skeletal information, A server means equipped with a control device that evaluates the user's playing form based on extracted skeletal information and generates a practice menu according to the evaluation results, A terminal device that presents the generated practice menu as a user interface and provides a learning environment that incorporates game elements, A server means equipped with a communication device that provides online instruction and feedback to users as they progress through their musical activities, A system that includes this.

2. The system according to claim 1, further comprising a function for providing performance support using sign language recognition technology.

3. The system according to claim 1, further comprising a function to compare the user's playing form with a model form and a function to provide guidance for improving playing technique.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A