System

The system addresses the challenges of learning musical instruments by recording and analyzing performance data to provide personalized feedback and practice modes, enhancing skill development and enjoyment.

JP2026021018APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122700
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Beginners face difficulties in learning musical instruments due to challenges with tuning, holding chords, and rhythm, and lack of professional instruction, leading to reduced enjoyment and efficiency in the learning process.

Method used

A system that allows users to record their musical instrument performance, analyze the video data using an analysis engine to extract finger movements, chord progressions, and rhythm, and provide personalized feedback through a generative AI model, enabling practice modes and collaborative videos to enhance learning.

Benefits of technology

Enables efficient skill improvement and maintains motivation by providing targeted feedback and practice opportunities, allowing users to progress effectively and enjoy the learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021018000001_ABST
    Figure 2026021018000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to record a performance of a musical instrument; means for the user to upload the recorded performance video; means for a server to receive the uploaded video data; means for the server to transmit the video data to an analysis engine and to receive analysis results; means for the server to generate feedback based on the analysis results; and means for the server to provide the generated feedback to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Many beginners to musical instruments often experience difficulties with tuning, holding chords, and rhythm during the self-learning process. This makes learning an instrument less enjoyable and often leads to abandonment. Furthermore, the lack of opportunities for professional instruction makes it difficult to progress efficiently. To address these issues, a system is needed to support beginners in effectively progressing with self-learning. [Means for solving the problem]

[0005] To solve the above problems, the following means are provided. A system is constructed that includes a means for a user to record their musical instrument performance and a means for uploading the recorded performance video. A server is provided with means for receiving the uploaded video data, transmitting the video data to an analysis engine, and receiving the analysis results. The server further generates specific feedback based on the analysis results and provides the feedback to the user. The analysis engine extracts finger movements, chord progressions, and rhythm from the video data. The system also includes a means for the user to select a practice mode based on a specific playing style, and the server provides a practice menu based on the selected practice mode. This allows beginners to learn efficiently and improve their musical instrument playing skills.

[0006] "User" refers to an individual who uses the system, particularly one who has the purpose of learning to play a musical instrument.

[0007] "Musical instrument" refers to any device or equipment used to play music, and in this invention refers specifically to the guitar as an example.

[0008] "Recording means" refers to a device, such as a smartphone or tablet, that a user uses to record their performance in video format.

[0009] "Means for uploading" refers to the application or software function for transmitting the recorded performance video to a server via the Internet.

[0010] "Server" refers to the central system that receives, stores, analyzes and processes data uploaded by Users.

[0011] "Video data" refers to a digital video file of a performance recorded by a user.

[0012] An "analysis engine" refers to software or algorithms that analyze video data and extract information about musical elements and performance techniques.

[0013] "Analysis results" refers to data regarding finger movements, chord progressions, rhythm, etc. obtained after the analysis engine processes the video data.

[0014] "Feedback" refers to guidance and advice about the user's playing technique that is generated by the server based on the analysis results.

[0015] "Practice Mode" refers to a customized practice program selected by a user based on a particular playing style or technique.

[0016] "Practice menu" refers to a series of practice tasks and instructional content provided based on the practice mode. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention relates to an online learning support system that allows beginners to learn musical instruments efficiently. Below, we will explain the program processing of this system in natural language and provide a concrete example.

[0039] Program processing explanation

[0040] 1. A user records and uploads their guitar performance

[0041] Users can record their guitar playing using a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" app and upload the video to the server through the app's interface. Users can also add tags and notes as needed for easy management.

[0042] 2. The server receives and analyzes the video data

[0043] The server receives video data uploaded by users. After receiving the video data, it sends it to an analysis engine for analysis. The analysis engine extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0044] 3. Generate feedback on analysis results

[0045] The server receives the analysis results sent from the analysis engine and uses the generative AI model to generate feedback on the user's performance, including specific tuning adjustments, how to hold chords, how to use a pick, and how to keep rhythm.

[0046] 4. Provide feedback to users

[0047] The server then provides the generated feedback to the user in the form of video, image, text, or audio, which the user can review through the "Pick Perfect" app. Providing feedback in a visually and audibly easy-to-understand format allows users to easily understand specific areas for improvement and how to improve them.

[0048] 5. Select a specific practice mode and practice

[0049] Users can select practice modes based on specific guitarist styles, such as "Home Mode" or "High-Level Mode." Based on the selected mode, the server provides instructional videos and practice menus for specific playing techniques and musical phrases. This allows users to efficiently practice specific styles and techniques.

[0050] 6. Creating and sharing collaborative videos

[0051] When users record and upload their practice results, the server generates a video of the user's performance and a collaboration with the AI ​​guitarist and vocalist. The collaboration video can be viewed on the user's app screen and shared on social media. This allows users to see their own progress and increase their motivation by sharing it with other users and the community.

[0052] Specific examples

[0053] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord. Once they've finished recording, they launch the Pick Perfect app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord, and detects any rhythm or pitch errors as appropriate.

[0054] The server then generates detailed feedback based on the analysis, including specific advice such as, "Your index finger is not positioned correctly, so not all the strings are sounding correctly." This feedback is provided with video, illustrations, and audio commentary, making it easy for users to understand and continue practicing.

[0055] When a user selects "Handy Mode," a practice menu based on the playing style of a specific guitarist is provided. The server transmits instructional videos on the guitarist's unique techniques and rhythm patterns, allowing the user to practice accordingly.

[0056] Furthermore, when a user reaches a certain level of skill, they can create a video in which they appear alongside their AI avatar and share it on social media, allowing users to see their skills improving and keeping their motivation high.

[0057] By implementing the system described above, beginners can learn to play musical instruments efficiently and experience the joy of playing musical instruments.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded file.

[0061] Step 2:

[0062] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[0063] Step 3:

[0064] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, the server prepares to send the video file to the analysis engine.

[0065] Step 4:

[0066] The server sends the video file to an analysis engine, which then analyzes the video data and extracts information about performance techniques such as finger movements, chord progressions, rhythm, and tempo.

[0067] Step 5:

[0068] The analysis engine generates the analysis results and sends them back to the server, including the finger position information for each frame and the accuracy of the sound.

[0069] Step 6:

[0070] The server receives the analysis results and generates feedback using a generative AI model based on the analysis data. The feedback consists of specific performance improvement points and advice.

[0071] Step 7:

[0072] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device, allowing the user to receive feedback in a variety of formats.

[0073] Step 8:

[0074] Users can check the feedback through the app and practice again according to the advice. The feedback includes specific points and methods for improvement, making it easy for users to understand and put into practice.

[0075] Step 9:

[0076] When a user selects a particular practice mode, such as "Home Mode" or "High Mode," the server provides a practice menu based on the selected mode, including explanations of the playing style and techniques of a particular guitarist.

[0077] Step 10:

[0078] Once the user reaches a certain level of skill, they can record and upload the performance video again. The server receives the newly uploaded video and re-analyzes it using the analysis engine.

[0079] Step 11:

[0080] Based on the new analysis results, the server generates a video of the user performing with the AI ​​guitarist and vocalist. The video is provided in a format where the user's performance and the AI's performance are synchronized.

[0081] Step 12:

[0082] Users can view their collaborative videos and share them on social media and other platforms through the "Pick Perfect" app, allowing them to share their progress with others and gain additional motivation.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] Conventional musical instrument practice systems do not provide sufficient feedback for users to efficiently improve their playing skills. Furthermore, they lack features such as practice modes based on specific playing styles or collaborative video generation to enjoy the results of practice, making it difficult for users to maintain their motivation to learn.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes means for receiving uploaded video data, means for transmitting the video data to an analysis engine and receiving the analysis results, means for generating feedback using a generative AI model based on the analysis results, means for the user to check the feedback and select a practice mode based on a specific playing style, means for providing a practice menu based on the selected practice mode, and means for generating and providing a collaborative video based on the user's practice results, which enables users to efficiently improve their playing skills and maintain their motivation to continue learning.

[0088] The "recording means" is a device or software function that allows a user to record a video of a musical instrument performance.

[0089] The "uploading means" is a function for transmitting a performance video recorded by a user to a server via a network.

[0090] The "receiving means" is a function that allows the server to receive video data uploaded by users via the network.

[0091] An "analysis engine" is software or a device for extracting performance information such as finger movements, chord progressions, and rhythm from video data.

[0092] The "feedback generation means" is a function that uses a generative AI model based on the analysis results to generate advice and adjustment instructions for the user's performance.

[0093] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to automatically generate feedback on a user's performance.

[0094] The "playing style selection means" is a function that allows the user to select a practice mode based on a specific guitarist or playing style.

[0095] The "practice menu providing means" is a function that provides the user with instructional videos and practice menus relating to specific performance techniques and musical phrases based on the selected performance style.

[0096] The "co-starring video generation means" is a function that generates a co-starring video with an AI avatar guitarist or vocalist based on the user's practice results.

[0097] "SNS sharing means" is a function for sharing the generated collaboration video on social networking services, etc.

[0098] This invention relates to an online learning support system that allows beginners to efficiently learn musical instruments. The system allows users to record and upload their musical instrument performances. The server then analyzes the performance data using an analysis engine and provides feedback using a generative AI model. It also includes a practice mode selection function based on specific playing styles and a collaborative video generation function.

[0099] First, users can record their guitar playing using a mobile device such as a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" application and select the recorded video. This video file is then uploaded to the server through the application's interface. Users can also manage their videos by adding tags and notes.

[0100] The server receives video data uploaded by users. This data is temporarily saved and stored in a database for the next analysis step. The analysis engine uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0101] Once the analysis is complete, the server uses a generative AI model, such as GPT-3 or T5, to generate feedback based on the analysis results. This feedback may include specific tuning adjustments, how to hold chords, how to use a pick, and rhythm. The generated feedback is provided in various formats, including video, images, text, and audio. This makes it easier for users to understand the feedback visually and audibly, and identify specific areas for improvement.

[0102] The user can also select practice modes based on specific guitarist styles, such as "Hotei Mode" and "Takanaka Mode." In this case, the server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the user's selection. This allows the user to efficiently practice specific styles and techniques.

[0103] Furthermore, by recording the results of the user's practice and uploading them back to the server, the server generates a video of the user's performance and a collaboration video with the AI ​​guitarist or vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. This allows users to realize their own progress and increase their motivation by sharing it with other users and the community.

[0104] As a concrete example, consider the case where a user is practicing an "F chord." After recording themselves playing an F chord using their smartphone, they launch the "Pick Perfect" app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord to detect errors in rhythm and pitch. Next, the server provides prompts to the generative AI model based on the analysis results, generating detailed feedback. For example, this could include specific advice such as, "Your index finger is not positioned correctly, so not all of the strings are sounding correctly."

[0105] When a user selects "Handy Mode," the server provides instructional videos on the guitarist's unique techniques and rhythm patterns. Furthermore, once the user reaches a certain level of skill, the server generates a video of the user performing with their AI avatar, which can then be shared on social media.

[0106] An example prompt for using a generative AI model is:

[0107] "Generate feedback on how to play the F chord in this video."

[0108] "Please analyze whether the user is playing with accurate rhythm."

[0109] "Create a practice routine based on a specific guitarist's style."

[0110] In this way, the present invention is a system that enables users to efficiently improve their playing skills and maintain a desire to continue learning.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1: Record and upload

[0113] Users use a smartphone or tablet to record their guitar playing. Once recording is complete, they launch the "instrument learning app" and select the recorded video file. Then, they upload it to the server through the application's interface. Users can add tags and notes to the video as needed. The input is the user's performance video, and the output is the video data uploaded to the server.

[0114] Specific behavior:

[0115] Record your guitar playing using the camera app on your smartphone.

[0116] Launch the "Instrument Learning App" and tap the "Upload Video" button.

[0117] Select the recording file and tap the "Upload" button.

[0118] Step 2: Receiving video data

[0119] The server receives video data uploaded by users. This data is temporarily stored and stored in a database for the next analysis step. The input is the video data sent by the user, and the output is the video data stored on the server.

[0120] Specific behavior:

[0121] The server receives the HTTP request and stores the video file.

[0122] Store the video file path and metadata in the database.

[0123] Step 3: Analyze the video data

[0124] The server sends the stored video data to an analysis engine, which uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data. The input is the video data stored on the server, and the output is the analyzed performance data.

[0125] Specific behavior:

[0126] The server calls the analysis engine API and sends the path to the video file.

[0127] An analysis engine processes the data and extracts information such as finger position and chord progression.

[0128] The extraction results are sent back to the server.

[0129] Step 4: Generate feedback from the analysis results

[0130] The server generates feedback using a generative AI model based on the analysis results sent from the analysis engine. For example, it uses GPT-3 or T5 to generate feedback on the user's performance, such as specific tuning adjustments, how to hold chords, and how to keep rhythm. The input is the analyzed performance data, and the output is the generated feedback.

[0131] Specific behavior:

[0132] The server inputs the analysis results into the generative AI model and provides prompt sentences.

[0133] The generative AI model generates feedback and sends it back to the server.

[0134] Step 5: Provide feedback

[0135] The server provides the generated feedback to the user. This feedback can be in the form of video, image, text, or audio, and the user can check the feedback through the "instrument learning app." The input is the generated feedback, and the output is the feedback provided to the user.

[0136] Specific behavior:

[0137] The server converts the generated feedback into an appropriate format.

[0138] The feedback will be associated with the user's account and made available for viewing in the "instrument learning app."

[0139] Step 6: Select Practice Mode and Practice

[0140] The user selects a practice mode based on a specific playing style, such as "Guitarist A Mode" or "Guitarist B Mode," within the app. The server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the selected mode. The input is the user's practice mode selection, and the output is a specific practice menu or instructional video.

[0141] Specific behavior:

[0142] The user selects practice mode in the "instrument learning app."

[0143] The server provides the user with content based on the selected mode.

[0144] Step 7: Create and share a collaborative video

[0145] The user records the results of their practice and uploads them to the server, which then generates a video of the user's performance and a collaboration video with the AI ​​guitarist and vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. The input is the user's new performance video, and the output is the created collaboration video.

[0146] Specific behavior:

[0147] The user then records their performance again and uploads the video.

[0148] The server combines the new video with footage of the AI ​​avatar to create a collaborative video.

[0149] The generated collaboration video is associated with the user's account and can be displayed within the app.

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] Traditional methods for learning equipment operation and maintenance make it difficult to master skills efficiently and accurately, and there is a risk of reduced productivity and safety issues due to operational errors and improper use of tools. Furthermore, there is a lack of sharing functions that allow employees to realize their own improvement in their skills and increase their motivation.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes: a means for a user to record equipment operation; a means for the user to upload the recorded operation video; a means for the server to receive the uploaded video data; a means for the server to transmit the video data to an analysis engine and receive the analysis results; a means for the server to generate feedback based on the analysis results; a means for the server to provide the generated feedback to the user; a means for the user to select a specific training mode; a means for the server to provide a training menu based on the selected training mode; and a means for generating a video featuring the user's operation video and an AI model. This allows employees to efficiently and accurately learn operation and maintenance techniques, reducing operational errors and improper tool use. Furthermore, through the generated feedback and video, employees can realize their own improvement in their skills and share them with the factory community, thereby increasing their motivation.

[0155] A "means for recording device operation" is a device or method that a user uses to record the operation of a machine or device.

[0156] "Means for uploading recorded operation videos" refers to the device or method used by a user to transmit or transfer the operation videos recorded by the user to a server.

[0157] The "means for the server to receive uploaded video data" refers to a device or method for the server to store and check video data received via the network.

[0158] The "means for the server to transmit video data to the analysis engine and receive the analysis results" refers to a device or method for the server to transmit video data to an engine for analyzing the data and obtain the analysis results.

[0159] The "means for the server to generate feedback based on the analysis results" refers to a device or method that allows the server to automatically generate advice or evaluations for the user based on data obtained from the analysis engine.

[0160] "Means by which the server provides generated feedback to the user" refers to a device or method for transmitting and displaying the generated feedback to the user.

[0161] A "means for user selection of a particular training mode" is a device or method by which a user selects a pre-defined training or practice mode.

[0162] The "means for the server to provide a training menu based on the selected training mode" refers to a device or method for providing specific training or practice content in accordance with the selected training mode.

[0163] "Means for generating a video of a user's operation and an AI model together" refers to a device or method for generating a new video of a user's operation and an AI model together by combining the video of the user's operation recorded with the video of the virtual model generated by AI.

[0164] This invention provides an online system that supports efficient learning of equipment operation and maintenance techniques. Below, the program processing of this system will be explained in natural language, along with specific examples.

[0165] System Overview

[0166] The system starts with employees recording their equipment operations and uploading the video data to a server. The server receives the uploaded video data, sends it to an analysis engine, and receives the analysis results. Based on the analysis results, the server generates feedback and provides it to the user. Furthermore, users can receive special training menus by selecting specific training modes. A means is also provided to generate videos in which the user's operation video is combined with an AI model.

[0167] Hardware and software used

[0168] 1. Hardware

[0169] Smart glasses or smartphones: used by users to record device operations.

[0170] Server: Receives, stores, and analyzes video data, and generates and provides feedback.

[0171] Factory Robots: Used to perform specific tasks.

[0172] 2. Software

[0173] Python: The primary programming language in which the entire program runs.

[0174] OpenCV: A library for loading video data and splitting frames.

[0175] numpy: A library for splitting frames and processing data.

[0176] Dedicated analysis engine (AI backend): An engine that analyzes operation procedures, accuracy of actions, and tools used from video data.

[0177] Generative AI model: A model that generates feedback based on analysis results.

[0178] Data processing and calculation

[0179] 1. Loading and processing video data

[0180] The server receives video files uploaded by users and splits them into frames using OpenCV, which are then sent to the analysis engine using the numpy library.

[0181] 2. Analysis and feedback generation

[0182] The analytics engine extracts information from each frame of the video, such as operational procedures, accuracy of movements, and tools used. Based on this, the server uses a generative AI model to generate feedback, including specific improvements, efficiency suggestions, and safety precautions.

[0183] 3. Providing Feedback

[0184] The generated feedback is provided in the form of video, images, text, and audio, and can be viewed by users through smart glasses or a smartphone application.

[0185] Specific examples

[0186] For example, an employee can record a machine maintenance operation and upload it to a server. This video is then analyzed by an analysis engine to identify problems with the operating procedures and how to properly use tools. The server then generates specific feedback based on the analysis results and provides it to the user. The user can then perform the operation again according to the feedback, thereby improving their skills.

[0187] Prompt Sentence Examples

[0188] Capture user operation videos and use the analytics engine to analyze operating procedures, tools used, and accuracy, generating feedback with specific suggestions for improvements and efficiencies, as well as safety precautions.

[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0190] Step 1:

[0191] The user records their device operations using smart glasses or a smartphone. The recorded video data is saved on the device. The input of this step is the user's operation behavior, and the output is the recorded video data.

[0192] Step 2:

[0193] The user launches the "Pick Perfect for Factory" app on their device, selects the recorded video data, and uploads it to the server. The video data is transferred from the user's device to the server. The input of this step is the recorded video data, and the output is the video data stored on the server.

[0194] Step 3:

[0195] The server receives the uploaded video data, divides it into frames using OpenCV and numpy, and sends them to the analysis engine. When the server divides the video data into frames, the input is the video data, and the output is image data for each frame. The analysis engine receives this and extracts information such as operation procedures, accuracy of actions, and tools used.

[0196] Step 4:

[0197] The server receives the analysis results sent by the analysis engine and uses the generative AI model to generate feedback including specific improvements, efficiency suggestions, and safety precautions. The input for this step is the analysis results, and the output is the generated feedback.

[0198] Step 5:

[0199] The server provides the generated feedback to the user. The feedback can be in the form of video, image, text, or audio. The user can view it through the app. The input of this step is the generated feedback, and the output is the feedback data that the user can view.

[0200] Step 6:

[0201] The user selects a specific training mode through the application and receives a training menu based on that mode. The server provides appropriate educational content based on the selected training mode. The input of this step is the user's selected training mode, and the output is the training menu and educational content.

[0202] Step 7:

[0203] The user's operation video is recorded again, and a video is generated that combines the re-uploaded video with the AI ​​model. The input for this step is the user's operation video and the AI ​​model, and the output is the combined video. The generated combined video can be viewed on the app and shared with the factory community.

[0204] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0205] This invention relates to an online learning support system that enables beginners to learn musical instruments efficiently and independently, and further enhances the user's learning experience by incorporating an emotion engine. Below, the program processing of this system is explained in natural language, and specific examples are given.

[0206] Program processing explanation

[0207] 1. Users record and upload their musical performances

[0208] Users use their smartphones or tablets to record their musical instrument performances. After recording is complete, they launch the "Pick Perfect" app, select the recorded file, and upload it to the server. At this time, users can enter tags and notes to make it easier to manage the recorded data.

[0209] 2. The server receives and analyzes the video data

[0210] The server receives the uploaded video data and stores it in temporary storage. After storage is complete, the video data is sent to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[0211] 3. The server uses the emotion engine to analyze the user's emotions.

[0212] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect their emotional state. The emotion engine identifies different emotions such as joy, sadness, surprise, and concentration.

[0213] 4. Generate feedback from analysis results and emotion data

[0214] The server combines the performance data received from the analysis engine and the emotional data received from the emotion engine, and generates feedback using a generative AI model. Specific feedback includes advice on improving performance technique and motivational messages.

[0215] 5. Provide feedback to users

[0216] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device. The user can then review the feedback through the "Pick Perfect" app and practice again according to the advice.

[0217] 6. Select a specific practice mode and practice

[0218] When a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style, including instructional videos on specific playing techniques and musical phrases.

[0219] 7. Creating and sharing collaborative videos

[0220] Once the user reaches a certain level of skill, they can record and upload the performance video again for re-analysis. The server then re-integrates the analysis results with the emotional data and generates a video of the user performing with their AI avatar, the guitarist and vocalist. The generated video can be viewed by the user on the app screen and shared on social media.

[0221] Specific examples

[0222] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord, and then upload the recording to the server via the "Pick Perfect" app. The server receives the video, and the analysis engine analyzes the intonation and sound accuracy, and then the emotion engine detects the user's facial expressions and tone of voice.

[0223] The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going." This not only helps users improve their playing technique, but also provides emotional support.

[0224] Furthermore, if the user selects "Home Mode," a practice menu tailored to the playing style of a specific guitarist is provided. Based on this practice mode, the server provides instructional videos on specific techniques and phrases.

[0225] Once a user reaches a certain level of skill, they can record and upload another video of their performance. The server then reanalyzes the newly received video and, based on the analysis results, generates a video of the user performing with their AI guitarist and vocalist. The generated video can be viewed by the user on the app screen and easily shared on social media and other platforms.

[0226] The above system further improves the user's learning experience, allowing them to continue practicing their instrument while improving their performance skills and maintaining their motivation.

[0227] The processing flow will be explained below.

[0228] Step 1:

[0229] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded video file.

[0230] Step 2:

[0231] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[0232] Step 3:

[0233] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, it automatically prepares to start the next analysis process.

[0234] Step 4:

[0235] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0236] Step 5:

[0237] The analysis engine generates analysis results and sends them to the server, which include detailed performance data for each frame.

[0238] Step 6:

[0239] The server receives the analysis results and sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[0240] Step 7:

[0241] The emotion engine generates emotion data and sends the data to the server, which includes different emotions such as joy, sadness, surprise, and concentration.

[0242] Step 8:

[0243] The server combines the analysis results with the emotional data and uses a generative AI model to generate feedback, including advice on improving performance technique and motivational messages.

[0244] Step 9:

[0245] The server sends the generated feedback to the user's device, which can be in the form of video, images, text, or audio.

[0246] Step 10:

[0247] The user reviews the feedback through the app, follows the advice on specific areas for improvement and how to improve, and practices again.

[0248] Step 11:

[0249] When a user selects a specific practice mode, the app configures the settings and sends the information to the server, which then provides a dedicated practice menu based on the selected practice mode.

[0250] Step 12:

[0251] The user practices according to a specific practice menu, then records the performance video again and uploads it to the server via the app.

[0252] Step 13:

[0253] The server receives the newly uploaded video and analyzes it again using the analysis engine and emotion engine. Based on the results, a video of the user performing together with an AI guitarist and vocalist is generated.

[0254] Step 14:

[0255] The server generates the video and sends it to the user's device, where the user can view it on the app screen and share it on social media or other platforms.

[0256] These are the specific processing steps of the present invention. This system allows users to practice while receiving appropriate feedback, and by receiving emotional support as well as improving their performance skills, they can maintain continuous motivation.

[0257] Example 2

[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0259] Existing instrument learning systems make it difficult for users to effectively improve their playing skills and maintain their motivation. Specifically, feedback is limited to improving playing technique, and comprehensive support that takes into account the user's emotional state and motivation is lacking. Furthermore, they lack the provision of personalized practice menus based on specific playing styles, making it difficult to progress efficiently.

[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0261] In this invention, the server includes means for an analysis engine to extract finger movements, chord progressions, and rhythms from video data, means for an emotion analysis engine to analyze facial expressions and tone of voice to detect the user's emotional state, and means for generating feedback using a generative AI model based on the analysis results and emotion data. This allows the user to receive advice that takes into account their emotional state along with technical feedback, allowing them to efficiently progress through musical instrument learning while receiving comprehensive support.

[0262] "User" refers to an individual who utilizes the system to practice playing an instrument and receive feedback.

[0263] "Instrument" refers to a tool or device that produces music for a user to play.

[0264] "Recording" refers to recording the user's performance as video data.

[0265] "Upload" refers to the process by which a user sends recorded video data to a server.

[0266] "Server" refers to a computer system that receives video data from users, stores it, analyzes it, and generates feedback.

[0267] "Video data" refers to video files of musical instrument performances recorded by users.

[0268] An "analysis engine" refers to software or algorithms used to extract technical information about musical instrument performance (finger movements, chord progressions, rhythm, etc.) from video data.

[0269] An "emotion analysis engine" refers to software or algorithms that analyze a user's facial expressions and tone of voice from video data to detect their emotional state.

[0270] "Generative AI model" refers to an artificial intelligence model that generates feedback appropriate for the user based on data from the analysis engine and sentiment analysis engine.

[0271] "Feedback" refers to messages that are generated based on the analysis results, such as advice on improving the user's playing technique or messages that increase motivation.

[0272] "Terminal" refers to the device (smartphone, tablet, etc.) that a user uses to record and check feedback.

[0273] "Practice mode" refers to a practice menu provided by the server that the user selects based on a particular playing style.

[0274] "Practice Menu" refers to content that provides instructional videos and advice on specific performance techniques or musical phrases.

[0275] This invention relates to an online learning support system for enabling beginners to efficiently self-study musical instruments. The system includes: means for a user to record their musical instrument performance; means for the user to upload the recorded performance video; means for a server to receive the uploaded video data; means for the server to transmit the video data to an analysis engine and receive analysis results that extract finger movements, chord progressions, and rhythm; means for the server to transmit the video data to an emotion analysis engine and analyze the user's emotional state; means for the server to generate feedback using a generative AI model based on the analysis results and the emotion data; and means for the server to provide the generated feedback to the user's terminal.

[0276] The specific process of this system is as follows: The user uses a device such as a smartphone or tablet to record themselves playing an instrument. A regular camera app is used for this recording. After recording is complete, the user launches a dedicated application, selects the recorded file, and uploads it to the server. When uploading, tags and notes are entered to make the recorded data easier to manage. The server receives the uploaded video data and saves it in temporary storage. A cloud service (e.g., Amazon S3 or Google Cloud Storage) is often used as this temporary storage.

[0277] The server then sends the saved video data to an analysis engine, which uses TensorFlow, OpenCV, and other tools to extract performance information such as finger movements, chord progressions, and rhythm. The server also sends the video data to an emotion analysis engine, which analyzes the user's facial expressions and tone of voice. The emotion analysis engine uses tools such as Microsoft Azure Emotion API and IBM Watson Tone Analyzer. This identifies the user's emotional state, such as joy, sadness, surprise, or concentration.

[0278] The analysis results and emotional data are integrated on the server, and feedback is generated using a generative AI model (e.g., GPT-4). This feedback includes advice on improving performance technique and motivational messages. The generated feedback is prepared in various formats (video, image, text, audio) and sent to the user's device. The user can view this feedback through a dedicated application.

[0279] Furthermore, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style. The practice menu includes instructional videos on specific playing techniques and musical phrases. These videos are provided using the YouTube API and Vimeo API.

[0280] For example, if a user is practicing an F chord, they can record themselves playing the F chord using their smartphone and upload the recording to a server via a dedicated app. The server receives the video, analyzes the articulation and sound accuracy using an analysis engine (TensorFlow or OpenCV), and then detects the user's facial expression and tone of voice using an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer). The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going!"

[0281] Here are some example prompts for a generative AI model:

[0282] "I'm practicing an F chord. Could you please analyze the video I recorded? Give me advice on finger placement and chord progression accuracy. I'd also like some motivational messages."

[0283] "The user selected a practice mode based on a specific playing style. Please provide a practice program tailored to the style of the specific guitarist. Also provide instructional videos on specific techniques and phrases."

[0284] "The user has reached a certain level of skill. Please analyze the new performance video and generate a collaboration video. Please generate a collaboration video between the user and the AI ​​and provide it in a format that can be shared on social media, etc."

[0285] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0286] Step 1:

[0287] A user records an instrument performance

[0288] Input: A user starts playing an instrument using a smartphone or tablet.

[0289] How it works: The user uses a regular camera app to record their performance as video data.

[0290] Output: Video data of the recorded instrument performance (video file)

[0291] Step 2:

[0292] Users upload recorded performance videos

[0293] Input: Recorded performance video (video file), launching the dedicated application

[0294] How it works: After recording is complete, the user launches the dedicated application, selects the recorded file, optionally enters tags and notes, and uploads the video to the server.

[0295] Output: Uploaded performance video (video file), management information (tags, notes)

[0296] Step 3:

[0297] The server receives the uploaded video data and stores it in temporary storage.

[0298] Input: Uploaded performance video (video file), management information (tags, notes)

[0299] How it works: The server receives video data uploaded by users and temporarily stores it in cloud storage (e.g., Amazon S3 or Google Cloud Storage).

[0300] Output: Video data saved in temporary storage

[0301] Step 4:

[0302] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[0303] Input: Saved video data

[0304] Operation: The server sends the video data to an analysis engine (TensorFlow or OpenCV), which analyzes the data to extract performance information such as finger movements, chord progressions, and rhythm.

[0305] Output: Performance analysis results (data on finger movements, chord progressions, rhythm, etc.)

[0306] Step 5:

[0307] The server sends the video data to the emotion analysis engine to analyze the user's emotional state.

[0308] Input: Saved video data

[0309] How it works: The server sends video data to an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer), which analyzes the user's facial expressions and tone of voice to identify their emotional state (happiness, sadness, surprise, concentration, etc.).

[0310] Output: Sentiment analysis results (data about the user's emotional state)

[0311] Step 6:

[0312] The server generates feedback using a generative AI model based on the analysis results and emotion data.

[0313] Input: Performance analysis results, emotion analysis results

[0314] How it works: The server integrates the analysis results with the emotional data and uses a generative AI model (e.g., GPT-4) to generate feedback for the user (such as advice on improving playing skills or messages to motivate them).

[0315] Output: Generated feedback (technical advice, motivational messages)

[0316] Step 7:

[0317] The server provides the generated feedback to the user's device.

[0318] Input: Generated feedback (technical advice, motivational messages)

[0319] How it works: The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device.

[0320] Output: Feedback is displayed in the user's dedicated application (technical advice, motivational messages)

[0321] Step 8:

[0322] The user selects a practice mode based on a specific playing style and receives a practice menu.

[0323] Input: The user selects a specific playing style in a dedicated application.

[0324] How it works: When a user selects a practice mode based on a specific playing style, the server provides instructional videos on specific playing techniques or musical phrases based on that selection. This is done using the YouTube API and Vimeo API.

[0325] Output: Practice menu (explanatory videos, technique information)

[0326] Step 9:

[0327] When a user reaches a certain level of skill, they can upload a new performance video.

[0328] Input: A device for users to re-record and upload performance videos

[0329] What it does: When a user reaches a certain skill level, it records a new performance video and uploads it to the server.

[0330] Output: Newly uploaded performance video (video file)

[0331] Step 10:

[0332] The server reanalyzes the newly received video and generates a co-video

[0333] Input: Newly uploaded performance video, analysis engine for reanalysis, and emotion analysis engine results

[0334] Operation: The server retransmits the newly received video to the analysis engine and emotion analysis engine, which reanalyzes the data. Then, based on the analysis results, it generates a video of the user performing with the AI ​​guitarist and vocalist.

[0335] Output: Collaborative video (a video of the user and their AI avatar)

[0336] Step 11:

[0337] The server provides the generated collaborative video to users, making it possible to share it.

[0338] Input: Generated collaboration video

[0339] Operation: The server sends the generated video to the user's dedicated application, allowing the user to view the video and share it on social media or other platforms.

[0340] Output: Shareable collaboration video (can be shared on social media)

[0341] (Application example 2)

[0342] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0343] Beginners learning musical instruments require accurate understanding of their progress and effective feedback when self-learning. However, conventional online learning systems often provide feedback based solely on the user's technical performance data, and lack support that takes into account the user's emotional state and motivation. The present invention aims to provide an online learning support system that enables beginners to efficiently self-learn musical instruments, improving their skills and maintaining their motivation at the same time.

[0344] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for users to upload recorded performance videos, means for analyzing the user's emotional state using an emotion engine, and means for generating feedback by integrating the analysis results and emotion data based on a generative AI model. This makes it possible to provide not only feedback based on the analysis results of the user's technical performance, but also comprehensive feedback that takes the user's emotional state into consideration in real time.

[0345] A "user" is an individual who is learning to play a musical instrument and who uses the online learning support system.

[0346] "Musical instrument" refers to tools used to play music, especially stringed or keyboard instruments such as guitars, pianos, and violins.

[0347] "Recording" refers to the act of saving a user's performance in video format on a digital device.

[0348] A "performance video" is a video file that records a user playing a musical instrument.

[0349] "Uploading" refers to the act of a user transmitting a recorded performance video to a server via the Internet.

[0350] A "server" is a computer system that receives, analyzes, and processes uploaded data.

[0351] "Video data" refers to digital data of a recorded video of a performance.

[0352] An "analysis engine" is software that analyzes video data and extracts technical information such as finger movements, chord progressions, and rhythm.

[0353] "Analysis results" refers to the technical information extracted from the video data by the analysis engine.

[0354] The "emotion engine" is software that analyzes the user's facial expressions and tone of voice to identify emotions such as joy, sadness, surprise, and concentration.

[0355] "Emotion data" is information about the user's emotional state analyzed by the emotion engine.

[0356] A "generative AI model" is an algorithm that uses generative adversarial networks (GANs) and other AI techniques to integrate analytical results with emotional data and generate feedback.

[0357] "Feedback" refers to advice and messages generated based on analysis results and emotional data that help users improve their performance skills.

[0358] A "wearable device" is a computing device that can be worn by a user, such as smart glasses or a head-mounted display.

[0359] "Real-time" refers to providing instant feedback at the moment the performance is taking place.

[0360] "Practice Mode" is a personalized practice program that a user selects based on a particular playing style or technique.

[0361] A "practice menu" is a series of practice exercises and instructional videos provided by the server according to the selected practice mode.

[0362] The present invention is an online learning support system for helping beginners to learn musical instruments efficiently. The system provides a means for users to record their musical instrument performances and upload the recorded performance videos. The uploaded video data is received by a server, where it is then processed by an analysis engine and an emotion engine. The analysis engine extracts information such as finger movements, chord progressions, and rhythm from the performance data, while the emotion engine analyzes the user's facial expressions and vocal tone to identify their emotional state.

[0363] The server integrates the performance data received from the analysis engine and the emotion data received from the emotion engine, and generates feedback using a generative AI model. This feedback includes advice on improving performance technique and messages to motivate the user.

[0364] Users can receive real-time feedback using wearable devices (e.g., smart glasses or head-mounted displays). This allows them to get feedback at the moment they are playing and immediately try to improve their technique. In addition, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on the user's selection. This practice menu includes instructional videos on specific techniques or musical phrases.

[0365] Furthermore, once the user reaches a certain level of skill, they can record another performance video, upload it to the server, and reanalyze it. The server then reintegrates the new analysis results with the emotional data to generate a video of the user performing with the virtual performer. This video can be viewed on the user's device screen and easily shared on social media.

[0366] Examples:

[0367] For example, let's say a user is practicing an "F chord." First, they use the smart glasses to record themselves playing the F chord, and then upload the recording to a server via an application. The server receives the video data, analyzes the finger movements and rhythm using an analysis engine, and then analyzes the user's facial expressions and tone of voice using an emotion engine to detect their emotional state. The analysis results include technical feedback such as "Your finger position is not accurate, so the notes are out of sync," as well as encouraging feedback such as "You're making a good learning pace, so keep going."

[0368] When a user selects the "Specific Guitarist Mode" as their practice mode, the server generates a practice menu based on this mode. The practice menu includes instructional videos on specific techniques and phrases. Once the user reaches a certain level, they record and upload the performance video again for further analysis. Based on the obtained data, the server generates a video of the user performing with a virtual musician and provides it to the user in a shareable format.

[0369] Example prompt sentence:

[0370] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[0371] This provides a system that allows users to improve their performance skills while continuing to learn while receiving emotional support.

[0372] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0373] (Processing step flow)

[0374] Step 1:

[0375] A user records their performance using a wearable device (e.g., smart glasses or a head-mounted display). The recorded video data is stored on the device.

[0376] Input: User's performance video

[0377] Output: Recorded video data

[0378] Specific behavior: The user puts on the device and starts playing, presses the record button to record the performance, and presses the button again to stop recording when finished.

[0379] Step 2:

[0380] The user uploads the recorded performance video to the server using the application. The user selects the video file in the application and presses the upload button.

[0381] Input: Recorded video data

[0382] Output: Video data uploaded to the server

[0383] Specific operation: Open the application, select the recording file, and press the upload button. The video data will be sent to the server via the Internet.

[0384] Step 3:

[0385] The server receives the uploaded video data and stores it in temporary storage.

[0386] Input: Uploaded video data

[0387] Output: Video data stored in temporary storage

[0388] Specific operation: The server receives data sent via the Internet and saves it in the specified directory.

[0389] Step 4:

[0390] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm. The analysis engine then generates the extracted technical information.

[0391] Input: Video data stored in temporary storage

[0392] Output: Technical analysis results

[0393] How it works: The server passes the video data to the analysis engine, which then uses image recognition algorithms to extract performance technique information, such as finger position and movement.

[0394] Step 5:

[0395] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[0396] Input: Video data stored in temporary storage

[0397] Output: Emotion analysis results

[0398] How it works: The server passes video data to the emotion engine, which then uses emotion recognition algorithms to analyze the user's facial expressions and tone of voice.

[0399] Step 6:

[0400] The server integrates the technical analysis results from the analysis engine and the emotional analysis results from the emotion engine and generates feedback using a generative AI model.

[0401] Input: Technical analysis results, sentiment analysis results

[0402] Output: Feedback from a generative AI model

[0403] Specific operation: The server inputs the two analysis results using a generative AI model such as Python or TensorFlow and generates an appropriate feedback message.

[0404] Step 7:

[0405] The server sends the generated feedback to the user's terminal, where it is displayed in real time on the user's wearable device.

[0406] Input: Feedback from a generative AI model

[0407] Output: Feedback provided to the user

[0408] Specific operation: The server sends a feedback message in text, audio, or image format to the user's device and displays it on the device.

[0409] (Example prompt)

[0410] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[0411] This allows users to receive real-time feedback on areas to improve their playing technique and maintain their motivation, enabling effective self-learning.

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0413] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0415] [Second embodiment]

[0416] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0417] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0418] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0420] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0422] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0423] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0424] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0427] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0428] The present invention relates to an online learning support system that allows beginners to learn musical instruments efficiently. Below, we will explain the program processing of this system in natural language and provide a concrete example.

[0429] Program processing explanation

[0430] 1. A user records and uploads their guitar performance

[0431] Users can record their guitar playing using a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" app and upload the video to the server through the app's interface. Users can also add tags and notes as needed for easy management.

[0432] 2. The server receives and analyzes the video data

[0433] The server receives video data uploaded by users. After receiving the video data, it sends it to an analysis engine for analysis. The analysis engine extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0434] 3. Generate feedback on analysis results

[0435] The server receives the analysis results sent from the analysis engine and uses the generative AI model to generate feedback on the user's performance, including specific tuning adjustments, how to hold chords, how to use a pick, and how to keep rhythm.

[0436] 4. Provide feedback to users

[0437] The server then provides the generated feedback to the user in the form of video, image, text, or audio, which the user can review through the "Pick Perfect" app. Providing feedback in a visually and audibly easy-to-understand format allows users to easily understand specific areas for improvement and how to improve them.

[0438] 5. Select a specific practice mode and practice

[0439] Users can select practice modes based on specific guitarist styles, such as "Home Mode" or "High-Level Mode." Based on the selected mode, the server provides instructional videos and practice menus for specific playing techniques and musical phrases. This allows users to efficiently practice specific styles and techniques.

[0440] 6. Creating and sharing collaborative videos

[0441] When users record and upload their practice results, the server generates a video of the user's performance and a collaboration with the AI ​​guitarist and vocalist. The collaboration video can be viewed on the user's app screen and shared on social media. This allows users to see their own progress and increase their motivation by sharing it with other users and the community.

[0442] Specific examples

[0443] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord. Once they've finished recording, they launch the Pick Perfect app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord, and detects any rhythm or pitch errors as appropriate.

[0444] The server then generates detailed feedback based on the analysis, including specific advice such as, "Your index finger is not positioned correctly, so not all the strings are sounding correctly." This feedback is provided with video, illustrations, and audio commentary, making it easy for users to understand and continue practicing.

[0445] When a user selects "Handy Mode," a practice menu based on the playing style of a specific guitarist is provided. The server transmits instructional videos on the guitarist's unique techniques and rhythm patterns, allowing the user to practice accordingly.

[0446] Furthermore, when a user reaches a certain level of skill, they can create a video in which they appear alongside their AI avatar and share it on social media, allowing users to see their skills improving and keeping their motivation high.

[0447] By implementing the system described above, beginners can learn to play musical instruments efficiently and experience the joy of playing musical instruments.

[0448] The processing flow will be explained below.

[0449] Step 1:

[0450] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded file.

[0451] Step 2:

[0452] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[0453] Step 3:

[0454] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, the server prepares to send the video file to the analysis engine.

[0455] Step 4:

[0456] The server sends the video file to an analysis engine, which then analyzes the video data and extracts information about performance techniques such as finger movements, chord progressions, rhythm, and tempo.

[0457] Step 5:

[0458] The analysis engine generates the analysis results and sends them back to the server, including the finger position information for each frame and the accuracy of the sound.

[0459] Step 6:

[0460] The server receives the analysis results and generates feedback using a generative AI model based on the analysis data. The feedback consists of specific performance improvement points and advice.

[0461] Step 7:

[0462] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device, allowing the user to receive feedback in a variety of formats.

[0463] Step 8:

[0464] Users can check the feedback through the app and practice again according to the advice. The feedback includes specific points and methods for improvement, making it easy for users to understand and put into practice.

[0465] Step 9:

[0466] When a user selects a particular practice mode, such as "Home Mode" or "High Mode," the server provides a practice menu based on the selected mode, including explanations of the playing style and techniques of a particular guitarist.

[0467] Step 10:

[0468] Once the user reaches a certain level of skill, they can record and upload the performance video again. The server receives the newly uploaded video and re-analyzes it using the analysis engine.

[0469] Step 11:

[0470] Based on the new analysis results, the server generates a video of the user performing with the AI ​​guitarist and vocalist. The video is provided in a format where the user's performance and the AI's performance are synchronized.

[0471] Step 12:

[0472] Users can view their collaborative videos and share them on social media and other platforms through the "Pick Perfect" app, allowing them to share their progress with others and gain additional motivation.

[0473] Example 1

[0474] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0475] Conventional musical instrument practice systems do not provide sufficient feedback for users to efficiently improve their playing skills. Furthermore, they lack features such as practice modes based on specific playing styles or collaborative video generation to enjoy the results of practice, making it difficult for users to maintain their motivation to learn.

[0476] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0477] In this invention, the server includes means for receiving uploaded video data, means for transmitting the video data to an analysis engine and receiving the analysis results, means for generating feedback using a generative AI model based on the analysis results, means for the user to check the feedback and select a practice mode based on a specific playing style, means for providing a practice menu based on the selected practice mode, and means for generating and providing a collaborative video based on the user's practice results, which enables users to efficiently improve their playing skills and maintain their motivation to continue learning.

[0478] The "recording means" is a device or software function that allows a user to record a video of a musical instrument performance.

[0479] The "uploading means" is a function for transmitting a performance video recorded by a user to a server via a network.

[0480] The "receiving means" is a function that allows the server to receive video data uploaded by users via the network.

[0481] An "analysis engine" is software or a device for extracting performance information such as finger movements, chord progressions, and rhythm from video data.

[0482] The "feedback generation means" is a function that uses a generative AI model based on the analysis results to generate advice and adjustment instructions for the user's performance.

[0483] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to automatically generate feedback on a user's performance.

[0484] The "playing style selection means" is a function that allows the user to select a practice mode based on a specific guitarist or playing style.

[0485] The "practice menu providing means" is a function that provides the user with instructional videos and practice menus relating to specific performance techniques and musical phrases based on the selected performance style.

[0486] The "co-starring video generation means" is a function that generates a co-starring video with an AI avatar guitarist or vocalist based on the user's practice results.

[0487] "SNS sharing means" is a function for sharing the generated collaboration video on social networking services, etc.

[0488] This invention relates to an online learning support system that allows beginners to efficiently learn musical instruments. The system allows users to record and upload their musical instrument performances. The server then analyzes the performance data using an analysis engine and provides feedback using a generative AI model. It also includes a practice mode selection function based on specific playing styles and a collaborative video generation function.

[0489] First, users can record their guitar playing using a mobile device such as a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" application and select the recorded video. This video file is then uploaded to the server through the application's interface. Users can also manage their videos by adding tags and notes.

[0490] The server receives video data uploaded by users. This data is temporarily saved and stored in a database for the next analysis step. The analysis engine uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0491] Once the analysis is complete, the server uses a generative AI model, such as GPT-3 or T5, to generate feedback based on the analysis results. This feedback may include specific tuning adjustments, how to hold chords, how to use a pick, and rhythm. The generated feedback is provided in various formats, including video, images, text, and audio. This makes it easier for users to understand the feedback visually and audibly, and identify specific areas for improvement.

[0492] The user can also select practice modes based on specific guitarist styles, such as "Hotei Mode" and "Takanaka Mode." In this case, the server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the user's selection. This allows the user to efficiently practice specific styles and techniques.

[0493] Furthermore, by recording the results of the user's practice and uploading them back to the server, the server generates a video of the user's performance and a collaboration video with the AI ​​guitarist or vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. This allows users to realize their own progress and increase their motivation by sharing it with other users and the community.

[0494] As a concrete example, consider the case where a user is practicing an "F chord." After recording themselves playing an F chord using their smartphone, they launch the "Pick Perfect" app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord to detect errors in rhythm and pitch. Next, the server provides prompts to the generative AI model based on the analysis results, generating detailed feedback. For example, this could include specific advice such as, "Your index finger is not positioned correctly, so not all of the strings are sounding correctly."

[0495] When a user selects "Handy Mode," the server provides instructional videos on the guitarist's unique techniques and rhythm patterns. Furthermore, once the user reaches a certain level of skill, the server generates a video of the user performing with their AI avatar, which can then be shared on social media.

[0496] An example prompt for using a generative AI model is:

[0497] "Generate feedback on how to play the F chord in this video."

[0498] "Please analyze whether the user is playing with accurate rhythm."

[0499] "Create a practice routine based on a specific guitarist's style."

[0500] In this way, the present invention is a system that enables users to efficiently improve their playing skills and maintain a desire to continue learning.

[0501] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0502] Step 1: Record and upload

[0503] Users use a smartphone or tablet to record their guitar playing. Once recording is complete, they launch the "instrument learning app" and select the recorded video file. Then, they upload it to the server through the application's interface. Users can add tags and notes to the video as needed. The input is the user's performance video, and the output is the video data uploaded to the server.

[0504] Specific behavior:

[0505] Record your guitar playing using the camera app on your smartphone.

[0506] Launch the "Instrument Learning App" and tap the "Upload Video" button.

[0507] Select the recording file and tap the "Upload" button.

[0508] Step 2: Receiving video data

[0509] The server receives video data uploaded by users. This data is temporarily stored and stored in a database for the next analysis step. The input is the video data sent by the user, and the output is the video data stored on the server.

[0510] Specific behavior:

[0511] The server receives the HTTP request and stores the video file.

[0512] Store the video file path and metadata in the database.

[0513] Step 3: Analyze the video data

[0514] The server sends the stored video data to an analysis engine, which uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data. The input is the video data stored on the server, and the output is the analyzed performance data.

[0515] Specific behavior:

[0516] The server calls the analysis engine API and sends the path to the video file.

[0517] An analysis engine processes the data and extracts information such as finger position and chord progression.

[0518] The extraction results are sent back to the server.

[0519] Step 4: Generate feedback from the analysis results

[0520] The server generates feedback using a generative AI model based on the analysis results sent from the analysis engine. For example, it uses GPT-3 or T5 to generate feedback on the user's performance, such as specific tuning adjustments, how to hold chords, and how to keep rhythm. The input is the analyzed performance data, and the output is the generated feedback.

[0521] Specific behavior:

[0522] The server inputs the analysis results into the generative AI model and provides prompt sentences.

[0523] The generative AI model generates feedback and sends it back to the server.

[0524] Step 5: Provide feedback

[0525] The server provides the generated feedback to the user. This feedback can be in the form of video, image, text, or audio, and the user can check the feedback through the "instrument learning app." The input is the generated feedback, and the output is the feedback provided to the user.

[0526] Specific behavior:

[0527] The server converts the generated feedback into an appropriate format.

[0528] The feedback will be associated with the user's account and made available for viewing in the "instrument learning app."

[0529] Step 6: Select Practice Mode and Practice

[0530] The user selects a practice mode based on a specific playing style, such as "Guitarist A Mode" or "Guitarist B Mode," within the app. The server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the selected mode. The input is the user's practice mode selection, and the output is a specific practice menu or instructional video.

[0531] Specific behavior:

[0532] The user selects practice mode in the "instrument learning app."

[0533] The server provides the user with content based on the selected mode.

[0534] Step 7: Create and share a collaborative video

[0535] The user records the results of their practice and uploads them to the server, which then generates a video of the user's performance and a collaboration video with the AI ​​guitarist and vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. The input is the user's new performance video, and the output is the created collaboration video.

[0536] Specific behavior:

[0537] The user then records their performance again and uploads the video.

[0538] The server combines the new video with footage of the AI ​​avatar to create a collaborative video.

[0539] The generated collaboration video is associated with the user's account and can be displayed within the app.

[0540] (Application example 1)

[0541] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0542] Traditional methods for learning equipment operation and maintenance make it difficult to master skills efficiently and accurately, and there is a risk of reduced productivity and safety issues due to operational errors and improper use of tools. Furthermore, there is a lack of sharing functions that allow employees to realize their own improvement in their skills and increase their motivation.

[0543] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0544] In this invention, the server includes: a means for a user to record equipment operation; a means for the user to upload the recorded operation video; a means for the server to receive the uploaded video data; a means for the server to transmit the video data to an analysis engine and receive the analysis results; a means for the server to generate feedback based on the analysis results; a means for the server to provide the generated feedback to the user; a means for the user to select a specific training mode; a means for the server to provide a training menu based on the selected training mode; and a means for generating a video featuring the user's operation video and an AI model. This allows employees to efficiently and accurately learn operation and maintenance techniques, reducing operational errors and improper tool use. Furthermore, through the generated feedback and video, employees can realize their own improvement in their skills and share them with the factory community, thereby increasing their motivation.

[0545] A "means for recording device operation" is a device or method that a user uses to record the operation of a machine or device.

[0546] "Means for uploading recorded operation videos" refers to the device or method used by a user to transmit or transfer the operation videos recorded by the user to a server.

[0547] The "means for the server to receive uploaded video data" refers to a device or method for the server to store and check video data received via the network.

[0548] The "means for the server to transmit video data to the analysis engine and receive the analysis results" refers to a device or method for the server to transmit video data to an engine for analyzing the data and obtain the analysis results.

[0549] The "means for the server to generate feedback based on the analysis results" refers to a device or method that allows the server to automatically generate advice or evaluations for the user based on data obtained from the analysis engine.

[0550] "Means by which the server provides generated feedback to the user" refers to a device or method for transmitting and displaying the generated feedback to the user.

[0551] A "means for user selection of a particular training mode" is a device or method by which a user selects a pre-defined training or practice mode.

[0552] The "means for the server to provide a training menu based on the selected training mode" refers to a device or method for providing specific training or practice content in accordance with the selected training mode.

[0553] "Means for generating a video of a user's operation and an AI model together" refers to a device or method for generating a new video of a user's operation and an AI model together by combining the video of the user's operation recorded with the video of the virtual model generated by AI.

[0554] This invention provides an online system that supports efficient learning of equipment operation and maintenance techniques. Below, the program processing of this system will be explained in natural language, along with specific examples.

[0555] System Overview

[0556] The system starts with employees recording their equipment operations and uploading the video data to a server. The server receives the uploaded video data, sends it to an analysis engine, and receives the analysis results. Based on the analysis results, the server generates feedback and provides it to the user. Furthermore, users can receive special training menus by selecting specific training modes. A means is also provided to generate videos in which the user's operation video is combined with an AI model.

[0557] Hardware and software used

[0558] 1. Hardware

[0559] Smart glasses or smartphones: used by users to record device operations.

[0560] Server: Receives, stores, and analyzes video data, and generates and provides feedback.

[0561] Factory Robots: Used to perform specific tasks.

[0562] 2. Software

[0563] Python: The primary programming language in which the entire program runs.

[0564] OpenCV: A library for loading video data and splitting frames.

[0565] numpy: A library for splitting frames and processing data.

[0566] Dedicated analysis engine (AI backend): An engine that analyzes operation procedures, accuracy of actions, and tools used from video data.

[0567] Generative AI model: A model that generates feedback based on analysis results.

[0568] Data processing and calculation

[0569] 1. Loading and processing video data

[0570] The server receives video files uploaded by users and splits them into frames using OpenCV, which are then sent to the analysis engine using the numpy library.

[0571] 2. Analysis and feedback generation

[0572] The analytics engine extracts information from each frame of the video, such as operational procedures, accuracy of movements, and tools used. Based on this, the server uses a generative AI model to generate feedback, including specific improvements, efficiency suggestions, and safety precautions.

[0573] 3. Providing Feedback

[0574] The generated feedback is provided in the form of video, images, text, and audio, and can be viewed by users through smart glasses or a smartphone application.

[0575] Specific examples

[0576] For example, an employee can record a machine maintenance operation and upload it to a server. This video is then analyzed by an analysis engine to identify problems with the operating procedures and how to properly use tools. The server then generates specific feedback based on the analysis results and provides it to the user. The user can then perform the operation again according to the feedback, thereby improving their skills.

[0577] Prompt Sentence Examples

[0578] Capture user operation videos and use the analytics engine to analyze operating procedures, tools used, and accuracy, generating feedback with specific suggestions for improvements and efficiencies, as well as safety precautions.

[0579] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0580] Step 1:

[0581] The user records their device operations using smart glasses or a smartphone. The recorded video data is saved on the device. The input of this step is the user's operation behavior, and the output is the recorded video data.

[0582] Step 2:

[0583] The user launches the "Pick Perfect for Factory" app on their device, selects the recorded video data, and uploads it to the server. The video data is transferred from the user's device to the server. The input of this step is the recorded video data, and the output is the video data stored on the server.

[0584] Step 3:

[0585] The server receives the uploaded video data, divides it into frames using OpenCV and numpy, and sends them to the analysis engine. When the server divides the video data into frames, the input is the video data, and the output is image data for each frame. The analysis engine receives this and extracts information such as operation procedures, accuracy of actions, and tools used.

[0586] Step 4:

[0587] The server receives the analysis results sent by the analysis engine and uses the generative AI model to generate feedback including specific improvements, efficiency suggestions, and safety precautions. The input for this step is the analysis results, and the output is the generated feedback.

[0588] Step 5:

[0589] The server provides the generated feedback to the user. The feedback can be in the form of video, image, text, or audio. The user can view it through the app. The input of this step is the generated feedback, and the output is the feedback data that the user can view.

[0590] Step 6:

[0591] The user selects a specific training mode through the application and receives a training menu based on that mode. The server provides appropriate educational content based on the selected training mode. The input of this step is the user's selected training mode, and the output is the training menu and educational content.

[0592] Step 7:

[0593] The user's operation video is recorded again, and a video is generated that combines the re-uploaded video with the AI ​​model. The input for this step is the user's operation video and the AI ​​model, and the output is the combined video. The generated combined video can be viewed on the app and shared with the factory community.

[0594] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0595] This invention relates to an online learning support system that enables beginners to learn musical instruments efficiently and independently, and further enhances the user's learning experience by incorporating an emotion engine. Below, the program processing of this system is explained in natural language, and specific examples are given.

[0596] Program processing explanation

[0597] 1. Users record and upload their musical performances

[0598] Users use their smartphones or tablets to record their musical instrument performances. After recording is complete, they launch the "Pick Perfect" app, select the recorded file, and upload it to the server. At this time, users can enter tags and notes to make it easier to manage the recorded data.

[0599] 2. The server receives and analyzes the video data

[0600] The server receives the uploaded video data and stores it in temporary storage. After storage is complete, the video data is sent to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[0601] 3. The server uses the emotion engine to analyze the user's emotions.

[0602] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect their emotional state. The emotion engine identifies different emotions such as joy, sadness, surprise, and concentration.

[0603] 4. Generate feedback from analysis results and emotion data

[0604] The server combines the performance data received from the analysis engine and the emotional data received from the emotion engine, and generates feedback using a generative AI model. Specific feedback includes advice on improving performance technique and motivational messages.

[0605] 5. Provide feedback to users

[0606] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device. The user can then review the feedback through the "Pick Perfect" app and practice again according to the advice.

[0607] 6. Select a specific practice mode and practice

[0608] When a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style, including instructional videos on specific playing techniques and musical phrases.

[0609] 7. Creating and sharing collaborative videos

[0610] Once the user reaches a certain level of skill, they can record and upload the performance video again for re-analysis. The server then re-integrates the analysis results with the emotional data and generates a video of the user performing with their AI avatar, the guitarist and vocalist. The generated video can be viewed by the user on the app screen and shared on social media.

[0611] Specific examples

[0612] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord, and then upload the recording to the server via the "Pick Perfect" app. The server receives the video, and the analysis engine analyzes the intonation and sound accuracy, and then the emotion engine detects the user's facial expressions and tone of voice.

[0613] The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going." This not only helps users improve their playing technique, but also provides emotional support.

[0614] Furthermore, if the user selects "Home Mode," a practice menu tailored to the playing style of a specific guitarist is provided. Based on this practice mode, the server provides instructional videos on specific techniques and phrases.

[0615] Once a user reaches a certain level of skill, they can record and upload another video of their performance. The server then reanalyzes the newly received video and, based on the analysis results, generates a video of the user performing with their AI guitarist and vocalist. The generated video can be viewed by the user on the app screen and easily shared on social media and other platforms.

[0616] The above system further improves the user's learning experience, allowing them to continue practicing their instrument while improving their performance skills and maintaining their motivation.

[0617] The processing flow will be explained below.

[0618] Step 1:

[0619] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded video file.

[0620] Step 2:

[0621] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[0622] Step 3:

[0623] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, it automatically prepares to start the next analysis process.

[0624] Step 4:

[0625] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0626] Step 5:

[0627] The analysis engine generates analysis results and sends them to the server, which include detailed performance data for each frame.

[0628] Step 6:

[0629] The server receives the analysis results and sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[0630] Step 7:

[0631] The emotion engine generates emotion data and sends the data to the server, which includes different emotions such as joy, sadness, surprise, and concentration.

[0632] Step 8:

[0633] The server combines the analysis results with the emotional data and uses a generative AI model to generate feedback, including advice on improving performance technique and motivational messages.

[0634] Step 9:

[0635] The server sends the generated feedback to the user's device, which can be in the form of video, images, text, or audio.

[0636] Step 10:

[0637] The user reviews the feedback through the app, follows the advice on specific areas for improvement and how to improve, and practices again.

[0638] Step 11:

[0639] When a user selects a specific practice mode, the app configures the settings and sends the information to the server, which then provides a dedicated practice menu based on the selected practice mode.

[0640] Step 12:

[0641] The user practices according to a specific practice menu, then records the performance video again and uploads it to the server via the app.

[0642] Step 13:

[0643] The server receives the newly uploaded video and analyzes it again using the analysis engine and emotion engine. Based on the results, a video of the user performing together with an AI guitarist and vocalist is generated.

[0644] Step 14:

[0645] The server generates the video and sends it to the user's device, where the user can view it on the app screen and share it on social media or other platforms.

[0646] These are the specific processing steps of the present invention. This system allows users to practice while receiving appropriate feedback, and by receiving emotional support as well as improving their performance skills, they can maintain continuous motivation.

[0647] Example 2

[0648] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0649] Existing instrument learning systems make it difficult for users to effectively improve their playing skills and maintain their motivation. Specifically, feedback is limited to improving playing technique, and comprehensive support that takes into account the user's emotional state and motivation is lacking. Furthermore, they lack the provision of personalized practice menus based on specific playing styles, making it difficult to progress efficiently.

[0650] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0651] In this invention, the server includes means for an analysis engine to extract finger movements, chord progressions, and rhythms from video data, means for an emotion analysis engine to analyze facial expressions and tone of voice to detect the user's emotional state, and means for generating feedback using a generative AI model based on the analysis results and emotion data. This allows the user to receive advice that takes into account their emotional state along with technical feedback, allowing them to efficiently progress through musical instrument learning while receiving comprehensive support.

[0652] "User" refers to an individual who utilizes the system to practice playing an instrument and receive feedback.

[0653] "Instrument" refers to a tool or device that produces music for a user to play.

[0654] "Recording" refers to recording the user's performance as video data.

[0655] "Upload" refers to the process by which a user sends recorded video data to a server.

[0656] "Server" refers to a computer system that receives video data from users, stores it, analyzes it, and generates feedback.

[0657] "Video data" refers to video files of musical instrument performances recorded by users.

[0658] An "analysis engine" refers to software or algorithms used to extract technical information about musical instrument performance (finger movements, chord progressions, rhythm, etc.) from video data.

[0659] An "emotion analysis engine" refers to software or algorithms that analyze a user's facial expressions and tone of voice from video data to detect their emotional state.

[0660] "Generative AI model" refers to an artificial intelligence model that generates feedback appropriate for the user based on data from the analysis engine and sentiment analysis engine.

[0661] "Feedback" refers to messages that are generated based on the analysis results, such as advice on improving the user's playing technique or messages that increase motivation.

[0662] "Terminal" refers to the device (smartphone, tablet, etc.) that a user uses to record and check feedback.

[0663] "Practice mode" refers to a practice menu provided by the server that the user selects based on a particular playing style.

[0664] "Practice Menu" refers to content that provides instructional videos and advice on specific performance techniques or musical phrases.

[0665] This invention relates to an online learning support system for enabling beginners to efficiently self-study musical instruments. The system includes: means for a user to record their musical instrument performance; means for the user to upload the recorded performance video; means for a server to receive the uploaded video data; means for the server to transmit the video data to an analysis engine and receive analysis results that extract finger movements, chord progressions, and rhythm; means for the server to transmit the video data to an emotion analysis engine and analyze the user's emotional state; means for the server to generate feedback using a generative AI model based on the analysis results and the emotion data; and means for the server to provide the generated feedback to the user's terminal.

[0666] The specific process of this system is as follows: The user uses a device such as a smartphone or tablet to record themselves playing an instrument. A regular camera app is used for this recording. After recording is complete, the user launches a dedicated application, selects the recorded file, and uploads it to the server. When uploading, tags and notes are entered to make the recorded data easier to manage. The server receives the uploaded video data and saves it in temporary storage. A cloud service (e.g., Amazon S3 or Google Cloud Storage) is often used as this temporary storage.

[0667] The server then sends the saved video data to an analysis engine, which uses TensorFlow, OpenCV, and other tools to extract performance information such as finger movements, chord progressions, and rhythm. The server also sends the video data to an emotion analysis engine, which analyzes the user's facial expressions and tone of voice. The emotion analysis engine uses tools such as Microsoft Azure Emotion API and IBM Watson Tone Analyzer. This identifies the user's emotional state, such as joy, sadness, surprise, or concentration.

[0668] The analysis results and emotional data are integrated on the server, and feedback is generated using a generative AI model (e.g., GPT-4). This feedback includes advice on improving performance technique and motivational messages. The generated feedback is prepared in various formats (video, image, text, audio) and sent to the user's device. The user can view this feedback through a dedicated application.

[0669] Furthermore, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style. The practice menu includes instructional videos on specific playing techniques and musical phrases. These videos are provided using the YouTube API and Vimeo API.

[0670] For example, if a user is practicing an F chord, they can record themselves playing the F chord using their smartphone and upload the recording to a server via a dedicated app. The server receives the video, analyzes the articulation and sound accuracy using an analysis engine (TensorFlow or OpenCV), and then detects the user's facial expression and tone of voice using an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer). The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going!"

[0671] Here are some example prompts for a generative AI model:

[0672] "I'm practicing an F chord. Could you please analyze the video I recorded? Give me advice on finger placement and chord progression accuracy. I'd also like some motivational messages."

[0673] "The user selected a practice mode based on a specific playing style. Please provide a practice program tailored to the style of the specific guitarist. Also provide instructional videos on specific techniques and phrases."

[0674] "The user has reached a certain level of skill. Please analyze the new performance video and generate a collaboration video. Please generate a collaboration video between the user and the AI ​​and provide it in a format that can be shared on social media, etc."

[0675] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0676] Step 1:

[0677] A user records an instrument performance

[0678] Input: A user starts playing an instrument using a smartphone or tablet.

[0679] How it works: The user uses a regular camera app to record their performance as video data.

[0680] Output: Video data of the recorded instrument performance (video file)

[0681] Step 2:

[0682] Users upload recorded performance videos

[0683] Input: Recorded performance video (video file), launching the dedicated application

[0684] How it works: After recording is complete, the user launches the dedicated application, selects the recorded file, optionally enters tags and notes, and uploads the video to the server.

[0685] Output: Uploaded performance video (video file), management information (tags, notes)

[0686] Step 3:

[0687] The server receives the uploaded video data and stores it in temporary storage.

[0688] Input: Uploaded performance video (video file), management information (tags, notes)

[0689] How it works: The server receives video data uploaded by users and temporarily stores it in cloud storage (e.g., Amazon S3 or Google Cloud Storage).

[0690] Output: Video data saved in temporary storage

[0691] Step 4:

[0692] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[0693] Input: Saved video data

[0694] Operation: The server sends the video data to an analysis engine (TensorFlow or OpenCV), which analyzes the data to extract performance information such as finger movements, chord progressions, and rhythm.

[0695] Output: Performance analysis results (data on finger movements, chord progressions, rhythm, etc.)

[0696] Step 5:

[0697] The server sends the video data to the emotion analysis engine to analyze the user's emotional state.

[0698] Input: Saved video data

[0699] How it works: The server sends video data to an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer), which analyzes the user's facial expressions and tone of voice to identify their emotional state (happiness, sadness, surprise, concentration, etc.).

[0700] Output: Sentiment analysis results (data about the user's emotional state)

[0701] Step 6:

[0702] The server generates feedback using a generative AI model based on the analysis results and emotion data.

[0703] Input: Performance analysis results, emotion analysis results

[0704] How it works: The server integrates the analysis results with the emotional data and uses a generative AI model (e.g., GPT-4) to generate feedback for the user (such as advice on improving playing skills or messages to motivate them).

[0705] Output: Generated feedback (technical advice, motivational messages)

[0706] Step 7:

[0707] The server provides the generated feedback to the user's device.

[0708] Input: Generated feedback (technical advice, motivational messages)

[0709] How it works: The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device.

[0710] Output: Feedback is displayed in the user's dedicated application (technical advice, motivational messages)

[0711] Step 8:

[0712] The user selects a practice mode based on a specific playing style and receives a practice menu.

[0713] Input: The user selects a specific playing style in a dedicated application.

[0714] How it works: When a user selects a practice mode based on a specific playing style, the server provides instructional videos on specific playing techniques or musical phrases based on that selection. This is done using the YouTube API and Vimeo API.

[0715] Output: Practice menu (explanatory videos, technique information)

[0716] Step 9:

[0717] When a user reaches a certain level of skill, they can upload a new performance video.

[0718] Input: A device for users to re-record and upload performance videos

[0719] What it does: When a user reaches a certain skill level, it records a new performance video and uploads it to the server.

[0720] Output: Newly uploaded performance video (video file)

[0721] Step 10:

[0722] The server reanalyzes the newly received video and generates a co-video

[0723] Input: Newly uploaded performance video, analysis engine for reanalysis, and emotion analysis engine results

[0724] Operation: The server retransmits the newly received video to the analysis engine and emotion analysis engine, which reanalyzes the data. Then, based on the analysis results, it generates a video of the user performing with the AI ​​guitarist and vocalist.

[0725] Output: Collaborative video (a video of the user and their AI avatar)

[0726] Step 11:

[0727] The server provides the generated collaborative video to users, making it possible to share it.

[0728] Input: Generated collaboration video

[0729] Operation: The server sends the generated video to the user's dedicated application, allowing the user to view the video and share it on social media or other platforms.

[0730] Output: Shareable collaboration video (can be shared on social media)

[0731] (Application example 2)

[0732] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0733] Beginners learning musical instruments require accurate understanding of their progress and effective feedback when self-learning. However, conventional online learning systems often provide feedback based solely on the user's technical performance data, and lack support that takes into account the user's emotional state and motivation. The present invention aims to provide an online learning support system that enables beginners to efficiently self-learn musical instruments, improving their skills and maintaining their motivation at the same time.

[0734] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for users to upload recorded performance videos, means for analyzing the user's emotional state using an emotion engine, and means for generating feedback by integrating the analysis results and emotion data based on a generative AI model. This makes it possible to provide not only feedback based on the analysis results of the user's technical performance, but also comprehensive feedback that takes the user's emotional state into consideration in real time.

[0735] A "user" is an individual who is learning to play a musical instrument and who uses the online learning support system.

[0736] "Musical instrument" refers to tools used to play music, especially stringed or keyboard instruments such as guitars, pianos, and violins.

[0737] "Recording" refers to the act of saving a user's performance in video format on a digital device.

[0738] A "performance video" is a video file that records a user playing a musical instrument.

[0739] "Uploading" refers to the act of a user transmitting a recorded performance video to a server via the Internet.

[0740] A "server" is a computer system that receives, analyzes, and processes uploaded data.

[0741] "Video data" refers to digital data of a recorded video of a performance.

[0742] An "analysis engine" is software that analyzes video data and extracts technical information such as finger movements, chord progressions, and rhythm.

[0743] "Analysis results" refers to the technical information extracted from the video data by the analysis engine.

[0744] The "emotion engine" is software that analyzes the user's facial expressions and tone of voice to identify emotions such as joy, sadness, surprise, and concentration.

[0745] "Emotion data" is information about the user's emotional state analyzed by the emotion engine.

[0746] A "generative AI model" is an algorithm that uses generative adversarial networks (GANs) and other AI techniques to integrate analytical results with emotional data and generate feedback.

[0747] "Feedback" refers to advice and messages generated based on analysis results and emotional data that help users improve their performance skills.

[0748] A "wearable device" is a computing device that can be worn by a user, such as smart glasses or a head-mounted display.

[0749] "Real-time" refers to providing instant feedback at the moment the performance is taking place.

[0750] "Practice Mode" is a personalized practice program that a user selects based on a particular playing style or technique.

[0751] A "practice menu" is a series of practice exercises and instructional videos provided by the server according to the selected practice mode.

[0752] The present invention is an online learning support system for helping beginners to learn musical instruments efficiently. The system provides a means for users to record their musical instrument performances and upload the recorded performance videos. The uploaded video data is received by a server, where it is then processed by an analysis engine and an emotion engine. The analysis engine extracts information such as finger movements, chord progressions, and rhythm from the performance data, while the emotion engine analyzes the user's facial expressions and vocal tone to identify their emotional state.

[0753] The server integrates the performance data received from the analysis engine and the emotion data received from the emotion engine, and generates feedback using a generative AI model. This feedback includes advice on improving performance technique and messages to motivate the user.

[0754] Users can receive real-time feedback using wearable devices (e.g., smart glasses or head-mounted displays). This allows them to get feedback at the moment they are playing and immediately try to improve their technique. In addition, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on the user's selection. This practice menu includes instructional videos on specific techniques or musical phrases.

[0755] Furthermore, once the user reaches a certain level of skill, they can record another performance video, upload it to the server, and reanalyze it. The server then reintegrates the new analysis results with the emotional data to generate a video of the user performing with the virtual performer. This video can be viewed on the user's device screen and easily shared on social media.

[0756] Examples:

[0757] For example, let's say a user is practicing an "F chord." First, they use the smart glasses to record themselves playing the F chord, and then upload the recording to a server via an application. The server receives the video data, analyzes the finger movements and rhythm using an analysis engine, and then analyzes the user's facial expressions and tone of voice using an emotion engine to detect their emotional state. The analysis results include technical feedback such as "Your finger position is not accurate, so the notes are out of sync," as well as encouraging feedback such as "You're making a good learning pace, so keep going."

[0758] When a user selects the "Specific Guitarist Mode" as their practice mode, the server generates a practice menu based on this mode. The practice menu includes instructional videos on specific techniques and phrases. Once the user reaches a certain level, they record and upload the performance video again for further analysis. Based on the obtained data, the server generates a video of the user performing with a virtual musician and provides it to the user in a shareable format.

[0759] Example prompt sentence:

[0760] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[0761] This provides a system that allows users to improve their performance skills while continuing to learn while receiving emotional support.

[0762] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0763] (Processing step flow)

[0764] Step 1:

[0765] A user records their performance using a wearable device (e.g., smart glasses or a head-mounted display). The recorded video data is stored on the device.

[0766] Input: User's performance video

[0767] Output: Recorded video data

[0768] Specific behavior: The user puts on the device and starts playing, presses the record button to record the performance, and presses the button again to stop recording when finished.

[0769] Step 2:

[0770] The user uploads the recorded performance video to the server using the application. The user selects the video file in the application and presses the upload button.

[0771] Input: Recorded video data

[0772] Output: Video data uploaded to the server

[0773] Specific operation: Open the application, select the recording file, and press the upload button. The video data will be sent to the server via the Internet.

[0774] Step 3:

[0775] The server receives the uploaded video data and stores it in temporary storage.

[0776] Input: Uploaded video data

[0777] Output: Video data stored in temporary storage

[0778] Specific operation: The server receives data sent via the Internet and saves it in the specified directory.

[0779] Step 4:

[0780] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm. The analysis engine then generates the extracted technical information.

[0781] Input: Video data stored in temporary storage

[0782] Output: Technical analysis results

[0783] How it works: The server passes the video data to the analysis engine, which then uses image recognition algorithms to extract performance technique information, such as finger position and movement.

[0784] Step 5:

[0785] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[0786] Input: Video data stored in temporary storage

[0787] Output: Emotion analysis results

[0788] How it works: The server passes video data to the emotion engine, which then uses emotion recognition algorithms to analyze the user's facial expressions and tone of voice.

[0789] Step 6:

[0790] The server integrates the technical analysis results from the analysis engine and the emotional analysis results from the emotion engine and generates feedback using a generative AI model.

[0791] Input: Technical analysis results, sentiment analysis results

[0792] Output: Feedback from a generative AI model

[0793] Specific operation: The server inputs the two analysis results using a generative AI model such as Python or TensorFlow and generates an appropriate feedback message.

[0794] Step 7:

[0795] The server sends the generated feedback to the user's terminal, where it is displayed in real time on the user's wearable device.

[0796] Input: Feedback from a generative AI model

[0797] Output: Feedback provided to the user

[0798] Specific operation: The server sends a feedback message in text, audio, or image format to the user's device and displays it on the device.

[0799] (Example prompt)

[0800] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[0801] This allows users to receive real-time feedback on areas to improve their playing technique and maintain their motivation, enabling effective self-learning.

[0802] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0803] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0804] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0805] [Third embodiment]

[0806] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0807] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0808] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0809] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0810] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0811] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0812] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0813] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0814] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0815] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0816] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0817] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0818] The present invention relates to an online learning support system that allows beginners to learn musical instruments efficiently. Below, we will explain the program processing of this system in natural language and provide a concrete example.

[0819] Program processing explanation

[0820] 1. A user records and uploads their guitar performance

[0821] Users can record their guitar playing using a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" app and upload the video to the server through the app's interface. Users can also add tags and notes as needed for easy management.

[0822] 2. The server receives and analyzes the video data

[0823] The server receives video data uploaded by users. After receiving the video data, it sends it to an analysis engine for analysis. The analysis engine extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0824] 3. Generate feedback on analysis results

[0825] The server receives the analysis results sent from the analysis engine and uses the generative AI model to generate feedback on the user's performance, including specific tuning adjustments, how to hold chords, how to use a pick, and how to keep rhythm.

[0826] 4. Provide feedback to users

[0827] The server then provides the generated feedback to the user in the form of video, image, text, or audio, which the user can review through the "Pick Perfect" app. Providing feedback in a visually and audibly easy-to-understand format allows users to easily understand specific areas for improvement and how to improve them.

[0828] 5. Select a specific practice mode and practice

[0829] Users can select practice modes based on specific guitarist styles, such as "Home Mode" or "High-Level Mode." Based on the selected mode, the server provides instructional videos and practice menus for specific playing techniques and musical phrases. This allows users to efficiently practice specific styles and techniques.

[0830] 6. Creating and sharing collaborative videos

[0831] When users record and upload their practice results, the server generates a video of the user's performance and a collaboration with the AI ​​guitarist and vocalist. The collaboration video can be viewed on the user's app screen and shared on social media. This allows users to see their own progress and increase their motivation by sharing it with other users and the community.

[0832] Specific examples

[0833] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord. Once they've finished recording, they launch the Pick Perfect app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord, and detects any rhythm or pitch errors as appropriate.

[0834] The server then generates detailed feedback based on the analysis, including specific advice such as, "Your index finger is not positioned correctly, so not all the strings are sounding correctly." This feedback is provided with video, illustrations, and audio commentary, making it easy for users to understand and continue practicing.

[0835] When a user selects "Handy Mode," a practice menu based on the playing style of a specific guitarist is provided. The server transmits instructional videos on the guitarist's unique techniques and rhythm patterns, allowing the user to practice accordingly.

[0836] Furthermore, when a user reaches a certain level of skill, they can create a video in which they appear alongside their AI avatar and share it on social media, allowing users to see their skills improving and keeping their motivation high.

[0837] By implementing the system described above, beginners can learn to play musical instruments efficiently and experience the joy of playing musical instruments.

[0838] The processing flow will be explained below.

[0839] Step 1:

[0840] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded file.

[0841] Step 2:

[0842] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[0843] Step 3:

[0844] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, the server prepares to send the video file to the analysis engine.

[0845] Step 4:

[0846] The server sends the video file to an analysis engine, which then analyzes the video data and extracts information about performance techniques such as finger movements, chord progressions, rhythm, and tempo.

[0847] Step 5:

[0848] The analysis engine generates the analysis results and sends them back to the server, including the finger position information for each frame and the accuracy of the sound.

[0849] Step 6:

[0850] The server receives the analysis results and generates feedback using a generative AI model based on the analysis data. The feedback consists of specific performance improvement points and advice.

[0851] Step 7:

[0852] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device, allowing the user to receive feedback in a variety of formats.

[0853] Step 8:

[0854] Users can check the feedback through the app and practice again according to the advice. The feedback includes specific points and methods for improvement, making it easy for users to understand and put into practice.

[0855] Step 9:

[0856] When a user selects a particular practice mode, such as "Home Mode" or "High Mode," the server provides a practice menu based on the selected mode, including explanations of the playing style and techniques of a particular guitarist.

[0857] Step 10:

[0858] Once the user reaches a certain level of skill, they can record and upload the performance video again. The server receives the newly uploaded video and re-analyzes it using the analysis engine.

[0859] Step 11:

[0860] Based on the new analysis results, the server generates a video of the user performing with the AI ​​guitarist and vocalist. The video is provided in a format where the user's performance and the AI's performance are synchronized.

[0861] Step 12:

[0862] Users can view their collaborative videos and share them on social media and other platforms through the "Pick Perfect" app, allowing them to share their progress with others and gain additional motivation.

[0863] Example 1

[0864] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0865] Conventional musical instrument practice systems do not provide sufficient feedback for users to efficiently improve their playing skills. Furthermore, they lack features such as practice modes based on specific playing styles or collaborative video generation to enjoy the results of practice, making it difficult for users to maintain their motivation to learn.

[0866] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0867] In this invention, the server includes means for receiving uploaded video data, means for transmitting the video data to an analysis engine and receiving the analysis results, means for generating feedback using a generative AI model based on the analysis results, means for the user to check the feedback and select a practice mode based on a specific playing style, means for providing a practice menu based on the selected practice mode, and means for generating and providing a collaborative video based on the user's practice results, which enables users to efficiently improve their playing skills and maintain their motivation to continue learning.

[0868] The "recording means" is a device or software function that allows a user to record a video of a musical instrument performance.

[0869] The "uploading means" is a function for transmitting a performance video recorded by a user to a server via a network.

[0870] The "receiving means" is a function that allows the server to receive video data uploaded by users via the network.

[0871] An "analysis engine" is software or a device for extracting performance information such as finger movements, chord progressions, and rhythm from video data.

[0872] The "feedback generation means" is a function that uses a generative AI model based on the analysis results to generate advice and adjustment instructions for the user's performance.

[0873] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to automatically generate feedback on a user's performance.

[0874] The "playing style selection means" is a function that allows the user to select a practice mode based on a specific guitarist or playing style.

[0875] The "practice menu providing means" is a function that provides the user with instructional videos and practice menus relating to specific performance techniques and musical phrases based on the selected performance style.

[0876] The "co-starring video generation means" is a function that generates a co-starring video with an AI avatar guitarist or vocalist based on the user's practice results.

[0877] "SNS sharing means" is a function for sharing the generated collaboration video on social networking services, etc.

[0878] This invention relates to an online learning support system that allows beginners to efficiently learn musical instruments. The system allows users to record and upload their musical instrument performances. The server then analyzes the performance data using an analysis engine and provides feedback using a generative AI model. It also includes a practice mode selection function based on specific playing styles and a collaborative video generation function.

[0879] First, users can record their guitar playing using a mobile device such as a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" application and select the recorded video. This video file is then uploaded to the server through the application's interface. Users can also manage their videos by adding tags and notes.

[0880] The server receives video data uploaded by users. This data is temporarily saved and stored in a database for the next analysis step. The analysis engine uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[0881] Once the analysis is complete, the server uses a generative AI model, such as GPT-3 or T5, to generate feedback based on the analysis results. This feedback may include specific tuning adjustments, how to hold chords, how to use a pick, and rhythm. The generated feedback is provided in various formats, including video, images, text, and audio. This makes it easier for users to understand the feedback visually and audibly, and identify specific areas for improvement.

[0882] The user can also select practice modes based on specific guitarist styles, such as "Hotei Mode" and "Takanaka Mode." In this case, the server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the user's selection. This allows the user to efficiently practice specific styles and techniques.

[0883] Furthermore, by recording the results of the user's practice and uploading them back to the server, the server generates a video of the user's performance and a collaboration video with the AI ​​guitarist or vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. This allows users to realize their own progress and increase their motivation by sharing it with other users and the community.

[0884] As a concrete example, consider the case where a user is practicing an "F chord." After recording themselves playing an F chord using their smartphone, they launch the "Pick Perfect" app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord to detect errors in rhythm and pitch. Next, the server provides prompts to the generative AI model based on the analysis results, generating detailed feedback. For example, this could include specific advice such as, "Your index finger is not positioned correctly, so not all of the strings are sounding correctly."

[0885] When a user selects "Handy Mode," the server provides instructional videos on the guitarist's unique techniques and rhythm patterns. Furthermore, once the user reaches a certain level of skill, the server generates a video of the user performing with their AI avatar, which can then be shared on social media.

[0886] An example prompt for using a generative AI model is:

[0887] "Generate feedback on how to play the F chord in this video."

[0888] "Please analyze whether the user is playing with accurate rhythm."

[0889] "Create a practice routine based on a specific guitarist's style."

[0890] In this way, the present invention is a system that enables users to efficiently improve their playing skills and maintain a desire to continue learning.

[0891] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0892] Step 1: Record and upload

[0893] Users use a smartphone or tablet to record their guitar playing. Once recording is complete, they launch the "instrument learning app" and select the recorded video file. Then, they upload it to the server through the application's interface. Users can add tags and notes to the video as needed. The input is the user's performance video, and the output is the video data uploaded to the server.

[0894] Specific behavior:

[0895] Record your guitar playing using the camera app on your smartphone.

[0896] Launch the "Instrument Learning App" and tap the "Upload Video" button.

[0897] Select the recording file and tap the "Upload" button.

[0898] Step 2: Receiving video data

[0899] The server receives video data uploaded by users. This data is temporarily stored and stored in a database for the next analysis step. The input is the video data sent by the user, and the output is the video data stored on the server.

[0900] Specific behavior:

[0901] The server receives the HTTP request and stores the video file.

[0902] Store the video file path and metadata in the database.

[0903] Step 3: Analyze the video data

[0904] The server sends the stored video data to an analysis engine, which uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data. The input is the video data stored on the server, and the output is the analyzed performance data.

[0905] Specific behavior:

[0906] The server calls the analysis engine API and sends the path to the video file.

[0907] An analysis engine processes the data and extracts information such as finger position and chord progression.

[0908] The extraction results are sent back to the server.

[0909] Step 4: Generate feedback from the analysis results

[0910] The server generates feedback using a generative AI model based on the analysis results sent from the analysis engine. For example, it uses GPT-3 or T5 to generate feedback on the user's performance, such as specific tuning adjustments, how to hold chords, and how to keep rhythm. The input is the analyzed performance data, and the output is the generated feedback.

[0911] Specific behavior:

[0912] The server inputs the analysis results into the generative AI model and provides prompt sentences.

[0913] The generative AI model generates feedback and sends it back to the server.

[0914] Step 5: Provide feedback

[0915] The server provides the generated feedback to the user. This feedback can be in the form of video, image, text, or audio, and the user can check the feedback through the "instrument learning app." The input is the generated feedback, and the output is the feedback provided to the user.

[0916] Specific behavior:

[0917] The server converts the generated feedback into an appropriate format.

[0918] The feedback will be associated with the user's account and made available for viewing in the "instrument learning app."

[0919] Step 6: Select Practice Mode and Practice

[0920] The user selects a practice mode based on a specific playing style, such as "Guitarist A Mode" or "Guitarist B Mode," within the app. The server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the selected mode. The input is the user's practice mode selection, and the output is a specific practice menu or instructional video.

[0921] Specific behavior:

[0922] The user selects practice mode in the "instrument learning app."

[0923] The server provides the user with content based on the selected mode.

[0924] Step 7: Create and share a collaborative video

[0925] The user records the results of their practice and uploads them to the server, which then generates a video of the user's performance and a collaboration video with the AI ​​guitarist and vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. The input is the user's new performance video, and the output is the created collaboration video.

[0926] Specific behavior:

[0927] The user then records their performance again and uploads the video.

[0928] The server combines the new video with footage of the AI ​​avatar to create a collaborative video.

[0929] The generated collaboration video is associated with the user's account and can be displayed within the app.

[0930] (Application example 1)

[0931] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0932] Traditional methods for learning equipment operation and maintenance make it difficult to master skills efficiently and accurately, and there is a risk of reduced productivity and safety issues due to operational errors and improper use of tools. Furthermore, there is a lack of sharing functions that allow employees to realize their own improvement in their skills and increase their motivation.

[0933] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0934] In this invention, the server includes: a means for a user to record equipment operation; a means for the user to upload the recorded operation video; a means for the server to receive the uploaded video data; a means for the server to transmit the video data to an analysis engine and receive the analysis results; a means for the server to generate feedback based on the analysis results; a means for the server to provide the generated feedback to the user; a means for the user to select a specific training mode; a means for the server to provide a training menu based on the selected training mode; and a means for generating a video featuring the user's operation video and an AI model. This allows employees to efficiently and accurately learn operation and maintenance techniques, reducing operational errors and improper tool use. Furthermore, through the generated feedback and video, employees can realize their own improvement in their skills and share them with the factory community, thereby increasing their motivation.

[0935] A "means for recording device operation" is a device or method that a user uses to record the operation of a machine or device.

[0936] "Means for uploading recorded operation videos" refers to the device or method used by a user to transmit or transfer the operation videos recorded by the user to a server.

[0937] The "means for the server to receive uploaded video data" refers to a device or method for the server to store and check video data received via the network.

[0938] The "means for the server to transmit video data to the analysis engine and receive the analysis results" refers to a device or method for the server to transmit video data to an engine for analyzing the data and obtain the analysis results.

[0939] The "means for the server to generate feedback based on the analysis results" refers to a device or method that allows the server to automatically generate advice or evaluations for the user based on data obtained from the analysis engine.

[0940] "Means by which the server provides generated feedback to the user" refers to a device or method for transmitting and displaying the generated feedback to the user.

[0941] A "means for user selection of a particular training mode" is a device or method by which a user selects a pre-defined training or practice mode.

[0942] The "means for the server to provide a training menu based on the selected training mode" refers to a device or method for providing specific training or practice content in accordance with the selected training mode.

[0943] "Means for generating a video of a user's operation and an AI model together" refers to a device or method for generating a new video of a user's operation and an AI model together by combining the video of the user's operation recorded with the video of the virtual model generated by AI.

[0944] This invention provides an online system that supports efficient learning of equipment operation and maintenance techniques. Below, the program processing of this system will be explained in natural language, along with specific examples.

[0945] System Overview

[0946] The system starts with employees recording their equipment operations and uploading the video data to a server. The server receives the uploaded video data, sends it to an analysis engine, and receives the analysis results. Based on the analysis results, the server generates feedback and provides it to the user. Furthermore, users can receive special training menus by selecting specific training modes. A means is also provided to generate videos in which the user's operation video is combined with an AI model.

[0947] Hardware and software used

[0948] 1. Hardware

[0949] Smart glasses or smartphones: used by users to record device operations.

[0950] Server: Receives, stores, and analyzes video data, and generates and provides feedback.

[0951] Factory Robots: Used to perform specific tasks.

[0952] 2. Software

[0953] Python: The primary programming language in which the entire program runs.

[0954] OpenCV: A library for loading video data and splitting frames.

[0955] numpy: A library for splitting frames and processing data.

[0956] Dedicated analysis engine (AI backend): An engine that analyzes operation procedures, accuracy of actions, and tools used from video data.

[0957] Generative AI model: A model that generates feedback based on analysis results.

[0958] Data processing and calculation

[0959] 1. Loading and processing video data

[0960] The server receives video files uploaded by users and splits them into frames using OpenCV, which are then sent to the analysis engine using the numpy library.

[0961] 2. Analysis and feedback generation

[0962] The analytics engine extracts information from each frame of the video, such as operational procedures, accuracy of movements, and tools used. Based on this, the server uses a generative AI model to generate feedback, including specific improvements, efficiency suggestions, and safety precautions.

[0963] 3. Providing Feedback

[0964] The generated feedback is provided in the form of video, images, text, and audio, and can be viewed by users through smart glasses or a smartphone application.

[0965] Specific examples

[0966] For example, an employee can record a machine maintenance operation and upload it to a server. This video is then analyzed by an analysis engine to identify problems with the operating procedures and how to properly use tools. The server then generates specific feedback based on the analysis results and provides it to the user. The user can then perform the operation again according to the feedback, thereby improving their skills.

[0967] Prompt Sentence Examples

[0968] Capture user operation videos and use the analytics engine to analyze operating procedures, tools used, and accuracy, generating feedback with specific suggestions for improvements and efficiencies, as well as safety precautions.

[0969] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0970] Step 1:

[0971] The user records their device operations using smart glasses or a smartphone. The recorded video data is saved on the device. The input of this step is the user's operation behavior, and the output is the recorded video data.

[0972] Step 2:

[0973] The user launches the "Pick Perfect for Factory" app on their device, selects the recorded video data, and uploads it to the server. The video data is transferred from the user's device to the server. The input of this step is the recorded video data, and the output is the video data stored on the server.

[0974] Step 3:

[0975] The server receives the uploaded video data, divides it into frames using OpenCV and numpy, and sends them to the analysis engine. When the server divides the video data into frames, the input is the video data, and the output is image data for each frame. The analysis engine receives this and extracts information such as operation procedures, accuracy of actions, and tools used.

[0976] Step 4:

[0977] The server receives the analysis results sent by the analysis engine and uses the generative AI model to generate feedback including specific improvements, efficiency suggestions, and safety precautions. The input for this step is the analysis results, and the output is the generated feedback.

[0978] Step 5:

[0979] The server provides the generated feedback to the user. The feedback can be in the form of video, image, text, or audio. The user can view it through the app. The input of this step is the generated feedback, and the output is the feedback data that the user can view.

[0980] Step 6:

[0981] The user selects a specific training mode through the application and receives a training menu based on that mode. The server provides appropriate educational content based on the selected training mode. The input of this step is the user's selected training mode, and the output is the training menu and educational content.

[0982] Step 7:

[0983] The user's operation video is recorded again, and a video is generated that combines the re-uploaded video with the AI ​​model. The input for this step is the user's operation video and the AI ​​model, and the output is the combined video. The generated combined video can be viewed on the app and shared with the factory community.

[0984] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0985] This invention relates to an online learning support system that enables beginners to learn musical instruments efficiently and independently, and further enhances the user's learning experience by incorporating an emotion engine. Below, the program processing of this system is explained in natural language, and specific examples are given.

[0986] Program processing explanation

[0987] 1. Users record and upload their musical performances

[0988] Users use their smartphones or tablets to record their musical instrument performances. After recording is complete, they launch the "Pick Perfect" app, select the recorded file, and upload it to the server. At this time, users can enter tags and notes to make it easier to manage the recorded data.

[0989] 2. The server receives and analyzes the video data

[0990] The server receives the uploaded video data and stores it in temporary storage. After storage is complete, the video data is sent to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[0991] 3. The server uses the emotion engine to analyze the user's emotions.

[0992] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect their emotional state. The emotion engine identifies different emotions such as joy, sadness, surprise, and concentration.

[0993] 4. Generate feedback from analysis results and emotion data

[0994] The server combines the performance data received from the analysis engine and the emotional data received from the emotion engine, and generates feedback using a generative AI model. Specific feedback includes advice on improving performance technique and motivational messages.

[0995] 5. Provide feedback to users

[0996] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device. The user can then review the feedback through the "Pick Perfect" app and practice again according to the advice.

[0997] 6. Select a specific practice mode and practice

[0998] When a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style, including instructional videos on specific playing techniques and musical phrases.

[0999] 7. Creating and sharing collaborative videos

[1000] Once the user reaches a certain level of skill, they can record and upload the performance video again for re-analysis. The server then re-integrates the analysis results with the emotional data and generates a video of the user performing with their AI avatar, the guitarist and vocalist. The generated video can be viewed by the user on the app screen and shared on social media.

[1001] Specific examples

[1002] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord, and then upload the recording to the server via the "Pick Perfect" app. The server receives the video, and the analysis engine analyzes the intonation and sound accuracy, and then the emotion engine detects the user's facial expressions and tone of voice.

[1003] The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going." This not only helps users improve their playing technique, but also provides emotional support.

[1004] Furthermore, if the user selects "Home Mode," a practice menu tailored to the playing style of a specific guitarist is provided. Based on this practice mode, the server provides instructional videos on specific techniques and phrases.

[1005] Once a user reaches a certain level of skill, they can record and upload another video of their performance. The server then reanalyzes the newly received video and, based on the analysis results, generates a video of the user performing with their AI guitarist and vocalist. The generated video can be viewed by the user on the app screen and easily shared on social media and other platforms.

[1006] The above system further improves the user's learning experience, allowing them to continue practicing their instrument while improving their performance skills and maintaining their motivation.

[1007] The processing flow will be explained below.

[1008] Step 1:

[1009] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded video file.

[1010] Step 2:

[1011] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[1012] Step 3:

[1013] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, it automatically prepares to start the next analysis process.

[1014] Step 4:

[1015] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[1016] Step 5:

[1017] The analysis engine generates analysis results and sends them to the server, which include detailed performance data for each frame.

[1018] Step 6:

[1019] The server receives the analysis results and sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[1020] Step 7:

[1021] The emotion engine generates emotion data and sends the data to the server, which includes different emotions such as joy, sadness, surprise, and concentration.

[1022] Step 8:

[1023] The server combines the analysis results with the emotional data and uses a generative AI model to generate feedback, including advice on improving performance technique and motivational messages.

[1024] Step 9:

[1025] The server sends the generated feedback to the user's device, which can be in the form of video, images, text, or audio.

[1026] Step 10:

[1027] The user reviews the feedback through the app, follows the advice on specific areas for improvement and how to improve, and practices again.

[1028] Step 11:

[1029] When a user selects a specific practice mode, the app configures the settings and sends the information to the server, which then provides a dedicated practice menu based on the selected practice mode.

[1030] Step 12:

[1031] The user practices according to a specific practice menu, then records the performance video again and uploads it to the server via the app.

[1032] Step 13:

[1033] The server receives the newly uploaded video and analyzes it again using the analysis engine and emotion engine. Based on the results, a video of the user performing together with an AI guitarist and vocalist is generated.

[1034] Step 14:

[1035] The server generates the video and sends it to the user's device, where the user can view it on the app screen and share it on social media or other platforms.

[1036] These are the specific processing steps of the present invention. This system allows users to practice while receiving appropriate feedback, and by receiving emotional support as well as improving their performance skills, they can maintain continuous motivation.

[1037] Example 2

[1038] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1039] Existing instrument learning systems make it difficult for users to effectively improve their playing skills and maintain their motivation. Specifically, feedback is limited to improving playing technique, and comprehensive support that takes into account the user's emotional state and motivation is lacking. Furthermore, they lack the provision of personalized practice menus based on specific playing styles, making it difficult to progress efficiently.

[1040] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1041] In this invention, the server includes means for an analysis engine to extract finger movements, chord progressions, and rhythms from video data, means for an emotion analysis engine to analyze facial expressions and tone of voice to detect the user's emotional state, and means for generating feedback using a generative AI model based on the analysis results and emotion data. This allows the user to receive advice that takes into account their emotional state along with technical feedback, allowing them to efficiently progress through musical instrument learning while receiving comprehensive support.

[1042] "User" refers to an individual who utilizes the system to practice playing an instrument and receive feedback.

[1043] "Instrument" refers to a tool or device that produces music for a user to play.

[1044] "Recording" refers to recording the user's performance as video data.

[1045] "Upload" refers to the process by which a user sends recorded video data to a server.

[1046] "Server" refers to a computer system that receives video data from users, stores it, analyzes it, and generates feedback.

[1047] "Video data" refers to video files of musical instrument performances recorded by users.

[1048] An "analysis engine" refers to software or algorithms used to extract technical information about musical instrument performance (finger movements, chord progressions, rhythm, etc.) from video data.

[1049] An "emotion analysis engine" refers to software or algorithms that analyze a user's facial expressions and tone of voice from video data to detect their emotional state.

[1050] "Generative AI model" refers to an artificial intelligence model that generates feedback appropriate for the user based on data from the analysis engine and sentiment analysis engine.

[1051] "Feedback" refers to messages that are generated based on the analysis results, such as advice on improving the user's playing technique or messages that increase motivation.

[1052] "Terminal" refers to the device (smartphone, tablet, etc.) that a user uses to record and check feedback.

[1053] "Practice mode" refers to a practice menu provided by the server that the user selects based on a particular playing style.

[1054] "Practice Menu" refers to content that provides instructional videos and advice on specific performance techniques or musical phrases.

[1055] This invention relates to an online learning support system for enabling beginners to efficiently self-study musical instruments. The system includes: means for a user to record their musical instrument performance; means for the user to upload the recorded performance video; means for a server to receive the uploaded video data; means for the server to transmit the video data to an analysis engine and receive analysis results that extract finger movements, chord progressions, and rhythm; means for the server to transmit the video data to an emotion analysis engine and analyze the user's emotional state; means for the server to generate feedback using a generative AI model based on the analysis results and the emotion data; and means for the server to provide the generated feedback to the user's terminal.

[1056] The specific process of this system is as follows: The user uses a device such as a smartphone or tablet to record themselves playing an instrument. A regular camera app is used for this recording. After recording is complete, the user launches a dedicated application, selects the recorded file, and uploads it to the server. When uploading, tags and notes are entered to make the recorded data easier to manage. The server receives the uploaded video data and saves it in temporary storage. A cloud service (e.g., Amazon S3 or Google Cloud Storage) is often used as this temporary storage.

[1057] The server then sends the saved video data to an analysis engine, which uses TensorFlow, OpenCV, and other tools to extract performance information such as finger movements, chord progressions, and rhythm. The server also sends the video data to an emotion analysis engine, which analyzes the user's facial expressions and tone of voice. The emotion analysis engine uses tools such as Microsoft Azure Emotion API and IBM Watson Tone Analyzer. This identifies the user's emotional state, such as joy, sadness, surprise, or concentration.

[1058] The analysis results and emotional data are integrated on the server, and feedback is generated using a generative AI model (e.g., GPT-4). This feedback includes advice on improving performance technique and motivational messages. The generated feedback is prepared in various formats (video, image, text, audio) and sent to the user's device. The user can view this feedback through a dedicated application.

[1059] Furthermore, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style. The practice menu includes instructional videos on specific playing techniques and musical phrases. These videos are provided using the YouTube API and Vimeo API.

[1060] For example, if a user is practicing an F chord, they can record themselves playing the F chord using their smartphone and upload the recording to a server via a dedicated app. The server receives the video, analyzes the articulation and sound accuracy using an analysis engine (TensorFlow or OpenCV), and then detects the user's facial expression and tone of voice using an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer). The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going!"

[1061] Here are some example prompts for a generative AI model:

[1062] "I'm practicing an F chord. Could you please analyze the video I recorded? Give me advice on finger placement and chord progression accuracy. I'd also like some motivational messages."

[1063] "The user selected a practice mode based on a specific playing style. Please provide a practice program tailored to the style of the specific guitarist. Also provide instructional videos on specific techniques and phrases."

[1064] "The user has reached a certain level of skill. Please analyze the new performance video and generate a collaboration video. Please generate a collaboration video between the user and the AI ​​and provide it in a format that can be shared on social media, etc."

[1065] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1066] Step 1:

[1067] A user records an instrument performance

[1068] Input: A user starts playing an instrument using a smartphone or tablet.

[1069] How it works: The user uses a regular camera app to record their performance as video data.

[1070] Output: Video data of the recorded instrument performance (video file)

[1071] Step 2:

[1072] Users upload recorded performance videos

[1073] Input: Recorded performance video (video file), launching the dedicated application

[1074] How it works: After recording is complete, the user launches the dedicated application, selects the recorded file, optionally enters tags and notes, and uploads the video to the server.

[1075] Output: Uploaded performance video (video file), management information (tags, notes)

[1076] Step 3:

[1077] The server receives the uploaded video data and stores it in temporary storage.

[1078] Input: Uploaded performance video (video file), management information (tags, notes)

[1079] How it works: The server receives video data uploaded by users and temporarily stores it in cloud storage (e.g., Amazon S3 or Google Cloud Storage).

[1080] Output: Video data saved in temporary storage

[1081] Step 4:

[1082] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[1083] Input: Saved video data

[1084] Operation: The server sends the video data to an analysis engine (TensorFlow or OpenCV), which analyzes the data to extract performance information such as finger movements, chord progressions, and rhythm.

[1085] Output: Performance analysis results (data on finger movements, chord progressions, rhythm, etc.)

[1086] Step 5:

[1087] The server sends the video data to the emotion analysis engine to analyze the user's emotional state.

[1088] Input: Saved video data

[1089] How it works: The server sends video data to an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer), which analyzes the user's facial expressions and tone of voice to identify their emotional state (happiness, sadness, surprise, concentration, etc.).

[1090] Output: Sentiment analysis results (data about the user's emotional state)

[1091] Step 6:

[1092] The server generates feedback using a generative AI model based on the analysis results and emotion data.

[1093] Input: Performance analysis results, emotion analysis results

[1094] How it works: The server integrates the analysis results with the emotional data and uses a generative AI model (e.g., GPT-4) to generate feedback for the user (such as advice on improving playing skills or messages to motivate them).

[1095] Output: Generated feedback (technical advice, motivational messages)

[1096] Step 7:

[1097] The server provides the generated feedback to the user's device.

[1098] Input: Generated feedback (technical advice, motivational messages)

[1099] How it works: The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device.

[1100] Output: Feedback is displayed in the user's dedicated application (technical advice, motivational messages)

[1101] Step 8:

[1102] The user selects a practice mode based on a specific playing style and receives a practice menu.

[1103] Input: The user selects a specific playing style in a dedicated application.

[1104] How it works: When a user selects a practice mode based on a specific playing style, the server provides instructional videos on specific playing techniques or musical phrases based on that selection. This is done using the YouTube API and Vimeo API.

[1105] Output: Practice menu (explanatory videos, technique information)

[1106] Step 9:

[1107] When a user reaches a certain level of skill, they can upload a new performance video.

[1108] Input: A device for users to re-record and upload performance videos

[1109] What it does: When a user reaches a certain skill level, it records a new performance video and uploads it to the server.

[1110] Output: Newly uploaded performance video (video file)

[1111] Step 10:

[1112] The server reanalyzes the newly received video and generates a co-video

[1113] Input: Newly uploaded performance video, analysis engine for reanalysis, and emotion analysis engine results

[1114] Operation: The server retransmits the newly received video to the analysis engine and emotion analysis engine, which reanalyzes the data. Then, based on the analysis results, it generates a video of the user performing with the AI ​​guitarist and vocalist.

[1115] Output: Collaborative video (a video of the user and their AI avatar)

[1116] Step 11:

[1117] The server provides the generated collaborative video to users, making it possible to share it.

[1118] Input: Generated collaboration video

[1119] Operation: The server sends the generated video to the user's dedicated application, allowing the user to view the video and share it on social media or other platforms.

[1120] Output: Shareable collaboration video (can be shared on social media)

[1121] (Application example 2)

[1122] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1123] Beginners learning musical instruments require accurate understanding of their progress and effective feedback when self-learning. However, conventional online learning systems often provide feedback based solely on the user's technical performance data, and lack support that takes into account the user's emotional state and motivation. The present invention aims to provide an online learning support system that enables beginners to efficiently self-learn musical instruments, improving their skills and maintaining their motivation at the same time.

[1124] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for users to upload recorded performance videos, means for analyzing the user's emotional state using an emotion engine, and means for generating feedback by integrating the analysis results and emotion data based on a generative AI model. This makes it possible to provide not only feedback based on the analysis results of the user's technical performance, but also comprehensive feedback that takes the user's emotional state into consideration in real time.

[1125] A "user" is an individual who is learning to play a musical instrument and who uses the online learning support system.

[1126] "Musical instrument" refers to tools used to play music, especially stringed or keyboard instruments such as guitars, pianos, and violins.

[1127] "Recording" refers to the act of saving a user's performance in video format on a digital device.

[1128] A "performance video" is a video file that records a user playing a musical instrument.

[1129] "Uploading" refers to the act of a user transmitting a recorded performance video to a server via the Internet.

[1130] A "server" is a computer system that receives, analyzes, and processes uploaded data.

[1131] "Video data" refers to digital data of a recorded video of a performance.

[1132] An "analysis engine" is software that analyzes video data and extracts technical information such as finger movements, chord progressions, and rhythm.

[1133] "Analysis results" refers to the technical information extracted from the video data by the analysis engine.

[1134] The "emotion engine" is software that analyzes the user's facial expressions and tone of voice to identify emotions such as joy, sadness, surprise, and concentration.

[1135] "Emotion data" is information about the user's emotional state analyzed by the emotion engine.

[1136] A "generative AI model" is an algorithm that uses generative adversarial networks (GANs) and other AI techniques to integrate analytical results with emotional data and generate feedback.

[1137] "Feedback" refers to advice and messages generated based on analysis results and emotional data that help users improve their performance skills.

[1138] A "wearable device" is a computing device that can be worn by a user, such as smart glasses or a head-mounted display.

[1139] "Real-time" refers to providing instant feedback at the moment the performance is taking place.

[1140] "Practice Mode" is a personalized practice program that a user selects based on a particular playing style or technique.

[1141] A "practice menu" is a series of practice exercises and instructional videos provided by the server according to the selected practice mode.

[1142] The present invention is an online learning support system for helping beginners to learn musical instruments efficiently. The system provides a means for users to record their musical instrument performances and upload the recorded performance videos. The uploaded video data is received by a server, where it is then processed by an analysis engine and an emotion engine. The analysis engine extracts information such as finger movements, chord progressions, and rhythm from the performance data, while the emotion engine analyzes the user's facial expressions and vocal tone to identify their emotional state.

[1143] The server integrates the performance data received from the analysis engine and the emotion data received from the emotion engine, and generates feedback using a generative AI model. This feedback includes advice on improving performance technique and messages to motivate the user.

[1144] Users can receive real-time feedback using wearable devices (e.g., smart glasses or head-mounted displays). This allows them to get feedback at the moment they are playing and immediately try to improve their technique. In addition, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on the user's selection. This practice menu includes instructional videos on specific techniques or musical phrases.

[1145] Furthermore, once the user reaches a certain level of skill, they can record another performance video, upload it to the server, and reanalyze it. The server then reintegrates the new analysis results with the emotional data to generate a video of the user performing with the virtual performer. This video can be viewed on the user's device screen and easily shared on social media.

[1146] Examples:

[1147] For example, let's say a user is practicing an "F chord." First, they use the smart glasses to record themselves playing the F chord, and then upload the recording to a server via an application. The server receives the video data, analyzes the finger movements and rhythm using an analysis engine, and then analyzes the user's facial expressions and tone of voice using an emotion engine to detect their emotional state. The analysis results include technical feedback such as "Your finger position is not accurate, so the notes are out of sync," as well as encouraging feedback such as "You're making a good learning pace, so keep going."

[1148] When a user selects the "Specific Guitarist Mode" as their practice mode, the server generates a practice menu based on this mode. The practice menu includes instructional videos on specific techniques and phrases. Once the user reaches a certain level, they record and upload the performance video again for further analysis. Based on the obtained data, the server generates a video of the user performing with a virtual musician and provides it to the user in a shareable format.

[1149] Example prompt sentence:

[1150] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[1151] This provides a system that allows users to improve their performance skills while continuing to learn while receiving emotional support.

[1152] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1153] (Processing step flow)

[1154] Step 1:

[1155] A user records their performance using a wearable device (e.g., smart glasses or a head-mounted display). The recorded video data is stored on the device.

[1156] Input: User's performance video

[1157] Output: Recorded video data

[1158] Specific behavior: The user puts on the device and starts playing, presses the record button to record the performance, and presses the button again to stop recording when finished.

[1159] Step 2:

[1160] The user uploads the recorded performance video to the server using the application. The user selects the video file in the application and presses the upload button.

[1161] Input: Recorded video data

[1162] Output: Video data uploaded to the server

[1163] Specific operation: Open the application, select the recording file, and press the upload button. The video data will be sent to the server via the Internet.

[1164] Step 3:

[1165] The server receives the uploaded video data and stores it in temporary storage.

[1166] Input: Uploaded video data

[1167] Output: Video data stored in temporary storage

[1168] Specific operation: The server receives data sent via the Internet and saves it in the specified directory.

[1169] Step 4:

[1170] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm. The analysis engine then generates the extracted technical information.

[1171] Input: Video data stored in temporary storage

[1172] Output: Technical analysis results

[1173] How it works: The server passes the video data to the analysis engine, which then uses image recognition algorithms to extract performance technique information, such as finger position and movement.

[1174] Step 5:

[1175] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[1176] Input: Video data stored in temporary storage

[1177] Output: Emotion analysis results

[1178] How it works: The server passes video data to the emotion engine, which then uses emotion recognition algorithms to analyze the user's facial expressions and tone of voice.

[1179] Step 6:

[1180] The server integrates the technical analysis results from the analysis engine and the emotional analysis results from the emotion engine and generates feedback using a generative AI model.

[1181] Input: Technical analysis results, sentiment analysis results

[1182] Output: Feedback from a generative AI model

[1183] Specific operation: The server inputs the two analysis results using a generative AI model such as Python or TensorFlow and generates an appropriate feedback message.

[1184] Step 7:

[1185] The server sends the generated feedback to the user's terminal, where it is displayed in real time on the user's wearable device.

[1186] Input: Feedback from a generative AI model

[1187] Output: Feedback provided to the user

[1188] Specific operation: The server sends a feedback message in text, audio, or image format to the user's device and displays it on the device.

[1189] (Example prompt)

[1190] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[1191] This allows users to receive real-time feedback on areas to improve their playing technique and maintain their motivation, enabling effective self-learning.

[1192] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1193] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1194] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1195] [Fourth embodiment]

[1196] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1197] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1198] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1199] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1200] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1201] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1202] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1203] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1204] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1205] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1206] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1207] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1208] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1209] The present invention relates to an online learning support system that allows beginners to learn musical instruments efficiently. Below, we will explain the program processing of this system in natural language and provide a concrete example.

[1210] Program processing explanation

[1211] 1. A user records and uploads their guitar performance

[1212] Users can record their guitar playing using a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" app and upload the video to the server through the app's interface. Users can also add tags and notes as needed for easy management.

[1213] 2. The server receives and analyzes the video data

[1214] The server receives video data uploaded by users. After receiving the video data, it sends it to an analysis engine for analysis. The analysis engine extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[1215] 3. Generate feedback on analysis results

[1216] The server receives the analysis results sent from the analysis engine and uses the generative AI model to generate feedback on the user's performance, including specific tuning adjustments, how to hold chords, how to use a pick, and how to keep rhythm.

[1217] 4. Provide feedback to users

[1218] The server then provides the generated feedback to the user in the form of video, image, text, or audio, which the user can review through the "Pick Perfect" app. Providing feedback in a visually and audibly easy-to-understand format allows users to easily understand specific areas for improvement and how to improve them.

[1219] 5. Select a specific practice mode and practice

[1220] Users can select practice modes based on specific guitarist styles, such as "Home Mode" or "High-Level Mode." Based on the selected mode, the server provides instructional videos and practice menus for specific playing techniques and musical phrases. This allows users to efficiently practice specific styles and techniques.

[1221] 6. Creating and sharing collaborative videos

[1222] When users record and upload their practice results, the server generates a video of the user's performance and a collaboration with the AI ​​guitarist and vocalist. The collaboration video can be viewed on the user's app screen and shared on social media. This allows users to see their own progress and increase their motivation by sharing it with other users and the community.

[1223] Specific examples

[1224] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord. Once they've finished recording, they launch the Pick Perfect app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord, and detects any rhythm or pitch errors as appropriate.

[1225] The server then generates detailed feedback based on the analysis, including specific advice such as, "Your index finger is not positioned correctly, so not all the strings are sounding correctly." This feedback is provided with video, illustrations, and audio commentary, making it easy for users to understand and continue practicing.

[1226] When a user selects "Handy Mode," a practice menu based on the playing style of a specific guitarist is provided. The server transmits instructional videos on the guitarist's unique techniques and rhythm patterns, allowing the user to practice accordingly.

[1227] Furthermore, when a user reaches a certain level of skill, they can create a video in which they appear alongside their AI avatar and share it on social media, allowing users to see their skills improving and keeping their motivation high.

[1228] By implementing the system described above, beginners can learn to play musical instruments efficiently and experience the joy of playing musical instruments.

[1229] The processing flow will be explained below.

[1230] Step 1:

[1231] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded file.

[1232] Step 2:

[1233] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[1234] Step 3:

[1235] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, the server prepares to send the video file to the analysis engine.

[1236] Step 4:

[1237] The server sends the video file to an analysis engine, which then analyzes the video data and extracts information about performance techniques such as finger movements, chord progressions, rhythm, and tempo.

[1238] Step 5:

[1239] The analysis engine generates the analysis results and sends them back to the server, including the finger position information for each frame and the accuracy of the sound.

[1240] Step 6:

[1241] The server receives the analysis results and generates feedback using a generative AI model based on the analysis data. The feedback consists of specific performance improvement points and advice.

[1242] Step 7:

[1243] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device, allowing the user to receive feedback in a variety of formats.

[1244] Step 8:

[1245] Users can check the feedback through the app and practice again according to the advice. The feedback includes specific points and methods for improvement, making it easy for users to understand and put into practice.

[1246] Step 9:

[1247] When a user selects a particular practice mode, such as "Home Mode" or "High Mode," the server provides a practice menu based on the selected mode, including explanations of the playing style and techniques of a particular guitarist.

[1248] Step 10:

[1249] Once the user reaches a certain level of skill, they can record and upload the performance video again. The server receives the newly uploaded video and re-analyzes it using the analysis engine.

[1250] Step 11:

[1251] Based on the new analysis results, the server generates a video of the user performing with the AI ​​guitarist and vocalist. The video is provided in a format where the user's performance and the AI's performance are synchronized.

[1252] Step 12:

[1253] Users can view their collaborative videos and share them on social media and other platforms through the "Pick Perfect" app, allowing them to share their progress with others and gain additional motivation.

[1254] Example 1

[1255] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1256] Conventional musical instrument practice systems do not provide sufficient feedback for users to efficiently improve their playing skills. Furthermore, they lack features such as practice modes based on specific playing styles or collaborative video generation to enjoy the results of practice, making it difficult for users to maintain their motivation to learn.

[1257] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1258] In this invention, the server includes means for receiving uploaded video data, means for transmitting the video data to an analysis engine and receiving the analysis results, means for generating feedback using a generative AI model based on the analysis results, means for the user to check the feedback and select a practice mode based on a specific playing style, means for providing a practice menu based on the selected practice mode, and means for generating and providing a collaborative video based on the user's practice results, which enables users to efficiently improve their playing skills and maintain their motivation to continue learning.

[1259] The "recording means" is a device or software function that allows a user to record a video of a musical instrument performance.

[1260] The "uploading means" is a function for transmitting a performance video recorded by a user to a server via a network.

[1261] The "receiving means" is a function that allows the server to receive video data uploaded by users via the network.

[1262] An "analysis engine" is software or a device for extracting performance information such as finger movements, chord progressions, and rhythm from video data.

[1263] The "feedback generation means" is a function that uses a generative AI model based on the analysis results to generate advice and adjustment instructions for the user's performance.

[1264] A "generative AI model" is an artificial intelligence model that uses machine learning algorithms to automatically generate feedback on a user's performance.

[1265] The "playing style selection means" is a function that allows the user to select a practice mode based on a specific guitarist or playing style.

[1266] The "practice menu providing means" is a function that provides the user with instructional videos and practice menus relating to specific performance techniques and musical phrases based on the selected performance style.

[1267] The "co-starring video generation means" is a function that generates a co-starring video with an AI avatar guitarist or vocalist based on the user's practice results.

[1268] "SNS sharing means" is a function for sharing the generated collaboration video on social networking services, etc.

[1269] This invention relates to an online learning support system that allows beginners to efficiently learn musical instruments. The system allows users to record and upload their musical instrument performances. The server then analyzes the performance data using an analysis engine and provides feedback using a generative AI model. It also includes a practice mode selection function based on specific playing styles and a collaborative video generation function.

[1270] First, users can record their guitar playing using a mobile device such as a smartphone or tablet. Once recording is complete, they launch the "Pick Perfect" application and select the recorded video. This video file is then uploaded to the server through the application's interface. Users can also manage their videos by adding tags and notes.

[1271] The server receives video data uploaded by users. This data is temporarily saved and stored in a database for the next analysis step. The analysis engine uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[1272] Once the analysis is complete, the server uses a generative AI model, such as GPT-3 or T5, to generate feedback based on the analysis results. This feedback may include specific tuning adjustments, how to hold chords, how to use a pick, and rhythm. The generated feedback is provided in various formats, including video, images, text, and audio. This makes it easier for users to understand the feedback visually and audibly, and identify specific areas for improvement.

[1273] The user can also select practice modes based on specific guitarist styles, such as "Hotei Mode" and "Takanaka Mode." In this case, the server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the user's selection. This allows the user to efficiently practice specific styles and techniques.

[1274] Furthermore, by recording the results of the user's practice and uploading them back to the server, the server generates a video of the user's performance and a collaboration video with the AI ​​guitarist or vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. This allows users to realize their own progress and increase their motivation by sharing it with other users and the community.

[1275] As a concrete example, consider the case where a user is practicing an "F chord." After recording themselves playing an F chord using their smartphone, they launch the "Pick Perfect" app, select the recorded file, and upload it. When the server receives the video, the analysis engine analyzes the finger placement and fingering of the F chord to detect errors in rhythm and pitch. Next, the server provides prompts to the generative AI model based on the analysis results, generating detailed feedback. For example, this could include specific advice such as, "Your index finger is not positioned correctly, so not all of the strings are sounding correctly."

[1276] When a user selects "Handy Mode," the server provides instructional videos on the guitarist's unique techniques and rhythm patterns. Furthermore, once the user reaches a certain level of skill, the server generates a video of the user performing with their AI avatar, which can then be shared on social media.

[1277] An example prompt for using a generative AI model is:

[1278] "Generate feedback on how to play the F chord in this video."

[1279] "Please analyze whether the user is playing with accurate rhythm."

[1280] "Create a practice routine based on a specific guitarist's style."

[1281] In this way, the present invention is a system that enables users to efficiently improve their playing skills and maintain a desire to continue learning.

[1282] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1283] Step 1: Record and upload

[1284] Users use a smartphone or tablet to record their guitar playing. Once recording is complete, they launch the "instrument learning app" and select the recorded video file. Then, they upload it to the server through the application's interface. Users can add tags and notes to the video as needed. The input is the user's performance video, and the output is the video data uploaded to the server.

[1285] Specific behavior:

[1286] Record your guitar playing using the camera app on your smartphone.

[1287] Launch the "Instrument Learning App" and tap the "Upload Video" button.

[1288] Select the recording file and tap the "Upload" button.

[1289] Step 2: Receiving video data

[1290] The server receives video data uploaded by users. This data is temporarily stored and stored in a database for the next analysis step. The input is the video data sent by the user, and the output is the video data stored on the server.

[1291] Specific behavior:

[1292] The server receives the HTTP request and stores the video file.

[1293] Store the video file path and metadata in the database.

[1294] Step 3: Analyze the video data

[1295] The server sends the stored video data to an analysis engine, which uses technologies such as "OpenPose" and "MediaPipe" to extract information such as finger movements, chord progressions, rhythm, and tempo from the video data. The input is the video data stored on the server, and the output is the analyzed performance data.

[1296] Specific behavior:

[1297] The server calls the analysis engine API and sends the path to the video file.

[1298] An analysis engine processes the data and extracts information such as finger position and chord progression.

[1299] The extraction results are sent back to the server.

[1300] Step 4: Generate feedback from the analysis results

[1301] The server generates feedback using a generative AI model based on the analysis results sent from the analysis engine. For example, it uses GPT-3 or T5 to generate feedback on the user's performance, such as specific tuning adjustments, how to hold chords, and how to keep rhythm. The input is the analyzed performance data, and the output is the generated feedback.

[1302] Specific behavior:

[1303] The server inputs the analysis results into the generative AI model and provides prompt sentences.

[1304] The generative AI model generates feedback and sends it back to the server.

[1305] Step 5: Provide feedback

[1306] The server provides the generated feedback to the user. This feedback can be in the form of video, image, text, or audio, and the user can check the feedback through the "instrument learning app." The input is the generated feedback, and the output is the feedback provided to the user.

[1307] Specific behavior:

[1308] The server converts the generated feedback into an appropriate format.

[1309] The feedback will be associated with the user's account and made available for viewing in the "instrument learning app."

[1310] Step 6: Select Practice Mode and Practice

[1311] The user selects a practice mode based on a specific playing style, such as "Guitarist A Mode" or "Guitarist B Mode," within the app. The server provides instructional videos and practice menus for specific playing techniques and musical phrases based on the selected mode. The input is the user's practice mode selection, and the output is a specific practice menu or instructional video.

[1312] Specific behavior:

[1313] The user selects practice mode in the "instrument learning app."

[1314] The server provides the user with content based on the selected mode.

[1315] Step 7: Create and share a collaborative video

[1316] The user records the results of their practice and uploads them to the server, which then generates a video of the user's performance and a collaboration video with the AI ​​guitarist and vocalist. The created collaboration video can be viewed on the user's app screen and can be shared on social media. The input is the user's new performance video, and the output is the created collaboration video.

[1317] Specific behavior:

[1318] The user then records their performance again and uploads the video.

[1319] The server combines the new video with footage of the AI ​​avatar to create a collaborative video.

[1320] The generated collaboration video is associated with the user's account and can be displayed within the app.

[1321] (Application example 1)

[1322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1323] Traditional methods for learning equipment operation and maintenance make it difficult to master skills efficiently and accurately, and there is a risk of reduced productivity and safety issues due to operational errors and improper use of tools. Furthermore, there is a lack of sharing functions that allow employees to realize their own improvement in their skills and increase their motivation.

[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1325] In this invention, the server includes: a means for a user to record equipment operation; a means for the user to upload the recorded operation video; a means for the server to receive the uploaded video data; a means for the server to transmit the video data to an analysis engine and receive the analysis results; a means for the server to generate feedback based on the analysis results; a means for the server to provide the generated feedback to the user; a means for the user to select a specific training mode; a means for the server to provide a training menu based on the selected training mode; and a means for generating a video featuring the user's operation video and an AI model. This allows employees to efficiently and accurately learn operation and maintenance techniques, reducing operational errors and improper tool use. Furthermore, through the generated feedback and video, employees can realize their own improvement in their skills and share them with the factory community, thereby increasing their motivation.

[1326] A "means for recording device operation" is a device or method that a user uses to record the operation of a machine or device.

[1327] "Means for uploading recorded operation videos" refers to the device or method used by a user to transmit or transfer the operation videos recorded by the user to a server.

[1328] The "means for the server to receive uploaded video data" refers to a device or method for the server to store and check video data received via the network.

[1329] The "means for the server to transmit video data to the analysis engine and receive the analysis results" refers to a device or method for the server to transmit video data to an engine for analyzing the data and obtain the analysis results.

[1330] The "means for the server to generate feedback based on the analysis results" refers to a device or method that allows the server to automatically generate advice or evaluations for the user based on data obtained from the analysis engine.

[1331] "Means by which the server provides generated feedback to the user" refers to a device or method for transmitting and displaying the generated feedback to the user.

[1332] A "means for user selection of a particular training mode" is a device or method by which a user selects a pre-defined training or practice mode.

[1333] The "means for the server to provide a training menu based on the selected training mode" refers to a device or method for providing specific training or practice content in accordance with the selected training mode.

[1334] "Means for generating a video of a user's operation and an AI model together" refers to a device or method for generating a new video of a user's operation and an AI model together by combining the video of the user's operation recorded with the video of the virtual model generated by AI.

[1335] This invention provides an online system that supports efficient learning of equipment operation and maintenance techniques. Below, the program processing of this system will be explained in natural language, along with specific examples.

[1336] System Overview

[1337] The system starts with employees recording their equipment operations and uploading the video data to a server. The server receives the uploaded video data, sends it to an analysis engine, and receives the analysis results. Based on the analysis results, the server generates feedback and provides it to the user. Furthermore, users can receive special training menus by selecting specific training modes. A means is also provided to generate videos in which the user's operation video is combined with an AI model.

[1338] Hardware and software used

[1339] 1. Hardware

[1340] Smart glasses or smartphones: used by users to record device operations.

[1341] Server: Receives, stores, and analyzes video data, and generates and provides feedback.

[1342] Factory Robots: Used to perform specific tasks.

[1343] 2. Software

[1344] Python: The primary programming language in which the entire program runs.

[1345] OpenCV: A library for loading video data and splitting frames.

[1346] numpy: A library for splitting frames and processing data.

[1347] Dedicated analysis engine (AI backend): An engine that analyzes operation procedures, accuracy of actions, and tools used from video data.

[1348] Generative AI model: A model that generates feedback based on analysis results.

[1349] Data processing and calculation

[1350] 1. Loading and processing video data

[1351] The server receives video files uploaded by users and splits them into frames using OpenCV, which are then sent to the analysis engine using the numpy library.

[1352] 2. Analysis and feedback generation

[1353] The analytics engine extracts information from each frame of the video, such as operational procedures, accuracy of movements, and tools used. Based on this, the server uses a generative AI model to generate feedback, including specific improvements, efficiency suggestions, and safety precautions.

[1354] 3. Providing Feedback

[1355] The generated feedback is provided in the form of video, images, text, and audio, and can be viewed by users through smart glasses or a smartphone application.

[1356] Specific examples

[1357] For example, an employee can record a machine maintenance operation and upload it to a server. This video is then analyzed by an analysis engine to identify problems with the operating procedures and how to properly use tools. The server then generates specific feedback based on the analysis results and provides it to the user. The user can then perform the operation again according to the feedback, thereby improving their skills.

[1358] Prompt Sentence Examples

[1359] Capture user operation videos and use the analytics engine to analyze operating procedures, tools used, and accuracy, generating feedback with specific suggestions for improvements and efficiencies, as well as safety precautions.

[1360] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1361] Step 1:

[1362] The user records their device operations using smart glasses or a smartphone. The recorded video data is saved on the device. The input of this step is the user's operation behavior, and the output is the recorded video data.

[1363] Step 2:

[1364] The user launches the "Pick Perfect for Factory" app on their device, selects the recorded video data, and uploads it to the server. The video data is transferred from the user's device to the server. The input of this step is the recorded video data, and the output is the video data stored on the server.

[1365] Step 3:

[1366] The server receives the uploaded video data, divides it into frames using OpenCV and numpy, and sends them to the analysis engine. When the server divides the video data into frames, the input is the video data, and the output is image data for each frame. The analysis engine receives this and extracts information such as operation procedures, accuracy of actions, and tools used.

[1367] Step 4:

[1368] The server receives the analysis results sent by the analysis engine and uses the generative AI model to generate feedback including specific improvements, efficiency suggestions, and safety precautions. The input for this step is the analysis results, and the output is the generated feedback.

[1369] Step 5:

[1370] The server provides the generated feedback to the user. The feedback can be in the form of video, image, text, or audio. The user can view it through the app. The input of this step is the generated feedback, and the output is the feedback data that the user can view.

[1371] Step 6:

[1372] The user selects a specific training mode through the application and receives a training menu based on that mode. The server provides appropriate educational content based on the selected training mode. The input of this step is the user's selected training mode, and the output is the training menu and educational content.

[1373] Step 7:

[1374] The user's operation video is recorded again, and a video is generated that combines the re-uploaded video with the AI ​​model. The input for this step is the user's operation video and the AI ​​model, and the output is the combined video. The generated combined video can be viewed on the app and shared with the factory community.

[1375] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1376] This invention relates to an online learning support system that enables beginners to learn musical instruments efficiently and independently, and further enhances the user's learning experience by incorporating an emotion engine. Below, the program processing of this system is explained in natural language, and specific examples are given.

[1377] Program processing explanation

[1378] 1. Users record and upload their musical performances

[1379] Users use their smartphones or tablets to record their musical instrument performances. After recording is complete, they launch the "Pick Perfect" app, select the recorded file, and upload it to the server. At this time, users can enter tags and notes to make it easier to manage the recorded data.

[1380] 2. The server receives and analyzes the video data

[1381] The server receives the uploaded video data and stores it in temporary storage. After storage is complete, the video data is sent to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[1382] 3. The server uses the emotion engine to analyze the user's emotions.

[1383] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect their emotional state. The emotion engine identifies different emotions such as joy, sadness, surprise, and concentration.

[1384] 4. Generate feedback from analysis results and emotion data

[1385] The server combines the performance data received from the analysis engine and the emotional data received from the emotion engine, and generates feedback using a generative AI model. Specific feedback includes advice on improving performance technique and motivational messages.

[1386] 5. Provide feedback to users

[1387] The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device. The user can then review the feedback through the "Pick Perfect" app and practice again according to the advice.

[1388] 6. Select a specific practice mode and practice

[1389] When a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style, including instructional videos on specific playing techniques and musical phrases.

[1390] 7. Creating and sharing collaborative videos

[1391] Once the user reaches a certain level of skill, they can record and upload the performance video again for re-analysis. The server then re-integrates the analysis results with the emotional data and generates a video of the user performing with their AI avatar, the guitarist and vocalist. The generated video can be viewed by the user on the app screen and shared on social media.

[1392] Specific examples

[1393] For example, let's say a user is practicing an F chord. First, they use their smartphone to record themselves playing the F chord, and then upload the recording to the server via the "Pick Perfect" app. The server receives the video, and the analysis engine analyzes the intonation and sound accuracy, and then the emotion engine detects the user's facial expressions and tone of voice.

[1394] The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going." This not only helps users improve their playing technique, but also provides emotional support.

[1395] Furthermore, if the user selects "Home Mode," a practice menu tailored to the playing style of a specific guitarist is provided. Based on this practice mode, the server provides instructional videos on specific techniques and phrases.

[1396] Once a user reaches a certain level of skill, they can record and upload another video of their performance. The server then reanalyzes the newly received video and, based on the analysis results, generates a video of the user performing with their AI guitarist and vocalist. The generated video can be viewed by the user on the app screen and easily shared on social media and other platforms.

[1397] The above system further improves the user's learning experience, allowing them to continue practicing their instrument while improving their performance skills and maintaining their motivation.

[1398] The processing flow will be explained below.

[1399] Step 1:

[1400] Users record their musical instrument performance using a smartphone or tablet. After recording is complete, users launch the "Pick Perfect" app and select the recorded video file.

[1401] Step 2:

[1402] Users upload the recorded video files to the server through the "Pick Perfect" app, which displays the upload progress and notifies users when it is complete.

[1403] Step 3:

[1404] The server receives the uploaded video file and saves it in temporary storage. After saving is complete, it automatically prepares to start the next analysis process.

[1405] Step 4:

[1406] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, rhythm, and tempo from the video data.

[1407] Step 5:

[1408] The analysis engine generates analysis results and sends them to the server, which include detailed performance data for each frame.

[1409] Step 6:

[1410] The server receives the analysis results and sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[1411] Step 7:

[1412] The emotion engine generates emotion data and sends the data to the server, which includes different emotions such as joy, sadness, surprise, and concentration.

[1413] Step 8:

[1414] The server combines the analysis results with the emotional data and uses a generative AI model to generate feedback, including advice on improving performance technique and motivational messages.

[1415] Step 9:

[1416] The server sends the generated feedback to the user's device, which can be in the form of video, images, text, or audio.

[1417] Step 10:

[1418] The user reviews the feedback through the app, follows the advice on specific areas for improvement and how to improve, and practices again.

[1419] Step 11:

[1420] When a user selects a specific practice mode, the app configures the settings and sends the information to the server, which then provides a dedicated practice menu based on the selected practice mode.

[1421] Step 12:

[1422] The user practices according to a specific practice menu, then records the performance video again and uploads it to the server via the app.

[1423] Step 13:

[1424] The server receives the newly uploaded video and analyzes it again using the analysis engine and emotion engine. Based on the results, a video of the user performing together with an AI guitarist and vocalist is generated.

[1425] Step 14:

[1426] The server generates the video and sends it to the user's device, where the user can view it on the app screen and share it on social media or other platforms.

[1427] These are the specific processing steps of the present invention. This system allows users to practice while receiving appropriate feedback, and by receiving emotional support as well as improving their performance skills, they can maintain continuous motivation.

[1428] Example 2

[1429] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1430] Existing instrument learning systems make it difficult for users to effectively improve their playing skills and maintain their motivation. Specifically, feedback is limited to improving playing technique, and comprehensive support that takes into account the user's emotional state and motivation is lacking. Furthermore, they lack the provision of personalized practice menus based on specific playing styles, making it difficult to progress efficiently.

[1431] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1432] In this invention, the server includes means for an analysis engine to extract finger movements, chord progressions, and rhythms from video data, means for an emotion analysis engine to analyze facial expressions and tone of voice to detect the user's emotional state, and means for generating feedback using a generative AI model based on the analysis results and emotion data. This allows the user to receive advice that takes into account their emotional state along with technical feedback, allowing them to efficiently progress through musical instrument learning while receiving comprehensive support.

[1433] "User" refers to an individual who utilizes the system to practice playing an instrument and receive feedback.

[1434] "Instrument" refers to a tool or device that produces music for a user to play.

[1435] "Recording" refers to recording the user's performance as video data.

[1436] "Upload" refers to the process by which a user sends recorded video data to a server.

[1437] "Server" refers to a computer system that receives video data from users, stores it, analyzes it, and generates feedback.

[1438] "Video data" refers to video files of musical instrument performances recorded by users.

[1439] An "analysis engine" refers to software or algorithms used to extract technical information about musical instrument performance (finger movements, chord progressions, rhythm, etc.) from video data.

[1440] An "emotion analysis engine" refers to software or algorithms that analyze a user's facial expressions and tone of voice from video data to detect their emotional state.

[1441] "Generative AI model" refers to an artificial intelligence model that generates feedback appropriate for the user based on data from the analysis engine and sentiment analysis engine.

[1442] "Feedback" refers to messages that are generated based on the analysis results, such as advice on improving the user's playing technique or messages that increase motivation.

[1443] "Terminal" refers to the device (smartphone, tablet, etc.) that a user uses to record and check feedback.

[1444] "Practice mode" refers to a practice menu provided by the server that the user selects based on a particular playing style.

[1445] "Practice Menu" refers to content that provides instructional videos and advice on specific performance techniques or musical phrases.

[1446] This invention relates to an online learning support system for enabling beginners to efficiently self-study musical instruments. The system includes: means for a user to record their musical instrument performance; means for the user to upload the recorded performance video; means for a server to receive the uploaded video data; means for the server to transmit the video data to an analysis engine and receive analysis results that extract finger movements, chord progressions, and rhythm; means for the server to transmit the video data to an emotion analysis engine and analyze the user's emotional state; means for the server to generate feedback using a generative AI model based on the analysis results and the emotion data; and means for the server to provide the generated feedback to the user's terminal.

[1447] The specific process of this system is as follows: The user uses a device such as a smartphone or tablet to record themselves playing an instrument. A regular camera app is used for this recording. After recording is complete, the user launches a dedicated application, selects the recorded file, and uploads it to the server. When uploading, tags and notes are entered to make the recorded data easier to manage. The server receives the uploaded video data and saves it in temporary storage. A cloud service (e.g., Amazon S3 or Google Cloud Storage) is often used as this temporary storage.

[1448] The server then sends the saved video data to an analysis engine, which uses TensorFlow, OpenCV, and other tools to extract performance information such as finger movements, chord progressions, and rhythm. The server also sends the video data to an emotion analysis engine, which analyzes the user's facial expressions and tone of voice. The emotion analysis engine uses tools such as Microsoft Azure Emotion API and IBM Watson Tone Analyzer. This identifies the user's emotional state, such as joy, sadness, surprise, or concentration.

[1449] The analysis results and emotional data are integrated on the server, and feedback is generated using a generative AI model (e.g., GPT-4). This feedback includes advice on improving performance technique and motivational messages. The generated feedback is prepared in various formats (video, image, text, audio) and sent to the user's device. The user can view this feedback through a dedicated application.

[1450] Furthermore, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on that style. The practice menu includes instructional videos on specific playing techniques and musical phrases. These videos are provided using the YouTube API and Vimeo API.

[1451] For example, if a user is practicing an F chord, they can record themselves playing the F chord using their smartphone and upload the recording to a server via a dedicated app. The server receives the video, analyzes the articulation and sound accuracy using an analysis engine (TensorFlow or OpenCV), and then detects the user's facial expression and tone of voice using an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer). The analysis results include technical feedback such as "Your finger position is not perfect, so the sound is not accurate," as well as motivational feedback such as "Your efforts are paying off, so keep going!"

[1452] Here are some example prompts for a generative AI model:

[1453] "I'm practicing an F chord. Could you please analyze the video I recorded? Give me advice on finger placement and chord progression accuracy. I'd also like some motivational messages."

[1454] "The user selected a practice mode based on a specific playing style. Please provide a practice program tailored to the style of the specific guitarist. Also provide instructional videos on specific techniques and phrases."

[1455] "The user has reached a certain level of skill. Please analyze the new performance video and generate a collaboration video. Please generate a collaboration video between the user and the AI ​​and provide it in a format that can be shared on social media, etc."

[1456] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1457] Step 1:

[1458] A user records an instrument performance

[1459] Input: A user starts playing an instrument using a smartphone or tablet.

[1460] How it works: The user uses a regular camera app to record their performance as video data.

[1461] Output: Video data of the recorded instrument performance (video file)

[1462] Step 2:

[1463] Users upload recorded performance videos

[1464] Input: Recorded performance video (video file), launching the dedicated application

[1465] How it works: After recording is complete, the user launches the dedicated application, selects the recorded file, optionally enters tags and notes, and uploads the video to the server.

[1466] Output: Uploaded performance video (video file), management information (tags, notes)

[1467] Step 3:

[1468] The server receives the uploaded video data and stores it in temporary storage.

[1469] Input: Uploaded performance video (video file), management information (tags, notes)

[1470] How it works: The server receives video data uploaded by users and temporarily stores it in cloud storage (e.g., Amazon S3 or Google Cloud Storage).

[1471] Output: Video data saved in temporary storage

[1472] Step 4:

[1473] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm.

[1474] Input: Saved video data

[1475] Operation: The server sends the video data to an analysis engine (TensorFlow or OpenCV), which analyzes the data to extract performance information such as finger movements, chord progressions, and rhythm.

[1476] Output: Performance analysis results (data on finger movements, chord progressions, rhythm, etc.)

[1477] Step 5:

[1478] The server sends the video data to the emotion analysis engine to analyze the user's emotional state.

[1479] Input: Saved video data

[1480] How it works: The server sends video data to an emotion analysis engine (Microsoft Azure Emotion API or IBM Watson Tone Analyzer), which analyzes the user's facial expressions and tone of voice to identify their emotional state (happiness, sadness, surprise, concentration, etc.).

[1481] Output: Sentiment analysis results (data about the user's emotional state)

[1482] Step 6:

[1483] The server generates feedback using a generative AI model based on the analysis results and emotion data.

[1484] Input: Performance analysis results, emotion analysis results

[1485] How it works: The server integrates the analysis results with the emotional data and uses a generative AI model (e.g., GPT-4) to generate feedback for the user (such as advice on improving playing skills or messages to motivate them).

[1486] Output: Generated feedback (technical advice, motivational messages)

[1487] Step 7:

[1488] The server provides the generated feedback to the user's device.

[1489] Input: Generated feedback (technical advice, motivational messages)

[1490] How it works: The server prepares the generated feedback in various formats (video, image, text, audio) and sends it to the user's device.

[1491] Output: Feedback is displayed in the user's dedicated application (technical advice, motivational messages)

[1492] Step 8:

[1493] The user selects a practice mode based on a specific playing style and receives a practice menu.

[1494] Input: The user selects a specific playing style in a dedicated application.

[1495] How it works: When a user selects a practice mode based on a specific playing style, the server provides instructional videos on specific playing techniques or musical phrases based on that selection. This is done using the YouTube API and Vimeo API.

[1496] Output: Practice menu (explanatory videos, technique information)

[1497] Step 9:

[1498] When a user reaches a certain level of skill, they can upload a new performance video.

[1499] Input: A device for users to re-record and upload performance videos

[1500] What it does: When a user reaches a certain skill level, it records a new performance video and uploads it to the server.

[1501] Output: Newly uploaded performance video (video file)

[1502] Step 10:

[1503] The server reanalyzes the newly received video and generates a co-video

[1504] Input: Newly uploaded performance video, analysis engine for reanalysis, and emotion analysis engine results

[1505] Operation: The server retransmits the newly received video to the analysis engine and emotion analysis engine, which reanalyzes the data. Then, based on the analysis results, it generates a video of the user performing with the AI ​​guitarist and vocalist.

[1506] Output: Collaborative video (a video of the user and their AI avatar)

[1507] Step 11:

[1508] The server provides the generated collaborative video to users, making it possible to share it.

[1509] Input: Generated collaboration video

[1510] Operation: The server sends the generated video to the user's dedicated application, allowing the user to view the video and share it on social media or other platforms.

[1511] Output: Shareable collaboration video (can be shared on social media)

[1512] (Application example 2)

[1513] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1514] Beginners learning musical instruments require accurate understanding of their progress and effective feedback when self-learning. However, conventional online learning systems often provide feedback based solely on the user's technical performance data, and lack support that takes into account the user's emotional state and motivation. The present invention aims to provide an online learning support system that enables beginners to efficiently self-learn musical instruments, improving their skills and maintaining their motivation at the same time.

[1515] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for users to upload recorded performance videos, means for analyzing the user's emotional state using an emotion engine, and means for generating feedback by integrating the analysis results and emotion data based on a generative AI model. This makes it possible to provide not only feedback based on the analysis results of the user's technical performance, but also comprehensive feedback that takes the user's emotional state into consideration in real time.

[1516] A "user" is an individual who is learning to play a musical instrument and who uses the online learning support system.

[1517] "Musical instrument" refers to tools used to play music, especially stringed or keyboard instruments such as guitars, pianos, and violins.

[1518] "Recording" refers to the act of saving a user's performance in video format on a digital device.

[1519] A "performance video" is a video file that records a user playing a musical instrument.

[1520] "Uploading" refers to the act of a user transmitting a recorded performance video to a server via the Internet.

[1521] A "server" is a computer system that receives, analyzes, and processes uploaded data.

[1522] "Video data" refers to digital data of a recorded video of a performance.

[1523] An "analysis engine" is software that analyzes video data and extracts technical information such as finger movements, chord progressions, and rhythm.

[1524] "Analysis results" refers to the technical information extracted from the video data by the analysis engine.

[1525] The "emotion engine" is software that analyzes the user's facial expressions and tone of voice to identify emotions such as joy, sadness, surprise, and concentration.

[1526] "Emotion data" is information about the user's emotional state analyzed by the emotion engine.

[1527] A "generative AI model" is an algorithm that uses generative adversarial networks (GANs) and other AI techniques to integrate analytical results with emotional data and generate feedback.

[1528] "Feedback" refers to advice and messages generated based on analysis results and emotional data that help users improve their performance skills.

[1529] A "wearable device" is a computing device that can be worn by a user, such as smart glasses or a head-mounted display.

[1530] "Real-time" refers to providing instant feedback at the moment the performance is taking place.

[1531] "Practice Mode" is a personalized practice program that a user selects based on a particular playing style or technique.

[1532] A "practice menu" is a series of practice exercises and instructional videos provided by the server according to the selected practice mode.

[1533] The present invention is an online learning support system for helping beginners to learn musical instruments efficiently. The system provides a means for users to record their musical instrument performances and upload the recorded performance videos. The uploaded video data is received by a server, where it is then processed by an analysis engine and an emotion engine. The analysis engine extracts information such as finger movements, chord progressions, and rhythm from the performance data, while the emotion engine analyzes the user's facial expressions and vocal tone to identify their emotional state.

[1534] The server integrates the performance data received from the analysis engine and the emotion data received from the emotion engine, and generates feedback using a generative AI model. This feedback includes advice on improving performance technique and messages to motivate the user.

[1535] Users can receive real-time feedback using wearable devices (e.g., smart glasses or head-mounted displays). This allows them to get feedback at the moment they are playing and immediately try to improve their technique. In addition, when a user selects a practice mode based on a specific playing style, the server provides a practice menu based on the user's selection. This practice menu includes instructional videos on specific techniques or musical phrases.

[1536] Furthermore, once the user reaches a certain level of skill, they can record another performance video, upload it to the server, and reanalyze it. The server then reintegrates the new analysis results with the emotional data to generate a video of the user performing with the virtual performer. This video can be viewed on the user's device screen and easily shared on social media.

[1537] Examples:

[1538] For example, let's say a user is practicing an "F chord." First, they use the smart glasses to record themselves playing the F chord, and then upload the recording to a server via an application. The server receives the video data, analyzes the finger movements and rhythm using an analysis engine, and then analyzes the user's facial expressions and tone of voice using an emotion engine to detect their emotional state. The analysis results include technical feedback such as "Your finger position is not accurate, so the notes are out of sync," as well as encouraging feedback such as "You're making a good learning pace, so keep going."

[1539] When a user selects the "Specific Guitarist Mode" as their practice mode, the server generates a practice menu based on this mode. The practice menu includes instructional videos on specific techniques and phrases. Once the user reaches a certain level, they record and upload the performance video again for further analysis. Based on the obtained data, the server generates a video of the user performing with a virtual musician and provides it to the user in a shareable format.

[1540] Example prompt sentence:

[1541] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[1542] This provides a system that allows users to improve their performance skills while continuing to learn while receiving emotional support.

[1543] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1544] (Processing step flow)

[1545] Step 1:

[1546] A user records their performance using a wearable device (e.g., smart glasses or a head-mounted display). The recorded video data is stored on the device.

[1547] Input: User's performance video

[1548] Output: Recorded video data

[1549] Specific behavior: The user puts on the device and starts playing, presses the record button to record the performance, and presses the button again to stop recording when finished.

[1550] Step 2:

[1551] The user uploads the recorded performance video to the server using the application. The user selects the video file in the application and presses the upload button.

[1552] Input: Recorded video data

[1553] Output: Video data uploaded to the server

[1554] Specific operation: Open the application, select the recording file, and press the upload button. The video data will be sent to the server via the Internet.

[1555] Step 3:

[1556] The server receives the uploaded video data and stores it in temporary storage.

[1557] Input: Uploaded video data

[1558] Output: Video data stored in temporary storage

[1559] Specific operation: The server receives data sent via the Internet and saves it in the specified directory.

[1560] Step 4:

[1561] The server sends the video data to an analysis engine, which extracts information such as finger movements, chord progressions, and rhythm. The analysis engine then generates the extracted technical information.

[1562] Input: Video data stored in temporary storage

[1563] Output: Technical analysis results

[1564] How it works: The server passes the video data to the analysis engine, which then uses image recognition algorithms to extract performance technique information, such as finger position and movement.

[1565] Step 5:

[1566] The server sends the video data to the emotion engine, which analyzes the user's facial expressions and tone of voice to identify their emotional state.

[1567] Input: Video data stored in temporary storage

[1568] Output: Emotion analysis results

[1569] How it works: The server passes video data to the emotion engine, which then uses emotion recognition algorithms to analyze the user's facial expressions and tone of voice.

[1570] Step 6:

[1571] The server integrates the technical analysis results from the analysis engine and the emotional analysis results from the emotion engine and generates feedback using a generative AI model.

[1572] Input: Technical analysis results, sentiment analysis results

[1573] Output: Feedback from a generative AI model

[1574] Specific operation: The server inputs the two analysis results using a generative AI model such as Python or TensorFlow and generates an appropriate feedback message.

[1575] Step 7:

[1576] The server sends the generated feedback to the user's terminal, where it is displayed in real time on the user's wearable device.

[1577] Input: Feedback from a generative AI model

[1578] Output: Feedback provided to the user

[1579] Specific operation: The server sends a feedback message in text, audio, or image format to the user's device and displays it on the device.

[1580] (Example prompt)

[1581] "Record yourself playing an instrument. Once you've finished recording, upload the data to our server to see how your playing has improved. If you have a particular playing style, please select that style. Check the feedback from the server."

[1582] This allows users to receive real-time feedback on areas to improve their playing technique and maintain their motivation, enabling effective self-learning.

[1583] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1584] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1585] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1586] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1587] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1588] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1589] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1590] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1591] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1592] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1593] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1594] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1595] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1596] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1597] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1598] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1599] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1600] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1601] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1602] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1603] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1604] The following is further disclosed regarding the above embodiment.

[1605] (Claim 1)

[1606] means for a user to record a performance of a musical instrument;

[1607] means for users to upload recorded performance videos;

[1608] A means for the server to receive the uploaded video data;

[1609] A means for the server to transmit video data to the analysis engine and receive the analysis results;

[1610] a means for the server to generate feedback based on the analysis results;

[1611] means for the server to provide generated feedback to the user;

[1612] A system including:

[1613] (Claim 2)

[1614] 2. The system of claim 1, wherein the analysis engine includes means for extracting finger movements, chord progressions, and rhythms from the video data.

[1615] (Claim 3)

[1616] means for a user to select a practice mode based on a particular playing style;

[1617] 10. The system of claim 1, wherein the server includes means for providing a practice menu based on a selected practice mode.

[1618] "Example 1"

[1619] (Claim 1)

[1620] means for a user to record a performance of a musical instrument;

[1621] means for users to upload recorded performance videos;

[1622] A means for the server to receive the uploaded video data;

[1623] A means for the server to transmit video data to the analysis engine and receive the analysis results;

[1624] A means for the server to generate feedback using a generative AI model based on the analysis results;

[1625] a means for the user to review the feedback and select a practice mode based on a particular playing style;

[1626] a means for the server to provide a practice menu based on the selected practice mode;

[1627] A means for the server to generate a collaborative video based on the user's practice results and provide the video to the user;

[1628] A system including:

[1629] (Claim 2)

[1630] 2. The system of claim 1, wherein the analysis engine includes means for extracting finger movements, chord progressions, and rhythms from the video data.

[1631] (Claim 3)

[1632] The system according to claim 1, further comprising a means for providing the co-starring video in a format that can be shared on social media or the like.

[1633] "Application Example 1"

[1634] (Claim 1)

[1635] A means for a user to record device operations;

[1636] means for a user to upload a recorded operation video;

[1637] A means for the server to receive the uploaded video data;

[1638] A means for the server to transmit video data to the analysis engine and receive the analysis results;

[1639] a means for the server to generate feedback based on the analysis results;

[1640] means for the server to provide generated feedback to the user;

[1641] a means for a user to select a particular training mode;

[1642] A means for the server to provide a training menu based on the selected training mode;

[1643] A means for generating a video of a user's operation and an AI model together;

[1644] A system including:

[1645] (Claim 2)

[1646] The system of claim 1, wherein the analysis engine includes means for extracting operation procedures, tools used, and accuracy of actions from the video data.

[1647] (Claim 3)

[1648] means for a user to select a practice mode based on a particular skill level;

[1649] 10. The system of claim 1, wherein the server includes means for providing educational content based on the selected practice mode.

[1650] "Example 2: Combining Emotion Engines"

[1651] (Claim 1)

[1652] means for a user to record a performance of a musical instrument;

[1653] means for users to upload recorded performance videos;

[1654] A means for the server to receive the uploaded video data;

[1655] The server sends the video data to the analysis engine and receives the analysis results that extract finger movements, chord progressions, and rhythm.

[1656] A server transmits the video data to an emotion analysis engine to analyze the user's emotional state;

[1657] A means for the server to generate feedback using a generative AI model based on the analysis results and emotion data;

[1658] means for the server to provide the generated feedback to the user's terminal;

[1659] A system including:

[1660] (Claim 2)

[1661] The analysis engine extracts finger movements, chord progressions, and rhythm from video data,

[1662] 10. The system of claim 1, wherein the emotion analysis engine includes means for analyzing facial expressions and tone of voice to detect the user's emotional state.

[1663] (Claim 3)

[1664] means for a user to select a practice mode based on a particular playing style;

[1665] a means for the server to provide a practice menu based on the selected practice mode;

[1666] 2. The system of claim 1, wherein the feedback includes advice on improving playing technique and / or motivational messages.

[1667] "Application example 2 when combining emotion engines"

[1668] (Claim 1)

[1669] means for a user to record a performance of a musical instrument;

[1670] means for users to upload recorded performance videos;

[1671] A means for the server to receive the uploaded video data;

[1672] A means for the server to transmit video data to the analysis engine and receive the analysis results;

[1673] a means for the server to generate feedback based on the analysis results;

[1674] means for using an emotion engine to analyze an emotional state of a user;

[1675] A means for generating feedback by integrating the analysis results and emotion data based on a generative AI model; and

[1676] means for the server to provide generated feedback to the user;

[1677] A system including a means for a user to receive real-time feedback using a wearable device.

[1678] (Claim 2)

[1679] 2. The system of claim 1, wherein the analysis engine includes means for extracting finger movements, chord progressions, and rhythms from the video data.

[1680] (Claim 3)

[1681] means for a user to select a practice mode based on a particular playing style;

[1682] 10. The system of claim 1, wherein the server includes means for providing a practice menu based on a selected practice mode. [Explanation of symbols]

[1683] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for a user to record a performance of a musical instrument; means for users to upload recorded performance videos; A means for the server to receive the uploaded video data; A means for the server to transmit video data to the analysis engine and receive the analysis results; a means for the server to generate feedback based on the analysis results; means for the server to provide generated feedback to the user; A system including:

2. 2. The system according to claim 1, wherein the analysis engine includes means for extracting finger movements, chord progressions, and rhythms from the video data.

3. means for a user to select a practice mode based on a particular playing style; 2. The system of claim 1, wherein the server includes means for providing a practice menu based on a selected practice mode.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A