System

A system objectively evaluates musical performance by analyzing pitch, rhythm, and dynamics, providing specific feedback to enhance learning and motivation in musical performance.

JP2026014918APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116392
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Musical performers, especially children in early stages of learning, face challenges in objectively evaluating their performance and receiving specific feedback to improve, leading to a decline in motivation due to subjective evaluation and resistance to feedback from parents or teachers.

Method used

A system that receives performance data, analyzes pitch, rhythm, and dynamics, evaluates performance expression, generates specific feedback, and transmits it visually and audibly to provide objective evaluation and improve musical performance.

Benefits of technology

Enables users to objectively assess their performance, identify specific areas for improvement, and maintain motivation through detailed and personalized feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014918000001_ABST
    Figure 2026014918000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving performance data; means for analyzing the performance data to evaluate pitch, rhythm, dynamics; means for evaluating a performance expression based on the evaluation results; means for generating feedback based on the evaluation results and the evaluation of the performance expression; and means for transmitting the feedback.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Improving musical performance requires not only accurate pitch, rhythm, and dynamic understanding, but also the ability to express a piece of music. However, it can be difficult for performers and children, especially those in their early stages of learning, to objectively evaluate their own performance and receive specific feedback to improve. Repeating the same practice routine can also be problematic, leading to a decline in motivation. Furthermore, students often resist feedback from parents or teachers, and a lack of objectivity often prevents effective instruction. This invention aims to solve these problems and provide an enjoyable and effective environment for learning music. [Means for solving the problem]

[0005] This invention relates to a system that receives performance data and evaluates pitch, rhythm, and dynamics. Specifically, the system includes a means for receiving performance data, a means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, a means for evaluating performance expression based on the evaluation results, a means for generating feedback based on the evaluation results and the evaluation of performance expression, and a means for transmitting the feedback. This system allows users to objectively evaluate their own performance and identify specific areas for improvement. Furthermore, the feedback is provided visually and audibly, making it easy to understand and contributing to maintaining motivation.

[0006] "Performance data" is data that includes audio information of the music played by the user.

[0007] "Means" are devices or software elements that allow the system to realize specific functions or operations.

[0008] "Analysis" is the process of conducting a detailed analysis of performance data and evaluating pitch, rhythm, and dynamics.

[0009] "Pitch" is a measure of pitch and is an important factor in evaluating accurate performance.

[0010] "Rhythm" is an element that evaluates the temporal placement and pattern of sounds and checks whether the performance is in tune with the rhythm of the music.

[0011] "Dynamics" is an element used to evaluate changes in volume and judge the dynamics of a performance.

[0012] "Evaluation" is the process of assigning scores and comments to each element of the performance based on the analysis results.

[0013] "Performance expression" is an element used to evaluate the emotions and expressive intentions that a performer puts into a piece of music.

[0014] "Feedback" is information including suggestions for improvement and advice provided to the user based on the analysis results of the performance data and the evaluation of the performance expression.

[0015] "Sending" is the process of transmitting the generated feedback to the user's terminal. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance. The system's purpose is to objectively evaluate the music performed by the user and provide specific feedback. Characteristic elements of this system include receiving performance data, analyzing the performance data (evaluating pitch, rhythm, and dynamics), evaluating performance expression, and generating and transmitting feedback.

[0038] overview

[0039] Users record their performances using devices such as smartphones or tablets and send the audio data to the system. The server analyzes the performance data, evaluating pitch, rhythm, and dynamics, and also assesses the expressiveness of the performance using a large-scale language model. Specific feedback is then generated and provided to the user based on the evaluation results.

[0040] User-side processing

[0041] 1. Recording your performance

[0042] The user can record their own performance using the device's recording function, and start and stop recording by user operation.

[0043] 2. Sending recording data

[0044] Once the recording is complete, the device will automatically send the audio data to the server, or the user can manually send the audio.

[0045] Server-side processing

[0046] 1. Receiving music data

[0047] The server receives the voice data sent from the user's device and stores it in an appropriate format for subsequent analysis.

[0048] 2. Pitch, rhythm, and dynamic analysis

[0049] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0050] 3. Evaluation of performance expression

[0051] The server uses a large-scale language model to evaluate the expressiveness of the performance, analyzing it based on specific instructions (e.g., "play it happily" or "let it rain") and generating text comments as a result.

[0052] 4. Generate feedback

[0053] The server generates comprehensive feedback based on the analysis results, including evaluations of pitch, rhythm, and dynamics, as well as comments about the performance expression.

[0054] 5. Submitting Feedback

[0055] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0056] Specific examples

[0057] As a concrete example, consider a user playing a piano piece. First, the user records their performance on their smartphone and sends the audio data to the system. The server receives the audio data and then evaluates the accuracy of pitch (whether each note is played at the correct pitch in the score), rhythm (whether the tempo is correct), and dynamics. If the user attempts to express their performance in a "playful" way, this expressiveness is also analyzed using a large-scale language model. Finally, specific feedback is generated and sent to the user, along with the evaluation results of pitch, rhythm, and dynamics, such as "Overall, your performance sounds playful, but some phrases lack dynamic control." This feedback allows the user to specifically understand areas for improvement in their performance and use this information for their next practice session.

[0058] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0059] The processing flow will be explained below.

[0060] User-side processing

[0061] Step 1:

[0062] The user launches the recording app on their device and taps the button to start playing.

[0063] Your device will begin recording your performance.

[0064] Step 2:

[0065] The user finishes playing.

[0066] The device stops recording and saves the audio data.

[0067] Step 3:

[0068] The user taps the send button for the recording.

[0069] The device transmits the stored voice data to the server.

[0070] Server-side processing

[0071] Step 1:

[0072] The server receives the voice data transmitted from the terminal.

[0073] Store the received data in the appropriate format.

[0074] Step 2:

[0075] The server analyzes the audio data and evaluates the pitch.

[0076] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[0077] Step 3:

[0078] The server analyzes the audio data for each rhythm pattern.

[0079] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[0080] Step 4:

[0081] The server analyzes the volume of the sound from the audio data.

[0082] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[0083] Step 5:

[0084] The server evaluates the performance expression using a large-scale language model.

[0085] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[0086] Step 6:

[0087] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, and performance expression.

[0088] Feedback includes a numerical rating, a graphical display, and text comments.

[0089] Step 7:

[0090] The server transmits the generated feedback data to the terminal.

[0091] Log the completion of the transmission.

[0092] Terminal side processing

[0093] Step 1:

[0094] The terminal receives the feedback data sent from the server.

[0095] Verify the integrity of the received data and check for any invalid data.

[0096] Step 2:

[0097] The terminal analyzes the feedback data.

[0098] The analysis results are displayed visually and audibly to the user.

[0099] Specific examples

[0100] Take the example of a user playing a piano piece. The user launches a recording app on their smartphone, starts playing, and records. When recording is finished, the audio data is sent to the server. The server analyzes the received audio data and evaluates pitch, rhythm, and dynamics. At the same time, it evaluates the performance expression using a large-scale language model. Feedback is then generated based on the evaluation results and sent to the user's device. The user receives the feedback on their device, understands which parts of their performance need improvement, and can use this information to improve their next practice.

[0101] Example 1

[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0103] Conventional systems supporting the improvement of musical performance lack sufficient analysis of performance data, making it difficult to provide specific feedback. Furthermore, evaluation of expressiveness is subjective, making it difficult to provide specific advice. As a result, it is difficult for users to accurately grasp areas for improvement in their own performance.

[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0105] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, and means for transmitting the feedback, thereby enabling detailed analysis of the user's performance data and providing specific and objective feedback.

[0106] "Performance data" is audio information generated during musical performance.

[0107] "Means for receiving" refers to a device or method for receiving performance data from an external source.

[0108] "Means for analysis" refers to technology for breaking down received performance data into elements such as pitch, rhythm, and dynamics and evaluating them.

[0109] "Means for evaluation" refers to a device or method for determining the quality of a performance based on the analysis results.

[0110] "Performance expression" is a method of musical expression that reflects the performer's emotions and nuances.

[0111] A "large-scale language model" is a type of artificial intelligence that is trained based on large amounts of data in natural language processing.

[0112] A "prompt" is a sentence used to give instructions to an AI model.

[0113] "Feedback" refers to specific advice or comments provided to the user based on the results of analysis and evaluation.

[0114] The "means for transmitting" refers to a device or method for sending the generated feedback to the user's terminal.

[0115] "Visual display means" refers to a device or method that displays feedback on a screen in the form of text or graphs.

[0116] "Auditory display means" refers to a device or method for audibly conveying feedback to the user.

[0117] "Karaoke scoring technology" is a technique for evaluating musical performances by comparing them with standard pitch and rhythm.

[0118] The present invention relates to a system that efficiently supports the improvement of musical performance. The purpose of this system is to objectively evaluate the music performed by a user and provide specific feedback. Detailed embodiments of this system are described below.

[0119] User-side processing

[0120] Users record their performances using devices such as smartphones and tablets. Recording can be started and stopped by user operation. For example, a user can record a performance using a voice recorder app on their smartphone. After recording is finished, the device automatically sends the audio data to the server. If the user has disabled automatic sending, they can also select the recorded file and send it manually.

[0121] Server-side processing

[0122] The server receives the voice data sent from the user's device. This data is saved in an appropriate format (e.g., WAV format). Specifically, the data is received using the HTTP protocol or the FTP protocol.

[0123] The server then analyzes the received audio data using karaoke scoring technology. For pitch analysis, it uses the Python music analysis library "LibROSA," comparing each note with a reference pitch. For rhythm analysis, it uses "pYin," an audio pitch detector tool, to confirm whether it matches the original pattern of the song. For dynamic analysis, it evaluates the amplitude of the audio data.

[0124] The performance expression is evaluated using a large-scale language model. Specifically, "GPT-4" is used, and the prompt is "Please rate the expressiveness of this piano piece played with enjoyment." Based on this, the large-scale language model analyzes the expressiveness of the performance and generates a text comment.

[0125] Generate and send feedback

[0126] The server generates feedback by integrating the results of pitch, rhythm, and dynamic analysis with the evaluation of performance expression using a large-scale language model. The generated feedback is designed to include specific advice. For example, it might say, "Overall, your playing is enjoyable, but some phrases lack dynamic control. Your tempo is off in some places, so we recommend practicing with a metronome."

[0127] Finally, the server sends the generated feedback to the user's device. The feedback is provided both visually and audibly, and the user can check the feedback within the app. This allows the user to specifically understand areas for improvement in their performance and use this information to improve their next practice.

[0128] Specific examples

[0129] For example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on their smartphone. The recording is automatically sent to the server. The server analyzes the received data and checks the accuracy of pitch and rhythm using the "LibROSA" library and the "pYin" tool. Furthermore, to evaluate the performance expression, the server inputs a prompt sentence into the large-scale language model "GPT-4" saying, "Please rate the expressiveness of this enjoyable piano performance," and generates comments about the expressiveness. Based on the analysis results, the server generates feedback such as, "Overall, the performance is enjoyable, but there is a lack of dynamic control in some phrases," and sends it to the user.

[0130] This allows the user to understand specific areas for improvement and use this information in their next practice session.

[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0132] Step 1: Record your performance

[0133] Users use the recording function of their smartphones or tablets to record their own performances. The input is the music played by the user, and the output is audio data (e.g., a WAV file). Specifically, recording begins when the user presses the "Start Recording" button on the device, and ends when the user presses the "Stop Recording" button.

[0134] Step 2: Send your recordings

[0135] The device sends the audio data to the server after recording is complete. The input is the recorded audio data, and the output is a notification that transmission is complete. Specifically, the device uploads the data to the server using the HTTP or FTP protocol. If the user has disabled automatic transmission, they can select the recorded file and press the "Send" button to send it manually.

[0136] Step 3: Receiving music data

[0137] The server receives the voice data sent from the user's device. The input is the voice data sent from the device, and the output is the voice data stored on the server. Specifically, the server saves the received voice data in an appropriate format (e.g., WAV format).

[0138] Step 4: Analyze pitch, rhythm and dynamics

[0139] The server analyzes audio data using karaoke scoring technology. The input is the saved audio data, and the output is the evaluation results of pitch, rhythm, and dynamics. Specifically, the server analyzes pitch using the Python music analysis library "LibROSA" and compares each note with a reference pitch. It also performs rhythm analysis using the audio pitch detector tool "pYin" to confirm whether it matches the original pattern of the song. It also evaluates the dynamics of the entire song through amplitude analysis.

[0140] Step 5: Evaluate the performance

[0141] The server evaluates the performance expression using a large-scale language model. The input is audio data and analysis results, and the output is text feedback about the performance expression. Specifically, the server inputs a prompt sentence, "Please rate the expressiveness of this piano piece played with enjoyment," into the large-scale language model (e.g., "GPT-4") and obtains the generated text comment.

[0142] Step 6: Generate feedback

[0143] The server generates feedback by integrating the analysis results and the evaluation results of performance expression. The inputs are the evaluation results of pitch, rhythm, and dynamics, and the evaluation results of performance expression, and the output is comprehensive feedback. Specifically, the server generates feedback comments containing specific advice based on these results. For example, it may generate a comment such as, "Overall, the performance appears enjoyable, but there is a lack of dynamic control in some phrases."

[0144] Step 7: Submit your feedback

[0145] The server sends the generated feedback to the user's terminal. The input is the generated feedback comment, and the output is the feedback displayed on the user's terminal. In specific operations, the server sends feedback data to the terminal so that the user can visually and audibly confirm the feedback.

[0146] Through the above steps, users can receive detailed analysis results of their own performance and specific advice for improvement, which can be used to improve their performance skills.

[0147] (Application example 1)

[0148] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0149] Current factory robots lack systems that evaluate the accuracy and efficiency of their movements and provide feedback. This leads to accumulated inaccuracies and inefficiencies in robot movements, resulting in reduced productivity. Especially for robots performing complex tasks, even minute errors in movement can cause serious problems, and specific and effective measures to resolve this are needed.

[0150] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0151] In this invention, the server includes means for receiving performance data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, means for transmitting the feedback, means for acquiring and transmitting robot movement data, means for analyzing the movement data to evaluate accuracy and efficiency, means for generating movement improvement feedback based on the analysis results, and means for transmitting the feedback to the robot and the operator terminal. This makes it possible to evaluate the accuracy and efficiency of the robot's movements in real time and provide specific and effective feedback.

[0152] "Performance data" is audio information of the music played by the user.

[0153] "Means for receiving" refers to a device or module for receiving specific data from the outside and transferring it to the inside.

[0154] "Means for analyzing" refers to devices or modules that process and analyze the received data and evaluate and extract information from it.

[0155] "Means for evaluation" refers to a device or module that judges the performance and accuracy of the analyzed data.

[0156] A "means for generating feedback" is a device or module that provides improvements and advice to users and systems based on the evaluation results.

[0157] A "transmitting means" is a device or module for sending internal data to an external terminal or system.

[0158] "Robot operation data" refers to information relating to the operation of a robot operating in a factory range, such as controlling the position, speed, and force.

[0159] The "means for acquiring and transmitting operation data" refers to a device or module that collects the operation data of the robot and transmits it to a server or the like.

[0160] The "means for evaluating accuracy and efficiency" is a device or module for analyzing the robot's motion data and determining how accurate and efficient the motion is.

[0161] The "means for generating motion improvement feedback" is a device or module that presents improvements to the robot's motion based on the evaluation results of the motion data.

[0162] An "operator terminal" is a computer or mobile device used by a worker in a factory to receive feedback and operational information from the robot.

[0163] This invention relates to a system called "Robot Work Buddy" that accurately evaluates the behavior of robots used in factories and provides efficient feedback. The system collects and analyzes robot behavior data, generates specific improvement advice, and sends it to the robot and its operator.

[0164] Hardware used

[0165] Factory robots: Equipped with various sensors (position, speed, force, etc.) to collect operational data.

[0166] Server: A central processing unit for analyzing received data. It uses Python's SciPy and TensorFlow.

[0167] Operator terminal: A computer or mobile device on which an operator receives feedback.

[0168] Software used

[0169] Data collection module: Has the function of collecting robot operation data and sending it to the server.

[0170] Dynamic control algorithm: Analyzes the robot's motion data and compares it with ideal motion patterns.

[0171] Machine learning algorithms: Using TensorFlow and other algorithms, patterns of motion data are learned to improve accuracy.

[0172] Evaluation and feedback generation module: Generates specific feedback based on the analysis results and sends it to the robot or operator.

[0173] Specific example of operation procedure

[0174] 1. Acquisition and transmission of motion data

[0175] First, robots operating in factories collect data on their position, speed, and force for each task. This data is acquired through sensors built into the robots and then sent to a server.

[0176] 2. Data Analysis

[0177] The server analyzes the received motion data. It uses Python's SciPy to calculate errors in the position data and TensorFlow to learn the motion patterns. Specifically, it compares them with ideal motion patterns and identifies where there are errors.

[0178] 3. Evaluation and feedback generation

[0179] Based on the analysis results, the evaluation and feedback generation module evaluates the accuracy and efficiency of the robot's movements. Based on the evaluation results, it generates specific feedback to help improve the robot's movements. For example, it generates a text comment such as, "The position was off during the bolt tightening operation, so please correct coordinate X."

[0180] 4. Submitting Feedback

[0181] The generated feedback is sent to the robot's control system and to the operator's terminal, where it can be visually confirmed using a visualization tool (e.g., Grafana).

[0182] Adding specific examples

[0183] When a factory robot tightens a bolt, it uses sensors to collect real-time data on the position, speed, and tightening force of each bolt. This data is then sent to a server, which analyzes it using SciPy and TensorFlow. Based on the analysis, specific feedback is generated, such as "The position was off, so please correct coordinate X next time." This feedback is sent to the robot's control system and operator terminal and reflected in the next task.

[0184] Prompt Sentence Examples

[0185] Use a generative AI model to generate feedback using the following prompt:

[0186] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0187] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0188] Step 1:

[0189] The robot acquires operational data for each task. Specifically, it measures data such as position, speed, and force using internal sensors. This data is collected for each task and temporarily stored in the robot's internal memory.

[0190] (Input: Robot movement, Output: Position data, velocity data, force data)

[0191] Step 2:

[0192] The robot sends the collected operation data to a server using the HTTP protocol, and the data stored in the robot's internal memory is uploaded to the server.

[0193] (Input: position data, velocity data, force data, output: motion data sent to the server)

[0194] Step 3:

[0195] The server analyzes the received motion data, calculates the error in the position data using Python's SciPy, and uses TensorFlow to learn and compare motion patterns. Specifically, it compares the ideal motion pattern with the actual data and identifies where the error exists.

[0196] (Input: movement data, output: position data error, comparison result with movement pattern)

[0197] Step 4:

[0198] The server generates evaluation and feedback based on the analysis results. Using the evaluation and feedback generation module, it evaluates the accuracy and efficiency of the action and generates specific advice for improvement. For example, it generates a text comment such as "The position was off, so please correct coordinate X next time."

[0199] (Input: position data error, comparison result with movement pattern, output: feedback)

[0200] Step 5:

[0201] The server sends the generated feedback to the robot and the operator's terminal, where it is displayed as a specific text comment and also sent to the robot's control system to be reflected in the robot's next task.

[0202] (Input: Feedback, Output: Display on operator terminal, transmission to robot control system)

[0203] Specifically, when a robot working in a factory tightens a bolt, the necessary data is acquired at each step, processed, and the results are fed back to improve the robot's behavior. An example of a generated prompt is as follows:

[0204] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0206] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. This system supports the improvement of the user's performance technique by objectively evaluating the music played by the user and providing specific feedback. It is also possible to analyze the user's emotional state and provide more personalized feedback.

[0207] overview

[0208] Users record their performance using a device such as a smartphone or tablet, and simultaneously use the device's camera and microphone to record their facial expressions and vocal tone while they play. The server analyzes the transmitted performance and emotional data, evaluating pitch, rhythm, dynamics, and performance expression, while also recognizing the user's emotional state. Based on the evaluation results, specific feedback is generated and provided to the user.

[0209] User-side processing

[0210] 1. Recording your performance

[0211] The user launches the recording app on their device and taps the button to start playing. At the same time, the device's camera and microphone are activated, recording the user's facial expressions and voice tone. Recording can be started and stopped by the user.

[0212] 2. Sending recording data and emotion data

[0213] Once the recording of the voice and emotion data is complete, the device automatically sends the data to the server, or the user can manually send it.

[0214] Server-side processing

[0215] 1. Receiving Data

[0216] The server receives the voice data and emotion data sent from the device and stores the received data in an appropriate format for subsequent analysis.

[0217] 2. Pitch, rhythm, and dynamic analysis

[0218] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0219] 3. Evaluation of performance expression

[0220] The server evaluates the performance expression using a large-scale language model. Based on the expression instructions entered by the user (e.g., "playing it happily"), it analyzes the audio data and saves the evaluation results.

[0221] 4. Emotion Analysis

[0222] The server uses an emotion engine to analyze the user's emotions. It analyzes facial expressions and vocal tone during the performance to recognize the user's emotional state. The emotion data is used as supplementary information for music evaluation.

[0223] 5. Generate feedback

[0224] The server generates feedback based on the evaluation of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, and text comments. It also includes personalized advice based on the user's emotional state.

[0225] 6. Submitting Feedback

[0226] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0227] Specific examples

[0228] As a concrete example, consider a user playing a piano piece. The user launches a recording app on their smartphone and taps the start recording button. The device's camera and microphone record the performance while capturing the user's facial expressions and vocal tone. After the performance and recording are complete, the audio and emotional data are sent to the server. The server analyzes the received data, evaluating pitch, rhythm, and dynamics, and using a large-scale language model to evaluate the performance. The emotion engine also analyzes the user's emotional state. Based on the evaluation results, specific feedback is generated and sent to the user's device, such as "Your pitch is generally accurate, but your rhythm is a little fast. Emotion analysis indicates that you seem to be enjoying your performance, but you seem a little nervous." This feedback allows the user to understand which aspects of their performance need improvement and use it to improve their next practice. The user can use the emotional state feedback to play more relaxed.

[0229] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0230] The processing flow will be explained below.

[0231] User-side processing

[0232] Step 1:

[0233] The user launches the recording app on their device and taps the button to start playing.

[0234] As the device begins recording your performance, the camera and microphone begin capturing your facial expressions and tone of voice.

[0235] Step 2:

[0236] The user finishes playing and taps the stop recording button.

[0237] The device stops recording and saves the performance data and emotional data.

[0238] Step 3:

[0239] The user taps the send button for the recording.

[0240] The device transmits the stored voice data and emotion data to the server.

[0241] Server-side processing

[0242] Step 1:

[0243] The server receives the voice data and emotion data transmitted from the terminal.

[0244] The received data is stored in an appropriate format for analysis.

[0245] Step 2:

[0246] The server analyzes the audio data and evaluates the pitch.

[0247] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[0248] Step 3:

[0249] The server analyzes the audio data for each rhythm pattern.

[0250] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[0251] Step 4:

[0252] The server analyzes the volume of the sound from the audio data.

[0253] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[0254] Step 5:

[0255] The server evaluates the performance expression using a large-scale language model.

[0256] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[0257] Step 6:

[0258] The server uses an emotion engine to analyze the user's emotions.

[0259] The system analyzes facial expressions and vocal tone during performance, recognizes the user's emotional state, and saves the evaluation results.

[0260] Step 7:

[0261] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0262] Feedback includes numerical ratings, graphical displays, and text comments, as well as personalized advice based on the user's emotional state.

[0263] Step 8:

[0264] The server transmits the generated feedback data to the user's terminal.

[0265] Log the completion of the transmission.

[0266] Terminal side processing

[0267] Step 1:

[0268] The terminal receives the feedback data sent from the server.

[0269] Verify the integrity of the received data and check for any invalid data.

[0270] Step 2:

[0271] The terminal analyzes the feedback data.

[0272] The analysis results are displayed visually and audibly to the user.

[0273] Specific examples

[0274] Let's take the example of a user playing a piano piece. First, the user launches the recording app on their smartphone and taps the start recording button. The device starts recording, and at the same time, the camera and microphone begin to capture the user's facial expressions and vocal tone. When the performance is finished, the user taps the stop recording button, and the recording data and emotional data are saved. Then, the user taps the send button for the recording data, and the voice data and emotional data are sent to the server.

[0275] The server receives the audio data and emotion data, analyzing pitch, rhythm, dynamics, and performance expression, respectively. At the same time, the emotion engine analyzes the user's facial expressions and vocal tone to evaluate their emotional state. Based on the evaluation results, specific feedback such as "The pitch is good, but the rhythm is a little fast. The performance seems enjoyable, but there is also a sense of tension" is generated and sent to the user's device. The user can receive the feedback on their device, understand where their performance needs improvement, and use the feedback to improve their next practice.

[0276] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0277] Example 2

[0278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0279] Conventional music performance improvement support systems have difficulty not only evaluating a user's performance technique but also providing personalized feedback based on the user's emotional state. Furthermore, in order to improve the accuracy of the evaluation results, they lack the functionality to evaluate performance expression using large-scale language models and analyze the user's emotional state.

[0280] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the user's performance, means for receiving performance data and emotional data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the performance data, means for analyzing the emotional data to evaluate the user's emotional state, means for generating feedback based on the evaluation result and the emotional state, and means for transmitting the feedback. This makes it possible to evaluate the user's performance technique as well as provide detailed and personalized feedback based on the user's emotional state.

[0281] "Means for recording a user's performance" refers to a device or software that digitally records the audio of a user playing music.

[0282] "Means for receiving performance data and emotional data" refers to a function for receiving audio data and emotional data sent from a user's terminal via a network to a device such as a server.

[0283] "Means for analyzing performance data to evaluate pitch, rhythm, and dynamics" refers to algorithms and software that automatically analyze the basic elements of music—pitch, rhythm, and dynamics—and evaluate their accuracy and expressiveness.

[0284] "Means for evaluating performance expression based on performance data" refers to technologies such as large-scale language models for analyzing and evaluating the expressive intent behind a user's performance.

[0285] "Means for analyzing emotional data and assessing a user's emotional state" refers to an engine or software that identifies the user's emotional state from the tone of voice and facial expressions in a recording, and analyzes and assesses it.

[0286] "Means for generating feedback based on evaluation results and emotional state" refers to systems and algorithms for generating feedback to be provided to a user based on the analyzed performance data and emotional data.

[0287] "Means for sending feedback" refers to communication functions and protocols for sending the generated feedback data to the user's terminal.

[0288] The present invention is a system that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. The system configuration includes a user terminal, a server, and a communication function for exchanging data between them. Specific embodiments of the system are described below.

[0289] Users record their performances using devices such as smartphones or tablets. When they launch the recording app and tap the start button, the device's camera and microphone are activated. This allows the user's performance audio, as well as their facial expressions and vocal tone, to be recorded. Recording can be started and stopped by the user.

[0290] Once recording is complete, the device sends the voice data and emotion data to the server. The server receives the data and stores it in the appropriate format. The server then analyzes the received voice data in terms of pitch, rhythm, and dynamics. Karaoke scoring technology (e.g., JOYSOUND's technology) is used for this analysis. Pitch analysis uses a MIDI analysis library, and analysis is performed for each note. Rhythm analysis is performed using a tempo analysis library, and dynamics analysis is performed using waveform analysis.

[0291] Furthermore, the server evaluates the performance expression using a large-scale language model (e.g., OpenAI's GPT-4). Based on the expression instructions specified by the user (e.g., "enjoyably"), the audio data is converted into text and analyzed by the model. The evaluation results of the performance expression are saved.

[0292] In addition, the server uses an emotion engine to analyze the user's emotional state while playing. It uses facial and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. Emotional data is used as supplementary information for music evaluation.

[0293] Based on the evaluation results, the server generates specific feedback. The feedback may include numerical evaluations, graph displays, text comments, etc. In addition, personalized advice based on the user's emotional state may also be provided. The generated feedback is sent from the server to the device. The device displays the sent feedback, providing it visually and audibly. This makes it easier for the user to understand the evaluation results of their performance. Specific examples of feedback that may be considered include the following:

[0294] "The pitch is generally accurate, but the rhythm is a little fast. Emotional analysis indicates that the performance is enjoyable, but with some tension."

[0295] Through this feedback, users can understand which parts of their performance need improvement and use that information for their next practice. Users can also use the feedback on their emotional state to help them play more relaxed.

[0296] An example of a prompt sentence that could be input to a generative AI model would be, "Please rate my piano playing."

[0297] This system allows users to effectively improve their musical performance skills and continue practicing while maintaining motivation and having fun.

[0298] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0299] Step 1:

[0300] The user launches the recording app on their device and taps the recording preparation button. This action causes the device to check the status of the camera and microphone and check that they are working properly. It notifies the user that they are ready. The input is the user's operation, and the output is a notification that the camera and microphone are ready to start.

[0301] Step 2:

[0302] The user taps the "Start Recording" button to begin playing. At this point, the device simultaneously activates its camera and microphone, recording the user's performance, facial expressions, and vocal tone. The input is the "Start Recording" button and the user's performance, and the output is the recorded audio and video data.

[0303] Step 3:

[0304] When the recording is finished, the user taps the "Stop Recording" button. The device responds by stopping the recording and saving the data. The input is the "Stop Recording" operation, and the output is the saved audio and video data.

[0305] Step 4:

[0306] The device prepares to send the recorded data (audio files) and emotion data (video files and audio analysis results) to the server. When the user presses the "Send" button in the confirmation dialog, the data is sent to the server. The input is the user's confirmation of sending, and the output is the data being sent to the server.

[0307] Step 5:

[0308] The server receives the voice data and emotion data sent from the device and stores them in the respective databases. The input is the received data, and the output is the stored data.

[0309] Step 6:

[0310] The server analyzes the audio data and evaluates pitch, rhythm, and dynamics. It uses a MIDI analysis library for pitch analysis, a tempo analysis library for rhythm analysis, and waveform analysis for dynamics analysis. The input is the audio data, and the output is the analysis results for each evaluation item.

[0311] Step 7:

[0312] The server evaluates performance expressions using a large-scale language model. Based on the user's instructions, the audio data is converted into text, which is then analyzed by the large-scale language model. The input is the audio data and instructions, and the output is the evaluation result of the performance expression.

[0313] Step 8:

[0314] The server uses an emotion engine to analyze the user's emotional state. It uses face and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. The input is video and audio data, and the output is the emotion analysis results.

[0315] Step 9:

[0316] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, text comments, and personalized advice. The input is each evaluation result, and the output is feedback data.

[0317] Step 10:

[0318] The server sends the generated feedback data to the terminal. The input is the feedback data, and the output is the data sent to the user's terminal.

[0319] Step 11:

[0320] The device receives the transmitted feedback data and provides it to the user visually and audibly, allowing the user to understand the feedback and use it for their next practice. The input is the feedback data, and the output is the provision of feedback to the user.

[0321] (Application example 2)

[0322] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0323] Conventional systems for supporting the improvement of musical performance focus on technical evaluations such as pitch, rhythm, and dynamics, but have the problem of being unable to provide feedback that takes into account the performer's emotional state. Another problem is that the feedback users receive is uniform for each individual performance, making it difficult to obtain specific and personalized advice that adapts to the user's emotional state. This makes it difficult for users to understand areas for improvement that correspond to their own emotional state, resulting in the problem of inefficient improvement of performance techniques.

[0324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0325] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for analyzing the user's emotional state, means for personalizing feedback based on the emotional state, means for generating the feedback, and means for transmitting the feedback, thereby enabling the provision of specific and personalized feedback that takes into account not only technical evaluation but also the user's emotional state.

[0326] The "means for receiving performance data" is a mechanism for recording the music played by the user and transmitting that data to the server.

[0327] The "means for evaluating pitch" is a mechanism that analyzes the pitch of the performance data and measures how well it matches a reference pitch.

[0328] The "means for evaluating rhythm" is a mechanism for analyzing the timing of the sounds in the performance data and measuring the degree of agreement with a specified rhythm pattern.

[0329] The "means for evaluating dynamics" is a system that analyzes the changes in volume of performance data and evaluates the dynamics of the entire performance.

[0330] The "means for evaluating performance expression" is a system for evaluating the expressiveness and technique of a user's performance based on the evaluation results of pitch, rhythm, and dynamics.

[0331] The "means for analyzing emotional state" is a mechanism for analyzing the user's facial expressions and vocal tone while playing and recognizing the emotional state at that time.

[0332] The "means for generating feedback" is a mechanism that provides the user with specific and personalized improvements based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0333] The "means for transmitting feedback" is a mechanism for transmitting the generated feedback to the user's terminal and displaying it visually and audibly.

[0334] Overall system overview

[0335] This invention is a system to help users improve their musical performance by recording their performance and analyzing their emotions through recording their facial expressions and vocal timbre. It is realized using the following hardware and software.

[0336] Hardware and software used

[0337] Smartphone or tablet (e.g. iPhone, Android device)

[0338] PC or server (e.g. cloud server)

[0339] Devices with built-in cameras and microphones

[0340] Programming language: Swift (iOS), Kotlin (Android)

[0341] API: Google Cloud Speech-to-Text, Google Vision API (sentiment analysis)

[0342] Backend: Node.js

[0343] Database: Firebase

[0344] Main features of the system

[0345] Users first record their performance using an application installed on their smartphone or tablet. During recording, the device's camera and microphone simultaneously operate to collect the user's facial expressions and vocal tone. The collected data is then sent to a cloud server in real time or after recording.

[0346] Music evaluation and emotion analysis process

[0347] The server receives the transmitted performance data and evaluates the pitch, rhythm, and dynamics. Pitch evaluation involves analyzing how closely the pitch of each note matches the reference pitch. Rhythm evaluation involves checking synchronization with the specified pattern. Dynamic evaluation involves analyzing the dynamics of the entire performance. In addition to these technical evaluations, the server uses the Google Vision API to analyze facial expression data during performance and recognize the user's emotional state.

[0348] Generating and Providing Feedback

[0349] Based on the evaluation results, specific and personalized feedback is generated for the user, including evaluation results on pitch, rhythm, and dynamics, as well as advice based on emotional state. The generated feedback is sent to the user's device in the form of visual and audio.

[0350] Examples and prompts

[0351] As a concrete example, let's consider a situation where a user plays the guitar and then receives feedback from an app. First, the user launches the app and taps the start recording button. After finishing playing, the user taps the stop recording button, which sends the data to the server. After analysis is performed on the server side, the following feedback is displayed on the user's device:

[0352] "The pitch is generally accurate, but the rhythm is a little fast. The emotional analysis shows that you are playing with enjoyment, but there are some parts that seem tense. Please try playing a little more relaxed."

[0353] Example prompt to launch the app:

[0354] "Select your instrument and start recording. Relax and play, as your facial expressions will be recorded."

[0355] In this way, the system of the present invention allows users to understand areas for improvement in their performance, both technically and emotionally, through specific and personalized feedback.

[0356] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0357] Step 1:

[0358] Recording of performance

[0359] The user launches the application on their smartphone or tablet and taps the recording start button.

[0360] Input: User's music, facial expressions, and vocal tone

[0361] How it works: The device's microphone records the music, and the camera records the user's facial expressions while they play. The audio and facial expression data are then stored on the device.

[0362] Output: Recorded performance data and emotional data (facial expressions and vocal timbre)

[0363] Step 2:

[0364] Sending recording data and emotional data

[0365] When the user taps the recording stop button, the recorded performance data and emotion data are sent to the server.

[0366] Input: Recorded performance data, emotional data

[0367] How it works: Data is sent from the device to a server over the internet.

[0368] Output: Performance data and emotion data stored on the server

[0369] Step 3:

[0370] Pitch, rhythm and dynamic analysis

[0371] The server evaluates the pitch, rhythm, and dynamics based on the received performance data.

[0372] Input: Received performance data

[0373] How it works: A server-based analysis program compares pitch with a reference value and analyzes rhythm and dynamics. Specifically, it analyzes audio data using the Google Cloud Speech-to-Text API.

[0374] Output: Pitch, rhythm, and dynamics evaluation results

[0375] Step 4:

[0376] Evaluation of performance expression

[0377] The server evaluates the performance expression based on the evaluation results of pitch, rhythm, and dynamics.

[0378] Input: Pitch, rhythm, and dynamic evaluation results

[0379] How it works: The performance expression evaluation algorithm in the server analyzes the data and performs a comprehensive evaluation of the performance expression.

[0380] Output: Evaluation results of performance expression

[0381] Step 5:

[0382] Emotion Analysis

[0383] The server analyzes the received emotion data and evaluates the user's emotional state.

[0384] Input: Received emotional data (facial expressions, tone of voice)

[0385] How it works: It uses the Google Vision API to analyze facial expression data and recognize the user's emotional state.

[0386] Output: Evaluation result of the user's emotional state

[0387] Step 6:

[0388] Generate feedback

[0389] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0390] Input: Pitch, rhythm, dynamics, performance expression, emotional state evaluation results

[0391] How it works: The server's feedback generator combines the results of each assessment to create specific, personalized feedback.

[0392] Output: Generated feedback data

[0393] Step 7:

[0394] Send Feedback

[0395] The server transmits the generated feedback to the user's terminal.

[0396] Input: Generated feedback data

[0397] Operation: The server sends feedback data to the terminal, which displays it visually and audibly to the user.

[0398] Output: Feedback displayed on the user's device

[0399] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0400] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0401] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0402] [Second embodiment]

[0403] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0404] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0405] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0406] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0407] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0409] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0410] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0411] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0412] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0413] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0414] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0415] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance. The system's purpose is to objectively evaluate the music performed by the user and provide specific feedback. Characteristic elements of this system include receiving performance data, analyzing the performance data (evaluating pitch, rhythm, and dynamics), evaluating performance expression, and generating and transmitting feedback.

[0416] overview

[0417] Users record their performances using devices such as smartphones or tablets and send the audio data to the system. The server analyzes the performance data, evaluating pitch, rhythm, and dynamics, and also assesses the expressiveness of the performance using a large-scale language model. Specific feedback is then generated and provided to the user based on the evaluation results.

[0418] User-side processing

[0419] 1. Recording your performance

[0420] The user can record their own performance using the device's recording function, and start and stop recording by user operation.

[0421] 2. Sending recording data

[0422] Once the recording is complete, the device will automatically send the audio data to the server, or the user can manually send the audio.

[0423] Server-side processing

[0424] 1. Receiving music data

[0425] The server receives the voice data sent from the user's device and stores it in an appropriate format for subsequent analysis.

[0426] 2. Pitch, rhythm, and dynamic analysis

[0427] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0428] 3. Evaluation of performance expression

[0429] The server uses a large-scale language model to evaluate the expressiveness of the performance, analyzing it based on specific instructions (e.g., "play it happily" or "let it rain") and generating text comments as a result.

[0430] 4. Generate feedback

[0431] The server generates comprehensive feedback based on the analysis results, including evaluations of pitch, rhythm, and dynamics, as well as comments about the performance expression.

[0432] 5. Submitting Feedback

[0433] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0434] Specific examples

[0435] As a concrete example, consider a user playing a piano piece. First, the user records their performance on their smartphone and sends the audio data to the system. The server receives the audio data and then evaluates the accuracy of pitch (whether each note is played at the correct pitch in the score), rhythm (whether the tempo is correct), and dynamics. If the user attempts to express their performance in a "playful" way, this expressiveness is also analyzed using a large-scale language model. Finally, specific feedback is generated and sent to the user, along with the evaluation results of pitch, rhythm, and dynamics, such as "Overall, your performance sounds playful, but some phrases lack dynamic control." This feedback allows the user to specifically understand areas for improvement in their performance and use this information for their next practice session.

[0436] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0437] The processing flow will be explained below.

[0438] User-side processing

[0439] Step 1:

[0440] The user launches the recording app on their device and taps the button to start playing.

[0441] Your device will begin recording your performance.

[0442] Step 2:

[0443] The user finishes playing.

[0444] The device stops recording and saves the audio data.

[0445] Step 3:

[0446] The user taps the send button for the recording.

[0447] The device transmits the stored voice data to the server.

[0448] Server-side processing

[0449] Step 1:

[0450] The server receives the voice data transmitted from the terminal.

[0451] Store the received data in the appropriate format.

[0452] Step 2:

[0453] The server analyzes the audio data and evaluates the pitch.

[0454] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[0455] Step 3:

[0456] The server analyzes the audio data for each rhythm pattern.

[0457] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[0458] Step 4:

[0459] The server analyzes the volume of the sound from the audio data.

[0460] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[0461] Step 5:

[0462] The server evaluates the performance expression using a large-scale language model.

[0463] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[0464] Step 6:

[0465] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, and performance expression.

[0466] Feedback includes a numerical rating, a graphical display, and text comments.

[0467] Step 7:

[0468] The server transmits the generated feedback data to the terminal.

[0469] Log the completion of the transmission.

[0470] Terminal side processing

[0471] Step 1:

[0472] The terminal receives the feedback data sent from the server.

[0473] Verify the integrity of the received data and check for any invalid data.

[0474] Step 2:

[0475] The terminal analyzes the feedback data.

[0476] The analysis results are displayed visually and audibly to the user.

[0477] Specific examples

[0478] Take the example of a user playing a piano piece. The user launches a recording app on their smartphone, starts playing, and records. When recording is finished, the audio data is sent to the server. The server analyzes the received audio data and evaluates pitch, rhythm, and dynamics. At the same time, it evaluates the performance expression using a large-scale language model. Feedback is then generated based on the evaluation results and sent to the user's device. The user receives the feedback on their device, understands which parts of their performance need improvement, and can use this information to improve their next practice.

[0479] Example 1

[0480] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0481] Conventional systems supporting the improvement of musical performance lack sufficient analysis of performance data, making it difficult to provide specific feedback. Furthermore, evaluation of expressiveness is subjective, making it difficult to provide specific advice. As a result, it is difficult for users to accurately grasp areas for improvement in their own performance.

[0482] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0483] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, and means for transmitting the feedback, thereby enabling detailed analysis of the user's performance data and providing specific and objective feedback.

[0484] "Performance data" is audio information generated during musical performance.

[0485] "Means for receiving" refers to a device or method for receiving performance data from an external source.

[0486] "Means for analysis" refers to technology for breaking down received performance data into elements such as pitch, rhythm, and dynamics and evaluating them.

[0487] "Means for evaluation" refers to a device or method for determining the quality of a performance based on the analysis results.

[0488] "Performance expression" is a method of musical expression that reflects the performer's emotions and nuances.

[0489] A "large-scale language model" is a type of artificial intelligence that is trained based on large amounts of data in natural language processing.

[0490] A "prompt" is a sentence used to give instructions to an AI model.

[0491] "Feedback" refers to specific advice or comments provided to the user based on the results of analysis and evaluation.

[0492] The "means for transmitting" refers to a device or method for sending the generated feedback to the user's terminal.

[0493] "Visual display means" refers to a device or method that displays feedback on a screen in the form of text or graphs.

[0494] "Auditory display means" refers to a device or method for audibly conveying feedback to the user.

[0495] "Karaoke scoring technology" is a technique for evaluating musical performances by comparing them with standard pitch and rhythm.

[0496] The present invention relates to a system that efficiently supports the improvement of musical performance. The purpose of this system is to objectively evaluate the music performed by a user and provide specific feedback. Detailed embodiments of this system are described below.

[0497] User-side processing

[0498] Users record their performances using devices such as smartphones and tablets. Recording can be started and stopped by user operation. For example, a user can record a performance using a voice recorder app on their smartphone. After recording is finished, the device automatically sends the audio data to the server. If the user has disabled automatic sending, they can also select the recorded file and send it manually.

[0499] Server-side processing

[0500] The server receives the voice data sent from the user's device. This data is saved in an appropriate format (e.g., WAV format). Specifically, the data is received using the HTTP protocol or the FTP protocol.

[0501] The server then analyzes the received audio data using karaoke scoring technology. For pitch analysis, it uses the Python music analysis library "LibROSA," comparing each note with a reference pitch. For rhythm analysis, it uses "pYin," an audio pitch detector tool, to confirm whether it matches the original pattern of the song. For dynamic analysis, it evaluates the amplitude of the audio data.

[0502] The performance expression is evaluated using a large-scale language model. Specifically, "GPT-4" is used, and the prompt is "Please rate the expressiveness of this piano piece played with enjoyment." Based on this, the large-scale language model analyzes the expressiveness of the performance and generates a text comment.

[0503] Generate and send feedback

[0504] The server generates feedback by integrating the results of pitch, rhythm, and dynamic analysis with the evaluation of performance expression using a large-scale language model. The generated feedback is designed to include specific advice. For example, it might say, "Overall, your playing is enjoyable, but some phrases lack dynamic control. Your tempo is off in some places, so we recommend practicing with a metronome."

[0505] Finally, the server sends the generated feedback to the user's device. The feedback is provided both visually and audibly, and the user can check the feedback within the app. This allows the user to specifically understand areas for improvement in their performance and use this information to improve their next practice.

[0506] Specific examples

[0507] For example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on their smartphone. The recording is automatically sent to the server. The server analyzes the received data and checks the accuracy of pitch and rhythm using the "LibROSA" library and the "pYin" tool. Furthermore, to evaluate the performance expression, the server inputs a prompt sentence into the large-scale language model "GPT-4" saying, "Please rate the expressiveness of this enjoyable piano performance," and generates comments about the expressiveness. Based on the analysis results, the server generates feedback such as, "Overall, the performance is enjoyable, but there is a lack of dynamic control in some phrases," and sends it to the user.

[0508] This allows the user to understand specific areas for improvement and use this information in their next practice session.

[0509] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0510] Step 1: Record your performance

[0511] Users use the recording function of their smartphones or tablets to record their own performances. The input is the music played by the user, and the output is audio data (e.g., a WAV file). Specifically, recording begins when the user presses the "Start Recording" button on the device, and ends when the user presses the "Stop Recording" button.

[0512] Step 2: Send your recordings

[0513] The device sends the audio data to the server after recording is complete. The input is the recorded audio data, and the output is a notification that transmission is complete. Specifically, the device uploads the data to the server using the HTTP or FTP protocol. If the user has disabled automatic transmission, they can select the recorded file and press the "Send" button to send it manually.

[0514] Step 3: Receiving music data

[0515] The server receives the voice data sent from the user's device. The input is the voice data sent from the device, and the output is the voice data stored on the server. Specifically, the server saves the received voice data in an appropriate format (e.g., WAV format).

[0516] Step 4: Analyze pitch, rhythm and dynamics

[0517] The server analyzes audio data using karaoke scoring technology. The input is the saved audio data, and the output is the evaluation results of pitch, rhythm, and dynamics. Specifically, the server analyzes pitch using the Python music analysis library "LibROSA" and compares each note with a reference pitch. It also performs rhythm analysis using the audio pitch detector tool "pYin" to confirm whether it matches the original pattern of the song. It also evaluates the dynamics of the entire song through amplitude analysis.

[0518] Step 5: Evaluate the performance

[0519] The server evaluates the performance expression using a large-scale language model. The input is audio data and analysis results, and the output is text feedback about the performance expression. Specifically, the server inputs a prompt sentence, "Please rate the expressiveness of this piano piece played with enjoyment," into the large-scale language model (e.g., "GPT-4") and obtains the generated text comment.

[0520] Step 6: Generate feedback

[0521] The server generates feedback by integrating the analysis results and the evaluation results of performance expression. The inputs are the evaluation results of pitch, rhythm, and dynamics, and the evaluation results of performance expression, and the output is comprehensive feedback. Specifically, the server generates feedback comments containing specific advice based on these results. For example, it may generate a comment such as, "Overall, the performance appears enjoyable, but there is a lack of dynamic control in some phrases."

[0522] Step 7: Submit your feedback

[0523] The server sends the generated feedback to the user's terminal. The input is the generated feedback comment, and the output is the feedback displayed on the user's terminal. In specific operations, the server sends feedback data to the terminal so that the user can visually and audibly confirm the feedback.

[0524] Through the above steps, users can receive detailed analysis results of their own performance and specific advice for improvement, which can be used to improve their performance skills.

[0525] (Application example 1)

[0526] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0527] Current factory robots lack systems that evaluate the accuracy and efficiency of their movements and provide feedback. This leads to accumulated inaccuracies and inefficiencies in robot movements, resulting in reduced productivity. Especially for robots performing complex tasks, even minute errors in movement can cause serious problems, and specific and effective measures to resolve this are needed.

[0528] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0529] In this invention, the server includes means for receiving performance data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, means for transmitting the feedback, means for acquiring and transmitting robot movement data, means for analyzing the movement data to evaluate accuracy and efficiency, means for generating movement improvement feedback based on the analysis results, and means for transmitting the feedback to the robot and the operator terminal. This makes it possible to evaluate the accuracy and efficiency of the robot's movements in real time and provide specific and effective feedback.

[0530] "Performance data" is audio information of the music played by the user.

[0531] "Means for receiving" refers to a device or module for receiving specific data from the outside and transferring it to the inside.

[0532] "Means for analyzing" refers to devices or modules that process and analyze the received data and evaluate and extract information from it.

[0533] "Means for evaluation" refers to a device or module that judges the performance and accuracy of the analyzed data.

[0534] A "means for generating feedback" is a device or module that provides improvements and advice to users and systems based on the evaluation results.

[0535] A "transmitting means" is a device or module for sending internal data to an external terminal or system.

[0536] "Robot operation data" refers to information relating to the operation of a robot operating in a factory range, such as controlling the position, speed, and force.

[0537] The "means for acquiring and transmitting operation data" refers to a device or module that collects the operation data of the robot and transmits it to a server or the like.

[0538] The "means for evaluating accuracy and efficiency" is a device or module for analyzing the robot's motion data and determining how accurate and efficient the motion is.

[0539] The "means for generating motion improvement feedback" is a device or module that presents improvements to the robot's motion based on the evaluation results of the motion data.

[0540] An "operator terminal" is a computer or mobile device used by a worker in a factory to receive feedback and operational information from the robot.

[0541] This invention relates to a system called "Robot Work Buddy" that accurately evaluates the behavior of robots used in factories and provides efficient feedback. The system collects and analyzes robot behavior data, generates specific improvement advice, and sends it to the robot and its operator.

[0542] Hardware used

[0543] Factory robots: Equipped with various sensors (position, speed, force, etc.) to collect operational data.

[0544] Server: A central processing unit for analyzing received data. It uses Python's SciPy and TensorFlow.

[0545] Operator terminal: A computer or mobile device on which an operator receives feedback.

[0546] Software used

[0547] Data collection module: Has the function of collecting robot operation data and sending it to the server.

[0548] Dynamic control algorithm: Analyzes the robot's motion data and compares it with ideal motion patterns.

[0549] Machine learning algorithms: Using TensorFlow and other algorithms, patterns of motion data are learned to improve accuracy.

[0550] Evaluation and feedback generation module: Generates specific feedback based on the analysis results and sends it to the robot or operator.

[0551] Specific example of operation procedure

[0552] 1. Acquisition and transmission of motion data

[0553] First, robots operating in factories collect data on their position, speed, and force for each task. This data is acquired through sensors built into the robots and then sent to a server.

[0554] 2. Data Analysis

[0555] The server analyzes the received motion data. It uses Python's SciPy to calculate errors in the position data and TensorFlow to learn the motion patterns. Specifically, it compares them with ideal motion patterns and identifies where there are errors.

[0556] 3. Evaluation and feedback generation

[0557] Based on the analysis results, the evaluation and feedback generation module evaluates the accuracy and efficiency of the robot's movements. Based on the evaluation results, it generates specific feedback to help improve the robot's movements. For example, it generates a text comment such as, "The position was off during the bolt tightening operation, so please correct coordinate X."

[0558] 4. Submitting Feedback

[0559] The generated feedback is sent to the robot's control system and to the operator's terminal, where it can be visually confirmed using a visualization tool (e.g., Grafana).

[0560] Adding specific examples

[0561] When a factory robot tightens a bolt, it uses sensors to collect real-time data on the position, speed, and tightening force of each bolt. This data is then sent to a server, which analyzes it using SciPy and TensorFlow. Based on the analysis, specific feedback is generated, such as "The position was off, so please correct coordinate X next time." This feedback is sent to the robot's control system and operator terminal and reflected in the next task.

[0562] Prompt Sentence Examples

[0563] Use a generative AI model to generate feedback using the following prompt:

[0564] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0565] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0566] Step 1:

[0567] The robot acquires operational data for each task. Specifically, it measures data such as position, speed, and force using internal sensors. This data is collected for each task and temporarily stored in the robot's internal memory.

[0568] (Input: Robot movement, Output: Position data, velocity data, force data)

[0569] Step 2:

[0570] The robot sends the collected operation data to a server using the HTTP protocol, and the data stored in the robot's internal memory is uploaded to the server.

[0571] (Input: position data, velocity data, force data, output: motion data sent to the server)

[0572] Step 3:

[0573] The server analyzes the received motion data, calculates the error in the position data using Python's SciPy, and uses TensorFlow to learn and compare motion patterns. Specifically, it compares the ideal motion pattern with the actual data and identifies where the error exists.

[0574] (Input: movement data, output: position data error, comparison result with movement pattern)

[0575] Step 4:

[0576] The server generates evaluation and feedback based on the analysis results. Using the evaluation and feedback generation module, it evaluates the accuracy and efficiency of the action and generates specific advice for improvement. For example, it generates a text comment such as "The position was off, so please correct coordinate X next time."

[0577] (Input: position data error, comparison result with movement pattern, output: feedback)

[0578] Step 5:

[0579] The server sends the generated feedback to the robot and the operator's terminal, where it is displayed as a specific text comment and also sent to the robot's control system to be reflected in the robot's next task.

[0580] (Input: Feedback, Output: Display on operator terminal, transmission to robot control system)

[0581] Specifically, when a robot working in a factory tightens a bolt, the necessary data is acquired at each step, processed, and the results are fed back to improve the robot's behavior. An example of a generated prompt is as follows:

[0582] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0583] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0584] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. This system supports the improvement of the user's performance technique by objectively evaluating the music played by the user and providing specific feedback. It is also possible to analyze the user's emotional state and provide more personalized feedback.

[0585] overview

[0586] Users record their performance using a device such as a smartphone or tablet, and simultaneously use the device's camera and microphone to record their facial expressions and vocal tone while they play. The server analyzes the transmitted performance and emotional data, evaluating pitch, rhythm, dynamics, and performance expression, while also recognizing the user's emotional state. Based on the evaluation results, specific feedback is generated and provided to the user.

[0587] User-side processing

[0588] 1. Recording your performance

[0589] The user launches the recording app on their device and taps the button to start playing. At the same time, the device's camera and microphone are activated, recording the user's facial expressions and voice tone. Recording can be started and stopped by the user.

[0590] 2. Sending recording data and emotion data

[0591] Once the recording of the voice and emotion data is complete, the device automatically sends the data to the server, or the user can manually send it.

[0592] Server-side processing

[0593] 1. Receiving Data

[0594] The server receives the voice data and emotion data sent from the device and stores the received data in an appropriate format for subsequent analysis.

[0595] 2. Pitch, rhythm, and dynamic analysis

[0596] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0597] 3. Evaluation of performance expression

[0598] The server evaluates the performance expression using a large-scale language model. Based on the expression instructions entered by the user (e.g., "playing it happily"), it analyzes the audio data and saves the evaluation results.

[0599] 4. Emotion Analysis

[0600] The server uses an emotion engine to analyze the user's emotions. It analyzes facial expressions and vocal tone during the performance to recognize the user's emotional state. The emotion data is used as supplementary information for music evaluation.

[0601] 5. Generate feedback

[0602] The server generates feedback based on the evaluation of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, and text comments. It also includes personalized advice based on the user's emotional state.

[0603] 6. Submitting Feedback

[0604] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0605] Specific examples

[0606] As a concrete example, consider a user playing a piano piece. The user launches a recording app on their smartphone and taps the start recording button. The device's camera and microphone record the performance while capturing the user's facial expressions and vocal tone. After the performance and recording are complete, the audio and emotional data are sent to the server. The server analyzes the received data, evaluating pitch, rhythm, and dynamics, and using a large-scale language model to evaluate the performance. The emotion engine also analyzes the user's emotional state. Based on the evaluation results, specific feedback is generated and sent to the user's device, such as "Your pitch is generally accurate, but your rhythm is a little fast. Emotion analysis indicates that you seem to be enjoying your performance, but you seem a little nervous." This feedback allows the user to understand which aspects of their performance need improvement and use it to improve their next practice. The user can use the emotional state feedback to play more relaxed.

[0607] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0608] The processing flow will be explained below.

[0609] User-side processing

[0610] Step 1:

[0611] The user launches the recording app on their device and taps the button to start playing.

[0612] As the device begins recording your performance, the camera and microphone begin capturing your facial expressions and tone of voice.

[0613] Step 2:

[0614] The user finishes playing and taps the stop recording button.

[0615] The device stops recording and saves the performance data and emotional data.

[0616] Step 3:

[0617] The user taps the send button for the recording.

[0618] The device transmits the stored voice data and emotion data to the server.

[0619] Server-side processing

[0620] Step 1:

[0621] The server receives the voice data and emotion data transmitted from the terminal.

[0622] The received data is stored in an appropriate format for analysis.

[0623] Step 2:

[0624] The server analyzes the audio data and evaluates the pitch.

[0625] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[0626] Step 3:

[0627] The server analyzes the audio data for each rhythm pattern.

[0628] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[0629] Step 4:

[0630] The server analyzes the volume of the sound from the audio data.

[0631] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[0632] Step 5:

[0633] The server evaluates the performance expression using a large-scale language model.

[0634] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[0635] Step 6:

[0636] The server uses an emotion engine to analyze the user's emotions.

[0637] The system analyzes facial expressions and vocal tone during performance, recognizes the user's emotional state, and saves the evaluation results.

[0638] Step 7:

[0639] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0640] Feedback includes numerical ratings, graphical displays, and text comments, as well as personalized advice based on the user's emotional state.

[0641] Step 8:

[0642] The server transmits the generated feedback data to the user's terminal.

[0643] Log the completion of the transmission.

[0644] Terminal side processing

[0645] Step 1:

[0646] The terminal receives the feedback data sent from the server.

[0647] Verify the integrity of the received data and check for any invalid data.

[0648] Step 2:

[0649] The terminal analyzes the feedback data.

[0650] The analysis results are displayed visually and audibly to the user.

[0651] Specific examples

[0652] Let's take the example of a user playing a piano piece. First, the user launches the recording app on their smartphone and taps the start recording button. The device starts recording, and at the same time, the camera and microphone begin to capture the user's facial expressions and vocal tone. When the performance is finished, the user taps the stop recording button, and the recording data and emotional data are saved. Then, the user taps the send button for the recording data, and the voice data and emotional data are sent to the server.

[0653] The server receives the audio data and emotion data, analyzing pitch, rhythm, dynamics, and performance expression, respectively. At the same time, the emotion engine analyzes the user's facial expressions and vocal tone to evaluate their emotional state. Based on the evaluation results, specific feedback such as "The pitch is good, but the rhythm is a little fast. The performance seems enjoyable, but there is also a sense of tension" is generated and sent to the user's device. The user can receive the feedback on their device, understand where their performance needs improvement, and use the feedback to improve their next practice.

[0654] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0655] Example 2

[0656] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0657] Conventional music performance improvement support systems have difficulty not only evaluating a user's performance technique but also providing personalized feedback based on the user's emotional state. Furthermore, in order to improve the accuracy of the evaluation results, they lack the functionality to evaluate performance expression using large-scale language models and analyze the user's emotional state.

[0658] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the user's performance, means for receiving performance data and emotional data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the performance data, means for analyzing the emotional data to evaluate the user's emotional state, means for generating feedback based on the evaluation result and the emotional state, and means for transmitting the feedback. This makes it possible to evaluate the user's performance technique as well as provide detailed and personalized feedback based on the user's emotional state.

[0659] "Means for recording a user's performance" refers to a device or software that digitally records the audio of a user playing music.

[0660] "Means for receiving performance data and emotional data" refers to a function for receiving audio data and emotional data sent from a user's terminal via a network to a device such as a server.

[0661] "Means for analyzing performance data to evaluate pitch, rhythm, and dynamics" refers to algorithms and software that automatically analyze the basic elements of music—pitch, rhythm, and dynamics—and evaluate their accuracy and expressiveness.

[0662] "Means for evaluating performance expression based on performance data" refers to technologies such as large-scale language models for analyzing and evaluating the expressive intent behind a user's performance.

[0663] "Means for analyzing emotional data and assessing a user's emotional state" refers to an engine or software that identifies the user's emotional state from the tone of voice and facial expressions in a recording, and analyzes and assesses it.

[0664] "Means for generating feedback based on evaluation results and emotional state" refers to systems and algorithms for generating feedback to be provided to a user based on the analyzed performance data and emotional data.

[0665] "Means for sending feedback" refers to communication functions and protocols for sending the generated feedback data to the user's terminal.

[0666] The present invention is a system that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. The system configuration includes a user terminal, a server, and a communication function for exchanging data between them. Specific embodiments of the system are described below.

[0667] Users record their performances using devices such as smartphones or tablets. When they launch the recording app and tap the start button, the device's camera and microphone are activated. This allows the user's performance audio, as well as their facial expressions and vocal tone, to be recorded. Recording can be started and stopped by the user.

[0668] Once recording is complete, the device sends the voice data and emotion data to the server. The server receives the data and stores it in the appropriate format. The server then analyzes the received voice data in terms of pitch, rhythm, and dynamics. Karaoke scoring technology (e.g., JOYSOUND's technology) is used for this analysis. Pitch analysis uses a MIDI analysis library, and analysis is performed for each note. Rhythm analysis is performed using a tempo analysis library, and dynamics analysis is performed using waveform analysis.

[0669] Furthermore, the server evaluates the performance expression using a large-scale language model (e.g., OpenAI's GPT-4). Based on the expression instructions specified by the user (e.g., "enjoyably"), the audio data is converted into text and analyzed by the model. The evaluation results of the performance expression are saved.

[0670] In addition, the server uses an emotion engine to analyze the user's emotional state while playing. It uses facial and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. Emotional data is used as supplementary information for music evaluation.

[0671] Based on the evaluation results, the server generates specific feedback. The feedback may include numerical evaluations, graph displays, text comments, etc. In addition, personalized advice based on the user's emotional state may also be provided. The generated feedback is sent from the server to the device. The device displays the sent feedback, providing it visually and audibly. This makes it easier for the user to understand the evaluation results of their performance. Specific examples of feedback that may be considered include the following:

[0672] "The pitch is generally accurate, but the rhythm is a little fast. Emotional analysis indicates that the performance is enjoyable, but with some tension."

[0673] Through this feedback, users can understand which parts of their performance need improvement and use that information for their next practice. Users can also use the feedback on their emotional state to help them play more relaxed.

[0674] An example of a prompt sentence that could be input to a generative AI model would be, "Please rate my piano playing."

[0675] This system allows users to effectively improve their musical performance skills and continue practicing while maintaining motivation and having fun.

[0676] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0677] Step 1:

[0678] The user launches the recording app on their device and taps the recording preparation button. This action causes the device to check the status of the camera and microphone and check that they are working properly. It notifies the user that they are ready. The input is the user's operation, and the output is a notification that the camera and microphone are ready to start.

[0679] Step 2:

[0680] The user taps the "Start Recording" button to begin playing. At this point, the device simultaneously activates its camera and microphone, recording the user's performance, facial expressions, and vocal tone. The input is the "Start Recording" button and the user's performance, and the output is the recorded audio and video data.

[0681] Step 3:

[0682] When the recording is finished, the user taps the "Stop Recording" button. The device responds by stopping the recording and saving the data. The input is the "Stop Recording" operation, and the output is the saved audio and video data.

[0683] Step 4:

[0684] The device prepares to send the recorded data (audio files) and emotion data (video files and audio analysis results) to the server. When the user presses the "Send" button in the confirmation dialog, the data is sent to the server. The input is the user's confirmation of sending, and the output is the data being sent to the server.

[0685] Step 5:

[0686] The server receives the voice data and emotion data sent from the device and stores them in the respective databases. The input is the received data, and the output is the stored data.

[0687] Step 6:

[0688] The server analyzes the audio data and evaluates pitch, rhythm, and dynamics. It uses a MIDI analysis library for pitch analysis, a tempo analysis library for rhythm analysis, and waveform analysis for dynamics analysis. The input is the audio data, and the output is the analysis results for each evaluation item.

[0689] Step 7:

[0690] The server evaluates performance expressions using a large-scale language model. Based on the user's instructions, the audio data is converted into text, which is then analyzed by the large-scale language model. The input is the audio data and instructions, and the output is the evaluation result of the performance expression.

[0691] Step 8:

[0692] The server uses an emotion engine to analyze the user's emotional state. It uses face and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. The input is video and audio data, and the output is the emotion analysis results.

[0693] Step 9:

[0694] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, text comments, and personalized advice. The input is each evaluation result, and the output is feedback data.

[0695] Step 10:

[0696] The server sends the generated feedback data to the terminal. The input is the feedback data, and the output is the data sent to the user's terminal.

[0697] Step 11:

[0698] The device receives the transmitted feedback data and provides it to the user visually and audibly, allowing the user to understand the feedback and use it for their next practice. The input is the feedback data, and the output is the provision of feedback to the user.

[0699] (Application example 2)

[0700] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0701] Conventional systems for supporting the improvement of musical performance focus on technical evaluations such as pitch, rhythm, and dynamics, but have the problem of being unable to provide feedback that takes into account the performer's emotional state. Another problem is that the feedback users receive is uniform for each individual performance, making it difficult to obtain specific and personalized advice that adapts to the user's emotional state. This makes it difficult for users to understand areas for improvement that correspond to their own emotional state, resulting in the problem of inefficient improvement of performance techniques.

[0702] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0703] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for analyzing the user's emotional state, means for personalizing feedback based on the emotional state, means for generating the feedback, and means for transmitting the feedback, thereby enabling the provision of specific and personalized feedback that takes into account not only technical evaluation but also the user's emotional state.

[0704] The "means for receiving performance data" is a mechanism for recording the music played by the user and transmitting that data to the server.

[0705] The "means for evaluating pitch" is a mechanism that analyzes the pitch of the performance data and measures how well it matches a reference pitch.

[0706] The "means for evaluating rhythm" is a mechanism for analyzing the timing of the sounds in the performance data and measuring the degree of agreement with a specified rhythm pattern.

[0707] The "means for evaluating dynamics" is a system that analyzes the changes in volume of performance data and evaluates the dynamics of the entire performance.

[0708] The "means for evaluating performance expression" is a system for evaluating the expressiveness and technique of a user's performance based on the evaluation results of pitch, rhythm, and dynamics.

[0709] The "means for analyzing emotional state" is a mechanism for analyzing the user's facial expressions and vocal tone while playing and recognizing the emotional state at that time.

[0710] The "means for generating feedback" is a mechanism that provides the user with specific and personalized improvements based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0711] The "means for transmitting feedback" is a mechanism for transmitting the generated feedback to the user's terminal and displaying it visually and audibly.

[0712] Overall system overview

[0713] This invention is a system to help users improve their musical performance by recording their performance and analyzing their emotions through recording their facial expressions and vocal timbre. It is realized using the following hardware and software.

[0714] Hardware and software used

[0715] Smartphone or tablet (e.g. iPhone, Android device)

[0716] PC or server (e.g. cloud server)

[0717] Devices with built-in cameras and microphones

[0718] Programming language: Swift (iOS), Kotlin (Android)

[0719] API: Google Cloud Speech-to-Text, Google Vision API (sentiment analysis)

[0720] Backend: Node.js

[0721] Database: Firebase

[0722] Main features of the system

[0723] Users first record their performance using an application installed on their smartphone or tablet. During recording, the device's camera and microphone simultaneously operate to collect the user's facial expressions and vocal tone. The collected data is then sent to a cloud server in real time or after recording.

[0724] Music evaluation and emotion analysis process

[0725] The server receives the transmitted performance data and evaluates the pitch, rhythm, and dynamics. Pitch evaluation involves analyzing how closely the pitch of each note matches the reference pitch. Rhythm evaluation involves checking synchronization with the specified pattern. Dynamic evaluation involves analyzing the dynamics of the entire performance. In addition to these technical evaluations, the server uses the Google Vision API to analyze facial expression data during performance and recognize the user's emotional state.

[0726] Generating and Providing Feedback

[0727] Based on the evaluation results, specific and personalized feedback is generated for the user, including evaluation results on pitch, rhythm, and dynamics, as well as advice based on emotional state. The generated feedback is sent to the user's device in the form of visual and audio.

[0728] Examples and prompts

[0729] As a concrete example, let's consider a situation where a user plays the guitar and then receives feedback from an app. First, the user launches the app and taps the start recording button. After finishing playing, the user taps the stop recording button, which sends the data to the server. After analysis is performed on the server side, the following feedback is displayed on the user's device:

[0730] "The pitch is generally accurate, but the rhythm is a little fast. The emotional analysis shows that you are playing with enjoyment, but there are some parts that seem tense. Please try playing a little more relaxed."

[0731] Example prompt to launch the app:

[0732] "Select your instrument and start recording. Relax and play, as your facial expressions will be recorded."

[0733] In this way, the system of the present invention allows users to understand areas for improvement in their performance, both technically and emotionally, through specific and personalized feedback.

[0734] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0735] Step 1:

[0736] Recording of performance

[0737] The user launches the application on their smartphone or tablet and taps the recording start button.

[0738] Input: User's music, facial expressions, and vocal tone

[0739] How it works: The device's microphone records the music, and the camera records the user's facial expressions while they play. The audio and facial expression data are then stored on the device.

[0740] Output: Recorded performance data and emotional data (facial expressions and vocal timbre)

[0741] Step 2:

[0742] Sending recording data and emotional data

[0743] When the user taps the recording stop button, the recorded performance data and emotion data are sent to the server.

[0744] Input: Recorded performance data, emotional data

[0745] How it works: Data is sent from the device to a server over the internet.

[0746] Output: Performance data and emotion data stored on the server

[0747] Step 3:

[0748] Pitch, rhythm and dynamic analysis

[0749] The server evaluates the pitch, rhythm, and dynamics based on the received performance data.

[0750] Input: Received performance data

[0751] How it works: A server-based analysis program compares pitch with a reference value and analyzes rhythm and dynamics. Specifically, it analyzes audio data using the Google Cloud Speech-to-Text API.

[0752] Output: Pitch, rhythm, and dynamics evaluation results

[0753] Step 4:

[0754] Evaluation of performance expression

[0755] The server evaluates the performance expression based on the evaluation results of pitch, rhythm, and dynamics.

[0756] Input: Pitch, rhythm, and dynamic evaluation results

[0757] How it works: The performance expression evaluation algorithm in the server analyzes the data and performs a comprehensive evaluation of the performance expression.

[0758] Output: Evaluation results of performance expression

[0759] Step 5:

[0760] Emotion Analysis

[0761] The server analyzes the received emotion data and evaluates the user's emotional state.

[0762] Input: Received emotional data (facial expressions, tone of voice)

[0763] How it works: It uses the Google Vision API to analyze facial expression data and recognize the user's emotional state.

[0764] Output: Evaluation result of the user's emotional state

[0765] Step 6:

[0766] Generate feedback

[0767] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[0768] Input: Pitch, rhythm, dynamics, performance expression, emotional state evaluation results

[0769] How it works: The server's feedback generator combines the results of each assessment to create specific, personalized feedback.

[0770] Output: Generated feedback data

[0771] Step 7:

[0772] Send Feedback

[0773] The server transmits the generated feedback to the user's terminal.

[0774] Input: Generated feedback data

[0775] Operation: The server sends feedback data to the terminal, which displays it visually and audibly to the user.

[0776] Output: Feedback displayed on the user's device

[0777] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0778] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0779] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0780] [Third embodiment]

[0781] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0782] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0783] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0784] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0785] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0786] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0787] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0788] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0789] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0790] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0791] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0792] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0793] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance. The system's purpose is to objectively evaluate the music performed by the user and provide specific feedback. Characteristic elements of this system include receiving performance data, analyzing the performance data (evaluating pitch, rhythm, and dynamics), evaluating performance expression, and generating and transmitting feedback.

[0794] overview

[0795] Users record their performances using devices such as smartphones or tablets and send the audio data to the system. The server analyzes the performance data, evaluating pitch, rhythm, and dynamics, and also assesses the expressiveness of the performance using a large-scale language model. Specific feedback is then generated and provided to the user based on the evaluation results.

[0796] User-side processing

[0797] 1. Recording your performance

[0798] The user can record their own performance using the device's recording function, and start and stop recording by user operation.

[0799] 2. Sending recording data

[0800] Once the recording is complete, the device will automatically send the audio data to the server, or the user can manually send the audio.

[0801] Server-side processing

[0802] 1. Receiving music data

[0803] The server receives the voice data sent from the user's device and stores it in an appropriate format for subsequent analysis.

[0804] 2. Pitch, rhythm, and dynamic analysis

[0805] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0806] 3. Evaluation of performance expression

[0807] The server uses a large-scale language model to evaluate the expressiveness of the performance, analyzing it based on specific instructions (e.g., "play it happily" or "let it rain") and generating text comments as a result.

[0808] 4. Generate feedback

[0809] The server generates comprehensive feedback based on the analysis results, including evaluations of pitch, rhythm, and dynamics, as well as comments about the performance expression.

[0810] 5. Submitting Feedback

[0811] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0812] Specific examples

[0813] As a concrete example, consider a user playing a piano piece. First, the user records their performance on their smartphone and sends the audio data to the system. The server receives the audio data and then evaluates the accuracy of pitch (whether each note is played at the correct pitch in the score), rhythm (whether the tempo is correct), and dynamics. If the user attempts to express their performance in a "playful" way, this expressiveness is also analyzed using a large-scale language model. Finally, specific feedback is generated and sent to the user, along with the evaluation results of pitch, rhythm, and dynamics, such as "Overall, your performance sounds playful, but some phrases lack dynamic control." This feedback allows the user to specifically understand areas for improvement in their performance and use this information for their next practice session.

[0814] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0815] The processing flow will be explained below.

[0816] User-side processing

[0817] Step 1:

[0818] The user launches the recording app on their device and taps the button to start playing.

[0819] Your device will begin recording your performance.

[0820] Step 2:

[0821] The user finishes playing.

[0822] The device stops recording and saves the audio data.

[0823] Step 3:

[0824] The user taps the send button for the recording.

[0825] The device transmits the stored voice data to the server.

[0826] Server-side processing

[0827] Step 1:

[0828] The server receives the voice data transmitted from the terminal.

[0829] Store the received data in the appropriate format.

[0830] Step 2:

[0831] The server analyzes the audio data and evaluates the pitch.

[0832] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[0833] Step 3:

[0834] The server analyzes the audio data for each rhythm pattern.

[0835] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[0836] Step 4:

[0837] The server analyzes the volume of the sound from the audio data.

[0838] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[0839] Step 5:

[0840] The server evaluates the performance expression using a large-scale language model.

[0841] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[0842] Step 6:

[0843] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, and performance expression.

[0844] Feedback includes a numerical rating, a graphical display, and text comments.

[0845] Step 7:

[0846] The server transmits the generated feedback data to the terminal.

[0847] Log the completion of the transmission.

[0848] Terminal side processing

[0849] Step 1:

[0850] The terminal receives the feedback data sent from the server.

[0851] Verify the integrity of the received data and check for any invalid data.

[0852] Step 2:

[0853] The terminal analyzes the feedback data.

[0854] The analysis results are displayed visually and audibly to the user.

[0855] Specific examples

[0856] Take the example of a user playing a piano piece. The user launches a recording app on their smartphone, starts playing, and records. When recording is finished, the audio data is sent to the server. The server analyzes the received audio data and evaluates pitch, rhythm, and dynamics. At the same time, it evaluates the performance expression using a large-scale language model. Feedback is then generated based on the evaluation results and sent to the user's device. The user receives the feedback on their device, understands which parts of their performance need improvement, and can use this information to improve their next practice.

[0857] Example 1

[0858] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0859] Conventional systems supporting the improvement of musical performance lack sufficient analysis of performance data, making it difficult to provide specific feedback. Furthermore, evaluation of expressiveness is subjective, making it difficult to provide specific advice. As a result, it is difficult for users to accurately grasp areas for improvement in their own performance.

[0860] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0861] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, and means for transmitting the feedback, thereby enabling detailed analysis of the user's performance data and providing specific and objective feedback.

[0862] "Performance data" is audio information generated during musical performance.

[0863] "Means for receiving" refers to a device or method for receiving performance data from an external source.

[0864] "Means for analysis" refers to technology for breaking down received performance data into elements such as pitch, rhythm, and dynamics and evaluating them.

[0865] "Means for evaluation" refers to a device or method for determining the quality of a performance based on the analysis results.

[0866] "Performance expression" is a method of musical expression that reflects the performer's emotions and nuances.

[0867] A "large-scale language model" is a type of artificial intelligence that is trained based on large amounts of data in natural language processing.

[0868] A "prompt" is a sentence used to give instructions to an AI model.

[0869] "Feedback" refers to specific advice or comments provided to the user based on the results of analysis and evaluation.

[0870] The "means for transmitting" refers to a device or method for sending the generated feedback to the user's terminal.

[0871] "Visual display means" refers to a device or method that displays feedback on a screen in the form of text or graphs.

[0872] "Auditory display means" refers to a device or method for audibly conveying feedback to the user.

[0873] "Karaoke scoring technology" is a technique for evaluating musical performances by comparing them with standard pitch and rhythm.

[0874] The present invention relates to a system that efficiently supports the improvement of musical performance. The purpose of this system is to objectively evaluate the music performed by a user and provide specific feedback. Detailed embodiments of this system are described below.

[0875] User-side processing

[0876] Users record their performances using devices such as smartphones and tablets. Recording can be started and stopped by user operation. For example, a user can record a performance using a voice recorder app on their smartphone. After recording is finished, the device automatically sends the audio data to the server. If the user has disabled automatic sending, they can also select the recorded file and send it manually.

[0877] Server-side processing

[0878] The server receives the voice data sent from the user's device. This data is saved in an appropriate format (e.g., WAV format). Specifically, the data is received using the HTTP protocol or the FTP protocol.

[0879] The server then analyzes the received audio data using karaoke scoring technology. For pitch analysis, it uses the Python music analysis library "LibROSA," comparing each note with a reference pitch. For rhythm analysis, it uses "pYin," an audio pitch detector tool, to confirm whether it matches the original pattern of the song. For dynamic analysis, it evaluates the amplitude of the audio data.

[0880] The performance expression is evaluated using a large-scale language model. Specifically, "GPT-4" is used, and the prompt is "Please rate the expressiveness of this piano piece played with enjoyment." Based on this, the large-scale language model analyzes the expressiveness of the performance and generates a text comment.

[0881] Generate and send feedback

[0882] The server generates feedback by integrating the results of pitch, rhythm, and dynamic analysis with the evaluation of performance expression using a large-scale language model. The generated feedback is designed to include specific advice. For example, it might say, "Overall, your playing is enjoyable, but some phrases lack dynamic control. Your tempo is off in some places, so we recommend practicing with a metronome."

[0883] Finally, the server sends the generated feedback to the user's device. The feedback is provided both visually and audibly, and the user can check the feedback within the app. This allows the user to specifically understand areas for improvement in their performance and use this information to improve their next practice.

[0884] Specific examples

[0885] For example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on their smartphone. The recording is automatically sent to the server. The server analyzes the received data and checks the accuracy of pitch and rhythm using the "LibROSA" library and the "pYin" tool. Furthermore, to evaluate the performance expression, the server inputs a prompt sentence into the large-scale language model "GPT-4" saying, "Please rate the expressiveness of this enjoyable piano performance," and generates comments about the expressiveness. Based on the analysis results, the server generates feedback such as, "Overall, the performance is enjoyable, but there is a lack of dynamic control in some phrases," and sends it to the user.

[0886] This allows the user to understand specific areas for improvement and use this information in their next practice session.

[0887] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0888] Step 1: Record your performance

[0889] Users use the recording function of their smartphones or tablets to record their own performances. The input is the music played by the user, and the output is audio data (e.g., a WAV file). Specifically, recording begins when the user presses the "Start Recording" button on the device, and ends when the user presses the "Stop Recording" button.

[0890] Step 2: Send your recordings

[0891] The device sends the audio data to the server after recording is complete. The input is the recorded audio data, and the output is a notification that transmission is complete. Specifically, the device uploads the data to the server using the HTTP or FTP protocol. If the user has disabled automatic transmission, they can select the recorded file and press the "Send" button to send it manually.

[0892] Step 3: Receiving music data

[0893] The server receives the voice data sent from the user's device. The input is the voice data sent from the device, and the output is the voice data stored on the server. Specifically, the server saves the received voice data in an appropriate format (e.g., WAV format).

[0894] Step 4: Analyze pitch, rhythm and dynamics

[0895] The server analyzes audio data using karaoke scoring technology. The input is the saved audio data, and the output is the evaluation results of pitch, rhythm, and dynamics. Specifically, the server analyzes pitch using the Python music analysis library "LibROSA" and compares each note with a reference pitch. It also performs rhythm analysis using the audio pitch detector tool "pYin" to confirm whether it matches the original pattern of the song. It also evaluates the dynamics of the entire song through amplitude analysis.

[0896] Step 5: Evaluate the performance

[0897] The server evaluates the performance expression using a large-scale language model. The input is audio data and analysis results, and the output is text feedback about the performance expression. Specifically, the server inputs a prompt sentence, "Please rate the expressiveness of this piano piece played with enjoyment," into the large-scale language model (e.g., "GPT-4") and obtains the generated text comment.

[0898] Step 6: Generate feedback

[0899] The server generates feedback by integrating the analysis results and the evaluation results of performance expression. The inputs are the evaluation results of pitch, rhythm, and dynamics, and the evaluation results of performance expression, and the output is comprehensive feedback. Specifically, the server generates feedback comments containing specific advice based on these results. For example, it may generate a comment such as, "Overall, the performance appears enjoyable, but there is a lack of dynamic control in some phrases."

[0900] Step 7: Submit your feedback

[0901] The server sends the generated feedback to the user's terminal. The input is the generated feedback comment, and the output is the feedback displayed on the user's terminal. In specific operations, the server sends feedback data to the terminal so that the user can visually and audibly confirm the feedback.

[0902] Through the above steps, users can receive detailed analysis results of their own performance and specific advice for improvement, which can be used to improve their performance skills.

[0903] (Application example 1)

[0904] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0905] Current factory robots lack systems that evaluate the accuracy and efficiency of their movements and provide feedback. This leads to accumulated inaccuracies and inefficiencies in robot movements, resulting in reduced productivity. Especially for robots performing complex tasks, even minute errors in movement can cause serious problems, and specific and effective measures to resolve this are needed.

[0906] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0907] In this invention, the server includes means for receiving performance data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, means for transmitting the feedback, means for acquiring and transmitting robot movement data, means for analyzing the movement data to evaluate accuracy and efficiency, means for generating movement improvement feedback based on the analysis results, and means for transmitting the feedback to the robot and the operator terminal. This makes it possible to evaluate the accuracy and efficiency of the robot's movements in real time and provide specific and effective feedback.

[0908] "Performance data" is audio information of the music played by the user.

[0909] "Means for receiving" refers to a device or module for receiving specific data from the outside and transferring it to the inside.

[0910] "Means for analyzing" refers to devices or modules that process and analyze the received data and evaluate and extract information from it.

[0911] "Means for evaluation" refers to a device or module that judges the performance and accuracy of the analyzed data.

[0912] A "means for generating feedback" is a device or module that provides improvements and advice to users and systems based on the evaluation results.

[0913] A "transmitting means" is a device or module for sending internal data to an external terminal or system.

[0914] "Robot operation data" refers to information relating to the operation of a robot operating in a factory range, such as controlling the position, speed, and force.

[0915] The "means for acquiring and transmitting operation data" refers to a device or module that collects the operation data of the robot and transmits it to a server or the like.

[0916] The "means for evaluating accuracy and efficiency" is a device or module for analyzing the robot's motion data and determining how accurate and efficient the motion is.

[0917] The "means for generating motion improvement feedback" is a device or module that presents improvements to the robot's motion based on the evaluation results of the motion data.

[0918] An "operator terminal" is a computer or mobile device used by a worker in a factory to receive feedback and operational information from the robot.

[0919] This invention relates to a system called "Robot Work Buddy" that accurately evaluates the behavior of robots used in factories and provides efficient feedback. The system collects and analyzes robot behavior data, generates specific improvement advice, and sends it to the robot and its operator.

[0920] Hardware used

[0921] Factory robots: Equipped with various sensors (position, speed, force, etc.) to collect operational data.

[0922] Server: A central processing unit for analyzing received data. It uses Python's SciPy and TensorFlow.

[0923] Operator terminal: A computer or mobile device on which an operator receives feedback.

[0924] Software used

[0925] Data collection module: Has the function of collecting robot operation data and sending it to the server.

[0926] Dynamic control algorithm: Analyzes the robot's motion data and compares it with ideal motion patterns.

[0927] Machine learning algorithms: Using TensorFlow and other algorithms, patterns of motion data are learned to improve accuracy.

[0928] Evaluation and feedback generation module: Generates specific feedback based on the analysis results and sends it to the robot or operator.

[0929] Specific example of operation procedure

[0930] 1. Acquisition and transmission of motion data

[0931] First, robots operating in factories collect data on their position, speed, and force for each task. This data is acquired through sensors built into the robots and then sent to a server.

[0932] 2. Data Analysis

[0933] The server analyzes the received motion data. It uses Python's SciPy to calculate errors in the position data and TensorFlow to learn the motion patterns. Specifically, it compares them with ideal motion patterns and identifies where there are errors.

[0934] 3. Evaluation and feedback generation

[0935] Based on the analysis results, the evaluation and feedback generation module evaluates the accuracy and efficiency of the robot's movements. Based on the evaluation results, it generates specific feedback to help improve the robot's movements. For example, it generates a text comment such as, "The position was off during the bolt tightening operation, so please correct coordinate X."

[0936] 4. Submitting Feedback

[0937] The generated feedback is sent to the robot's control system and to the operator's terminal, where it can be visually confirmed using a visualization tool (e.g., Grafana).

[0938] Adding specific examples

[0939] When a factory robot tightens a bolt, it uses sensors to collect real-time data on the position, speed, and tightening force of each bolt. This data is then sent to a server, which analyzes it using SciPy and TensorFlow. Based on the analysis, specific feedback is generated, such as "The position was off, so please correct coordinate X next time." This feedback is sent to the robot's control system and operator terminal and reflected in the next task.

[0940] Prompt Sentence Examples

[0941] Use a generative AI model to generate feedback using the following prompt:

[0942] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0943] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0944] Step 1:

[0945] The robot acquires operational data for each task. Specifically, it measures data such as position, speed, and force using internal sensors. This data is collected for each task and temporarily stored in the robot's internal memory.

[0946] (Input: Robot movement, Output: Position data, velocity data, force data)

[0947] Step 2:

[0948] The robot sends the collected operation data to a server using the HTTP protocol, and the data stored in the robot's internal memory is uploaded to the server.

[0949] (Input: position data, velocity data, force data, output: motion data sent to the server)

[0950] Step 3:

[0951] The server analyzes the received motion data, calculates the error in the position data using Python's SciPy, and uses TensorFlow to learn and compare motion patterns. Specifically, it compares the ideal motion pattern with the actual data and identifies where the error exists.

[0952] (Input: movement data, output: position data error, comparison result with movement pattern)

[0953] Step 4:

[0954] The server generates evaluation and feedback based on the analysis results. Using the evaluation and feedback generation module, it evaluates the accuracy and efficiency of the action and generates specific advice for improvement. For example, it generates a text comment such as "The position was off, so please correct coordinate X next time."

[0955] (Input: position data error, comparison result with movement pattern, output: feedback)

[0956] Step 5:

[0957] The server sends the generated feedback to the robot and the operator's terminal, where it is displayed as a specific text comment and also sent to the robot's control system to be reflected in the robot's next task.

[0958] (Input: Feedback, Output: Display on operator terminal, transmission to robot control system)

[0959] Specifically, when a robot working in a factory tightens a bolt, the necessary data is acquired at each step, processed, and the results are fed back to improve the robot's behavior. An example of a generated prompt is as follows:

[0960] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[0961] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0962] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. This system supports the improvement of the user's performance technique by objectively evaluating the music played by the user and providing specific feedback. It is also possible to analyze the user's emotional state and provide more personalized feedback.

[0963] overview

[0964] Users record their performance using a device such as a smartphone or tablet, and simultaneously use the device's camera and microphone to record their facial expressions and vocal tone while they play. The server analyzes the transmitted performance and emotional data, evaluating pitch, rhythm, dynamics, and performance expression, while also recognizing the user's emotional state. Based on the evaluation results, specific feedback is generated and provided to the user.

[0965] User-side processing

[0966] 1. Recording your performance

[0967] The user launches the recording app on their device and taps the button to start playing. At the same time, the device's camera and microphone are activated, recording the user's facial expressions and voice tone. Recording can be started and stopped by the user.

[0968] 2. Sending recording data and emotion data

[0969] Once the recording of the voice and emotion data is complete, the device automatically sends the data to the server, or the user can manually send it.

[0970] Server-side processing

[0971] 1. Receiving Data

[0972] The server receives the voice data and emotion data sent from the device and stores the received data in an appropriate format for subsequent analysis.

[0973] 2. Pitch, rhythm, and dynamic analysis

[0974] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[0975] 3. Evaluation of performance expression

[0976] The server evaluates the performance expression using a large-scale language model. Based on the expression instructions entered by the user (e.g., "playing it happily"), it analyzes the audio data and saves the evaluation results.

[0977] 4. Emotion Analysis

[0978] The server uses an emotion engine to analyze the user's emotions. It analyzes facial expressions and vocal tone during the performance to recognize the user's emotional state. The emotion data is used as supplementary information for music evaluation.

[0979] 5. Generate feedback

[0980] The server generates feedback based on the evaluation of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, and text comments. It also includes personalized advice based on the user's emotional state.

[0981] 6. Submitting Feedback

[0982] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[0983] Specific examples

[0984] As a concrete example, consider a user playing a piano piece. The user launches a recording app on their smartphone and taps the start recording button. The device's camera and microphone record the performance while capturing the user's facial expressions and vocal tone. After the performance and recording are complete, the audio and emotional data are sent to the server. The server analyzes the received data, evaluating pitch, rhythm, and dynamics, and using a large-scale language model to evaluate the performance. The emotion engine also analyzes the user's emotional state. Based on the evaluation results, specific feedback is generated and sent to the user's device, such as "Your pitch is generally accurate, but your rhythm is a little fast. Emotion analysis indicates that you seem to be enjoying your performance, but you seem a little nervous." This feedback allows the user to understand which aspects of their performance need improvement and use it to improve their next practice. The user can use the emotional state feedback to play more relaxed.

[0985] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[0986] The processing flow will be explained below.

[0987] User-side processing

[0988] Step 1:

[0989] The user launches the recording app on their device and taps the button to start playing.

[0990] As the device begins recording your performance, the camera and microphone begin capturing your facial expressions and tone of voice.

[0991] Step 2:

[0992] The user finishes playing and taps the stop recording button.

[0993] The device stops recording and saves the performance data and emotional data.

[0994] Step 3:

[0995] The user taps the send button for the recording.

[0996] The device transmits the stored voice data and emotion data to the server.

[0997] Server-side processing

[0998] Step 1:

[0999] The server receives the voice data and emotion data transmitted from the terminal.

[1000] The received data is stored in an appropriate format for analysis.

[1001] Step 2:

[1002] The server analyzes the audio data and evaluates the pitch.

[1003] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[1004] Step 3:

[1005] The server analyzes the audio data for each rhythm pattern.

[1006] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[1007] Step 4:

[1008] The server analyzes the volume of the sound from the audio data.

[1009] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[1010] Step 5:

[1011] The server evaluates the performance expression using a large-scale language model.

[1012] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[1013] Step 6:

[1014] The server uses an emotion engine to analyze the user's emotions.

[1015] The system analyzes facial expressions and vocal tone during performance, recognizes the user's emotional state, and saves the evaluation results.

[1016] Step 7:

[1017] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1018] Feedback includes numerical ratings, graphical displays, and text comments, as well as personalized advice based on the user's emotional state.

[1019] Step 8:

[1020] The server transmits the generated feedback data to the user's terminal.

[1021] Log the completion of the transmission.

[1022] Terminal side processing

[1023] Step 1:

[1024] The terminal receives the feedback data sent from the server.

[1025] Verify the integrity of the received data and check for any invalid data.

[1026] Step 2:

[1027] The terminal analyzes the feedback data.

[1028] The analysis results are displayed visually and audibly to the user.

[1029] Specific examples

[1030] Let's take the example of a user playing a piano piece. First, the user launches the recording app on their smartphone and taps the start recording button. The device starts recording, and at the same time, the camera and microphone begin to capture the user's facial expressions and vocal tone. When the performance is finished, the user taps the stop recording button, and the recording data and emotional data are saved. Then, the user taps the send button for the recording data, and the voice data and emotional data are sent to the server.

[1031] The server receives the audio data and emotion data, analyzing pitch, rhythm, dynamics, and performance expression, respectively. At the same time, the emotion engine analyzes the user's facial expressions and vocal tone to evaluate their emotional state. Based on the evaluation results, specific feedback such as "The pitch is good, but the rhythm is a little fast. The performance seems enjoyable, but there is also a sense of tension" is generated and sent to the user's device. The user can receive the feedback on their device, understand where their performance needs improvement, and use the feedback to improve their next practice.

[1032] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[1033] Example 2

[1034] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1035] Conventional music performance improvement support systems have difficulty not only evaluating a user's performance technique but also providing personalized feedback based on the user's emotional state. Furthermore, in order to improve the accuracy of the evaluation results, they lack the functionality to evaluate performance expression using large-scale language models and analyze the user's emotional state.

[1036] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the user's performance, means for receiving performance data and emotional data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the performance data, means for analyzing the emotional data to evaluate the user's emotional state, means for generating feedback based on the evaluation result and the emotional state, and means for transmitting the feedback. This makes it possible to evaluate the user's performance technique as well as provide detailed and personalized feedback based on the user's emotional state.

[1037] "Means for recording a user's performance" refers to a device or software that digitally records the audio of a user playing music.

[1038] "Means for receiving performance data and emotional data" refers to a function for receiving audio data and emotional data sent from a user's terminal via a network to a device such as a server.

[1039] "Means for analyzing performance data to evaluate pitch, rhythm, and dynamics" refers to algorithms and software that automatically analyze the basic elements of music—pitch, rhythm, and dynamics—and evaluate their accuracy and expressiveness.

[1040] "Means for evaluating performance expression based on performance data" refers to technologies such as large-scale language models for analyzing and evaluating the expressive intent behind a user's performance.

[1041] "Means for analyzing emotional data and assessing a user's emotional state" refers to an engine or software that identifies the user's emotional state from the tone of voice and facial expressions in a recording, and analyzes and assesses it.

[1042] "Means for generating feedback based on evaluation results and emotional state" refers to systems and algorithms for generating feedback to be provided to a user based on the analyzed performance data and emotional data.

[1043] "Means for sending feedback" refers to communication functions and protocols for sending the generated feedback data to the user's terminal.

[1044] The present invention is a system that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. The system configuration includes a user terminal, a server, and a communication function for exchanging data between them. Specific embodiments of the system are described below.

[1045] Users record their performances using devices such as smartphones or tablets. When they launch the recording app and tap the start button, the device's camera and microphone are activated. This allows the user's performance audio, as well as their facial expressions and vocal tone, to be recorded. Recording can be started and stopped by the user.

[1046] Once recording is complete, the device sends the voice data and emotion data to the server. The server receives the data and stores it in the appropriate format. The server then analyzes the received voice data in terms of pitch, rhythm, and dynamics. Karaoke scoring technology (e.g., JOYSOUND's technology) is used for this analysis. Pitch analysis uses a MIDI analysis library, and analysis is performed for each note. Rhythm analysis is performed using a tempo analysis library, and dynamics analysis is performed using waveform analysis.

[1047] Furthermore, the server evaluates the performance expression using a large-scale language model (e.g., OpenAI's GPT-4). Based on the expression instructions specified by the user (e.g., "enjoyably"), the audio data is converted into text and analyzed by the model. The evaluation results of the performance expression are saved.

[1048] In addition, the server uses an emotion engine to analyze the user's emotional state while playing. It uses facial and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. Emotional data is used as supplementary information for music evaluation.

[1049] Based on the evaluation results, the server generates specific feedback. The feedback may include numerical evaluations, graph displays, text comments, etc. In addition, personalized advice based on the user's emotional state may also be provided. The generated feedback is sent from the server to the device. The device displays the sent feedback, providing it visually and audibly. This makes it easier for the user to understand the evaluation results of their performance. Specific examples of feedback that may be considered include the following:

[1050] "The pitch is generally accurate, but the rhythm is a little fast. Emotional analysis indicates that the performance is enjoyable, but with some tension."

[1051] Through this feedback, users can understand which parts of their performance need improvement and use that information for their next practice. Users can also use the feedback on their emotional state to help them play more relaxed.

[1052] An example of a prompt sentence that could be input to a generative AI model would be, "Please rate my piano playing."

[1053] This system allows users to effectively improve their musical performance skills and continue practicing while maintaining motivation and having fun.

[1054] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1055] Step 1:

[1056] The user launches the recording app on their device and taps the recording preparation button. This action causes the device to check the status of the camera and microphone and check that they are working properly. It notifies the user that they are ready. The input is the user's operation, and the output is a notification that the camera and microphone are ready to start.

[1057] Step 2:

[1058] The user taps the "Start Recording" button to begin playing. At this point, the device simultaneously activates its camera and microphone, recording the user's performance, facial expressions, and vocal tone. The input is the "Start Recording" button and the user's performance, and the output is the recorded audio and video data.

[1059] Step 3:

[1060] When the recording is finished, the user taps the "Stop Recording" button. The device responds by stopping the recording and saving the data. The input is the "Stop Recording" operation, and the output is the saved audio and video data.

[1061] Step 4:

[1062] The device prepares to send the recorded data (audio files) and emotion data (video files and audio analysis results) to the server. When the user presses the "Send" button in the confirmation dialog, the data is sent to the server. The input is the user's confirmation of sending, and the output is the data being sent to the server.

[1063] Step 5:

[1064] The server receives the voice data and emotion data sent from the device and stores them in the respective databases. The input is the received data, and the output is the stored data.

[1065] Step 6:

[1066] The server analyzes the audio data and evaluates pitch, rhythm, and dynamics. It uses a MIDI analysis library for pitch analysis, a tempo analysis library for rhythm analysis, and waveform analysis for dynamics analysis. The input is the audio data, and the output is the analysis results for each evaluation item.

[1067] Step 7:

[1068] The server evaluates performance expressions using a large-scale language model. Based on the user's instructions, the audio data is converted into text, which is then analyzed by the large-scale language model. The input is the audio data and instructions, and the output is the evaluation result of the performance expression.

[1069] Step 8:

[1070] The server uses an emotion engine to analyze the user's emotional state. It uses face and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. The input is video and audio data, and the output is the emotion analysis results.

[1071] Step 9:

[1072] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, text comments, and personalized advice. The input is each evaluation result, and the output is feedback data.

[1073] Step 10:

[1074] The server sends the generated feedback data to the terminal. The input is the feedback data, and the output is the data sent to the user's terminal.

[1075] Step 11:

[1076] The device receives the transmitted feedback data and provides it to the user visually and audibly, allowing the user to understand the feedback and use it for their next practice. The input is the feedback data, and the output is the provision of feedback to the user.

[1077] (Application example 2)

[1078] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1079] Conventional systems for supporting the improvement of musical performance focus on technical evaluations such as pitch, rhythm, and dynamics, but have the problem of being unable to provide feedback that takes into account the performer's emotional state. Another problem is that the feedback users receive is uniform for each individual performance, making it difficult to obtain specific and personalized advice that adapts to the user's emotional state. This makes it difficult for users to understand areas for improvement that correspond to their own emotional state, resulting in the problem of inefficient improvement of performance techniques.

[1080] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1081] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for analyzing the user's emotional state, means for personalizing feedback based on the emotional state, means for generating the feedback, and means for transmitting the feedback, thereby enabling the provision of specific and personalized feedback that takes into account not only technical evaluation but also the user's emotional state.

[1082] The "means for receiving performance data" is a mechanism for recording the music played by the user and transmitting that data to the server.

[1083] The "means for evaluating pitch" is a mechanism that analyzes the pitch of the performance data and measures how well it matches a reference pitch.

[1084] The "means for evaluating rhythm" is a mechanism for analyzing the timing of the sounds in the performance data and measuring the degree of agreement with a specified rhythm pattern.

[1085] The "means for evaluating dynamics" is a system that analyzes the changes in volume of performance data and evaluates the dynamics of the entire performance.

[1086] The "means for evaluating performance expression" is a system for evaluating the expressiveness and technique of a user's performance based on the evaluation results of pitch, rhythm, and dynamics.

[1087] The "means for analyzing emotional state" is a mechanism for analyzing the user's facial expressions and vocal tone while playing and recognizing the emotional state at that time.

[1088] The "means for generating feedback" is a mechanism that provides the user with specific and personalized improvements based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1089] The "means for transmitting feedback" is a mechanism for transmitting the generated feedback to the user's terminal and displaying it visually and audibly.

[1090] Overall system overview

[1091] This invention is a system to help users improve their musical performance by recording their performance and analyzing their emotions through recording their facial expressions and vocal timbre. It is realized using the following hardware and software.

[1092] Hardware and software used

[1093] Smartphone or tablet (e.g. iPhone, Android device)

[1094] PC or server (e.g. cloud server)

[1095] Devices with built-in cameras and microphones

[1096] Programming language: Swift (iOS), Kotlin (Android)

[1097] API: Google Cloud Speech-to-Text, Google Vision API (sentiment analysis)

[1098] Backend: Node.js

[1099] Database: Firebase

[1100] Main features of the system

[1101] Users first record their performance using an application installed on their smartphone or tablet. During recording, the device's camera and microphone simultaneously operate to collect the user's facial expressions and vocal tone. The collected data is then sent to a cloud server in real time or after recording.

[1102] Music evaluation and emotion analysis process

[1103] The server receives the transmitted performance data and evaluates the pitch, rhythm, and dynamics. Pitch evaluation involves analyzing how closely the pitch of each note matches the reference pitch. Rhythm evaluation involves checking synchronization with the specified pattern. Dynamic evaluation involves analyzing the dynamics of the entire performance. In addition to these technical evaluations, the server uses the Google Vision API to analyze facial expression data during performance and recognize the user's emotional state.

[1104] Generating and Providing Feedback

[1105] Based on the evaluation results, specific and personalized feedback is generated for the user, including evaluation results on pitch, rhythm, and dynamics, as well as advice based on emotional state. The generated feedback is sent to the user's device in the form of visual and audio.

[1106] Examples and prompts

[1107] As a concrete example, let's consider a situation where a user plays the guitar and then receives feedback from an app. First, the user launches the app and taps the start recording button. After finishing playing, the user taps the stop recording button, which sends the data to the server. After analysis is performed on the server side, the following feedback is displayed on the user's device:

[1108] "The pitch is generally accurate, but the rhythm is a little fast. The emotional analysis shows that you are playing with enjoyment, but there are some parts that seem tense. Please try playing a little more relaxed."

[1109] Example prompt to launch the app:

[1110] "Select your instrument and start recording. Relax and play, as your facial expressions will be recorded."

[1111] In this way, the system of the present invention allows users to understand areas for improvement in their performance, both technically and emotionally, through specific and personalized feedback.

[1112] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1113] Step 1:

[1114] Recording of performance

[1115] The user launches the application on their smartphone or tablet and taps the recording start button.

[1116] Input: User's music, facial expressions, and vocal tone

[1117] How it works: The device's microphone records the music, and the camera records the user's facial expressions while they play. The audio and facial expression data are then stored on the device.

[1118] Output: Recorded performance data and emotional data (facial expressions and vocal timbre)

[1119] Step 2:

[1120] Sending recording data and emotional data

[1121] When the user taps the recording stop button, the recorded performance data and emotion data are sent to the server.

[1122] Input: Recorded performance data, emotional data

[1123] How it works: Data is sent from the device to a server over the internet.

[1124] Output: Performance data and emotion data stored on the server

[1125] Step 3:

[1126] Pitch, rhythm and dynamic analysis

[1127] The server evaluates the pitch, rhythm, and dynamics based on the received performance data.

[1128] Input: Received performance data

[1129] How it works: A server-based analysis program compares pitch with a reference value and analyzes rhythm and dynamics. Specifically, it analyzes audio data using the Google Cloud Speech-to-Text API.

[1130] Output: Pitch, rhythm, and dynamics evaluation results

[1131] Step 4:

[1132] Evaluation of performance expression

[1133] The server evaluates the performance expression based on the evaluation results of pitch, rhythm, and dynamics.

[1134] Input: Pitch, rhythm, and dynamic evaluation results

[1135] How it works: The performance expression evaluation algorithm in the server analyzes the data and performs a comprehensive evaluation of the performance expression.

[1136] Output: Evaluation results of performance expression

[1137] Step 5:

[1138] Emotion Analysis

[1139] The server analyzes the received emotion data and evaluates the user's emotional state.

[1140] Input: Received emotional data (facial expressions, tone of voice)

[1141] How it works: It uses the Google Vision API to analyze facial expression data and recognize the user's emotional state.

[1142] Output: Evaluation result of the user's emotional state

[1143] Step 6:

[1144] Generate feedback

[1145] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1146] Input: Pitch, rhythm, dynamics, performance expression, emotional state evaluation results

[1147] How it works: The server's feedback generator combines the results of each assessment to create specific, personalized feedback.

[1148] Output: Generated feedback data

[1149] Step 7:

[1150] Send Feedback

[1151] The server transmits the generated feedback to the user's terminal.

[1152] Input: Generated feedback data

[1153] Operation: The server sends feedback data to the terminal, which displays it visually and audibly to the user.

[1154] Output: Feedback displayed on the user's device

[1155] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1156] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1157] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1158] [Fourth embodiment]

[1159] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1160] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1161] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1162] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1163] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1164] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1165] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1166] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1167] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1168] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1169] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1170] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1171] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1172] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance. The system's purpose is to objectively evaluate the music performed by the user and provide specific feedback. Characteristic elements of this system include receiving performance data, analyzing the performance data (evaluating pitch, rhythm, and dynamics), evaluating performance expression, and generating and transmitting feedback.

[1173] overview

[1174] Users record their performances using devices such as smartphones or tablets and send the audio data to the system. The server analyzes the performance data, evaluating pitch, rhythm, and dynamics, and also assesses the expressiveness of the performance using a large-scale language model. Specific feedback is then generated and provided to the user based on the evaluation results.

[1175] User-side processing

[1176] 1. Recording your performance

[1177] The user can record their own performance using the device's recording function, and start and stop recording by user operation.

[1178] 2. Sending recording data

[1179] Once the recording is complete, the device will automatically send the audio data to the server, or the user can manually send the audio.

[1180] Server-side processing

[1181] 1. Receiving music data

[1182] The server receives the voice data sent from the user's device and stores it in an appropriate format for subsequent analysis.

[1183] 2. Pitch, rhythm, and dynamic analysis

[1184] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[1185] 3. Evaluation of performance expression

[1186] The server uses a large-scale language model to evaluate the expressiveness of the performance, analyzing it based on specific instructions (e.g., "play it happily" or "let it rain") and generating text comments as a result.

[1187] 4. Generate feedback

[1188] The server generates comprehensive feedback based on the analysis results, including evaluations of pitch, rhythm, and dynamics, as well as comments about the performance expression.

[1189] 5. Submitting Feedback

[1190] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[1191] Specific examples

[1192] As a concrete example, consider a user playing a piano piece. First, the user records their performance on their smartphone and sends the audio data to the system. The server receives the audio data and then evaluates the accuracy of pitch (whether each note is played at the correct pitch in the score), rhythm (whether the tempo is correct), and dynamics. If the user attempts to express their performance in a "playful" way, this expressiveness is also analyzed using a large-scale language model. Finally, specific feedback is generated and sent to the user, along with the evaluation results of pitch, rhythm, and dynamics, such as "Overall, your performance sounds playful, but some phrases lack dynamic control." This feedback allows the user to specifically understand areas for improvement in their performance and use this information for their next practice session.

[1193] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[1194] The processing flow will be explained below.

[1195] User-side processing

[1196] Step 1:

[1197] The user launches the recording app on their device and taps the button to start playing.

[1198] Your device will begin recording your performance.

[1199] Step 2:

[1200] The user finishes playing.

[1201] The device stops recording and saves the audio data.

[1202] Step 3:

[1203] The user taps the send button for the recording.

[1204] The device transmits the stored voice data to the server.

[1205] Server-side processing

[1206] Step 1:

[1207] The server receives the voice data transmitted from the terminal.

[1208] Store the received data in the appropriate format.

[1209] Step 2:

[1210] The server analyzes the audio data and evaluates the pitch.

[1211] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[1212] Step 3:

[1213] The server analyzes the audio data for each rhythm pattern.

[1214] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[1215] Step 4:

[1216] The server analyzes the volume of the sound from the audio data.

[1217] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[1218] Step 5:

[1219] The server evaluates the performance expression using a large-scale language model.

[1220] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[1221] Step 6:

[1222] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, and performance expression.

[1223] Feedback includes a numerical rating, a graphical display, and text comments.

[1224] Step 7:

[1225] The server transmits the generated feedback data to the terminal.

[1226] Log the completion of the transmission.

[1227] Terminal side processing

[1228] Step 1:

[1229] The terminal receives the feedback data sent from the server.

[1230] Verify the integrity of the received data and check for any invalid data.

[1231] Step 2:

[1232] The terminal analyzes the feedback data.

[1233] The analysis results are displayed visually and audibly to the user.

[1234] Specific examples

[1235] Take the example of a user playing a piano piece. The user launches a recording app on their smartphone, starts playing, and records. When recording is finished, the audio data is sent to the server. The server analyzes the received audio data and evaluates pitch, rhythm, and dynamics. At the same time, it evaluates the performance expression using a large-scale language model. Feedback is then generated based on the evaluation results and sent to the user's device. The user receives the feedback on their device, understands which parts of their performance need improvement, and can use this information to improve their next practice.

[1236] Example 1

[1237] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1238] Conventional systems supporting the improvement of musical performance lack sufficient analysis of performance data, making it difficult to provide specific feedback. Furthermore, evaluation of expressiveness is subjective, making it difficult to provide specific advice. As a result, it is difficult for users to accurately grasp areas for improvement in their own performance.

[1239] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1240] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, and means for transmitting the feedback, thereby enabling detailed analysis of the user's performance data and providing specific and objective feedback.

[1241] "Performance data" is audio information generated during musical performance.

[1242] "Means for receiving" refers to a device or method for receiving performance data from an external source.

[1243] "Means for analysis" refers to technology for breaking down received performance data into elements such as pitch, rhythm, and dynamics and evaluating them.

[1244] "Means for evaluation" refers to a device or method for determining the quality of a performance based on the analysis results.

[1245] "Performance expression" is a method of musical expression that reflects the performer's emotions and nuances.

[1246] A "large-scale language model" is a type of artificial intelligence that is trained based on large amounts of data in natural language processing.

[1247] A "prompt" is a sentence used to give instructions to an AI model.

[1248] "Feedback" refers to specific advice or comments provided to the user based on the results of analysis and evaluation.

[1249] The "means for transmitting" refers to a device or method for sending the generated feedback to the user's terminal.

[1250] "Visual display means" refers to a device or method that displays feedback on a screen in the form of text or graphs.

[1251] "Auditory display means" refers to a device or method for audibly conveying feedback to the user.

[1252] "Karaoke scoring technology" is a technique for evaluating musical performances by comparing them with standard pitch and rhythm.

[1253] The present invention relates to a system that efficiently supports the improvement of musical performance. The purpose of this system is to objectively evaluate the music performed by a user and provide specific feedback. Detailed embodiments of this system are described below.

[1254] User-side processing

[1255] Users record their performances using devices such as smartphones and tablets. Recording can be started and stopped by user operation. For example, a user can record a performance using a voice recorder app on their smartphone. After recording is finished, the device automatically sends the audio data to the server. If the user has disabled automatic sending, they can also select the recorded file and send it manually.

[1256] Server-side processing

[1257] The server receives the voice data sent from the user's device. This data is saved in an appropriate format (e.g., WAV format). Specifically, the data is received using the HTTP protocol or the FTP protocol.

[1258] The server then analyzes the received audio data using karaoke scoring technology. For pitch analysis, it uses the Python music analysis library "LibROSA," comparing each note with a reference pitch. For rhythm analysis, it uses "pYin," an audio pitch detector tool, to confirm whether it matches the original pattern of the song. For dynamic analysis, it evaluates the amplitude of the audio data.

[1259] The performance expression is evaluated using a large-scale language model. Specifically, "GPT-4" is used, and the prompt is "Please rate the expressiveness of this piano piece played with enjoyment." Based on this, the large-scale language model analyzes the expressiveness of the performance and generates a text comment.

[1260] Generate and send feedback

[1261] The server generates feedback by integrating the results of pitch, rhythm, and dynamic analysis with the evaluation of performance expression using a large-scale language model. The generated feedback is designed to include specific advice. For example, it might say, "Overall, your playing is enjoyable, but some phrases lack dynamic control. Your tempo is off in some places, so we recommend practicing with a metronome."

[1262] Finally, the server sends the generated feedback to the user's device. The feedback is provided both visually and audibly, and the user can check the feedback within the app. This allows the user to specifically understand areas for improvement in their performance and use this information to improve their next practice.

[1263] Specific examples

[1264] For example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on their smartphone. The recording is automatically sent to the server. The server analyzes the received data and checks the accuracy of pitch and rhythm using the "LibROSA" library and the "pYin" tool. Furthermore, to evaluate the performance expression, the server inputs a prompt sentence into the large-scale language model "GPT-4" saying, "Please rate the expressiveness of this enjoyable piano performance," and generates comments about the expressiveness. Based on the analysis results, the server generates feedback such as, "Overall, the performance is enjoyable, but there is a lack of dynamic control in some phrases," and sends it to the user.

[1265] This allows the user to understand specific areas for improvement and use this information in their next practice session.

[1266] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1267] Step 1: Record your performance

[1268] Users use the recording function of their smartphones or tablets to record their own performances. The input is the music played by the user, and the output is audio data (e.g., a WAV file). Specifically, recording begins when the user presses the "Start Recording" button on the device, and ends when the user presses the "Stop Recording" button.

[1269] Step 2: Send your recordings

[1270] The device sends the audio data to the server after recording is complete. The input is the recorded audio data, and the output is a notification that transmission is complete. Specifically, the device uploads the data to the server using the HTTP or FTP protocol. If the user has disabled automatic transmission, they can select the recorded file and press the "Send" button to send it manually.

[1271] Step 3: Receiving music data

[1272] The server receives the voice data sent from the user's device. The input is the voice data sent from the device, and the output is the voice data stored on the server. Specifically, the server saves the received voice data in an appropriate format (e.g., WAV format).

[1273] Step 4: Analyze pitch, rhythm and dynamics

[1274] The server analyzes audio data using karaoke scoring technology. The input is the saved audio data, and the output is the evaluation results of pitch, rhythm, and dynamics. Specifically, the server analyzes pitch using the Python music analysis library "LibROSA" and compares each note with a reference pitch. It also performs rhythm analysis using the audio pitch detector tool "pYin" to confirm whether it matches the original pattern of the song. It also evaluates the dynamics of the entire song through amplitude analysis.

[1275] Step 5: Evaluate the performance

[1276] The server evaluates the performance expression using a large-scale language model. The input is audio data and analysis results, and the output is text feedback about the performance expression. Specifically, the server inputs a prompt sentence, "Please rate the expressiveness of this piano piece played with enjoyment," into the large-scale language model (e.g., "GPT-4") and obtains the generated text comment.

[1277] Step 6: Generate feedback

[1278] The server generates feedback by integrating the analysis results and the evaluation results of performance expression. The inputs are the evaluation results of pitch, rhythm, and dynamics, and the evaluation results of performance expression, and the output is comprehensive feedback. Specifically, the server generates feedback comments containing specific advice based on these results. For example, it may generate a comment such as, "Overall, the performance appears enjoyable, but there is a lack of dynamic control in some phrases."

[1279] Step 7: Submit your feedback

[1280] The server sends the generated feedback to the user's terminal. The input is the generated feedback comment, and the output is the feedback displayed on the user's terminal. In specific operations, the server sends feedback data to the terminal so that the user can visually and audibly confirm the feedback.

[1281] Through the above steps, users can receive detailed analysis results of their own performance and specific advice for improvement, which can be used to improve their performance skills.

[1282] (Application example 1)

[1283] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1284] Current factory robots lack systems that evaluate the accuracy and efficiency of their movements and provide feedback. This leads to accumulated inaccuracies and inefficiencies in robot movements, resulting in reduced productivity. Especially for robots performing complex tasks, even minute errors in movement can cause serious problems, and specific and effective measures to resolve this are needed.

[1285] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1286] In this invention, the server includes means for receiving performance data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for generating feedback based on the evaluation results and the evaluation of performance expression, means for transmitting the feedback, means for acquiring and transmitting robot movement data, means for analyzing the movement data to evaluate accuracy and efficiency, means for generating movement improvement feedback based on the analysis results, and means for transmitting the feedback to the robot and the operator terminal. This makes it possible to evaluate the accuracy and efficiency of the robot's movements in real time and provide specific and effective feedback.

[1287] "Performance data" is audio information of the music played by the user.

[1288] "Means for receiving" refers to a device or module for receiving specific data from the outside and transferring it to the inside.

[1289] "Means for analyzing" refers to devices or modules that process and analyze the received data and evaluate and extract information from it.

[1290] "Means for evaluation" refers to a device or module that judges the performance and accuracy of the analyzed data.

[1291] A "means for generating feedback" is a device or module that provides improvements and advice to users and systems based on the evaluation results.

[1292] A "transmitting means" is a device or module for sending internal data to an external terminal or system.

[1293] "Robot operation data" refers to information relating to the operation of a robot operating in a factory range, such as controlling the position, speed, and force.

[1294] The "means for acquiring and transmitting operation data" refers to a device or module that collects the operation data of the robot and transmits it to a server or the like.

[1295] The "means for evaluating accuracy and efficiency" is a device or module for analyzing the robot's motion data and determining how accurate and efficient the motion is.

[1296] The "means for generating motion improvement feedback" is a device or module that presents improvements to the robot's motion based on the evaluation results of the motion data.

[1297] An "operator terminal" is a computer or mobile device used by a worker in a factory to receive feedback and operational information from the robot.

[1298] This invention relates to a system called "Robot Work Buddy" that accurately evaluates the behavior of robots used in factories and provides efficient feedback. The system collects and analyzes robot behavior data, generates specific improvement advice, and sends it to the robot and its operator.

[1299] Hardware used

[1300] Factory robots: Equipped with various sensors (position, speed, force, etc.) to collect operational data.

[1301] Server: A central processing unit for analyzing received data. It uses Python's SciPy and TensorFlow.

[1302] Operator terminal: A computer or mobile device on which an operator receives feedback.

[1303] Software used

[1304] Data collection module: Has the function of collecting robot operation data and sending it to the server.

[1305] Dynamic control algorithm: Analyzes the robot's motion data and compares it with ideal motion patterns.

[1306] Machine learning algorithms: Using TensorFlow and other algorithms, patterns of motion data are learned to improve accuracy.

[1307] Evaluation and feedback generation module: Generates specific feedback based on the analysis results and sends it to the robot or operator.

[1308] Specific example of operation procedure

[1309] 1. Acquisition and transmission of motion data

[1310] First, robots operating in factories collect data on their position, speed, and force for each task. This data is acquired through sensors built into the robots and then sent to a server.

[1311] 2. Data Analysis

[1312] The server analyzes the received motion data. It uses Python's SciPy to calculate errors in the position data and TensorFlow to learn the motion patterns. Specifically, it compares them with ideal motion patterns and identifies where there are errors.

[1313] 3. Evaluation and feedback generation

[1314] Based on the analysis results, the evaluation and feedback generation module evaluates the accuracy and efficiency of the robot's movements. Based on the evaluation results, it generates specific feedback to help improve the robot's movements. For example, it generates a text comment such as, "The position was off during the bolt tightening operation, so please correct coordinate X."

[1315] 4. Submitting Feedback

[1316] The generated feedback is sent to the robot's control system and to the operator's terminal, where it can be visually confirmed using a visualization tool (e.g., Grafana).

[1317] Adding specific examples

[1318] When a factory robot tightens a bolt, it uses sensors to collect real-time data on the position, speed, and tightening force of each bolt. This data is then sent to a server, which analyzes it using SciPy and TensorFlow. Based on the analysis, specific feedback is generated, such as "The position was off, so please correct coordinate X next time." This feedback is sent to the robot's control system and operator terminal and reflected in the next task.

[1319] Prompt Sentence Examples

[1320] Use a generative AI model to generate feedback using the following prompt:

[1321] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[1322] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1323] Step 1:

[1324] The robot acquires operational data for each task. Specifically, it measures data such as position, speed, and force using internal sensors. This data is collected for each task and temporarily stored in the robot's internal memory.

[1325] (Input: Robot movement, Output: Position data, velocity data, force data)

[1326] Step 2:

[1327] The robot sends the collected operation data to a server using the HTTP protocol, and the data stored in the robot's internal memory is uploaded to the server.

[1328] (Input: position data, velocity data, force data, output: motion data sent to the server)

[1329] Step 3:

[1330] The server analyzes the received motion data, calculates the error in the position data using Python's SciPy, and uses TensorFlow to learn and compare motion patterns. Specifically, it compares the ideal motion pattern with the actual data and identifies where the error exists.

[1331] (Input: movement data, output: position data error, comparison result with movement pattern)

[1332] Step 4:

[1333] The server generates evaluation and feedback based on the analysis results. Using the evaluation and feedback generation module, it evaluates the accuracy and efficiency of the action and generates specific advice for improvement. For example, it generates a text comment such as "The position was off, so please correct coordinate X next time."

[1334] (Input: position data error, comparison result with movement pattern, output: feedback)

[1335] Step 5:

[1336] The server sends the generated feedback to the robot and the operator's terminal, where it is displayed as a specific text comment and also sent to the robot's control system to be reflected in the robot's next task.

[1337] (Input: Feedback, Output: Display on operator terminal, transmission to robot control system)

[1338] Specifically, when a robot working in a factory tightens a bolt, the necessary data is acquired at each step, processed, and the results are fed back to improve the robot's behavior. An example of a generated prompt is as follows:

[1339] "Show the following data as the robot tightens a bolt: position data, velocity data, and force data. Compare these data with the ideal values ​​and generate specific feedback for improvement."

[1340] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1341] This invention relates to a system called "Music Buddy" that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. This system supports the improvement of the user's performance technique by objectively evaluating the music played by the user and providing specific feedback. It is also possible to analyze the user's emotional state and provide more personalized feedback.

[1342] overview

[1343] Users record their performance using a device such as a smartphone or tablet, and simultaneously use the device's camera and microphone to record their facial expressions and vocal tone while they play. The server analyzes the transmitted performance and emotional data, evaluating pitch, rhythm, dynamics, and performance expression, while also recognizing the user's emotional state. Based on the evaluation results, specific feedback is generated and provided to the user.

[1344] User-side processing

[1345] 1. Recording your performance

[1346] The user launches the recording app on their device and taps the button to start playing. At the same time, the device's camera and microphone are activated, recording the user's facial expressions and voice tone. Recording can be started and stopped by the user.

[1347] 2. Sending recording data and emotion data

[1348] Once the recording of the voice and emotion data is complete, the device automatically sends the data to the server, or the user can manually send it.

[1349] Server-side processing

[1350] 1. Receiving Data

[1351] The server receives the voice data and emotion data sent from the device and stores the received data in an appropriate format for subsequent analysis.

[1352] 2. Pitch, rhythm, and dynamic analysis

[1353] The server uses karaoke scoring technology to evaluate the pitch, rhythm, and dynamics of the audio data. For pitch, each note is analyzed and compared with a reference pitch. For rhythm, it checks whether it matches the original pattern of the song. For dynamics, it evaluates the dynamics of the entire song.

[1354] 3. Evaluation of performance expression

[1355] The server evaluates the performance expression using a large-scale language model. Based on the expression instructions entered by the user (e.g., "playing it happily"), it analyzes the audio data and saves the evaluation results.

[1356] 4. Emotion Analysis

[1357] The server uses an emotion engine to analyze the user's emotions. It analyzes facial expressions and vocal tone during the performance to recognize the user's emotional state. The emotion data is used as supplementary information for music evaluation.

[1358] 5. Generate feedback

[1359] The server generates feedback based on the evaluation of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, and text comments. It also includes personalized advice based on the user's emotional state.

[1360] 6. Submitting Feedback

[1361] The server transmits the generated feedback data to the user's terminal, and the transmitted feedback is provided visually and audibly to help the user understand.

[1362] Specific examples

[1363] As a concrete example, consider a user playing a piano piece. The user launches a recording app on their smartphone and taps the start recording button. The device's camera and microphone record the performance while capturing the user's facial expressions and vocal tone. After the performance and recording are complete, the audio and emotional data are sent to the server. The server analyzes the received data, evaluating pitch, rhythm, and dynamics, and using a large-scale language model to evaluate the performance. The emotion engine also analyzes the user's emotional state. Based on the evaluation results, specific feedback is generated and sent to the user's device, such as "Your pitch is generally accurate, but your rhythm is a little fast. Emotion analysis indicates that you seem to be enjoying your performance, but you seem a little nervous." This feedback allows the user to understand which aspects of their performance need improvement and use it to improve their next practice. The user can use the emotional state feedback to play more relaxed.

[1364] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[1365] The processing flow will be explained below.

[1366] User-side processing

[1367] Step 1:

[1368] The user launches the recording app on their device and taps the button to start playing.

[1369] As the device begins recording your performance, the camera and microphone begin capturing your facial expressions and tone of voice.

[1370] Step 2:

[1371] The user finishes playing and taps the stop recording button.

[1372] The device stops recording and saves the performance data and emotional data.

[1373] Step 3:

[1374] The user taps the send button for the recording.

[1375] The device transmits the stored voice data and emotion data to the server.

[1376] Server-side processing

[1377] Step 1:

[1378] The server receives the voice data and emotion data transmitted from the terminal.

[1379] The received data is stored in an appropriate format for analysis.

[1380] Step 2:

[1381] The server analyzes the audio data and evaluates the pitch.

[1382] The audio data is divided into notes, compared with a reference pitch, and the evaluation results are saved.

[1383] Step 3:

[1384] The server analyzes the audio data for each rhythm pattern.

[1385] Check whether the rhythm matches the original pattern of the song and save the evaluation results.

[1386] Step 4:

[1387] The server analyzes the volume of the sound from the audio data.

[1388] Check the dynamic changes throughout the song, evaluate whether there are any unnatural fluctuations, and save the evaluation results.

[1389] Step 5:

[1390] The server evaluates the performance expression using a large-scale language model.

[1391] Based on the expression instructions entered by the user (e.g., "enjoyably"), the voice data is analyzed and the evaluation results are saved.

[1392] Step 6:

[1393] The server uses an emotion engine to analyze the user's emotions.

[1394] The system analyzes facial expressions and vocal tone during performance, recognizes the user's emotional state, and saves the evaluation results.

[1395] Step 7:

[1396] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1397] Feedback includes numerical ratings, graphical displays, and text comments, as well as personalized advice based on the user's emotional state.

[1398] Step 8:

[1399] The server transmits the generated feedback data to the user's terminal.

[1400] Log the completion of the transmission.

[1401] Terminal side processing

[1402] Step 1:

[1403] The terminal receives the feedback data sent from the server.

[1404] Verify the integrity of the received data and check for any invalid data.

[1405] Step 2:

[1406] The terminal analyzes the feedback data.

[1407] The analysis results are displayed visually and audibly to the user.

[1408] Specific examples

[1409] Let's take the example of a user playing a piano piece. First, the user launches the recording app on their smartphone and taps the start recording button. The device starts recording, and at the same time, the camera and microphone begin to capture the user's facial expressions and vocal tone. When the performance is finished, the user taps the stop recording button, and the recording data and emotional data are saved. Then, the user taps the send button for the recording data, and the voice data and emotional data are sent to the server.

[1410] The server receives the audio data and emotion data, analyzing pitch, rhythm, dynamics, and performance expression, respectively. At the same time, the emotion engine analyzes the user's facial expressions and vocal tone to evaluate their emotional state. Based on the evaluation results, specific feedback such as "The pitch is good, but the rhythm is a little fast. The performance seems enjoyable, but there is also a sense of tension" is generated and sent to the user's device. The user can receive the feedback on their device, understand where their performance needs improvement, and use the feedback to improve their next practice.

[1411] This allows users to effectively improve their musical performance skills while also maintaining motivation and continuing to practice in a fun way.

[1412] Example 2

[1413] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1414] Conventional music performance improvement support systems have difficulty not only evaluating a user's performance technique but also providing personalized feedback based on the user's emotional state. Furthermore, in order to improve the accuracy of the evaluation results, they lack the functionality to evaluate performance expression using large-scale language models and analyze the user's emotional state.

[1415] The specification processing by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording the user's performance, means for receiving performance data and emotional data, means for analyzing the performance data to evaluate pitch, rhythm, and dynamics, means for evaluating performance expression based on the performance data, means for analyzing the emotional data to evaluate the user's emotional state, means for generating feedback based on the evaluation result and the emotional state, and means for transmitting the feedback. This makes it possible to evaluate the user's performance technique as well as provide detailed and personalized feedback based on the user's emotional state.

[1416] "Means for recording a user's performance" refers to a device or software that digitally records the audio of a user playing music.

[1417] "Means for receiving performance data and emotional data" refers to a function for receiving audio data and emotional data sent from a user's terminal via a network to a device such as a server.

[1418] "Means for analyzing performance data to evaluate pitch, rhythm, and dynamics" refers to algorithms and software that automatically analyze the basic elements of music—pitch, rhythm, and dynamics—and evaluate their accuracy and expressiveness.

[1419] "Means for evaluating performance expression based on performance data" refers to technologies such as large-scale language models for analyzing and evaluating the expressive intent behind a user's performance.

[1420] "Means for analyzing emotional data and assessing a user's emotional state" refers to an engine or software that identifies the user's emotional state from the tone of voice and facial expressions in a recording, and analyzes and assesses it.

[1421] "Means for generating feedback based on evaluation results and emotional state" refers to systems and algorithms for generating feedback to be provided to a user based on the analyzed performance data and emotional data.

[1422] "Means for sending feedback" refers to communication functions and protocols for sending the generated feedback data to the user's terminal.

[1423] The present invention is a system that supports the improvement of musical performance, and is combined with an emotion engine that recognizes the user's emotions. The system configuration includes a user terminal, a server, and a communication function for exchanging data between them. Specific embodiments of the system are described below.

[1424] Users record their performances using devices such as smartphones or tablets. When they launch the recording app and tap the start button, the device's camera and microphone are activated. This allows the user's performance audio, as well as their facial expressions and vocal tone, to be recorded. Recording can be started and stopped by the user.

[1425] Once recording is complete, the device sends the voice data and emotion data to the server. The server receives the data and stores it in the appropriate format. The server then analyzes the received voice data in terms of pitch, rhythm, and dynamics. Karaoke scoring technology (e.g., JOYSOUND's technology) is used for this analysis. Pitch analysis uses a MIDI analysis library, and analysis is performed for each note. Rhythm analysis is performed using a tempo analysis library, and dynamics analysis is performed using waveform analysis.

[1426] Furthermore, the server evaluates the performance expression using a large-scale language model (e.g., OpenAI's GPT-4). Based on the expression instructions specified by the user (e.g., "enjoyably"), the audio data is converted into text and analyzed by the model. The evaluation results of the performance expression are saved.

[1427] In addition, the server uses an emotion engine to analyze the user's emotional state while playing. It uses facial and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. Emotional data is used as supplementary information for music evaluation.

[1428] Based on the evaluation results, the server generates specific feedback. The feedback may include numerical evaluations, graph displays, text comments, etc. In addition, personalized advice based on the user's emotional state may also be provided. The generated feedback is sent from the server to the device. The device displays the sent feedback, providing it visually and audibly. This makes it easier for the user to understand the evaluation results of their performance. Specific examples of feedback that may be considered include the following:

[1429] "The pitch is generally accurate, but the rhythm is a little fast. Emotional analysis indicates that the performance is enjoyable, but with some tension."

[1430] Through this feedback, users can understand which parts of their performance need improvement and use that information for their next practice. Users can also use the feedback on their emotional state to help them play more relaxed.

[1431] An example of a prompt sentence that could be input to a generative AI model would be, "Please rate my piano playing."

[1432] This system allows users to effectively improve their musical performance skills and continue practicing while maintaining motivation and having fun.

[1433] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1434] Step 1:

[1435] The user launches the recording app on their device and taps the recording preparation button. This action causes the device to check the status of the camera and microphone and check that they are working properly. It notifies the user that they are ready. The input is the user's operation, and the output is a notification that the camera and microphone are ready to start.

[1436] Step 2:

[1437] The user taps the "Start Recording" button to begin playing. At this point, the device simultaneously activates its camera and microphone, recording the user's performance, facial expressions, and vocal tone. The input is the "Start Recording" button and the user's performance, and the output is the recorded audio and video data.

[1438] Step 3:

[1439] When the recording is finished, the user taps the "Stop Recording" button. The device responds by stopping the recording and saving the data. The input is the "Stop Recording" operation, and the output is the saved audio and video data.

[1440] Step 4:

[1441] The device prepares to send the recorded data (audio files) and emotion data (video files and audio analysis results) to the server. When the user presses the "Send" button in the confirmation dialog, the data is sent to the server. The input is the user's confirmation of sending, and the output is the data being sent to the server.

[1442] Step 5:

[1443] The server receives the voice data and emotion data sent from the device and stores them in the respective databases. The input is the received data, and the output is the stored data.

[1444] Step 6:

[1445] The server analyzes the audio data and evaluates pitch, rhythm, and dynamics. It uses a MIDI analysis library for pitch analysis, a tempo analysis library for rhythm analysis, and waveform analysis for dynamics analysis. The input is the audio data, and the output is the analysis results for each evaluation item.

[1446] Step 7:

[1447] The server evaluates performance expressions using a large-scale language model. Based on the user's instructions, the audio data is converted into text, which is then analyzed by the large-scale language model. The input is the audio data and instructions, and the output is the evaluation result of the performance expression.

[1448] Step 8:

[1449] The server uses an emotion engine to analyze the user's emotional state. It uses face and voice recognition technologies to estimate emotions from the user's facial expressions and tone of voice. The input is video and audio data, and the output is the emotion analysis results.

[1450] Step 9:

[1451] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state. The feedback includes numerical evaluation, graph display, text comments, and personalized advice. The input is each evaluation result, and the output is feedback data.

[1452] Step 10:

[1453] The server sends the generated feedback data to the terminal. The input is the feedback data, and the output is the data sent to the user's terminal.

[1454] Step 11:

[1455] The device receives the transmitted feedback data and provides it to the user visually and audibly, allowing the user to understand the feedback and use it for their next practice. The input is the feedback data, and the output is the provision of feedback to the user.

[1456] (Application example 2)

[1457] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1458] Conventional systems for supporting the improvement of musical performance focus on technical evaluations such as pitch, rhythm, and dynamics, but have the problem of being unable to provide feedback that takes into account the performer's emotional state. Another problem is that the feedback users receive is uniform for each individual performance, making it difficult to obtain specific and personalized advice that adapts to the user's emotional state. This makes it difficult for users to understand areas for improvement that correspond to their own emotional state, resulting in the problem of inefficient improvement of performance techniques.

[1459] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1460] In this invention, the server includes means for receiving performance data, means for analyzing the performance data and evaluating pitch, rhythm, and dynamics, means for evaluating performance expression based on the evaluation results, means for analyzing the user's emotional state, means for personalizing feedback based on the emotional state, means for generating the feedback, and means for transmitting the feedback, thereby enabling the provision of specific and personalized feedback that takes into account not only technical evaluation but also the user's emotional state.

[1461] The "means for receiving performance data" is a mechanism for recording the music played by the user and transmitting that data to the server.

[1462] The "means for evaluating pitch" is a mechanism that analyzes the pitch of the performance data and measures how well it matches a reference pitch.

[1463] The "means for evaluating rhythm" is a mechanism for analyzing the timing of the sounds in the performance data and measuring the degree of agreement with a specified rhythm pattern.

[1464] The "means for evaluating dynamics" is a system that analyzes the changes in volume of performance data and evaluates the dynamics of the entire performance.

[1465] The "means for evaluating performance expression" is a system for evaluating the expressiveness and technique of a user's performance based on the evaluation results of pitch, rhythm, and dynamics.

[1466] The "means for analyzing emotional state" is a mechanism for analyzing the user's facial expressions and vocal tone while playing and recognizing the emotional state at that time.

[1467] The "means for generating feedback" is a mechanism that provides the user with specific and personalized improvements based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1468] The "means for transmitting feedback" is a mechanism for transmitting the generated feedback to the user's terminal and displaying it visually and audibly.

[1469] Overall system overview

[1470] This invention is a system to help users improve their musical performance by recording their performance and analyzing their emotions through recording their facial expressions and vocal timbre. It is realized using the following hardware and software.

[1471] Hardware and software used

[1472] Smartphone or tablet (e.g. iPhone, Android device)

[1473] PC or server (e.g. cloud server)

[1474] Devices with built-in cameras and microphones

[1475] Programming language: Swift (iOS), Kotlin (Android)

[1476] API: Google Cloud Speech-to-Text, Google Vision API (sentiment analysis)

[1477] Backend: Node.js

[1478] Database: Firebase

[1479] Main features of the system

[1480] Users first record their performance using an application installed on their smartphone or tablet. During recording, the device's camera and microphone simultaneously operate to collect the user's facial expressions and vocal tone. The collected data is then sent to a cloud server in real time or after recording.

[1481] Music evaluation and emotion analysis process

[1482] The server receives the transmitted performance data and evaluates the pitch, rhythm, and dynamics. Pitch evaluation involves analyzing how closely the pitch of each note matches the reference pitch. Rhythm evaluation involves checking synchronization with the specified pattern. Dynamic evaluation involves analyzing the dynamics of the entire performance. In addition to these technical evaluations, the server uses the Google Vision API to analyze facial expression data during performance and recognize the user's emotional state.

[1483] Generating and Providing Feedback

[1484] Based on the evaluation results, specific and personalized feedback is generated for the user, including evaluation results on pitch, rhythm, and dynamics, as well as advice based on emotional state. The generated feedback is sent to the user's device in the form of visual and audio.

[1485] Examples and prompts

[1486] As a concrete example, let's consider a situation where a user plays the guitar and then receives feedback from an app. First, the user launches the app and taps the start recording button. After finishing playing, the user taps the stop recording button, which sends the data to the server. After analysis is performed on the server side, the following feedback is displayed on the user's device:

[1487] "The pitch is generally accurate, but the rhythm is a little fast. The emotional analysis shows that you are playing with enjoyment, but there are some parts that seem tense. Please try playing a little more relaxed."

[1488] Example prompt to launch the app:

[1489] "Select your instrument and start recording. Relax and play, as your facial expressions will be recorded."

[1490] In this way, the system of the present invention allows users to understand areas for improvement in their performance, both technically and emotionally, through specific and personalized feedback.

[1491] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1492] Step 1:

[1493] Recording of performance

[1494] The user launches the application on their smartphone or tablet and taps the recording start button.

[1495] Input: User's music, facial expressions, and vocal tone

[1496] How it works: The device's microphone records the music, and the camera records the user's facial expressions while they play. The audio and facial expression data are then stored on the device.

[1497] Output: Recorded performance data and emotional data (facial expressions and vocal timbre)

[1498] Step 2:

[1499] Sending recording data and emotional data

[1500] When the user taps the recording stop button, the recorded performance data and emotion data are sent to the server.

[1501] Input: Recorded performance data, emotional data

[1502] How it works: Data is sent from the device to a server over the internet.

[1503] Output: Performance data and emotion data stored on the server

[1504] Step 3:

[1505] Pitch, rhythm and dynamic analysis

[1506] The server evaluates the pitch, rhythm, and dynamics based on the received performance data.

[1507] Input: Received performance data

[1508] How it works: A server-based analysis program compares pitch with a reference value and analyzes rhythm and dynamics. Specifically, it analyzes audio data using the Google Cloud Speech-to-Text API.

[1509] Output: Pitch, rhythm, and dynamics evaluation results

[1510] Step 4:

[1511] Evaluation of performance expression

[1512] The server evaluates the performance expression based on the evaluation results of pitch, rhythm, and dynamics.

[1513] Input: Pitch, rhythm, and dynamic evaluation results

[1514] How it works: The performance expression evaluation algorithm in the server analyzes the data and performs a comprehensive evaluation of the performance expression.

[1515] Output: Evaluation results of performance expression

[1516] Step 5:

[1517] Emotion Analysis

[1518] The server analyzes the received emotion data and evaluates the user's emotional state.

[1519] Input: Received emotional data (facial expressions, tone of voice)

[1520] How it works: It uses the Google Vision API to analyze facial expression data and recognize the user's emotional state.

[1521] Output: Evaluation result of the user's emotional state

[1522] Step 6:

[1523] Generate feedback

[1524] The server generates feedback based on the evaluation results of pitch, rhythm, dynamics, performance expression, and emotional state.

[1525] Input: Pitch, rhythm, dynamics, performance expression, emotional state evaluation results

[1526] How it works: The server's feedback generator combines the results of each assessment to create specific, personalized feedback.

[1527] Output: Generated feedback data

[1528] Step 7:

[1529] Send Feedback

[1530] The server transmits the generated feedback to the user's terminal.

[1531] Input: Generated feedback data

[1532] Operation: The server sends feedback data to the terminal, which displays it visually and audibly to the user.

[1533] Output: Feedback displayed on the user's device

[1534] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1535] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1536] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1537] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1538] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1539] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1540] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1541] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1542] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1543] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1544] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1545] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1546] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1547] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1548] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1549] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1550] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1551] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1552] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1553] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1554] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1555] The following is further disclosed regarding the above embodiment.

[1556] (Claim 1)

[1557] means for receiving performance data;

[1558] means for analyzing the performance data and evaluating pitch, rhythm, and dynamics;

[1559] a means for evaluating performance expression based on the evaluation results;

[1560] a means for generating feedback based on the evaluation result and the evaluation of the performance expression;

[1561] means for transmitting said feedback;

[1562] A system including:

[1563] (Claim 2)

[1564] 2. The system according to claim 1, wherein the means for receiving the performance data is means for recording the user's performance.

[1565] (Claim 3)

[1566] 2. The system according to claim 1, wherein the means for sending the feedback comprises means for displaying the feedback visually and audibly on the user's terminal.

[1567] "Example 1"

[1568] (Claim 1)

[1569] means for receiving performance data;

[1570] means for analyzing the performance data and evaluating pitch, rhythm, and dynamics;

[1571] a means for evaluating performance expression based on the evaluation results;

[1572] a means for generating feedback based on the evaluation result and the evaluation of the performance expression;

[1573] means for transmitting said feedback;

[1574] A system including:

[1575] (Claim 2)

[1576] 2. The system according to claim 1, wherein the means for receiving the performance data is means for recording the performance of the user.

[1577] (Claim 3)

[1578] 2. The system according to claim 1, wherein the means for sending the feedback is means for displaying the feedback visually and audibly on the user's terminal.

[1579] (Claim 4)

[1580] 2. The system according to claim 1, wherein the performance data is analyzed for pitch, rhythm, and dynamics using karaoke scoring technology.

[1581] (Claim 5)

[1582] 2. The system according to claim 1, wherein the evaluation of the performance expression is performed using a large-scale language model and a prompt sentence.

[1583] "Application Example 1"

[1584] (Claim 1)

[1585] means for receiving performance data;

[1586] means for analyzing the performance data and evaluating pitch, rhythm, and dynamics;

[1587] a means for evaluating performance expression based on the evaluation results;

[1588] a means for generating feedback based on the evaluation result and the evaluation of the performance expression;

[1589] means for transmitting said feedback;

[1590] means for acquiring and transmitting robot motion data;

[1591] means for analyzing said operational data to assess accuracy and efficiency;

[1592] means for generating behavior improvement feedback based on the analysis results;

[1593] means for transmitting said feedback to the robot and to an operator terminal;

[1594] A system including:

[1595] (Claim 2)

[1596] 2. The system according to claim 1, wherein the means for receiving the performance data is means for recording the user's performance.

[1597] (Claim 3)

[1598] 2. The system according to claim 1, wherein the means for sending the feedback comprises means for displaying the feedback visually and audibly on the user's terminal.

[1599] "Example 2: Combining Emotion Engines"

[1600] (Claim 1)

[1601] means for recording a user's performance;

[1602] means for receiving the performance data and emotion data;

[1603] means for analyzing the performance data and evaluating pitch, rhythm, and dynamics;

[1604] means for evaluating performance expressions based on the performance data;

[1605] means for analyzing the emotion data to assess the user's emotional state;

[1606] means for generating feedback based on the evaluation results and the emotional state;

[1607] means for transmitting said feedback;

[1608] A system including:

[1609] (Claim 2)

[1610] 10. The system of claim 1, wherein the means for sending feedback comprises means for providing visual and audio feedback to the user's computing device.

[1611] (Claim 3)

[1612] 2. The system of claim 1, wherein the means for evaluating the performance expression uses a large-scale language model.

[1613] "Application example 2 when combining emotion engines"

[1614] (Claim 1)

[1615] means for receiving performance data;

[1616] means for analyzing the performance data and evaluating pitch, rhythm, and dynamics;

[1617] a means for evaluating performance expression based on the evaluation results;

[1618] a means for generating feedback based on the evaluation result and the evaluation of the performance expression;

[1619] means for analyzing the emotional state of a user;

[1620] means for personalizing feedback based on said emotional state;

[1621] means for transmitting said feedback;

[1622] A system including:

[1623] (Claim 2)

[1624] 2. The system according to claim 1, wherein the means for receiving the performance data is means for recording the user's performance and collecting emotion data using a face recognition camera.

[1625] (Claim 3)

[1626] 2. The system according to claim 1, wherein the means for sending the feedback is means for visually and audibly displaying the feedback on the user's terminal and means for indicating areas for improvement while emphasizing specific aspects. [Explanation of symbols]

[1627] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving performance data; means for analyzing the performance data and evaluating pitch, rhythm, and dynamics; a means for evaluating performance expression based on the evaluation results; a means for generating feedback based on the evaluation result and the evaluation of the performance expression; means for transmitting said feedback; A system including:

2. 2. The system according to claim 1, wherein the means for receiving the performance data is means for recording the user's performance.

3. 2. The system of claim 1, wherein the means for sending the feedback comprises means for visually and audibly displaying the feedback on the user's terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A