System

The system objectively evaluates musical performance by analyzing pitch, rhythm, and expressiveness, offering detailed feedback to enhance practice effectiveness.

JP2026022553APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124070
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Children find it difficult to objectively evaluate their musical instrument performance, leading to tedious practice and lost motivation due to lack of effective feedback methods.

Method used

A system that records and analyzes a user's musical performance, evaluating pitch, rhythm, dynamics, and expressiveness using frequency analysis, note timing detection, volume fluctuations, and a generative AI model, providing objective feedback and practice suggestions.

Benefits of technology

Enables users to receive accurate and detailed evaluations, improving their musical skills by identifying specific areas for improvement and enhancing practice efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022553000001_ABST
    Figure 2026022553000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for recording a sound of a musical instrument played by a user with high sound quality; means for transmitting the recorded sound data to a server; means for analyzing the sound data and evaluating a pitch, rhythm, dynamics, and expressiveness; means for proposing a phrase necessary for improvement based on an analysis result; and means for transmitting an evaluation result to a terminal of the user and visually displaying the evaluation result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Accurate pitch, rhythm, dynamics, and musical expression are essential for improving musical instrument performance. However, children, in particular, are still developing the ability to objectively listen to and evaluate their own performance, and often find it difficult to accept feedback from parents or teachers. As a result, repeated practice becomes tedious and motivation is lost. Conventional techniques lack effective solutions to these problems. The present invention aims to solve these problems and provide users with an enjoyable musical experience. [Means for solving the problem]

[0005] The present invention provides a means for recording a user's voice with high quality while playing an instrument and a means for transmitting the recorded voice data to a server. The server further includes a means for analyzing the voice data and accurately evaluating pitch, rhythm, dynamics, and expressiveness. The system also includes a means for suggesting phrases necessary for improvement based on the analysis results, and a means for transmitting the evaluation results to the user's device and visually displaying them. Frequency analysis is used to detect pitch, note start and end timing is analyzed to detect rhythm, and volume fluctuations are analyzed to detect dynamics. Expressiveness is also evaluated using a generative AI model to evaluate the emotional expression of the performance. The system also includes a means for the user to replay the performance, record new voice data, and re-evaluate. This allows users to receive objective evaluations and specific practice suggestions, enabling them to improve efficiently.

[0006] Below are definitions of important words.

[0007] "User" refers to a person who plays a musical instrument and uses the system of the present invention to evaluate and attempt to improve their performance.

[0008] "Device" refers to the device used by the User to record their performance, transmit the data to the Server, and receive the evaluation results, including smartphones, tablets, and PCs.

[0009] The term "server" refers to a computer system that receives voice data sent from a user, analyzes it, and sends the results to a terminal.

[0010] "Recording means" refers to a function that records the user's musical instrument performance in high quality and saves it as digital audio data.

[0011] "Transmission means" refers to the communication function for transmitting the recorded voice data from the terminal to the server, which is usually data transmission via the Internet.

[0012] "Analysis means" refers to algorithms and processes for analyzing the transmitted audio data and evaluating pitch, rhythm, dynamics, and expressiveness.

[0013] "Pitch" refers to the element that compares the played note with the pitch written in the musical score and evaluates the degree of match.

[0014] "Rhythm" refers to the element that evaluates the accuracy of the start and end timing of played notes by comparing them with a reference rhythm.

[0015] "Dynamics" refers to an element that analyzes fluctuations in the volume of the sound being played and evaluates whether the appropriate dynamics are being expressed as instructed in the musical score.

[0016] "Expressiveness" refers to an element that evaluates the ability to convey emotion and nuance throughout the performance.

[0017] "Generative AI model" refers to the artificial intelligence technology used to evaluate the emotional expression and consistency of a performance.

[0018] "Evaluation results" refers to evaluation data regarding pitch, rhythm, dynamics, and expressiveness generated by the analysis means.

[0019] "Suggestion means" refers to a function that suggests phrases and practice methods necessary for the user to improve based on the evaluation results.

[0020] "Means for visual display" refers to a function that displays the evaluation results as graphs or numerical values ​​on the terminal screen and provides feedback to the user.

[0021] The "re-evaluation means" refers to a process in which the user performs the performance again, records new audio data, and re-evaluates the performance using the analysis means. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] MODE FOR CARRYING OUT THE INVENTION

[0044] The present invention is a system that records a user's performance of a musical instrument, transmits the audio data to a server, and analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness of the user. The system of the present invention allows users to objectively evaluate their own performance and find effective practice methods. The following describes in detail an embodiment of the present invention.

[0045] Overview of program processing

[0046] 1. Receiving and recording audio input

[0047] When the user starts playing, the device launches a recording application. The device records the performance in high quality and saves the recorded data as a digital audio file. During recording, real-time audio data is stored in memory.

[0048] 2. Sending audio data

[0049] After the recording is complete, the device uploads the recorded audio file to a specific server URL via an HTTP request, and the data is securely encoded and transmitted.

[0050] 3. Analysis of audio data

[0051] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially performs pitch analysis, rhythm analysis, dynamic analysis, and expressiveness analysis.

[0052] Pitch analysis:

[0053] The server uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compare it with the pitch written in the sheet music.

[0054] Rhythm analysis:

[0055] The server detects the start and end timing of notes on the time axis and evaluates the degree of agreement with the reference rhythm.

[0056] Dynamic analysis:

[0057] The server detects fluctuations in audio volume and analyzes whether the dynamics of each note are as specified in the musical score.

[0058] Expressiveness analysis:

[0059] The server uses a generative AI model to evaluate the emotional expression of a performance, which is then compared with existing data on famous performances to calculate the degree of match in expressiveness.

[0060] 4. Generating evaluation results

[0061] Based on the analysis results, the server calculates scores for pitch, rhythm, dynamics, and expressiveness, and generates an overall evaluation, which includes numerical scores and graphs for each element.

[0062] The server also lists suggestions for specific phrase practice to help the user improve and adds them to the evaluation results.

[0063] 5. Submitting and displaying evaluation results

[0064] The server then sends the generated evaluation results to the terminal, securely using HTTPS.

[0065] The device analyzes the received evaluation results and displays them visually on the user interface, providing feedback in the form of graphs, numbers, and text, allowing the user to check the detailed evaluation content.

[0066] Specific examples

[0067] If a user were to play "Twinkle Twinkle Little Star" on the piano, the performance would be recorded and analyzed and evaluated using the following procedure.

[0068] The user starts playing the piano and the device records the audio.

[0069] After the performance is finished, the device will stop recording and upload the audio file to the server.

[0070] The server analyzes the received audio data and evaluates it for pitch, rhythm, dynamics, and expressiveness. For example, it may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good.

[0071] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[0072] The server sends the evaluation results and suggestions to the device, which displays them on the user interface. The user can then visually check the feedback, replay the performance, and self-evaluate and try to improve.

[0073] In this way, the system of the present invention provides the user with objective evaluations and specific practice suggestions, and supports improvement in musical instrument playing.

[0074] The processing flow will be explained below.

[0075] Step 1:

[0076] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording is performed in real time, and the audio data is stored in the device's memory. During recording, the device displays the audio waveform and provides a visual indication of the progress.

[0077] Step 2:

[0078] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After saving is complete, the user selects the "Send" button to send the recorded data to the server.

[0079] Step 3:

[0080] In response to user operations, the device sends an HTTP request and uploads the recorded audio file to the server. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[0081] Step 4:

[0082] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of the analysis.

[0083] Step 5:

[0084] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[0085] Step 6:

[0086] Next, the server performs rhythm analysis, detecting the start and end timing of notes on the time axis and comparing them with a reference rhythm template. This evaluates the rhythmic accuracy of the performance. The analysis results are recorded as data including timestamps.

[0087] Step 7:

[0088] The server continues dynamic analysis, analyzing the fluctuations in audio volume to detect the sound pressure level of each note. It compares this with the dynamics instructions in the score to evaluate the appropriateness of the dynamics. The analysis results are recorded as a numerical sound pressure level.

[0089] Step 8:

[0090] Finally, the server analyzes the performance's expressiveness. It uses a generative AI model to evaluate the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data and calculates the degree of agreement in expressiveness. The analysis results are recorded as an emotional score.

[0091] Step 9:

[0092] The server combines all the analysis results to generate an overall evaluation. It also combines scores for pitch, rhythm, dynamics, and expressiveness to create visualizations (e.g., radar charts). It also lists specific phrase practice suggestions to help the user improve and adds them to the evaluation results.

[0093] Step 10:

[0094] The server then sends the generated evaluation results to the device. The transmission is secure using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as specific suggestions for improvement.

[0095] Step 11:

[0096] The device analyzes the received evaluation results and displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check the detailed evaluation. The device also provides "replay" and "re-evaluation" buttons, allowing users to replay and receive a new evaluation.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] Conventional methods for evaluating musical instrument performance require the presence of an instructor with specialized knowledge, making it difficult to obtain objective and quantitative evaluations. Furthermore, when users self-evaluate, the evaluation tends to be subjective, making it difficult to find effective practice methods. Furthermore, it is difficult to evaluate the emotional expression and subtle dynamics of a performance, limiting the overall improvement of performance skills.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes a means for analyzing audio data and evaluating scale, rhythm, volume changes, and emotional expression; a means for suggesting practice content necessary for improvement based on the analysis results; and a means for transmitting the evaluation results to the user's device and visually displaying them. This enables objective and quantitative evaluation of the user's performance and provides effective practice methods. Furthermore, using a generative AI model to evaluate expressiveness enables comprehensive evaluation, including emotional expression, helping to improve the user's performance skills.

[0102] "Audio data" is information that digitally records the sound of an instrument played by a user.

[0103] A "communication device" is a device for transmitting and receiving acoustic data via a network.

[0104] "Scale" refers to the reference frequency of each note written on a musical score.

[0105] "Rhythm" is a temporal pattern based on the timing of the start and end of musical notes.

[0106] "Volume change" refers to fluctuations in the volume of the audio during performance.

[0107] "Emotional expression" is an indicator that evaluates the emotions and nuances in a performance.

[0108] A "generative AI model" is an artificial intelligence model that has been trained in advance using machine learning algorithms.

[0109] "Visually displaying" means providing analysis results and feedback to the user in a form that can be seen as graphs or text.

[0110] "Practice content" refers to specific playing phrases and tasks that the user needs to improve.

[0111] MODE FOR CARRYING OUT THE INVENTION

[0112] The present invention is a system that supports users in improving their musical performance skills by recording, analyzing, and evaluating the sounds of musical instruments played by the user with high accuracy. The following describes in detail a method for specifically implementing the system of the present invention.

[0113] System Configuration

[0114] Hardware:

[0115] 1. Terminal: A device operated by a user, such as a smartphone, tablet, or computer.

[0116] 2. Server: A remote server for performing analytical processing, using a computer system with high-performance processing capabilities.

[0117] software:

[0118] 1. Recording application: An application that starts when the user starts playing and records high-quality audio. For example, audio recording applications such as "Audacity" or "Pro Tools" can be used.

[0119] 2. Audio analysis program: A program that runs on the server and analyzes audio data. This analysis uses audio processing tools such as "sox" or "FFmpeg."

[0120] 3. Generative AI model: An AI model equipped with a machine learning algorithm for analyzing emotional expressions. For example, a model pre-trained using "TensorFlow" or "PyTorch" is used.

[0121] Processing Overview

[0122] The user plays an instrument and records the sound on the device. The recorded sound data is then uploaded to the server by the user's operation. The server analyzes the received sound data and evaluates the scale, rhythm, volume changes, and emotional expression. Based on the evaluation results, the system suggests specific practice content and sends the evaluation results to the device, where they are visually displayed.

[0123] Specific actions

[0124] 1. Receiving and recording audio input:

[0125] When the user starts playing, an audio recording application is launched on the device. The sound is recorded in high quality, and the audio data is temporarily stored in memory as it is being recorded. When recording is finished, the audio data is saved as a digital audio file. For example, it may be saved with a file name such as "recording.wav."

[0126] 2. Sending audio data:

[0127] After recording is complete, the device uploads the recorded audio file to a specific server URL via a user-initiated HTTP POST request, with the data encoded and securely transmitted. For example, a request can be sent to the URL "https: / / example.com / upload."

[0128] 3. Analysis of audio data:

[0129] The server stores the received audio data in storage and first performs noise reduction, using audio processing tools such as "sox" for noise filtering.

[0130] Next, the scale is analyzed using FFT (Fast Fourier Transform) to detect the reference frequency of each note.

[0131] The system analyzes rhythm by detecting the start and end timing of notes, and analyzes dynamics by detecting fluctuations in voice volume.

[0132] Emotional expression is evaluated using a pre-trained generative AI model, which compares the performance with existing master performance data to calculate the degree of agreement between the performance's emotional expression.

[0133] 4. Generating evaluation results:

[0134] The server then calculates scores for scale, rhythm, volume, and emotional expression based on the analysis results, generating an overall evaluation that includes a numerical score and graphs, and suggests specific phrases for the user to practice.

[0135] 5. Submitting and viewing the evaluation results:

[0136] The server sends the evaluation results to the device, which then visually displays them on the user interface, for example, by displaying the scale and rhythm evaluation scores in bar graphs or pie charts, allowing the user to check the evaluation of their own performance.

[0137] Example operation

[0138] When a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following steps:

[0139] The user starts playing the piano and the device records the audio.

[0140] After the performance is finished, the device stops recording and uploads the audio file "kirakira_sei.wav" to the server.

[0141] The server analyzes the received audio data and scores it with 85 points for pitch, 70 points for rhythm, 90 points for dynamics, and 80 points for emotional expression.

[0142] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[0143] The server sends the evaluation results and suggestions to the device, which displays them on the user interface, allowing the user to visually confirm the feedback and try playing again.

[0144] This system allows users to easily obtain evaluations of their instrument performance and areas for improvement, effectively improving their own playing skills.

[0145] Prompt Sentence Examples

[0146] "Please analyze the following audio data and rate it for pitch, rhythm, volume changes, and emotional expression:"

[0147] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0148] Step 1: Receiving and Recording Audio Input

[0149] The device launches a recording application to receive audio input, and when the user starts playing an instrument, it records the audio in real time with high quality.

[0150] Input: Audio of the user playing an instrument.

[0151] Data processing: The device temporarily stores the audio data being recorded in memory.

[0152] Output: Once recording is complete, the audio data is saved as a digital audio file (e.g., recording.wav).

[0153] Step 2: Sending audio data

[0154] After the recording is completed, the device uploads the recorded audio file to a specific server URL according to the user's operation.

[0155] Input: Recorded digital audio files.

[0156] Data processing: Encode the audio data and send it to the server using an HTTP POST request.

[0157] Output: The audio data is securely uploaded to the server. For example, send a request to the URL "https: / / example.com / upload".

[0158] Step 3: Analyzing the audio data

[0159] The server first stores the received audio data in storage.

[0160] Input: Uploaded digital audio files.

[0161] Acoustic analysis:

[0162] Data processing: Perform noise reduction and use audio processing tools such as "sox" to remove unwanted noise.

[0163] Data calculation: Analyzes the musical scale using FFT (Fast Fourier Transform) and detects the reference frequency of each note. Detects the start and end timing of notes on the time axis and analyzes rhythm. Detects fluctuations in voice volume and analyzes dynamics.

[0164] Emotional expression rating:

[0165] Data computation: Evaluate the emotional expression of a performance using a generative AI model (e.g., a pre-trained model using TensorFlow or PyTorch). Compare with existing renowned performance data to calculate the degree of consistency of expressiveness.

[0166] Output: The analysis results include scores for pitch, rhythm, volume changes, and emotional expression.

[0167] Step 4: Generate evaluation results

[0168] Based on the analysis results, the server calculates scores for scale, rhythm, volume changes and emotional expression to generate an overall rating.

[0169] Input: Audio analysis results data (scale, rhythm, volume changes, emotional expressions).

[0170] Data calculation: A score is calculated based on each analysis result, and an overall evaluation is made. For example, the scale score is evaluated as 85 points, the rhythm score as 70 points, the dynamics score as 90 points, and the expressiveness score as 80 points.

[0171] Practice suggestions:

[0172] Data calculation: Suggests specific phrase practice to help users improve. These suggestions are automatically generated based on the analysis results.

[0173] Output: Evaluation results and specific practice suggestions are generated.

[0174] Step 5: Submit and view the evaluation results

[0175] The server transmits the generated evaluation results to the user's terminal.

[0176] Input: Assessment results and practice suggestion data.

[0177] Data processing: Formatting the evaluation results for visual presentation.

[0178] Output: Send data securely using HTTPS.

[0179] The terminal visually displays the received evaluation results on a user interface.

[0180] Specific operation: The evaluation results are displayed visually in bar graphs, pie charts, etc., allowing users to check their own performance evaluation.

[0181] (Application example 1)

[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0183] When users play an instrument, it is difficult to obtain objective evaluations of their performance. Furthermore, because it is difficult to evaluate oneself, it is difficult to find effective practice methods. Existing voice analysis systems lack detailed feedback and suggestions, making them inadequate to support users' improvement.

[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0185] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for suggesting phrases necessary for improvement based on the analysis results, means for transmitting the evaluation results to the user's terminal and visually displaying them, and means for displaying the evaluation results on a user interface so that the user can check detailed feedback. This allows the user to receive an objective evaluation of their performance and practice effectively while identifying specific areas for improvement.

[0186] "High-quality recording" refers to technology or means that can record audio signals clearly with low noise.

[0187] "Transmitting audio data" refers to transferring recorded digital data to a server via a network.

[0188] "Pitch analysis" is a technique that detects the frequency of each note in a performance and evaluates the exact pitch.

[0189] "Rhythm analysis" is a technique for analyzing the temporal structure of a performance and evaluating the accuracy of timing.

[0190] "Dynamic analysis" is a technology that detects changes in the volume of a performance and evaluates dynamics and expressiveness.

[0191] A "generative AI model" refers to a model of artificial intelligence that has been trained to perform a specific task using machine learning algorithms.

[0192] "Noise reduction" refers to techniques and methods for removing unwanted background noise from recorded audio data.

[0193] "User interface" refers to the screen and operation method used by the system and the user to exchange information.

[0194] "Overall scoring" is the process of calculating an overall evaluation score based on the results of the analyzed performance.

[0195] The purpose of this invention is to provide a system that allows users to obtain objective evaluations of their musical instrument performance and specific practice methods. Specifically, this system transmits audio data played by the user to a server, analyzes and evaluates pitch, rhythm, dynamics, and expressiveness, and visually displays the results on the user's terminal.

[0196] The system is configured as follows:

[0197] 1. Receiving and recording audio input:

[0198] When a user plays an instrument, a smartphone or other device records the sound in high quality using the smartphone's microphone, and the recorded data is saved as a digital audio file.

[0199] 2. Sending audio data:

[0200] After recording is complete, the device uploads the recorded audio file to a server via HTTPS, where the transmission is encoded to ensure security.

[0201] 3. Analysis of audio data:

[0202] The server stores the received audio data in storage and begins analysis. The analysis performed by the server first includes noise reduction. Then, pitch analysis uses FFT (Fast Fourier Transform), and rhythm analysis involves detecting the start and end timing of notes. Dynamic analysis detects fluctuations in audio volume, and expressiveness is evaluated using a generative AI model. For example, TensorFlow or PyTorch are used.

[0203] 4. Generating evaluation results:

[0204] The server calculates scores for each element based on the analysis results and generates an overall evaluation. Using noise reduction and various analysis methods, it provides detailed scores for pitch, rhythm, dynamics, and expressiveness. It also generates suggestions for specific phrase practice to improve your proficiency based on the evaluation results.

[0205] 5. Submitting and viewing the evaluation results:

[0206] The server then sends the generated evaluation results to the device. This transmission is also securely secured using HTTPS. The device analyzes the received evaluation results and displays them on the user interface. This allows the user to visually check the feedback and confirm the detailed evaluation content.

[0207] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on a smartphone. The recording data is then uploaded to a server, which analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness. For example, the server may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good. The evaluation results are sent to the smartphone and displayed visually on a user interface. Feedback includes suggestions for "short phrases to improve rhythm."

[0208] The system provides users with more accurate and detailed feedback, supporting more effective instrument practice.

[0209] Example prompt sentence:

[0210] "Analyze this performance data and evaluate pitch, rhythm, dynamics, and expressiveness."

[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0212] Step 1: Receiving and Recording Audio Input

[0213] When a user plays an instrument, the device (smartphone) uses a microphone to record the performance sound in high quality. When recording starts, a recording application is launched and real-time audio data is stored in memory. When recording ends, this audio data is saved as a digital audio file. The input is the user's performance sound, and the output is the saved digital audio file.

[0214] Step 2: Sending audio data

[0215] After the recording is complete, the device uploads the saved audio file to the server via HTTPS via user interaction. The audio data is properly encoded before transmission to ensure security. The input of this step is the recorded digital audio file, and the output is the audio data uploaded to the server.

[0216] Step 3: Noise reduction of audio data

[0217] The server stores the received audio data in storage and first performs noise reduction. Digital signal processing technology is used to remove unwanted background noise from the audio data. The input is the audio data stored on the server, and the output is the audio data with the noise removed.

[0218] Step 4: Analyzing the pitch

[0219] The server analyzes the pitch of the noise-removed audio data using FFT (Fast Fourier Transform). It detects the reference frequency of each note and compares it with the pitch written in the music score. The input is the noise-removed audio data, and the output is the result of the pitch analysis.

[0220] Step 5: Analyze the rhythm

[0221] After the pitch analysis is complete, the server performs rhythm analysis. It detects the start and end timing of notes on the time axis and evaluates their correspondence with the reference rhythm. The input is the FFT analysis result, and the output is the result of rhythm analysis.

[0222] Step 6: Dynamic analysis

[0223] The server then performs dynamic analysis, detecting volume fluctuations for each note in the audio data and analyzing whether the dynamic range is as specified in the score. The input is the rhythm analysis result, and the output is the dynamic analysis result.

[0224] Step 7: Expressive Analysis

[0225] The server uses a generative AI model to evaluate the emotional expressiveness of a performance. The model compares it with existing famous performance data to calculate the degree of match in expressiveness. The input is the dynamic analysis result, and the output is the result of expressiveness analysis.

[0226] Step 8: Generate evaluation results

[0227] The server generates an overall score based on the analysis results of pitch, rhythm, dynamics, and expressiveness. It then integrates the analysis results to create numerical scores and graphs for each element. It also lists suggestions for specific phrase practice to help the user improve. The input is all the analysis results, and the output is an overall evaluation result and practice suggestions.

[0228] Step 9: Submit and view the evaluation results

[0229] The server sends the assessment results to the device using HTTPS. The device analyzes the received assessment results and provides them to the user by visually displaying them in a user interface. The display includes graphs, numerical values, and text feedback. The input is the assessment results and practice suggestions, and the output is the feedback displayed in the user interface.

[0230] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it in step 1. The performance is uploaded to the server in step 2, and various analyses are performed in steps 3 to 7. An overall evaluation is created in step 8, and the results are displayed on the user's device in step 9. An example of a prompt sentence is "Analyze this performance data and evaluate its pitch, rhythm, dynamics, and expressiveness."

[0231] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0232] MODE FOR CARRYING OUT THE INVENTION

[0233] The present invention is a system for supporting the improvement of musical instrument performance. It not only records, analyzes, and evaluates a user's performance data, but also reflects the user's emotional state in the evaluation. This system allows users to objectively evaluate their own performance and find appropriate practice methods. Specific embodiments for implementing the present invention are described below.

[0234] Overview of program processing

[0235] 1. Receiving and recording audio input

[0236] When the user starts playing, the device launches a recording application and records the performance sound in high quality through the microphone. Recording is done in real time, and the audio data is saved in the device's memory. The audio waveform and recording time are displayed during recording.

[0237] 2. Sending audio data

[0238] When the performance is finished and the user stops recording, the device saves the recording in the specified audio format. When the user selects the "Send" button, the recording is uploaded to the server.

[0239] 3. Analysis of audio data

[0240] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially analyzes pitch, rhythm, dynamics, and expressiveness.

[0241] Pitch analysis: Uses FFT to identify the base frequency of a note and compares it to the pitch written in the sheet music.

[0242] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[0243] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[0244] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[0245] 4. Emotion Recognition by Emotion Engine

[0246] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them using the emotion engine.

[0247] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[0248] 5. Generating evaluation results

[0249] The server integrates the results of the voice data analysis and the emotion engine to generate a comprehensive evaluation. It calculates a score based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and creates visualization data.

[0250] The server also generates advice and messages to help maintain motivation that are tailored to the user's emotional state and adds them to the suggestion list.

[0251] 6. Submitting and displaying evaluation results

[0252] The server then sends the generated evaluation results to the device. The transmission is secure and done using HTTPS.

[0253] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also provided.

[0254] Specific examples

[0255] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[0256] The user starts playing the piano and the device records the audio.

[0257] After the performance is over, the device saves the recording data and uploads it to the server.

[0258] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[0259] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[0260] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0261] The server sends the evaluation results and suggestions to the device, which then displays them on the user interface, allowing the user to receive visual feedback and advice based on their emotional state.

[0262] In this way, the system of the present invention provides the user with objective performance evaluation as well as feedback according to their emotional state, supporting enjoyable and effective improvement in musical instrument playing.

[0263] The processing flow will be explained below.

[0264] MODE FOR CARRYING OUT THE INVENTION

[0265] Specific flow of program processing

[0266] Step 1:

[0267] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording occurs in real time, and the audio data is stored in the device's memory. During recording, the audio waveform and recording time are displayed on the screen, allowing the user to check the progress.

[0268] Step 2:

[0269] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After the recording data has been saved, the user selects the "Send" button to send the recording data to the server.

[0270] Step 3:

[0271] The device uploads the recorded audio file to the server via an HTTP request in response to user input. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[0272] Step 4:

[0273] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of subsequent analysis.

[0274] Step 5:

[0275] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[0276] Step 6:

[0277] The server then performs rhythm analysis, detecting the start and end times of notes on the time axis and comparing them with a reference rhythm template. The analysis results are recorded as data including timestamps.

[0278] Step 7:

[0279] The server also performs dynamic analysis, analyzing fluctuations in audio volume and assessing whether the dynamics of each note are in line with the musical notation. The results are recorded as a numerical value representing the sound pressure level.

[0280] Step 8:

[0281] Next, the server performs expressiveness analysis. Using a generative AI model, it evaluates the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data to calculate the degree of agreement in expressiveness, and records the analysis results as an emotional score.

[0282] Step 9:

[0283] The device captures the user's facial expressions and tone of voice during and after performance, and activates the emotion engine, which analyzes the user's facial expressions and tone of voice to evaluate the user's emotional state in real time.

[0284] Step 10:

[0285] The server combines the results of the voice data analysis with the emotion analysis results from the emotion engine to generate an overall evaluation. It also combines scores based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state to create visualization data. It also generates motivational advice and messages according to the user's emotional state.

[0286] Step 11:

[0287] The server then sends the generated evaluation results and motivation advice to the device. The transmission is securely performed using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as advice based on the user's emotional state.

[0288] Step 12:

[0289] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check detailed evaluations and motivational advice. The device also has "replay" and "reevaluation" buttons, allowing users to play again and receive a new evaluation.

[0290] Specific examples

[0291] For example, if a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and the following process is performed.

[0292] The user starts playing and the terminal starts recording.

[0293] After the performance is finished, the recording data is saved on the device and sent to the server.

[0294] The server analyzes the received data and evaluates it for pitch, rhythm, strength, and expressiveness. At the same time, the device analyzes the user's facial expressions and voice and evaluates their emotional state using an emotion engine.

[0295] The server integrates the overall analysis results and generates an overall evaluation such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotional evaluation."

[0296] The server suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0297] The device displays the evaluation results and suggestions on a user interface, allowing the user to see visual feedback and motivational advice.

[0298] As described above, the system of the present invention provides the user with objective performance evaluation and feedback based on emotional state, enabling enjoyable and effective improvement in musical instrument playing.

[0299] Example 2

[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] Conventional systems for supporting the improvement of musical instrument playing skills are specialized in evaluating the user's performance technique, but are unable to provide comprehensive feedback that takes into account the user's emotional state. This makes it difficult to support maintaining motivation or improving the emotional expressiveness of performance. Furthermore, conventional systems only evaluate performance technique criteria such as pitch, rhythm, and dynamics, and are limited in their assessment of expressiveness.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness, means for capturing the user's facial expression and tone of voice and analyzing their emotional state, and means for integrating the emotional state into the analysis results and generating an overall evaluation. This makes it possible to provide comprehensive feedback that integrates the user's performance technique and emotional state, and to support maintaining motivation and improving the emotional expressiveness of performance.

[0303] A "user" is an individual who plays an instrument and inputs performance data into the system.

[0304] "Terminal" refers to an electronic device that a user uses to record musical instrument performances and display analysis results.

[0305] "Server" refers to a computer system that receives and analyzes voice data sent from a terminal, generates evaluation results, and sends them to the terminal.

[0306] "Performing" refers to the act of a user making sounds using an instrument.

[0307] "Audio data" refers to data that is a digital recording of music played by a user.

[0308] "Noise reduction" refers to the process of removing unnecessary noise from audio data.

[0309] "FFT" stands for Fast Fourier Transform, a transformation technique frequently used in signal processing.

[0310] A "generative AI model" is a model that uses artificial intelligence to generate and analyze data.

[0311] A "prompt sentence" is an input sentence given to a generative AI model.

[0312] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[0313] "Evaluation result" refers to comprehensive feedback generated based on analysis of speech data and emotional state.

[0314] "Visualized data" refers to data that displays the evaluation results in graphs, charts, etc. so that users can intuitively understand them.

[0315] "Advice" refers to specific guidelines suggested to help users improve their playing skills and motivation.

[0316] This invention is a system to support the improvement of musical instrument performance by recording, analyzing, and evaluating the user's performance data, and also reflecting the user's emotional state in the evaluation. This system allows the user to objectively evaluate their own performance and find appropriate practice methods.

[0317] Recording and sending audio

[0318] When a user starts playing an instrument, the device launches a recording application and records the performance through the microphone in high quality. The recorded audio data is stored in the device's memory and sent to a server as needed. During recording, the audio waveform and recording time are displayed, allowing the user to understand their performance in real time.

[0319] Analysis of audio data

[0320] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it performs the following analysis steps in sequence.

[0321] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of the note and compare it to the pitch written in the sheet music.

[0322] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[0323] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[0324] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[0325] emotion recognition

[0326] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The camera and microphone capture the user's facial expressions and tone of voice and analyze them using the emotion engine. The server integrates the analysis results of the emotion engine with the analysis results of the voice data and reflects the user's emotional state in an overall evaluation.

[0327] Generating and displaying evaluation results

[0328] The server combines the results of the voice data analysis with the emotion engine's results to generate an overall evaluation. A score is calculated based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and visualization data is created. Furthermore, advice and messages to maintain motivation according to the user's emotional state are generated and added to a suggestion list. The server sends the generated evaluation results to the device, which then visually displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also displayed.

[0329] Examples of prompt statements

[0330] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[0331] The user starts playing the piano and the device records the audio.

[0332] After the performance is over, the device saves the recording data and uploads it to the server.

[0333] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[0334] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[0335] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0336] The server sends the evaluation results and suggestions to the terminal, which displays them on the user interface.

[0337] An example of a prompt for input to a generative AI model is as follows:

[0338] "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and score them on pitch, rhythm, dynamics, and expressiveness. Also, analyze the user's facial expressions and tone of voice to assess their emotional state and create comprehensive feedback."

[0339] In this way, the system of the present invention can provide the user with objective performance evaluation as well as feedback according to their emotional state, thereby supporting enjoyable and effective improvement in musical instrument playing.

[0340] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0341] Step 1: Receiving and Recording Audio Input

[0342] When the user starts playing an instrument, the device launches a recording application and records the performance in high quality through the microphone.

[0343] Input: User's playing sound, recording application

[0344] Specific operation: The recording application captures the audio input from the microphone and saves it as audio data in the device's memory. During recording, the audio waveform and recording time are displayed in real time.

[0345] Output: Recorded audio data

[0346] Step 2: Save and send audio data

[0347] When the user finishes playing and presses the "Send" button, the device stops recording and saves the recorded data.

[0348] Input: User click on "Submit" button, recorded voice data

[0349] Specific operation: Calls the StopRecording() method after recording is complete. The saved audio data is saved as a file in the specified audio format (e.g. WAV, MP3).

[0350] Output: Audio data saved in the specified format

[0351] The terminal uploads the audio data to the server.

[0352] Input: Saved audio data

[0353] Specific behavior: Generates an HTTP POST request to send the saved audio data to the server.

[0354] Output: Audio data uploaded to the server

[0355] Step 3: Analyzing the audio data

[0356] The server stores the received audio data in temporary storage and begins analysis.

[0357] Input: Uploaded audio data

[0358] Specific operation: Save the audio data as a temporary file and perform the following analysis steps.

[0359] Noise reduction: Removes unwanted noise from audio data.

[0360] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of a note and compare it with a reference value.

[0361] Rhythm analysis: Detects the start and end times of notes and compares them with a reference rhythm.

[0362] Dynamics analysis: Analyzes fluctuations in voice volume and evaluates the appropriateness of dynamics.

[0363] Expressiveness analysis: Using a generative AI model, the emotional expression of a performance is evaluated based on a prompt.

[0364] Prompt: "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and rate it on pitch, rhythm, dynamics, and expressiveness. Also, rate the user's emotional state by analyzing their facial expressions and tone of voice, and create overall feedback."

[0365] Output: Pitch analysis results, rhythm analysis results, dynamic analysis results, expressiveness analysis results

[0366] Step 4: Emotion Recognition

[0367] The device captures the user's facial expressions and tone of voice to activate the emotion engine.

[0368] Input: User's facial expression (camera image), tone of voice (audio data)

[0369] Specific operation: Uses a camera and microphone to capture facial expression data and voice tone data in real time, and analyzes the data using an emotion engine.

[0370] Output: Parsed emotion data

[0371] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[0372] Input: Voice analysis results, emotion analysis results

[0373] Specific operation: The results of speech analysis and sentiment analysis are integrated, and an evaluation algorithm is applied to generate an overall evaluation.

[0374] Output: Integrated overall rating

[0375] Step 5: Generate evaluation results

[0376] The server integrates the results of the voice data analysis and the emotion engine results to generate an overall evaluation.

[0377] Input: Integrated overall rating

[0378] Specific actions: Based on the integrated overall evaluation, the system calculates a score based on pitch, rhythm, dynamics, expressiveness, and emotional state, creates visualization data, and generates advice and messages to maintain motivation and adds them to a suggestion list.

[0379] Output: Overall evaluation data, visualization data (graphs, charts, etc.), motivation advice

[0380] Step 6: Submit and view the results of the evaluation

[0381] The server transmits the generated evaluation results to the terminal.

[0382] Input: Overall evaluation data, Visualization data, Advice

[0383] Specific operation: Evaluation results, visualization data, and advice are securely sent to the device using HTTPS.

[0384] Output: Data sent to the terminal

[0385] The terminal receives the evaluation results and visually displays them on a user interface.

[0386] Input: Overall evaluation data, Visualization data, Advice

[0387] Specific operation: Analyzes the evaluation results and visualizes the feedback as graphs, numbers, and text. Users can check the detailed evaluation results and appropriate motivation advice.

[0388] Output: Visual feedback to the user

[0389] Through the above steps, the system of the present invention can provide the user with detailed performance evaluations and appropriate feedback, and support the user in improving their musical instrument playing in an enjoyable and effective way.

[0390] (Application example 2)

[0391] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0392] Current systems for supporting the improvement of musical instrument performance only provide objective feedback on the user's performance, and are unable to provide comprehensive evaluations or suggestions that take into account the user's emotional state. As a result, they are unable to suggest phrases or products that the user truly needs, resulting in a lack of motivation and personalized support.

[0393] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0394] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for analyzing the user's emotions from their facial expressions and tone of voice, means for suggesting phrases and products necessary for improvement based on the analysis results and emotion analysis results, and means for transmitting the evaluation results to the user's terminal and visually displaying them, thereby enabling comprehensive and personalized feedback on the user's performance.

[0395] The "recording means" is a device or software for recording the sound played by the user in high quality.

[0396] The "transmission means" is a communication device or software for transmitting the recorded voice data to the server.

[0397] "Analysis means" refers to a device or software for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness.

[0398] "Emotion analysis means" refers to a device or software that analyzes a user's facial expression and tone of voice to evaluate their emotional state.

[0399] The "suggestion means" is a device or software that suggests phrases and products necessary for improvement based on the analysis results and emotion analysis results.

[0400] The "display means" is a device or software for transmitting the evaluation results to the user's terminal and visually displaying them.

[0401] The "frequency analysis means" is a device or software for analyzing the frequency of audio data in order to detect pitch.

[0402] "Rhythm analysis means" refers to a device or software for analyzing the start and end timing of notes in order to detect rhythm.

[0403] A "volume analysis means" is a device or software for analyzing variations in voice volume to detect strength and weakness.

[0404] A "generative AI model" is a machine learning model used to evaluate expressiveness.

[0405] A "prompt sentence" is an input sentence for generating product suggestions from the user's emotional state.

[0406] This invention is a virtual store shopping assistant system that makes personalized product recommendations taking into account the user's emotional state. This system records the sound of a musical instrument played by the user in high quality, analyzes the data, and analyzes the user's emotions from their facial expressions and tone of voice to provide comprehensive feedback.

[0407] Hardware and Software

[0408] This system is implemented using the following hardware and software:

[0409] Hardware:

[0410] Smartphone or smart glasses

[0411] Camera (for facial expression analysis)

[0412] Microphone (for recording audio)

[0413] software:

[0414] Python

[0415] OpenCV (for facial expression analysis)

[0416] TensorFlow (for facial expression analysis model)

[0417] Sounddevice (for audio recording)

[0418] Soundfile (for manipulating audio data)

[0419] Requests (for sending data)

[0420] Processing flow

[0421] Audio recording and analysis

[0422] The device records the user's voice when they ask a question or make a comment about a product. The recorded voice data is sent to a server, where it is analyzed, including frequency analysis. Specific analysis items include pitch, rhythm, dynamics, and expressiveness.

[0423] facial expression analysis

[0424] The device's built-in camera captures the user's facial expression data, which is then processed by a generative AI model using TensorFlow for emotion recognition. The resulting data is used to evaluate the user's level of interest or anxiety.

[0425] Suggestions and Displays

[0426] The server combines the results of voice analysis and facial expression analysis to generate prompts and make appropriate product suggestions. Based on the generated prompts, the system makes personalized suggestions to the user. The suggestions are displayed on the device.

[0427] Specific examples

[0428] If a user is browsing home decor products in a virtual shopping application and asks, "How big is this sofa?", the following occurs:

[0429] 1. Audio recording: The device's microphone records the audio.

[0430] 2. Audio analysis: Analyze pitch, rhythm, dynamics, and expressiveness from recorded data.

[0431] 3. Facial Expression Analysis: The camera captures the user's facial expressions and evaluates the user's emotional state.

[0432] 4. Proposal generation: Based on the results of voice analysis and facial expression analysis, a generative AI model is used to generate optimal product proposals for the user.

[0433] Prompt Sentence Examples

[0434] The prompt for "How big is this sofa?" is as follows:

[0435] Analyze the customer's interest and anxiety based on their questions while they are browsing products, and provide optimal product suggestions. Also, analyze their facial expressions to see their level of tension or interest, and provide explanations with kind words.

[0436] This allows users to receive personalized feedback based on their emotional state, resulting in a better shopping experience.

[0437] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0438] Step 1:

[0439] The device records the user's voice when they ask a question or make a comment about a product. Specifically, it captures high-quality audio data through a microphone and saves it in an audio format such as WAV. The recorded data is used as is in the next step.

[0440] Input: User's voice

[0441] Output: High-quality audio data (WAV format)

[0442] Step 2:

[0443] The device sends the recorded audio data to the server. The data is transferred using a secure communication protocol such as HTTPS. The server then stores the received audio data in its storage.

[0444] Input: High-quality audio data

[0445] Output: Upload audio data to the server

[0446] Step 3:

[0447] The server analyzes the transmitted audio data. The analysis items are pitch, rhythm, dynamics, and expressiveness. Specific processing includes frequency analysis using FFT, time-domain note analysis, detection of voice volume fluctuations, and evaluation of emotional expression using a generative AI model.

[0448] Input: High-quality audio data

[0449] Output: Evaluation results for pitch, rhythm, dynamics, and expressiveness

[0450] Step 4:

[0451] The device uses a camera to capture the user's facial expressions. The captured video data is input into a generative AI model using TensorFlow to recognize the user's emotional state from their facial expressions. Specifically, it extracts facial features and classifies emotions based on them.

[0452] Input: User facial expression video data

[0453] Output: User's emotional state (interest, anxiety, etc.)

[0454] Step 5:

[0455] The server combines the results of voice analysis and facial expression analysis to generate prompts, which are then used by a generative AI model to provide optimal product recommendations to the user.

[0456] Input: Pitch, rhythm, dynamics, expressiveness evaluation results, and the user's emotional state

[0457] Output: prompt statement

[0458] Step 6:

[0459] The server creates specific product suggestions based on the generated prompt sentences, and provides the suggestions to the user as personalized feedback.

[0460] Input: prompt statement

[0461] Output: Personalized product recommendations

[0462] Step 7:

[0463] The terminal receives the suggestions sent from the server and visually displays them on the user interface. Specifically, the suggestions are displayed as text, graphs, or images, providing visual feedback to the user.

[0464] Input: Proposal from the server

[0465] Output: Visual feedback on the user interface

[0466] These steps allow users to receive personalized feedback based on their emotional state and enjoy a better shopping experience.

[0467] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0468] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0469] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0470] [Second embodiment]

[0471] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0472] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0473] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0474] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0475] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0477] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0478] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0479] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0480] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0481] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0482] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0483] MODE FOR CARRYING OUT THE INVENTION

[0484] The present invention is a system that records a user's performance of a musical instrument, transmits the audio data to a server, and analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness of the user. The system of the present invention allows users to objectively evaluate their own performance and find effective practice methods. The following describes in detail an embodiment of the present invention.

[0485] Overview of program processing

[0486] 1. Receiving and recording audio input

[0487] When the user starts playing, the device launches a recording application. The device records the performance in high quality and saves the recorded data as a digital audio file. During recording, real-time audio data is stored in memory.

[0488] 2. Sending audio data

[0489] After the recording is complete, the device uploads the recorded audio file to a specific server URL via an HTTP request, and the data is securely encoded and transmitted.

[0490] 3. Analysis of audio data

[0491] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially performs pitch analysis, rhythm analysis, dynamic analysis, and expressiveness analysis.

[0492] Pitch analysis:

[0493] The server uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compare it with the pitch written in the sheet music.

[0494] Rhythm analysis:

[0495] The server detects the start and end timing of notes on the time axis and evaluates the degree of agreement with the reference rhythm.

[0496] Dynamic analysis:

[0497] The server detects fluctuations in audio volume and analyzes whether the dynamics of each note are as specified in the musical score.

[0498] Expressiveness analysis:

[0499] The server uses a generative AI model to evaluate the emotional expression of a performance, which is then compared with existing data on famous performances to calculate the degree of match in expressiveness.

[0500] 4. Generating evaluation results

[0501] Based on the analysis results, the server calculates scores for pitch, rhythm, dynamics, and expressiveness, and generates an overall evaluation, which includes numerical scores and graphs for each element.

[0502] The server also lists suggestions for specific phrase practice to help the user improve and adds them to the evaluation results.

[0503] 5. Submitting and displaying evaluation results

[0504] The server then sends the generated evaluation results to the terminal, securely using HTTPS.

[0505] The device analyzes the received evaluation results and displays them visually on the user interface, providing feedback in the form of graphs, numbers, and text, allowing the user to check the detailed evaluation content.

[0506] Specific examples

[0507] If a user were to play "Twinkle Twinkle Little Star" on the piano, the performance would be recorded and analyzed and evaluated using the following procedure.

[0508] The user starts playing the piano and the device records the audio.

[0509] After the performance is finished, the device will stop recording and upload the audio file to the server.

[0510] The server analyzes the received audio data and evaluates it for pitch, rhythm, dynamics, and expressiveness. For example, it may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good.

[0511] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[0512] The server sends the evaluation results and suggestions to the device, which displays them on the user interface. The user can then visually check the feedback, replay the performance, and self-evaluate and try to improve.

[0513] In this way, the system of the present invention provides the user with objective evaluations and specific practice suggestions, and supports improvement in musical instrument playing.

[0514] The processing flow will be explained below.

[0515] Step 1:

[0516] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording is performed in real time, and the audio data is stored in the device's memory. During recording, the device displays the audio waveform and provides a visual indication of the progress.

[0517] Step 2:

[0518] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After saving is complete, the user selects the "Send" button to send the recorded data to the server.

[0519] Step 3:

[0520] In response to user operations, the device sends an HTTP request and uploads the recorded audio file to the server. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[0521] Step 4:

[0522] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of the analysis.

[0523] Step 5:

[0524] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[0525] Step 6:

[0526] Next, the server performs rhythm analysis, detecting the start and end timing of notes on the time axis and comparing them with a reference rhythm template. This evaluates the rhythmic accuracy of the performance. The analysis results are recorded as data including timestamps.

[0527] Step 7:

[0528] The server continues dynamic analysis, analyzing the fluctuations in audio volume to detect the sound pressure level of each note. It compares this with the dynamics instructions in the score to evaluate the appropriateness of the dynamics. The analysis results are recorded as a numerical sound pressure level.

[0529] Step 8:

[0530] Finally, the server analyzes the performance's expressiveness. It uses a generative AI model to evaluate the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data and calculates the degree of agreement in expressiveness. The analysis results are recorded as an emotional score.

[0531] Step 9:

[0532] The server combines all the analysis results to generate an overall evaluation. It also combines scores for pitch, rhythm, dynamics, and expressiveness to create visualizations (e.g., radar charts). It also lists specific phrase practice suggestions to help the user improve and adds them to the evaluation results.

[0533] Step 10:

[0534] The server then sends the generated evaluation results to the device. The transmission is secure using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as specific suggestions for improvement.

[0535] Step 11:

[0536] The device analyzes the received evaluation results and displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check the detailed evaluation. The device also provides "replay" and "re-evaluation" buttons, allowing users to replay and receive a new evaluation.

[0537] Example 1

[0538] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0539] Conventional methods for evaluating musical instrument performance require the presence of an instructor with specialized knowledge, making it difficult to obtain objective and quantitative evaluations. Furthermore, when users self-evaluate, the evaluation tends to be subjective, making it difficult to find effective practice methods. Furthermore, it is difficult to evaluate the emotional expression and subtle dynamics of a performance, limiting the overall improvement of performance skills.

[0540] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0541] In this invention, the server includes a means for analyzing audio data and evaluating scale, rhythm, volume changes, and emotional expression; a means for suggesting practice content necessary for improvement based on the analysis results; and a means for transmitting the evaluation results to the user's device and visually displaying them. This enables objective and quantitative evaluation of the user's performance and provides effective practice methods. Furthermore, using a generative AI model to evaluate expressiveness enables comprehensive evaluation, including emotional expression, helping to improve the user's performance skills.

[0542] "Audio data" is information that digitally records the sound of an instrument played by a user.

[0543] A "communication device" is a device for transmitting and receiving acoustic data via a network.

[0544] "Scale" refers to the reference frequency of each note written on a musical score.

[0545] "Rhythm" is a temporal pattern based on the timing of the start and end of musical notes.

[0546] "Volume change" refers to fluctuations in the volume of the audio during performance.

[0547] "Emotional expression" is an indicator that evaluates the emotions and nuances in a performance.

[0548] A "generative AI model" is an artificial intelligence model that has been trained in advance using machine learning algorithms.

[0549] "Visually displaying" means providing analysis results and feedback to the user in a form that can be seen as graphs or text.

[0550] "Practice content" refers to specific playing phrases and tasks that the user needs to improve.

[0551] MODE FOR CARRYING OUT THE INVENTION

[0552] The present invention is a system that supports users in improving their musical performance skills by recording, analyzing, and evaluating the sounds of musical instruments played by the user with high accuracy. The following describes in detail a method for specifically implementing the system of the present invention.

[0553] System Configuration

[0554] Hardware:

[0555] 1. Terminal: A device operated by a user, such as a smartphone, tablet, or computer.

[0556] 2. Server: A remote server for performing analytical processing, using a computer system with high-performance processing capabilities.

[0557] software:

[0558] 1. Recording application: An application that starts when the user starts playing and records high-quality audio. For example, audio recording applications such as "Audacity" or "Pro Tools" can be used.

[0559] 2. Audio analysis program: A program that runs on the server and analyzes audio data. This analysis uses audio processing tools such as "sox" or "FFmpeg."

[0560] 3. Generative AI model: An AI model equipped with a machine learning algorithm for analyzing emotional expressions. For example, a model pre-trained using "TensorFlow" or "PyTorch" is used.

[0561] Processing Overview

[0562] The user plays an instrument and records the sound on the device. The recorded sound data is then uploaded to the server by the user's operation. The server analyzes the received sound data and evaluates the scale, rhythm, volume changes, and emotional expression. Based on the evaluation results, the system suggests specific practice content and sends the evaluation results to the device, where they are visually displayed.

[0563] Specific actions

[0564] 1. Receiving and recording audio input:

[0565] When the user starts playing, an audio recording application is launched on the device. The sound is recorded in high quality, and the audio data is temporarily stored in memory as it is being recorded. When recording is finished, the audio data is saved as a digital audio file. For example, it may be saved with a file name such as "recording.wav."

[0566] 2. Sending audio data:

[0567] After recording is complete, the device uploads the recorded audio file to a specific server URL via a user-initiated HTTP POST request, with the data encoded and securely transmitted. For example, a request can be sent to the URL "https: / / example.com / upload."

[0568] 3. Analysis of audio data:

[0569] The server stores the received audio data in storage and first performs noise reduction, using audio processing tools such as "sox" for noise filtering.

[0570] Next, the scale is analyzed using FFT (Fast Fourier Transform) to detect the reference frequency of each note.

[0571] The system analyzes rhythm by detecting the start and end timing of notes, and analyzes dynamics by detecting fluctuations in voice volume.

[0572] Emotional expression is evaluated using a pre-trained generative AI model, which compares the performance with existing master performance data to calculate the degree of agreement between the performance's emotional expression.

[0573] 4. Generating evaluation results:

[0574] The server then calculates scores for scale, rhythm, volume, and emotional expression based on the analysis results, generating an overall evaluation that includes a numerical score and graphs, and suggests specific phrases for the user to practice.

[0575] 5. Submitting and viewing the evaluation results:

[0576] The server sends the evaluation results to the device, which then visually displays them on the user interface, for example, by displaying the scale and rhythm evaluation scores in bar graphs or pie charts, allowing the user to check the evaluation of their own performance.

[0577] Example operation

[0578] When a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following steps:

[0579] The user starts playing the piano and the device records the audio.

[0580] After the performance is finished, the device stops recording and uploads the audio file "kirakira_sei.wav" to the server.

[0581] The server analyzes the received audio data and scores it with 85 points for pitch, 70 points for rhythm, 90 points for dynamics, and 80 points for emotional expression.

[0582] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[0583] The server sends the evaluation results and suggestions to the device, which displays them on the user interface, allowing the user to visually confirm the feedback and try playing again.

[0584] This system allows users to easily obtain evaluations of their instrument performance and areas for improvement, effectively improving their own playing skills.

[0585] Prompt Sentence Examples

[0586] "Please analyze the following audio data and rate it for pitch, rhythm, volume changes, and emotional expression:"

[0587] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0588] Step 1: Receiving and Recording Audio Input

[0589] The device launches a recording application to receive audio input, and when the user starts playing an instrument, it records the audio in real time with high quality.

[0590] Input: Audio of the user playing an instrument.

[0591] Data processing: The device temporarily stores the audio data being recorded in memory.

[0592] Output: Once recording is complete, the audio data is saved as a digital audio file (e.g., recording.wav).

[0593] Step 2: Sending audio data

[0594] After the recording is completed, the device uploads the recorded audio file to a specific server URL according to the user's operation.

[0595] Input: Recorded digital audio files.

[0596] Data processing: Encode the audio data and send it to the server using an HTTP POST request.

[0597] Output: The audio data is securely uploaded to the server. For example, send a request to the URL "https: / / example.com / upload".

[0598] Step 3: Analyzing the audio data

[0599] The server first stores the received audio data in storage.

[0600] Input: Uploaded digital audio files.

[0601] Acoustic analysis:

[0602] Data processing: Perform noise reduction and use audio processing tools such as "sox" to remove unwanted noise.

[0603] Data calculation: Analyzes the musical scale using FFT (Fast Fourier Transform) and detects the reference frequency of each note. Detects the start and end timing of notes on the time axis and analyzes rhythm. Detects fluctuations in voice volume and analyzes dynamics.

[0604] Emotional expression rating:

[0605] Data computation: Evaluate the emotional expression of a performance using a generative AI model (e.g., a pre-trained model using TensorFlow or PyTorch). Compare with existing renowned performance data to calculate the degree of consistency of expressiveness.

[0606] Output: The analysis results include scores for pitch, rhythm, volume changes, and emotional expression.

[0607] Step 4: Generate evaluation results

[0608] Based on the analysis results, the server calculates scores for scale, rhythm, volume changes and emotional expression to generate an overall rating.

[0609] Input: Audio analysis results data (scale, rhythm, volume changes, emotional expressions).

[0610] Data calculation: A score is calculated based on each analysis result, and an overall evaluation is made. For example, the scale score is evaluated as 85 points, the rhythm score as 70 points, the dynamics score as 90 points, and the expressiveness score as 80 points.

[0611] Practice suggestions:

[0612] Data calculation: Suggests specific phrase practice to help users improve. These suggestions are automatically generated based on the analysis results.

[0613] Output: Evaluation results and specific practice suggestions are generated.

[0614] Step 5: Submit and view the evaluation results

[0615] The server transmits the generated evaluation results to the user's terminal.

[0616] Input: Assessment results and practice suggestion data.

[0617] Data processing: Formatting the evaluation results for visual presentation.

[0618] Output: Send data securely using HTTPS.

[0619] The terminal visually displays the received evaluation results on a user interface.

[0620] Specific operation: The evaluation results are displayed visually in bar graphs, pie charts, etc., allowing users to check their own performance evaluation.

[0621] (Application example 1)

[0622] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0623] When users play an instrument, it is difficult to obtain objective evaluations of their performance. Furthermore, because it is difficult to evaluate oneself, it is difficult to find effective practice methods. Existing voice analysis systems lack detailed feedback and suggestions, making them inadequate to support users' improvement.

[0624] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0625] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for suggesting phrases necessary for improvement based on the analysis results, means for transmitting the evaluation results to the user's terminal and visually displaying them, and means for displaying the evaluation results on a user interface so that the user can check detailed feedback. This allows the user to receive an objective evaluation of their performance and practice effectively while identifying specific areas for improvement.

[0626] "High-quality recording" refers to technology or means that can record audio signals clearly with low noise.

[0627] "Transmitting audio data" refers to transferring recorded digital data to a server via a network.

[0628] "Pitch analysis" is a technique that detects the frequency of each note in a performance and evaluates the exact pitch.

[0629] "Rhythm analysis" is a technique for analyzing the temporal structure of a performance and evaluating the accuracy of timing.

[0630] "Dynamic analysis" is a technology that detects changes in the volume of a performance and evaluates dynamics and expressiveness.

[0631] A "generative AI model" refers to a model of artificial intelligence that has been trained to perform a specific task using machine learning algorithms.

[0632] "Noise reduction" refers to techniques and methods for removing unwanted background noise from recorded audio data.

[0633] "User interface" refers to the screen and operation method used by the system and the user to exchange information.

[0634] "Overall scoring" is the process of calculating an overall evaluation score based on the results of the analyzed performance.

[0635] The purpose of this invention is to provide a system that allows users to obtain objective evaluations of their musical instrument performance and specific practice methods. Specifically, this system transmits audio data played by the user to a server, analyzes and evaluates pitch, rhythm, dynamics, and expressiveness, and visually displays the results on the user's terminal.

[0636] The system is configured as follows:

[0637] 1. Receiving and recording audio input:

[0638] When a user plays an instrument, a smartphone or other device records the sound in high quality using the smartphone's microphone, and the recorded data is saved as a digital audio file.

[0639] 2. Sending audio data:

[0640] After recording is complete, the device uploads the recorded audio file to a server via HTTPS, where the transmission is encoded to ensure security.

[0641] 3. Analysis of audio data:

[0642] The server stores the received audio data in storage and begins analysis. The analysis performed by the server first includes noise reduction. Then, pitch analysis uses FFT (Fast Fourier Transform), and rhythm analysis involves detecting the start and end timing of notes. Dynamic analysis detects fluctuations in audio volume, and expressiveness is evaluated using a generative AI model. For example, TensorFlow or PyTorch are used.

[0643] 4. Generating evaluation results:

[0644] The server calculates scores for each element based on the analysis results and generates an overall evaluation. Using noise reduction and various analysis methods, it provides detailed scores for pitch, rhythm, dynamics, and expressiveness. It also generates suggestions for specific phrase practice to improve your proficiency based on the evaluation results.

[0645] 5. Submitting and viewing the evaluation results:

[0646] The server then sends the generated evaluation results to the device. This transmission is also securely secured using HTTPS. The device analyzes the received evaluation results and displays them on the user interface. This allows the user to visually check the feedback and confirm the detailed evaluation content.

[0647] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on a smartphone. The recording data is then uploaded to a server, which analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness. For example, the server may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good. The evaluation results are sent to the smartphone and displayed visually on a user interface. Feedback includes suggestions for "short phrases to improve rhythm."

[0648] The system provides users with more accurate and detailed feedback, supporting more effective instrument practice.

[0649] Example prompt sentence:

[0650] "Analyze this performance data and evaluate pitch, rhythm, dynamics, and expressiveness."

[0651] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0652] Step 1: Receiving and Recording Audio Input

[0653] When a user plays an instrument, the device (smartphone) uses a microphone to record the performance sound in high quality. When recording starts, a recording application is launched and real-time audio data is stored in memory. When recording ends, this audio data is saved as a digital audio file. The input is the user's performance sound, and the output is the saved digital audio file.

[0654] Step 2: Sending audio data

[0655] After the recording is complete, the device uploads the saved audio file to the server via HTTPS via user interaction. The audio data is properly encoded before transmission to ensure security. The input of this step is the recorded digital audio file, and the output is the audio data uploaded to the server.

[0656] Step 3: Noise reduction of audio data

[0657] The server stores the received audio data in storage and first performs noise reduction. Digital signal processing technology is used to remove unwanted background noise from the audio data. The input is the audio data stored on the server, and the output is the audio data with the noise removed.

[0658] Step 4: Analyzing the pitch

[0659] The server analyzes the pitch of the noise-removed audio data using FFT (Fast Fourier Transform). It detects the reference frequency of each note and compares it with the pitch written in the music score. The input is the noise-removed audio data, and the output is the result of the pitch analysis.

[0660] Step 5: Analyze the rhythm

[0661] After the pitch analysis is complete, the server performs rhythm analysis. It detects the start and end timing of notes on the time axis and evaluates their correspondence with the reference rhythm. The input is the FFT analysis result, and the output is the result of rhythm analysis.

[0662] Step 6: Dynamic analysis

[0663] The server then performs dynamic analysis, detecting volume fluctuations for each note in the audio data and analyzing whether the dynamic range is as specified in the score. The input is the rhythm analysis result, and the output is the dynamic analysis result.

[0664] Step 7: Expressive Analysis

[0665] The server uses a generative AI model to evaluate the emotional expressiveness of a performance. The model compares it with existing famous performance data to calculate the degree of match in expressiveness. The input is the dynamic analysis result, and the output is the result of expressiveness analysis.

[0666] Step 8: Generate evaluation results

[0667] The server generates an overall score based on the analysis results of pitch, rhythm, dynamics, and expressiveness. It then integrates the analysis results to create numerical scores and graphs for each element. It also lists suggestions for specific phrase practice to help the user improve. The input is all the analysis results, and the output is an overall evaluation result and practice suggestions.

[0668] Step 9: Submit and view the evaluation results

[0669] The server sends the assessment results to the device using HTTPS. The device analyzes the received assessment results and provides them to the user by visually displaying them in a user interface. The display includes graphs, numerical values, and text feedback. The input is the assessment results and practice suggestions, and the output is the feedback displayed in the user interface.

[0670] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it in step 1. The performance is uploaded to the server in step 2, and various analyses are performed in steps 3 to 7. An overall evaluation is created in step 8, and the results are displayed on the user's device in step 9. An example of a prompt sentence is "Analyze this performance data and evaluate its pitch, rhythm, dynamics, and expressiveness."

[0671] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0672] MODE FOR CARRYING OUT THE INVENTION

[0673] The present invention is a system for supporting the improvement of musical instrument performance. It not only records, analyzes, and evaluates a user's performance data, but also reflects the user's emotional state in the evaluation. This system allows users to objectively evaluate their own performance and find appropriate practice methods. Specific embodiments for implementing the present invention are described below.

[0674] Overview of program processing

[0675] 1. Receiving and recording audio input

[0676] When the user starts playing, the device launches a recording application and records the performance sound in high quality through the microphone. Recording is done in real time, and the audio data is saved in the device's memory. The audio waveform and recording time are displayed during recording.

[0677] 2. Sending audio data

[0678] When the performance is finished and the user stops recording, the device saves the recording in the specified audio format. When the user selects the "Send" button, the recording is uploaded to the server.

[0679] 3. Analysis of audio data

[0680] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially analyzes pitch, rhythm, dynamics, and expressiveness.

[0681] Pitch analysis: Uses FFT to identify the base frequency of a note and compares it to the pitch written in the sheet music.

[0682] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[0683] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[0684] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[0685] 4. Emotion Recognition by Emotion Engine

[0686] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them using the emotion engine.

[0687] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[0688] 5. Generating evaluation results

[0689] The server integrates the results of the voice data analysis and the emotion engine to generate a comprehensive evaluation. It calculates a score based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and creates visualization data.

[0690] The server also generates advice and messages to help maintain motivation that are tailored to the user's emotional state and adds them to the suggestion list.

[0691] 6. Submitting and displaying evaluation results

[0692] The server then sends the generated evaluation results to the device. The transmission is secure and done using HTTPS.

[0693] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also provided.

[0694] Specific examples

[0695] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[0696] The user starts playing the piano and the device records the audio.

[0697] After the performance is over, the device saves the recording data and uploads it to the server.

[0698] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[0699] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[0700] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0701] The server sends the evaluation results and suggestions to the device, which then displays them on the user interface, allowing the user to receive visual feedback and advice based on their emotional state.

[0702] In this way, the system of the present invention provides the user with objective performance evaluation as well as feedback according to their emotional state, supporting enjoyable and effective improvement in musical instrument playing.

[0703] The processing flow will be explained below.

[0704] MODE FOR CARRYING OUT THE INVENTION

[0705] Specific flow of program processing

[0706] Step 1:

[0707] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording occurs in real time, and the audio data is stored in the device's memory. During recording, the audio waveform and recording time are displayed on the screen, allowing the user to check the progress.

[0708] Step 2:

[0709] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After the recording data has been saved, the user selects the "Send" button to send the recording data to the server.

[0710] Step 3:

[0711] The device uploads the recorded audio file to the server via an HTTP request in response to user input. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[0712] Step 4:

[0713] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of subsequent analysis.

[0714] Step 5:

[0715] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[0716] Step 6:

[0717] The server then performs rhythm analysis, detecting the start and end times of notes on the time axis and comparing them with a reference rhythm template. The analysis results are recorded as data including timestamps.

[0718] Step 7:

[0719] The server also performs dynamic analysis, analyzing fluctuations in audio volume and assessing whether the dynamics of each note are in line with the musical notation. The results are recorded as a numerical value representing the sound pressure level.

[0720] Step 8:

[0721] Next, the server performs expressiveness analysis. Using a generative AI model, it evaluates the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data to calculate the degree of agreement in expressiveness, and records the analysis results as an emotional score.

[0722] Step 9:

[0723] The device captures the user's facial expressions and tone of voice during and after performance, and activates the emotion engine, which analyzes the user's facial expressions and tone of voice to evaluate the user's emotional state in real time.

[0724] Step 10:

[0725] The server combines the results of the voice data analysis with the emotion analysis results from the emotion engine to generate an overall evaluation. It also combines scores based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state to create visualization data. It also generates motivational advice and messages according to the user's emotional state.

[0726] Step 11:

[0727] The server then sends the generated evaluation results and motivation advice to the device. The transmission is securely performed using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as advice based on the user's emotional state.

[0728] Step 12:

[0729] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check detailed evaluations and motivational advice. The device also has "replay" and "reevaluation" buttons, allowing users to play again and receive a new evaluation.

[0730] Specific examples

[0731] For example, if a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and the following process is performed.

[0732] The user starts playing and the terminal starts recording.

[0733] After the performance is finished, the recording data is saved on the device and sent to the server.

[0734] The server analyzes the received data and evaluates it for pitch, rhythm, strength, and expressiveness. At the same time, the device analyzes the user's facial expressions and voice and evaluates their emotional state using an emotion engine.

[0735] The server integrates the overall analysis results and generates an overall evaluation such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotional evaluation."

[0736] The server suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0737] The device displays the evaluation results and suggestions on a user interface, allowing the user to see visual feedback and motivational advice.

[0738] As described above, the system of the present invention provides the user with objective performance evaluation and feedback based on emotional state, enabling enjoyable and effective improvement in musical instrument playing.

[0739] Example 2

[0740] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0741] Conventional systems for supporting the improvement of musical instrument playing skills are specialized in evaluating the user's performance technique, but are unable to provide comprehensive feedback that takes into account the user's emotional state. This makes it difficult to support maintaining motivation or improving the emotional expressiveness of performance. Furthermore, conventional systems only evaluate performance technique criteria such as pitch, rhythm, and dynamics, and are limited in their assessment of expressiveness.

[0742] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness, means for capturing the user's facial expression and tone of voice and analyzing their emotional state, and means for integrating the emotional state into the analysis results and generating an overall evaluation. This makes it possible to provide comprehensive feedback that integrates the user's performance technique and emotional state, and to support maintaining motivation and improving the emotional expressiveness of performance.

[0743] A "user" is an individual who plays an instrument and inputs performance data into the system.

[0744] "Terminal" refers to an electronic device that a user uses to record musical instrument performances and display analysis results.

[0745] "Server" refers to a computer system that receives and analyzes voice data sent from a terminal, generates evaluation results, and sends them to the terminal.

[0746] "Performing" refers to the act of a user making sounds using an instrument.

[0747] "Audio data" refers to data that is a digital recording of music played by a user.

[0748] "Noise reduction" refers to the process of removing unnecessary noise from audio data.

[0749] "FFT" stands for Fast Fourier Transform, a transformation technique frequently used in signal processing.

[0750] A "generative AI model" is a model that uses artificial intelligence to generate and analyze data.

[0751] A "prompt sentence" is an input sentence given to a generative AI model.

[0752] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[0753] "Evaluation result" refers to comprehensive feedback generated based on analysis of speech data and emotional state.

[0754] "Visualized data" refers to data that displays the evaluation results in graphs, charts, etc. so that users can intuitively understand them.

[0755] "Advice" refers to specific guidelines suggested to help users improve their playing skills and motivation.

[0756] This invention is a system to support the improvement of musical instrument performance by recording, analyzing, and evaluating the user's performance data, and also reflecting the user's emotional state in the evaluation. This system allows the user to objectively evaluate their own performance and find appropriate practice methods.

[0757] Recording and sending audio

[0758] When a user starts playing an instrument, the device launches a recording application and records the performance through the microphone in high quality. The recorded audio data is stored in the device's memory and sent to a server as needed. During recording, the audio waveform and recording time are displayed, allowing the user to understand their performance in real time.

[0759] Analysis of audio data

[0760] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it performs the following analysis steps in sequence.

[0761] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of the note and compare it to the pitch written in the sheet music.

[0762] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[0763] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[0764] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[0765] emotion recognition

[0766] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The camera and microphone capture the user's facial expressions and tone of voice and analyze them using the emotion engine. The server integrates the analysis results of the emotion engine with the analysis results of the voice data and reflects the user's emotional state in an overall evaluation.

[0767] Generating and displaying evaluation results

[0768] The server combines the results of the voice data analysis with the emotion engine's results to generate an overall evaluation. A score is calculated based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and visualization data is created. Furthermore, advice and messages to maintain motivation according to the user's emotional state are generated and added to a suggestion list. The server sends the generated evaluation results to the device, which then visually displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also displayed.

[0769] Examples of prompt statements

[0770] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[0771] The user starts playing the piano and the device records the audio.

[0772] After the performance is over, the device saves the recording data and uploads it to the server.

[0773] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[0774] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[0775] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[0776] The server sends the evaluation results and suggestions to the terminal, which displays them on the user interface.

[0777] An example of a prompt for input to a generative AI model is as follows:

[0778] "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and score them on pitch, rhythm, dynamics, and expressiveness. Also, analyze the user's facial expressions and tone of voice to assess their emotional state and create comprehensive feedback."

[0779] In this way, the system of the present invention can provide the user with objective performance evaluation as well as feedback according to their emotional state, thereby supporting enjoyable and effective improvement in musical instrument playing.

[0780] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0781] Step 1: Receiving and Recording Audio Input

[0782] When the user starts playing an instrument, the device launches a recording application and records the performance in high quality through the microphone.

[0783] Input: User's playing sound, recording application

[0784] Specific operation: The recording application captures the audio input from the microphone and saves it as audio data in the device's memory. During recording, the audio waveform and recording time are displayed in real time.

[0785] Output: Recorded audio data

[0786] Step 2: Save and send audio data

[0787] When the user finishes playing and presses the "Send" button, the device stops recording and saves the recorded data.

[0788] Input: User click on "Submit" button, recorded voice data

[0789] Specific operation: Calls the StopRecording() method after recording is complete. The saved audio data is saved as a file in the specified audio format (e.g. WAV, MP3).

[0790] Output: Audio data saved in the specified format

[0791] The terminal uploads the audio data to the server.

[0792] Input: Saved audio data

[0793] Specific behavior: Generates an HTTP POST request to send the saved audio data to the server.

[0794] Output: Audio data uploaded to the server

[0795] Step 3: Analyzing the audio data

[0796] The server stores the received audio data in temporary storage and begins analysis.

[0797] Input: Uploaded audio data

[0798] Specific operation: Save the audio data as a temporary file and perform the following analysis steps.

[0799] Noise reduction: Removes unwanted noise from audio data.

[0800] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of a note and compare it with a reference value.

[0801] Rhythm analysis: Detects the start and end times of notes and compares them with a reference rhythm.

[0802] Dynamics analysis: Analyzes fluctuations in voice volume and evaluates the appropriateness of dynamics.

[0803] Expressiveness analysis: Using a generative AI model, the emotional expression of a performance is evaluated based on a prompt.

[0804] Prompt: "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and rate it on pitch, rhythm, dynamics, and expressiveness. Also, rate the user's emotional state by analyzing their facial expressions and tone of voice, and create overall feedback."

[0805] Output: Pitch analysis results, rhythm analysis results, dynamic analysis results, expressiveness analysis results

[0806] Step 4: Emotion Recognition

[0807] The device captures the user's facial expressions and tone of voice to activate the emotion engine.

[0808] Input: User's facial expression (camera image), tone of voice (audio data)

[0809] Specific operation: Uses a camera and microphone to capture facial expression data and voice tone data in real time, and analyzes the data using an emotion engine.

[0810] Output: Parsed emotion data

[0811] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[0812] Input: Voice analysis results, emotion analysis results

[0813] Specific operation: The results of speech analysis and sentiment analysis are integrated, and an evaluation algorithm is applied to generate an overall evaluation.

[0814] Output: Integrated overall rating

[0815] Step 5: Generate evaluation results

[0816] The server integrates the results of the voice data analysis and the emotion engine results to generate an overall evaluation.

[0817] Input: Integrated overall rating

[0818] Specific actions: Based on the integrated overall evaluation, the system calculates a score based on pitch, rhythm, dynamics, expressiveness, and emotional state, creates visualization data, and generates advice and messages to maintain motivation and adds them to a suggestion list.

[0819] Output: Overall evaluation data, visualization data (graphs, charts, etc.), motivation advice

[0820] Step 6: Submit and view the results of the evaluation

[0821] The server transmits the generated evaluation results to the terminal.

[0822] Input: Overall evaluation data, Visualization data, Advice

[0823] Specific operation: Evaluation results, visualization data, and advice are securely sent to the device using HTTPS.

[0824] Output: Data sent to the terminal

[0825] The terminal receives the evaluation results and visually displays them on a user interface.

[0826] Input: Overall evaluation data, Visualization data, Advice

[0827] Specific operation: Analyzes the evaluation results and visualizes the feedback as graphs, numbers, and text. Users can check the detailed evaluation results and appropriate motivation advice.

[0828] Output: Visual feedback to the user

[0829] Through the above steps, the system of the present invention can provide the user with detailed performance evaluations and appropriate feedback, and support the user in improving their musical instrument playing in an enjoyable and effective way.

[0830] (Application example 2)

[0831] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0832] Current systems for supporting the improvement of musical instrument performance only provide objective feedback on the user's performance, and are unable to provide comprehensive evaluations or suggestions that take into account the user's emotional state. As a result, they are unable to suggest phrases or products that the user truly needs, resulting in a lack of motivation and personalized support.

[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0834] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for analyzing the user's emotions from their facial expressions and tone of voice, means for suggesting phrases and products necessary for improvement based on the analysis results and emotion analysis results, and means for transmitting the evaluation results to the user's terminal and visually displaying them, thereby enabling comprehensive and personalized feedback on the user's performance.

[0835] The "recording means" is a device or software for recording the sound played by the user in high quality.

[0836] The "transmission means" is a communication device or software for transmitting the recorded voice data to the server.

[0837] "Analysis means" refers to a device or software for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness.

[0838] "Emotion analysis means" refers to a device or software that analyzes a user's facial expression and tone of voice to evaluate their emotional state.

[0839] The "suggestion means" is a device or software that suggests phrases and products necessary for improvement based on the analysis results and emotion analysis results.

[0840] The "display means" is a device or software for transmitting the evaluation results to the user's terminal and visually displaying them.

[0841] The "frequency analysis means" is a device or software for analyzing the frequency of audio data in order to detect pitch.

[0842] "Rhythm analysis means" refers to a device or software for analyzing the start and end timing of notes in order to detect rhythm.

[0843] A "volume analysis means" is a device or software for analyzing variations in voice volume to detect strength and weakness.

[0844] A "generative AI model" is a machine learning model used to evaluate expressiveness.

[0845] A "prompt sentence" is an input sentence for generating product suggestions from the user's emotional state.

[0846] This invention is a virtual store shopping assistant system that makes personalized product recommendations taking into account the user's emotional state. This system records the sound of a musical instrument played by the user in high quality, analyzes the data, and analyzes the user's emotions from their facial expressions and tone of voice to provide comprehensive feedback.

[0847] Hardware and Software

[0848] This system is implemented using the following hardware and software:

[0849] Hardware:

[0850] Smartphone or smart glasses

[0851] Camera (for facial expression analysis)

[0852] Microphone (for recording audio)

[0853] software:

[0854] Python

[0855] OpenCV (for facial expression analysis)

[0856] TensorFlow (for facial expression analysis model)

[0857] Sounddevice (for audio recording)

[0858] Soundfile (for manipulating audio data)

[0859] Requests (for sending data)

[0860] Processing flow

[0861] Audio recording and analysis

[0862] The device records the user's voice when they ask a question or make a comment about a product. The recorded voice data is sent to a server, where it is analyzed, including frequency analysis. Specific analysis items include pitch, rhythm, dynamics, and expressiveness.

[0863] facial expression analysis

[0864] The device's built-in camera captures the user's facial expression data, which is then processed by a generative AI model using TensorFlow for emotion recognition. The resulting data is used to evaluate the user's level of interest or anxiety.

[0865] Suggestions and Displays

[0866] The server combines the results of voice analysis and facial expression analysis to generate prompts and make appropriate product suggestions. Based on the generated prompts, the system makes personalized suggestions to the user. The suggestions are displayed on the device.

[0867] Specific examples

[0868] If a user is browsing home decor products in a virtual shopping application and asks, "How big is this sofa?", the following occurs:

[0869] 1. Audio recording: The device's microphone records the audio.

[0870] 2. Audio analysis: Analyze pitch, rhythm, dynamics, and expressiveness from recorded data.

[0871] 3. Facial Expression Analysis: The camera captures the user's facial expressions and evaluates the user's emotional state.

[0872] 4. Proposal generation: Based on the results of voice analysis and facial expression analysis, a generative AI model is used to generate optimal product proposals for the user.

[0873] Prompt Sentence Examples

[0874] The prompt for "How big is this sofa?" is as follows:

[0875] Analyze the customer's interest and anxiety based on their questions while they are browsing products, and provide optimal product suggestions. Also, analyze their facial expressions to see their level of tension or interest, and provide explanations with kind words.

[0876] This allows users to receive personalized feedback based on their emotional state, resulting in a better shopping experience.

[0877] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0878] Step 1:

[0879] The device records the user's voice when they ask a question or make a comment about a product. Specifically, it captures high-quality audio data through a microphone and saves it in an audio format such as WAV. The recorded data is used as is in the next step.

[0880] Input: User's voice

[0881] Output: High-quality audio data (WAV format)

[0882] Step 2:

[0883] The device sends the recorded audio data to the server. The data is transferred using a secure communication protocol such as HTTPS. The server then stores the received audio data in its storage.

[0884] Input: High-quality audio data

[0885] Output: Upload audio data to the server

[0886] Step 3:

[0887] The server analyzes the transmitted audio data. The analysis items are pitch, rhythm, dynamics, and expressiveness. Specific processing includes frequency analysis using FFT, time-domain note analysis, detection of voice volume fluctuations, and evaluation of emotional expression using a generative AI model.

[0888] Input: High-quality audio data

[0889] Output: Evaluation results for pitch, rhythm, dynamics, and expressiveness

[0890] Step 4:

[0891] The device uses a camera to capture the user's facial expressions. The captured video data is input into a generative AI model using TensorFlow to recognize the user's emotional state from their facial expressions. Specifically, it extracts facial features and classifies emotions based on them.

[0892] Input: User facial expression video data

[0893] Output: User's emotional state (interest, anxiety, etc.)

[0894] Step 5:

[0895] The server combines the results of voice analysis and facial expression analysis to generate prompts, which are then used by a generative AI model to provide optimal product recommendations to the user.

[0896] Input: Pitch, rhythm, dynamics, expressiveness evaluation results, and the user's emotional state

[0897] Output: prompt statement

[0898] Step 6:

[0899] The server creates specific product suggestions based on the generated prompt sentences, and provides the suggestions to the user as personalized feedback.

[0900] Input: prompt statement

[0901] Output: Personalized product recommendations

[0902] Step 7:

[0903] The terminal receives the suggestions sent from the server and visually displays them on the user interface. Specifically, the suggestions are displayed as text, graphs, or images, providing visual feedback to the user.

[0904] Input: Proposal from the server

[0905] Output: Visual feedback on the user interface

[0906] These steps allow users to receive personalized feedback based on their emotional state and enjoy a better shopping experience.

[0907] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0908] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0909] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0910] [Third embodiment]

[0911] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0912] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0913] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0914] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0915] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0916] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0917] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0918] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0919] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0920] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0921] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0922] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0923] MODE FOR CARRYING OUT THE INVENTION

[0924] The present invention is a system that records a user's performance of a musical instrument, transmits the audio data to a server, and analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness of the user. The system of the present invention allows users to objectively evaluate their own performance and find effective practice methods. The following describes in detail an embodiment of the present invention.

[0925] Overview of program processing

[0926] 1. Receiving and recording audio input

[0927] When the user starts playing, the device launches a recording application. The device records the performance in high quality and saves the recorded data as a digital audio file. During recording, real-time audio data is stored in memory.

[0928] 2. Sending audio data

[0929] After the recording is complete, the device uploads the recorded audio file to a specific server URL via an HTTP request, and the data is securely encoded and transmitted.

[0930] 3. Analysis of audio data

[0931] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially performs pitch analysis, rhythm analysis, dynamic analysis, and expressiveness analysis.

[0932] Pitch analysis:

[0933] The server uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compare it with the pitch written in the sheet music.

[0934] Rhythm analysis:

[0935] The server detects the start and end timing of notes on the time axis and evaluates the degree of agreement with the reference rhythm.

[0936] Dynamic analysis:

[0937] The server detects fluctuations in audio volume and analyzes whether the dynamics of each note are as specified in the musical score.

[0938] Expressiveness analysis:

[0939] The server uses a generative AI model to evaluate the emotional expression of a performance, which is then compared with existing data on famous performances to calculate the degree of match in expressiveness.

[0940] 4. Generating evaluation results

[0941] Based on the analysis results, the server calculates scores for pitch, rhythm, dynamics, and expressiveness, and generates an overall evaluation, which includes numerical scores and graphs for each element.

[0942] The server also lists suggestions for specific phrase practice to help the user improve and adds them to the evaluation results.

[0943] 5. Submitting and displaying evaluation results

[0944] The server then sends the generated evaluation results to the terminal, securely using HTTPS.

[0945] The device analyzes the received evaluation results and displays them visually on the user interface, providing feedback in the form of graphs, numbers, and text, allowing the user to check the detailed evaluation content.

[0946] Specific examples

[0947] If a user were to play "Twinkle Twinkle Little Star" on the piano, the performance would be recorded and analyzed and evaluated using the following procedure.

[0948] The user starts playing the piano and the device records the audio.

[0949] After the performance is finished, the device will stop recording and upload the audio file to the server.

[0950] The server analyzes the received audio data and evaluates it for pitch, rhythm, dynamics, and expressiveness. For example, it may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good.

[0951] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[0952] The server sends the evaluation results and suggestions to the device, which displays them on the user interface. The user can then visually check the feedback, replay the performance, and self-evaluate and try to improve.

[0953] In this way, the system of the present invention provides the user with objective evaluations and specific practice suggestions, and supports improvement in musical instrument playing.

[0954] The processing flow will be explained below.

[0955] Step 1:

[0956] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording is performed in real time, and the audio data is stored in the device's memory. During recording, the device displays the audio waveform and provides a visual indication of the progress.

[0957] Step 2:

[0958] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After saving is complete, the user selects the "Send" button to send the recorded data to the server.

[0959] Step 3:

[0960] In response to user operations, the device sends an HTTP request and uploads the recorded audio file to the server. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[0961] Step 4:

[0962] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of the analysis.

[0963] Step 5:

[0964] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[0965] Step 6:

[0966] Next, the server performs rhythm analysis, detecting the start and end timing of notes on the time axis and comparing them with a reference rhythm template. This evaluates the rhythmic accuracy of the performance. The analysis results are recorded as data including timestamps.

[0967] Step 7:

[0968] The server continues dynamic analysis, analyzing the fluctuations in audio volume to detect the sound pressure level of each note. It compares this with the dynamics instructions in the score to evaluate the appropriateness of the dynamics. The analysis results are recorded as a numerical sound pressure level.

[0969] Step 8:

[0970] Finally, the server analyzes the performance's expressiveness. It uses a generative AI model to evaluate the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data and calculates the degree of agreement in expressiveness. The analysis results are recorded as an emotional score.

[0971] Step 9:

[0972] The server combines all the analysis results to generate an overall evaluation. It also combines scores for pitch, rhythm, dynamics, and expressiveness to create visualizations (e.g., radar charts). It also lists specific phrase practice suggestions to help the user improve and adds them to the evaluation results.

[0973] Step 10:

[0974] The server then sends the generated evaluation results to the device. The transmission is secure using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as specific suggestions for improvement.

[0975] Step 11:

[0976] The device analyzes the received evaluation results and displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check the detailed evaluation. The device also provides "replay" and "re-evaluation" buttons, allowing users to replay and receive a new evaluation.

[0977] Example 1

[0978] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0979] Conventional methods for evaluating musical instrument performance require the presence of an instructor with specialized knowledge, making it difficult to obtain objective and quantitative evaluations. Furthermore, when users self-evaluate, the evaluation tends to be subjective, making it difficult to find effective practice methods. Furthermore, it is difficult to evaluate the emotional expression and subtle dynamics of a performance, limiting the overall improvement of performance skills.

[0980] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0981] In this invention, the server includes a means for analyzing audio data and evaluating scale, rhythm, volume changes, and emotional expression; a means for suggesting practice content necessary for improvement based on the analysis results; and a means for transmitting the evaluation results to the user's device and visually displaying them. This enables objective and quantitative evaluation of the user's performance and provides effective practice methods. Furthermore, using a generative AI model to evaluate expressiveness enables comprehensive evaluation, including emotional expression, helping to improve the user's performance skills.

[0982] "Audio data" is information that digitally records the sound of an instrument played by a user.

[0983] A "communication device" is a device for transmitting and receiving acoustic data via a network.

[0984] "Scale" refers to the reference frequency of each note written on a musical score.

[0985] "Rhythm" is a temporal pattern based on the timing of the start and end of musical notes.

[0986] "Volume change" refers to fluctuations in the volume of the audio during performance.

[0987] "Emotional expression" is an indicator that evaluates the emotions and nuances in a performance.

[0988] A "generative AI model" is an artificial intelligence model that has been trained in advance using machine learning algorithms.

[0989] "Visually displaying" means providing analysis results and feedback to the user in a form that can be seen as graphs or text.

[0990] "Practice content" refers to specific playing phrases and tasks that the user needs to improve.

[0991] MODE FOR CARRYING OUT THE INVENTION

[0992] The present invention is a system that supports users in improving their musical performance skills by recording, analyzing, and evaluating the sounds of musical instruments played by the user with high accuracy. The following describes in detail a method for specifically implementing the system of the present invention.

[0993] System Configuration

[0994] Hardware:

[0995] 1. Terminal: A device operated by a user, such as a smartphone, tablet, or computer.

[0996] 2. Server: A remote server for performing analytical processing, using a computer system with high-performance processing capabilities.

[0997] software:

[0998] 1. Recording application: An application that starts when the user starts playing and records high-quality audio. For example, audio recording applications such as "Audacity" or "Pro Tools" can be used.

[0999] 2. Audio analysis program: A program that runs on the server and analyzes audio data. This analysis uses audio processing tools such as "sox" or "FFmpeg."

[1000] 3. Generative AI model: An AI model equipped with a machine learning algorithm for analyzing emotional expressions. For example, a model pre-trained using "TensorFlow" or "PyTorch" is used.

[1001] Processing Overview

[1002] The user plays an instrument and records the sound on the device. The recorded sound data is then uploaded to the server by the user's operation. The server analyzes the received sound data and evaluates the scale, rhythm, volume changes, and emotional expression. Based on the evaluation results, the system suggests specific practice content and sends the evaluation results to the device, where they are visually displayed.

[1003] Specific actions

[1004] 1. Receiving and recording audio input:

[1005] When the user starts playing, an audio recording application is launched on the device. The sound is recorded in high quality, and the audio data is temporarily stored in memory as it is being recorded. When recording is finished, the audio data is saved as a digital audio file. For example, it may be saved with a file name such as "recording.wav."

[1006] 2. Sending audio data:

[1007] After recording is complete, the device uploads the recorded audio file to a specific server URL via a user-initiated HTTP POST request, with the data encoded and securely transmitted. For example, a request can be sent to the URL "https: / / example.com / upload."

[1008] 3. Analysis of audio data:

[1009] The server stores the received audio data in storage and first performs noise reduction, using audio processing tools such as "sox" for noise filtering.

[1010] Next, the scale is analyzed using FFT (Fast Fourier Transform) to detect the reference frequency of each note.

[1011] The system analyzes rhythm by detecting the start and end timing of notes, and analyzes dynamics by detecting fluctuations in voice volume.

[1012] Emotional expression is evaluated using a pre-trained generative AI model, which compares the performance with existing master performance data to calculate the degree of agreement between the performance's emotional expression.

[1013] 4. Generating evaluation results:

[1014] The server then calculates scores for scale, rhythm, volume, and emotional expression based on the analysis results, generating an overall evaluation that includes a numerical score and graphs, and suggests specific phrases for the user to practice.

[1015] 5. Submitting and viewing the evaluation results:

[1016] The server sends the evaluation results to the device, which then visually displays them on the user interface, for example, by displaying the scale and rhythm evaluation scores in bar graphs or pie charts, allowing the user to check the evaluation of their own performance.

[1017] Example operation

[1018] When a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following steps:

[1019] The user starts playing the piano and the device records the audio.

[1020] After the performance is finished, the device stops recording and uploads the audio file "kirakira_sei.wav" to the server.

[1021] The server analyzes the received audio data and scores it with 85 points for pitch, 70 points for rhythm, 90 points for dynamics, and 80 points for emotional expression.

[1022] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[1023] The server sends the evaluation results and suggestions to the device, which displays them on the user interface, allowing the user to visually confirm the feedback and try playing again.

[1024] This system allows users to easily obtain evaluations of their instrument performance and areas for improvement, effectively improving their own playing skills.

[1025] Prompt Sentence Examples

[1026] "Please analyze the following audio data and rate it for pitch, rhythm, volume changes, and emotional expression:"

[1027] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1028] Step 1: Receiving and Recording Audio Input

[1029] The device launches a recording application to receive audio input, and when the user starts playing an instrument, it records the audio in real time with high quality.

[1030] Input: Audio of the user playing an instrument.

[1031] Data processing: The device temporarily stores the audio data being recorded in memory.

[1032] Output: Once recording is complete, the audio data is saved as a digital audio file (e.g., recording.wav).

[1033] Step 2: Sending audio data

[1034] After the recording is completed, the device uploads the recorded audio file to a specific server URL according to the user's operation.

[1035] Input: Recorded digital audio files.

[1036] Data processing: Encode the audio data and send it to the server using an HTTP POST request.

[1037] Output: The audio data is securely uploaded to the server. For example, send a request to the URL "https: / / example.com / upload".

[1038] Step 3: Analyzing the audio data

[1039] The server first stores the received audio data in storage.

[1040] Input: Uploaded digital audio files.

[1041] Acoustic analysis:

[1042] Data processing: Perform noise reduction and use audio processing tools such as "sox" to remove unwanted noise.

[1043] Data calculation: Analyzes the musical scale using FFT (Fast Fourier Transform) and detects the reference frequency of each note. Detects the start and end timing of notes on the time axis and analyzes rhythm. Detects fluctuations in voice volume and analyzes dynamics.

[1044] Emotional expression rating:

[1045] Data computation: Evaluate the emotional expression of a performance using a generative AI model (e.g., a pre-trained model using TensorFlow or PyTorch). Compare with existing renowned performance data to calculate the degree of consistency of expressiveness.

[1046] Output: The analysis results include scores for pitch, rhythm, volume changes, and emotional expression.

[1047] Step 4: Generate evaluation results

[1048] Based on the analysis results, the server calculates scores for scale, rhythm, volume changes and emotional expression to generate an overall rating.

[1049] Input: Audio analysis results data (scale, rhythm, volume changes, emotional expressions).

[1050] Data calculation: A score is calculated based on each analysis result, and an overall evaluation is made. For example, the scale score is evaluated as 85 points, the rhythm score as 70 points, the dynamics score as 90 points, and the expressiveness score as 80 points.

[1051] Practice suggestions:

[1052] Data calculation: Suggests specific phrase practice to help users improve. These suggestions are automatically generated based on the analysis results.

[1053] Output: Evaluation results and specific practice suggestions are generated.

[1054] Step 5: Submit and view the evaluation results

[1055] The server transmits the generated evaluation results to the user's terminal.

[1056] Input: Assessment results and practice suggestion data.

[1057] Data processing: Formatting the evaluation results for visual presentation.

[1058] Output: Send data securely using HTTPS.

[1059] The terminal visually displays the received evaluation results on a user interface.

[1060] Specific operation: The evaluation results are displayed visually in bar graphs, pie charts, etc., allowing users to check their own performance evaluation.

[1061] (Application example 1)

[1062] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1063] When users play an instrument, it is difficult to obtain objective evaluations of their performance. Furthermore, because it is difficult to evaluate oneself, it is difficult to find effective practice methods. Existing voice analysis systems lack detailed feedback and suggestions, making them inadequate to support users' improvement.

[1064] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1065] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for suggesting phrases necessary for improvement based on the analysis results, means for transmitting the evaluation results to the user's terminal and visually displaying them, and means for displaying the evaluation results on a user interface so that the user can check detailed feedback. This allows the user to receive an objective evaluation of their performance and practice effectively while identifying specific areas for improvement.

[1066] "High-quality recording" refers to technology or means that can record audio signals clearly with low noise.

[1067] "Transmitting audio data" refers to transferring recorded digital data to a server via a network.

[1068] "Pitch analysis" is a technique that detects the frequency of each note in a performance and evaluates the exact pitch.

[1069] "Rhythm analysis" is a technique for analyzing the temporal structure of a performance and evaluating the accuracy of timing.

[1070] "Dynamic analysis" is a technology that detects changes in the volume of a performance and evaluates dynamics and expressiveness.

[1071] A "generative AI model" refers to a model of artificial intelligence that has been trained to perform a specific task using machine learning algorithms.

[1072] "Noise reduction" refers to techniques and methods for removing unwanted background noise from recorded audio data.

[1073] "User interface" refers to the screen and operation method used by the system and the user to exchange information.

[1074] "Overall scoring" is the process of calculating an overall evaluation score based on the results of the analyzed performance.

[1075] The purpose of this invention is to provide a system that allows users to obtain objective evaluations of their musical instrument performance and specific practice methods. Specifically, this system transmits audio data played by the user to a server, analyzes and evaluates pitch, rhythm, dynamics, and expressiveness, and visually displays the results on the user's terminal.

[1076] The system is configured as follows:

[1077] 1. Receiving and recording audio input:

[1078] When a user plays an instrument, a smartphone or other device records the sound in high quality using the smartphone's microphone, and the recorded data is saved as a digital audio file.

[1079] 2. Sending audio data:

[1080] After recording is complete, the device uploads the recorded audio file to a server via HTTPS, where the transmission is encoded to ensure security.

[1081] 3. Analysis of audio data:

[1082] The server stores the received audio data in storage and begins analysis. The analysis performed by the server first includes noise reduction. Then, pitch analysis uses FFT (Fast Fourier Transform), and rhythm analysis involves detecting the start and end timing of notes. Dynamic analysis detects fluctuations in audio volume, and expressiveness is evaluated using a generative AI model. For example, TensorFlow or PyTorch are used.

[1083] 4. Generating evaluation results:

[1084] The server calculates scores for each element based on the analysis results and generates an overall evaluation. Using noise reduction and various analysis methods, it provides detailed scores for pitch, rhythm, dynamics, and expressiveness. It also generates suggestions for specific phrase practice to improve your proficiency based on the evaluation results.

[1085] 5. Submitting and viewing the evaluation results:

[1086] The server then sends the generated evaluation results to the device. This transmission is also securely secured using HTTPS. The device analyzes the received evaluation results and displays them on the user interface. This allows the user to visually check the feedback and confirm the detailed evaluation content.

[1087] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on a smartphone. The recording data is then uploaded to a server, which analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness. For example, the server may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good. The evaluation results are sent to the smartphone and displayed visually on a user interface. Feedback includes suggestions for "short phrases to improve rhythm."

[1088] The system provides users with more accurate and detailed feedback, supporting more effective instrument practice.

[1089] Example prompt sentence:

[1090] "Analyze this performance data and evaluate pitch, rhythm, dynamics, and expressiveness."

[1091] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1092] Step 1: Receiving and Recording Audio Input

[1093] When a user plays an instrument, the device (smartphone) uses a microphone to record the performance sound in high quality. When recording starts, a recording application is launched and real-time audio data is stored in memory. When recording ends, this audio data is saved as a digital audio file. The input is the user's performance sound, and the output is the saved digital audio file.

[1094] Step 2: Sending audio data

[1095] After the recording is complete, the device uploads the saved audio file to the server via HTTPS via user interaction. The audio data is properly encoded before transmission to ensure security. The input of this step is the recorded digital audio file, and the output is the audio data uploaded to the server.

[1096] Step 3: Noise reduction of audio data

[1097] The server stores the received audio data in storage and first performs noise reduction. Digital signal processing technology is used to remove unwanted background noise from the audio data. The input is the audio data stored on the server, and the output is the audio data with the noise removed.

[1098] Step 4: Analyzing the pitch

[1099] The server analyzes the pitch of the noise-removed audio data using FFT (Fast Fourier Transform). It detects the reference frequency of each note and compares it with the pitch written in the music score. The input is the noise-removed audio data, and the output is the result of the pitch analysis.

[1100] Step 5: Analyze the rhythm

[1101] After the pitch analysis is complete, the server performs rhythm analysis. It detects the start and end timing of notes on the time axis and evaluates their correspondence with the reference rhythm. The input is the FFT analysis result, and the output is the result of rhythm analysis.

[1102] Step 6: Dynamic analysis

[1103] The server then performs dynamic analysis, detecting volume fluctuations for each note in the audio data and analyzing whether the dynamic range is as specified in the score. The input is the rhythm analysis result, and the output is the dynamic analysis result.

[1104] Step 7: Expressive Analysis

[1105] The server uses a generative AI model to evaluate the emotional expressiveness of a performance. The model compares it with existing famous performance data to calculate the degree of match in expressiveness. The input is the dynamic analysis result, and the output is the result of expressiveness analysis.

[1106] Step 8: Generate evaluation results

[1107] The server generates an overall score based on the analysis results of pitch, rhythm, dynamics, and expressiveness. It then integrates the analysis results to create numerical scores and graphs for each element. It also lists suggestions for specific phrase practice to help the user improve. The input is all the analysis results, and the output is an overall evaluation result and practice suggestions.

[1108] Step 9: Submit and view the evaluation results

[1109] The server sends the assessment results to the device using HTTPS. The device analyzes the received assessment results and provides them to the user by visually displaying them in a user interface. The display includes graphs, numerical values, and text feedback. The input is the assessment results and practice suggestions, and the output is the feedback displayed in the user interface.

[1110] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it in step 1. The performance is uploaded to the server in step 2, and various analyses are performed in steps 3 to 7. An overall evaluation is created in step 8, and the results are displayed on the user's device in step 9. An example of a prompt sentence is "Analyze this performance data and evaluate its pitch, rhythm, dynamics, and expressiveness."

[1111] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1112] MODE FOR CARRYING OUT THE INVENTION

[1113] The present invention is a system for supporting the improvement of musical instrument performance. It not only records, analyzes, and evaluates a user's performance data, but also reflects the user's emotional state in the evaluation. This system allows users to objectively evaluate their own performance and find appropriate practice methods. Specific embodiments for implementing the present invention are described below.

[1114] Overview of program processing

[1115] 1. Receiving and recording audio input

[1116] When the user starts playing, the device launches a recording application and records the performance sound in high quality through the microphone. Recording is done in real time, and the audio data is saved in the device's memory. The audio waveform and recording time are displayed during recording.

[1117] 2. Sending audio data

[1118] When the performance is finished and the user stops recording, the device saves the recording in the specified audio format. When the user selects the "Send" button, the recording is uploaded to the server.

[1119] 3. Analysis of audio data

[1120] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially analyzes pitch, rhythm, dynamics, and expressiveness.

[1121] Pitch analysis: Uses FFT to identify the base frequency of a note and compares it to the pitch written in the sheet music.

[1122] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[1123] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[1124] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[1125] 4. Emotion Recognition by Emotion Engine

[1126] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them using the emotion engine.

[1127] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[1128] 5. Generating evaluation results

[1129] The server integrates the results of the voice data analysis and the emotion engine to generate a comprehensive evaluation. It calculates a score based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and creates visualization data.

[1130] The server also generates advice and messages to help maintain motivation that are tailored to the user's emotional state and adds them to the suggestion list.

[1131] 6. Submitting and displaying evaluation results

[1132] The server then sends the generated evaluation results to the device. The transmission is secure and done using HTTPS.

[1133] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also provided.

[1134] Specific examples

[1135] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[1136] The user starts playing the piano and the device records the audio.

[1137] After the performance is over, the device saves the recording data and uploads it to the server.

[1138] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[1139] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[1140] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1141] The server sends the evaluation results and suggestions to the device, which then displays them on the user interface, allowing the user to receive visual feedback and advice based on their emotional state.

[1142] In this way, the system of the present invention provides the user with objective performance evaluation as well as feedback according to their emotional state, supporting enjoyable and effective improvement in musical instrument playing.

[1143] The processing flow will be explained below.

[1144] MODE FOR CARRYING OUT THE INVENTION

[1145] Specific flow of program processing

[1146] Step 1:

[1147] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording occurs in real time, and the audio data is stored in the device's memory. During recording, the audio waveform and recording time are displayed on the screen, allowing the user to check the progress.

[1148] Step 2:

[1149] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After the recording data has been saved, the user selects the "Send" button to send the recording data to the server.

[1150] Step 3:

[1151] The device uploads the recorded audio file to the server via an HTTP request in response to user input. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[1152] Step 4:

[1153] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of subsequent analysis.

[1154] Step 5:

[1155] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[1156] Step 6:

[1157] The server then performs rhythm analysis, detecting the start and end times of notes on the time axis and comparing them with a reference rhythm template. The analysis results are recorded as data including timestamps.

[1158] Step 7:

[1159] The server also performs dynamic analysis, analyzing fluctuations in audio volume and assessing whether the dynamics of each note are in line with the musical notation. The results are recorded as a numerical value representing the sound pressure level.

[1160] Step 8:

[1161] Next, the server performs expressiveness analysis. Using a generative AI model, it evaluates the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data to calculate the degree of agreement in expressiveness, and records the analysis results as an emotional score.

[1162] Step 9:

[1163] The device captures the user's facial expressions and tone of voice during and after performance, and activates the emotion engine, which analyzes the user's facial expressions and tone of voice to evaluate the user's emotional state in real time.

[1164] Step 10:

[1165] The server combines the results of the voice data analysis with the emotion analysis results from the emotion engine to generate an overall evaluation. It also combines scores based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state to create visualization data. It also generates motivational advice and messages according to the user's emotional state.

[1166] Step 11:

[1167] The server then sends the generated evaluation results and motivation advice to the device. The transmission is securely performed using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as advice based on the user's emotional state.

[1168] Step 12:

[1169] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check detailed evaluations and motivational advice. The device also has "replay" and "reevaluation" buttons, allowing users to play again and receive a new evaluation.

[1170] Specific examples

[1171] For example, if a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and the following process is performed.

[1172] The user starts playing and the terminal starts recording.

[1173] After the performance is finished, the recording data is saved on the device and sent to the server.

[1174] The server analyzes the received data and evaluates it for pitch, rhythm, strength, and expressiveness. At the same time, the device analyzes the user's facial expressions and voice and evaluates their emotional state using an emotion engine.

[1175] The server integrates the overall analysis results and generates an overall evaluation such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotional evaluation."

[1176] The server suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1177] The device displays the evaluation results and suggestions on a user interface, allowing the user to see visual feedback and motivational advice.

[1178] As described above, the system of the present invention provides the user with objective performance evaluation and feedback based on emotional state, enabling enjoyable and effective improvement in musical instrument playing.

[1179] Example 2

[1180] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1181] Conventional systems for supporting the improvement of musical instrument playing skills are specialized in evaluating the user's performance technique, but are unable to provide comprehensive feedback that takes into account the user's emotional state. This makes it difficult to support maintaining motivation or improving the emotional expressiveness of performance. Furthermore, conventional systems only evaluate performance technique criteria such as pitch, rhythm, and dynamics, and are limited in their assessment of expressiveness.

[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness, means for capturing the user's facial expression and tone of voice and analyzing their emotional state, and means for integrating the emotional state into the analysis results and generating an overall evaluation. This makes it possible to provide comprehensive feedback that integrates the user's performance technique and emotional state, and to support maintaining motivation and improving the emotional expressiveness of performance.

[1183] A "user" is an individual who plays an instrument and inputs performance data into the system.

[1184] "Terminal" refers to an electronic device that a user uses to record musical instrument performances and display analysis results.

[1185] "Server" refers to a computer system that receives and analyzes voice data sent from a terminal, generates evaluation results, and sends them to the terminal.

[1186] "Performing" refers to the act of a user making sounds using an instrument.

[1187] "Audio data" refers to data that is a digital recording of music played by a user.

[1188] "Noise reduction" refers to the process of removing unnecessary noise from audio data.

[1189] "FFT" stands for Fast Fourier Transform, a transformation technique frequently used in signal processing.

[1190] A "generative AI model" is a model that uses artificial intelligence to generate and analyze data.

[1191] A "prompt sentence" is an input sentence given to a generative AI model.

[1192] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[1193] "Evaluation result" refers to comprehensive feedback generated based on analysis of speech data and emotional state.

[1194] "Visualized data" refers to data that displays the evaluation results in graphs, charts, etc. so that users can intuitively understand them.

[1195] "Advice" refers to specific guidelines suggested to help users improve their playing skills and motivation.

[1196] This invention is a system to support the improvement of musical instrument performance by recording, analyzing, and evaluating the user's performance data, and also reflecting the user's emotional state in the evaluation. This system allows the user to objectively evaluate their own performance and find appropriate practice methods.

[1197] Recording and sending audio

[1198] When a user starts playing an instrument, the device launches a recording application and records the performance through the microphone in high quality. The recorded audio data is stored in the device's memory and sent to a server as needed. During recording, the audio waveform and recording time are displayed, allowing the user to understand their performance in real time.

[1199] Analysis of audio data

[1200] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it performs the following analysis steps in sequence.

[1201] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of the note and compare it to the pitch written in the sheet music.

[1202] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[1203] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[1204] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[1205] emotion recognition

[1206] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The camera and microphone capture the user's facial expressions and tone of voice and analyze them using the emotion engine. The server integrates the analysis results of the emotion engine with the analysis results of the voice data and reflects the user's emotional state in an overall evaluation.

[1207] Generating and displaying evaluation results

[1208] The server combines the results of the voice data analysis with the emotion engine's results to generate an overall evaluation. A score is calculated based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and visualization data is created. Furthermore, advice and messages to maintain motivation according to the user's emotional state are generated and added to a suggestion list. The server sends the generated evaluation results to the device, which then visually displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also displayed.

[1209] Examples of prompt statements

[1210] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[1211] The user starts playing the piano and the device records the audio.

[1212] After the performance is over, the device saves the recording data and uploads it to the server.

[1213] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[1214] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[1215] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1216] The server sends the evaluation results and suggestions to the terminal, which displays them on the user interface.

[1217] An example of a prompt for input to a generative AI model is as follows:

[1218] "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and score them on pitch, rhythm, dynamics, and expressiveness. Also, analyze the user's facial expressions and tone of voice to assess their emotional state and create comprehensive feedback."

[1219] In this way, the system of the present invention can provide the user with objective performance evaluation as well as feedback according to their emotional state, thereby supporting enjoyable and effective improvement in musical instrument playing.

[1220] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1221] Step 1: Receiving and Recording Audio Input

[1222] When the user starts playing an instrument, the device launches a recording application and records the performance in high quality through the microphone.

[1223] Input: User's playing sound, recording application

[1224] Specific operation: The recording application captures the audio input from the microphone and saves it as audio data in the device's memory. During recording, the audio waveform and recording time are displayed in real time.

[1225] Output: Recorded audio data

[1226] Step 2: Save and send audio data

[1227] When the user finishes playing and presses the "Send" button, the device stops recording and saves the recorded data.

[1228] Input: User click on "Submit" button, recorded voice data

[1229] Specific operation: Calls the StopRecording() method after recording is complete. The saved audio data is saved as a file in the specified audio format (e.g. WAV, MP3).

[1230] Output: Audio data saved in the specified format

[1231] The terminal uploads the audio data to the server.

[1232] Input: Saved audio data

[1233] Specific behavior: Generates an HTTP POST request to send the saved audio data to the server.

[1234] Output: Audio data uploaded to the server

[1235] Step 3: Analyzing the audio data

[1236] The server stores the received audio data in temporary storage and begins analysis.

[1237] Input: Uploaded audio data

[1238] Specific operation: Save the audio data as a temporary file and perform the following analysis steps.

[1239] Noise reduction: Removes unwanted noise from audio data.

[1240] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of a note and compare it with a reference value.

[1241] Rhythm analysis: Detects the start and end times of notes and compares them with a reference rhythm.

[1242] Dynamics analysis: Analyzes fluctuations in voice volume and evaluates the appropriateness of dynamics.

[1243] Expressiveness analysis: Using a generative AI model, the emotional expression of a performance is evaluated based on a prompt.

[1244] Prompt: "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and rate it on pitch, rhythm, dynamics, and expressiveness. Also, rate the user's emotional state by analyzing their facial expressions and tone of voice, and create overall feedback."

[1245] Output: Pitch analysis results, rhythm analysis results, dynamic analysis results, expressiveness analysis results

[1246] Step 4: Emotion Recognition

[1247] The device captures the user's facial expressions and tone of voice to activate the emotion engine.

[1248] Input: User's facial expression (camera image), tone of voice (audio data)

[1249] Specific operation: Uses a camera and microphone to capture facial expression data and voice tone data in real time, and analyzes the data using an emotion engine.

[1250] Output: Parsed emotion data

[1251] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[1252] Input: Voice analysis results, emotion analysis results

[1253] Specific operation: The results of speech analysis and sentiment analysis are integrated, and an evaluation algorithm is applied to generate an overall evaluation.

[1254] Output: Integrated overall rating

[1255] Step 5: Generate evaluation results

[1256] The server integrates the results of the voice data analysis and the emotion engine results to generate an overall evaluation.

[1257] Input: Integrated overall rating

[1258] Specific actions: Based on the integrated overall evaluation, the system calculates a score based on pitch, rhythm, dynamics, expressiveness, and emotional state, creates visualization data, and generates advice and messages to maintain motivation and adds them to a suggestion list.

[1259] Output: Overall evaluation data, visualization data (graphs, charts, etc.), motivation advice

[1260] Step 6: Submit and view the results of the evaluation

[1261] The server transmits the generated evaluation results to the terminal.

[1262] Input: Overall evaluation data, Visualization data, Advice

[1263] Specific operation: Evaluation results, visualization data, and advice are securely sent to the device using HTTPS.

[1264] Output: Data sent to the terminal

[1265] The terminal receives the evaluation results and visually displays them on a user interface.

[1266] Input: Overall evaluation data, Visualization data, Advice

[1267] Specific operation: Analyzes the evaluation results and visualizes the feedback as graphs, numbers, and text. Users can check the detailed evaluation results and appropriate motivation advice.

[1268] Output: Visual feedback to the user

[1269] Through the above steps, the system of the present invention can provide the user with detailed performance evaluations and appropriate feedback, and support the user in improving their musical instrument playing in an enjoyable and effective way.

[1270] (Application example 2)

[1271] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1272] Current systems for supporting the improvement of musical instrument performance only provide objective feedback on the user's performance, and are unable to provide comprehensive evaluations or suggestions that take into account the user's emotional state. As a result, they are unable to suggest phrases or products that the user truly needs, resulting in a lack of motivation and personalized support.

[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1274] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for analyzing the user's emotions from their facial expressions and tone of voice, means for suggesting phrases and products necessary for improvement based on the analysis results and emotion analysis results, and means for transmitting the evaluation results to the user's terminal and visually displaying them, thereby enabling comprehensive and personalized feedback on the user's performance.

[1275] The "recording means" is a device or software for recording the sound played by the user in high quality.

[1276] The "transmission means" is a communication device or software for transmitting the recorded voice data to the server.

[1277] "Analysis means" refers to a device or software for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness.

[1278] "Emotion analysis means" refers to a device or software that analyzes a user's facial expression and tone of voice to evaluate their emotional state.

[1279] The "suggestion means" is a device or software that suggests phrases and products necessary for improvement based on the analysis results and emotion analysis results.

[1280] The "display means" is a device or software for transmitting the evaluation results to the user's terminal and visually displaying them.

[1281] The "frequency analysis means" is a device or software for analyzing the frequency of audio data in order to detect pitch.

[1282] "Rhythm analysis means" refers to a device or software for analyzing the start and end timing of notes in order to detect rhythm.

[1283] A "volume analysis means" is a device or software for analyzing variations in voice volume to detect strength and weakness.

[1284] A "generative AI model" is a machine learning model used to evaluate expressiveness.

[1285] A "prompt sentence" is an input sentence for generating product suggestions from the user's emotional state.

[1286] This invention is a virtual store shopping assistant system that makes personalized product recommendations taking into account the user's emotional state. This system records the sound of a musical instrument played by the user in high quality, analyzes the data, and analyzes the user's emotions from their facial expressions and tone of voice to provide comprehensive feedback.

[1287] Hardware and Software

[1288] This system is implemented using the following hardware and software:

[1289] Hardware:

[1290] Smartphone or smart glasses

[1291] Camera (for facial expression analysis)

[1292] Microphone (for recording audio)

[1293] software:

[1294] Python

[1295] OpenCV (for facial expression analysis)

[1296] TensorFlow (for facial expression analysis model)

[1297] Sounddevice (for audio recording)

[1298] Soundfile (for manipulating audio data)

[1299] Requests (for sending data)

[1300] Processing flow

[1301] Audio recording and analysis

[1302] The device records the user's voice when they ask a question or make a comment about a product. The recorded voice data is sent to a server, where it is analyzed, including frequency analysis. Specific analysis items include pitch, rhythm, dynamics, and expressiveness.

[1303] facial expression analysis

[1304] The device's built-in camera captures the user's facial expression data, which is then processed by a generative AI model using TensorFlow for emotion recognition. The resulting data is used to evaluate the user's level of interest or anxiety.

[1305] Suggestions and Displays

[1306] The server combines the results of voice analysis and facial expression analysis to generate prompts and make appropriate product suggestions. Based on the generated prompts, the system makes personalized suggestions to the user. The suggestions are displayed on the device.

[1307] Specific examples

[1308] If a user is browsing home decor products in a virtual shopping application and asks, "How big is this sofa?", the following occurs:

[1309] 1. Audio recording: The device's microphone records the audio.

[1310] 2. Audio analysis: Analyze pitch, rhythm, dynamics, and expressiveness from recorded data.

[1311] 3. Facial Expression Analysis: The camera captures the user's facial expressions and evaluates the user's emotional state.

[1312] 4. Proposal generation: Based on the results of voice analysis and facial expression analysis, a generative AI model is used to generate optimal product proposals for the user.

[1313] Prompt Sentence Examples

[1314] The prompt for "How big is this sofa?" is as follows:

[1315] Analyze the customer's interest and anxiety based on their questions while they are browsing products, and provide optimal product suggestions. Also, analyze their facial expressions to see their level of tension or interest, and provide explanations with kind words.

[1316] This allows users to receive personalized feedback based on their emotional state, resulting in a better shopping experience.

[1317] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1318] Step 1:

[1319] The device records the user's voice when they ask a question or make a comment about a product. Specifically, it captures high-quality audio data through a microphone and saves it in an audio format such as WAV. The recorded data is used as is in the next step.

[1320] Input: User's voice

[1321] Output: High-quality audio data (WAV format)

[1322] Step 2:

[1323] The device sends the recorded audio data to the server. The data is transferred using a secure communication protocol such as HTTPS. The server then stores the received audio data in its storage.

[1324] Input: High-quality audio data

[1325] Output: Upload audio data to the server

[1326] Step 3:

[1327] The server analyzes the transmitted audio data. The analysis items are pitch, rhythm, dynamics, and expressiveness. Specific processing includes frequency analysis using FFT, time-domain note analysis, detection of voice volume fluctuations, and evaluation of emotional expression using a generative AI model.

[1328] Input: High-quality audio data

[1329] Output: Evaluation results for pitch, rhythm, dynamics, and expressiveness

[1330] Step 4:

[1331] The device uses a camera to capture the user's facial expressions. The captured video data is input into a generative AI model using TensorFlow to recognize the user's emotional state from their facial expressions. Specifically, it extracts facial features and classifies emotions based on them.

[1332] Input: User facial expression video data

[1333] Output: User's emotional state (interest, anxiety, etc.)

[1334] Step 5:

[1335] The server combines the results of voice analysis and facial expression analysis to generate prompts, which are then used by a generative AI model to provide optimal product recommendations to the user.

[1336] Input: Pitch, rhythm, dynamics, expressiveness evaluation results, and the user's emotional state

[1337] Output: prompt statement

[1338] Step 6:

[1339] The server creates specific product suggestions based on the generated prompt sentences, and provides the suggestions to the user as personalized feedback.

[1340] Input: prompt statement

[1341] Output: Personalized product recommendations

[1342] Step 7:

[1343] The terminal receives the suggestions sent from the server and visually displays them on the user interface. Specifically, the suggestions are displayed as text, graphs, or images, providing visual feedback to the user.

[1344] Input: Proposal from the server

[1345] Output: Visual feedback on the user interface

[1346] These steps allow users to receive personalized feedback based on their emotional state and enjoy a better shopping experience.

[1347] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1348] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1349] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1350] [Fourth embodiment]

[1351] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1352] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1353] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1354] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1355] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1356] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1357] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1358] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1359] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1360] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1361] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1362] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1363] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1364] MODE FOR CARRYING OUT THE INVENTION

[1365] The present invention is a system that records a user's performance of a musical instrument, transmits the audio data to a server, and analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness of the user. The system of the present invention allows users to objectively evaluate their own performance and find effective practice methods. The following describes in detail an embodiment of the present invention.

[1366] Overview of program processing

[1367] 1. Receiving and recording audio input

[1368] When the user starts playing, the device launches a recording application. The device records the performance in high quality and saves the recorded data as a digital audio file. During recording, real-time audio data is stored in memory.

[1369] 2. Sending audio data

[1370] After the recording is complete, the device uploads the recorded audio file to a specific server URL via an HTTP request, and the data is securely encoded and transmitted.

[1371] 3. Analysis of audio data

[1372] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially performs pitch analysis, rhythm analysis, dynamic analysis, and expressiveness analysis.

[1373] Pitch analysis:

[1374] The server uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compare it with the pitch written in the sheet music.

[1375] Rhythm analysis:

[1376] The server detects the start and end timing of notes on the time axis and evaluates the degree of agreement with the reference rhythm.

[1377] Dynamic analysis:

[1378] The server detects fluctuations in audio volume and analyzes whether the dynamics of each note are as specified in the musical score.

[1379] Expressiveness analysis:

[1380] The server uses a generative AI model to evaluate the emotional expression of a performance, which is then compared with existing data on famous performances to calculate the degree of match in expressiveness.

[1381] 4. Generating evaluation results

[1382] Based on the analysis results, the server calculates scores for pitch, rhythm, dynamics, and expressiveness, and generates an overall evaluation, which includes numerical scores and graphs for each element.

[1383] The server also lists suggestions for specific phrase practice to help the user improve and adds them to the evaluation results.

[1384] 5. Submitting and displaying evaluation results

[1385] The server then sends the generated evaluation results to the terminal, securely using HTTPS.

[1386] The device analyzes the received evaluation results and displays them visually on the user interface, providing feedback in the form of graphs, numbers, and text, allowing the user to check the detailed evaluation content.

[1387] Specific examples

[1388] If a user were to play "Twinkle Twinkle Little Star" on the piano, the performance would be recorded and analyzed and evaluated using the following procedure.

[1389] The user starts playing the piano and the device records the audio.

[1390] After the performance is finished, the device will stop recording and upload the audio file to the server.

[1391] The server analyzes the received audio data and evaluates it for pitch, rhythm, dynamics, and expressiveness. For example, it may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good.

[1392] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[1393] The server sends the evaluation results and suggestions to the device, which displays them on the user interface. The user can then visually check the feedback, replay the performance, and self-evaluate and try to improve.

[1394] In this way, the system of the present invention provides the user with objective evaluations and specific practice suggestions, and supports improvement in musical instrument playing.

[1395] The processing flow will be explained below.

[1396] Step 1:

[1397] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording is performed in real time, and the audio data is stored in the device's memory. During recording, the device displays the audio waveform and provides a visual indication of the progress.

[1398] Step 2:

[1399] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After saving is complete, the user selects the "Send" button to send the recorded data to the server.

[1400] Step 3:

[1401] In response to user operations, the device sends an HTTP request and uploads the recorded audio file to the server. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[1402] Step 4:

[1403] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of the analysis.

[1404] Step 5:

[1405] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[1406] Step 6:

[1407] Next, the server performs rhythm analysis, detecting the start and end timing of notes on the time axis and comparing them with a reference rhythm template. This evaluates the rhythmic accuracy of the performance. The analysis results are recorded as data including timestamps.

[1408] Step 7:

[1409] The server continues dynamic analysis, analyzing the fluctuations in audio volume to detect the sound pressure level of each note. It compares this with the dynamics instructions in the score to evaluate the appropriateness of the dynamics. The analysis results are recorded as a numerical sound pressure level.

[1410] Step 8:

[1411] Finally, the server analyzes the performance's expressiveness. It uses a generative AI model to evaluate the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data and calculates the degree of agreement in expressiveness. The analysis results are recorded as an emotional score.

[1412] Step 9:

[1413] The server combines all the analysis results to generate an overall evaluation. It also combines scores for pitch, rhythm, dynamics, and expressiveness to create visualizations (e.g., radar charts). It also lists specific phrase practice suggestions to help the user improve and adds them to the evaluation results.

[1414] Step 10:

[1415] The server then sends the generated evaluation results to the device. The transmission is secure using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as specific suggestions for improvement.

[1416] Step 11:

[1417] The device analyzes the received evaluation results and displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check the detailed evaluation. The device also provides "replay" and "re-evaluation" buttons, allowing users to replay and receive a new evaluation.

[1418] Example 1

[1419] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1420] Conventional methods for evaluating musical instrument performance require the presence of an instructor with specialized knowledge, making it difficult to obtain objective and quantitative evaluations. Furthermore, when users self-evaluate, the evaluation tends to be subjective, making it difficult to find effective practice methods. Furthermore, it is difficult to evaluate the emotional expression and subtle dynamics of a performance, limiting the overall improvement of performance skills.

[1421] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1422] In this invention, the server includes a means for analyzing audio data and evaluating scale, rhythm, volume changes, and emotional expression; a means for suggesting practice content necessary for improvement based on the analysis results; and a means for transmitting the evaluation results to the user's device and visually displaying them. This enables objective and quantitative evaluation of the user's performance and provides effective practice methods. Furthermore, using a generative AI model to evaluate expressiveness enables comprehensive evaluation, including emotional expression, helping to improve the user's performance skills.

[1423] "Audio data" is information that digitally records the sound of an instrument played by a user.

[1424] A "communication device" is a device for transmitting and receiving acoustic data via a network.

[1425] "Scale" refers to the reference frequency of each note written on a musical score.

[1426] "Rhythm" is a temporal pattern based on the timing of the start and end of musical notes.

[1427] "Volume change" refers to fluctuations in the volume of the audio during performance.

[1428] "Emotional expression" is an indicator that evaluates the emotions and nuances in a performance.

[1429] A "generative AI model" is an artificial intelligence model that has been trained in advance using machine learning algorithms.

[1430] "Visually displaying" means providing analysis results and feedback to the user in a form that can be seen as graphs or text.

[1431] "Practice content" refers to specific playing phrases and tasks that the user needs to improve.

[1432] MODE FOR CARRYING OUT THE INVENTION

[1433] The present invention is a system that supports users in improving their musical performance skills by recording, analyzing, and evaluating the sounds of musical instruments played by the user with high accuracy. The following describes in detail a method for specifically implementing the system of the present invention.

[1434] System Configuration

[1435] Hardware:

[1436] 1. Terminal: A device operated by a user, such as a smartphone, tablet, or computer.

[1437] 2. Server: A remote server for performing analytical processing, using a computer system with high-performance processing capabilities.

[1438] software:

[1439] 1. Recording application: An application that starts when the user starts playing and records high-quality audio. For example, audio recording applications such as "Audacity" or "Pro Tools" can be used.

[1440] 2. Audio analysis program: A program that runs on the server and analyzes audio data. This analysis uses audio processing tools such as "sox" or "FFmpeg."

[1441] 3. Generative AI model: An AI model equipped with a machine learning algorithm for analyzing emotional expressions. For example, a model pre-trained using "TensorFlow" or "PyTorch" is used.

[1442] Processing Overview

[1443] The user plays an instrument and records the sound on the device. The recorded sound data is then uploaded to the server by the user's operation. The server analyzes the received sound data and evaluates the scale, rhythm, volume changes, and emotional expression. Based on the evaluation results, the system suggests specific practice content and sends the evaluation results to the device, where they are visually displayed.

[1444] Specific actions

[1445] 1. Receiving and recording audio input:

[1446] When the user starts playing, an audio recording application is launched on the device. The sound is recorded in high quality, and the audio data is temporarily stored in memory as it is being recorded. When recording is finished, the audio data is saved as a digital audio file. For example, it may be saved with a file name such as "recording.wav."

[1447] 2. Sending audio data:

[1448] After recording is complete, the device uploads the recorded audio file to a specific server URL via a user-initiated HTTP POST request, with the data encoded and securely transmitted. For example, a request can be sent to the URL "https: / / example.com / upload."

[1449] 3. Analysis of audio data:

[1450] The server stores the received audio data in storage and first performs noise reduction, using audio processing tools such as "sox" for noise filtering.

[1451] Next, the scale is analyzed using FFT (Fast Fourier Transform) to detect the reference frequency of each note.

[1452] The system analyzes rhythm by detecting the start and end timing of notes, and analyzes dynamics by detecting fluctuations in voice volume.

[1453] Emotional expression is evaluated using a pre-trained generative AI model, which compares the performance with existing master performance data to calculate the degree of agreement between the performance's emotional expression.

[1454] 4. Generating evaluation results:

[1455] The server then calculates scores for scale, rhythm, volume, and emotional expression based on the analysis results, generating an overall evaluation that includes a numerical score and graphs, and suggests specific phrases for the user to practice.

[1456] 5. Submitting and viewing the evaluation results:

[1457] The server sends the evaluation results to the device, which then visually displays them on the user interface, for example, by displaying the scale and rhythm evaluation scores in bar graphs or pie charts, allowing the user to check the evaluation of their own performance.

[1458] Example operation

[1459] When a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following steps:

[1460] The user starts playing the piano and the device records the audio.

[1461] After the performance is finished, the device stops recording and uploads the audio file "kirakira_sei.wav" to the server.

[1462] The server analyzes the received audio data and scores it with 85 points for pitch, 70 points for rhythm, 90 points for dynamics, and 80 points for emotional expression.

[1463] The server integrates the analysis results and suggests "short phrases to improve rhythm" to the user.

[1464] The server sends the evaluation results and suggestions to the device, which displays them on the user interface, allowing the user to visually confirm the feedback and try playing again.

[1465] This system allows users to easily obtain evaluations of their instrument performance and areas for improvement, effectively improving their own playing skills.

[1466] Prompt Sentence Examples

[1467] "Please analyze the following audio data and rate it for pitch, rhythm, volume changes, and emotional expression:"

[1468] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1469] Step 1: Receiving and Recording Audio Input

[1470] The device launches a recording application to receive audio input, and when the user starts playing an instrument, it records the audio in real time with high quality.

[1471] Input: Audio of the user playing an instrument.

[1472] Data processing: The device temporarily stores the audio data being recorded in memory.

[1473] Output: Once recording is complete, the audio data is saved as a digital audio file (e.g., recording.wav).

[1474] Step 2: Sending audio data

[1475] After the recording is completed, the device uploads the recorded audio file to a specific server URL according to the user's operation.

[1476] Input: Recorded digital audio files.

[1477] Data processing: Encode the audio data and send it to the server using an HTTP POST request.

[1478] Output: The audio data is securely uploaded to the server. For example, send a request to the URL "https: / / example.com / upload".

[1479] Step 3: Analyzing the audio data

[1480] The server first stores the received audio data in storage.

[1481] Input: Uploaded digital audio files.

[1482] Acoustic analysis:

[1483] Data processing: Perform noise reduction and use audio processing tools such as "sox" to remove unwanted noise.

[1484] Data calculation: Analyzes the musical scale using FFT (Fast Fourier Transform) and detects the reference frequency of each note. Detects the start and end timing of notes on the time axis and analyzes rhythm. Detects fluctuations in voice volume and analyzes dynamics.

[1485] Emotional expression rating:

[1486] Data computation: Evaluate the emotional expression of a performance using a generative AI model (e.g., a pre-trained model using TensorFlow or PyTorch). Compare with existing renowned performance data to calculate the degree of consistency of expressiveness.

[1487] Output: The analysis results include scores for pitch, rhythm, volume changes, and emotional expression.

[1488] Step 4: Generate evaluation results

[1489] Based on the analysis results, the server calculates scores for scale, rhythm, volume changes and emotional expression to generate an overall rating.

[1490] Input: Audio analysis results data (scale, rhythm, volume changes, emotional expressions).

[1491] Data calculation: A score is calculated based on each analysis result, and an overall evaluation is made. For example, the scale score is evaluated as 85 points, the rhythm score as 70 points, the dynamics score as 90 points, and the expressiveness score as 80 points.

[1492] Practice suggestions:

[1493] Data calculation: Suggests specific phrase practice to help users improve. These suggestions are automatically generated based on the analysis results.

[1494] Output: Evaluation results and specific practice suggestions are generated.

[1495] Step 5: Submit and view the evaluation results

[1496] The server transmits the generated evaluation results to the user's terminal.

[1497] Input: Assessment results and practice suggestion data.

[1498] Data processing: Formatting the evaluation results for visual presentation.

[1499] Output: Send data securely using HTTPS.

[1500] The terminal visually displays the received evaluation results on a user interface.

[1501] Specific operation: The evaluation results are displayed visually in bar graphs, pie charts, etc., allowing users to check their own performance evaluation.

[1502] (Application example 1)

[1503] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1504] When users play an instrument, it is difficult to obtain objective evaluations of their performance. Furthermore, because it is difficult to evaluate oneself, it is difficult to find effective practice methods. Existing voice analysis systems lack detailed feedback and suggestions, making them inadequate to support users' improvement.

[1505] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1506] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for suggesting phrases necessary for improvement based on the analysis results, means for transmitting the evaluation results to the user's terminal and visually displaying them, and means for displaying the evaluation results on a user interface so that the user can check detailed feedback. This allows the user to receive an objective evaluation of their performance and practice effectively while identifying specific areas for improvement.

[1507] "High-quality recording" refers to technology or means that can record audio signals clearly with low noise.

[1508] "Transmitting audio data" refers to transferring recorded digital data to a server via a network.

[1509] "Pitch analysis" is a technique that detects the frequency of each note in a performance and evaluates the exact pitch.

[1510] "Rhythm analysis" is a technique for analyzing the temporal structure of a performance and evaluating the accuracy of timing.

[1511] "Dynamic analysis" is a technology that detects changes in the volume of a performance and evaluates dynamics and expressiveness.

[1512] A "generative AI model" refers to a model of artificial intelligence that has been trained to perform a specific task using machine learning algorithms.

[1513] "Noise reduction" refers to techniques and methods for removing unwanted background noise from recorded audio data.

[1514] "User interface" refers to the screen and operation method used by the system and the user to exchange information.

[1515] "Overall scoring" is the process of calculating an overall evaluation score based on the results of the analyzed performance.

[1516] The purpose of this invention is to provide a system that allows users to obtain objective evaluations of their musical instrument performance and specific practice methods. Specifically, this system transmits audio data played by the user to a server, analyzes and evaluates pitch, rhythm, dynamics, and expressiveness, and visually displays the results on the user's terminal.

[1517] The system is configured as follows:

[1518] 1. Receiving and recording audio input:

[1519] When a user plays an instrument, a smartphone or other device records the sound in high quality using the smartphone's microphone, and the recorded data is saved as a digital audio file.

[1520] 2. Sending audio data:

[1521] After recording is complete, the device uploads the recorded audio file to a server via HTTPS, where the transmission is encoded to ensure security.

[1522] 3. Analysis of audio data:

[1523] The server stores the received audio data in storage and begins analysis. The analysis performed by the server first includes noise reduction. Then, pitch analysis uses FFT (Fast Fourier Transform), and rhythm analysis involves detecting the start and end timing of notes. Dynamic analysis detects fluctuations in audio volume, and expressiveness is evaluated using a generative AI model. For example, TensorFlow or PyTorch are used.

[1524] 4. Generating evaluation results:

[1525] The server calculates scores for each element based on the analysis results and generates an overall evaluation. Using noise reduction and various analysis methods, it provides detailed scores for pitch, rhythm, dynamics, and expressiveness. It also generates suggestions for specific phrase practice to improve your proficiency based on the evaluation results.

[1526] 5. Submitting and viewing the evaluation results:

[1527] The server then sends the generated evaluation results to the device. This transmission is also securely secured using HTTPS. The device analyzes the received evaluation results and displays them on the user interface. This allows the user to visually check the feedback and confirm the detailed evaluation content.

[1528] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it on a smartphone. The recording data is then uploaded to a server, which analyzes and evaluates the pitch, rhythm, dynamics, and expressiveness. For example, the server may evaluate the pitch as accurate, the rhythm as slightly unstable, the dynamics as appropriate, and the expressiveness as good. The evaluation results are sent to the smartphone and displayed visually on a user interface. Feedback includes suggestions for "short phrases to improve rhythm."

[1529] The system provides users with more accurate and detailed feedback, supporting more effective instrument practice.

[1530] Example prompt sentence:

[1531] "Analyze this performance data and evaluate pitch, rhythm, dynamics, and expressiveness."

[1532] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1533] Step 1: Receiving and Recording Audio Input

[1534] When a user plays an instrument, the device (smartphone) uses a microphone to record the performance sound in high quality. When recording starts, a recording application is launched and real-time audio data is stored in memory. When recording ends, this audio data is saved as a digital audio file. The input is the user's performance sound, and the output is the saved digital audio file.

[1535] Step 2: Sending audio data

[1536] After the recording is complete, the device uploads the saved audio file to the server via HTTPS via user interaction. The audio data is properly encoded before transmission to ensure security. The input of this step is the recorded digital audio file, and the output is the audio data uploaded to the server.

[1537] Step 3: Noise reduction of audio data

[1538] The server stores the received audio data in storage and first performs noise reduction. Digital signal processing technology is used to remove unwanted background noise from the audio data. The input is the audio data stored on the server, and the output is the audio data with the noise removed.

[1539] Step 4: Analyzing the pitch

[1540] The server analyzes the pitch of the noise-removed audio data using FFT (Fast Fourier Transform). It detects the reference frequency of each note and compares it with the pitch written in the music score. The input is the noise-removed audio data, and the output is the result of the pitch analysis.

[1541] Step 5: Analyze the rhythm

[1542] After the pitch analysis is complete, the server performs rhythm analysis. It detects the start and end timing of notes on the time axis and evaluates their correspondence with the reference rhythm. The input is the FFT analysis result, and the output is the result of rhythm analysis.

[1543] Step 6: Dynamic analysis

[1544] The server then performs dynamic analysis, detecting volume fluctuations for each note in the audio data and analyzing whether the dynamic range is as specified in the score. The input is the rhythm analysis result, and the output is the dynamic analysis result.

[1545] Step 7: Expressive Analysis

[1546] The server uses a generative AI model to evaluate the emotional expressiveness of a performance. The model compares it with existing famous performance data to calculate the degree of match in expressiveness. The input is the dynamic analysis result, and the output is the result of expressiveness analysis.

[1547] Step 8: Generate evaluation results

[1548] The server generates an overall score based on the analysis results of pitch, rhythm, dynamics, and expressiveness. It then integrates the analysis results to create numerical scores and graphs for each element. It also lists suggestions for specific phrase practice to help the user improve. The input is all the analysis results, and the output is an overall evaluation result and practice suggestions.

[1549] Step 9: Submit and view the evaluation results

[1550] The server sends the assessment results to the device using HTTPS. The device analyzes the received assessment results and provides them to the user by visually displaying them in a user interface. The display includes graphs, numerical values, and text feedback. The input is the assessment results and practice suggestions, and the output is the feedback displayed in the user interface.

[1551] As a concrete example, a user plays "Twinkle Twinkle Little Star" on the piano and records it in step 1. The performance is uploaded to the server in step 2, and various analyses are performed in steps 3 to 7. An overall evaluation is created in step 8, and the results are displayed on the user's device in step 9. An example of a prompt sentence is "Analyze this performance data and evaluate its pitch, rhythm, dynamics, and expressiveness."

[1552] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1553] MODE FOR CARRYING OUT THE INVENTION

[1554] The present invention is a system for supporting the improvement of musical instrument performance. It not only records, analyzes, and evaluates a user's performance data, but also reflects the user's emotional state in the evaluation. This system allows users to objectively evaluate their own performance and find appropriate practice methods. Specific embodiments for implementing the present invention are described below.

[1555] Overview of program processing

[1556] 1. Receiving and recording audio input

[1557] When the user starts playing, the device launches a recording application and records the performance sound in high quality through the microphone. Recording is done in real time, and the audio data is saved in the device's memory. The audio waveform and recording time are displayed during recording.

[1558] 2. Sending audio data

[1559] When the performance is finished and the user stops recording, the device saves the recording in the specified audio format. When the user selects the "Send" button, the recording is uploaded to the server.

[1560] 3. Analysis of audio data

[1561] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it sequentially analyzes pitch, rhythm, dynamics, and expressiveness.

[1562] Pitch analysis: Uses FFT to identify the base frequency of a note and compares it to the pitch written in the sheet music.

[1563] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[1564] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[1565] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[1566] 4. Emotion Recognition by Emotion Engine

[1567] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The device uses a camera and microphone to capture the user's facial expressions and tone of voice, and analyzes them using the emotion engine.

[1568] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[1569] 5. Generating evaluation results

[1570] The server integrates the results of the voice data analysis and the emotion engine to generate a comprehensive evaluation. It calculates a score based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and creates visualization data.

[1571] The server also generates advice and messages to help maintain motivation that are tailored to the user's emotional state and adds them to the suggestion list.

[1572] 6. Submitting and displaying evaluation results

[1573] The server then sends the generated evaluation results to the device. The transmission is secure and done using HTTPS.

[1574] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also provided.

[1575] Specific examples

[1576] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[1577] The user starts playing the piano and the device records the audio.

[1578] After the performance is over, the device saves the recording data and uploads it to the server.

[1579] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[1580] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[1581] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1582] The server sends the evaluation results and suggestions to the device, which then displays them on the user interface, allowing the user to receive visual feedback and advice based on their emotional state.

[1583] In this way, the system of the present invention provides the user with objective performance evaluation as well as feedback according to their emotional state, supporting enjoyable and effective improvement in musical instrument playing.

[1584] The processing flow will be explained below.

[1585] MODE FOR CARRYING OUT THE INVENTION

[1586] Specific flow of program processing

[1587] Step 1:

[1588] When a user starts playing an instrument, the device launches a recording application and captures the performance through the microphone in high quality. Recording occurs in real time, and the audio data is stored in the device's memory. During recording, the audio waveform and recording time are displayed on the screen, allowing the user to check the progress.

[1589] Step 2:

[1590] When the user finishes playing and stops recording, the device saves the recorded audio data in the specified audio format (e.g., WAV, MP3). After the recording data has been saved, the user selects the "Send" button to send the recording data to the server.

[1591] Step 3:

[1592] The device uploads the recorded audio file to the server via an HTTP request in response to user input. The data is encoded and sent using a secure protocol (HTTPS). The progress of the transfer is displayed on the device, and the user is notified when the transfer is complete.

[1593] Step 4:

[1594] The server temporarily stores the received audio data in storage, then applies a noise reduction algorithm to remove unwanted noise, improving the accuracy of subsequent analysis.

[1595] Step 5:

[1596] The server begins analyzing the audio data, first performing pitch analysis. It uses FFT (Fast Fourier Transform) to detect the reference frequency of each note and compares it with the pitch written in the sheet music. The analysis results are recorded as numerical data.

[1597] Step 6:

[1598] The server then performs rhythm analysis, detecting the start and end times of notes on the time axis and comparing them with a reference rhythm template. The analysis results are recorded as data including timestamps.

[1599] Step 7:

[1600] The server also performs dynamic analysis, analyzing fluctuations in audio volume and assessing whether the dynamics of each note are in line with the musical notation. The results are recorded as a numerical value representing the sound pressure level.

[1601] Step 8:

[1602] Next, the server performs expressiveness analysis. Using a generative AI model, it evaluates the emotional expression of the performance. The AI ​​model compares the performance with existing famous performance data to calculate the degree of agreement in expressiveness, and records the analysis results as an emotional score.

[1603] Step 9:

[1604] The device captures the user's facial expressions and tone of voice during and after performance, and activates the emotion engine, which analyzes the user's facial expressions and tone of voice to evaluate the user's emotional state in real time.

[1605] Step 10:

[1606] The server combines the results of the voice data analysis with the emotion analysis results from the emotion engine to generate an overall evaluation. It also combines scores based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state to create visualization data. It also generates motivational advice and messages according to the user's emotional state.

[1607] Step 11:

[1608] The server then sends the generated evaluation results and motivation advice to the device. The transmission is securely performed using HTTPS. The evaluation results include scores for analyzed pitch, rhythm, dynamics, and expressiveness, as well as advice based on the user's emotional state.

[1609] Step 12:

[1610] The device analyzes the received evaluation results and displays them visually on the user interface. The feedback is visualized as graphs, numbers, and text, allowing users to check detailed evaluations and motivational advice. The device also has "replay" and "reevaluation" buttons, allowing users to play again and receive a new evaluation.

[1611] Specific examples

[1612] For example, if a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and the following process is performed.

[1613] The user starts playing and the terminal starts recording.

[1614] After the performance is finished, the recording data is saved on the device and sent to the server.

[1615] The server analyzes the received data and evaluates it for pitch, rhythm, strength, and expressiveness. At the same time, the device analyzes the user's facial expressions and voice and evaluates their emotional state using an emotion engine.

[1616] The server integrates the overall analysis results and generates an overall evaluation such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotional evaluation."

[1617] The server suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1618] The device displays the evaluation results and suggestions on a user interface, allowing the user to see visual feedback and motivational advice.

[1619] As described above, the system of the present invention provides the user with objective performance evaluation and feedback based on emotional state, enabling enjoyable and effective improvement in musical instrument playing.

[1620] Example 2

[1621] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1622] Conventional systems for supporting the improvement of musical instrument playing skills are specialized in evaluating the user's performance technique, but are unable to provide comprehensive feedback that takes into account the user's emotional state. This makes it difficult to support maintaining motivation or improving the emotional expressiveness of performance. Furthermore, conventional systems only evaluate performance technique criteria such as pitch, rhythm, and dynamics, and are limited in their assessment of expressiveness.

[1623] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness, means for capturing the user's facial expression and tone of voice and analyzing their emotional state, and means for integrating the emotional state into the analysis results and generating an overall evaluation. This makes it possible to provide comprehensive feedback that integrates the user's performance technique and emotional state, and to support maintaining motivation and improving the emotional expressiveness of performance.

[1624] A "user" is an individual who plays an instrument and inputs performance data into the system.

[1625] "Terminal" refers to an electronic device that a user uses to record musical instrument performances and display analysis results.

[1626] "Server" refers to a computer system that receives and analyzes voice data sent from a terminal, generates evaluation results, and sends them to the terminal.

[1627] "Performing" refers to the act of a user making sounds using an instrument.

[1628] "Audio data" refers to data that is a digital recording of music played by a user.

[1629] "Noise reduction" refers to the process of removing unnecessary noise from audio data.

[1630] "FFT" stands for Fast Fourier Transform, a transformation technique frequently used in signal processing.

[1631] A "generative AI model" is a model that uses artificial intelligence to generate and analyze data.

[1632] A "prompt sentence" is an input sentence given to a generative AI model.

[1633] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[1634] "Evaluation result" refers to comprehensive feedback generated based on analysis of speech data and emotional state.

[1635] "Visualized data" refers to data that displays the evaluation results in graphs, charts, etc. so that users can intuitively understand them.

[1636] "Advice" refers to specific guidelines suggested to help users improve their playing skills and motivation.

[1637] This invention is a system to support the improvement of musical instrument performance by recording, analyzing, and evaluating the user's performance data, and also reflecting the user's emotional state in the evaluation. This system allows the user to objectively evaluate their own performance and find appropriate practice methods.

[1638] Recording and sending audio

[1639] When a user starts playing an instrument, the device launches a recording application and records the performance through the microphone in high quality. The recorded audio data is stored in the device's memory and sent to a server as needed. During recording, the audio waveform and recording time are displayed, allowing the user to understand their performance in real time.

[1640] Analysis of audio data

[1641] The server stores the received audio data in storage and begins analysis. First, it performs noise reduction to remove unnecessary noise. Next, it performs the following analysis steps in sequence.

[1642] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of the note and compare it to the pitch written in the sheet music.

[1643] Rhythm analysis: Analyze the start and end timing of notes and compare them with a reference rhythm.

[1644] Dynamic analysis: Detects fluctuations in speech volume and assesses appropriate dynamics.

[1645] Expressiveness analysis: Using generative AI models to assess the emotional expression of a performance.

[1646] emotion recognition

[1647] The device runs an emotion engine that recognizes emotions from the user's facial expressions and tone of voice. The camera and microphone capture the user's facial expressions and tone of voice and analyze them using the emotion engine. The server integrates the analysis results of the emotion engine with the analysis results of the voice data and reflects the user's emotional state in an overall evaluation.

[1648] Generating and displaying evaluation results

[1649] The server combines the results of the voice data analysis with the emotion engine's results to generate an overall evaluation. A score is calculated based on pitch, rhythm, dynamics, expressiveness, and the user's emotional state, and visualization data is created. Furthermore, advice and messages to maintain motivation according to the user's emotional state are generated and added to a suggestion list. The server sends the generated evaluation results to the device, which then visually displays them on the user interface. The feedback is visualized as graphs, numbers, and text, allowing the user to check the detailed evaluation. Appropriate motivation advice is also displayed.

[1650] Examples of prompt statements

[1651] If a user plays "Twinkle Twinkle Little Star" on the piano, the performance is recorded and analyzed and evaluated using the following procedure.

[1652] The user starts playing the piano and the device records the audio.

[1653] After the performance is over, the device saves the recording data and uploads it to the server.

[1654] The server analyzes the received voice data and evaluates pitch, rhythm, strength, and expressiveness. At the same time, the device runs an emotion engine to analyze emotions based on the user's facial expressions and tone of voice.

[1655] The server combines the results of the voice analysis and emotion analysis to generate an overall evaluation, such as "good pitch, slightly unstable rhythm, appropriate dynamics, good expressiveness, and good emotion evaluation."

[1656] The server aggregates the analysis results and suggests to the user "phrases to improve rhythm" and "messages to increase motivation."

[1657] The server sends the evaluation results and suggestions to the terminal, which displays them on the user interface.

[1658] An example of a prompt for input to a generative AI model is as follows:

[1659] "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and score them on pitch, rhythm, dynamics, and expressiveness. Also, analyze the user's facial expressions and tone of voice to assess their emotional state and create comprehensive feedback."

[1660] In this way, the system of the present invention can provide the user with objective performance evaluation as well as feedback according to their emotional state, thereby supporting enjoyable and effective improvement in musical instrument playing.

[1661] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1662] Step 1: Receiving and Recording Audio Input

[1663] When the user starts playing an instrument, the device launches a recording application and records the performance in high quality through the microphone.

[1664] Input: User's playing sound, recording application

[1665] Specific operation: The recording application captures the audio input from the microphone and saves it as audio data in the device's memory. During recording, the audio waveform and recording time are displayed in real time.

[1666] Output: Recorded audio data

[1667] Step 2: Save and send audio data

[1668] When the user finishes playing and presses the "Send" button, the device stops recording and saves the recorded data.

[1669] Input: User click on "Submit" button, recorded voice data

[1670] Specific operation: Calls the StopRecording() method after recording is complete. The saved audio data is saved as a file in the specified audio format (e.g. WAV, MP3).

[1671] Output: Audio data saved in the specified format

[1672] The terminal uploads the audio data to the server.

[1673] Input: Saved audio data

[1674] Specific behavior: Generates an HTTP POST request to send the saved audio data to the server.

[1675] Output: Audio data uploaded to the server

[1676] Step 3: Analyzing the audio data

[1677] The server stores the received audio data in temporary storage and begins analysis.

[1678] Input: Uploaded audio data

[1679] Specific operation: Save the audio data as a temporary file and perform the following analysis steps.

[1680] Noise reduction: Removes unwanted noise from audio data.

[1681] Pitch analysis: Using FFT (Fast Fourier Transform) to identify the base frequency of a note and compare it with a reference value.

[1682] Rhythm analysis: Detects the start and end times of notes and compares them with a reference rhythm.

[1683] Dynamics analysis: Analyzes fluctuations in voice volume and evaluates the appropriateness of dynamics.

[1684] Expressiveness analysis: Using a generative AI model, the emotional expression of a performance is evaluated based on a prompt.

[1685] Prompt: "After a user plays 'Twinkle Twinkle Little Star' on the piano, analyze the recording and rate it on pitch, rhythm, dynamics, and expressiveness. Also, rate the user's emotional state by analyzing their facial expressions and tone of voice, and create overall feedback."

[1686] Output: Pitch analysis results, rhythm analysis results, dynamic analysis results, expressiveness analysis results

[1687] Step 4: Emotion Recognition

[1688] The device captures the user's facial expressions and tone of voice to activate the emotion engine.

[1689] Input: User's facial expression (camera image), tone of voice (audio data)

[1690] Specific operation: Uses a camera and microphone to capture facial expression data and voice tone data in real time, and analyzes the data using an emotion engine.

[1691] Output: Parsed emotion data

[1692] The server integrates the analysis results of the emotion engine with the analysis results of the voice data, and reflects the user's emotional state in the evaluation.

[1693] Input: Voice analysis results, emotion analysis results

[1694] Specific operation: The results of speech analysis and sentiment analysis are integrated, and an evaluation algorithm is applied to generate an overall evaluation.

[1695] Output: Integrated overall rating

[1696] Step 5: Generate evaluation results

[1697] The server integrates the results of the voice data analysis and the emotion engine results to generate an overall evaluation.

[1698] Input: Integrated overall rating

[1699] Specific actions: Based on the integrated overall evaluation, the system calculates a score based on pitch, rhythm, dynamics, expressiveness, and emotional state, creates visualization data, and generates advice and messages to maintain motivation and adds them to a suggestion list.

[1700] Output: Overall evaluation data, visualization data (graphs, charts, etc.), motivation advice

[1701] Step 6: Submit and view the results of the evaluation

[1702] The server transmits the generated evaluation results to the terminal.

[1703] Input: Overall evaluation data, Visualization data, Advice

[1704] Specific operation: Evaluation results, visualization data, and advice are securely sent to the device using HTTPS.

[1705] Output: Data sent to the terminal

[1706] The terminal receives the evaluation results and visually displays them on a user interface.

[1707] Input: Overall evaluation data, Visualization data, Advice

[1708] Specific operation: Analyzes the evaluation results and visualizes the feedback as graphs, numbers, and text. Users can check the detailed evaluation results and appropriate motivation advice.

[1709] Output: Visual feedback to the user

[1710] Through the above steps, the system of the present invention can provide the user with detailed performance evaluations and appropriate feedback, and support the user in improving their musical instrument playing in an enjoyable and effective way.

[1711] (Application example 2)

[1712] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1713] Current systems for supporting the improvement of musical instrument performance only provide objective feedback on the user's performance, and are unable to provide comprehensive evaluations or suggestions that take into account the user's emotional state. As a result, they are unable to suggest phrases or products that the user truly needs, resulting in a lack of motivation and personalized support.

[1714] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1715] In this invention, the server includes means for recording the sound of the instrument played by the user in high quality, means for transmitting the recorded sound data to the server, means for analyzing the sound data and evaluating pitch, rhythm, dynamics, and expressiveness, means for analyzing the user's emotions from their facial expressions and tone of voice, means for suggesting phrases and products necessary for improvement based on the analysis results and emotion analysis results, and means for transmitting the evaluation results to the user's terminal and visually displaying them, thereby enabling comprehensive and personalized feedback on the user's performance.

[1716] The "recording means" is a device or software for recording the sound played by the user in high quality.

[1717] The "transmission means" is a communication device or software for transmitting the recorded voice data to the server.

[1718] "Analysis means" refers to a device or software for analyzing audio data and evaluating pitch, rhythm, dynamics, and expressiveness.

[1719] "Emotion analysis means" refers to a device or software that analyzes a user's facial expression and tone of voice to evaluate their emotional state.

[1720] The "suggestion means" is a device or software that suggests phrases and products necessary for improvement based on the analysis results and emotion analysis results.

[1721] The "display means" is a device or software for transmitting the evaluation results to the user's terminal and visually displaying them.

[1722] The "frequency analysis means" is a device or software for analyzing the frequency of audio data in order to detect pitch.

[1723] "Rhythm analysis means" refers to a device or software for analyzing the start and end timing of notes in order to detect rhythm.

[1724] A "volume analysis means" is a device or software for analyzing variations in voice volume to detect strength and weakness.

[1725] A "generative AI model" is a machine learning model used to evaluate expressiveness.

[1726] A "prompt sentence" is an input sentence for generating product suggestions from the user's emotional state.

[1727] This invention is a virtual store shopping assistant system that makes personalized product recommendations taking into account the user's emotional state. This system records the sound of a musical instrument played by the user in high quality, analyzes the data, and analyzes the user's emotions from their facial expressions and tone of voice to provide comprehensive feedback.

[1728] Hardware and Software

[1729] This system is implemented using the following hardware and software:

[1730] Hardware:

[1731] Smartphone or smart glasses

[1732] Camera (for facial expression analysis)

[1733] Microphone (for recording audio)

[1734] software:

[1735] Python

[1736] OpenCV (for facial expression analysis)

[1737] TensorFlow (for facial expression analysis model)

[1738] Sounddevice (for audio recording)

[1739] Soundfile (for manipulating audio data)

[1740] Requests (for sending data)

[1741] Processing flow

[1742] Audio recording and analysis

[1743] The device records the user's voice when they ask a question or make a comment about a product. The recorded voice data is sent to a server, where it is analyzed, including frequency analysis. Specific analysis items include pitch, rhythm, dynamics, and expressiveness.

[1744] facial expression analysis

[1745] The device's built-in camera captures the user's facial expression data, which is then processed by a generative AI model using TensorFlow for emotion recognition. The resulting data is used to evaluate the user's level of interest or anxiety.

[1746] Suggestions and Displays

[1747] The server combines the results of voice analysis and facial expression analysis to generate prompts and make appropriate product suggestions. Based on the generated prompts, the system makes personalized suggestions to the user. The suggestions are displayed on the device.

[1748] Specific examples

[1749] If a user is browsing home decor products in a virtual shopping application and asks, "How big is this sofa?", the following occurs:

[1750] 1. Audio recording: The device's microphone records the audio.

[1751] 2. Audio analysis: Analyze pitch, rhythm, dynamics, and expressiveness from recorded data.

[1752] 3. Facial Expression Analysis: The camera captures the user's facial expressions and evaluates the user's emotional state.

[1753] 4. Proposal generation: Based on the results of voice analysis and facial expression analysis, a generative AI model is used to generate optimal product proposals for the user.

[1754] Prompt Sentence Examples

[1755] The prompt for "How big is this sofa?" is as follows:

[1756] Analyze the customer's interest and anxiety based on their questions while they are browsing products, and provide optimal product suggestions. Also, analyze their facial expressions to see their level of tension or interest, and provide explanations with kind words.

[1757] This allows users to receive personalized feedback based on their emotional state, resulting in a better shopping experience.

[1758] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1759] Step 1:

[1760] The device records the user's voice when they ask a question or make a comment about a product. Specifically, it captures high-quality audio data through a microphone and saves it in an audio format such as WAV. The recorded data is used as is in the next step.

[1761] Input: User's voice

[1762] Output: High-quality audio data (WAV format)

[1763] Step 2:

[1764] The device sends the recorded audio data to the server. The data is transferred using a secure communication protocol such as HTTPS. The server then stores the received audio data in its storage.

[1765] Input: High-quality audio data

[1766] Output: Upload audio data to the server

[1767] Step 3:

[1768] The server analyzes the transmitted audio data. The analysis items are pitch, rhythm, dynamics, and expressiveness. Specific processing includes frequency analysis using FFT, time-domain note analysis, detection of voice volume fluctuations, and evaluation of emotional expression using a generative AI model.

[1769] Input: High-quality audio data

[1770] Output: Evaluation results for pitch, rhythm, dynamics, and expressiveness

[1771] Step 4:

[1772] The device uses a camera to capture the user's facial expressions. The captured video data is input into a generative AI model using TensorFlow to recognize the user's emotional state from their facial expressions. Specifically, it extracts facial features and classifies emotions based on them.

[1773] Input: User facial expression video data

[1774] Output: User's emotional state (interest, anxiety, etc.)

[1775] Step 5:

[1776] The server combines the results of voice analysis and facial expression analysis to generate prompts, which are then used by a generative AI model to provide optimal product recommendations to the user.

[1777] Input: Pitch, rhythm, dynamics, expressiveness evaluation results, and the user's emotional state

[1778] Output: prompt statement

[1779] Step 6:

[1780] The server creates specific product suggestions based on the generated prompt sentences, and provides the suggestions to the user as personalized feedback.

[1781] Input: prompt statement

[1782] Output: Personalized product recommendations

[1783] Step 7:

[1784] The terminal receives the suggestions sent from the server and visually displays them on the user interface. Specifically, the suggestions are displayed as text, graphs, or images, providing visual feedback to the user.

[1785] Input: Proposal from the server

[1786] Output: Visual feedback on the user interface

[1787] These steps allow users to receive personalized feedback based on their emotional state and enjoy a better shopping experience.

[1788] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1789] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1790] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1791] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1792] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1793] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1794] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1795] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1796] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1797] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1798] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1799] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1800] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1801] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1802] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1803] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1804] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1805] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1806] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1807] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1808] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1809] The following is further disclosed regarding the above embodiment.

[1810] (Claim 1)

[1811] A means for recording the sound of a musical instrument played by a user in high quality;

[1812] means for transmitting the recorded voice data to a server;

[1813] means for analyzing the audio data and evaluating pitch, rhythm, dynamics, and expressiveness;

[1814] A method to suggest phrases necessary for improvement based on the analysis results, and

[1815] The evaluation results are sent to the user's terminal and displayed visually.

[1816] A system including:

[1817] (Claim 2)

[1818] a means for performing frequency analysis to detect pitch in analyzing audio data;

[1819] means for analyzing the start and end timing of notes to detect rhythm;

[1820] means for analyzing variations in speech volume to detect dynamics;

[1821] A means of using generative AI models to evaluate expressiveness and

[1822] The system of claim 1 further comprising:

[1823] (Claim 3)

[1824] The user can then replay the piece based on the analysis results, record new audio data, and reassess it.

[1825] The system of claim 1 further comprising:

[1826] "Example 1"

[1827] (Claim 1)

[1828] A means for recording with high accuracy the sound data played by the user;

[1829] means for transmitting the recorded acoustic data to a communication device;

[1830] means for analyzing the acoustic data to assess pitch, rhythm, volume changes, and emotional expression;

[1831] A method to suggest the necessary practice content for improvement based on the analysis results, and

[1832] A means for transmitting the evaluation results to the user's device and visually displaying them.

[1833] A system including:

[1834] (Claim 2)

[1835] a means for performing frequency analysis to detect a musical scale in analyzing acoustic data;

[1836] means for analyzing the start and end timing of sounds to detect rhythm;

[1837] means for analyzing fluctuations in the volume data to detect changes in volume;

[1838] Using generative AI models to evaluate emotional expressions

[1839] The system of claim 1 further comprising:

[1840] (Claim 3)

[1841] The user can then replay the piece based on the analysis results, record new acoustic data, and re-evaluate it.

[1842] The system of claim 1 further comprising:

[1843] "Application Example 1"

[1844] (Claim 1)

[1845] A means for recording the sound of a musical instrument played by a user in high quality;

[1846] means for transmitting the recorded voice data to a server;

[1847] means for analyzing the audio data and evaluating pitch, rhythm, dynamics, and expressiveness;

[1848] A method to suggest phrases necessary for improvement based on the analysis results, and

[1849] means for transmitting the evaluation results to a user's terminal and visually displaying them;

[1850] A means for displaying the evaluation results on a user interface so that the user can check detailed feedback;

[1851] A system including:

[1852] (Claim 2)

[1853] a means for performing frequency analysis to detect pitch in analyzing audio data;

[1854] means for analyzing the start and end timing of notes to detect rhythm;

[1855] means for analyzing variations in speech volume to detect dynamics;

[1856] a means for using a generative AI model to evaluate expressiveness;

[1857] a means for performing noise reduction;

[1858] 10. The system of claim 1.

[1859] (Claim 3)

[1860] A means for the user to replay the performance based on the analysis results, record new audio data, and re-evaluate the performance;

[1861] A method for providing a comprehensive score based on the evaluation of pitch, rhythm, dynamics, and expressiveness, and

[1862] 10. The system of claim 1.

[1863] "Example 2: Combining Emotion Engines"

[1864] (Claim 1)

[1865] A means for recording the sound of a musical instrument played by a user in high quality;

[1866] means for transmitting the recorded voice data to a server;

[1867] means for analyzing the audio data and evaluating pitch, rhythm, dynamics, and expressiveness;

[1868] A method to suggest phrases necessary for improvement based on the analysis results, and

[1869] means for transmitting the evaluation results to a user's terminal and visually displaying them;

[1870] A means for capturing a user's facial expressions and tone of voice to analyze their emotional state;

[1871] means for integrating the emotional state into the analysis results to generate an overall assessment;

[1872] A system including:

[1873] (Claim 2)

[1874] a means for performing frequency analysis to detect pitch in analyzing audio data;

[1875] means for analyzing the start and end timing of notes to detect rhythm;

[1876] means for analyzing variations in speech volume to detect dynamics;

[1877] a means for using a generative AI model to evaluate expressiveness;

[1878] means for using an emotion engine to analyze an emotional state of a user;

[1879] The system of claim 1 further comprising:

[1880] (Claim 3)

[1881] A means for the user to replay the performance based on the analysis results, record new audio data, and re-evaluate the performance;

[1882] The system of claim 1 further comprising:

[1883] "Application example 2 when combining emotion engines"

[1884] (Claim 1)

[1885] A means for recording the sound of a musical instrument played by a user in high quality;

[1886] means for transmitting the recorded voice data to a server;

[1887] means for analyzing the audio data and evaluating pitch, rhythm, dynamics, and expressiveness;

[1888] A means of analyzing emotions from the user's facial e...

Claims

1. A means for recording the sound of a musical instrument played by a user in high quality; means for transmitting the recorded voice data to a server; means for analyzing the audio data and evaluating pitch, rhythm, dynamics, and expressiveness; A method to suggest phrases necessary for improvement based on the analysis results, and The evaluation results are sent to the user's terminal and displayed visually. A system including:

2. a means for performing frequency analysis to detect pitch in analyzing audio data; means for analyzing the start and end timing of notes to detect rhythm; means for analyzing variations in speech volume to detect dynamics; A means of using generative AI models to evaluate expressiveness and The system of claim 1 further comprising:

3. The user can then replay the piece based on the analysis results, record new audio data, and reassess it. The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A