Server terminal, playback terminal, program, data processing method, and musical instrument practice system for practicing musical instruments at home
The musical instrument practice system addresses the lack of instructor feedback by capturing and editing video and audio data from lessons, enabling learners to correct mistakes and improve their skills through targeted feedback.
Patent Information
- Application Number
- JP2024011370
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-01-29
AI Technical Summary
Existing musical instrument practice systems fail to provide effective feedback on errors made by learners during practice, as they lack instructor feedback and do not address the causes of errors in the learner's performance.
A musical instrument practice system comprising a video acquisition terminal, server terminal, and playback terminal that captures and edits video and audio data from lessons, allowing learners to practice with instructor feedback by playing back corrected performances and instructions.
Enables learners to receive targeted feedback on their performance, correcting mistakes made during practice and improving their playing skills by referring to instructor feedback even when practicing alone.
Smart Images

Figure 2025116751000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a server terminal, a playback terminal, a program, a data processing method, and a musical instrument practice system for practicing a musical instrument at home, which detects incorrect playing parts when a practicer is practicing a musical instrument at home and provides re-instruction to the practicer using edited video containing the instructor's comments. [Background technology]
[0002] In recent years, the number of people practicing musical instruments has been increasing, and these students are now attending music schools and music classes to learn how to play their instruments. Typically, instructors at music schools and music classes set practice assignments before each lesson, and students then play and practice pieces at home in response to the assignments. Students then show their practice results to their instructors in the next lesson. At this time, the instructors can identify problems with the students' performance and provide appropriate guidance accordingly. However, when students try to review their performance at home after a lesson, various factors (e.g., not fully understanding the instructor's comments during class, low motivation to practice at home, forgetting what they were taught) can affect their ability to practice efficiently.
[0003] Therefore, to solve the above problems, it is essential to develop a home practice and support system for learners. Patent Document 1 proposes a pronunciation practice system that automatically arranges audio data in accordance with the tempo of background music by utilizing the human characteristic of spontaneous behavior. The pronunciation practice system disclosed in Patent Document 1 also provides a pronunciation practice function based on background music, signal sounds, model pronunciations, and images corresponding to the model pronunciations. Learners practice pronunciation in an environment where background music is constantly playing. In one embodiment, the pronunciation practice system stores multiple model pronunciations and videos or images corresponding to the model pronunciations as learning content. The learner inputs pronunciation practice instructions into an input device, and the control unit reads and arranges the stored audio data (background music, signal sounds, and model pronunciations). The control unit also plays a signal sound (referred to as an audio signal for warning the learner) over the beat of the background music for each audio data to be heard by the learner. The control unit of the pronunciation practice system repeatedly plays the model pronunciations and signals while also displaying a distributed image corresponding to the model pronunciation on a display. At this time, the learner may feel the rhythm of the background music and synchronize with its tempo, while listening to the model pronunciation and effectively predicting the timing of overlapping by the cue sound.
[0004] Patent Document 2 discloses a system that supports a learner when practicing a piece of music. In this practice support system, musical score data is stored in the system and the musical score is displayed on a display terminal. The practice support system also determines the performance position corresponding to the musical score based on the learner's performance data, and can play music from any specified position. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2021-67844 [Patent Document 2] Japanese Patent Application Laid-Open No. 2013-200455 Summary of the Invention [Problem to be solved by the invention]
[0006] According to the above Patent Documents 1 and 2, while a learner is playing an instrument, the learner's own performance data or standard performance data (i.e., model performance data used as a model for practice, or performance data used by an instructor for instruction) is played back and presented, but there is no mention of the causes of errors that may occur while the learner is playing, or how to correct errors once they are discovered. Furthermore, the support systems described in the prior art documents have the learner actively practice by themselves in response to the model, which poses the problem that the learner is unable to receive feedback from the instructor.
[0007] The present disclosure has been made in consideration of these problems, and aims to provide a musical instrument practice system that allows a practicer to repeatedly practice at home using video footage of lessons. [Means for solving the problem]
[0008] One embodiment of the present invention is a musical instrument practice system including a video acquisition terminal that acquires practice video, a server terminal that edits the practice video, and a playback terminal that plays back the edited practice video received from the server terminal, wherein the video acquisition terminal includes a first image capture unit that captures images of a practicer playing and their associated audio data, and images of an instructor instructing and their associated audio data, a first storage unit that stores the images captured by the first image capture unit and the acquired audio data, and a first transmission unit that transmits the video, including the images and audio data stored in the first storage unit, to the server terminal. The server terminal comprises a first receiving unit that receives the video transmitted from the first transmitting unit, an editing unit that automatically edits the video received by the first receiving unit according to the sound contained in the video, and a second transmitting unit that transmits the video edited by the editing unit to the playback terminal, and the playback terminal comprises a second receiving unit that receives the edited video, a second storage unit that stores the video received from the second receiving unit, a recording unit that records the sound being practiced by the practicer, and a playback unit that plays back the sound of the instructor that corresponds to the sound being practiced by the practicer, based on the sound recorded by the recording unit and the edited video.
[0009] In the above-mentioned musical instrument practice system, the editing unit may be characterized by extracting the sounds of the practicer and the instructor's performance and speech from the video received by the first receiving unit, each as multiple pieces of performance data and speech data, and automatically clustering the performance data according to the similarity of the sounds made by the practicer during performance. In the above-mentioned musical instrument practice system, the editing unit may further be characterized by having a second memory unit that allocates utterance data based on the time period of the video to corresponding clusters obtained by clustering the performance data and stores them as a performance practice set. In the above musical instrument practice system, the playback terminal may include a third storage unit that stores sample sounds uploaded by the practicer, and may be characterized in that the sounds of the instructor's speech are replaced according to the performance practice set based on the sample sounds stored in the third storage unit. In the above-mentioned musical instrument practice system, the playback terminal may include a setting unit that sets the overall practice time based on the performance practice set, and the setting unit may further be characterized in that it sets the practice time individually depending on the performance proficiency of the learner. In the above-mentioned musical instrument practice system, the playback terminal may be characterized by having a second image capture unit that captures images of the practicer playing and the sound thereof, and a detection unit that detects whether the practicer is playing correctly based on the actions of playing the musical instrument and the sound thereof, based on the video captured by the second image capture unit. In the above-mentioned musical instrument practice system, the playback terminal may be characterized in that, when the detection unit detects an incorrectly played portion, it estimates the similarity between the video captured by the second image capture unit and the performance practice set, and if the similarity is equal to or greater than a predetermined threshold, it plays back the performance practice set. The above-described musical instrument practice system may be characterized in that, before practicing a performance, the learner refers to a performance practice set, and specifies and plays back the performance practice set.
[0010] The server terminal includes a first receiving unit that receives video of a lesson, an editing unit that automatically edits the video received by the first receiving unit according to the sound contained in the video, and a second transmitting unit that transmits the video edited by the editing unit to a playback terminal.
[0011] This is a data processing method that includes executing a first receiving step of receiving video of a lesson, an editing step of automatically editing the video received from the first receiving step in accordance with the sound contained in the video, and a second transmitting step of transmitting the video edited by the editing step to a playback terminal.
[0012] This is a program for realizing a first receiving function that receives video of a lesson, an editing function that automatically edits the video received by the first receiving function according to the sound contained in the video, and a second transmitting function that transmits the video edited by the editing function to a playback terminal.
[0013] This playback terminal includes a second receiving unit that receives edited lesson video, a second storage unit that stores the video received from the second receiving unit, a recording unit that records the sound that the learner is practicing, and a playback unit that plays back the instructor's sound that corresponds to the sound that the learner is practicing based on the sound recorded by the recording unit and the edited video.
[0014] This is a data processing method that includes executing a second receiving step of receiving edited lesson video, a second storing step of storing the video received from the second receiving step, a recording step of recording the sound being practiced by the learner, and a playback step of playing back the instructor's sound that corresponds to the sound being practiced by the learner based on the sound recorded by the recording unit step and the edited video.
[0015] This is a program for realizing a second receiving function that receives edited lesson video, a second storage function that stores the video received from the second receiving function, a recording function that records the sounds that the learner is practicing, and a playback function that plays back the instructor's sounds that correspond to the sounds that the learner is practicing, based on the sounds recorded by the recording function and the edited video. [Effects of the Invention]
[0016] According to the instrument practice system of the present invention, the state of a learner receiving instruction from an instructor is photographed and recorded. When the learner practices an instrument by himself, the instrument practice system records the sound, and if the learner makes a mistake similar to one made in class, the system can play back the content of the instruction given by the instructor during class, so that the learner can recall the content of the instruction (including praise and scolding) given by the instructor even when practicing alone. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a schematic diagram illustrating a configuration of a practice assistance system according to an embodiment of the present disclosure. [Figure 2] 2 is a block diagram showing the configuration and functions of the video acquisition terminal of FIG. 1. FIG. [Figure 3] 2 is a block diagram showing the configuration and functions of the server terminal of FIG. 1. FIG. [Figure 4] FIG. 10 is a diagram showing a portion of possible video data used in the practice support system in an embodiment of the present disclosure. [Figure 5] 4 is a flowchart showing an example of automatic editing of data in the server terminal of FIG. 3. [Figure 6] 2 is a flowchart showing an example of re-editing the video shown in FIG. 1 and obtaining a performance practice set. [Figure 7] 2 is a block diagram showing the configuration and functions of the playback terminal of FIG. 1. FIG. [Figure 8] 8 is an image showing an example of image replacement in the setting section of FIG. 7. [Figure 9] 8 is a flowchart showing an example of setting a practice time in the playback terminal of FIG. 7. [Figure 10] 8 is a flowchart showing an example of a process for extracting similar performance content and playing back the points pointed out by the instructor during the lesson in association with the practice video on the playback terminal of FIG. 7. [Figure 11]8 is a flowchart showing an example of presentation of the atmosphere of a practice piece on the playback terminal of FIG. 7. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, the same or equivalent components, processes, and signals shown in each drawing will be assigned the same reference numerals, and redundant explanations will be omitted where appropriate. Also, in each drawing, some components that are not important for explanation will be omitted.
[0019] In the following description, "photographing" may mean simultaneously capturing image data showing the appearance and movements of a subject and audio data of the subject using a camera and microphone. "Video" captured by photography in this embodiment may include both image data and audio data. "Capture" refers to capturing only image data showing the appearance and movements of a subject via a camera. "Recording" refers to recording the audio of a subject. "Sound source" may include speech data of a learner, speech data of an instructor, performance data of a learner (instrument sound), and performance data of an instructor (instrument sound). Meanwhile, "musical instrument" in this embodiment is an instrument used to play music, and may include not only keyboard instruments but also wind instruments, string instruments, percussion instruments, etc.
[0020] In addition, in order to record image data and audio data according to this embodiment, the camera and microphone may be installed separately. Specifically, the camera according to this embodiment is configured to acquire image data of an image of the practice, and the microphone is configured to acquire audio data of the audio of the practice and the audio during instruction. However, the functions of the camera and microphone are not limited to those of the above embodiment, and although the camera and microphone are described separately in this specification, audio data may be collected by a microphone attached to the camera.
[0021] <Summary> As shown in FIG. 1, the musical instrument practice system according to this embodiment comprises a video capture terminal 100, a server terminal 200, and a playback terminal 300, all of which are communicatively connected via a network NW. In the musical instrument practice system, the video capture terminal 100 acquires practice video (hereinafter referred to as video data) including audio and video data, transmits the video data to the server terminal 200, and the server terminal 200 automatically edits the video data. The server terminal 200 then transmits the edited video data to the playback terminal 300. Based on the received edited video data, the playback terminal 300 can detect whether the learner's performance matches the musical score when practicing at home, or whether the learner is following the instructor's instructions in class. If the learner performs incorrectly, the playback terminal 300 plays the corresponding edited video. "Incorrect performance" includes performance that does not follow the notes or beat of the music, or performance that does not follow the instructor's instructions, such as dynamics, hand shape, fingering, or how the fingers play the keys. In addition, the video data of the instructor can be replaced with a virtual character, and the image and voice of the virtual character can be used. In this case, it is expected that the trainee will become more interested in the practice and will be more motivated to practice. The instrument practice system may also have a function for determining the movements of the practicer or the way the instrument is played. The video capture terminal 100 acquires image data of the practicer and stores it in the second storage unit 220 of the server terminal 200. When the practicer re-practices the instrument at home, it becomes possible to determine based on the image data whether the practicer is performing the movements of the instrument correctly in order to correct incorrect playing techniques in accordance with the instructor's instructions.
[0022] 1, the musical instrument practice system in this embodiment comprises a video acquisition terminal 100, a server terminal 200, and a playback terminal 300, and each function may be realized through data sharing on a network NW. The configuration of the musical instrument practice system is not limited to the above, and the video acquired from the video acquisition terminal 100 may be uploaded directly to the server terminal 200 without going through the network NW (for example, the video data acquired by the video acquisition terminal 100 may be stored in a flash memory, and the flash memory may be connected to the server terminal 200 for uploading).
[0023] Furthermore, the video data acquired by the video acquisition terminal 100 may be divided into (1) performance data of the learner in class (sound of the learner's instrument), (2) image data of the learner in class, (3) performance data of the instructor in class (sound of the instructor's instrument), (4) image data of the instructor in class, (5) speech data of the learner in class, and (6) speech data of the instructor in class. In addition to acquiring the above six types of data, if there are multiple learners, the video acquisition terminal 100 may also acquire voice data (including the instrument sounds of the attendees and the questions of the attendees) and image data of the attendees' accompanying learners, as well as speech data of the parents who are present at the lesson.
[0024] In this embodiment, the server terminal 200 analyzes the audio and image data based on the video data acquired from the video acquisition terminal 100, dividing the data into three groups: (i) the instrumental sounds of the learner and the instructor, (ii) the questioning voice of the learner and the instruction voice of the instructor, and (iii) the movements of the learner and the movements of the instructor. The learner's instrument sound is the sound of the instrument when the learner is playing. The instructor's instrument sound is the sound of the instrument that the instructor is playing for instruction during a lesson. The learner's question voice is the voice of the learner asking the instructor about a problem they are having while playing. The instructor's instruction voice is the voice made by the instructor during a lesson (including voices pointing out and voices giving praise). The learner's actions are the actions of the learner playing the instrument. The instructor's actions are the actions of the instructor playing the instrument to instruct the learner.
[0025] The learner's instrument sound and the learner's movement, and the instructor's instrument sound and the instructor's movement may be extracted as a set depending on time. For example, after acquiring the video data, the server terminal 200 receives a set of video data including the learner's instrument sound and the learner's movement (or the instructor's instrument sound and the instructor's movement at the time of the instructor). In addition, the server terminal 200 transmits the audio signals (practice participant's question voice, instructor's instruction voice, practice participant's instrument sound and instructor's instrument sound) recorded via the microphone of the video acquisition terminal 100 and the image information (practice participant's actions and instructor's actions) captured via the camera of the video acquisition terminal 100 to the playback terminal 300 separately, and it is not necessary to combine them into one set in the video acquisition terminal 100.
[0026] The playback terminal 300 mainly functions to allow the learner to review the content of the lesson at home. Each piece of data compiled in the server terminal 200 serves as detection material for determining whether the learner is playing correctly when practicing at home, and also as training data for the learner practicing at home. In other words, if the learner is playing the instrument correctly at home (i.e., no instruction from the instructor is required), there is no need to play back video containing the instructor's instruction. If the learner plays incorrectly and repeats the mistakes made in class, the playback terminal 300 can detect the part where the performance is incorrect and plays back video showing the instruction given when the learner played incorrectly during practice with the instructor.
[0027] The playback terminal 300 according to this embodiment may be equipped with the automatic editing function described above, or may further edit the image data and audio data played back by the playback terminal 300. Specifically, the playback terminal 300 may replace each image data and audio data acquired from the server terminal 200 with a new image and voice tone acquired from an external terminal or the Internet, without changing the audio content. As an example, the original instructor image and instructor voice may be replaced with an animated character image and voice without changing the audio content. In this case, the trainee can relieve stress during practice, and if the trainee uses a voice and character that appeals to them, improvement in the trainee's practice efficiency can be expected.
[0028] The instrument practice system provides a function that allows the learner to re-learn and review parts that they were unable to understand during class. Furthermore, when practicing an instrument at home, the learner can efficiently identify common practice problems based on the data obtained from the class and receive re-guidance. The following describes in detail the mechanisms of the various structures that realize the instrument practice system and home practice function according to this embodiment.
[0029] <Configuration> <Video Acquisition Terminal 100> FIG. 2 is a block diagram showing an example of the configuration of a video capture terminal 100 that captures the voice data of a learner, the image data of the learner, the voice data of the instructor, and the image data of the instructor. As shown in FIG. 2, the video capture terminal 100 has a RAM (Random Access Memory, not shown), a ROM (Read Only Memory, not shown), a first image capture unit 110, a first storage unit 120, and a first transmission unit 130. The video capture terminal 100 is a computer system that includes a processor and memory, and the processor executes a program to realize a function of capturing images of a performance and transmitting the images to a server terminal 200. In order to capture the voice data and image data of the learner and the instructor, the first image capture unit 110 is further provided with a right-facing camera 111, a left-facing camera 112, a foot-facing camera 113, a downward-facing camera 114, a whole-view capturing camera 115, and a microphone 116. Each component is connected to each other so as to be able to communicate with each other via the Internet.
[0030] The first image capturing unit 110 has a function of capturing images of the learner playing and their audio data, and images of the instructor teaching and their audio data. The first image capturing unit 110 is only required to capture images of the lesson, and is composed of at least one camera and microphone. In this embodiment, the instrument is described as a piano, but the instrument is not limited to a piano and may be various other instruments, such as a guitar, ukulele, violin, or saxophone. When the instrument is a piano, the first image capturing unit 110 may include a right-facing camera 111, a left-facing camera 112, a foot-facing camera 113, and a downward-facing camera 114. The piano keyboard's right-facing camera 111 is installed on the left side of the practicer (the left side of the piano keyboard) when the practicer is facing the keyboard so that he or she can play the piano, so that it can capture images of the right side (the keyboard side). The piano keyboard's left-facing camera 112 is installed on the practicer's right side (the right side of the piano keyboard) when the practicer is facing the keyboard so that it can capture images of the left side (the keyboard side). The piano keyboard's foot-facing camera 113 is installed below the practicer (the piano pedals) when the practicer is facing the keyboard so that it can capture images above. The piano keyboard's downward-facing camera 114 is installed above the practicer (the piano's top roof) when the practicer is facing the keyboard so that it can capture images below. By filming the way the learner plays, it is possible to identify mistakes that occur when playing the piano. In addition, by providing a foot-facing camera 113 that films how the learner presses the pedals and a downward-facing camera 114 that films from above the keyboard, it is possible to comprehensively obtain image data of the learner using the piano.
[0031] Meanwhile, right-facing camera 111, left-facing camera 112, foot-facing camera 113, and downward-facing camera 114 can not only acquire image data of the learner, but can also acquire image data of the learner playing the piano for the instructor to provide instruction when the learner plays the piano without following the music score. In other words, it is possible to acquire image data of the person playing the musical instrument.
[0032] The first photographing unit 110 may also include an overall imaging camera 115 that captures an image of the entire practitioner and the instructor. That is, while the practitioner is taking a lesson, image data of the practitioner and the instructor can all be captured from other angles. In this case, the overall imaging camera 115 may be able to supplement image data of angles of view that cannot be captured using the right-facing camera 111, left-facing camera 112, foot-facing camera 113, and downward-facing camera 114.
[0033] The number and positions of the cameras provided in the first image capturing unit 110 may be changed depending on the type of instrument, the practitioner, etc. Furthermore, depending on the type and brand of camera, the camera may have not only a photographing function but also a sound recording function. In this case, the microphone 116 does not need to be installed. Furthermore, a camera for photographing the instructor may be provided in the first image capturing unit 110. This camera may have a function for automatically tracking and photographing the instructor so that the instructor is captured within a predetermined range in the captured image.
[0034] Image data acquired from each camera mounted on first photographing unit 110 and audio data recorded by microphone 116 are stored in an information storage database of first storage unit 120. The audio data and image data may be sent to first storage unit 120 as a set, or may be sent to first storage unit 120 separately.
[0035] The first storage unit 120 stores programs and data required for the operation of the video acquisition terminal 100. The first storage unit 120 may be configured mechanically from a ROM (Read Only Memory) or a RAM (Random Access Memory). The first storage unit 120 may store each piece of data using a cloud storage service.
[0036] The audio data and image data are stored in the information storage database of the first storage unit 120, and then transmitted to the server terminal 200 via the first transmission unit 130. The transmission method for each type of data is not limited to one, and the audio data and image data may be transmitted together to the server terminal 200 (i.e., the audio data acquired from the microphone and the image data acquired from the camera are transmitted as two types of data without being integrated), or the audio data and image data may be transmitted together to the server terminal 200 while corresponding to each other in chronological order (i.e., the audio data and image data are transmitted as a set, as one type of data).
[0037] <Server terminal 200> 3, the server terminal 200 is a cloud server and is made up of a first receiving unit 210, a second storage unit 220, an editing unit 230, and a second transmitting unit 240. The configuration is not limited to the above, and the editing unit 230 may be subdivided into an audio data editing unit, an image data editing unit, and an integration unit.
[0038] The first receiving unit 210 installed in the server terminal 200 has a function of receiving video data including audio data and image data from the video acquisition terminal 100 as a communication path for performing processing related to data exchange.
[0039] Furthermore, the first receiving unit 210 temporarily stores each piece of data received from the first transmitting unit 130 in the second storage unit 220 of the server terminal 200, and outputs it to the editing unit 230. The configuration of the first receiving unit 210 is not limited to this, and the received video data may be output directly to the editing unit 230 without being temporarily stored in the second storage unit 220.
[0040] The editing unit 230 receives the data from the first receiving unit 210, and then divides the received data into features and uses, and performs individual or integrated analysis.
[0041] Specifically, in this embodiment, the data used for editing is video data. As shown in Figure 4, the video data is divided into audio data and image data based on the audio, actions, and the time at which each audio or action was recorded or photographed, and each audio data and image data is accompanied by a timestamp. Next, each data may be further classified according to its intended use. For example, image data is divided into performer image data showing the actions of a performer, instructor image data showing the actions of an instructor, and lesson image data showing lesson images (i.e., images including both the performer and the instructor).
[0042] Furthermore, since the audio data includes both human voices (i.e., including performers and instructors) and instrument sounds, the audio data is divided into performance data and speech data, and the server terminal 200 can link the data to each piece of editing image data using a timestamp. In this case, the performance data and speech data may be categorized in more detail. For example, the performance data may be divided into performance data of the learner and performance data of the instructor, and the speech data may be divided into speech data of the learner and speech data of the instructor.
[0043] When receiving video data from the video acquisition terminal 100, the editing unit 230 can classify the video data into audio data and image data. As an example, the editing unit 230 selects and extracts the portion of the video data in which audio is recorded, the so-called sound track. After the sound track is extracted, the remaining portion of the video data becomes image data. In this case, the image data may be linked to the audio data using a timestamp.
[0044] <Sound source separation> The editing unit 230 can distinguish who is speaking by separating the sound sources according to the speaker's frequency. Specifically, the sound of an instrument or the voice of a learner or instructor each has its own characteristics, and each piece of audio data also represents its own signal (hereinafter referred to as "audio signal"). The editing unit 230 may separate multiple sound sources according to frequency, etc., based on these audio signals. Furthermore, the audio signals in this embodiment may include a learner speech signal, an instructor speech signal, a learner performance signal, and an instructor performance signal.
[0045] The frequencies of the trainee speech signal and the instructor speech signal may be determined based on factors such as the function of the human body (e.g., gender, age, height, weight, etc.) and the space (e.g., recording room, soundproof room, etc.). For example, if the trainee is a child, the trainee's voice tone is usually higher than that of an adult instructor, so the frequency is relatively higher and therefore higher frequency than the instructor's voice tone. Conversely, the voice tone of an instructor who is an adult is usually lower than that of a trainee who is a child, so the frequency is relatively lower and therefore is shown as a low frequency.
[0046] Since the same instrument is used in the performances by the student and the instructor, the frequency of the sound produced by both is essentially the same. Therefore, it is not possible to distinguish between the performance of the student and the performance of the instructor based on the frequency of the sound alone. In this case, it is possible to identify who is performing by also referencing the image data.
[0047] Specifically, the sound source separation method relates to a technology that uses data such as changes in wavelength and amplitude due to frequency, as well as time data from image data. That is, in this sound source separation method, the editing unit 230 uses wavelength and amplitude due to frequency to separate a mixed audio signal (i.e., an audio signal acquired via a microphone, which typically includes an audio signal of the learner's speech, an audio signal of the instructor's speech, an audio signal of the learner's performance, and an audio signal of the instructor's performance) into learner's speech data, instructor's speech data, and performance data. Then, the editing unit 230 uses timestamps to identify the learner's speech data, instructor's speech data, and performance data as image data, and determines who is performing.
[0048] 5 is a flowchart showing the classification of video data and the process of combining the classified data. Before the editing unit 230 extracts specific audio data from the mixed audio signal to be separated, the method for separating a plurality of sound waves according to this embodiment includes the steps of: A1, the editing department 230 divides the mixed audio signal according to the frequency of the mixed audio signal, and specifies a time period (the real time corresponding to the time of shooting) for the divided audio section (the audio section is a section in which sound exists in the recorded audio signal, specifically, a section in which there is a fluctuation in the amplitude of the frequency indicating audio), and assigns a time stamp corresponding to this time period (step S101); A2, the editing department 230 divides the image data into a plurality of video sections according to time periods corresponding to the audio sections and assigns time stamps to the video sections (step S102); A3, the editing unit 230 matches the audio section and the image section using the time stamp (step S103). If the time when each sound or each action was recorded or photographed is the same, the audio section belongs to the image section, and the audio and the image sections are treated as one set (hereinafter referred to as edited video). A4, the editing department 230 classifies the learner speech data, the instructor speech data, and the performance data included in the edited video according to the physical characteristics of the audio signal (step S104); A5 further includes the editing unit 230 determining who the performer is by detecting the performance actions of the learner and the instructor based on the image section (step S105).
[0049] In the aforementioned step A1, the editing unit 230 can extract target audio data from a mixed audio signal containing multiple sound sources using the features of the audio data. Specifically, audio signals are expressed by frequency fluctuations over time, and sound sources can basically be separated by the frequency band of the sound. In addition, for speech data, the intonation of the voice and the tempo (speed) of the conversation can also be separated by the timing of frequency vibrations along the time axis. As an example, the waveform of an audio signal from a musical instrument being played is different from that of an audio signal from a human speaking, so if the audio signal changes significantly, it is considered that the sound source has changed.
[0050] Furthermore, the editing unit 230 divides the mixed audio signal into a plurality of audio segments based on frequency, wavelength, and amplitude, and identifies a time period corresponding to each audio segment in the real time (hereinafter referred to as time series) recorded during shooting, and sets a timestamp for each audio segment. Specifically, the editing unit 230 divides the mixed audio signal into a plurality of audio segments based on frequency, thereby generating audio segments including any of the speech of the learner, the speech of the instructor, and the audio of the performance. The editing unit 230 sets the time of the audio segment before division as the time of each divided audio segment. For example, if the audio segment before division is located in the time period from 10 minutes to 11 minutes on the original time axis, after division, the editing unit 230 sets the timestamp for the corresponding time period (from 10 minutes to 11 minutes) as the time of the audio segment after division.
[0051] In step A2 described above, the editing unit 230 may divide the image data into a plurality of image segments according to time periods corresponding to the audio segments. As an example, the editing unit 230 extracts images corresponding to a time period (for example, from 10 minutes to 11 minutes) corresponding to the audio segments from the image data, and sets the extracted images as image segments. The editing unit 230 then simultaneously assigns a time stamp corresponding to the time period to the image segments.
[0052] In step A3 described above, the editing unit 230 matches audio segments and image segments using timestamps. If the timestamps are the same, the audio segment belongs to the image segment, and so they may be treated as one set (hereinafter referred to as edited video). If the timestamps do not match, the audio segment and its image segment are separate segments and are not treated as one set. Furthermore, because the image segments are divided according to the time zone of the audio segment, the final data set (i.e., a data set including audio data and image data matched by timestamps) always includes one audio segment and one image segment.
[0053] In step A4 described above, the editing unit 230 determines whether the audio data is human speech data or performance data based on the physical features of the audio signal. The edited video, which has been matched using timestamps, includes learner speech data, instructor speech data, learner performance data, and instructor performance data. While it is possible to distinguish between the learner speech data and the instructor speech data based on the characteristics of the subjects (e.g., age and gender), the learner performance data and the instructor performance data have the same frequency because the same instrument is used, making direct separation difficult.
[0054] Generally, it is possible to determine whether someone is speaking in edited video based on motion information such as human mouth movement. As an example, the editing unit 230 has a function for detecting human mouth movement from edited video, and when mouth movement is detected, the person making the sound in the edited video is considered to be the person whose mouth is moving, making it possible to determine who is making the sound. If sound is included as audio data and no person whose mouth is moving is detected in the corresponding image data, the editing unit 230 determines that the sound is the sound of a musical instrument.
[0055] However, although it is possible to determine who is making the sound by following the mouth movements, there is a risk that the mouth movements cannot be detected if the trainees and instructors in the video data (or image data) are wearing masks or the like to cover their faces, or if the video loaded on the video acquisition terminal 100 has blind spots.
[0056] In this case, the editing unit 230 may identify the sound source based on the physical features of each audio signal. As an example, the editing unit 230 classifies the instructor's speech data, the learner's speech data, and the performance data by frequency. Because the frequency of human audio signals differs from the frequency of musical instrument audio signals, there are bands where the audio signals do not overlap. Therefore, the editing unit 230 may set the frequency of the non-overlapping band as the reference threshold value. Specifically, if the frequency of an audio signal in the edited video is higher than a certain threshold, the audio signal in the edited video may be considered performance data, and if the frequency of an audio signal in the edited video is lower than a certain threshold, the audio signal in the edited video may be considered instructor's speech data or learner's speech data.
[0057] However, due to individual differences between people, the frequency bands of a child learner and an adult instructor may overlap, and instrument performance data may also overlap with human voice data (for example, the frequency band of an adult's normal speaking voice is 150Hz to 500Hz, the frequency band of a child's speaking voice is approximately 1000Hz to 2000Hz, and the frequency band of a piano is approximately 27Hz to 4186Hz, so the frequency bands of instruments overlap with the human frequency band). In this case, it may not be possible to separate sound sources based on frequency alone.
[0058] Therefore, after performing the first audio separation using frequency, the editing unit 230 introduces a second audio separation technique that identifies and quantifies the waveforms (wavelength, amplitude, etc.) of the audio sources. The higher the tone of the audio data, the higher the amplitude of the waveforms of the audio sources contained in the mixed audio signal, but the shorter the wavelength. Conversely, the lower the tone of the audio data, the lower the amplitude and the longer the wavelength, making it possible to determine whose voice it is. However, the physical feature amounts used for audio separation in this embodiment are not limited to these.
[0059] As another method for distinguishing between the voice of a learner, the voice of an instructor, and the sound of an instrument, the editing unit 230 may use a voice recognition model to distinguish between the voices of a learner, the voice of an instructor, and the sound of an instrument. The voice recognition model is a model that has learned the voices of a learner, the voice of an instructor, and the sound of an instrument. That is, the voice recognition model is a learning model that has learned teacher data in which voice data of a learner's voice is annotated with information indicating that the learner is a learner, teacher data in which voice data of an instructor's voice is annotated with information indicating that the instructor is a trainer, and teacher data in which voice data of an instrument's voice is annotated with information indicating that the instrument is a trainee. The model accepts input voice data and estimates whether the voice data is the voice of a learner, the voice of an instructor, or the voice of an instrument. By inputting speech data into the voice recognition model, it is possible to identify who is making the voice, allowing the editing unit 230 to classify the voice data more efficiently.
[0060] In step A5 described above, after identifying the speech data and performance data between the learner and instructor, the editing unit 230 must further separate the performance data. Unlike speech data, performance data is data acquired while playing the same instrument, and it may not be possible to determine who played the performance data using frequency.
[0061] As one method for identifying performance data according to this embodiment, the editing unit 230 may use acquired information about the movements of the instrument being played in the edited video to identify the subject of the performance data. Specifically, the editing unit 230 detects the person playing the instrument in each edited video, and the person is the performer in that edited video. Furthermore, if the performer in the edited video is a practicing musician and there is no speech data, that edited video may be deleted because it does not contain any instructor's comments. Furthermore, the detection target is not limited to the above. Performers may also be identified by detecting the movements of the hands playing the instrument or the amount of force applied to the instrument. In other words, the more physical features that can represent the instructor's performance data, the easier it is to identify the performer and the more accurate the identification.
[0062] <Processing each data after separation> Next, the processing of each data after audio separation according to this embodiment will be described with reference to Fig. 6. Based on the edited video obtained according to the physical feature amounts of each data, the editing unit 230 extracts edited video having only instructor utterance data, edited video having only instructor performance data, and edited video having both instructor utterance data and instructor performance (hereinafter referred to as non-silence sections) (step S201). Edited video that is not extracted may be treated as a silent space and deleted.
[0063] In order to determine whether the extracted non-silence sections are similar to each other, in this embodiment, each non-silence section is analyzed as a "language of performance," and the similarity of each non-silence section is measured using a similarity calculation method for natural language processing. The "language of performance" refers to information in which sound is expressed in a form that can be expressed in language, and more specifically, it refers to information in which the pitch, duration, importance of each note (i.e., the number of times each note appears), etc. of each note are expressed as vectors after the sound is converted into a mel spectrogram. A mel spectrogram is information that expresses the amplitude of sound on a time axis and a frequency axis in the mel scale.
[0064] Specifically, in this embodiment, the audio data of each non-silence section may be treated like a natural language. The editing unit 230 can generate a mel spectrogram based on each non-silence section. A mel spectrogram is a spectrogram whose frequency axis is in the mel scale and is typically used in speech recognition. The editing unit 230 creates a feature vector representing the "language of performance" in the non-silence section from the generated mel spectrogram. The feature vector may include the pitch of the voice (hereinafter referred to as "pitch"), the continuity of the voice (i.e., the longest vocalization duration), etc.
[0065] As an example of generating a feature vector, if there are five pieces of audio data in one non-silence section, the editing unit 230 treats each of the five pieces of audio data as independent data and creates a feature vector based on audio characteristics such as the pitch of each piece of audio data or the continuity of the audio. The method for creating the feature vector is not limited, and the editing unit 230 may calculate the pitch difference between adjacent pieces of audio data based on the order of performance to create a feature vector. In a non-silence section, the order of performance is a1, a2, and a3, and if a feature vector is to be created based on the pitch difference, the editing unit 230 can calculate the pitch difference using the calculation methods of (a2-a1) and (a3-a2). The editing unit 230 may then use the calculation result as the feature vector ((a2-a1), (a3-a2)) for that non-silence section.
[0066] The editing unit 230 can then use the created feature vectors to calculate the relative similarity between the non-silence segments. The method for calculating the similarity in this embodiment is not limited. Each method will be described below.
[0067] <Similarity calculation method 1> Specifically, the editing unit 230 calculates the cosine similarity between each non-quiet section based on the feature vector of each quiet section, and determines the similarity between the non-quiet sections. As an example, the editing unit 230 generates a mel spectrogram from a non-silence section A having five pieces of audio data, and then creates a feature vector A[a1, a2, a3, a4, a5] based on the pitch of the audio. The editing unit 230 also generates a mel spectrogram from a non-silence section B having five pieces of audio data, and then creates a feature vector B[b1, b2, b3, b4, b5] based on the pitch of the audio. The editing unit 230 then calculates the cosine similarity between the non-silence section A and the non-silence section B based on the feature vector A and the feature vector B.
[0068]
number
[0069] Feature vector A = (a1, a2, , a n ), feature vector A∈R n Feature vector B=(b1,b2,...,b n ), feature vector B∈R n Feature vector A and feature vector B form a certain angle, and the editing unit 230 can calculate the relative similarity between non-quiet section A and non-quiet section B based on the cosine value of the angle between feature vector A and feature vector B. Generally, when the value indicating cos(a, b) is close to 1, the relative similarity between non-quiet section A and non-quiet section B is considered to be high, and when it is close to -1, the relative similarity between non-quiet section A and non-quiet section B is considered to be low.
[0070] <Similarity calculation method 2> Specifically, the editing unit 230 converts the feature vector of each quiet section into a TF-IDF value (short for term frequency-inverse document frequency), calculates the cosine similarity between each non-quiet section based on the TF-IDF value, and determines the similarity between the non-quiet sections. Among these, the TF value (short for Term Frequency) is an index representing the frequency of occurrence of a word, and the IDF value (short for Inverse Document Frequency) is an index representing the inverse document frequency.
[0071]
number
[0072] In Equation 2 of this embodiment, tf(t, d) is the TF value of certain audio data t in the non-silence section d. t,d is the number of times that the audio data t appears in the non-quiet section d. (S∈d) n (S,d) is the sum of the occurrence counts of all sounds recognized as different sounds in the non-quiet section d.
[0073]
number
[0074] In Equation 3 of this embodiment, idf(t) is an index that indicates the importance of each piece of audio data relative to the non-quiet section, and is called the inverse document frequency. N is the number of pieces of audio data in the non-quiet section, and df(t) is the number of times that audio data t appears. As an example of similarity calculation method 2, the editing unit 230 generates a mel spectrogram from a non-silence section A containing five pieces of audio data, and then creates a feature vector A[a1, a2, a3, a4, a5] based on the pitch of the audio. The editing unit 230 also generates a mel spectrogram from a non-silence section B containing five pieces of audio data, and then creates a feature vector B[b1, b2, b3, b4, b5] based on the pitch of the audio. Next, the editing unit 230 calculates a TF-IDF value for each pitch of feature vector A, and also calculates a TF-IDF value for each pitch of feature vector B. Then, it calculates cosine similarity between the vectors with the TF-IDF values for each pitch. Similarity calculation method 2, which calculates cosine similarity using TF-IDF values, has higher accuracy than similarity calculation method 1. This is because the same musical score often contains multiple sections with similar rhythms due to modulation (e.g., transposition or modulation). In this case, similarity calculation method 1 may conversely determine that portions with similar rhythms are dissimilar. As a result, if the cosine similarity based on the TF-IDF value is close to 1, the relative similarity between non-quiet section A and non-quiet section B is considered to be high, and if it is close to -1, the relative similarity between non-quiet section A and non-quiet section B is considered to be low.
[0075] Furthermore, a threshold value for the similarity between each non-silence section may be freely set according to the feature vector. The set threshold value is stored in the second storage unit 220. If the threshold value is higher than the set threshold, the similarity between multiple non-silence sections is high, and the multiple non-silence sections may be clustered into one group (cluster). Conversely, if the threshold value is lower than the set threshold, the multiple non-silence sections may be considered to be different, and the multiple non-silence sections may be clustered into separate groups.
[0076] After calculating the similarity of each non-silence section, the editing unit 230 clusters the non-silence sections according to a threshold value (step S202). The non-silence sections may not be perfectly identical due to external factors. For example, the similarity calculated by the feature vectors will differ depending on whether a performer (including a learner and an instructor) plays the same musical score quickly or slowly (or modulates the performance). Therefore, by comparing the similarities of multiple input performance audio signals, if the deviation between the similarities is within a certain range, the multiple non-silence sections can be determined to be similar. However, the threshold value for determining similarity need not be a fixed value, but may be a relative value based on the similarity of each non-silence section.
[0077] The editing unit 230 may calculate the deviation between the similarities of each non-silence section according to the physical feature of the audio, or may calculate the deviation between the similarities using an externally connected audio similarity determination program.
[0078] The editing unit 230 ranks the clustered non-silence sections along a time axis (step S203). Specifically, the editing unit 230 ranks the non-silence sections according to the time at which each sound or action was recorded or photographed. Even in the case of overlapping performances (for example, when there are multiple identical sound sections in a musical score, repeatedly playing those sections is called overlapping performance), the editing unit 230 can identify the ranking of the non-silence sections using the timestamps.
[0079] Furthermore, the editing unit 230 may combine edited video having only instructor utterance data and edited video having only instructor performance data based on the time when each sound or each action was recorded or photographed. Specifically, the editing unit 230 uses time stamps to compare edited video having only instructor utterance data with edited video having only instructor performance data, and if there is a time correlation, the edited video having only instructor data is combined into one set with the edited video having only instructor performance data. For example, if the editing unit 230 has three edited videos (edited video A (10 minutes 10 seconds to 10 minutes 20 seconds) containing only instructor utterance data, edited video B (10 minutes 20 seconds to 10 minutes 40 seconds) containing only instructor performance data, and edited video C (15 minutes 24 seconds to 15 minutes 30 seconds) containing both instructor utterance data and instructor performance data), the editing unit 230 can compare the time information of edited video A and the second edited video B using timestamps, and since the compared edited videos A and B have a relationship on the time axis (i.e., time continuity), they can be integrated into one set (hereinafter referred to as the matched video). In this case, edited video C, which was not compared, already contains instructor utterance data and instructor performance data, so it can be stored as an independent edited video as review material for the student without being compared. The final edited footage and the collated footage are recorded as a performance practice set. The performance practice set is information that associates the learner's performance data for a predetermined period of time where the learner is playing incorrectly in the practice video with the instruction content (instructor's speech data and image data) of the instructor's instruction when the learner plays in that manner (step S204).
[0080] The instructor's scolding and praise included in the performance practice set may be summarized. As an example of summarizing, if there is multiple instructor utterance data for the same part, the editing unit 230 may organize the multiple instructor utterance data in the performance practice set that includes that part and combine them into a single performance practice set. Also, if there are multiple performance practice sets, the editing unit 230 may organize the instructor utterance data and combine them into a single performance practice set. As another example of summarization, if the editing unit 230 detects speech data with the same meaning in multiple performance practice sets at the same point, for example, if the instructor has speech data such as "Play more slowly!" and "Slow down here!", the editing unit 230 may use a speech recognition method, such as a morphological analysis of speech, or connect to an external speech recognition app to detect adjectives and verbs in the two pieces of speech data. Specifically, the editing unit 230 uses morphological analysis to extract the verb "play" and the adjective "slowly" from the two pieces of speech data. Then, if the editing unit 230 detects the adjective "slowly" multiple times and determines that the two pieces of speech data have the same meaning, it may combine the two pieces of speech data into one piece, "Play slowly!". Conversely, if the editing unit 230 does not detect speech data with the same meaning, i.e., if the same verb or adjective is not detected, the editing unit 230 may extract the multiple pieces of speech data and list them together in a single performance practice set.
[0081] The editing unit 230 may also convert the instructor's speech data in the performance practice set into text to organize it. For example, the editing unit 230 extracts the instructor's speech data from the performance practice set, adds punctuation, and converts it into text (i.e., speech-to-text conversion). The editing unit 230 may then use morphological analysis to identify frequently appearing keywords in the text and group phrases found therein into a single group. While the student is practicing, they can refer to the compiled results to review the points pointed out during the lesson.
[0082] <Playback terminal 300> As shown in FIG. 7, the playback terminal 300 includes a reception processing unit 310 including a second receiving unit 311, a third storage unit 312, and a setting unit 313, and a playback processing unit 320 including a second image capturing unit 321, a fourth storage unit 324, a detection unit 325, and a playback unit 326. The playback terminal 300 has the functions of re-editing, detecting, and playing each musical performance practice set received from the server terminal 200. Specifically, the playback terminal 300 includes a processor and memory. The setting unit 313 may correspond to the processor, and the third storage unit 312 and the fourth storage unit 324 may correspond to memory. The setting unit 313, which functions as a processor, may replace the image data and audio data in the musical performance practice set with different image data and audio data based on the musical performance practice set stored in the third storage unit 312. In this case, the obtained video data may be called post-setting video data.
[0083] On the other hand, when the practicer is practicing an instrument at home, the second image capturing unit 321 installed in the playback terminal 300 may record video of the practice (hereinafter referred to as home practice video data). The home practice video data acquired by the second image capturing unit 321 is stored in the fourth storage unit 324 as detection data. If the practicer makes a mistake while practicing at home, the detection unit 325 may detect the mistake from the fourth storage unit 324 and cause the playback unit 326 to play the video data after setting.
[0084] The receiving and processing unit 310 according to this embodiment can reprocess each piece of data received from the second transmitting unit 240. Specifically, the setting unit 313 may replace the instructor's image and voice that appear in the musical performance practice set with a different image and voice, based on the musical performance practice set received by the second receiving unit 311. The replacement image or GIF (Graphics Interchange Format) image and voice may be stored in advance in an information storage database in the third storage unit 312.
[0085] As shown in Figure 8, each figure appearing in the image (including a figure 801 of a learner and a figure 802 of an instructor; i.e., a part of the image) may be changed to a replacement image 804 (as shown in the upper right figure (1)). The video played as a performance practice set may include the instructor in class, the learners themselves, instruments, etc., but this video may be changed so that only the image of the character shown in the replacement image 804 (as shown in the lower right figure (2)) is an image of the instructor speaking the content of the instruction. The data relating to the replacement image 804 may be data stored in the third storage unit 312, or the practicer may upload a preferred replacement image 804 to the playback terminal 300. When the practicer uploads an image, the receiving processing unit 310 may be provided with an image control unit for changing the size, color, etc. of the image. A replacement image is an image displayed in place of the instructor, and changing the instructor to the replacement image 804 means that when the playback terminal 300 plays back the instruction content, it performs image processing to change the image of the instructor to the replacement image 804 specified by the learner. By changing the replacement image 804 to an anime or game character, a celebrity, a virtual idol, or the like that the learner likes, the musical instrument practice system (playback terminal 300) can increase the learner's motivation to practice.
[0086] Meanwhile, the playback terminal 300 can also reset the instructor's speech. Specifically, the reception processing unit 310 may connect to an external audio processing device, audio processing application, audio processing program, etc., and change the speech data in the performance practice set to replacement audio (hereinafter, "sample sound") according to the frequency. Specifically, the reception processing unit 310 inputs the sample sound uploaded by the learner based on the speech data in the performance practice set into a sound quality modification model such as ChatGPT (registered trademark) or S0-VITS-SVC-V4 (registered trademark) connected to the setting unit 313. The setting unit 313 trains the sound quality modification model and outputs the learned speech data. Furthermore, the more the playback terminal 300 collects the instructor's speech data, the higher the accuracy of reproducing the learned speech data.
[0087] Alternatively, the speech data in the performance practice set received from the second receiving unit 311 may be read by extracting words and phrases contained in the speech data and using the audio stored in the third storage unit 312. As an example, the words and phrases extracted from the speech data may be introduced into a voice changer app and played back in a different voice (for example, replacing the original instructor's voice with a child's voice). The above image and audio replacement method can reduce the tension felt by practitioners when practicing at home and increase their passion during practice.
[0088] As shown in Figure 9, when a practicer is practicing playing an instrument at home, he or she may set various indicators for practice according to his or her own needs (for example, setting practice time, selecting songs to practice, playing lesson videos, etc.). Specifically, before practicing at home, the student can play back the video of the lesson and can set in advance the music to be practiced from now on (step 301). The total practice time the learner intends to practice can then be set via the terminal (step 302). For example, the learner can freely input their practice time using a smartphone app. Alternatively, the playback terminal 300 can set the learner's practice time through multiple voice conversations with the learner. This setting can be performed by the learner inputting a number indicating the practice time, or it can be achieved through an automatic response function implemented by the playback terminal 300. The automatic response function can be designed fuzzy and programmed to negotiate the practice time with the learner if there is a discrepancy between the time the learner intends to set and the desired practice time. In this case, the playback terminal 300 can be equipped with a microphone 323 and connected to an external voice recognition system (e.g., ChatGPT (registered trademark)) to recognize the learner's voice. The learner can then freely select and play the song they wish to practice.
[0089] After the learner has decided on the piece of music he or she has selected, the playback terminal 300 presents a performance practice set corresponding to the piece of music, allowing the learner to reconfirm the points pointed out in the lesson before practicing at home (step 303).
[0090] Alternatively, the learner may set the difficulty level of each practice piece based on his or her own subjective judgment (step 304). Specifically, the playback terminal 300 may set each piece in order of difficulty, from highest to lowest: S>A>B>C>D, based on the learner's own level of proficiency. For a highly difficult piece, the speed may be slower, the practice time may be relatively longer, and the learner may play the instrument with one hand. Conversely, for a less difficult piece, the speed may be the same as the standard performance data, the practice time may be relatively shorter, and the learner may play with both hands. However, the method of setting the difficulty level is not limited to this, and the receiving processor 310 may also determine the difficulty level of a piece based on the number of pointed out points in the performance practice set.
[0091] Since the time required for each piece of music varies depending on the difficulty level, the practicer sets a final practice plan according to his or her needs (e.g., which pieces he or she is not good at practicing) (step 305). After the practice plan is determined, the practicer may begin practicing at home.
[0092] The playback processing unit 320 according to this embodiment can detect errors from the video of the practicer practicing at home, and can simultaneously play back the performance practice set. Taking the piano as an example, the second image capturing unit 321 captures the way the learner presses the black keys and white keys and the pedals, while the microphone 323 captures the performance voice of the learner. The obtained video data including the performance data and voice data is then stored in an information storage database of the fourth storage unit 324.
[0093] Furthermore, the playback terminal 300 uses the musical performance practice sets received from the server terminal 200 as materials for home practice. Before passing the musical performance practice sets to the playback processing unit 320, the playback terminal 300 determines via the setting unit 313 whether or not to replace the images and sounds in the musical performance practice sets. If the learner intends to reconfigure the images and sounds in each musical performance practice set (i.e., to replace the instructor image in the musical performance practice set with another character, or to replace the instructor's audio information with someone else's audio information without changing the content of the instructions), the playback terminal 300 may pass the configured musical performance practice set to the detection unit 325 and use it as practice material. If the learner does not configure a musical performance practice set, the playback terminal 300 may pass the musical performance practice set received from the server terminal 200 directly to the playback processing unit 320 and use it as practice material.
[0094] When determining the correctness of the learner's movements, the detection unit 325 uses a performance practice set as evidence data and compares it with the image data of the learner captured by the second image capture unit 321. As an example, the detection unit 325 may determine the image data based on the movement of the learner's fingers or fluctuations in the keys or strings of the instrument. On the other hand, while the sound produced by an instrument varies depending on the force applied to the keys or strings when playing an instrument, the detection unit 325 according to this embodiment can detect whether the position of the learner's movements is correct, without measuring the applied force (for example, if the finger position is correct, the movements are considered correct. However, without being limited to this, the correctness of the performance may also be determined by detecting movements such as the hand shape, fingering, and how the fingers play the keys). Alternatively, the detection unit 325 may measure the sound pressure due to the force applied to the keys or strings and detect the correctness of the movements based on the volume of the collected sound.
[0095] When determining the accuracy of the learner's performance, the detection unit 325 can use the performance practice set as basis data and compare it with performance data obtained from the microphone 323 installed in the playback terminal 300. The detection unit 325 can also calculate the similarity between the performance data during practice at home and the performance data included in the performance practice set, and compare it with a predetermined threshold.
[0096] If the difference is higher than the predetermined threshold, it is determined that the home practice video data matches the performance data in the performance practice set, and it is assumed that the student has made the same mistake during class again. In this case, the playback terminal 300 plays back the speech data and the corresponding image data included in the performance practice set.
[0097] Conversely, if the difference is lower than the predetermined threshold, it is determined that the home practice video data does not match the performance data included in the performance practice set. In this case, a re-evaluation may be performed to determine whether the practicer performed correctly or made another error. One method for re-evaluating this is to load standard performance data into the playback terminal 300, compare the home practice video data with the standard performance data, and use the similarity to determine whether the practicer performed correctly. If the home practice video data and standard performance data are similar and the performance is correct, the playback unit 326 does not need to play back the performance practice set. If the home practice video data and the standard performance data are not similar, and the so-called practicer makes a different error, the reproducing unit 326 cannot provide any indication of points to be pointed out because there is no speech data from the instructor, and it may be possible to reproduce only the parts that do not match.
[0098] In short, when the detection unit 325 determines the performance data that the practicer is practicing at home, it first compares it with the performance data included in the performance practice set, and if the home practice video data and the performance practice set are similar to each other by a predetermined amount or more, it may play the speech data and image data included in the performance practice set. Next, if the home practice video data and the performance practice set do not match, the detection unit 325 introduces standard performance data and further compares it with the home practice video data. If the standard performance data matches, the performance data may be skipped. Finally, if the home practice video data and the standard performance data do not match, the audio data of the mismatched portion may be reproduced.
[0099] FIG. 10 is a flowchart showing the process of the playback terminal 300 when a learner practices a musical instrument at home in this embodiment. The playback terminal 300 detects the position where the learner is performing (the position where the learner is seen performing in the video captured by the second image capturing unit 321) (step 401). While the practicer is performing at home, the second image capturing unit 321 installed in the playback terminal 300 simultaneously captures home practice video data of the practicer practicing at home (step 402). At this time, the second image capturing unit 321 may capture both audio and video, or only audio or only video. The playback terminal 300 calculates the similarity for each piece of home practice video data by comparing it with the musical performance practice set and the standard performance data, and presents a list of information indicating the musical performance practice sets in order of high to low similarity to the practice user (step 403). In this case, the playback terminal 300 may present a list of multiple musical performance practice sets ranked in order of high to low similarity, or may present only the musical performance practice set with the highest similarity. Next, the playback terminal 300 determines whether the learner is having a conversation with the playback terminal 300 (step S404). The learner may input voice into the microphone 323 mounted on the playback terminal 300. The voice may be a question or an inquiry related to the practice. At this time, if the playback terminal 300 detects a frequency fluctuation, it may identify that the learner is speaking. If the playback terminal 300 can detect a voice signal, it may consider that the learner is speaking and proceed to the next step (YES in step S404). If the playback terminal 300 does not detect a frequency change and does not identify the learner's speech (NO in step 404), the learner has understood that part of the practice piece and may proceed to the next stage, and the playback terminal 300 determines whether the learner wants to continue practicing (step 407). If the learner is still trying to practice (YES in step 407), the process returns to the step where the learner is playing (step 401). If the learner does not want to continue practicing (NO in step 407), the current instrument practice ends (step 408). Conversely, if the playback terminal 300 detects frequency fluctuations and identifies what is called a learner's speech (YES in step 404), it determines whether the learner is having difficulty with the list of information showing the performance practice set, sorted by similarity (step 405). If the learner is not having difficulty with the video being played and understands it well (NO in step 405), it may proceed to the next stage. It determines whether the learner will continue practicing (step 407). If the learner is still trying to practice (YES in step 407), it returns to the step where the learner is playing (step 401). Conversely, if the learner is not yet trying to practice (NO in step 407), it ends the current instrument practice (step 408). On the other hand, if the learner has difficulty in determining the order based on similarity (YES in step 405), the playback terminal 300 may allow the learner to play the performance practice set that corresponds to that time based on similarity. As an example, the learner may specify and play the performance practice set themselves, or the playback terminal 300 may automatically play the performance practice set with the highest similarity.
[0100] The embodiment for determining whether or not the learner is having difficulty may be realized by morphological analysis, face recognition, or by adding functionality to the terminal.
[0101] <Example 1: Morphological analysis method> As an example, the playback terminal 300 can connect to an external terminal (usually a voice analysis application) to compare the performance data of the practicer at home with the standard performance data and then display the degree of similarity. The practice performance set with the highest degree of similarity may be displayed immediately before the practicer asks a question. If the practicer has questions about the practice performance set, they may ask questions or negations such as "why?", "how?", or "how?" and input them as speech data into the playback terminal 300. At this time, the playback terminal 300 divides the utterance data input by the learner into subjects, predicates, and interrogatives, and compares them with keywords stored in the third storage unit 312, the fourth storage unit 324, or a newly installed storage unit. For example, when the learner utters, "How do you play like this?", the playback terminal 300 divides the utterance into "how to play" as the subject, "to do" as the predicate, and "how...?" as the interrogative using morphological analysis. If the units representing the divided natural language, i.e., parts of the subjects and interrogatives, match or highly match the preset criteria, the playback unit 326 will play the performance practice set with the highest similarity. The playback terminal 300 may detect only interrogative words or negative words without detecting subjects and predicates. For example, if the learner directly utters "Why?" or "I don't understand!", the utterance does not contain a subject or predicate, and the playback terminal 300 cannot detect keywords. In this case, the playback terminal 300 determines whether the interrogative words and negative words match those stored in the storage unit. If there is a perfect match or a high degree of match, the playback unit 326 may directly play the performance practice set with the highest similarity from the list showing the order of similarity.
[0102] <Example 2: Face Recognition> For example, if a learner has questions about the performance practice set presented in order of similarity, they may show a gloomy or questioning expression, such as frowning, closing their eyes, or pursing their lips, which indicates dissatisfaction or doubt. The playback terminal 300 compares the facial expression captured by the overall image capturing camera 322 of the second image capturing unit 321. If an incongruity or unnaturalness is detected, the playback terminal 300 may play the musical performance practice set with the highest similarity.
[0103] <Example 3: Adding a function such as a button to freely select review points on the device> As an example, a selection button may be added to the playback terminal 300. The learner may select the musical performance practice set that he or she most wants to review from among the musical performance practice sets presented in order of similarity.
[0104] In addition, the system for practicing a musical instrument at home according to this embodiment may be connected to an external system such as ChatGPT (registered trademark) to realize a voice conversation function. Also, the practicer may play a specific performance practice set through the external system while having a conversation with the external system.
[0105] As shown in Figure 11, the instrument practice system of this embodiment is not limited to determining whether or not a performance is being performed correctly, but may also be able to introduce the atmosphere and background of the practice piece in one example to the practice person. Before practicing an instrument, the learner may check the practice piece in advance (step S501). If the learner practices an instrument at a specified time, the playback unit 326 of the instrument practice system according to this embodiment may automatically distribute information about the piece to the learner at the specified time. If the learner specifies the practice time taking into consideration his or her own performance proficiency, the learner can select a piece stored in the playback terminal 300.
[0106] After the etude is selected, it is checked whether there is an image that shows the atmosphere of the selected etude (step 502). As an example, the playback terminal 300 according to this embodiment may connect to ChatGPT (registered trademark) (YES in step 502), so that the learner can select the etude and search for atmospheric data related to the etude online via ChatGPT (registered trademark) (step 503). If the atmosphere of the etude is not detected using ChatGPT (registered trademark) (NO in step 502), a commentary that shows the atmosphere of the piece may be compiled from the speech data of the instructor of the performance practice set that belongs to the etude stored in the third storage unit 312 or the fourth storage unit 324 (step 504).
[0107] <Supplementary information> Although the musical instrument practice system according to the present invention has been described in the above embodiment, it goes without saying that the present invention is not limited to the above embodiment. Various modifications will be described below.
[0108] (1) In the above embodiment, the editing unit 230 of the server terminal 200 may be installed in the playback terminal 300. As an example, after the first image capturing unit 110 and microphone 116 installed in the video capturing terminal 100 capture the video and audio of the practicer, they may not store the video and audio in the first storage unit 120 and second storage unit 220, but may instead pass the video and audio directly to the editing unit 230 via the Internet. Furthermore, the video capturing terminal 100 may be configured integrally with the server terminal 200. In this case, the integrated terminal may automatically edit the captured audio data and image data as the lesson progresses.
[0109] (2) Furthermore, the embodiment of acquiring video including images and audio of the learner during the lesson is not limited to this, and the video may be acquired via a smartphone, tablet, or the like.
[0110] (3) In the embodiment, the description has been given using frequency fluctuations and nT vibrations per unit time, but this is not limiting. For example, the playback terminal 300 may convert performance data into text as it receives it, or may recognize notes including the title of the piece and convert the performance data into text using machine learning. In this case, the playback terminal 300 may compare performance data obtained when the learner is performing in class with performance data obtained when the learner is performing at home when converting the data into text. If the converted text results differ, there is a possibility that the part was played incorrectly, and the playback terminal 300 may detect this and play it back.
[0111] (4) In the example according to this embodiment, the playback terminal 300 may search for and play back videos of performances by famous musicians, commentary from music experts, and the like.
[0112] (5) In addition, in the functions and control procedures of each component described in this specification, and particularly in the processing procedures described using flowcharts, the processing methods and procedures may omit some parts, or new parts may be added, or procedures may be substituted or changed in order, and such omissions, additions, or changes in order are also included in the scope of this disclosure as long as they do not deviate from the spirit of this disclosure. [Explanation of symbols]
[0113] 100 Video acquisition terminal 110 First Filming Section 111 Right-facing camera 112 Left-facing camera 113 Foot Camera 114 Downward Camera 115 Overall Imaging Camera 116 Mike 120 1st memory section 130 First Transmission Unit 200 Server terminal 210 First receiving unit 220 2nd memory section 230 Editorial Department 240 Second Transmission Unit 300 playback devices 310 Receiving processing unit 311 Second Receiving Unit 312 Third memory section 313 Settings 320 Reproduction Processing Unit 321 2nd Filming Department 322 Overall Imaging Camera 323 Mike 324 4th memory section 325 Detector 326 Playback Department 800 images 801 Practitioner Images 802 Leader Images 803 Audio Data 804 Replacement Image
Claims
1. A musical instrument practice system including a video acquisition terminal that acquires practice videos, a server terminal that edits the practice videos, and a playback terminal that plays back the edited practice videos received from the server terminal, The video acquisition terminal a first photographing unit that photographs an image of a learner playing and its audio data, and an image of an instructor teaching and its audio data; a first storage unit that stores the image captured by the first imaging unit and the acquired audio data; a first transmission unit that transmits the video including the image and audio data stored in the first storage unit to a server terminal; The server terminal a first receiving unit that receives the video transmitted from the first transmitting unit; an editing unit that automatically edits the video received by the first receiving unit in accordance with the sound included in the video; a second transmission unit that transmits the video edited by the editing unit to the playback terminal; The playback terminal a second receiving unit that receives the edited video; a second storage unit that stores the video received from the second receiving unit; a recording unit for recording the sound being practiced by the person practicing; A musical instrument practice system comprising: a playback unit that plays back the instructor's sound corresponding to the sound being practiced by the learner based on the sound recorded by the recording unit and the edited video.
2. The editorial department From the video received by the first receiving unit, extracting the sounds of the learner and the sounds of the instructor's performance and speech for each of a plurality of pieces of performance data and speech data; The performance data is automatically clustered according to the similarity of the sounds being played by the practicer.
2. The musical instrument practice system according to claim 1.
3. The editorial department further: The speech data based on the time period of the video, a second storage unit that assigns corresponding utterances to the clusters obtained by clustering the performance data and stores the utterances as a performance practice set; 3. The musical instrument practice system according to claim 2.
4. The playback terminal a third storage unit for storing sample sounds uploaded by the learner; The sound of the instructor's speech is replaced in accordance with the performance practice set based on the sample sound stored in the third storage unit.
4. The musical instrument practice system according to claim 3.
5. The playback terminal a setting unit that sets an overall practice time based on the performance practice set, The setting unit further sets the practice time depending on the performance proficiency of the learner.
4. The musical instrument practice system according to claim 3.
6. The playback terminal a second photographing unit that photographs an image of the person practicing and the sound thereof; a detection unit that detects whether the learner is playing the instrument correctly based on the motion and sound of the learner playing the instrument, based on the video captured by the second image capture unit; 4. The musical instrument practice system according to claim 3.
7. The playback terminal To detect incorrectly played parts, A similarity between the video captured by the second image capturing unit and the musical performance practice set is estimated, and if the similarity is equal to or greater than a predetermined threshold, the musical performance practice set is played back.
7. The musical instrument practice system according to claim 6.
8. Before the learner practices a performance, the learner refers to the performance practice set, and Select the relevant practice set and play it.
4. The musical instrument practice system according to claim 3.
9. a first receiving unit that receives the lesson video; an editing unit that automatically edits the video received by the first receiving unit in accordance with the sound included in the video; a second transmitting unit that transmits the video edited by the editing unit to a playback terminal;
10. a first receiving step of receiving a lesson video; an editing step of automatically editing the video received in the first receiving step in accordance with the sound included in the video; a second transmission step of transmitting the video edited by the editing step to a playback terminal. Data processing methods.
11. a first receiving function for receiving a lesson video; an editing function that automatically edits the video received by the first receiving function in accordance with the sound included in the video; A second transmission function is realized to transmit the video edited by the editing function to a playback terminal. program.
12. a second receiving unit that receives the edited lesson video; a second storage unit that stores the video received from the second receiving unit; a recording unit that records the sounds that the practicer is practicing; a playback unit that plays back the instructor's sound corresponding to the sound being practiced by the learner based on the sound recorded by the recording unit and the edited video. Playback device.
13. a second receiving step of receiving the edited lesson video; a second storage step of storing the video received from the second receiving step; a recording step of recording the sound being practiced by the practicer; a reproducing step of reproducing the instructor's sound corresponding to the sound being practiced by the learner based on the sound recorded in the recording step and the edited video. Data processing methods.
14. a second receiving function for receiving the edited lesson video; a second storage function for storing the video received by the second receiving function; A recording function that records the sounds that the practicer is practicing, and a playback function for playing back the instructor's sound corresponding to the sound being practiced by the learner based on the sound recorded by the recording function and the edited video. program.
Citation Information
Patent Citations
Playing training apparatus and playing training method
JP2002182553A
Video recording / playback device
JP2017032693A
Information processing method and information processing system
WO2022070769A1
Musical performance training support system
JP2013200455A
Pronunciation practice system
JP2021067844A