Method and apparatus for evaluating online musical instrument performance based on virtual reality

The virtual reality-based method addresses the limitations of existing online education by using 3D virtual reality and AI to enhance musical instrument learning, offering immersive and effective education with personalized feedback.

WO2026034655A1PCT designated stage Publication Date: 2026-02-12STUDIO VR CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/011559
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing non-face-to-face or virtual reality-based online education solutions for musical instrument learning are limited to theory and lecture-based education, resulting in low educational immersion and effectiveness, particularly for students in rural areas with limited access to education.

Method used

A virtual reality-based method that films a famous musician playing an instrument, recognizes their hand movements, generates a 3D virtual hand in a virtual space, and transmits it to a student's terminal for learning, allowing the student to practice and receive feedback on their instrument playing through a performance evaluation system using artificial intelligence for accurate hand shape recognition.

Benefits of technology

Enhances educational immersion and effectiveness by providing immersive online musical instrument education, enabling students to improve their skills quickly by identifying common errors and receiving personalized feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024011559_12022026_PF_FP_ABST
    Figure KR2024011559_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for evaluating an online musical instrument performance based on virtual reality, the method comprising the steps of: transmitting a three-dimensional virtual space including a three-dimensional virtual musical instrument and three-dimensional virtual hands to an educational terminal; acquiring a performance practice video, showing motions of a student practicing a musical instrument, from the educational terminal; dividing the performance practice video into a plurality of performance practice image frames; calculating an evaluation score for the musical instrument performance motions of the student on the basis of the plurality of performance practice image frames and a plurality of music performance image frames; determining the moment at which the performer's performance is most frequently inaccurate on the basis of the performance practice video and music performance video; generating a performance feedback report including the calculated evaluation score and the determined most frequently inaccurate moment; and transmitting the performance feedback report to the educational terminal. The method enables the student to learn how to play the musical instrument online and receive feedback on the student's performance.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for evaluating online musical instrument performance based on virtual reality

[0001] This relates to a method and device for evaluating online musical instrument performance based on virtual reality.

[0002] Virtual Reality (VR) technology simulates objects and backgrounds of the real world using a computer and provides the simulated virtual objects and backgrounds to the user, Augmented Reality (AR) technology provides virtual objects created virtually on top of images of the real world, and Mixed Reality (MR) technology provides an environment where virtual reality is grafted onto the real world, allowing physical objects in the real world to interact with virtual objects.

[0003] Virtual reality (VR) technology renders both objects and backgrounds because it converts both real-world objects and backgrounds into virtual objects and provides an image composed solely of virtual objects. In contrast, augmented reality (AR) technology renders only the target object and does not render the rest because it provides an image in which only the target object is rendered and the other objects and background are not rendered. In other words, while a virtual reality (VR) screen displays both objects and backgrounds as virtual objects generated through 3D rendering, an augmented reality (AR) screen displays only some objects as virtual objects generated through 3D rendering, while the rest display actual objects and spaces in the real world captured by a camera module.

[0004] With the recent advancements in virtual reality and communication technologies, a growing number of services utilizing the Metaverse, a 3D virtual world online, are emerging. Various services enabling meetings, education, gatherings, and gaming within the Metaverse are emerging and being utilized across diverse fields.

[0005] In particular, with the emergence of metaverse technology in the education field, non-face-to-face classes and education utilizing virtual reality technology are attracting attention as a way to resolve the educational gap for students in rural areas with lower access to education compared to urban areas.

[0006] However, existing non-face-to-face or virtual reality-based online education solutions are limited to theory and lecture-based education, which has the problem of low educational immersion and very low educational effectiveness.

[0007] The purpose of this invention is to provide a method and device for evaluating online musical instrument performance based on virtual reality. Furthermore, the technical challenges described above are not limited to these, and additional technical challenges may be derived from the following description.

[0008] According to one embodiment of the present invention, there is provided a method for transmitting a three-dimensional virtual space including a three-dimensional virtual instrument and a three-dimensional virtual hand to an educational terminal, wherein the three-dimensional virtual hand and the three-dimensional virtual instrument in the three-dimensional virtual space correspond to the hand and the instrument of the player in a music performance video representing the player's musical instrument playing motions, and the three-dimensional virtual instrument moves along the hand shape that the player moves on the instrument, and the music performance video includes a plurality of music performance image frames; a method for obtaining a performance practice video representing the motions of a student practicing a musical instrument from the educational terminal; a method for dividing the performance practice video into a plurality of performance practice image frames; a method for calculating an evaluation score for the student's musical instrument playing motions based on the plurality of performance practice image frames and the plurality of music performance image frames; a method for determining a moment at which the player's performance makes the most mistakes based on the performance practice video and the music performance video; a method for generating a performance feedback report including the calculated evaluation score and the determined moment at which the player makes the most mistakes; and a step of transmitting the performance feedback report to an educational terminal.

[0009] This virtual reality-based online music teaching method involves filming a famous musician playing an instrument and recognizing their hand movements from the recorded performance video. Based on the recognized hand movements, the online music teaching method generates a virtual hand in a 3D virtual space that moves identically to the famous musician's hand movements. This hand is then transmitted to an educational terminal and provided to the student. Students can learn how to play the instrument, specifically fingering, by observing and imitating the virtual hand movements in the provided 3D virtual space.

[0010] In addition, the online music teaching method according to the present invention provides a student with a method of playing an instrument through a three-dimensional virtual space, obtains a performance practice video showing the student's instrument playing movements from the student, compares the obtained performance practice video with a music performance video of a famous performer, and calculates an evaluation score for the student's instrument playing, thereby calculating a score for the student's instrument playing. The online music teaching method can determine the student's own instrument playing status by providing the calculated score to the student.

[0011] In addition, the online music teaching method according to the present invention provides students with the most common errors in playing an instrument, thereby providing them with points to be aware of when playing the instrument. Consequently, students can improve their instrumental skills more quickly.

[0012] Furthermore, the online music teaching method according to the present invention utilizes artificial intelligence to recognize the hand shapes of performers and students, enabling more accurate recognition of hand shapes. When recognizing hand shapes, the method utilizes a greater number of recognition points corresponding to the hands of performers and students compared to conventional hand motion recognition methods, enabling more accurate hand shape recognition.

[0013] Additionally, the online music teaching method according to the present invention utilizes virtual reality technology to teach musical instrument playing online, thereby enabling students to receive more immersive online musical instrument education.

[0014] FIG. 1 is a diagram illustrating an online music education system according to one embodiment of the present invention.

[0015] Figure 2 is a configuration diagram of the education server shown in Figure 1.

[0016] Figure 3 is a flowchart of a virtual reality-based online music teaching method according to an embodiment of the present invention.

[0017] FIG. 4 is a detailed flowchart of a step for creating a virtual space including a 3D virtual instrument and a 3D virtual hand as shown in FIG. 3.

[0018] Figure 5 is an example drawing showing the recognition points of the performer's hand.

[0019] Figure 6 is a detailed flowchart of the steps for calculating an evaluation score for a student's musical instrument playing movements illustrated in Figure 3.

[0020] Figure 7 is a flowchart of an online musical instrument performance evaluation method according to one embodiment of the present invention.

[0021] According to one embodiment of the present invention, there is provided a method for transmitting a three-dimensional virtual space including a three-dimensional virtual instrument and a three-dimensional virtual hand to an educational terminal, wherein the three-dimensional virtual hand and the three-dimensional virtual instrument in the three-dimensional virtual space correspond to the hand and the instrument of the player in a music performance video representing the player's musical instrument playing motions, and the three-dimensional virtual instrument moves along the hand shape that the player moves on the instrument, and the music performance video includes a plurality of music performance image frames; a method for obtaining a performance practice video representing the motions of a student practicing a musical instrument from the educational terminal; a method for dividing the performance practice video into a plurality of performance practice image frames; a method for calculating an evaluation score for the student's musical instrument playing motions based on the plurality of performance practice image frames and the plurality of music performance image frames; a method for determining a moment at which the player's performance makes the most mistakes based on the performance practice video and the music performance video; a method for generating a performance feedback report including the calculated evaluation score and the determined moment at which the player makes the most mistakes; and a step of transmitting the performance feedback report to an educational terminal.

[0022] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the technical idea of ​​the present disclosure is not limited to the embodiments described below and may be implemented in various different forms. The following embodiments are provided only to complete the technical idea of ​​the present disclosure and to fully inform those skilled in the art of the present disclosure of the scope of the present disclosure, and the technical idea of ​​the present disclosure is defined only by the scope of the claims.

[0023] When assigning reference numerals to components in each drawing, it should be noted that identical components are assigned the same numerals whenever possible, even if they appear on different drawings. Furthermore, when describing the present disclosure, if a detailed description of a related known configuration or function is deemed likely to obscure the gist of the present disclosure, such detailed description will be omitted.

[0024] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in the same sense as commonly understood by those of ordinary skill in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. The terminology used herein is for the purpose of describing embodiments and is not intended to limit the disclosure. In this specification, singular forms also include plural forms, unless specifically stated otherwise.

[0025] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.

[0026] As used herein, the terms “comprises” and / or “comprising” do not exclude the presence or addition of one or more other components, steps, operations and / or elements.

[0027] The terms used in the detailed description of the embodiments of the present invention below have the following meanings. “Real world” refers to a physical real space, and “virtual reality” refers to a three-dimensional virtual space that mirrors part or all of the space of real reality. “Real object” refers to a physical object existing in a physical real space, and “virtual object” refers to a non-physical object existing in a virtual space that is a replica of a physical object. “Augmented reality (AR)” is a space where real reality and virtual reality are mixed, and a three-dimensional virtual object is superimposed on an image or background of real reality to provide a single image to the user.

[0028] Components included in one embodiment and components with common functions may be described using the same designations in other embodiments. Unless otherwise stated, the descriptions provided in one embodiment may also apply to other embodiments, and specific descriptions may be omitted to the extent that they overlap or would be readily apparent to those skilled in the art.

[0029] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0030] FIG. 1 is a diagram illustrating an online music education system according to one embodiment of the present invention. Referring to FIG. 1, a music education system (1) according to one embodiment of the present invention includes an education server (11), a performer terminal (12), and an education terminal (13). The education server (11) is a server that generates non-face-to-face music education content and provides non-face-to-face music education services. It transmits and receives data and signals related to online music education from the performer terminal (12) and the education terminal (13). The education server (11) receives data and signals for generating online music education content from the performer terminal (13), and generates music education content including a three-dimensional virtual object based on the data and signals transmitted from the performer terminal (13). The education server (11) transmits the generated music education content to the education terminal (13).

[0031] The performer terminal (12) serves as a performer's terminal, capturing the performer's movements while playing an instrument and utilizing this as basic data for music education content. The performer, who is the user of the performer terminal (12), generates performance data representing the performance of the instrument and transmits the generated performance data to the education server (11).

[0032] The educational terminal (13) is a terminal for students learning to play a musical instrument, and outputs music education content transmitted from the educational server (11). The educational terminal (13) records the student's musical instrument playing movements according to the output music education content, and transmits the recorded video data to the educational server (11).

[0033] Additionally, the education server (11) generates feedback on the student's performance based on video data transmitted from the education terminal (13). The education server (11) then transmits the generated feedback back to the education terminal (13).

[0034] Here, the education server (11) may be a single computer or a collection of multiple computers. The performer terminal (12) and the education terminal (13) are computing devices with communication functions, which can execute various applications and can execute online music lesson applications. Examples of the performer terminal (12) and the education terminal (13) include a head-mounted display (HMD), a VR / AR headset, smartglass, a smartphone, a tablet PC, etc. In addition to the exemplary devices described above, the performer terminal (12) and the education terminal (13) include a display module capable of outputting the generated composite image.

[0035] The steps of constructing an online music teaching method according to an embodiment of the present invention can be performed in the education server (11), performer terminal (12), and education terminal (13) of the music teaching system (1). Below, a method for generating a composite image including a three-dimensional virtual object will be described in detail.

[0036] Figure 2 is a configuration diagram of the education server (11) illustrated in Figure 1. Referring to Figure 2, the education server (11) includes a processor (1101), a communication module (1102), storage (1103), a virtual space creation module (1104), a performance practice analysis module (1105), and a report creation module (1106).

[0037] The processor (1101) of the education server (11) processes general tasks performed on the education server (11).

[0038] The communication module (1102) of the education server (11) can transmit and receive data, messages, and signals to and from a performer terminal (12), an education terminal (13), and an external server (not shown) through a wide area network such as the Internet by connecting to a mobile communication base station or a Wi-Fi repeater.

[0039] The storage (1103) of the education server (11) stores data required to execute an application or program that implements a virtual reality-based online music teaching method. For example, the storage (1103) stores generated virtual reality music education content. In addition, the storage (1103) stores a training set for training at least one artificial neural network included in the education server (11).

[0040] The virtual space creation module (1104) of the education server (11) analyzes a music performance video in which a performer plays an instrument, and creates a 3D virtual space including 3D virtual objects (3D virtual hands and 3D virtual instruments) corresponding to the shape of the performer's hands and the shape of the instrument. The virtual space creation module (1104) divides the music performance video into frames, recognizes the shape of the performer's hands and the shape of the instrument represented by each divided frame, renders a 3D virtual object corresponding to the recognized shape of the hand and the shape of the instrument, and places it on the 3D virtual space, thereby creating a 3D virtual space. A specific method for creating a 3D virtual space including a 3D virtual hand and a virtual instrument in the virtual space creation module (1104) will be described in detail below.

[0041] The performance practice analysis module (1105) of the education server (11) analyzes a performance practice video of a student playing an instrument input from an education terminal (13), and calculates an evaluation score for the student's instrument playing movements. The performance practice analysis module (1105) compares the performance practice video with a music performance video, and calculates an evaluation score for the student's instrument playing movements based on the comparison result. More specifically, the performance practice analysis module (1105) compares the hand shapes of the student and the performer in the performance practice video and the music performance video, and calculates a score for the student's instrument playing movements based on the degree to which the student's hand shapes are similar to the performer's hand shapes. A specific method for calculating an evaluation score for the student's instrument playing movements in the performance practice analysis module (1105) will be described in detail below.

[0042] The report generation module (1106) of the education server (11) generates a performance feedback report including evaluation scores for the performance practice video produced by the performance practice analysis module (1105). The report generation module (1106) generates a performance feedback mock-up to be provided to students who are users of the education terminal (13). The performance feedback report includes scores for the performance practice date and time, the title of the performance practice song, the evaluation score, and the most common moment of error.

[0043] In the education server (11), the virtual space creation module (1104), the performance practice analysis module (1105), and the report creation module (1106) may be implemented by a separate dedicated processor different from the processor (1101), or may be implemented by a computer program running on the processor (1101).

[0044] The education server (11) may include additional components in addition to the components described above. For example, as illustrated in FIG. 2, the education server (11) includes a bus for transmitting data between components, and, although omitted from FIG. 2, further includes a power module for supplying driving power to each component.

[0045] Thus, descriptions of components that are obvious to those skilled in the art to which this embodiment pertains will be omitted as they may obscure the specific features of this embodiment. Below, each component of the education server (11) will be described in detail in the course of explaining a virtual reality-based online music teaching method according to one embodiment of the present invention.

[0046] FIG. 3 is a flowchart of a virtual reality-based online music teaching method according to an embodiment of the present invention. Referring to FIG. 3, in step 301, the education server (11) obtains a music performance video showing the movements of a performer playing an instrument. The communication module (1102) of the education server (11) receives the music performance video from the performance terminal (12). The performance terminal (12) generates a music performance video showing the movements of the performer playing the instrument. The camera module (1204) of the performance terminal (12) photographs the movements of the performer playing the instrument and generates a music performance video showing the movements of the performer playing the instrument. The performance terminal (12) transmits the generated music performance video to the education server (11) through the communication module (1202).

[0047] The education server (11) can obtain a performance video by searching and loading a music performance video stored in storage (1103). The education server (11) inputs the obtained music performance video into the virtual space creation module (1104).

[0048] In step 302, the education server (11) creates a three-dimensional virtual space including a virtual instrument and a virtual hand of a performer based on the acquired music performance video. The virtual space creation module (1104) of the education server (11) analyzes the music performance video acquired in step 301 and creates a three-dimensional virtual space including a three-dimensional virtual hand representing the hand of a performer included in the music performance video. The virtual space creation module (1104) extracts the finger movements of the instrument and the performer included in the music performance video, creates a three-dimensional virtual hand that moves on a three-dimensional virtual instrument according to the extracted instrument and finger movements, and creates a three-dimensional virtual space including the three-dimensional virtual hand that moves on the virtual instrument.

[0049] The steps for creating 3D educational content including 3D virtual hands moving on a virtual instrument will be described below with reference to FIG. 4.

[0050] FIG. 4 is a detailed flowchart of a step of creating a virtual space including a three-dimensional virtual instrument and a three-dimensional virtual hand illustrated in FIG. 3. Referring to FIG. 4, in step 3021, the virtual space creation module (1104) of the education server (11) divides the acquired music performance video into a plurality of music performance image frames. The virtual space creation module (1104) divides the music performance video into a plurality of music performance image frames included in the music performance video. For example, if the music performance video is 24 FPS (Frames Per Second) and the total playback time is 60 seconds, the virtual space creation module (1104) divides the music performance video into 1,440 frames.

[0051] Here, a music performance video is a video depicting the movements of a performer playing an instrument. A music performance image frame is an image depicting the moment the performer plays the instrument. A music performance video is composed of multiple music performance image frames.

[0052] At step 3022, the virtual space creation module (1104) of the education server (11) recognizes the hand shape and the instrument shape of the performer included in the plurality of segmented music performance image frames, and determines the position of each finger of the performer on the instrument based on the recognized hand shape and instrument shape. The virtual space creation module (1104) recognizes the instrument and hand shape represented by each of the plurality of music performance image frames. More specifically, the virtual space creation module (1104) recognizes and determines the hand shape and the instrument shape represented by each music performance image frame.

[0053] For example, if the instrument is a piano, the virtual space generation module (1104) recognizes the keys of the piano represented by each of a plurality of music performance image frames. The virtual space generation module (1104) recognizes the keys of the piano included in each music performance image frame. The virtual space generation module (1104) includes a model determination artificial neural network that determines a piano model from a portion representing the piano keys included in the music performance image frame. The model determination artificial neural network is an artificial neural network that is pre-trained to output a piano model included in an image representing the piano keys from an image representing the piano keys. The model determination artificial neural network is an artificial neural network that is pre-trained to output a model of the piano included in an input image representing the piano keys by a training set that includes training images representing the piano keys and the piano models included in each training image. The virtual space generation module (1104) recognizes the piano keys included in the music performance image frame based on the determined piano model. Here, the virtual space creation module (1104) recognizes the piano keys included in each music performance image frame based on the number, size, arrangement, etc. of white and black keys for each piano model stored in the storage (1103).

[0054] In addition, the virtual space generation module (1104) recognizes the shape of the player's hand included in each music performance image frame. The virtual space generation module (1104) determines recognition points of the player's hand based on the hand shape represented by each music performance image frame, and recognizes the shape of the player's hand based on the recognition points. The virtual space generation module (1104) further includes a hand recognition artificial neural network that is trained in advance to output a plurality of recognition points corresponding to the shape of the player's hand from an image frame representing the shape of the player's hand. The hand recognition artificial neural network is an artificial neural network that is trained in advance to output a plurality of recognition points corresponding to each part of the hand included in an image representing the shape of the hand. The hand recognition artificial neural network is an artificial neural network that is trained in advance to output a model of a piano included in an input image representing a piano keyboard by means of a training set including training images representing the shape of the player's hand and a plurality of recognition points corresponding to the shape of the hand included in each training image.

[0055] The virtual space creation module (1104) determines recognition points representing specific parts of the hand, and recognizes the shape of the performer's hand included in each music performance image frame based on the determined recognition points.

[0056] In this regard, FIG. 5 of the present invention is an exemplary drawing showing recognition points of a performer's hand. Referring to FIG. 5, (a) of FIG. 5 is an example showing the number of recognition points corresponding to a palm when recognizing a conventional hand shape, and (b) of FIG. 5 is an example showing the number of recognition points utilized in a hand shape recognition method according to the present invention. The conventional hand shape recognition method exemplified in FIG. 5 (a) generally recognizes a hand shape using 21 recognition points. The hand shape recognition method of the present invention exemplified in FIG. 5 (b) recognizes a hand shape using 63 recognition points.

[0057] The present invention can accurately recognize hand shapes by recognizing the hand shapes of a performer or student using more recognition points compared to conventional hand shape recognition methods.

[0058] The virtual space creation module (1104) determines the position of each finger of the player's hand on the piano keyboard in each music performance image frame based on the shape of the piano keyboard and the player's hand recognized from each music performance image frame.

[0059] In step 3023, the virtual space generation module (1104) of the education server (11) renders a virtual instrument and a virtual hand based on the recognized hand shape and instrument shape, and generates a three-dimensional virtual space including the virtual instrument and the virtual hand. The virtual space generation module (1104) renders a virtual instrument corresponding to an instrument included in a music performance image frame and a virtual hand corresponding to the performer's hand based on the hand shape, instrument shape, and finger positions recognized in step 3022. More specifically, the virtual space generation module (1104) generates a three-dimensional virtual space. The virtual space generation module (1104) places the virtual instrument rendered on the generated three-dimensional virtual space at the center of the three-dimensional virtual space, and places the virtual hand rendered on the placed virtual instrument according to the finger positions determined in step 3022. The virtual space generation module (1104) changes the position and shape of a virtual hand rendered in a three-dimensional virtual space according to the shape of the hand and the shape of the instrument in each of a plurality of music performance image frames in step 3022. In the three-dimensional virtual space, the virtual hand moves along the shape of the player's hand moving on the virtual instrument on the real instrument. On the virtual instrument, the virtual space generation module (1104) generates a three-dimensional virtual instrument and virtual hand corresponding to the player's performance movements shown in the performance video in the above-described manner.

[0060] The virtual space generation module (1104) distinguishes a plurality of music performance image frames segmented from a music performance video into music performance image frames having the same hand shape and instrument shape. The virtual space generation module (1104) distinguishes a plurality of music performance image frame groups having the same hand shape and instrument shape. The virtual space generation module (1104) determines at least one music performance image frame having the same hand shape and instrument shape as the same music performance image frame group. For example, the virtual space generation module (1104) can distinguish 1,440 music performance image frames into 100 music performance image frame groups. A method for distinguishing a plurality of music performance image frame groups will be described in detail below.

[0061] The virtual space generation module (1104) compares any two music performance image frames among a plurality of music performance image frames and calculates the similarity between the two music performance image frames. The virtual space generation module (1104) includes a similarity artificial neural network that calculates the similarity between the two music performance image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output the similarity between the images represented by the two input image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output the similarity between the two input image frames by a training set that includes a plurality of training image frames representing an appearance of a musician playing an instrument and the training similarity between the two training image frames. Here, the similarity is expressed as a numerical score from 0 to 100. When the similarity is 100, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely identical, and when the similarity is 0, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely different.

[0062] The virtual space generation module (1104) determines music performance image frames having a similarity higher than a preset threshold similarity score as a music performance image frame group. Here, the threshold similarity score is a similarity score that serves as a criterion for considering hand shapes and instrument shapes represented by the music performance image frames as the same, and can be preset by the user. The threshold similarity score is proportional to the resolution and illumination of the music performance video, respectively. When the resolution or illumination of the music performance video is high, the threshold similarity score is set high. Conversely, when the resolution or illumination of the music performance video is low, the threshold similarity score is set low.

[0063] The virtual space creation module (1104) creates a three-dimensional virtual instrument and a virtual hand for one music performance image frame for each of a plurality of distinct music performance image frame groups, and associates the created three-dimensional virtual instrument and virtual hand with all music performance image frames included in the plurality of music performance image frame groups.

[0064] The virtual space generation module (1104) recognizes hand shapes and instrument shapes for multiple music performance image frames having the same hand shape and instrument shape, and performs the process of rendering a 3D virtual hand and virtual instrument only once based on the recognized hand shape and instrument shape, and associates it with all music performance image frames within the same group, thereby omitting the recognition step and the rendering step for all music performance image frames. Accordingly, a 3D virtual space including a 3D virtual hand and virtual instrument can be generated more quickly.

[0065] The virtual space creation module (1104) inputs a three-dimensional virtual space including the created virtual instrument and virtual hand into the communication module (1102).

[0066] In step 303, the education server (11) transmits the generated three-dimensional virtual space to the education terminal (13). The communication module (1102) of the education server (11) transmits the three-dimensional virtual space including the three-dimensional virtual instrument and virtual hand generated in step 302 to the education terminal (13). Here, the communication module (1102) transmits data representing the three-dimensional virtual space to the education terminal (13) using the communication module (1102).

[0067] An educational terminal (13) that receives data representing a three-dimensional virtual space outputs the three-dimensional virtual space. More specifically, the educational terminal (13) outputs a virtual instrument and a virtual hand on the three-dimensional virtual space through an output module (130X). The educational terminal (13) outputs a virtual instrument and a virtual hand on the three-dimensional virtual space from the viewpoint of a virtual camera at a predetermined position and direction on the three-dimensional virtual space. The educational terminal (13) shows to the user of the educational terminal (13), that is, the student, that the shape and position of the virtual hand on the virtual instrument in the three-dimensional virtual space change in time series. For example, the educational terminal (13) shows the movement of the virtual hand on the virtual piano keyboard in the three-dimensional virtual space to the student of the educational terminal (13). The student learns how to play the instrument by watching the movement of the virtual piano keyboard and the virtual hand on the three-dimensional virtual space output through the educational terminal (13) and imitating the hand shape of the virtual hand.

[0068] The educational terminal (13) generates a performance practice video showing the movements of a student practicing an instrument by following the movements of a virtual hand on a virtual instrument in a three-dimensional virtual space. The camera module of the educational terminal (13) captures the movements of the student practicing the instrument and generates a performance practice video showing the movements of the student practicing the instrument.

[0069] In step 304, the education server (11) obtains a performance practice video from the education terminal (13) in response to the transmission of the 3D virtual space. The communication module (1102) of the education server (11) receives the performance practice video from the education terminal (13). The communication module (1102) of the education server (11) inputs the obtained performance practice video into the performance practice analysis module (1105).

[0070] In step 305, the education server (11) calculates an evaluation score for the student's instrument playing based on the performance practice video and the music performance video. The performance practice analysis module (1105) of the education server (11) compares the performance practice video and the music performance video, and calculates an evaluation score for the student's instrument playing movements based on the comparison result. The performance practice analysis module (1105) compares the student's instrument playing movements in the performance practice video with the player's instrument playing movements in the music performance video, calculates a similarity between the student's instrument playing movements and the player's instrument playing movements based on the comparison result, and calculates an evaluation score for the student's instrument playing movements based on the calculated similarity.

[0071] The steps for calculating the evaluation score for a student's musical instrument playing movements will be described below with reference to Figure 6.

[0072] Fig. 6 is a detailed flowchart of a step for calculating an evaluation score for a student's musical instrument playing movements illustrated in Fig. 3. Referring to Fig. 6, in step 3051, the performance practice analysis module (1105) of the education server (11) divides the acquired performance practice video into a plurality of performance practice image frames. The performance practice analysis module (1105) divides the performance practice video into a plurality of performance practice image frames included in the performance practice video. For example, if the performance practice video is 24 FPS and the total playback time is 60 seconds, the performance practice analysis module (1105) divides the performance practice video into 1,440 frames.

[0073] Here, the performance practice video has the same FPS and playback time as the music performance video. If the FPS and playback time of the performance practice video are different from those of the music performance video, the performance practice analysis module (1105) converts the performance practice video so that the FPS and playback time of the performance practice video are the same as those of the music performance video. The specific process of converting the FPS and playback time is omitted as it obscures the features of the present invention.

[0074] In step 3052, the performance practice analysis module (1105) of the education server (11) calculates an evaluation score for the student's musical instrument performance movements based on a plurality of performance practice image frames and a plurality of music performance image frames. The performance practice analysis module (1105) compares the performance practice image frames with the corresponding music performance image frames and calculates an evaluation score for the student's musical instrument performance movements. For example, when comparing 1,440 performance practice image frames and 1,440 music performance image frames, the performance practice analysis module (1105) compares the first performance practice image frame with the first music performance image frame, compares the second performance practice image frame with the second music performance image frame, and compares the nth performance practice image frame with the nth music performance image frame.

[0075] The performance practice analysis module (1105) according to an embodiment of the present invention recognizes the hand shape and the instrument shape of the student included in each of a plurality of performance practice image frames, and determines the position of each finger of the student on the instrument based on the recognized hand shape and instrument shape. The performance practice analysis module (1105) recognizes the instrument and hand shape represented by each of the plurality of performance practice image frames. More specifically, the performance practice analysis module (1105) recognizes and determines the hand shape and the instrument shape represented by each performance practice image frame.

[0076] For example, if the instrument is a piano, the performance practice analysis module (1105) recognizes the piano keys represented by each of a plurality of performance practice image frames. The performance practice analysis module (1105) recognizes the piano keys included in each performance practice image frame. The performance practice analysis module (1105) includes a model determination artificial neural network that determines a piano model from a portion representing the piano keys included in the performance practice image frames. The model determination artificial neural network is an artificial neural network that is pre-trained to output a piano model included in an image representing a piano keyboard from an image representing the piano keyboard. The model determination artificial neural network is an artificial neural network that is pre-trained to output a model of a piano included in an input image representing the piano keyboard by a training set that includes training images representing the piano keyboard and the piano models included in each training image. The performance practice analysis module (1105) recognizes the piano keys included in the performance practice image frames based on the determined piano models. Here, the performance practice analysis module (1105) recognizes the piano keys included in each performance practice image frame based on the number, size, arrangement, etc. of white keys and black keys for each piano model stored in the storage (1103).

[0077] In addition, the performance practice analysis module (1105) recognizes the hand shape of the performer included in each performance practice image frame. The performance practice analysis module (1105) determines recognition points of the student's hand based on the hand shape indicated by each performance practice image frame, and recognizes the hand shape of the student based on the recognition points. The performance practice analysis module (1105) determines recognition points indicating the position of each hand, and recognizes the hand shape of the student included in each music performance image frame based on the determined recognition points.

[0078] The performance practice analysis module (1105) determines the position of each finger of the player's hand on the piano keys in each performance practice image frame based on the piano keys and the player's hand shape recognized from each performance practice image frame. The performance practice analysis module (1105) determines the position of each finger of the student on the instrument based on the hand shape and the instrument shape recognized from each performance practice image frame.

[0079] The performance practice analysis module (1105) compares the positions of the student's fingers on the instrument with the positions of the player's fingers on the instrument, and determines the number of the student's fingers that match the positions of the player's fingers. The performance practice analysis module (1105) compares the positions of each of the student's fingers with the positions of each of the player's fingers, and determines the number of fingers that are positioned at the same position. For example, the performance practice analysis module (1105) determines the number as 10 if the positions of the student's fingers and the player's fingers all match, determines the number as 8 if the positions of 8 fingers out of 10 match, and determines the number as 0 if the positions of all the fingers do not match.

[0080] The performance practice analysis module (1105) compares the finger positions of the student and the performer in each of a plurality of performance practice image frames and each corresponding music performance image frame, and determines the number of fingers whose positions match. For example, when comparing 1,440 performance practice image frames and 1,440 music performance image frames, the performance practice analysis module (1105) determines the number of fingers whose positions match in the first performance practice image frame and the first music performance image frame, determines the number of fingers whose positions match in the second performance practice image frame and the second music performance image frame, and determines the number of fingers whose positions match in the nth performance practice image frame and the nth music performance image frame. The performance practice analysis module (1105) determines the number of the student's fingers whose positions match in all of the performance practice image frames with the performer's finger positions.

[0081] The performance practice analysis module (1105) calculates an evaluation score for the student's instrument performance based on the number of fingers that match the determined positions. More specifically, the performance practice analysis module (1105) calculates an evaluation score for the student's instrument performance based on the number of multiple performance practice frames and the determined number of matching fingers. The performance practice analysis module (1105) calculates an evaluation score based on the finger matching rate, which is the ratio of the number of finger positions of the student in the entire performance practice frames to the total number of fingers that match the finger positions of the performer. For example, if there are 1,440 performance practice frames, the number of finger positions of the student in the entire performance practice frames is 1,440 × 10 (the number of fingers on both hands) = 14,440. Here, if the number of the student's fingers that match the finger positions of the performer is determined to be 9,780, the finger matching rate is The performance practice analysis module (1105) determines the calculated finger coincidence rate as an evaluation score. In the example described above, the performance practice analysis module (1105) calculates the calculated finger coincidence rate of 67.91 as an evaluation score for the student's musical instrument playing movements.

[0082] A performance practice analysis module (1105) according to another embodiment of the present invention compares a performance practice image frame with a corresponding music performance image frame, and calculates a similarity between the performance practice image frame and the corresponding music performance image frame. The performance practice analysis module (1105) includes a similarity artificial neural network that calculates a similarity between a performance practice image frame and a corresponding music performance image frame. The similarity artificial neural network is an artificial neural network that is pre-trained to output a similarity between images represented by two input image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output a similarity between two input image frames by a training set that includes a plurality of training image frames representing an appearance of a performer playing a musical instrument and a training similarity between the two training image frames. Here, the similarity is expressed as a numerical score from 0 to 100. When the similarity is 100, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely identical, and when the similarity is 0, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely different.

[0083] The performance practice analysis module (1105) calculates the similarity between each of a plurality of performance practice image frames and the corresponding music performance image frame. For example, when comparing 1,440 performance practice image frames and 1,440 music performance image frames, the performance practice analysis module (1105) calculates the similarity between the first performance practice image frame and the first music performance image frame, calculates the similarity between the second performance practice image frame and the second music performance image frame, and calculates the similarity between the nth performance practice image frame and the nth music performance image frame.

[0084] The performance practice analysis module (1105) calculates an evaluation score based on the similarity calculated from all performance practice image frames. The performance practice analysis module (1105) calculates the average value of the overall similarity calculated from all performance practice image frames as the evaluation score.

[0085] In step 306, the education server (11) determines the moment at which the performer makes the most mistakes in his / her performance based on the performance practice video and the music performance video. The performance practice analysis module (1105) of the education server (11) compares each of the plurality of performance practice image frames with the corresponding instrument performance image frame, and, based on the comparison result, determines the performance practice image frame that is most different from the corresponding music performance image frame among the plurality of performance practice image frames of the performance practice video.

[0086] According to one embodiment of the present invention, the performance practice analysis module (1105) determines a performance practice image frame having the smallest number of fingers where the student's finger positions on the instrument of the performance practice image frame and the performer's finger positions on the instrument of the music performance image frame match, and determines the determined performance practice image frame as the moment with the most mistakes.

[0087] According to another embodiment of the present invention, the performance practice analysis module (1105) determines a performance practice image frame having the lowest similarity between each performance practice image frame and a corresponding music performance image frame among a plurality of performance practice image frames, and determines the determined performance practice image frame as the moment with the most mistakes.

[0088] The performance practice analysis module (1105) inputs the evaluation score calculated in step 305 and the most frequently wrong moment determined in step 306 into the report generation module (1106).

[0089] In step 307, the training server (11) generates a performance feedback report including the calculated evaluation scores and the determined most frequently erroneous moments. The report generation module (1106) of the training server (11) generates a performance feedback report including the calculated evaluation scores in step 305 and the determined most frequently erroneous moments in step 306. The report generation module (1106) inputs the generated performance feedback report to the communication module (1102).

[0090] At step 308, the training server (11) transmits a performance feedback report to the training terminal (13). The communication module (1102) of the training server (11) transmits the generated performance feedback report to the training terminal (13).

[0091] Figure 7 is a flowchart of an online musical instrument performance evaluation method according to an embodiment of the present invention. Referring to Figure 7, in step 701, the communication module (1102) of the education server (11) transmits a 3D virtual space including a 3D virtual instrument and a 3D virtual hand generated in step 302 to the education terminal (13). Here, the communication module (1102) transmits data representing the 3D virtual space to the education terminal (13) using the communication module (1102).

[0092] An educational terminal (13) that receives data representing a three-dimensional virtual space outputs the three-dimensional virtual space. More specifically, the educational terminal (13) outputs a virtual instrument and a virtual hand on the three-dimensional virtual space through an output module (130X). The educational terminal (13) outputs a virtual instrument and a virtual hand on the three-dimensional virtual space from the viewpoint of a virtual camera at a predetermined position and direction on the three-dimensional virtual space. The educational terminal (13) shows to the user of the educational terminal (13), that is, the student, that the shape and position of the virtual hand on the virtual instrument in the three-dimensional virtual space change in time series. For example, the educational terminal (13) shows the movement of the virtual hand on the virtual piano keyboard in the three-dimensional virtual space to the student of the educational terminal (13). The student learns how to play the instrument by watching the movement of the virtual piano keyboard and the virtual hand on the three-dimensional virtual space output through the educational terminal (13) and imitating the hand shape of the virtual hand.

[0093] In step 702, the education server (11) obtains a performance practice video from the education terminal (13) in response to the transmission of the three-dimensional virtual space. The communication module (1102) of the education server (11) receives the performance practice video from the education terminal (13). The communication module (1102) of the education server (11) inputs the obtained performance practice video into the performance practice analysis module (1105).

[0094] The educational terminal (13) generates a performance practice video showing the movements of a student practicing an instrument by following the movements of a virtual hand on a virtual instrument in a three-dimensional virtual space. The camera module of the educational terminal (13) captures the movements of the student practicing the instrument and generates a performance practice video showing the movements of the student practicing the instrument. The educational terminal (13) transmits the generated performance practice video to the educational server (11).

[0095] In step 703, the performance practice analysis module (1105) of the education server (11) divides the acquired performance practice video into multiple performance practice image frames. The performance practice analysis module (1105) divides the performance practice video into multiple performance practice image frames. For example, if the performance practice video is 24 FPS and the total playback time is 60 seconds, the performance practice analysis module (1105) divides the performance practice video into 1,440 frames.

[0096] Here, the performance practice video has the same FPS and playback time as the music performance video. If the FPS and playback time of the performance practice video are different from those of the music performance video, the performance practice analysis module (1105) converts the performance practice video so that the FPS and playback time of the performance practice video are the same as those of the music performance video. The specific process of converting the FPS and playback time is omitted as it obscures the features of the present invention.

[0097] In step 704, the performance practice analysis module (1105) of the education server (11) calculates an evaluation score for the student's musical instrument performance movements based on a plurality of performance practice image frames and a plurality of music performance image frames. The performance practice analysis module (1105) compares the performance practice image frames with the corresponding music performance image frames and calculates an evaluation score for the student's musical instrument performance movements. For example, when comparing 1,440 performance practice image frames and 1,440 music performance image frames, the performance practice analysis module (1105) compares the first performance practice image frame with the first music performance image frame, compares the second performance practice image frame with the second music performance image frame, and compares the nth performance practice image frame with the nth music performance image frame. The specific method for calculating the evaluation score will be replaced with the content described in step 3052.

[0098] In step 705, the education server (11) determines the moment at which the performer makes the most mistakes in his / her performance based on the performance practice video and the music performance video. The performance practice analysis module (1105) of the education server (11) compares each of the plurality of performance practice image frames with the corresponding instrument performance image frame, and, based on the comparison result, determines the performance practice image frame that is most different from the corresponding music performance image frame among the plurality of performance practice image frames of the performance practice video.

[0099] The performance practice analysis module (1105) inputs the evaluation score calculated in step 703 and the most frequently wrong moment determined in step 704 into the report generation module (1106).

[0100] In step 706, the training server (11) generates a performance feedback report including the calculated evaluation scores and the determined most frequently erroneous moments. The report generation module (1106) of the training server (11) generates a performance feedback report including the calculated evaluation scores in step 305 and the determined most frequently erroneous moments in step 306. The report generation module (1106) inputs the generated performance feedback report to the communication module (1102).

[0101] At step 707, the training server (11) transmits a performance feedback report to the training terminal (13). The communication module (1102) of the training server (11) transmits the generated performance feedback report to the training terminal (13).

[0102] The virtual reality-based online music teaching method according to the embodiments of the present invention described above films a famous musician playing an instrument, and recognizes the hand movements of the famous musician playing the instrument from the filmed performance video. Based on the recognized hand movements, the online music teaching method generates a virtual hand that moves identically to the hand movements of the famous musician in a three-dimensional virtual space, and transmits this to an educational terminal to provide it to a student. The student can learn how to play the instrument, i.e., fingering, by observing the movements of the virtual hand in the provided three-dimensional virtual space and imitating the movements of the virtual hand.

[0103] In addition, the online music teaching method according to the present invention provides a student with a method of playing an instrument through a three-dimensional virtual space, obtains a performance practice video showing the student's instrument playing movements from the student, compares the obtained performance practice video with a music performance video of a famous performer, and calculates an evaluation score for the student's instrument playing, thereby calculating a score for the student's instrument playing. The online music teaching method can determine the student's own instrument playing status by providing the calculated score to the student.

[0104] In addition, the online music teaching method according to the present invention provides students with the most common errors in playing an instrument, thereby providing them with points to be aware of when playing the instrument. Consequently, students can improve their instrumental skills more quickly.

[0105] Furthermore, the online music teaching method according to the present invention utilizes artificial intelligence to recognize the hand shapes of performers and students, enabling more accurate recognition of hand shapes. When recognizing hand shapes, the method utilizes a greater number of recognition points corresponding to the hands of performers and students compared to conventional hand motion recognition methods, enabling more accurate hand shape recognition.

[0106] Additionally, the online music teaching method according to the present invention utilizes virtual reality technology to teach musical instrument playing online, thereby enabling students to receive more immersive online musical instrument education.

[0107] Meanwhile, the embodiments of the present invention described above can be written as a program that can be executed on a computer, and can be implemented in a general-purpose digital computer that runs the program using a computer-readable recording medium. In addition, the structure of the data used in the embodiments of the present invention described above can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as a magnetic storage medium (e.g., a ROM floppy disk, a hard disk, etc.) and an optical reading medium (e.g., a CD-ROM, a DVD, etc.). A program that performs a virtual reality-based online music teaching method according to the embodiments of the present invention is recorded on the computer-readable recording medium.

[0108] The present invention has been described with a focus on preferred embodiments. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

[0109] 11: Education Server

[0110] 1101: Processor 1102: Communication Module

[0111] 1103: Storage 1104: Virtual space creation module

[0112] 1105: Performance Practice Analysis Module 1106: Report Generation Module

[0113] 12: Player terminal

[0114] 13: Educational terminal

Claims

1. In the virtual reality-based online musical instrument performance evaluation method, A step of transmitting a three-dimensional virtual space including a three-dimensional virtual instrument and a three-dimensional virtual hand to an educational terminal, wherein the three-dimensional virtual hand and the three-dimensional virtual instrument in the three-dimensional virtual space correspond to the hand and the instrument of a performer in a music performance video representing the performer's musical instrument playing motion, and the three-dimensional virtual instrument moves along the shape of the hand moving on the instrument, and the music performance video includes a plurality of music performance image frames; A step of obtaining a performance practice video showing the movements of a student practicing a musical instrument from the above educational terminal; A step of dividing the above performance practice video into multiple performance practice image frames; A step of calculating an evaluation score for a student's musical instrument playing movements based on the plurality of performance practice image frames and the plurality of music performance image frames; A step of determining the moment at which the performer makes the most mistakes in his / her performance based on the above performance practice video and the above music performance video; A step of generating a performance feedback report including the above-mentioned calculated evaluation scores and the most frequently wrong moments determined; and An online musical instrument performance evaluation method, characterized in that it comprises a step of transmitting the above performance feedback report to an educational terminal.

Citation Information

Patent Citations

  • Musical instrument practicing device

    JP2019053170A

  • Apparatus and method for performing virtual musical instrument on the basis of finger-motion

    KR1020130067856A

  • CCTV camera pole

    KR1020210141067A

  • Locking device that can remove the lock from the door in case of emergency release

    KR102503421B1

  • Online music teaching method and apparatus based on virtual reality

    KR102622163B1