Virtual reality-based online music teaching method and device having improved object recognition rate

The method generates a 3D virtual space with precise finger recognition and feedback to enhance immersion and accuracy in virtual reality-based online music education, addressing the limitations of existing systems.

WO2026018958A1PCT designated stage Publication Date: 2026-01-22STUDIO VR CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/011554
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2024-08-06
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing virtual reality-based online music education solutions suffer from low educational immersion and ineffective instrumental technique learning due to inaccurate finger position recognition in virtual reality-based online instrumental lessons.

Method used

A method involving the generation of a three-dimensional virtual space with a virtual hand and instrument based on a music performance video, utilizing advanced image frame transformation and artificial intelligence for precise finger position recognition, followed by evaluation and feedback on student performance.

Benefits of technology

Enhances educational immersion and accuracy in learning instrumental techniques by providing precise finger position recognition and performance feedback, enabling students to improve their skills more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024011554_22012026_PF_FP_ABST
    Figure KR2024011554_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a virtual reality-based online music teaching method and device, in which the hand movements of a performer, extracted from a video of the performer playing an instrument, are converted into a three-dimensional (3D) virtual hand within a 3D virtual space, and the converted virtual hand is then provided to a learner, enabling the learner to observe and imitate the instrument-playing movements of the 3D virtual hand through an educational terminal, thereby allowing the learner to learn to play an instrument online.
Need to check novelty before this filing date? Find Prior Art

Description

A virtual reality-based online music teaching method and device with improved object recognition rate

[0001] The present invention relates to a virtual reality-based online music teaching method and device having an improved object recognition rate.

[0002] Virtual Reality (VR) technology simulates objects and backgrounds of the real world using a computer and provides the simulated virtual objects and backgrounds to the user, Augmented Reality (AR) technology provides virtual objects created virtually on top of images of the real world, and Mixed Reality (MR) technology provides an environment where virtual reality is grafted onto the real world, allowing physical objects in the real world to interact with virtual objects.

[0003] Virtual reality (VR) technology renders both objects and backgrounds because it converts both real-world objects and backgrounds into virtual objects and provides an image composed solely of virtual objects. In contrast, augmented reality (AR) technology renders only the target object and does not render the rest because it provides an image in which only the target object is rendered and the other objects and background are not rendered. In other words, while a virtual reality (VR) screen displays both objects and backgrounds as virtual objects generated through 3D rendering, an augmented reality (AR) screen displays only some objects as virtual objects generated through 3D rendering, while the rest display actual objects and spaces in the real world captured by a camera module.

[0004] With the recent advancements in virtual reality and communication technologies, a growing number of services utilizing the Metaverse, a 3D virtual world online, are emerging. Various services enabling meetings, education, gatherings, and gaming within the Metaverse are emerging and being utilized across diverse fields.

[0005] In particular, with the emergence of metaverse technology in the education field, non-face-to-face classes and education utilizing virtual reality technology are attracting attention as a way to resolve the educational gap for students in rural areas with lower access to education compared to urban areas.

[0006] However, existing non-face-to-face or virtual reality-based online education solutions are limited to theory and lecture-based education, which has the problem of low educational immersion and very low educational effectiveness.

[0007] To address these issues, online instrumental lessons utilizing virtual reality technology were developed. However, they suffered from difficulties accurately identifying the finger positions of the player. Consequently, existing virtual reality-based online instrumental lessons hindered users from accurately learning instrument fingering techniques.

[0008] The purpose of this invention is to provide a virtual reality-based online music teaching method and device with improved recognition rates. Furthermore, the technical challenges described above are not limited to these challenges, and other technical challenges may be derived from the following description.

[0009] A virtual reality-based online music teaching method according to one embodiment of the present invention comprises the steps of: obtaining a music performance video showing a movement of a performer playing a musical instrument; generating a three-dimensional virtual space including a three-dimensional virtual hand and a virtual instrument corresponding to the performer's hand and the musical instrument based on the music performance video; and transmitting the three-dimensional virtual space including the three-dimensional virtual hand and the virtual instrument to an educational terminal, wherein the virtual hand of the three-dimensional virtual space moves on the virtual instrument following the hand shape of the performer moving on the musical instrument, and the step of generating the three-dimensional virtual space comprises the steps of: dividing the music performance video into a plurality of music performance image frames; generating a transformed music performance image frame in which a shooting point of a camera that captured the performer and the musical instrument is transformed based on the plurality of music performance image frames; recognizing the hand shape and the musical instrument shape of the performer included in the plurality of transformed music performance image frames, and determining the position of each finger of the performer on the musical instrument based on the recognized hand shape and the musical instrument shape; And, based on the recognized hand shape and instrument shape, a step of rendering a 3D virtual hand corresponding to the player's hand and a 3D virtual instrument corresponding to the instrument, and generating a 3D virtual space including the rendered 3D virtual hand and 3D virtual instrument.

[0010] A virtual reality-based online music teaching method according to one embodiment of the present invention further includes the steps of: obtaining a performance practice video showing a movement of a student playing a musical instrument from the educational terminal; calculating an evaluation score for the student's musical instrument performance based on the performance practice video and the music performance video; generating a performance feedback report including the calculated evaluation score; and transmitting the generated performance feedback report to the educational terminal.

[0011] The step of generating the above-described converted music performance image frame calculates a reference position and a reference shooting angle of a camera that shoots the music performance video based on the shape of the musical instrument included in the plurality of music performance image frames, and generates a converted music performance image frame representing the musical instrument and the performer's hand shot from an arbitrary virtual camera viewpoint for each of the plurality of music performance image frames based on the calculated reference position and reference shooting angle.

[0012] Based on the above-described calculated reference position and reference shooting angle, generating a transformed music performance image frame representing a musical instrument and a performer's hands shot from an arbitrary virtual camera viewpoint for each of the plurality of music performance image frames includes calculating a reference position and a reference shooting angle of a camera that shot the music performance video in an arbitrary virtual space with the center of the musical instrument included in the music performance image frame as a reference point, setting the transformed position and the transformed shooting angle of the virtual camera to a position and shooting angle preset in the virtual space, calculating a transformation function between the reference camera and the virtual camera according to the following mathematical expression 1, and converting and generating the plurality of music performance image frames into the plurality of transformed music performance image frames based on the calculated transformation function.

[0013] (Equation 1)

[0014]

[0015] Here, σ is a correction constant for correcting the distortion of the camera module that captured the video, T is a transformation function, (x0, y0, z0) are coordinates indicating the position of the reference camera, and (x1, y1, z1) are coordinates indicating the position of the virtual camera.

[0016] According to another embodiment of the present invention, a computer-readable recording medium has recorded thereon a program for performing a method according to an embodiment of the present invention.

[0017] This virtual reality-based online music teaching method involves filming a famous musician playing an instrument and recognizing their hand movements from the recorded performance video. Based on the recognized hand movements, the online music teaching method generates a virtual hand in a 3D virtual space that moves identically to the famous musician's hand movements. This hand is then transmitted to an educational terminal and provided to the student. Students can learn how to play the instrument, specifically fingering, by observing and imitating the virtual hand movements in the provided 3D virtual space.

[0018] In addition, the online music teaching method according to the present invention provides a student with a method of playing an instrument through a three-dimensional virtual space, obtains a performance practice video showing the student's instrument playing movements from the student, compares the obtained performance practice video with a music performance video of a famous performer, and calculates an evaluation score for the student's instrument playing, thereby calculating a score for the student's instrument playing. The online music teaching method can determine the student's own instrument playing status by providing the calculated score to the student.

[0019] In addition, the online music teaching method according to the present invention provides students with the most common errors in playing an instrument, thereby providing them with pointers on what to pay attention to when playing. Consequently, students can improve their instrumental skills more quickly.

[0020] Furthermore, the online music teaching method according to the present invention utilizes artificial intelligence to recognize the hand shapes of performers and students, enabling more accurate recognition of hand shapes. When recognizing hand shapes, the method utilizes a greater number of recognition points corresponding to the hands of performers and students compared to conventional hand motion recognition methods, enabling more accurate hand shape recognition.

[0021] In addition, the online music teaching method according to the present invention compares the positions of the fingers of a performer playing an instrument with the positions of the fingers of a student playing an instrument, and when the student judges his or her own musical instrument playing skills, converts a music performance image included in a music performance video in which a performer plays an instrument into the same shooting position and shooting angle as a performance practice video in which a student plays an instrument, recognizes the finger positions of the performer and the student in the converted music performance image and the performance practice image, and compares the positions of the recognized fingers, thereby making it possible to more accurately determine whether the fingerings of the student and the performer match.

[0022] Additionally, the online music teaching method according to the present invention utilizes virtual reality technology to teach musical instrument playing online, thereby enabling students to receive more immersive online musical instrument education.

[0023] FIG. 1 is a diagram illustrating an online music education system according to one embodiment of the present invention.

[0024] Figure 2 is a configuration diagram of the education server shown in Figure 1.

[0025] FIG. 3 is a flowchart of a virtual reality-based online music teaching method with improved recognition rate according to an embodiment of the present invention.

[0026] FIG. 4 is a detailed flowchart of a step for creating a virtual space including a 3D virtual instrument and a 3D virtual hand as shown in FIG. 3.

[0027] Figure 5 is a drawing illustrating a process of converting a music performance image frame into a converted music performance image frame.

[0028] Figure 6 is an example drawing showing the recognition points of the performer's hand.

[0029] Figure 7 is a detailed flowchart of the steps for calculating an evaluation score for a student's musical instrument playing movements illustrated in Figure 3.

[0030] A virtual reality-based online music teaching method according to one embodiment of the present invention comprises the steps of: obtaining a music performance video showing a movement of a performer playing a musical instrument; generating a three-dimensional virtual space including a three-dimensional virtual hand and a virtual instrument corresponding to the performer's hand and the musical instrument based on the music performance video; and transmitting the three-dimensional virtual space including the three-dimensional virtual hand and the virtual instrument to an educational terminal, wherein the virtual hand of the three-dimensional virtual space moves on the virtual instrument following the hand shape of the performer moving on the musical instrument, and the step of generating the three-dimensional virtual space comprises the steps of: dividing the music performance video into a plurality of music performance image frames; generating a transformed music performance image frame in which a shooting point of a camera that captured the performer and the musical instrument is transformed based on the plurality of music performance image frames; recognizing the hand shape and the musical instrument shape of the performer included in the plurality of transformed music performance image frames, and determining the position of each finger of the performer on the musical instrument based on the recognized hand shape and the musical instrument shape; And, based on the recognized hand shape and instrument shape, a step of rendering a 3D virtual hand corresponding to the player's hand and a 3D virtual instrument corresponding to the instrument, and generating a 3D virtual space including the rendered 3D virtual hand and 3D virtual instrument.

[0031] A virtual reality-based online music teaching method according to one embodiment of the present invention further includes the steps of: obtaining a performance practice video showing a movement of a student playing a musical instrument from the educational terminal; calculating an evaluation score for the student's musical instrument performance based on the performance practice video and the music performance video; generating a performance feedback report including the calculated evaluation score; and transmitting the generated performance feedback report to the educational terminal.

[0032] The step of generating the above-described converted music performance image frame calculates a reference position and a reference shooting angle of a camera that shoots the music performance video based on the shape of the musical instrument included in the plurality of music performance image frames, and generates a converted music performance image frame representing the musical instrument and the performer's hand shot from an arbitrary virtual camera viewpoint for each of the plurality of music performance image frames based on the calculated reference position and reference shooting angle.

[0033] Based on the above-described calculated reference position and reference shooting angle, generating a transformed music performance image frame representing a musical instrument and a performer's hands shot from an arbitrary virtual camera viewpoint for each of the plurality of music performance image frames includes calculating a reference position and a reference shooting angle of a camera that shot the music performance video in an arbitrary virtual space with the center of the musical instrument included in the music performance image frame as a reference point, setting the transformed position and the transformed shooting angle of the virtual camera to a position and shooting angle preset in the virtual space, calculating a transformation function between the reference camera and the virtual camera according to the following mathematical expression 1, and converting and generating the plurality of music performance image frames into the plurality of transformed music performance image frames based on the calculated transformation function.

[0034] (Equation 1)

[0035]

[0036] Here, σ is a correction constant for correcting the distortion of the camera module that captured the video, T is a transformation function, (x0, y0, z0) are coordinates indicating the position of the reference camera, and (x1, y1, z1) are coordinates indicating the position of the virtual camera.

[0037] According to another embodiment of the present invention, a computer-readable recording medium has recorded thereon a program for performing a method according to an embodiment of the present invention.

[0038] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the technical idea of ​​the present disclosure is not limited to the embodiments described below and may be implemented in various different forms. The following embodiments are provided only to complete the technical idea of ​​the present disclosure and to fully inform those skilled in the art of the present disclosure of the scope of the present disclosure, and the technical idea of ​​the present disclosure is defined only by the scope of the claims.

[0039] When assigning reference numerals to components in each drawing, it should be noted that identical components are assigned the same numerals whenever possible, even if they appear on different drawings. Furthermore, when describing the present disclosure, if a detailed description of a related known configuration or function is deemed likely to obscure the gist of the present disclosure, such detailed description will be omitted.

[0040] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in the same sense as commonly understood by those of ordinary skill in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. The terminology used herein is for the purpose of describing embodiments and is not intended to limit the disclosure. In this specification, singular forms also include plural forms, unless specifically stated otherwise.

[0041] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.

[0042] As used herein, the terms “comprises” and / or “comprising” do not exclude the presence or addition of one or more other components, steps, operations and / or elements.

[0043] The terms used in the detailed description of the embodiments of the present invention below have the following meanings. “Real world” refers to a physical real space, and “virtual reality” refers to a three-dimensional virtual space that mirrors part or all of the space of real reality. “Real object” refers to a physical object existing in a physical real space, and “virtual object” refers to a non-physical object existing in a virtual space that is a replica of a physical object. “Augmented reality (AR)” is a space where real reality and virtual reality are mixed, and a three-dimensional virtual object is superimposed on an image or background of real reality to provide a single image to the user.

[0044] Components included in one embodiment and components with common functions may be described using the same designations in other embodiments. Unless otherwise stated, the descriptions provided in one embodiment may also apply to other embodiments, and specific descriptions may be omitted to the extent that they overlap or would be readily apparent to those skilled in the art.

[0045] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0046] FIG. 1 is a diagram illustrating an online music education system according to one embodiment of the present invention. Referring to FIG. 1, a music education system (1) according to one embodiment of the present invention includes an education server (11), a performer terminal (12), and an education terminal (13). The education server (11) is a server that generates non-face-to-face music education content and provides non-face-to-face music education services. It transmits and receives data and signals related to online music education from the performer terminal (12) and the education terminal (13). The education server (11) receives data and signals for generating online music education content from the performer terminal (13), and generates music education content including a three-dimensional virtual object based on the data and signals transmitted from the performer terminal (13). The education server (11) transmits the generated music education content to the education terminal (13).

[0047] The performer terminal (12) serves as a performer's terminal, capturing the performer's movements while playing an instrument and utilizing this as basic data for music education content. The performer, who is the user of the performer terminal (12), generates performance data representing the performance of the instrument and transmits the generated performance data to the education server (11).

[0048] The educational terminal (13) is a terminal for students learning to play a musical instrument, and outputs music education content transmitted from the educational server (11). The educational terminal (13) records the student's musical instrument playing movements according to the output music education content, and transmits the recorded video data to the educational server (11).

[0049] Additionally, the education server (11) generates feedback on the student's performance based on video data transmitted from the education terminal (13). The education server (11) then transmits the generated feedback back to the education terminal (13).

[0050] Here, the education server (11) may be a single computer or a collection of multiple computers. The performer terminal (12) and the education terminal (13) are computing devices with communication functions, which can execute various applications and can execute online music lesson applications. Examples of the performer terminal (12) and the education terminal (13) include a head-mounted display (HMD), a VR / AR headset, smartglass, a smartphone, a tablet PC, etc. In addition to the exemplary devices described above, the performer terminal (12) and the education terminal (13) include a display module capable of outputting the generated composite image.

[0051] The steps of constructing an online music teaching method according to an embodiment of the present invention can be performed in the education server (11), performer terminal (12), and education terminal (13) of the music teaching system (1). Below, a method for generating a composite image including a three-dimensional virtual object will be described in detail.

[0052] Figure 2 is a configuration diagram of the education server (11) illustrated in Figure 1. Referring to Figure 2, the education server (11) includes a processor (1101), a communication module (1102), storage (1103), a virtual space creation module (1104), a performance practice analysis module (1105), and a report creation module (1106).

[0053] The processor (1101) of the education server (11) processes general tasks performed on the education server (11).

[0054] The communication module (1102) of the education server (11) can transmit and receive data, messages, and signals to and from a performer terminal (12), an education terminal (13), and an external server (not shown) through a wide area network such as the Internet by connecting to a mobile communication base station or a Wi-Fi repeater.

[0055] The storage (1103) of the education server (11) stores data required to execute an application or program that implements a virtual reality-based online music teaching method. For example, the storage (1103) stores generated virtual reality music education content. In addition, the storage (1103) stores a training set for training at least one artificial neural network included in the education server (11).

[0056] The virtual space creation module (1104) of the education server (11) analyzes a music performance video in which a performer plays an instrument, and creates a 3D virtual space including 3D virtual objects (3D virtual hands and 3D virtual instruments) corresponding to the shape of the performer's hands and the shape of the instrument. The virtual space creation module (1104) divides the music performance video into frames, recognizes the shape of the performer's hands and the shape of the instrument represented by each divided frame, renders a 3D virtual object corresponding to the recognized shape of the hand and the shape of the instrument, and places it on the 3D virtual space, thereby creating a 3D virtual space. A specific method for creating a 3D virtual space including a 3D virtual hand and a virtual instrument in the virtual space creation module (1104) will be described in detail below.

[0057] The performance practice analysis module (1105) of the education server (11) analyzes a performance practice video of a student playing an instrument input from an education terminal (13), and calculates an evaluation score for the student's instrument playing movements. The performance practice analysis module (1105) compares the performance practice video with a music performance video, and calculates an evaluation score for the student's instrument playing movements based on the comparison result. More specifically, the performance practice analysis module (1105) compares the hand shapes of the student and the performer in the performance practice video and the music performance video, and calculates a score for the student's instrument playing movements based on the degree to which the student's hand shapes are similar to the performer's hand shapes. A specific method for calculating an evaluation score for the student's instrument playing movements in the performance practice analysis module (1105) will be described in detail below.

[0058] The report generation module (1106) of the education server (11) generates a performance feedback report including evaluation scores for the performance practice video produced by the performance practice analysis module (1105). The report generation module (1106) generates a performance feedback mock-up to be provided to students who are users of the education terminal (13). The performance feedback report includes scores for the performance practice date and time, the title of the performance practice song, the evaluation score, and the most common moment of error.

[0059] In the education server (11), the virtual space creation module (1104), the performance practice analysis module (1105), and the report creation module (1106) may be implemented by a separate dedicated processor different from the processor (1101), or may be implemented by a computer program running on the processor (1101).

[0060] The education server (11) may include additional components in addition to the components described above. For example, as illustrated in FIG. 2, the education server (11) includes a bus for transmitting data between components, and, although omitted from FIG. 2, further includes a power module for supplying driving power to each component.

[0061] Thus, descriptions of components that are obvious to those skilled in the art to which this embodiment pertains will be omitted as they may obscure the specific features of this embodiment. Below, each component of the education server (11) will be described in detail in the course of explaining a virtual reality-based online music teaching method according to one embodiment of the present invention.

[0062] FIG. 3 is a flowchart of a virtual reality-based online music teaching method according to an embodiment of the present invention. Referring to FIG. 3, in step 301, the education server (11) obtains a music performance video showing the movements of a performer playing an instrument. The communication module (1102) of the education server (11) receives the music performance video from the performance terminal (12). The performance terminal (12) generates a music performance video showing the movements of the performer playing the instrument. The camera module (1204) of the performance terminal (12) photographs the movements of the performer playing the instrument and generates a music performance video showing the movements of the performer playing the instrument. The performance terminal (12) transmits the generated music performance video to the education server (11) through the communication module (1202).

[0063] The education server (11) can obtain a performance video by searching and loading a music performance video stored in storage (1103). The education server (11) inputs the obtained music performance video into the virtual space creation module (1104).

[0064] In step 302, the education server (11) creates a three-dimensional virtual space including a virtual instrument and a virtual hand of a performer based on the acquired music performance video. The virtual space creation module (1104) of the education server (11) analyzes the music performance video acquired in step 301 and creates a three-dimensional virtual space including a three-dimensional virtual hand representing the hand of a performer included in the music performance video. The virtual space creation module (1104) extracts the finger movements of the instrument and the performer included in the music performance video, creates a three-dimensional virtual hand that moves on a three-dimensional virtual instrument according to the extracted instrument and finger movements, and creates a three-dimensional virtual space including the three-dimensional virtual hand that moves on the virtual instrument.

[0065] The steps for creating 3D educational content including 3D virtual hands moving on a virtual instrument will be described below with reference to FIG. 4.

[0066] FIG. 4 is a detailed flowchart of a step of creating a virtual space including a three-dimensional virtual instrument and a three-dimensional virtual hand illustrated in FIG. 3. Referring to FIG. 4, in step 3021, the virtual space creation module (1104) of the education server (11) divides the acquired music performance video into a plurality of music performance image frames. The virtual space creation module (1104) divides the music performance video into a plurality of music performance image frames included in the music performance video. For example, if the music performance video is 24 FPS (Frames Per Second) and the total playback time is 60 seconds, the virtual space creation module (1104) divides the music performance video into 1,440 frames.

[0067] Here, a music performance video is a video depicting the movements of a performer playing an instrument. A music performance image frame is an image depicting the moment the performer plays the instrument. A music performance video is composed of multiple music performance image frames.

[0068] In step 3022, the virtual space generation module (1104) of the education server (11) generates a converted music performance image frame that converts the position and viewpoint of the camera that captured the performer and the musical instrument based on the plurality of segmented music performance image frames. The converted music performance image frame means a converted music performance image frame in which the original music performance image frame representing the musical instrument and the performer's hands captured from the position of the camera in real reality is converted into a frame representing the musical instrument and the performer's hands captured from an arbitrary camera position. In other words, the converted music performance image frame is a music performance image frame in which the viewpoint of the camera that captured the original music performance image frame is converted. The virtual space generation module (1104) generates a music performance image frame in which the position and shooting angle of the camera are converted based on each of the plurality of music performance image frames.

[0069] More specifically, the virtual space generation module (1104) calculates the reference position and reference shooting angle of the camera that shoots the music performance video based on the shape of the instrument included in the multiple segmented music performance image frames. The virtual space generation module (1104) recognizes the instrument in the music performance image frame. For example, if the instrument is a piano, the virtual space generation module (1104) compares the basic shape of the piano keyboard with the shape of the piano keyboard shown by the music performance image frame, and calculates the reference position and reference shooting angle of the camera that shoots the music performance video based on the comparison result. Here, the reference position and the reference shooting angle of the camera mean the position and the direction in which the camera that shoots the music performance video is located, respectively, and mean the position and shooting angle of the camera in the virtual space generated with the center of the instrument included in the music performance image frame as the reference point. The virtual space creation module (1104) calculates the reference position and reference shooting angle of a camera that shoots a music performance video in an arbitrary virtual space based on the center of an instrument included in a music performance image frame as a reference point, and calculates a conversion function between the transformation position and the transformation shooting angle of the virtual camera. Here, the transformation position and the transformation shooting angle of the virtual camera may be any preset position and any shooting angle.

[0070] The virtual space generation module (1104) converts each of a plurality of music performance image frames into a converted music performance image frame representing a musical instrument and a performer's hands, captured from an arbitrary virtual camera viewpoint, based on the calculated reference position and reference shooting angle of the camera. The conversion of the music performance image frame into a converted music performance image frame in the virtual space generation module (1104) will be described with reference to FIG. 5.

[0071] Fig. 5 is a diagram illustrating a process of converting a music performance image frame into a converted music performance image frame. Referring to the example illustrated in Fig. 5, a camera that captures a music performance video in an arbitrary virtual space with the center of the musical instrument as a reference point is converted into a virtual camera positioned in front of the musical instrument in the arbitrary virtual space, and a plurality of music performance image frames included in the music performance video are converted into converted music performance image frames that appear to have been captured by the converted virtual camera.

[0072] More specifically, the virtual space generation module (1104) converts the coordinates (x0, y0, z0) indicating the position of the calculated reference camera into the coordinates (x1, y1, z1) indicating the transformation position of the virtual camera. Here, the virtual camera is a camera positioned to photograph the instrument from the front, and is determined at an arbitrary position in an arbitrary virtual space. The virtual space generation module (1104) generates a transformation function between the reference position of the reference camera and the transformation position of the virtual camera. The transformation function includes a transformation matrix. The relationship among the reference position of the reference camera, the position of the virtual camera, and the transformation function is as shown in the following mathematical expression 1.

[0073] (Equation 1)

[0074]

[0075] Here, σ is a correction constant for correcting the distortion of the camera module that captured the video, and T is a conversion function. The correction constant σ is a constant for correcting the distortion of the camera that captured the music performance video, and is a constant that varies depending on the camera that captured the music performance video and the reference position of the camera.

[0076] The virtual public production module (1104) includes a compensation constant artificial neural network that outputs a compensation constant for a camera from a music performance video captured by a camera. The compensation constant artificial neural network is an artificial neural network that is pre-trained to output a compensation constant for a camera that captured a music performance video from a music performance video. The compensation constant artificial neural network is an artificial neural network that is pre-trained to output a compensation constant for a camera that captured a music performance video from an input music performance video by a training set that includes training videos captured by any camera and training compensation constants for the cameras that captured each training video.

[0077] The virtual space generation module (1104) generates a plurality of transformed music performance image frames by converting a plurality of music performance image frames into a plurality of transformed music performance image frames based on a correction constant and a transformation function.

[0078] In embodiments of the present invention, the virtual space creation module (1104) can set the position of the virtual camera to be identical to the position of the camera module of the player terminal (12) based on the position of the instrument played by the player when taking pictures of the player's hands and instrument using the camera module of the player terminal (12) worn by the player.

[0079] In step 3023, the virtual space creation module (1104) of the education server (11) recognizes the hand shape and the instrument shape of the performer included in the plurality of converted music performance image frames, and determines the position of each finger of the performer on the instrument based on the recognized hand shape and instrument shape. The virtual space creation module (1104) recognizes the instrument and hand shape represented by each of the plurality of converted music performance image frames. More specifically, the virtual space creation module (1104) recognizes and determines the hand shape and the instrument shape represented by each converted music performance image frame.

[0080] For example, if the instrument is a piano, the virtual space generation module (1104) recognizes the keys of the piano represented by each of a plurality of converted music performance image frames. The virtual space generation module (1104) recognizes the keys of the piano included in each music performance image frame. The virtual space generation module (1104) includes a model determination artificial neural network that determines a piano model from a portion representing the piano keys included in the converted music performance image frame. The model determination artificial neural network is an artificial neural network that is pre-trained to output a piano model included in an image representing the piano keys from an image representing the piano keys. The model determination artificial neural network is an artificial neural network that is pre-trained to output a model of the piano included in an input image representing the piano keys by a training set including training images representing the piano keys and the piano models included in each training image. The virtual space generation module (1104) recognizes the piano keys included in the converted music performance image frame based on the determined piano model. Here, the virtual space creation module (1104) recognizes the piano keys included in each converted music performance image frame based on the number, size, arrangement, etc. of white keys and black keys for each piano model stored in the storage (1103).

[0081] In addition, the virtual space generation module (1104) recognizes the shape of the player's hand included in each converted music performance image frame. The virtual space generation module (1104) determines recognition points of the player's hand based on the hand shape represented by each music performance image frame, and recognizes the shape of the player's hand based on the recognition points. The virtual space generation module (1104) further includes a hand recognition artificial neural network that is trained in advance to output a plurality of recognition points corresponding to the shape of the player's hand from a converted music performance image frame representing the shape of the player's hand. The hand recognition artificial neural network is an artificial neural network that is trained in advance to output a plurality of recognition points corresponding to each part of the hand included in an image representing the shape of the hand. The hand recognition artificial neural network is an artificial neural network that is trained in advance to output a model of a piano included in an input image representing a piano keyboard by means of a training set including training images representing the shape of the player's hand and a plurality of recognition points corresponding to the shape of the hand included in each training image.

[0082] The virtual space creation module (1104) determines recognition points representing specific parts of the hand, and recognizes the shape of the performer's hand included in each converted music performance image frame based on the determined recognition points.

[0083] In this regard, FIG. 6 of the present invention is an exemplary drawing showing recognition points of a performer's hand. Referring to FIG. 6, (a) of FIG. 6 is an example showing the number of recognition points corresponding to a palm when recognizing a conventional hand shape, and (b) of FIG. 6 is an example showing the number of recognition points utilized in a hand shape recognition method according to the present invention. The conventional hand shape recognition method exemplified in FIG. 6 (a) generally recognizes a hand shape using 21 recognition points. The hand shape recognition method of the present invention exemplified in FIG. 6 (b) recognizes a hand shape using 63 recognition points.

[0084] The present invention can accurately recognize hand shapes by recognizing the hand shapes of a performer or student using more recognition points compared to conventional hand shape recognition methods.

[0085] The virtual space creation module (1104) determines the position of each finger of the player's hand on the piano keyboard in each converted music performance image frame based on the shape of the piano keyboard and the player's hand recognized from each converted music performance image frame.

[0086] In step 3024, the virtual space generation module (1104) of the education server (11) renders a virtual instrument and a virtual hand based on the recognized hand shape and instrument shape, and generates a three-dimensional virtual space including the virtual instrument and the virtual hand. The virtual space generation module (1104) renders a virtual instrument corresponding to an instrument included in a converted music performance image frame and a virtual hand corresponding to the performer's hand based on the hand shape, instrument shape, and finger positions recognized in step 3023. More specifically, the virtual space generation module (1104) generates a three-dimensional virtual space. The virtual space generation module (1104) places the virtual instrument rendered on the generated three-dimensional virtual space at the center of the three-dimensional virtual space, and places the virtual hand rendered on the placed virtual instrument according to the finger positions determined in step 3023. The virtual space generation module (1104) changes the position and shape of a virtual hand rendered in a three-dimensional virtual space according to the hand shape and instrument shape in each of a plurality of transformed music performance image frames in step 3023. In the three-dimensional virtual space, the virtual hand moves along the shape of the player's hand moving on the virtual instrument. On the virtual instrument, the virtual space generation module (1104) generates a three-dimensional virtual instrument and virtual hand corresponding to the player's performance movements shown in the performance video in the above-described manner.

[0087] The virtual space generation module (1104) distinguishes a plurality of transformed music performance image frames segmented from a music performance video into music performance image frames having the same hand shape and instrument shape. The virtual space generation module (1104) distinguishes a plurality of transformed music performance image frame groups having the same hand shape and instrument shape. The virtual space generation module (1104) determines at least one music performance image frame having the same hand shape and instrument shape as the same transformed music performance image frame group. For example, the virtual space generation module (1104) can distinguish 1,440 music performance image frames into 100 music performance image frame groups. A method for distinguishing a plurality of transformed music performance image frame groups will be described in detail below.

[0088] The virtual space generation module (1104) compares any two converted music performance image frames among a plurality of converted music performance image frames and calculates the similarity between the two converted music performance image frames. The virtual space generation module (1104) includes a similarity artificial neural network that calculates the similarity between the two converted music performance image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output the similarity between the images represented by the two input image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output the similarity between the two input image frames by a training set that includes a plurality of training image frames representing an appearance of a musician playing an instrument and the training similarity between the two training image frames. Here, the similarity is expressed as a numerical score from 0 to 100. When the similarity is 100, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely identical, and when the similarity is 0, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely different.

[0089] The virtual space generation module (1104) determines a converted music performance image frame having a similarity higher than a preset threshold similarity score as a music performance image frame group. Here, the threshold similarity score is a similarity score that serves as a criterion for considering the hand shape and instrument shape represented by the converted music performance image frame as the same, and can be preset by the user. The threshold similarity score is proportional to the resolution and illumination of the music performance video, respectively. When the resolution or illumination of the music performance video is high, the threshold similarity score is set high. Conversely, when the resolution or illumination of the music performance video is low, the threshold similarity score is set low.

[0090] The virtual space generation module (1104) generates a three-dimensional virtual instrument and a virtual hand for each of a plurality of distinct converted music performance image frame groups, and associates the generated three-dimensional virtual instrument and virtual hand with all converted music performance image frames included in the plurality of converted music performance image frame groups.

[0091] The virtual space generation module (1104) recognizes hand shapes and instrument shapes for multiple converted music performance image frames having the same hand shape and instrument shape, and performs the process of rendering a 3D virtual hand and virtual instrument only once based on the recognized hand shape and instrument shape, and associates it with all converted music performance image frames in the same group, thereby omitting the recognition step and rendering step for all converted music performance image frames. Accordingly, a 3D virtual space including a 3D virtual hand and virtual instrument can be generated more quickly.

[0092] The virtual space creation module (1104) inputs a three-dimensional virtual space including the created virtual instrument and virtual hand into the communication module (1102).

[0093] In step 303, the education server (11) transmits the generated three-dimensional virtual space to the education terminal (13). The communication module (1102) of the education server (11) transmits the three-dimensional virtual space including the virtual instrument and virtual hand generated in step 302 to the education terminal (13). Here, the communication module (1102) transmits data representing the three-dimensional virtual space to the education terminal (13) using the communication module (1102).

[0094] An educational terminal (13) that receives data representing a three-dimensional virtual space outputs the three-dimensional virtual space. More specifically, the educational terminal (13) outputs a virtual instrument and a virtual hand on the three-dimensional virtual space through an output module. The educational terminal (13) outputs the virtual instrument and the virtual hand on the three-dimensional virtual space from the viewpoint of a virtual camera at a predetermined position and direction on the three-dimensional virtual space. The educational terminal (13) shows the user of the educational terminal (13), i.e., the student, that the shape and position of the virtual hand on the virtual instrument in the three-dimensional virtual space change in time series. For example, the educational terminal (13) shows the movement of the virtual hand on the virtual piano keyboard in the three-dimensional virtual space to the student of the educational terminal (13). The student learns how to play the instrument by watching the movement of the virtual piano keyboard and the virtual hand on the three-dimensional virtual space output through the educational terminal (13) and imitating the hand shape of the virtual hand.

[0095] The educational terminal (13) generates a performance practice video showing the movements of a student practicing an instrument by following the movements of a virtual hand on a virtual instrument in a three-dimensional virtual space. The camera module of the educational terminal (13) captures the movements of the student practicing the instrument and generates a performance practice video showing the movements of the student practicing the instrument.

[0096] In step 304, the education server (11) obtains a performance practice video from the education terminal (13) in response to the transmission of the 3D virtual space. The communication module (1102) of the education server (11) receives the performance practice video from the education terminal (13). The communication module (1102) of the education server (11) inputs the obtained performance practice video into the performance practice analysis module (1105).

[0097] In step 305, the education server (11) calculates an evaluation score for the student's instrument playing based on the performance practice video and the music performance video. The performance practice analysis module (1105) of the education server (11) compares the performance practice video and the music performance video, and calculates an evaluation score for the student's instrument playing movements based on the comparison result. The performance practice analysis module (1105) compares the student's instrument playing movements in the performance practice video with the player's instrument playing movements in the music performance video, calculates a similarity between the student's instrument playing movements and the player's instrument playing movements based on the comparison result, and calculates an evaluation score for the student's instrument playing movements based on the calculated similarity.

[0098] The steps for calculating the evaluation score for a student's musical instrument playing movements will be described below with reference to Figure 7.

[0099] Fig. 7 is a detailed flowchart of a step for calculating an evaluation score for a student's musical instrument playing movements illustrated in Fig. 3. Referring to Fig. 7, in step 3051, the performance practice analysis module (1105) of the education server (11) divides the acquired performance practice video into a plurality of performance practice image frames. The performance practice analysis module (1105) divides the performance practice video into a plurality of performance practice image frames included in the performance practice video. For example, if the performance practice video is 24 FPS and the total playback time is 60 seconds, the performance practice analysis module (1105) divides the performance practice video into 1,440 frames.

[0100] Here, the performance practice video has the same FPS and playback time as the music performance video. If the FPS and playback time of the performance practice video are different from those of the music performance video, the performance practice analysis module (1105) converts the performance practice video so that the FPS and playback time of the performance practice video are the same as those of the music performance video. The specific process of converting the FPS and playback time is omitted as it obscures the features of the present invention.

[0101] In step 3052, the performance practice analysis module (1105) of the education server calculates an evaluation score for the student's musical instrument performance movements based on a plurality of performance practice image frames and a plurality of converted music performance image frames. The performance practice analysis module (1105) compares the performance practice image frames with the corresponding converted music performance image frames, and calculates an evaluation score for the student's musical instrument performance movements. For example, when comparing 1,440 performance practice image frames and 1,440 converted music performance image frames, the performance practice analysis module (1105) compares the first performance practice image frame with the first converted music performance image frame, compares the second performance practice image frame with the second converted music performance image frame, and compares the nth performance practice image frame with the nth converted music performance image frame.

[0102] The performance practice analysis module (1105) according to an embodiment of the present invention recognizes the hand shape and the instrument shape of the student included in each of a plurality of performance practice image frames, and determines the position of each finger of the student on the instrument based on the recognized hand shape and instrument shape. The performance practice analysis module (1105) recognizes the instrument and hand shape represented by each of the plurality of performance practice image frames. More specifically, the performance practice analysis module (1105) recognizes and determines the hand shape and the instrument shape represented by each performance practice image frame.

[0103] For example, if the instrument is a piano, the performance practice analysis module (1105) recognizes the piano keys represented by each of a plurality of performance practice image frames. The performance practice analysis module (1105) recognizes the piano keys included in each performance practice image frame. The performance practice analysis module (1105) includes a model determination artificial neural network that determines a piano model from a portion representing the piano keys included in the performance practice image frames. The model determination artificial neural network is an artificial neural network that is pre-trained to output a piano model included in an image representing a piano keyboard from an image representing the piano keyboard. The model determination artificial neural network is an artificial neural network that is pre-trained to output a model of a piano included in an input image representing the piano keyboard by a training set that includes training images representing the piano keyboard and the piano models included in each training image. The performance practice analysis module (1105) recognizes the piano keys included in the performance practice image frames based on the determined piano models. Here, the performance practice analysis module (1105) recognizes the piano keys included in each performance practice image frame based on the number, size, arrangement, etc. of white keys and black keys for each piano model stored in the storage (1103).

[0104] In addition, the performance practice analysis module (1105) recognizes the hand shape of the performer included in each performance practice image frame. The performance practice analysis module (1105) determines recognition points of the student's hand based on the hand shape indicated by each performance practice image frame, and recognizes the hand shape of the student based on the recognition points. The performance practice analysis module (1105) determines recognition points indicating the position of each hand, and recognizes the hand shape of the student included in each music performance image frame based on the determined recognition points.

[0105] The performance practice analysis module (1105) determines the position of each finger of the player's hand on the piano keys in each performance practice image frame based on the piano keys and the player's hand shape recognized from each performance practice image frame. The performance practice analysis module (1105) determines the position of each finger of the student on the instrument based on the hand shape and the instrument shape recognized from each performance practice image frame.

[0106] The performance practice analysis module (1105) compares the positions of the student's fingers on the instrument with the positions of the player's fingers on the instrument, and determines the number of the student's fingers that match the positions of the player's fingers. The performance practice analysis module (1105) compares the positions of each of the student's fingers with the positions of each of the player's fingers, and determines the number of fingers that are positioned at the same position. For example, the performance practice analysis module (1105) determines the number as 10 if the positions of the student's fingers and the player's fingers all match, determines the number as 8 if the positions of 8 fingers out of 10 match, and determines the number as 0 if the positions of all the fingers do not match.

[0107] The performance practice analysis module (1105) compares the finger positions of the student and the performer in each of the plurality of performance practice image frames and each corresponding converted music performance image frame, and determines the number of fingers whose positions match. For example, when comparing 1,440 performance practice image frames and 1,440 converted music performance image frames, the performance practice analysis module (1105) determines the number of fingers whose positions match in the first performance practice image frame and the first converted music performance image frame, determines the number of fingers whose positions match in the second performance practice image frame and the second converted music performance image frame, and determines the number of fingers whose positions match in the nth performance practice image frame and the nth converted music performance image frame. The performance practice analysis module (1105) determines the number of the student's fingers whose positions match in all of the performance practice image frames with the performer's finger positions.

[0108] The performance practice analysis module (1105) calculates an evaluation score for the student's instrument performance based on the number of fingers that match the determined positions. More specifically, the performance practice analysis module (1105) calculates an evaluation score for the student's instrument performance based on the number of multiple performance practice frames and the determined number of matching fingers. The performance practice analysis module (1105) calculates an evaluation score based on the finger matching rate, which is the ratio of the number of finger positions of the student in the entire performance practice frames to the total number of fingers that match the finger positions of the performer. For example, if there are 1,440 performance practice frames, the number of finger positions of the student in the entire performance practice frames is 1,440 × 10 (the number of fingers on both hands) = 14,440. Here, if the number of the student's fingers that match the finger positions of the performer is determined to be 9,780, the finger matching rate is The performance practice analysis module (1105) determines the calculated finger coincidence rate as an evaluation score. In the example described above, the performance practice analysis module (1105) calculates the calculated finger coincidence rate of 67.91 as an evaluation score for the student's musical instrument playing movements.

[0109] A performance practice analysis module (1105) according to another embodiment of the present invention compares a performance practice image frame with a corresponding transformed music performance image frame, and calculates a similarity between the performance practice image frame and the corresponding transformed music performance image frame. The performance practice analysis module (1105) includes a similarity artificial neural network that calculates a similarity between a performance practice image frame and a corresponding transformed music performance image frame. The similarity artificial neural network is an artificial neural network that is pre-trained to output a similarity between images represented by two input image frames. The similarity artificial neural network is an artificial neural network that is pre-trained to output a similarity between two input image frames by a training set that includes a plurality of training image frames representing an appearance of a performer playing a musical instrument and a training similarity between the two training image frames. Here, the similarity is expressed as a numerical score from 0 to 100. When the similarity is 100, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely identical, and when the similarity is 0, it means that the shape of the player's hand and the shape of the instrument expressed by the two input image frames are completely different.

[0110] The performance practice analysis module (1105) calculates the similarity between each of a plurality of performance practice image frames and the corresponding converted music performance image frame. For example, when comparing 1,440 performance practice image frames and 1,440 music performance image frames, the performance practice analysis module (1105) calculates the similarity between the first performance practice image frame and the first converted music performance image frame, calculates the similarity between the second performance practice image frame and the second converted music performance image frame, and calculates the similarity between the nth performance practice image frame and the nth converted music performance image frame.

[0111] The performance practice analysis module (1105) calculates an evaluation score based on the similarity calculated from all performance practice image frames. The performance practice analysis module (1105) calculates the average value of the overall similarity calculated from all performance practice image frames as the evaluation score.

[0112] In step 306, the education server (11) determines the moment at which the performer makes the most mistakes in his / her performance based on the performance practice video and the music performance video. The performance practice analysis module (1105) of the education server (11) compares each of the plurality of performance practice image frames with the corresponding converted music performance image frame, and, based on the comparison result, determines the performance practice image frame that is most different from the corresponding music performance image frame among the plurality of performance practice image frames of the performance practice video.

[0113] According to one embodiment of the present invention, the performance practice analysis module (1105) determines a performance practice image frame having the smallest number of fingers where the student's finger positions on the instrument of the performance practice image frame and the performer's finger positions on the instrument of the converted music performance image frame match, and determines the determined performance practice image frame as the moment with the most mistakes.

[0114] According to another embodiment of the present invention, the performance practice analysis module (1105) determines a performance practice image frame having the lowest similarity between each performance practice image frame and a corresponding converted music performance image frame among a plurality of performance practice image frames, and determines the determined performance practice image frame as the moment with the most mistakes.

[0115] The performance practice analysis module (1105) inputs the evaluation score calculated in step 305 and the most frequently wrong moment determined in step 306 into the report generation module (1106).

[0116] In step 307, the training server (11) generates a performance feedback report including the calculated evaluation scores and the determined most frequently erroneous moments. The report generation module (1106) of the training server (11) generates a performance feedback report including the calculated evaluation scores in step 305 and the determined most frequently erroneous moments in step 306. The report generation module (1106) inputs the generated performance feedback report to the communication module (1102).

[0117] At step 308, the training server (11) transmits a performance feedback report to the training terminal (13). The communication module (1102) of the training server (11) transmits the generated performance feedback report to the training terminal (13).

[0118] The virtual reality-based online music teaching method according to the embodiments of the present invention described above films a famous musician playing an instrument, and recognizes the hand movements of the famous musician playing the instrument from the filmed performance video. Based on the recognized hand movements, the online music teaching method generates a virtual hand that moves identically to the hand movements of the famous musician in a three-dimensional virtual space, and transmits this to an educational terminal to provide it to a student. The student can learn how to play the instrument, i.e., fingering, by observing the movements of the virtual hand in the provided three-dimensional virtual space and imitating the movements of the virtual hand.

[0119] In addition, the online music teaching method according to the present invention provides a student with a method of playing an instrument through a three-dimensional virtual space, obtains a performance practice video showing the student's instrument playing movements from the student, compares the obtained performance practice video with a music performance video of a famous performer, and calculates an evaluation score for the student's instrument playing, thereby calculating a score for the student's instrument playing. The online music teaching method can determine the student's own instrument playing status by providing the calculated score to the student.

[0120] In addition, the online music teaching method according to the present invention provides students with the most common errors in playing an instrument, thereby providing them with pointers on what to pay attention to when playing. Consequently, students can improve their instrumental skills more quickly.

[0121] Furthermore, the online music teaching method according to the present invention utilizes artificial intelligence to recognize the hand shapes of performers and students, enabling more accurate recognition of hand shapes. When recognizing hand shapes, the method utilizes a greater number of recognition points corresponding to the hands of performers and students compared to conventional hand motion recognition methods, enabling more accurate hand shape recognition.

[0122] In addition, the online music teaching method according to the present invention compares the positions of the fingers of a performer playing an instrument with the positions of the fingers of a student playing an instrument, and when the student judges his or her own musical instrument playing skills, converts a music performance image included in a music performance video in which a performer plays an instrument into the same shooting position and shooting angle as a performance practice video in which a student plays an instrument, recognizes the finger positions of the performer and the student in the converted music performance image and the performance practice image, and compares the positions of the recognized fingers, thereby making it possible to more accurately determine whether the fingerings of the student and the performer match.

[0123] Additionally, the online music teaching method according to the present invention utilizes virtual reality technology to teach musical instrument playing online, thereby enabling students to receive more immersive online musical instrument education.

[0124] Meanwhile, the embodiments of the present invention described above can be written as a program that can be executed on a computer, and can be implemented in a general-purpose digital computer that runs the program using a computer-readable recording medium. In addition, the structure of the data used in the embodiments of the present invention described above can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as a magnetic storage medium (e.g., a ROM floppy disk, a hard disk, etc.) and an optical reading medium (e.g., a CD-ROM, a DVD, etc.). A program that performs a virtual reality-based online music teaching method according to the embodiments of the present invention is recorded on the computer-readable recording medium.

[0125] The present invention has been described with a focus on preferred embodiments. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

[0126] 11: Image synthesis server

[0127] 1101: Processor 1102: Communication Module

[0128] 1103: Storage 1104: Depth map generation module

[0129] 1105: Virtual plane determination module 1106: Background synthesis module

[0130] 1107: 3D virtual object synthesis module

[0131] 12: User terminal

Claims

1. In the virtual reality-based online music teaching method, A step of obtaining a music performance video showing the movements of a performer playing a musical instrument; A step of generating a three-dimensional virtual space including three-dimensional virtual hands and virtual instruments corresponding to the performer's hands and instruments based on the above music performance video; and A step of transmitting a three-dimensional virtual space including the three-dimensional virtual hand and virtual instrument to an educational terminal, The virtual hand in the above 3D virtual space moves along the shape of the player's hand moving on the instrument on the virtual instrument, The step of creating the above 3D virtual space is A step of dividing the above music performance video into multiple music performance image frames; A step of generating a transformed music performance image frame in which the shooting time of the camera that captured the performer and the instrument is converted based on the plurality of music performance image frames; A step of recognizing the hand shape and the instrument shape of the performer included in a plurality of transformed music performance image frames, and determining the position of each finger of the performer on the instrument based on the recognized hand shape and the instrument shape; and An online music teaching method characterized by comprising the step of rendering a three-dimensional virtual hand corresponding to the player's hand and a three-dimensional virtual instrument corresponding to the instrument based on the recognized hand shape and instrument shape, and generating a three-dimensional virtual space including the rendered three-dimensional virtual hand and three-dimensional virtual instrument.

2. In paragraph 1, A step of obtaining a performance practice video showing the movements of a student playing a musical instrument from the above educational terminal; A step of calculating an evaluation score for the student's musical instrument performance based on the above performance practice video and the above music performance video; A step of generating a performance feedback report including the above-mentioned calculated evaluation scores; and An online music teaching method further comprising a step of transmitting the generated performance feedback report to the educational terminal.

3. In paragraph 1, The step of generating the above-mentioned converted music performance image frame is Based on the shape of the musical instrument included in the plurality of music performance image frames, the reference position and reference shooting angle of the camera that shoots the music performance video are calculated, An online music teaching method characterized in that, based on the above-described calculated reference position and reference shooting angle, a converted music performance image frame representing a musical instrument and a performer's hands is generated from each of the plurality of music performance image frames shot from an arbitrary virtual camera viewpoint.

4. In paragraph 3, Based on the above-described calculated reference position and reference shooting angle, generating a transformed music performance image frame representing a musical instrument and a performer's hands shot from an arbitrary virtual camera viewpoint for each of the plurality of music performance image frames. Calculate the reference position and reference shooting angle of the camera that filmed the music performance video in an arbitrary virtual space with the center of the instrument included in the music performance image frame as the reference point, Set the transformation position and transformation shooting angle of the above virtual camera to a preset position and shooting angle in the virtual space, The conversion function between the reference camera and the virtual camera is calculated according to the following mathematical expression 1, An online music teaching method characterized in that, based on the above-described conversion function, the plurality of music performance image frames are converted and generated into the plurality of converted music performance image frames. (Equation 1) Here, σ is a correction constant for correcting the distortion of the camera module that captured the video, T is a transformation function, (x0, y0, z0) are coordinates indicating the position of the reference camera, and (x1, y1, z1) are coordinates indicating the position of the virtual camera.

5. A computer-readable recording medium having recorded thereon a program for performing the method described in any one of paragraphs 1 to 4.

Citation Information

Patent Citations

  • Musical instrument practicing device

    JP2019053170A

  • System and method for virtual fitting based on augument reality

    KR102340904B1

  • Locking device that can remove the lock from the door in case of emergency release

    KR102503421B1

  • Online music teaching method and apparatus based on virtual reality

    KR102622163B1