A textbook interaction system based on AR augmented reality and pattern recognition

CN113126761BActive Publication Date: 2026-09-11HANGZHOU HANGGANG CHICHENG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110404405.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-15
Publication Date
2026-09-11
Estimated Expiration
2041-04-15

AI Technical Summary

Technical Problem

[0003]钢琴初学者的练习需要以钢琴为基础,并需要老师在旁进行指导,需要耗费的成本比较大,若自己进行练习无法指出不足并进行改正,学习效率不高

Benefits of technology

[0024] In this invention, the acquired musical score is converted into a corresponding note code sequence table, which is then converted into key gestures and compared with the target gestures. This enables a user-friendly, simplified piano learning experience, eliminating the need for a teacher and making it more beginner-friendly while saving costs. The invention also integrates the target gestures with the simulated physical piano, enhancing the simulation effect and providing an immersive experience. Furthermore, a practice mode guides practice, and once a certain level of proficiency is achieved, a performance mode allows for free playing, enabling progressive teaching and improving learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113126761B_ABST
    Figure CN113126761B_ABST
Patent Text Reader

Abstract

The application discloses a teaching material interaction system based on AR (Augmented Reality) and graphic recognition, a mixed reality module displays virtual images and / or real images; an image acquisition module acquires target gestures and music staves, the target gestures include first gestures and second gestures, a processing module identifies and matches the target gestures with the first gestures, if the matching is successful, a note code sequence list corresponding to the music staves is acquired, and the key positions and key times of a simulated real piano are determined based on the note codes; in an exercise mode, the mixed reality module displays virtual images and real images, the virtual images are the simulated real piano and the key positions of the current key times, the real images are the second gestures, and an audio output module outputs corresponding audio; in a performance mode, the mixed reality module displays virtual images and real images, the virtual images are only the simulated real piano, the real images are the second gestures, and an audio acquisition module acquires corresponding audio. Progressive teaching is realized, and learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interactive textbook systems, specifically relating to an interactive textbook system based on augmented reality (AR) and image recognition. Background Technology

[0002] In terms of a textbook, without a teacher's guidance, learners need to read and understand it repeatedly in order to master the knowledge. This is especially true for beginners in musical instruments. After learning the basics of music, they need to practice repeatedly before they can play songs, such as piano beginners.

[0003] Piano beginners need to practice with a piano as a foundation and require guidance from a teacher, which is quite costly. If they practice on their own, they cannot identify and correct their shortcomings, resulting in low learning efficiency. Summary of the Invention

[0004] The purpose of this invention is to provide a textbook interaction system based on AR (Augmented Reality) and image recognition to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] An interactive textbook system based on AR (Augmented Reality) and image recognition is characterized by comprising a mixed reality module, an image acquisition module, an audio output module, an audio acquisition module, a data storage module, and a processing module.

[0007] The mixed reality module is used to display virtual images and / or real images, wherein the virtual image is a simulated physical piano and the real image is an image captured by the image acquisition module;

[0008] The image acquisition module acquires images of target gestures and musical scores. The target gestures include a first gesture and a second gesture. The first gesture is used to scan the musical score, and the second gesture is applied to the keys of a simulated physical piano.

[0009] The processing module is used to identify and match the target gesture with the first gesture, and after a successful match, to perform deep learning recognition on the musical score acquired by the image acquisition module to obtain the corresponding note code sequence table; and to determine the key positions and key times of the simulated physical piano based on the note codes.

[0010] The audio output module is used to play the audio of musical notation through a note encoding sequence table;

[0011] The audio acquisition module is used to acquire the target audio;

[0012] The data storage module is used to store musical scores and corresponding note encoding sequence tables;

[0013] The interactive teaching material system includes a practice mode and a performance mode. In the practice mode, the mixed reality module displays virtual and real images. The virtual image is a simulated physical piano and the key position at the current key press time. The real image is a second gesture. The audio output module outputs the corresponding audio.

[0014] In performance mode, the mixed reality module displays virtual and real images. The virtual image is only a simulation of a physical piano, and the real image is a second gesture. The audio acquisition module acquires the corresponding audio.

[0015] Preferably, the note encoding sequence table includes several note codes, and any note code is an 8-bit binary code {A7,A6,A5,…,A0}, where A7 is used to represent the treble mark, A6A5 represents the number of underscores, A4 is used to represent the bass mark, and A3A2A1A0 is used to represent the note primitive.

[0016] Preferably, in the practice mode, the key position at the current key press time is marked with a first mark, and the key position at the next key press time is marked with a second mark.

[0017] Preferably, the system further includes a first discrimination module, used to discriminate whether the key position at the current key press time is consistent with the second gesture in practice mode. If they are inconsistent, the error count is incremented by 1; otherwise, it is not incremented by 1, and the error rate of the corresponding musical score is calculated. If the error rate is greater than the preset value, the performance mode cannot be entered.

[0018] Preferably, in the practice mode, if the key position of the current key press time is inconsistent with the second gesture, the timing of the current key press time is stopped and the audio is not played until the second gesture is consistent with the key position of the current key press time.

[0019] Preferably, the system further includes an audio-visual integration module and a network transmission module. In performance mode, the audio-visual integration module combines the audio from the audio acquisition module with the virtual and displayed images from the mixed reality module based on a time series to obtain video footage, which is then sent to a third party via the network transmission module.

[0020] Preferably, the target gesture further includes a third gesture, which is used to select a musical score fragment. The musical score fragment is at least one measure in the musical score. The musical score is detected by a target detection algorithm to obtain the target boxes of the musical score measures and number them. Each measure in the musical score fragment is matched with the musical score to determine the target number of any measure in the musical score fragment. Several corresponding note codes are determined by the target number and the note code sequence to obtain the note code sub-sequence. The key positions and key times of the simulated physical piano are determined by the note code sub-sequence, and audio playback is achieved.

[0021] The textbook interaction system based on AR augmented reality and image recognition as described in claim 1 is characterized in that the system further includes a scoring module for calculating the performance score in performance mode, including: determining whether the second gesture and the key position are consistent; if they are inconsistent, the number of key errors is incremented by 1, and the number of time errors is incremented by 1; otherwise, determining whether the duration of the second gesture is consistent with the key press time of the key position; if so, the number of time errors is incremented by 1, and calculating the key error rate and the time error rate respectively. The performance score is a weighted sum of key error rate and timing error rate.

[0022] Preferably, for measures in which key press errors occur in performance mode, a note code subsequence is extracted from the note code sequence and stored in a high-frequency practice database for single or multiple practice sessions in practice mode.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0024] In this invention, the acquired musical score is converted into a corresponding note code sequence table, which is then converted into key gestures and compared with the target gestures. This enables a user-friendly, simplified piano learning experience, eliminating the need for a teacher and making it more beginner-friendly while saving costs. The invention also integrates the target gestures with the simulated physical piano, enhancing the simulation effect and providing an immersive experience. Furthermore, a practice mode guides practice, and once a certain level of proficiency is achieved, a performance mode allows for free playing, enabling progressive teaching and improving learning efficiency. Attached Figure Description

[0025] Figure 1 This is a block diagram of the present invention.

[0026] Figure 2 This is a diagram of the basic encoding of musical notes. Detailed Implementation

[0027] The technical solution of the present invention will be further explained below with reference to the accompanying drawings and specific embodiments. The scope of protection of the present invention includes, but is not limited to, the contents described in the specific embodiments.

[0028] A textbook interaction system based on AR (Augmented Reality) and image recognition includes a mixed reality module, an image acquisition module, an audio output module, an audio acquisition module, a data storage module, and a processing module.

[0029] The mixed reality module is used to display virtual images and / or real images, wherein the virtual image is a simulated physical piano and the real image is an image captured by the image acquisition module.

[0030] The image acquisition module acquires images of target gestures and musical scores. The target gestures include a first gesture and a second gesture. The first gesture is used to scan the musical score, and the second gesture is applied to the keys of a simulated physical piano.

[0031] The mixed reality module here overlays the virtual world onto the real world on the screen and allows for interaction. In other words, it overlays a simulated physical piano onto a real image for interaction. Real target gestures and simulated physical pianos are superimposed on the same screen or space in real time and exist simultaneously.

[0032] The target gesture is a three-dimensional gesture. The "second gesture function and simulate the keys of a real piano" means that when the finger is pressed down, the corresponding key of the simulated real piano is pressed down until the finger is lifted and the corresponding key of the simulated real piano is reset.

[0033] The processing module is used to identify and match the target gesture with the first gesture, and after successful matching, to perform deep learning recognition on the musical score acquired by the image acquisition module to obtain the corresponding note code sequence table; and to determine the key positions and key times of the simulated physical piano based on the note codes.

[0034] In this invention, after the image acquisition module acquires the target gesture, it synchronizes it to the processing module. The processing module compares the target gesture with a first gesture. If the comparison is successful, the target gesture is considered to be the first gesture, and the music score needs to be acquired. In this invention, the acquired music score needs to be preprocessed to form a preprocessed image as a new music score. This music score removes irrelevant information such as lyrics and watermarks and undergoes rotation correction. After deep learning recognition of the music score, the corresponding note encoding sequence table is obtained. This is common knowledge in the field, and those skilled in the art can set it according to the actual situation. The resulting note encoding sequence table contains several note codes. Each note code is an 8-bit binary code {A7, A6, A5, ..., A0}. A7 represents the treble clef; A7 = 1 if there is a treble clef above the note, otherwise A7 = 0. A6 and A5 represent underscores; A6 and A5 = 01 if there is one underscore below the note, 10 if there are two underscores, and 11 if there are three underscores. A4 represents the bass clef; A4 = 1 if there is a bass clef below the note, otherwise A4 = 0. A3, A2, A1, and A0 represent note primitives. There are 14 types of note primitives, such as... Figure 2 As shown.

[0035] The audio output module is used to play the audio of musical notation using a note encoding sequence list. In this invention, the audio output module converts the digital signal—the note encoding list—into an analog signal to achieve audio playback of the musical notation. In this invention, the playback duration of a note can be automatically calculated based on the encoding queue; this is common knowledge in the field, and those skilled in the art can set it according to actual conditions.

[0036] The audio acquisition module is used to acquire the target audio.

[0037] The data storage module is used to store musical scores and corresponding note code sequence lists. In this invention, by setting up a data storage module, repeated recognition of musical scores is avoided, and the musical scores and corresponding note code lists can be used as training data for deep learning recognition to optimize the deep learning recognition model.

[0038] The interactive teaching material system includes a practice mode and a performance mode. In the practice mode, the mixed reality module displays virtual and real images. The virtual image is a simulated physical piano and the key position at the current key press time. The real image is a second gesture. The audio output module outputs the corresponding audio.

[0039] In performance mode, the mixed reality module displays virtual and real images. The virtual image is only a simulation of a physical piano, and the real image is a second gesture. The audio acquisition module acquires the corresponding audio.

[0040] In this invention, the acquired musical score is converted into a corresponding note code sequence table, which is then converted into key gestures and compared with the target gestures. This enables a user-friendly, simplified piano learning experience, eliminating the need for a teacher and making it more beginner-friendly while saving costs. The invention also integrates the target gestures with the simulated physical piano, enhancing the simulation effect and providing an immersive experience. Furthermore, a practice mode guides practice, and once a certain level of proficiency is achieved, a performance mode allows for free playing, enabling progressive teaching and improving learning efficiency.

[0041] In the practice mode, the key position at the current key press time is marked as a first mark, and the key position at the next key press time is marked as a second mark.

[0042] In this invention, by setting a second mark to enhance the learner's awareness of effectiveness, the learner's ability to memorize musical scores is gradually improved, and good piano playing habits are cultivated in the learner.

[0043] The system also includes a first discrimination module, used to discriminate whether the key position at the current key press time is consistent with the second gesture in practice mode. If they are inconsistent, the error count is incremented by 1; otherwise, it is not incremented by 1. The error rate of the corresponding musical score is then calculated. If the error rate is greater than the preset value, the performance mode cannot be entered.

[0044] In this invention, by setting a first discrimination module and a corresponding preset value, the player cannot enter the performance mode when their proficiency does not meet the requirements, forcing the player to repeat the practice without needing someone to supervise them.

[0045] In the practice mode, if the key position of the current key press time is inconsistent with the second gesture, the timing of the current key press time will stop and the audio will not be played until the second gesture is consistent with the key position of the current key press time.

[0046] In this invention, the method is used to correct the key presses of the trainee.

[0047] The system also includes an audio-visual integration module and a network transmission module. In performance mode, the audio-visual integration module combines the audio from the audio acquisition module with the virtual and displayed images from the mixed reality module based on a time series to obtain video footage, which is then sent to a third party via the network transmission module.

[0048] Sending the video footage to a third party, typically an educational institution or teacher, allows them to directly identify the problem and provide online guidance without the need for face-to-face instruction.

[0049] The system also includes a scoring module for calculating the performance score in performance mode. This includes: determining whether the second gesture matches the key position; if not, incrementing the key error count and timing error count by 1; otherwise, determining whether the duration of the second gesture matches the key press time; if so, incrementing the timing error count by 1. The system then calculates the key error rate and timing error rate separately. The performance score is a weighted sum of key error rate and timing error rate. The weights are preset, and the sum of the two weights equals 1. This scoring module allows learners to gain a clear understanding of their playing ability, facilitating subsequent practice.

[0050] For measures where errors occur in performance mode, note code subsequences are extracted from the note code sequence and stored in the high-frequency practice database for single or multiple practice sessions in practice mode. These errors include key press errors and timing errors.

[0051] The target gesture also includes a third gesture, which is used to select a musical score fragment. The musical score fragment is at least one measure in the musical score. The musical score is detected by a target detection algorithm to obtain the target boxes of the musical score measures and number them. Each measure in the musical score fragment is matched with the musical score to determine the target number of any measure in the musical score fragment. Several corresponding note codes are determined by the target number and the note code sequence table to obtain the note code sub-sequence. The key positions and key times of the simulated physical piano are determined by the note code sub-sequence, and audio playback is realized.

[0052] In this invention, a musical score segment is a series of consecutive measures in a musical score. The target detection algorithm is the YOLOv2 algorithm. The target boxes of each measure are numbered sequentially from left to right. The target number of the target box corresponding to any measure in the musical score segment is determined by matching. Since there are note codes for representing bar lines in the note code sequence table, the note codes of the bar lines are counted from top to bottom to obtain the note code of the bar line corresponding to the target number. The note codes between the note code of the current bar line and the note code of the previous bar line are the note code subsequence.

Claims

1. A textbook interactive system based on AR (Augmented Reality) and image recognition, characterized in that, It includes a mixed reality module, an image acquisition module, an audio output module, an audio acquisition module, a data storage module, and a processing module. The mixed reality module is used to display virtual images and / or real images, wherein the virtual image is a simulated physical piano and the real image is an image captured by the image acquisition module; The image acquisition module acquires images of target gestures and musical scores. The target gestures include a first gesture and a second gesture. The first gesture is used to scan the musical score, and the second gesture is applied to the keys of a simulated physical piano. The processing module is used to identify and match the target gesture with the first gesture, and after a successful match, to perform deep learning recognition on the musical score acquired by the image acquisition module to obtain the corresponding note code sequence table; and to determine the key positions and key times of the simulated physical piano based on the note codes. The audio output module is used to play the audio of musical notation through a note encoding sequence table; The audio acquisition module is used to acquire the target audio; The data storage module is used to store musical scores and corresponding note encoding sequence tables; The interactive teaching material system includes a practice mode and a performance mode. In the practice mode, the mixed reality module displays virtual and real images. The virtual image is a simulated physical piano and the key position at the current key press time. The real image is a second gesture. The audio output module outputs the corresponding audio. In performance mode, the mixed reality module displays virtual and real images. The virtual image is only a simulation of a physical piano, and the real image is a second gesture. The audio acquisition module acquires the corresponding audio.

2. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, The note encoding sequence table includes several note codes. Each note code is an 8-bit binary code {A7,A6,A5,…,A0}, where A7 is used to represent the treble mark, A6A5 represents the number of underscores, A4 is used to represent the bass mark, and A3A2A1A0 is used to represent the note primitive.

3. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, In the practice mode, the key position at the current key press time is marked as a first mark, and the key position at the next key press time is marked as a second mark.

4. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, The system also includes a first discrimination module, used to discriminate whether the key position at the current key press time is consistent with the second gesture in practice mode. If they are inconsistent, the error count is incremented by 1; otherwise, it is not incremented by 1. The error rate of the corresponding musical score is then calculated. If the error rate is greater than the preset value, the performance mode cannot be entered.

5. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, In the practice mode, if the key position of the current key press time is inconsistent with the second gesture, the timing of the current key press time will stop and the audio will not be played until the second gesture is consistent with the key position of the current key press time.

6. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, The system also includes an audio-visual integration module and a network transmission module. In performance mode, the audio-visual integration module combines the audio from the audio acquisition module with the virtual and displayed images from the mixed reality module based on a time series to obtain video footage, which is then sent to a third party via the network transmission module.

7. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, The target gesture also includes a third gesture, which is used to select a musical score fragment. The musical score fragment is at least one measure in the musical score. The musical score is detected by a target detection algorithm to obtain the target boxes of the musical score measures and number them. Each measure in the musical score fragment is matched with the musical score to determine the target number of any measure in the musical score fragment. Several corresponding note codes are determined by the target number and the note code sequence to obtain the note code sub-sequence. The key positions and key times of the simulated physical piano are determined by the note code sub-sequence, and audio playback is realized.

8. The textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, The system also includes a scoring module for calculating the performance score in performance mode. This includes: determining whether the second gesture matches the key position; if not, incrementing the key error count and timing error count by 1; otherwise, determining whether the duration of the second gesture matches the key press time; if so, incrementing the timing error count by 1. The system then calculates the key error rate and timing error rate separately. The performance score is a weighted sum of key error rate and timing error rate.

9. A textbook interaction system based on AR (Augmented Reality) and image recognition as described in claim 1, characterized in that, For measures where key press errors occur in performance mode, the note code subsequence is extracted from the note code sequence and stored in the high-frequency practice database for single or multiple practice sessions in practice mode.

Citation Information

Patent Citations

  • Augmented reality based piano teaching method and device and piano

    CN108205946A

  • Methods and systems for representing a pre-modeled object within virtual reality data

    US20190378333A1