Audio and music score matching method and system
Through the dual matching mechanism and the audio and music score matching method of time adjustment, users' self-judgment problems and automatic page turn error problems in piano practice are solved, and the matching accuracy and user experience are improved, especially suitable for beginners.
Patent Information
- Application Number
- CN202510279933.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2025-03-10
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the prior art, it is difficult for users to independently judge the difference between the performance effect and the demonstration performance during piano practice. The automatic page turning function is prone to errors or misplay among beginners, affecting performance coherence, and the existing matching methods have defects in computing reliability and real-timeness.
The double matching mechanism is adopted, combining real-time matching and test time switching methods, the similarity between the note data and the expected performance characteristics is judged through matching rules, dynamically adjust the matching order, and optimize the matching process between the audio and the score, including real-time single note matching and tentative multinote data sequence matching, and adjust the test duration based on the user's wrong performance frequency.
It improves the accuracy and real-timeness of audio and music score matching, reduces the amount of calculation, and improves the user experience. Especially when beginners have high errors, they reduce misjudgment through delay matching to ensure the consistency and accuracy of positioning marks.
Smart Images

Figure CN120508674A_ABST
Abstract
Description
[0001] Priority application This application claims priority to Chinese invention patent application [2025102324395] "[A method and system for matching audio and music scores]" filed on February 27, 2025, which is incorporated by reference in its entirety. Technical Field
[0002] The present invention relates to the technical field of real-time music score matching, and in particular to a method and system for matching audio and music scores. Background Art
[0003] Piano playing is a comprehensive activity that combines logical and visual thinking, mental and physical effort, and technical and artistic skills. Often, when practicing piano, especially for beginners, it's difficult to judge the accuracy of their practice. Learning an instrument is a muscle movement, and if practiced for a long time without professional guidance, mistakes, once established, can be difficult to correct.
[0004] Currently, several intelligent piano learning solutions have been proposed. For example, patent application CN114420074A discloses a method for synchronizing music scores with audio and computer-readable storage. This method can use a pre-acquired complete audio track (such as a piano recording of a master or demonstration performance) as a piano performance demonstration and match the audio track with the music score. This allows users to accurately compare the music score with the audio while playing the audio. This method of synchronizing music scores with audio can help users learn by comparing it to professional performance recordings.
[0005] However, it's difficult for users to independently judge the difference between their performance and the demonstration's performance. To address this, prior art has proposed some piano practice assistance functions that electronically display the user's performance status (such as playing rhythm, performance, etc.) to help users understand the effectiveness of their practice during self-practice. For example, patent application CN107767847A discloses an intelligent piano performance evaluation method and system, which includes: matching the user's performance time sequence with the master's performance time sequence on a time axis; performing intonation determination on the user's note groups and the master's note groups; performing rhythm determination on the relative duration of the user's notes and the master's notes; calculating the relative speed of the user's performance speed and the master's performance speed for speed determination; performing velocity determination on the user's note dynamics and the master's note dynamics; performing a comprehensive evaluation of the performance intonation, rhythm, speed, and dynamics; and marking the results of the comprehensive evaluation on the staff of the smart terminal. However, the applicant has noted that this timing matching method has certain drawbacks in evaluation reliability.
[0006] Furthermore, during piano practice, users often need to turn pages of music scores according to the performance structure. However, manual page turning can affect the user's performance continuity to a certain extent. Therefore, some existing technologies have attempted to use automatic page turning functions, but these functions may have the following drawbacks: 1) the page turning rhythm has a low synergy with the user's actual performance rhythm; 2) the user is required to manually issue page turning commands, which may still affect the performance continuity.
[0007] For example, patent application CN114417915A discloses a two-dimensional sequence similarity assessment system for music sheet flipping. This method uses similarity calculations on pitch and time dimensions to perform position matching. Specifically, it regards multiple notes at the same time as a whole and performs calculations, with multiple notes that are fully matched and compared as the benchmark. The user's latest playing position is thus calculated, and whether to perform a music sheet flipping operation is determined based on the playing position. For another example, patent application CN109961800A also discloses a music sheet page turning processing method and device, the method comprising: filtering preset audio to obtain page-turning notes of a target instrument; obtaining target audio in real time during the performance of the music sheet, and analyzing the target audio to obtain target performance notes; if it is determined that the target performance note matches the target page-turning note, controlling the music sheet to perform a page-turning operation. For another example, patent application CN115294591A also discloses a method and device, electronic device, and storage medium for intelligent page turning of music scores. The method includes: completely identifying paper music scores or electronic music scores, generating corresponding electronic audio, extracting time-frequency information features of the played music and matching them with the time-frequency information features of the electronic audio of the music scores, using the electronic audio corresponding to the music scores as a standard reference for music score matching, determining the current performance progress, and thus realizing intelligent page turning of paper music scores or electronic music scores.
[0008] It can be seen that the intelligent page turning route adopted in the prior art is mainly: the user sets the page turning node, and once the performance is monitored to the page turning node, the page turning is started. However, the applicant has noticed that although this automatic page turning function reduces the user's manual labor to a certain extent. However, for beginners, due to their low proficiency in music scores, they are more dependent on the content of the music scores, and are more likely to make mistakes or miss notes during practice. Therefore, when faced with some special situations, such as turning the page too early or turning the page by mistake, the traditional page turning mode is more likely to cause some adverse interference to the performer. Summary of the Invention
[0009] The purpose of the present invention is to provide a method and system for matching audio and music scores, which can partially solve or alleviate the above-mentioned deficiencies in the prior art and improve the matching accuracy of audio and music scores.
[0010] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: A first aspect of the present invention is to provide a method for matching audio and music scores, comprising the steps of: S101, providing a reference database of musical scores, the musical scores comprising: a plurality of markers for marking expected performance characteristics, at least one of the markers constituting a musical unit; correspondingly, the reference database comprising: a plurality of segments of expected performance characteristics corresponding to the plurality of musical units, wherein the expected performance characteristics include expected pitch, expected dynamics, and expected duration, and the musical units are further associated with expected performance orders; S102, obtaining first note data generated by a user playing at a first time, the note data being actual performance characteristics corresponding to at least one marker, the actual performance characteristics including: actual pitch, actual velocity, and actual duration; S103, determining whether second note data generated at a second time immediately before the first time is successfully matched to a first target music unit using a matching rule, wherein the matching rule is that when a similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit of the note data; If yes, then perform the following steps: S104, determining a first dynamic prediction order at the first time according to the first target music unit; S105, obtaining a first expected performance feature corresponding to the first dynamic prediction sequence; S106, using the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; If so, the second target music unit is marked as the current performance node.
[0011] In some embodiments, when the judgment result of S103 is no, the following steps are executed: When the lengths of the first time and the second time meet a preset first trial time, the following steps are executed: S108, generating a note data sequence based on the first note data and adjacent to-be-matched note data, wherein the to-be-matched note data refers to at least one note data that is not matched to the target music unit; S109, calculating the similarity between the note data sequence and the expected performance characteristics of at least one third music unit; wherein the third music unit is at least one syllable unit in a historical period extending forward by a first length along a second dynamic prediction order, and / or the third music unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data most recently successfully matched to the target music unit; S110: Mark the third music unit with the greatest similarity as the third target music unit of the note data sequence.
[0012] In some embodiments, the steps further include: When the determination result of S106 is no, obtaining third note data at a third time after the first time, and identifying the third note data as note data to be matched; When the first time and the third time satisfy the first trial time, steps S108 - S109 are executed.
[0013] In some embodiments, the steps further include: Calculating the frequency of incorrect performances by the user during the first period; wherein, when the determination result of S105 is negative once, the number of incorrect performances is incremented by one; When the frequency of incorrect performance is greater than a first set number of times, the first trial time is extended; When the erroneous playing frequency is greater than the second set number of times and less than or equal to the first set number of times, the first trial time is maintained; When the erroneous playing frequency is less than or equal to the second set number of times, the first trial time is reduced.
[0014] In some embodiments, the identifier is a musical note.
[0015] In some embodiments, one of the music units corresponds to one of the identifiers.
[0016] In some embodiments, the similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes: L=aX+bY+cZ; L is the similarity, X is the pitch similarity, Y is the dynamics similarity, Z is the duration similarity, a is the first weight, b is the second weight, and c is the third weight.
[0017] The present invention also provides an audio and music score matching system, comprising: a reference data module for providing a reference database of musical scores, wherein the musical scores include: a plurality of markers for marking expected performance characteristics, at least one of the markers forming a musical unit; and correspondingly, the reference database includes: a plurality of segments of expected performance characteristics corresponding to the plurality of musical units, wherein the expected performance characteristics include: expected pitch, expected dynamics, and expected duration, and the musical units are further associated with expected performance orders; a feature acquisition module configured to acquire first note data generated by a user's performance at a first time, wherein the note data is actual performance features corresponding to at least one marker, wherein the actual performance features include actual pitch, actual velocity, and actual duration; a first matching module, configured to determine, using a matching rule, whether second note data generated at a second time immediately before the first time is successfully matched to a first target music unit, wherein the matching rule is that when a similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit; If yes, then entry is allowed: a prediction module, configured to determine a first dynamic prediction order at the first time according to the first target music unit; An expected data acquisition module, configured to acquire a first expected performance feature corresponding to the first dynamic prediction sequence; The second matching module is used to use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; if so, mark the second target music unit as the current performance node.
[0018] In some embodiments, when the judgment result of the first matching module is negative, and when the lengths of the first time and the second time meet a preset first trial time, access to the following modules of the system is allowed: a data sequence generating module, configured to generate a note data sequence based on the first note data and adjacent to-be-matched note data, wherein the to-be-matched note data refers to at least one note data that is not matched to the target music unit; a similarity calculation module, configured to calculate a similarity between the note data sequence and an expected performance feature of at least one third music unit; wherein the third music unit is at least one syllable unit in a historical period extending forward by a first length along a second dynamic prediction sequence, and / or the third music unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction sequence; the second dynamic prediction sequence being determined based on the note data most recently successfully matched to the target music unit; The marking module is used to mark the third music unit with the greatest similarity as the third target music unit of the note data sequence.
[0019] In some embodiments, when the judgment result of the second matching module is no, the third note data at a third time after the first time is obtained, and the third note data is identified as the note data to be matched; and when the first time and the third time meet the first trial time, it is allowed to enter the similarity calculation module and the marking module.
[0020] Beneficial technical effects: The preferred dynamic matching method of audio and music scores in this application is optimized at least two levels: 1) A technical path that uses dual matching mechanisms to switch between two routes: matching mechanism 1 (prioritizing matching of notes one by one in the normal playing order) and matching mechanism 2 (setting a trial time to tentatively match data sequences of multiple notes) based on real-time matching results.
[0021] It is worth noting that the flexible switching of this dual matching mechanism can, on the one hand, control the amount of calculation in each matching process to a certain extent, that is, to impose a certain degree of restriction on the number of music units to be matched and the matching interval (that is, the length of the extension); on the other hand, in the process of applying matching mechanism 1, it is preferred to use a single note for matching, which can also improve the matching efficiency, thereby ensuring the real-time and continuity of the positioning mark display.
[0022] From another perspective, the switching mode of this dual-matching mechanism can comprehensively balance the marker update efficiency required for the display consistency of the positioning marker and the matching calculation required for display accuracy, thereby optimizing the user experience efficiency without increasing the processing pressure.
[0023] 2) The trial duration is flexibly set based on the user's actual playing level (such as the frequency of incorrect playing). This flexible setting of the trial duration can achieve a certain degree of balance between the real-time display and the display reliability of the positioning mark, especially taking into account the actual needs of users at different stages or of different types.
[0024] For example, when the user's proficiency in the current music score is low (for example, the user is a beginner, or the user is practicing a new song), the frequency of incorrect playing during the performance may be relatively high. At this time, once it is discovered that the user may have problems such as skipping and missing strings, the real-time matching will be appropriately slowed down to accumulate a data sequence of a certain trial length before matching, so as to avoid matching errors, which will have a counter-effect on the user.
[0025] At this point, the marker on the score can remain in its current location (for example, at the position or adjacent area of the last note successfully matched to the target musical unit), rather than continuing to move backward along the musical score. This temporary delay tracking setting also helps to attract the user's attention, allowing them to verify the current performance from a personal perspective.
[0026] It is worth noting that this application preferably selects a limited number of three types of data, namely pitch, duration, and dynamics, from the dual dimensions of basic techniques and emotional techniques to evaluate data similarity. Among them, pitch and duration are used to reflect the user's basic technical ability, while dynamics can reflect the user's grasp of emotion. Therefore, the selection of factors in this dual dimension can not only improve the reliability of data matching to a certain extent, but also control the amount of data matching calculations to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the various elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work.
[0028] Figure 1 1 is a flow chart of a matching method in an exemplary embodiment of the present invention; Figure 2 is a schematic diagram of the module structure of a matching system in an exemplary embodiment of the present invention; Figure 3 1 is a flow chart of a page turning method in an exemplary embodiment of the present invention; Figure 4 A schematic diagram of the effect of double-page display in an exemplary embodiment of the present invention; Figure 5 Schematic diagram of page turning effect under double-page display in an exemplary embodiment of the present invention; Figure 6 FIG. 4 is a schematic diagram of the module structure of a page turning system in an exemplary embodiment of the present invention.
[0029] Summary of reference numerals: positioning mark 01, display frame 02, first display page 03, historical score content 031, new score content 032, second display page 04, update mark 05. DETAILED DESCRIPTION
[0030] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0031] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate description of the present invention and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.
[0032] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate and simplify the description of the present invention and are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0033] As used herein, unless otherwise expressly specified or limited, the terms "installed," "provided with," and "connected" should be understood broadly. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection, a direct connection, an indirect connection via an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention on a case-by-case basis.
[0034] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.
[0035] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0036] As used in this specification, the term "about" typically means + / - 5% of the stated value, more typically + / - 4% of the stated value, more typically + / - 3% of the stated value, more typically + / - 2% of the stated value, even more typically + / - 1% of the stated value, and even more typically + / - 0.5% of the stated value.
[0037] In this specification, certain embodiments may be disclosed in a format that is within a range. It should be understood that this description of "within a range" is merely for convenience and brevity and should not be interpreted as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values within this range. For example, the description of a range of 1-6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numbers within this range, such as 1, 2, 3, 4, 5, and 6. Regardless of the breadth of the range, the above rules apply.
[0038] In this article, "music score" or "music notation" is a regular combination of various written symbols (or identifiers) that record musical pitch or rhythm. For example, common music scores include simple notation, five-line notation, etc.
[0039] Taking the piano as an example, the musical notation used for pianos is typically five-line staff. Five-line staff is a method of recording music using notes of varying durations and other symbols (also known as written notation) on five equally spaced parallel lines. Each line of the staff, and the spaces between them, are called, from bottom to top, the first line, second line, third line, fourth line, fifth line, and first space, second space, third space, and fourth space. If there are insufficient lines and spaces, additional lines and spaces can be added above or below the staff. These extra lines and spaces are called upper extra lines, upper extra spaces, lower extra lines, and lower extra spaces, respectively, and each represents a pitch. The fixed height of these pitches is determined by the clef used.
[0040] A syllable in musical notation is the smallest unit of sound that can be produced independently, representing the rhythm and melody of the music. In musical notation, syllables are represented by specific identifiers, such as notes, rests, and slurs. These symbols specifically represent the musical characteristics of each syllable, such as its length, pitch, and dynamics. The combination of these symbols constitutes a complete musical work. In this specification, a typical identifier is a note. A syllable can contain one or more notes.
[0041] Example 1 See also Figure 1 As shown, the present invention provides a method for matching audio and music scores, comprising the steps of: S101, providing a reference database of music scores, wherein the music scores include: a plurality of markers for marking expected performance features, at least one of the markers constitutes a music unit, and correspondingly, the reference database includes: a plurality of segments of expected performance features corresponding to the plurality of music units, wherein the music units are also associated with an expected performance order.
[0042] Preferably, in some embodiments, the expected performance characteristics include: expected pitch, expected strength, and expected duration.
[0043] Preferably, in some embodiments, one marker constitutes one music unit. An expected performance order is generated for each music unit according to the performance order of the markers in the music score.
[0044] S102, obtaining first note data generated by the user's performance at a first time t1, wherein the note data is actual performance characteristics corresponding to at least one marker. Preferably, the actual performance characteristics include: actual pitch, actual velocity, and actual duration; S103, determining whether the second note data generated at a second time t2 immediately before the first time t1 is successfully matched to the first target music unit using a matching rule, wherein the matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit; For example, in some embodiments, when the degree of overlap between the actual pitch, actual velocity and actual duration of the first note data and the expected pitch, expected velocity and expected duration of the musical unit is greater than the set overlap, the similarity is considered to be high (i.e., greater than the first set threshold).
[0045] If the judgment result of S103 is yes, then execute the following steps: S104, determining a first dynamic prediction order at the first time according to the first target music unit; For example, according to the performance order of the music score, the music unit that appears after the first target music unit is set as the unit to be performed at the next time (or the predicted performance unit), and its performance order is defined as the first dynamic prediction order.
[0046] S105, obtaining the first expected performance feature corresponding to the first dynamic prediction sequence; That is, the expected performance characteristics of the unit to be played are obtained.
[0047] S106, using the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; If so, the second target music unit is marked as the current performance node.
[0048] For example, in some embodiments, a positioning mark can be added to the second target music unit on the music score. Of course, the positioning mark can move according to the user's performance progress, so that the user can observe the performance progress on the music score, so that the user can self-correct during the piano practice process.
[0049] That is, in this embodiment, when the first target music unit is successfully matched at the current time (e.g., the second time t2) (i.e., the performance result at the previous time is preliminarily determined to be accurate), the performance information of the next node is directly predicted based on the current position of the first target music unit (i.e., the performance node at the second time) according to the performance order of the score. That is, the music unit located after and adjacent to the first target music unit is set as the unit to be played at the next node. Thus, the first dynamic prediction order for the next time is determined in real time from the time dimension.
[0050] In some embodiments, the marker is a musical note.
[0051] For example, in some embodiments, the audio file generated by the user's performance can be collected in real time, and the actual pitch, force, and duration corresponding to each note can be analyzed based on the audio file.
[0052] For another example, in some embodiments, data such as pitch and duration can be analyzed from audio files, and the force can be directly or indirectly collected using sensors.
[0053] For example, in some embodiments, velocity sensors, displacement sensors, or acceleration sensors are installed beneath piano keys to detect the velocity, displacement, or acceleration of the corresponding keys in real time, thereby indirectly calculating the force exerted on the keys. For another example, in some embodiments, an image sensor can be used to capture a hand image and calculate the finger pressure depth, which can then be converted into force.
[0054] In some embodiments, when the judgment result of S103 is no, that is, when the second note data at the time before the current moment cannot be matched with the music score according to the conventional playing order, the steps are executed: When the lengths of the first time and the second time meet a preset first trial time, the following steps are executed: S108, generating a note data sequence based on the first note data and adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that is not matched to the target music unit; For example, in some embodiments, the note data to be matched refers to the second note data, and the note data sequence is a data sequence formed by combining the expected performance features of at least two notes.
[0055] S109, calculating the similarity between the note data sequence and the expected performance characteristics of at least one third music unit; wherein the third music unit is at least one syllable unit covered by a historical period extending forward by a first length along a second dynamic prediction order, and / or the third music unit is at least one syllable unit covered by a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data most recently successfully matched to the target music unit; For example, in some embodiments, when the note data at both the first and second times cannot be successfully matched according to the conventional playing order, a fourth target music unit (i.e., the most recently successfully matched performance node) at the time immediately before the second time may be searched for, and a second dynamic prediction order may be generated based on the fourth target music unit. For example, the second dynamic prediction order may be the performance order of the next adjacent music unit of the fourth target music unit under the conventional playing order.
[0056] Preferably, in this embodiment, a method of extending a certain length forward and backward along the second dynamic prediction sequence is adopted to appropriately expand the note data sequence matching interval.
[0057] S110 , marking the third music unit with the greatest similarity as the third target music unit of the note data sequence, that is, determining the third music unit as the current performance node.
[0058] In this embodiment, the third music unit is a combination of at least two music units (therefore it can also be called a music combination).
[0059] For example, in some embodiments, similarity calculations are performed between a note data sequence and multiple music combinations to obtain multiple similarity values. The maximum similarity value is selected, and a determination is made as to whether the maximum similarity value is greater than a first set threshold. If so, the music combination corresponding to the maximum similarity value is considered to be successfully matched with the note data sequence. This means that the user is currently playing the position of the music combination.
[0060] In some embodiments, the steps further include: When the determination result of S106 is no, obtaining third note data at a third time after the first time, and identifying the third note data as note data to be matched; wherein the note data to be matched refers to at least one note data that is not matched to the target music unit; When the first time and the third time satisfy the first trial time, steps S108 - S109 are executed.
[0061] That is to say, in this embodiment, when the target music unit is successfully matched at the time before the first time t1 (such as the second time t2), but the note data at the first time cannot be successfully matched, a certain length of trial time will be reserved for the matching process, so that a similarity match can be performed when a note data sequence with a certain length is captured.
[0062] In some embodiments, the first trial time may have an initial value preset by a user.
[0063] In some embodiments, the method further includes the step of calculating the frequency of incorrect performances by the user during the first time period; wherein, when the judgment result of S105 is negative once, the number of incorrect performances is increased by one; wherein, the frequency of incorrect performances refers to the number of incorrect performances per unit time.
[0064] In some embodiments, when the erroneous playing frequency is greater than a first set number of times, the first trial time is extended; in some embodiments, when the erroneous playing frequency is greater than a second set number of times and less than or equal to the first set number of times, the first trial time is maintained; in some embodiments, when the erroneous playing frequency is less than or equal to the second set number of times, the first trial time is reduced.
[0065] For example, in some embodiments, when a note data or a sequence of note data is generated that cannot be matched with the music score according to the normal playing order, it is considered that the user has made a playing error, and an incorrect performance will be recorded.
[0066] Preferably, when the frequency of incorrect performances is high, the trial time can be appropriately extended to avoid a short data sequence and failure to match the correct point.
[0067] Applicants have noted that beginners are more likely to make mistakes during their performances, particularly out of order, such as skipping the current measure. Furthermore, different users handle these mistakes differently. For example, some may replay the incorrect section, while others may continue playing according to the normal sequence. Conventional time-dimension matching mechanisms are unable to address these issues.
[0068] Furthermore, from a user experience perspective, if the positioning mark on the score experiences a significant delay, the interface display loses consistency with the user's actual playing progress, impacting the user experience. Furthermore, if the positioning mark can be displayed continuously on the score, but a misjudgment occurs during the rapid judgment process, resulting in an incorrect display of the positioning mark, this display inconsistency will also impact the user experience. For beginners, given their limited familiarity with the score, this discontinuity or error in the display can easily lead to further inaccurate control of the user's own rhythm, further misleading the user and causing the error to escalate.
[0069] In this regard, the preferred method of dynamic matching of audio and music scores in this application is optimized at least in two aspects: 1) A technical path that uses dual matching mechanisms to switch between two routes: matching mechanism 1 (prioritizing matching of notes one by one in the normal playing order) and matching mechanism 2 (setting a trial time to tentatively match data sequences of multiple notes) based on real-time matching results.
[0070] It is worth noting that the flexible switching of this dual matching mechanism can, on the one hand, control the amount of calculation in each matching process to a certain extent, that is, to impose a certain degree of restriction on the number of music units to be matched and the matching interval (that is, the length of the extension); on the other hand, in the process of applying matching mechanism 1, it is preferred to use a single note for matching, which can also improve the matching efficiency, thereby ensuring the real-time and continuity of the positioning mark display.
[0071] From another perspective, the switching mode of this dual-matching mechanism can comprehensively balance the marker update efficiency required for the display consistency of the positioning marker and the matching calculation required for display accuracy, thereby optimizing the user experience efficiency without increasing the processing pressure.
[0072] 2) The trial duration is flexibly set based on the user's actual playing level (such as the frequency of incorrect playing). This flexible setting of the trial duration can achieve a certain degree of balance between the real-time display and the display reliability of the positioning mark, especially taking into account the actual needs of users at different stages or of different types.
[0073] For example, when the user's proficiency in the current music score is low (for example, the user is a beginner, or the user is practicing a new song), the frequency of incorrect playing during the performance may be relatively high. At this time, once it is discovered that the user may have problems such as skipping and missing strings, the real-time matching will be appropriately slowed down to accumulate a data sequence of a certain trial length before matching, so as to avoid matching errors, which will have a counter-effect on the user.
[0074] At this point, the marker on the score can remain in its current location (for example, at the position or adjacent area of the last note successfully matched to the target musical unit), rather than continuing to move backward along the musical score. This temporary delay tracking setting also helps to attract the user's attention, allowing them to verify the current performance from a personal perspective.
[0075] For example, when a user has a high level of proficiency in the current music score (for example, if the user is a professional pianist, or if the user has practiced the music score for a long time), the frequency of incorrect playing during the performance is generally low. In this case, even if a certain degree of error occurs, a small amount of real-time data (i.e., a sequence of note data) can be used to quickly complete the matching in the adjacent area, thereby ensuring the consistency of the positioning mark display and avoiding interference with the user's playing rhythm (the consistency display in this case is more likely to maintain a high degree of consistency with the user's playing rhythm, reducing the sense of interruption or conflict caused by errors on the display interface). At the same time, due to the user's high level of proficiency, their ability to self-correct is also relatively high. Even if an incorrect positioning occurs in special circumstances, the user can quickly adjust to the normal state at a relatively fast pace.
[0076] Furthermore, in some embodiments, the similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes: L=aX+bY+cZ; L is the similarity, X is the pitch similarity, Y is the dynamics similarity, Z is the duration similarity, a is the first weight, b is the second weight, and c is the third weight.
[0077] It's worth noting that in this embodiment, data similarity is assessed by selecting a limited set of data: pitch, duration, and dynamics, from a dual dimension of basic technique and emotional technique. Pitch and duration reflect a user's basic technical skills, while dynamics reflects their emotional grasp. This dual-dimensional selection of factors can improve the reliability of data matching to a certain extent while also providing some control over the computational complexity of data matching.
[0078] It should be understood that the step numbers in this specification are merely for ease of presentation and understanding of the solution, and are not intended to limit the order in which the method is executed. For example, in some embodiments, S102 may be executed first, followed by S103. Alternatively, in other embodiments, when the real-time display of the positioning mark is less demanding, S102-S103 may be executed simultaneously.
[0079] Example 2 See also Figure 2As shown, a system for matching audio and music scores is characterized by comprising: Reference data module 101 is configured to provide a reference database of musical scores, wherein the musical scores include: a plurality of markers for marking expected performance characteristics, wherein at least one of the markers constitutes a musical unit; and correspondingly, the reference database includes: a plurality of segments of expected performance characteristics corresponding to the plurality of musical units, wherein the expected performance characteristics include: expected pitch, expected dynamics, and expected duration, and the musical units are further associated with expected performance orders; The feature acquisition module 102 is configured to acquire first note data generated by a user's performance at a first time, wherein the note data is actual performance features corresponding to at least one marker, wherein the actual performance features include actual pitch, actual velocity, and actual duration; A first matching module 103 is configured to determine, using a matching rule, whether second note data generated at a second time immediately before the first time is successfully matched to a first target music unit, wherein the matching rule is that when a similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit of the note data; If yes, then entry is allowed: A prediction module 104, configured to determine a first dynamic prediction order at the first time according to the first target music unit; An expected data acquisition module 105 is configured to acquire the first expected performance feature corresponding to the first dynamic prediction sequence; The second matching module 106 is configured to use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; if so, mark the second target music unit as the current performance node.
[0080] In some embodiments, when the determination result of the first matching module 103 is negative, and when the lengths of the first time and the second time satisfy a preset first trial time, access to the following modules of the system is allowed: a data sequence generating module 107 for generating a note data sequence based on the first note data and adjacent to-be-matched note data, wherein the to-be-matched note data refers to at least one note data that is not matched to the target music unit; a similarity calculation module 108 for calculating a similarity between the note data sequence and an expected performance feature of at least one third music unit; wherein the third music unit is at least one syllable unit in a historical period extending forward by a first length along a second dynamic prediction sequence, and / or the third music unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction sequence; the second dynamic prediction sequence is determined based on the note data most recently successfully matched to the target music unit; The marking module 109 is configured to mark the third music unit with the greatest similarity as the third target music unit of the note data sequence.
[0081] In some embodiments, when the judgment result of the second matching module 106 is no, the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched; and when the first time and the third time meet the first trial time, it is allowed to enter the similarity calculation module 108 and the marking module 109.
[0082] In some embodiments, the expected performance characteristics include one or more of the following factors: expected pitch, expected dynamics, and expected duration. The actual performance characteristics may also include one or more of the following factors: actual pitch, actual dynamics, and actual duration. Correspondingly, when the performance characteristics are characterized by one factor, the corresponding similarity can be the similarity (or overlap) of the single factor. Alternatively, when the performance characteristics are characterized by two factors, the corresponding similarity can be the weighted sum of the similarities of the two single factors. It is understandable that the specific weighted proportion can be preset by the user based on actual conditions (such as the type of music score).
[0083] It is understandable that the system in this specification can be used to implement the methods or steps in any other method embodiment, which will not be described in detail here.
[0084] Example 3 The present invention provides a method for updating a music score page, comprising the steps of: Providing a display page in a first display mode, the display page being used to display partial score content of the music score, for example, the score content being the result of displaying multiple written symbols on the page according to general music theory rules; Determine whether the target music unit played by the user reaches a preset page update point, where the page update point corresponds to a music unit, and the music unit is composed of at least one identifier; wherein the target music unit is also the user's current playing point, or playing node.
[0085] If so, the display page is updated using the first page turning method.
[0086] In this embodiment, the user's current playing point can be located by the audio and music score matching method adopted in the present invention.
[0087] In some embodiments, the first display mode may include one or more of the following: single-page display, page display (eg, double-page display, multi-page display).
[0088] For example, in some embodiments, the double-page display can be two pages displayed side by side. For example, two pages displayed side by side can simulate the state of opening a music book.
[0089] For example, in some embodiments, the first page turning method may include one or more of the following: scrolling, simulated book turning (i.e., simulating the changing state of pages when a user manually flips through a paper sheet), sliding, and overall page update. For example, when turning pages by scrolling, the current page can be scrolled in a set page turning direction to display the score content of the next page in the area where the current page left off. For example, sliding can cause the current page to slide in a set page turning direction, and the score content of the next page to be displayed in the area where the current page left off. For another example, overall page turning can directly update the score content displayed on the current page.
[0090] In some embodiments, the page update method further includes: obtaining a trigger update signal from the user, and when the trigger update signal is received, defining a page update point according to the trigger update signal. For example, the page update point can also be a page turning time, thereby completing the page update.
[0091] In some embodiments, the page update point can be pre-set by the user. For example, the page update point can be a certain syllable unit. For another example, the page update point can also be defined by an update time. The update time can be a time point or a selectable time range.
[0092] Preferably, in some embodiments, especially in single-page display mode, in order to reduce the interference of page turning on the user's playing fluency, page turning will be selected between phrase gaps. Phrase gaps refer to the space or pauses between phrases in a musical work. In music creation and performance, the handling of phrase gaps is crucial to the fluency and expressiveness of music. A phrase is the basic structural unit that constitutes a piece of music and can express a relatively complete meaning. A phrase is similar to a sentence in an article and has the function of expressing a complete idea. In a musical work, a phrase usually consists of several bars, has a clear starting and ending point, and can express a specific emotion or idea. And a bar consists of at least one musical unit.
[0093] For example, in some embodiments, the method further comprises the steps of: Finding at least one phrase gap adjacent to the page update point (such as update time); Calculating the time difference between the phrase gap and the page update point; When the time difference is less than a set time threshold, the corresponding phrase gap is set to a corrected page update point; otherwise, the initially set / acquired page update point may be retained.
[0094] Furthermore, in some embodiments, when there are two phrase gaps, the phrase gap with the smaller time difference may be selected as the corrected page update point.
[0095] In this embodiment, the page update point is corrected by the phrase gap, which can reduce the interference to the user's performance continuity during the page turning process to a certain extent.
[0096] In some embodiments, practice data of a user practicing a music score multiple times may be collected, and the practice data may include: page turning time points of the user on different display pages during different practice sessions. In this regard, the page update point may be the page turning time point selected by the user during the last practice session.
[0097] Furthermore, in some embodiments, the practice data further includes: For a displayed page, record the frequency of incorrect performances in the immediate period after at least one page turning time point; When the erroneous playing frequency is greater than a preset first erroneous frequency threshold, the page updating point is selected to be corrected, and the corrected page updating point is the page turning time point plus the set delay time.
[0098] In this embodiment, to prevent users from misjudging the page-turning rhythm during early practice, the page-turning timing is optimized based on the performance quality after the page-turning to optimize the timing of page-turning. In this embodiment, for single-page display mode, this optimization of page-turning timing ensures performance continuity before and after page-turning, and to a certain extent, prevents incorrect playing during the page-turning process due to user inexperience with the music.
[0099] Furthermore, in some embodiments, the method further includes: After correcting the page update point, collecting new practice data of the user; And when the frequency of incorrect performance of the new practice data decreases in the immediate period after the page update point, you can choose to temporarily maintain the current page update point.
[0100] Specifically, if the frequency of incorrect performances decreases and is less than a set second error frequency threshold, the current page update point can be maintained. Alternatively, if the frequency of incorrect performances decreases but remains greater than or equal to the second error frequency threshold after at least two practice sessions, further corrections to the page update point can be attempted. For example, the corrected page update point is the previous page update point plus the set delay time.
[0101] For example, in some embodiments, the types of the first display mode and the first page turning mode can be selected by the user.
[0102] For example, in some embodiments, the practice data will also record the type of the first display mode and / or the type of the first page turning mode currently being used. In this case, if the user does not input a new adjustment instruction during the current performance phase, the first display mode and / or the first page turning mode used last time can be directly used as the current default setting mode.
[0103] Further, see Figure 3 As shown, the present invention also provides a method for updating a music score page, the method comprising the steps of: S201, providing a first display page and a second display page, wherein the first display page is used to display a first partial score content of the music score, and the second display page is used to display a second partial score content of the music score, and the first partial score content and the second partial score content are coherent content in the music score; Preferably, the first display page and the second display page can be displayed side by side on the interface, such as the first and second display pages can be displayed side by side up and down.
[0104] S202, selecting a current display page and a page to be updated from the first display page and the second display page according to the user's current playing point (i.e., the target music unit matched / located at the current moment), wherein the current display page is the display page where the current playing point is located; S204, determining whether the current playing point reaches a preset page update point, wherein the page update point corresponds to a music unit, and the music unit is composed of at least one identifier; If yes, then perform the following steps: S205, obtaining the remaining playing time of the currently displayed page and the playing speed of the user; For example, in some embodiments, the remaining playing duration can be calculated based on the remaining to-be-played markers (such as notes, rests, etc.) on the currently displayed page using music theory rules.
[0105] For example, in some embodiments, the personal style and musical understanding of the performer (equivalent to the user) may have an impact on the performance speed, so the performance speed of the user when playing the music score can be recorded in real time.
[0106] For example, in some embodiments, corresponding to music scores of different genres, their performance speeds may have a general recommended range (e.g., the recommended range may be determined based on exemplary performance data in an existing music library). Therefore, when the performer's own performance data is limited, the performance speed determined by the exemplary performance data may also be used as a reference.
[0107] S206, determining an expected playing time of the currently displayed page after the page update point according to the remaining playing time and the playing speed; S207: determining an update speed of the page to be updated according to the expected playing time, and performing a page update operation on the page to be updated within the expected playing time.
[0108] It's important to note that duration generally refers to the length of a note, or the amount of time it takes up. The unit of measurement is typically the beat, the smallest unit of time in music. Tempo refers to the number of beats per unit of time. For example, tempo can be measured in BPM (beats per minute). Therefore, the expected duration of the remaining notes can be estimated by dividing the remaining duration by the tempo.
[0109] For example, the actual playing time of each marker (such as a note) can be obtained by dividing the duration by the playing speed, and the expected playing time can be obtained by summing the actual playing times of multiple markers.
[0110] For example, in some embodiments, the update speed refers to the magnitude of the change in the page content per unit time. Preferably, under the expected performance duration, a relatively uniform update speed can be used to update the page.
[0111] In this embodiment, the page update operation is preferably completed for the page to be updated within the expected playing time.
[0112] For example, in some embodiments, at least one display page is further associated with / set with a latest update point, ie, the page should be updated before the latest update point at the latest.
[0113] For example, in this embodiment, the remaining playing duration can be the duration between the page update point and the latest update point, that is, the remaining playing duration is determined by calculating the duration of the marker between the page update point and the latest update point.
[0114] Preferably, in some embodiments, in order to maintain the continuity of the user's performance, especially for users who are not very proficient in the performance content, it is necessary to update the content of the current page after the performance of the current page is completed.
[0115] In some embodiments, the method further includes: S203, using a separate display mark to distinguish and display the currently displayed page and the page to be updated.
[0116] In some embodiments, the method further comprises the steps of: When the current playing point is switched to another display page, the selection of the current display page and the page to be updated is switched between the first display page and the second display page. In other words, the first display page and the second display page will alternate between the current display page and the page to be updated in sequence.
[0117] In some embodiments, before S204, the step further includes: Determining whether a trigger update signal sent by the user is received before entering the first response time of the page update point; If yes, defining the page update point according to the trigger update signal; If not, the historical update node of the user is obtained, and the historical update node is defined as the page update point.
[0118] In some embodiments, further comprising: A positioning mark for locating the target syllable unit played by the user is displayed in the currently displayed page.
[0119] In some embodiments, the positioning mark moves along a first update direction on the currently displayed page, and correspondingly, the updated score content on the page to be updated is updated along a second update direction on the page to be updated, and the first update direction is opposite to the second update direction.
[0120] by Figure 4 、 Figure 5 For example, the positioning mark can be dynamically moved along the direction of the line staff in the score (i.e., horizontally) following the performance node. At the same time, the update direction on the page to be updated can also be horizontal. This dual-page mode, achieved by horizontal movement and updating, has a certain visual coordination between the positioning mark movement and the update content change effects. This dual-page mode also ensures that the current performance data display and subsequent performance data updates on the page can be synchronized. As a result, even if the performer has a low level of score knowledge, it is less likely to make mistakes or miss notes due to incomplete score display.
[0121] In some embodiments, the separation display mark includes one or more of the following: brightness display, indicator mark.
[0122] For example, in some embodiments, different brightness levels may be used to display the currently displayed page and the page to be updated, such as enhancing the display brightness of the currently displayed page to display the two differently.
[0123] For another example, in some embodiments, the indicator mark may also be reflected in different colors and mark sizes of the written symbols of the currently displayed page and the page to be updated, so as to visually distinguish the display effects of the two.
[0124] Alternatively, in some embodiments, the indicator may also be embodied as a display box, an arrow, or other indication forms, and these indicator marks may indicate the front displayed page to prompt the user to focus on the core area.
[0125] See also Figure 4-Figure 5 As shown, it shows the effect of horizontal update changes based on double-page display. Figure 4 The first display page 03 and the second display page 04 are shown, and the music score contents are displayed on the display pages respectively. Figure 4 In the figure, the first display page 01 is the current display page. On the current display page, a positioning mark 01 is displayed at the target music unit played by the user to indicate the current performance progress. On the other hand, a display frame 02 (equivalent to a separator display mark) is set outside the current display page to mark the current display page, so that the user can more intuitively focus on the key content of the score.
[0126] Further, see Figure 5 As shown, when the target music unit played by the user reaches the second display page 04, the current display page and the page to be updated are switched in the first and second display pages. At this time, the positioning mark 01 moves to the new current display page (that is, the second display page 04), and the page to be updated (that is, the current first display page 03) will perform the page update operation.
[0127] At this time, the historical score content 031 of the page to be updated will be gradually replaced by the new score content 032.
[0128] Figure 5An exemplary display of a sliding update is shown, wherein historical score content 031 gradually moves along the second update direction, while new score content 032 gradually fills the area left by historical score content 031. Preferably, an update marker 05 (e.g., a scoreless area of a certain length) is used to separate historical score content 031 and new score content 032 to visually distinguish between the new and old content.
[0129] Example 4 See also Figure 6 As shown, the present invention also provides a music score page update system, comprising: The page display module 201 is configured to provide a first display page and a second display page, wherein the first display page is configured to display a first partial score content of the music score, and the second display page is configured to display a second partial score content of the music score, wherein the first partial score content and the second partial score content are coherent content in the music score; A page selection module 202 is configured to select a current display page and a page to be updated from the first display page and the second display page according to a target music unit currently being played by the user, wherein the current display page is the display page where the current playing point is located; A separate display module 203 is configured to distinguish and display the currently displayed page and the page to be updated by using a separate display mark; An update determination module 204 is configured to determine whether the target music unit has reached a preset page update point, wherein the page update point corresponds to a music unit, and the music unit is composed of at least one identifier; If the determination result of the update determination module 204 is yes, then entry is allowed: The performance data acquisition module 205 is used to obtain the remaining performance time of the currently displayed page and the user's performance speed; a duration determination module 206 for determining an expected performance duration of the currently displayed page after the page update point according to the remaining performance duration and the performance speed; The updating module 207 is configured to determine an updating speed of the page to be updated according to the expected playing duration, and perform a page updating operation on the page to be updated within the expected playing duration.
[0130] In some embodiments, the system further comprises: The page switching module is used to switch the selection of the current display page and the page to be updated in the first display page and the second display page when the target music unit switches to another display page.
[0131] The present invention further provides a computer-readable storage medium storing a computer program implemented in the form of computer-readable instructions for implementing the steps of the method described in any embodiment of the present invention. When the computer program is invoked and executed by a computer, the computer program executes the steps of the method described in any embodiment of the present invention. The present invention further provides a computer program product including the computer program. When the computer program is executed by a processor, the computer program implements the steps of the method described in any embodiment of the present invention.
[0132] It is understandable that the system in the present invention can implement the steps of the method described in any embodiment of the present invention, and will not be repeated here.
[0133] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0135] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for matching audio and music score, characterized in that: Including steps: S101, providing a reference database of musical scores, the musical scores comprising: a plurality of markers for marking expected performance characteristics, at least one of the markers constituting a musical unit; correspondingly, the reference database comprising: a plurality of segments of expected performance characteristics corresponding to the plurality of musical units, wherein the expected performance characteristics include expected pitch, expected dynamics, and expected duration, and the musical units are further associated with expected performance orders; S102, obtaining first note data generated by a user playing at a first time, the note data being actual performance characteristics corresponding to at least one marker, the actual performance characteristics including: actual pitch, actual velocity, and actual duration; S103, determining whether second note data generated at a second time immediately before the first time is successfully matched to a first target music unit using a matching rule, wherein the matching rule is that when a similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit of the note data; If yes, then perform the following steps: S104, determining a first dynamic prediction order at the first time according to the first target music unit; S105, acquiring a corresponding first expected performance feature according to the first dynamic prediction sequence; S106, using the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; If so, the second target music unit is marked as the current performance node.
2. The method according to claim 1, characterized in that When the judgment result of S103 is no, the following steps are executed: When the lengths of the first time and the second time meet a preset first trial time, the following steps are executed: S108, generating a note data sequence based on the first note data and adjacent to-be-matched note data, wherein the to-be-matched note data refers to at least one note data that is not matched to the target music unit; S109, calculating the similarity between the note data sequence and the expected performance characteristics of at least one third music unit; wherein the third music unit is at least one syllable unit in a historical period extending forward by a first length along a second dynamic prediction order, and / or the third music unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data most recently successfully matched to the target music unit; S110: Mark the third music unit with the greatest similarity as the third target music unit of the note data sequence.
3. The method according to claim 2, characterized in that Also includes the steps: When the determination result of S106 is no, obtaining third note data at a third time after the first time, and identifying the third note data as note data to be matched; When the first time and the third time satisfy the first trial time, steps S108 - S109 are executed.
4. The method according to claim 2, characterized in that Also includes the steps: Calculating the frequency of incorrect performances by the user during the first period; wherein, when the determination result of S106 is negative once, the number of incorrect performances is increased by one; When the frequency of incorrect performance is greater than a first set number of times, the first trial time is extended; When the erroneous playing frequency is greater than the second set number of times and less than or equal to the first set number of times, the first trial time is maintained; When the erroneous playing frequency is less than or equal to the second set number of times, the first trial time is reduced.
5. The method according to claim 1, wherein The identifier is a musical note.
6. The method according to claim 1, characterized in that One of the music units corresponds to one of the identifiers.
7. The method according to any one of claims 1 to 6, characterized in that The similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes: ; L is the similarity, X is the pitch similarity, Y is the dynamics similarity, Z is the duration similarity, a is the first weight, b is the second weight, and c is the third weight.
8. A system for matching audio and music scores, characterized in that: include: A reference data module (101) is used to provide a reference database of music scores, wherein the music scores include: a plurality of markers for marking expected performance features, at least one of the markers constitutes a music unit, and correspondingly, the reference database includes: a plurality of segments of expected performance features corresponding to the plurality of music units, wherein the expected performance features include: expected pitch, expected dynamics, and expected duration, and the music units are also associated with expected performance sequences; A feature acquisition module (102) is used to acquire first note data generated by a user playing at a first time, wherein the note data is an actual performance feature corresponding to at least one marker, and the actual performance feature includes: actual pitch, actual force, and actual duration; A first matching module (103) is configured to determine, using a matching rule, whether second note data generated at a second time adjacent to the first time successfully matches a first target music unit, wherein the matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the music unit corresponding to the note data is considered to be the target music unit; If yes, then entry is allowed: A prediction module (104), configured to determine a first dynamic prediction order at the first time according to the first target music unit; An expected data acquisition module (105) is used to acquire the corresponding first expected performance feature according to the first dynamic prediction sequence; The second matching module (106) is used to use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; if so, marking the second target music unit as the current performance node.
9. The system according to claim 8, characterized in that When the judgment result of the first matching module (103) is no, and when the lengths of the first time and the second time meet a preset first trial time, access to the following modules of the system is allowed: a data sequence generating module (107), configured to generate a note data sequence based on the first note data and adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that is not matched to the target music unit; A similarity calculation module (108) is used to calculate the similarity between the note data sequence and the expected performance characteristics of at least one third music unit; wherein the third music unit is at least one syllable unit in a historical period extending forward by a first length along a second dynamic prediction sequence, and / or, the third music unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction sequence; the second dynamic prediction sequence is determined based on the note data that has been most recently successfully matched to the target music unit; The marking module (109) is used to mark the third music unit with the greatest similarity as the third target music unit of the note data sequence.
10. The system according to claim 9, characterized in that When the judgment result of the second matching module (106) is no, the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched; and when the first time and the third time meet the first trial time, it is allowed to enter the similarity calculation module (108) and the marking module (109).
Citation Information
Patent Citations
Piano playing smart evaluation method and system
CN107767847A
Music score page turning processing method and device
CN109961800A
Music score intelligent page turning method and device, electronic equipment and storage medium
CN115294591A
Method for following musical tone playing in real time and related product
CN110111761A
Automatic page turning system applied to music score
CN113071243A
Cited By
Playing skill display method and system
CN120783712A
Method and system for inserting pedal data in playing file
CN121640970A
A method and system for inserting foot pedal data into a playback file
CN121640970B