A method and system for matching audio with sheet music

By employing a dual-matching mechanism and similarity assessment, the audio and sheet music matching strategy is dynamically adjusted, resolving the issues of user self-judgment difficulties and automatic page-turning interference during piano practice, thereby improving matching accuracy and user experience.

CN120508674BActive Publication Date: 2026-07-17WUXIAN HONGYIN (CHONGQING) TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUXIAN HONGYIN (CHONGQING) TECHNOLOGY CO LTD
Filing Date
2025-03-10
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, it is difficult for users to independently judge the difference between their performance and the demonstration performance during piano practice. Furthermore, the automatic page-turning function can easily cause adverse interference for beginners, affecting the continuity and accuracy of their performance.

Method used

A dual matching mechanism is adopted, combining real-time matching and trial-and-error time switching. By evaluating the similarity of pitch, velocity, and duration, the matching strategy is dynamically adjusted to optimize the matching process between audio and music score, including real-time matching of single notes and trial-and-error matching of multiple note data sequences.

Benefits of technology

It improves the accuracy and consistency of audio-to-score matching, reduces computational load, and enhances the user experience, especially improving beginners' self-correction capabilities and performance consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508674B_ABST
    Figure CN120508674B_ABST
Patent Text Reader

Abstract

This application relates to the field of music score matching technology, specifically to an audio-musical score matching method and system, comprising: providing a music score reference database, which includes: multiple segments of expected performance features of multiple musical units; acquiring first note data generated by a user's performance at a first time; determining whether second note data generated at an adjacent second time before the first time successfully matches a first target musical unit; if so, determining a first dynamic prediction order at the first time based on the first target musical unit; determining whether the musical unit corresponding to the first dynamic prediction order is the second target musical unit of the first note data; if so, marking the second target musical unit as the current performance node. The music score matching technology of this invention can dynamically adjust the matching time axis in conjunction with the performance state, thereby quickly achieving audio-musical score location matching with limited matching computation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority application

[0002] This application claims priority to Chinese Invention Patent Application No. [2025102324395] "A Method and System for Matching Audio and Musical Score", filed on February 27, 2025, which is incorporated herein by reference in its entirety. Technical Field

[0003] This invention relates to the field of real-time music score matching technology, specifically to a method and system for matching audio with music score. Background Technology

[0004] Piano playing is a comprehensive activity that combines logical and visual thinking, mental and physical exertion, and technical skill with artistry. Often, especially for beginners, it's difficult to judge the accuracy of their practice. Furthermore, learning an instrument is a muscular activity; without professional guidance, mistakes become ingrained and difficult to correct over time.

[0005] Currently, existing technologies have also proposed some intelligent piano learning solutions. For example, patent application CN114420074A discloses a method for synchronizing sheet music and audio, and computer-readable storage. This method can use a pre-acquired complete audio track (such as a recording of a master or demonstration piano performance) as a performance demonstration of the sheet music and match the audio track with the sheet music. Therefore, when playing the audio, the user can accurately compare and view the sheet music. This method of synchronizing sheet music and audio can help users learn by comparing with professional performance recordings.

[0006] However, it is difficult for users to independently judge the difference between their playing and the demonstrated performance. To address this, existing technologies have proposed some auxiliary practice functions to electronically display the user's playing status (such as rhythm, performance quality, etc.) to help users grasp the effectiveness of their practice. For example, patent application CN107767847A discloses an intelligent piano performance evaluation method and system, which includes: matching the user's playing time sequence and the master's playing time sequence on a time axis; determining the intonation of the note groups played by the user and the master; determining the rhythm by comparing the relative durations of the notes played by the user and the master; determining the tempo by calculating the relative tempo of the user and the master; determining the dynamics of the notes played by the user and the master; comprehensively evaluating the intonation, rhythm, tempo, and dynamics; and marking the comprehensive evaluation results on the staff of the intelligent terminal. However, the applicant noted that this time-series matching method has certain shortcomings in terms of evaluation reliability.

[0007] In addition, during practice, users often need to turn pages of the sheet music according to the performance structure. However, manually turning pages may affect the continuity of the user's performance to some extent. Therefore, some automatic page-turning functions have been tried in existing technologies, but these page-turning functions may have the following defects: 1) the coordination between the page-turning rhythm and the user's actual playing rhythm is low; 2) the user needs to manually issue the page-turning command, which may still affect the continuity of the performance.

[0008] For example, patent application CN114417915A discloses a two-dimensional sequence similarity evaluation system for sheet music turning. This method uses similarity calculations based on pitch and time dimensions for position matching. Specifically, it treats multiple notes at the same time as a whole for calculation, using a complete match of multiple matching notes as a benchmark. This calculates the user's latest playing position, and determines whether to perform a sheet music turning operation based on the playing position. As another example, patent application CN109961800A also discloses a sheet music turning method and apparatus. The method includes: filtering a preset audio to obtain the turning notes for the target instrument; acquiring the target audio during the performance of the sheet music in real time, and analyzing the target audio to obtain the target playing notes; if it is determined that the target playing notes match the target turning notes, then controlling the sheet music to perform a turning operation. For example, patent application CN115294591A also discloses a method and device for intelligent page turning of sheet music, an electronic device, and a storage medium. The method includes: fully recognizing paper sheet music or electronic sheet music, generating corresponding electronic audio, extracting the time-frequency information features of the music being played and matching them with the time-frequency information features of the electronic audio of the sheet music, using the electronic audio corresponding to the sheet music as a standard reference for matching the music score, and determining the current performance progress, thereby realizing intelligent page turning of paper sheet music or electronic sheet music.

[0009] It can be seen that the intelligent page-turning route used in existing technologies mainly involves the user setting page-turning nodes, and page turning is initiated once the performance reaches a page-turning node. However, the applicant notes that while this automatic page-turning function reduces the user's manual work to some extent, beginners, due to their lower familiarity with the score, rely heavily on the content of the score and are more prone to making mistakes or omissions during practice. Therefore, the traditional page-turning mode can actually cause some adverse interference to the performer in certain situations, such as turning a page too early or in the wrong direction. Summary of the Invention

[0010] The purpose of this invention is to provide a method and system for matching audio with sheet music, which partially solves or alleviates the above-mentioned shortcomings in the prior art and can improve the accuracy of matching audio with sheet music.

[0011] To solve the aforementioned technical problems, the present invention specifically adopts the following technical solution:

[0012] A first aspect of the present invention is to provide a method for matching audio with sheet music, comprising the steps of:

[0013] S101, a reference database of musical scores is provided, the musical scores including: multiple markers for marking expected performance features, at least one of the markers forming a musical unit, corresponding to: multiple expected performance features corresponding to multiple musical units, and the expected performance features including: expected pitch, expected dynamics, and expected duration, the musical unit is also associated with an expected performance order;

[0014] S102, acquire the first note data generated by the user's performance at the first moment, the note data being the actual performance features corresponding to at least one marker, the actual performance features including: actual pitch, actual dynamics, and actual duration;

[0015] S103, using a matching rule to determine whether the second note data generated at the adjacent second time before the first time successfully matches the first target music unit, wherein the matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data;

[0016] If so, proceed with the following steps:

[0017] S104, determine the first dynamic prediction order at the first time based on the first target music unit;

[0018] S105, Obtain the first expected performance feature corresponding to the first dynamic prediction order;

[0019] S106, use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data;

[0020] If so, mark the second target music unit as the current performance node.

[0021] In some embodiments, if the determination result of S103 is negative, then the following steps are performed:

[0022] When the lengths of the first time and the second time satisfy a preset first trial time, the following steps are executed:

[0023] S108, Generate a note data sequence based on the first note data and the adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target music unit;

[0024] S109, calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit;

[0025] S110, the third music unit with the highest similarity is marked as the third target music unit of the note data sequence.

[0026] In some embodiments, the steps further include:

[0027] If the judgment result of S106 is negative, then the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched.

[0028] When the first time and the third time satisfy the first trial time, then steps S108-S109 are executed.

[0029] In some embodiments, the steps further include:

[0030] Calculate the frequency of incorrect playing by the user during the first time period; wherein, when the judgment result of S105 is negative once, the number of incorrect playing is incremented by one;

[0031] If the frequency of incorrect playing exceeds the first set number of times, then the first trial time is extended;

[0032] When the frequency of incorrect playing is greater than the second set number and less than or equal to the first set number, the first trial time is maintained.

[0033] When the frequency of incorrect playing is less than or equal to the second set number of times, the first trial time is reduced.

[0034] In some embodiments, the identifier is a musical note.

[0035] In some embodiments, one music unit corresponds to one identifier.

[0036] In some embodiments, the similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes: L = aX + bY + cZ;

[0037] L represents similarity, X represents pitch similarity, Y represents dynamic similarity, Z represents tempo similarity, a represents the first weight, b represents the second weight, and c represents the third weight.

[0038] The present invention also provides an audio-to-musical score matching system, comprising:

[0039] A reference data module is used to provide a reference database of musical scores. The musical scores include: multiple markers for annotating expected performance features, at least one of the markers forming a musical unit. Correspondingly, the reference database includes: multiple expected performance features corresponding to multiple musical units, and the expected performance features include: expected pitch, expected dynamics, and expected duration. The musical units are also associated with expected performance order.

[0040] The feature acquisition module is used to acquire the first note data generated by the user's performance at the first moment. The note data is the actual performance feature corresponding to at least one marker. The actual performance feature includes: actual pitch, actual dynamics, and actual duration.

[0041] The first matching module is used to determine, using a matching rule, whether the second note data generated at the second time adjacent to the first time successfully matches the first target music unit. The matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data.

[0042] If so, then entry is permitted:

[0043] The prediction module is used to determine the first dynamic prediction order at the first time based on the first target music unit;

[0044] The expected data acquisition module is used to acquire the first expected performance features corresponding to the first dynamic prediction order.

[0045] The second matching module is used to determine, using the matching rules, whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; if so, the second target music unit is marked as the current performance node.

[0046] In some embodiments, when the judgment result of the first matching module is negative, and when the length of the first time and the second time meets a preset first trial time, the following modules of the system are allowed to enter:

[0047] A data sequence generation module is used to generate a note data sequence based on the first note data and adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target musical unit;

[0048] A similarity calculation module is used to calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit;

[0049] A tagging module is used to tag the third music unit with the highest similarity as the third target music unit of the note data sequence.

[0050] In some embodiments, when the judgment result of the second matching module is negative, the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched; and when the first time and the third time satisfy the first trial time, the similarity calculation module and the tagging module are allowed to enter.

[0051] Beneficial technical effects:

[0052] The preferred dynamic matching method for audio and sheet music in this application is optimized at least on two levels:

[0053] 1) A technical approach that uses a dual matching mechanism is adopted, which combines real-time matching results and switches back and forth between two routes: Matching Mechanism 1 (prioritizing matching one note at a time according to the conventional playing order) and Matching Mechanism 2 (setting a trial time to attempt to match the data sequence of multiple notes).

[0054] It is worth noting that the flexible switching of this dual matching mechanism can, on the one hand, control the amount of computation in each matching process to a certain extent, that is, limit the number of musical units to be matched and the matching interval (i.e. the length of the extension) to a certain extent; on the other hand, in the process of applying matching mechanism 1, it is preferable to use a single note for matching, which can also improve matching efficiency and thus ensure the real-time and continuous display of the positioning marker.

[0055] From another perspective, this dual-matching mechanism can comprehensively balance the marker update efficiency required for the continuity of the location marker display and the matching calculation requirements for display accuracy, thereby optimizing the user experience efficiency without increasing the processing pressure.

[0056] 2) The trial duration is flexibly set based on the user's actual playing level (such as the frequency of incorrect playing). This flexible setting of the trial duration can balance the real-time display and reliability of the positioning mark to a certain extent, and in particular, it can take into account the actual needs of users at different stages or of different types.

[0057] For example, when a user has a low level of familiarity with the current sheet music (e.g., the user is a beginner, or the user is practicing a new piece), the frequency of errors in playing the piece may be higher. If the user is found to be skipping or missing a piece, the real-time matching will be slowed down appropriately to accumulate a data sequence of a certain trial length before matching is performed, in order to avoid matching errors that could have a negative effect on the user.

[0058] At this point, the positioning marker on the score can remain in its current position area (e.g., at the position of the last note that successfully matched the target musical unit or in an adjacent area), instead of continuing to move backward along the staff lines. This temporary delay in following the music also helps to draw the user's attention, allowing them to verify the current performance from their own perspective.

[0059] It is worth noting that this application preferably selects a limited number of data types—pitch, duration, and intensity—from two dimensions: basic technique and emotional technique, to evaluate data similarity. Pitch and duration reflect the user's basic technical ability, while intensity reflects the user's grasp of emotion. Therefore, this dual-dimensional selection of factors can improve the reliability of data matching to a certain extent, while also controlling the computational load of data matching. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0061] Figure 1This is a schematic diagram of the matching method in an exemplary embodiment of the present invention;

[0062] Figure 2 This is a schematic diagram of the module structure of the matching system in an exemplary embodiment of the present invention;

[0063] Figure 3 This is a schematic diagram of a page-turning method in an exemplary embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram illustrating the effect of a double-page display in an exemplary embodiment of the present invention;

[0065] Figure 5 This is a schematic diagram illustrating the page-turning effect under a dual-page display in an exemplary embodiment of the present invention;

[0066] Figure 6 This is a schematic diagram of the module structure of a page-turning system in an exemplary embodiment of the present invention.

[0067] Summary of reference numerals in the attached chart: Positioning marker 01, Display box 02, First display page 03, Historical score content 031, New score content 032, Second display page 04, Update marker 05. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0069] In this document, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.

[0070] In this document, the terms "upper," "lower," "inner," "outer," "front," "rear," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the present invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0071] In this document, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0072] In this document, "and / or" includes any and all combinations of one or more of the listed related items.

[0073] In this article, "multiple" means two or more, that is, it includes two, three, four, five, etc.

[0074] As used in this specification, the term "about" typically means + / -5% of the value, more typically + / -4% of the value, more typically + / -3% of the value, more typically + / -2% of the value, even more typically + / -1% of the value, and even more typically + / -0.5% of the value.

[0075] In this specification, certain embodiments may be disclosed in a range-bound format. It should be understood that this "range-bound" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered as having specifically disclosed all possible subranges and the individual numerical values ​​within those ranges. For example, a description of the range 1-6 should be considered as having specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numbers within those ranges, such as 1, 2, 3, 4, 5, and 6. This rule applies regardless of the breadth of the range.

[0076] In this article, "musical score" or "musical notation" refers to a regular combination of written symbols (or identifiers) that record musical pitch or rhythm. Common musical scores include numbered musical notation, staff notation, and so on.

[0077] Taking the piano as an example, the sheet music used on the piano is usually in the form of staff notation. Staff notation is a method of recording music by marking notes of different durations and other symbols (i.e., written symbols) on five equally spaced parallel horizontal lines. Each line of the staff and the spaces between the lines are called, from bottom to top, the first line, second line, third line, fourth line, fifth line, and the first space, second space, third space, fourth space. If there are not enough lines and spaces, additional lines and spaces can be added above or below the staff. These ledger lines and spaces are respectively called the upper ledger first line, upper ledger first space, lower ledger first line, lower ledger first space, etc., each representing a pitch. The fixed pitch of these pitches is determined by the clef used.

[0078] In musical notation, a syllable is the smallest unit of sound that can be produced independently, used to represent rhythm and melody. Syllables are represented by specific symbols in sheet music, such as notes, rests, and slurs. The musical characteristics of each syllable, such as its length, pitch, and dynamics, are specifically represented by these symbols. The combination of these symbols constitutes a complete musical work. A typical symbol used in this specification is the note. A syllable may contain one or more notes.

[0079] Example 1

[0080] See Figure 1 As shown, this invention provides a method for matching audio with sheet music, including the following steps:

[0081] S101, a reference database of musical scores is provided, the musical scores including: multiple markers for marking expected performance features, at least one of the markers forming a musical unit, corresponding to: multiple expected performance features corresponding to multiple musical units, the musical units also being associated with an expected performance order.

[0082] Preferably, in some embodiments, the desired performance characteristics include: desired pitch, desired dynamics, and desired duration.

[0083] Preferably, in some embodiments, a marker constitutes a musical unit. A desired performance order is generated for each musical unit according to the order in which the markers are played in the score.

[0084] S102, acquire the first note data generated by the user's performance at the first time t1, wherein the note data is the actual performance feature corresponding to at least one marker. Preferably, the actual performance feature includes: actual pitch, actual dynamics, and actual duration;

[0085] S103, using a matching rule to determine whether the second note data generated at the adjacent second time t2 before the first time t1 successfully matches the first target music unit, wherein the matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data;

[0086] For example, in some embodiments, when the actual pitch, actual dynamics, and actual duration of the first note data all overlap with the expected pitch, expected dynamics, and expected duration of the musical tone unit, the similarity is considered to be high (i.e., greater than the first set threshold).

[0087] If the result of S103 is yes, then proceed with the following steps:

[0088] S104, determine the first dynamic prediction order at the first time based on the first target music unit;

[0089] For example, according to the performance order of the score, the music unit that appears after the first target music unit is set as the unit to be performed in the next time (or, the predicted performance unit), and its performance order is defined as the first dynamic prediction order.

[0090] S105, Obtain the first expected performance feature corresponding to the first dynamic prediction sequence;

[0091] That is, to obtain the desired performance characteristics of the unit to be performed.

[0092] S106, use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data;

[0093] If so, mark the second target music unit as the current performance node.

[0094] For example, in some embodiments, a display positioning mark can be added to the second target musical unit on the sheet music. Of course, the positioning mark can move according to the user's playing progress, allowing the user to observe the playing progress on the sheet music and thus enabling the user to self-correct during practice.

[0095] In other words, in this embodiment, when the first target music unit is successfully matched at the previous time (e.g., the second time t2) (i.e., the performance result of the previous time is initially considered accurate), the performance information of the next node is predicted directly according to the performance order of the score based on the current position of the first target music unit (i.e., the performance node at the second time). That is, the music unit located after and adjacent to the first target music unit is set as the unit to be performed at the next node. Thus, the first dynamic prediction order for the next time is determined in real time from the time dimension.

[0096] In some embodiments, the marker is a musical note.

[0097] For example, in some embodiments, audio files generated by the user's performance can be acquired in real time, and the actual pitch, dynamics, and duration of each note can be analyzed from the audio files.

[0098] For example, in some embodiments, data such as pitch and duration can be analyzed from audio files, while intensity can be obtained directly or indirectly using sensors.

[0099] For example, in some embodiments, a speed sensor, displacement sensor, or acceleration sensor is provided below the piano keys, which can detect the speed, displacement, or acceleration of the corresponding keys in real time, and thus indirectly calculate the force applied to the keys. As another example, in some embodiments, an image sensor can be used to capture a photograph of the hand, and the depth of finger pressure can be calculated; the force is then calculated by converting the pressure depth.

[0100] In some embodiments, when the determination result of S103 is negative, that is, when the second note data at the previous time cannot be matched with the score according to the normal performance order, the following steps are executed:

[0101] When the lengths of the first time and the second time satisfy a preset first trial time, the following steps are executed:

[0102] S108, Generate a note data sequence based on the first note data and the adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target music unit;

[0103] For example, in some embodiments, the note data to be matched refers to the second note data. The note data sequence is a data sequence formed by combining the expected performance features of at least two notes.

[0104] S109, calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit covered in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit covered in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit;

[0105] For example, in some embodiments, when the note data at the first and second times cannot be successfully matched according to the conventional performance order, a fourth target music unit (i.e., the most recently successfully matched performance node) at the time preceding the second time can be found, and a second dynamic prediction order can be generated based on this fourth target music unit. For example, the second dynamic prediction order is the performance order of the next adjacent music unit of the fourth target music unit under the conventional performance order.

[0106] Preferably, in this embodiment, a certain length is extended along the second dynamic prediction sequence to appropriately expand the matching interval of the note data sequence.

[0107] S110, the third music unit with the highest similarity is marked as the third target music unit of the note data sequence, that is, the third music unit is determined as the current performance node.

[0108] In this embodiment, the third music unit is a combination of at least two music units (therefore it can also be called a music combination).

[0109] For example, in some embodiments, the similarity of the note data sequence with multiple musical combinations is calculated to obtain multiple similarity values. The largest similarity value is selected, and it is determined whether the largest similarity value is greater than a first preset threshold. If so, the musical combination corresponding to the largest similarity value is considered to have successfully matched the note data sequence. That is, it is considered that the user is currently playing the position of that musical combination.

[0110] In some embodiments, the steps further include:

[0111] When the judgment result of S106 is negative, the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched; wherein, the note data to be matched refers to at least one note data that has not been matched with the target music unit;

[0112] When the first time and the third time satisfy the first trial time, then steps S108-S109 are executed.

[0113] In other words, in this embodiment, if the target music unit is successfully matched in the time before the first time t1 (such as the second time t2), but the note data in the first time cannot be successfully matched, a certain length of trial time will be reserved for the matching process so that when a note data sequence of a certain length is captured, similarity matching will be performed.

[0114] In some embodiments, the first trial time may have an initial value preset by the user.

[0115] In some embodiments, the method further includes the step of: calculating the frequency of incorrect performances by the user within a first time period; wherein, when the judgment result of S105 is negative once, the number of incorrect performances is incremented by one; wherein, the frequency of incorrect performances refers to the number of incorrect performances per unit time.

[0116] In some embodiments, when the frequency of incorrect playing is greater than a first set number, the first trial time is extended; in some embodiments, when the frequency of incorrect playing is greater than a second set number and less than or equal to the first set number, the first trial time is maintained; in some embodiments, when the frequency of incorrect playing is less than or equal to the second set number, the first trial time is reduced.

[0117] For example, in some embodiments, when a generated note data or a sequence of note data cannot be matched with the score in the usual playing order, it is considered that the user has made a playing error, and the wrong playing will be recorded.

[0118] Preferably, when the frequency of incorrect playing is too high, the trial time can be appropriately extended to avoid the data sequence being too short and failing to match the accurate point.

[0119] The applicant noted that beginners are more likely to make mistakes during performances, especially in terms of sequence, such as skipping a measure. Furthermore, different users handle these mistakes differently; some might replay the incorrect section, while others might continue playing in the usual order. In this case, traditional time-based matching mechanisms are insufficient to address the issue of incorrect playing.

[0120] Furthermore, from a user experience perspective, if there is a significant delay in the positioning marks on the sheet music, the continuity between the displayed content and the user's actual playing progress will be poor, impacting the user experience. Additionally, if the positioning marks can follow continuously on the sheet music, but misjudgments occur during rapid judgment, leading to incorrect display, this inconsistency will also affect the user experience. For beginners, whose familiarity with the sheet music is already limited, such discontinuities or errors can easily lead to less accurate rhythm control, potentially misleading the user and exacerbating the errors.

[0121] In response, the preferred dynamic matching method for audio and sheet music in this application is optimized at least in two aspects:

[0122] 1) A technical approach that uses a dual matching mechanism is adopted, which combines real-time matching results and switches back and forth between two routes: Matching Mechanism 1 (prioritizing matching one note at a time according to the conventional playing order) and Matching Mechanism 2 (setting a trial time to attempt to match the data sequence of multiple notes).

[0123] It is worth noting that the flexible switching of this dual matching mechanism can, on the one hand, control the amount of computation in each matching process to a certain extent, that is, limit the number of musical units to be matched and the matching interval (i.e. the length of the extension) to a certain extent; on the other hand, in the process of applying matching mechanism 1, it is preferable to use a single note for matching, which can also improve matching efficiency and thus ensure the real-time and continuous display of the positioning marker.

[0124] From another perspective, this dual-matching mechanism can comprehensively balance the marker update efficiency required for the continuity of the location marker display and the matching calculation requirements for display accuracy, thereby optimizing the user experience efficiency without increasing the processing pressure.

[0125] 2) The trial duration is flexibly set based on the user's actual playing level (such as the frequency of incorrect playing). This flexible setting of the trial duration can balance the real-time display and reliability of the positioning mark to a certain extent, and in particular, it can take into account the actual needs of users at different stages or of different types.

[0126] For example, when a user has a low level of familiarity with the current sheet music (e.g., the user is a beginner, or the user is practicing a new piece), the frequency of errors in playing the piece may be higher. If the user is found to be skipping or missing a piece, the real-time matching will be slowed down appropriately to accumulate a data sequence of a certain trial length before matching is performed, in order to avoid matching errors that could have a negative effect on the user.

[0127] At this point, the positioning marker on the score can remain in its current position area (e.g., at the position of the last note that successfully matched the target musical unit or in an adjacent area), instead of continuing to move backward along the staff lines. This temporary delay in following the music also helps to draw the user's attention, allowing them to verify the current performance from their own perspective.

[0128] For example, when a user is highly proficient in the current sheet music (e.g., the user is a professional pianist, or the user has practiced the sheet music for a long time), the frequency of errors during performance is generally low. In this case, even if a certain degree of error occurs, it can be quickly matched in the vicinity using a small amount of real-time data (i.e., note data sequence), thereby ensuring the continuity of the positioning marker display and avoiding interference with the user's playing rhythm (the continuous display at this time is more consistent with the user's playing rhythm, reducing the sense of interruption or conflict caused by errors on the display interface). At the same time, because the user is highly proficient, their self-correction ability is also relatively high. Even if an incorrect positioning occurs in special circumstances, the user can quickly adjust back to the normal state at a relatively fast pace.

[0129] Furthermore, in some embodiments, the similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes:

[0130] L = aX + bY + cZ;

[0131] L represents similarity, X represents pitch similarity, Y represents dynamic similarity, Z represents tempo similarity, a represents the first weight, b represents the second weight, and c represents the third weight.

[0132] It is worth noting that in this embodiment, a limited set of three types of data—pitch, duration, and intensity—are preferably selected from two dimensions: basic technique and emotional technique, to evaluate data similarity. Pitch and duration reflect the user's basic technical ability, while intensity reflects the user's grasp of emotion. Therefore, this dual-dimensional selection of factors can improve the reliability of data matching to a certain extent, while also controlling the computational load of data matching.

[0133] It is understood that the step numbers in this specification are only for the convenience of describing and understanding the solution, and are not intended to restrict the execution order of the method. For example, in some embodiments, S102 can be executed first, followed by S103. Alternatively, in other embodiments, when the real-time display requirements of the positioning marker are low, the order of S102-S103 can also be executed synchronously.

[0134] Example 2

[0135] See Figure 2 As shown, an audio-to-musical score matching system is characterized by comprising:

[0136] Reference data module 101 is used to provide a reference database of musical scores. The musical scores include: multiple markers for marking expected performance features. At least one of the markers constitutes a musical unit. Correspondingly, the reference database includes: multiple expected performance features corresponding to multiple musical units. The expected performance features include: expected pitch, expected dynamics, and expected duration. The musical unit is also associated with an expected performance order.

[0137] The feature acquisition module 102 is used to acquire the first note data generated by the user's performance at the first time. The note data is the actual performance feature corresponding to at least one marker. The actual performance feature includes: actual pitch, actual dynamics, and actual duration.

[0138] The first matching module 103 is used to determine, using a matching rule, whether the second note data generated at the second time adjacent to the first time successfully matches the first target music unit. The matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data.

[0139] If so, then entry is permitted:

[0140] Prediction module 104 is used to determine the first dynamic prediction order at the first time based on the first target music unit;

[0141] The expected data acquisition module 105 is used to acquire the first expected performance features corresponding to the first dynamic prediction sequence.

[0142] The second matching module 106 is used to determine, using the matching rules, whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; if so, the second target music unit is marked as the current performance node.

[0143] In some embodiments, when the determination result of the first matching module 103 is negative, and when the lengths of the first time and the second time satisfy a preset first trial time, the following modules of the system are allowed to enter:

[0144] The data sequence generation module 107 is used to generate a note data sequence based on the first note data and the adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target music unit;

[0145] A similarity calculation module 108 is used to calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit;

[0146] The tagging module 109 is used to tag the third music unit with the highest similarity as the third target music unit of the note data sequence.

[0147] In some embodiments, when the judgment result of the second matching module 106 is negative, the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched; and when the first time and the third time satisfy the first trial time, the similarity calculation module 108 and the tagging module 109 are allowed to enter.

[0148] In some embodiments, the desired performance characteristics include one or more of the following factors: desired pitch, desired dynamics, and desired duration. The actual performance characteristics may also include one or more of the following factors: actual pitch, actual dynamics, and actual duration. Correspondingly, when a performance characteristic is represented by one factor, the corresponding similarity can be the similarity (or overlap) of that single factor. Alternatively, when a performance characteristic is represented by two factors, the corresponding similarity can be the weighted sum of the similarities of the two single factors. It is understood that the specific weighting can be preset by the user based on the actual situation (such as the type of sheet music).

[0149] It is understood that the system described in this specification can be used to implement the methods or steps in any other method embodiment, which will not be repeated here.

[0150] Example 3

[0151] This invention provides a method for updating the page of a musical score, comprising the following steps:

[0152] A display page is provided in a first display mode. The display page is used to display partial content of the musical score. For example, the content of the score is the result of multiple written symbols being displayed on the page according to general music theory rules.

[0153] Determine whether the target music unit played by the user has reached the preset page update point. The page update point corresponds to a music unit, and the music unit consists of at least one identifier. The target music unit is also the user's current playing point, or the playing node.

[0154] If so, the displayed page will be updated using the first page-turning method.

[0155] In this embodiment, the user's current playing position can be located using the audio-musical score matching method employed in this invention.

[0156] In some embodiments, the first display method may include one or more of the following: single-page display, paginated display (such as two-page display, multi-page display).

[0157] For example, in some embodiments, a two-page display can be two pages displayed side by side. For example, displaying two pages side by side can simulate the state of an open music book.

[0158] For example, in some embodiments, the first page-turning method may include one or more of the following: scrolling, simulating book page turning (i.e., simulating the page changes when a user manually flips through a printed sheet music sheet), swiping, and overall page-turning updates. For example, when scrolling, the current page can scroll in a set direction to display the sheet music content of the next page in the area left by the current page. For example, swiping allows the current page to slide in a set direction, displaying the sheet music content of the next page in the area left by the current page. As another example, overall page-turning can directly update the entire sheet music content displayed on the current page.

[0159] In some embodiments, the page update method further includes: obtaining a user's trigger update signal; when the trigger update signal is received, defining a page update point based on the trigger update signal, for example, the page update point may also be the page turn time, thereby completing the page update.

[0160] In some embodiments, the page update point can be preset by the user. For example, the page update point can be a single syllable. Alternatively, the page update point can be defined by an update time. The update time can be a single point in time or an optional time range.

[0161] Preferably, in some embodiments, especially in single-page display mode, to reduce the interference of page turning on the user's playing smoothness, page turning will be selected between phrase intervals. A phrase interval refers to the space or pause between musical phrases in a musical work. In music composition and performance, the handling of phrase intervals is crucial to the smoothness and expressiveness of the music. A musical phrase is the basic structural unit that constitutes a piece of music, capable of expressing a relatively complete meaning. A musical phrase is similar to a sentence in an article, having the function of expressing a complete thought. In a musical work, a phrase usually consists of several measures, has a clear start and end point, and can express specific emotions or thoughts. A measure consists of at least one musical unit.

[0162] For example, in some embodiments, the method further includes the step of:

[0163] Find at least one musical phrase gap adjacent to the page update point (such as update time);

[0164] Calculate the time difference between the musical phrase interval and the page update point;

[0165] When the time difference is less than the set time threshold, the corresponding musical phrase gap will be set to the corrected page update point; otherwise, the initially set / obtained page update point can be retained.

[0166] Furthermore, in some embodiments, when there are two musical phrase gaps, the musical phrase gap with the smaller time difference can be selected as the corrected page update point.

[0167] In this embodiment, the page update points are corrected by adjusting the intervals between musical phrases, which can reduce the interference with the continuity of the user's performance during page turning to a certain extent.

[0168] In some embodiments, practice data from multiple practice sessions of the sheet music can be collected. This practice data includes page-turning timestamps on different display pages during different practice sessions. The page update point can be the page-turning timetamp selected by the user during the previous practice session.

[0169] Furthermore, in some embodiments, the practice data further includes:

[0170] For a given display page, record the frequency of incorrect performances in the immediate vicinity following at least one page turn.

[0171] When the frequency of incorrect performances exceeds a preset first error frequency threshold, the page update point is corrected. The corrected page update point is the page turn time plus a set delay time.

[0172] In this embodiment, to prevent users from having difficulty grasping the page-turning rhythm in the early stages of practicing, the page-turning time is optimized and corrected based on the performance quality after turning the page, thus optimizing the timing of page turns. In this embodiment, for single-page display mode, this optimization of page-turning time ensures the continuity of performance before and after turning the page, and to a certain extent avoids misplays during page-turning due to the user's unfamiliarity with the score.

[0173] Furthermore, in some embodiments, it also includes:

[0174] After correcting the page update points, new practice data from users is collected;

[0175] Furthermore, if the frequency of incorrect playing decreases in the immediate vicinity of the new practice data after the page update point, then you can choose to temporarily maintain the current page update point.

[0176] Specifically, if the frequency of incorrect playing decreases and falls below the set second error frequency threshold, the current page update point can be maintained. Alternatively, if the frequency of incorrect playing decreases, but after at least two practice sessions the frequency of incorrect playing is still greater than or equal to the second error frequency threshold, the page update point can be further adjusted. For example, the adjusted page update point in this case would be the previous page update point plus a set delay time.

[0177] For example, in some embodiments, the types of both the first display mode and the first page-turning mode can be selected by the user.

[0178] For example, in some embodiments, the practice data will also record the type of the currently used first display mode and / or the type of the first page-turning mode. In this case, if the user does not input new adjustment instructions during the current performance phase, the previously used first display mode and / or first page-turning mode can be directly adopted as the current default setting.

[0179] Further, see Figure 3 As shown, the present invention also provides a method for updating sheet music pages, the method comprising the steps of:

[0180] S201, a first display page and a second display page are provided, wherein the first display page is used to display a first partial score content of the score, and the second display page is used to display a second partial score content of the score, and the first partial score content and the second partial score content are continuous content in the score;

[0181] Preferably, the first display page and the second display page can be displayed side by side on the interface, such as the first and second display pages can be displayed vertically side by side.

[0182] S202, select the current display page and the page to be updated from the first display page and the second display page according to the user's current performance position (i.e. the target music unit matched / located at the current moment), wherein the current display page is the display page where the current performance position is located;

[0183] S204, determine whether the current performance point has reached the preset page update point, the page update point corresponds to a music unit, and the music unit consists of at least one identifier;

[0184] If so, proceed with the following steps:

[0185] S205, obtain the remaining performance time value of the currently displayed page and the user's performance speed;

[0186] For example, in some embodiments, the remaining performance time value can be calculated from the remaining unplayed markers (such as notes, rests, etc.) on the currently displayed page using music theory rules.

[0187] For example, in some embodiments, the performer's (equivalent to the user's) personal style and musical understanding may affect the playing speed, so the user's playing speed when playing the score can be recorded in real time.

[0188] For example, in some embodiments, the playing speed for sheet music corresponding to different musical styles may have a generally recommended range (such as a range that can be determined based on exemplary performance data in an existing music library). Therefore, when the performer has a limited amount of their own playing data, they can also refer to the playing speed determined by exemplary performance data.

[0189] S206, determine the expected performance duration of the currently displayed page after the page update point based on the remaining performance time value and the performance speed;

[0190] S207, determine the update speed of the page to be updated based on the expected performance duration, and perform a page update operation on the page to be updated within the expected performance duration.

[0191] It's important to note that note value generally refers to the length of time a note occupies, usually measured in beats. In music, a beat is the smallest unit of time used to measure the duration of a piece. Tempo, on the other hand, refers to the number of beats per unit of time; for example, tempo can be expressed as BPM (Beats Per Minute). Therefore, by dividing the remaining note value by the tempo, we can predict the expected playing time for the user to complete the remaining notes.

[0192] For example, dividing the time value by the playing speed gives the actual playing time of each symbol (such as a note), and the expected playing time is obtained by summing the actual playing times of multiple symbols.

[0193] For example, in some embodiments, the update speed refers to the magnitude of change in page content per unit time. Preferably, the page can be updated at a relatively uniform update speed for the expected performance duration.

[0194] In this embodiment, it is preferable to complete the page update operation on the page to be updated within the expected performance duration.

[0195] For example, in some embodiments, at least one display page is also associated with / set with a latest update point, meaning that the page should be updated no later than the latest update point.

[0196] For example, in this embodiment, the remaining performance time value can be the time interval between the page update point and the latest update point. That is, the remaining performance time value is determined by calculating the time value of the marker between the page update point and the latest update point.

[0197] Preferably, in some embodiments, in order to maintain the continuity of the user's performance, especially for users who are not very familiar with the performance content, the content of the current page needs to be updated only after the performance on the current page is completed.

[0198] In some embodiments, the method further includes: S203, using a separation display flag to distinguish between the currently displayed page and the page to be updated.

[0199] In some embodiments, the method further includes the step of:

[0200] When the current performance point switches to another display page, the selection between the current display page and the page to be updated is switched between the first display page and the second display page. That is, the first display page and the second display page will alternate sequentially between the current display page and the page to be updated.

[0201] In some embodiments, prior to S204, the following step is also included:

[0202] Determine whether a trigger update signal from the user is received before the first response time to the page update point;

[0203] If so, then define the page update point according to the trigger update signal;

[0204] If not, then obtain the user's historical update nodes and define the historical update nodes as the page update points.

[0205] In some embodiments, it also includes:

[0206] The current display page displays a positioning marker for locating the target syllable unit played by the user.

[0207] In some embodiments, the positioning marker moves along a first update direction on the current display page, and correspondingly, the updated spectrum content on the page to be updated is updated along a second update direction on the page to be updated, and the first update direction is opposite to the second update direction.

[0208] by Figure 4 , Figure 5 For example, the positioning marker can dynamically move along the direction of the staff lines in the score (i.e., the horizontal direction) to follow the performance nodes. Simultaneously, the update direction on the page to be updated can also be horizontal. This two-page mode, based on horizontal movement and updating, achieves a certain degree of visual harmony in the movement of the positioning marker and the changes in updated content. Furthermore, this two-page mode ensures that the current performance data display and subsequent performance data updates on the page can be synchronized. Therefore, even if the performer has a limited knowledge of the score, it is less likely to result in mistakes or omissions due to incomplete score display.

[0209] In some embodiments, the separating display mark includes one or more of the following: brightness display, indicator mark.

[0210] For example, in some embodiments, different brightness levels can be used to display the current display page and the page to be updated respectively, such as increasing the display brightness of the current display page to show the difference between the two.

[0211] For example, in some embodiments, the indicator mark may also be manifested as a difference in the color and mark size of the written symbols on the currently displayed page and the page to be updated, so as to visually distinguish the display effects of the two.

[0212] Alternatively, in some embodiments, the indicator may also be a display box, arrow, or other indicator form that can direct the previously displayed page to prompt the user to focus on key areas.

[0213] See Figures 4-5 As shown, it illustrates the effect of horizontal update changes based on a two-page display. Figure 4 The first display page 03 and the second display page 04 are shown, each displaying the musical score. It can be seen that... Figure 4In this display, the first display page 01 is the current display page. On the current display page, on the one hand, a positioning mark 01 is displayed at the target music unit played by the user to indicate the current playing progress to the user. On the other hand, a display frame 02 (equivalent to a separator display mark) is set outside the current display page to mark the current display page so that the user can more intuitively focus on the key score content.

[0214] Further, see Figure 5 As shown, when the target music unit played by the user reaches the second display page 04, the current display page and the page to be updated are switched between the first and second display pages. At this time, the positioning mark 01 moves to the new current display page (i.e., the second display page 04), while the page to be updated (i.e., the current first display page 03) will perform a page update operation.

[0215] At this point, the historical chart content 031 on the page to be updated will be gradually replaced with the new chart content 032.

[0216] Figure 5 An exemplary display of a sliding update is shown, wherein historical spectrum content 031 gradually moves along a second update direction, while new spectrum content 032 gradually fills the area left by the historical spectrum content 031. Preferably, an update marker 05 (such as a spectrum-free area of ​​a certain length) is used to separate the historical spectrum content 031 and the new spectrum content 032 to visually facilitate the user's distinction between the old and new content.

[0217] Example 4

[0218] See Figure 6 As shown, the present invention also provides a sheet music page updating system, comprising:

[0219] The page display module 201 is used to provide a first display page and a second display page, wherein the first display page is used to display a first partial score content of the musical score, and the second display page is used to display a second partial score content of the musical score, and the first partial score content and the second partial score content are continuous content in the musical score;

[0220] The page selection module 202 is used to select the current display page and the page to be updated from the first display page and the second display page according to the target music unit currently being played by the user. The current display page is the display page where the current playing point is located.

[0221] The separation display module 203 is used to distinguish between the currently displayed page and the page to be updated by using a separation display flag;

[0222] The update judgment module 204 is used to determine whether the target music unit has reached a preset page update point. The page update point corresponds to a music unit, and the music unit consists of at least one identifier.

[0223] If the result of the update judgment module 204 is yes, then entry is allowed:

[0224] The performance data acquisition module 205 is used to acquire the remaining performance time value of the currently displayed page and the user's performance speed;

[0225] The duration determination module 206 is used to determine the expected performance duration of the currently displayed page after the page update point based on the remaining performance time value and the performance speed.

[0226] The update module 207 is used to determine the update speed of the page to be updated based on the expected performance duration, and to perform page update operations on the page to be updated within the expected performance duration.

[0227] In some embodiments, the system further includes:

[0228] The page switching module is used to switch the selection of the current display page and the page to be updated between the first display page and the second display page when the target music unit switches to another display page.

[0229] The present invention also provides a computer-readable storage medium storing, in the form of computer-readable instructions, a computer program implemented according to the steps of the method described in any embodiment of the present invention, wherein the computer program, when invoked by a computer, performs the steps of the method described in any embodiment of the present invention. The present invention also provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described in any embodiment of the present invention.

[0230] It is understood that the system in this invention can implement the steps of the method described in any embodiment of this invention, which will not be repeated here.

[0231] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0232] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0233] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for matching audio with sheet music, characterized in that, Including the following steps: S101, a reference database of musical scores is provided, the musical scores including: multiple markers for marking expected performance features, at least one of the markers forming a musical unit, corresponding to: multiple expected performance features corresponding to multiple musical units, and the expected performance features including: expected pitch, expected dynamics, and expected duration, the musical unit is also associated with an expected performance order; S102, acquire the first note data generated by the user's performance at the first moment, the note data being the actual performance features corresponding to at least one marker, the actual performance features including: actual pitch, actual dynamics, and actual duration; S103, using a matching rule to determine whether the second note data generated at the adjacent second time before the first time successfully matches the first target music unit, wherein the matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data; If so, proceed with the following steps: S104, determine the first dynamic prediction order at the first time based on the first target music unit; S105, Obtain the corresponding first expected performance feature according to the first dynamic prediction sequence; S106, use the matching rule to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data; If so, mark the second target music unit as the current performance node; This also includes the following steps: Calculate the frequency of incorrect performances by the user during the first time period; wherein, when the judgment result of S106 is negative once, the number of incorrect performances is incremented by one; When the frequency of incorrect playing exceeds a first set number of times, the first trial time is extended; wherein, the first trial time has an initial value preset by the user; When the frequency of incorrect playing is greater than the second set number and less than or equal to the first set number, the first trial time is maintained. When the frequency of incorrect playing is less than or equal to the second set number of times, the first trial time is reduced.

2. The method according to claim 1, characterized in that, If the result of S103 is negative, then the following steps are executed: When the lengths of the first time and the second time satisfy the preset first trial time, then the following steps are executed: S108, Generate a note data sequence based on the first note data and the adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target music unit; S109, calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit; S110, the third music unit with the highest similarity is marked as the third target music unit of the note data sequence.

3. The method according to claim 2, characterized in that, It also includes the following steps: If the judgment result of S106 is negative, then the third note data at the third time after the first time is obtained, and the third note data is identified as the note data to be matched. When the first time and the third time satisfy the first trial time, then steps S108-S109 are executed.

4. The method according to claim 1, characterized in that, The marker is a musical note.

5. The method according to claim 1, characterized in that, One musical unit corresponds to one of the markers.

6. The method according to any one of claims 1-5, characterized in that, The similarity is calculated using a similarity evaluation model, wherein the similarity evaluation model includes: ; L represents similarity, X represents pitch similarity, Y represents dynamic similarity, Z represents tempo similarity, a represents the first weight, b represents the second weight, and c represents the third weight.

7. An audio-musical score matching system, characterized in that, include: The reference data module (101) is used to provide a reference database of musical scores, the musical scores including: multiple markers for marking expected performance features, at least one of the markers forming a musical unit, corresponding to the reference database including: multiple expected performance features corresponding to multiple musical units, and the expected performance features including: expected pitch, expected dynamics, and expected duration, the musical unit is also associated with an expected performance order; The feature acquisition module (102) is used to acquire the first note data generated by the user's performance at the first time. The note data is the actual performance feature corresponding to at least one marker. The actual performance feature includes: actual pitch, actual dynamics, and actual duration. The first matching module (103) is used to determine, by means of a matching rule, whether the second note data generated at the second time adjacent to the first time successfully matches the first target music unit. The matching rule is that when the similarity between the note data and the expected performance feature of the music unit is greater than a first set threshold, the corresponding music unit is considered to be the target music unit of the note data. If so, then entry is permitted: Prediction module (104) is used to determine the first dynamic prediction order at the first time based on the first target music unit; The expected data acquisition module (105) is used to acquire the corresponding first expected performance features according to the first dynamic prediction order; The second matching module (106) is used to determine whether the music unit corresponding to the first expected performance feature is the second target music unit of the first note data using the matching rules; if so, the second target music unit is marked as the current performance node. The system is also used to perform: Calculate the frequency of incorrect performances by the user during the first time period; wherein, when the judgment result of the second matching module is negative once, the number of incorrect performances is incremented by one; When the frequency of incorrect playing exceeds a first set number of times, the first trial time is extended; wherein, the first trial time has an initial value preset by the user; When the frequency of incorrect playing is greater than the second set number and less than or equal to the first set number, the first trial time is maintained. When the frequency of incorrect playing is less than or equal to the second set number of times, the first trial time is reduced.

8. The system according to claim 7, characterized in that, When the judgment result of the first matching module (103) is negative, and when the length of the first time and the second time meets the preset first trial time, then the following modules of the system are allowed to enter: The data sequence generation module (107) is used to generate a note data sequence based on the first note data and the adjacent note data to be matched, wherein the note data to be matched refers to at least one note data that has not been matched with the target music unit; A similarity calculation module (108) is used to calculate the similarity between the note data sequence and the expected performance features of at least one third musical unit; wherein, the third musical unit is at least one syllable unit in a historical period extending forward by a first length along the second dynamic prediction order, and / or, the third musical unit is at least one syllable unit in a future period extending backward by a second length along the second dynamic prediction order; the second dynamic prediction order is determined based on the note data of the most recently successfully matched target musical unit; The tagging module (109) is used to tag the third music unit with the highest similarity as the third target music unit of the note data sequence.

9. The system according to claim 8, characterized in that, When the judgment result of the second matching module (106) is negative, the third note data at the third time after the first time is obtained and the third note data is identified as the note data to be matched; and when the first time and the third time satisfy the first trial time, the similarity calculation module (108) and the tagging module (109) are allowed to enter.