Music score performance alignment for automatic performance evaluation

A predictive inference technique using a modified HMM aligns performed notes with score notes to provide qualitative feedback, addressing the limitations of existing music training applications by accurately identifying rhythm and pitch errors in real-time.

JP2025536678APending Publication Date: 2025-11-07MUSIC APP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025528585
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2023-11-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing computer-based music training applications fail to provide adequate qualitative feedback on a user's performance, often misinterpreting audio signals and lacking the ability to identify patterns in rhythm and pitch errors.

Method used

Implementing a predictive inference technique using a modified Hidden Markov Model (HMM) to estimate a maximum likelihood sequence of composite states, which includes pitch and timing deltas, to align performed notes with score notes and generate qualitative feedback.

Benefits of technology

Provides accurate and timely qualitative feedback on a user's performance, highlighting areas for improvement such as rhythm inaccuracies, pitch deviations, and missed notes, enhancing the learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536678000001_ABST
    Figure 2025536678000001_ABST
Patent Text Reader

Abstract

Techniques for performance alignment of musical scores for automatic performance evaluation are described. Embodiments receive performance note data defining the pitches and onset times of notes played as a user plays a musical score, where the musical score is computationally represented by score note data defining the pitches and onset times of the score notes. The played notes are automatically aligned to their respective score notes by calculating a maximum likelihood sequence of composite hand states for a sequence of time steps, each corresponding to the onset timing of a respective one of the score notes. Evaluation feedback, including qualitative evaluation feedback, is automatically generated by pattern matching the note-by-note alignment to an evaluation model. Some embodiments provide additional features, such as performing automatic alignment, evaluation, and feedback while the user is playing the musical score, in contexts such as two-handed playing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 425,715, filed November 16, 2022, which is incorporated herein by reference.

[0002]

[0002] Field Embodiments relate generally to music performance feedback applications, and more particularly to score performance alignment for automatic performance evaluation. [Background technology]

[0003]

[0003] A major value provided by any skilled teacher, including music teachers, is the ability to qualitatively evaluate a student's performance so that useful feedback can be clearly communicated. Existing computer-based applications for music training, such as learning to play the piano, are woefully deficient in this regard. Summary of the Invention

[0004]

[0004] Embodiments of the present invention relate to score performance alignment for automatic performance evaluation. For example, embodiments receive performance note data defining the pitches and onset times of notes played as a user plays a musical score, the musical score being computationally represented by score note data defining the pitches and onset times of the score notes. The performed notes are automatically aligned to their respective score notes by calculating a maximum likelihood sequence of composite hand states for a sequence of time steps, each corresponding to the onset timing of a respective one of the score notes. Evaluation feedback, including qualitative evaluation feedback, is automatically generated by pattern matching the note-by-note alignment to an evaluation model. Some embodiments provide additional features, such as performing automatic alignment, evaluation, and feedback while the user is playing the score, in contexts such as two-handed playing.

[0005]

[0005] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any and all drawings, and each claim.

[0006]

[0006] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0007]

[0007] The present disclosure is described in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1 illustrates an exemplary system supporting various embodiments described herein. [Figure 1B] FIG. 1 illustrates an exemplary system supporting various embodiments described herein. [Figure 2] FIG. 1 shows an example set of score notes and played notes in a "piano roll" layout, along with a table of associated time and pitch delta values. [Figure 3] FIG. 1 shows a plot of representative pitch and time delta of several played notes plotted on a two-dimensional plane against the pitch and timing of score notes. [Figure 4A] FIG. 1 illustrates an exemplary mode selection / discretization for exemplary data. [Figure 4B] FIG. 1 illustrates an exemplary mode selection / discretization for exemplary data. [Figure 5A] FIG. 1 illustrates an exemplary process flow for aligning performance notes to score notes according to embodiments described herein. [Figure 5B] FIG. 1 illustrates an exemplary process flow for aligning performance notes to score notes according to embodiments described herein. [Figure 6] FIG. 1 shows a piano roll format representation of a given time window, including a sequence of score notes (darker thin horizontal lines) and played notes (lighter thicker horizontal lines). [Figure 7A] FIG. 7 shows a kernel density plot on the timing-pitch delta hypothesis space for left-hand events, corresponding to time steps in the given time window of FIG. 6. [Figure 7B] FIG. 7 shows a kernel density plot on the timing-pitch delta hypothesis space for right-hand events, corresponding to time steps in the given time window of FIG. 6. [Figure 8] FIG. 1 is a schematic diagram of one embodiment of a computer system capable of implementing various system components and / or performing various steps of methods provided by various embodiments. [Figure 9A] FIG. 10 shows several plots representing an exemplary performance that is generally correct, except for a silence between approximately 5 and 7 seconds. [Figure 9B] FIG. 10 shows several plots representing an exemplary performance that is generally correct, except for a silence between approximately 5 and 7 seconds. [Figure 9C] FIG. 10 shows several plots representing an exemplary performance that is generally correct, except for a silence between approximately 5 and 7 seconds. [Figure 10A] FIG. 10 shows several plots representing an example performance with several problems, including the user playing an octave higher starting at 27 seconds, and the user playing several incorrect notes throughout the song. [Figure 10B] FIG. 10 shows several plots representing an example performance with several problems, including the user playing an octave higher starting at 27 seconds, and the user playing several incorrect notes throughout the song. [Figure 10C] FIG. 10 shows several plots representing an example performance with several problems, including the user playing an octave higher starting at 27 seconds, and the user playing several incorrect notes throughout the song. [Figure 11]FIG. 1 is a flow diagram of an exemplary method for score performance alignment for automatic performance evaluation, according to embodiments described herein.

[0009]

[0019] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a second label (e.g., lowercase) that distinguishes between the similar components. When only a first reference label is used herein, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION

[0010]

[0020] Embodiments of the disclosed technology will become clearer when considered in conjunction with the following description of the drawings. In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, those skilled in the art should recognize that the present invention may be practiced without these specific details. In some instances, circuits, structures, and techniques are not shown in detail to avoid obscuring the present invention.

[0011]

[0021] When learning an instrument such as the piano, it is beneficial to be able to receive feedback on one's performance. Some computer-based applications provide an environment in which a user can play a piece of music, and the application provides feedback. However, to date, these conventional applications have been inadequate in the amount and type of feedback provided. For example, some conventional applications provide the user with a sequence of notes to be played (written notes on a score). While and / or after the user plays the written notes on the score, the conventional application breaks down the performance into a sequence of played notes, each having at least an associated time and pitch. The conventional application determines whether each score note was played correctly by looking at it and determining whether there are any notes played with the same pitch and sufficiently close timing. The conventional application can then provide note-by-note feedback (e.g., by displaying correctly played written notes in green and incorrectly played score notes in red) and / or a quantitative measure of accuracy (e.g., the percentage of notes played correctly).

[0012]

[0022] While such note-by-note or quantitative feedback can be useful, it does not provide the user with any qualitative analysis. For example, a user may not be able to easily determine whether there is a problem with rhythm or pitch, or whether there is any pattern to the user's errors. As one example, a beginner user may play an entire passage perfectly, except that all the notes are an octave off from their correct position. As another example, a user may play all the notes with the correct pitch, but the rhythm is consistently incorrect. In both cases, conventional applications may mark all the notes as incorrect in a simplistic and unhelpful way, leaving the user with no clear guidance for improvement.

[0013]

[0023] Additional drawbacks arise in environments where an application "listens" to a user's performance using microphone input. Such applications traditionally perform signal processing on a received audio signal by decomposing the signal into discrete note events with pitch and timing, looking for changes in dominant frequency content and / or amplitude envelope. This process can be prone to transcription errors. For example, noise and / or other artifacts in the received audio signal can cause the application to recognize multiple played notes as a single note event, recognize a single played note as multiple note events, fail to recognize a note event, recognize phantom note events (e.g., by misinterpreting background noise as a note), etc.

[0014]

[0024] The techniques described herein use prediction techniques to estimate a maximum likelihood sequence of composite states given a sequence of written note events (in a presented musical score) and a sequence of played note events (received as an audio signal). Each composite state in the maximum likelihood sequence represents, at a corresponding time step (e.g., a score-based time reference), an estimated decision of whether the user is playing at that time step, and, if so, the delta between what the user is playing at that time step and what is represented in the musical score at that time step. The prediction technique can include applying a modified Hidden Markov Model (HMM) to generate and rank composite state hypotheses at each time step. Each composite state hypothesis can include a complex set of variables, including at least a played / unplayed variable and one or more continuous variables to represent pitch, timing, and / or other logical note characteristics (e.g., dynamics). In some embodiments, the Viterbi algorithm (e.g., or alternative methods such as particle filtering) is applied to calculate the most likely sequence of composite states given the score and the sequence of played notes. In some such embodiments, the HMM yields a very large number of complex-state hypotheses at each time step (e.g., theoretically infinite for continuous variables), and one or more clustering / mode discovery algorithms (e.g., mean shift, gradient-based hill climbing, etc.) are applied to reduce the number to a discrete set of N complex-state hypotheses, where N is a predetermined number (e.g., 10). In such embodiments, the Viterbi algorithm is applied to the discrete set of N complex-state hypotheses at each step. In some embodiments, N is a predetermined number (e.g., 10). In other embodiments, N varies from time step to time step but is bounded by a heuristic (e.g., by taking the 25th percentile of the maximum mode of the emission probability distribution at each time step). In such embodiments, the Viterbi algorithm is applied to a discrete and finite set of complex-state hypotheses. In some embodiments, the maximum likelihood sequence of estimated complex states is used to generate qualitative feedback on the user's performance.For example, information contained in the maximum likelihood sequence of estimated composite states directly conveys regions where the user did not play, regions where the user played consistently early or consistently late, or regions where the user played with systematic pitch deviations (e.g., when the user plays an entire sequence of notes in the wrong octave). The maximum likelihood sequence of estimated composite states also serves as a basis for further inference of individual note errors, more localized performance errors such as early or late note onsets, etc. In addition to highlighting errors, the maximum likelihood sequence of estimated composite states can also be used to formulate positive feedback (e.g., highlighting regions where the user played incorrectly).

[0015]

[0025] As known to those skilled in the art of musical performance, a "musical score" uses a form of music notation to represent a musical composition (i.e., a piece of music) as a sequence of scored events. In this specification, a musical score is assumed to provide a score-based time reference to the user. For example, a musical score may include a "4:4" time signature, indicating that each measure has four beats, with each beat corresponding to a quarter note. Thus, playing a piece of music with a metronome setting of 60 beats per minute involves playing the piece at a rate of 60 quarter notes per minute, or 15 measures per minute. Sounding events are notated in the musical score in the form of notes: these events are typically pitched and delimited in time by a start time and an offset time (the time positions where the pitched notes begin and end, respectively). Each score note can include an associated start time and offset time. As used herein, a "start time" is the time a score note begins, and an "offset time" is the time a score note ends. The relative horizontal placement of score notes on a staff can indicate the start time, and standard notation symbols specify the duration of the score notes, thereby indicating the offset time. Standard music notation includes large symbols for designating score notes and their durations. For example, various symbols are used to indicate that a note or rest occupies a duration such as a whole note, half note, quarter note, eighth note, dotted quarter note, double dotted quarter note, triplet eighth note, or quarter note tied to a sixteenth note. The vertical placement of a score note on a staff indicates its pitch. For example, a score note placed across the highest line of a staff marked with a treble clef sign has the pitch "F" in the fifth octave (F5). Note events that are notated in musical notation are referred to herein as "score notes."

[0016]

[0026] Proper performance of score notes involves applying notated rhythmic timing to a score-based time standard. For example, with a time signature of "4:4" and a metronome setting of 60 beats per minute, the correct rhythmic performance of the above example eighth note or rest would begin 7 seconds after the start of the piece (with a note onset at 7 seconds) and last a half second (i.e., with a note offset at 7.5 seconds). In some cases, the duration of the playing is determined by additional and / or alternative factors, such as the musical genre, the time signature, and additional musical notation (e.g., indicating that a particular passage should be played faster, that a passage should gradually slow down, etc.). By recognizing and understanding all musical notation, a musician can know when to play each score note and how long to hold each score note, thereby transforming notated music into played music.

[0017]

[0027] The ability to convert score notes into performed notes, especially in real time, can involve extensive training and practice. Even skilled musicians can struggle to play complex and / or unfamiliar sequences of score notes with consistently accurate timing and pitch. While several conventional approaches exist for automatically evaluating a user's performance, these approaches tend to be limited in several ways. One limitation is that such conventional performance evaluation systems tend to focus on an overall qualification score (good, bad) or provide note-by-note correct / incorrect feedback indications. Another limitation is that such conventional systems typically assume that the performance is a roughly accurate rendering of the score (e.g., entire passages are not transposed or skipped). Thus, these conventional systems tend to be unable to provide useful qualitative feedback to the performer.

[0018]

[0028] Embodiments herein provide novel techniques for using predictive inference of complex-state sequences to automatically generate accurate transcriptions of a user's performance of a sequence of score notes and / or to automatically generate qualitative assessments of the user's performance. For example, embodiments herein assume that at any given time during a user's performance, the state of the performance can be characterized by either complete silence (e.g., no notes played by either hand), partial silence in one or both hands, complete or partial pitch transposition in one or both hands, or complete or partial time shifts in one or both hands (within a relatively small window due to the in-time nature of the performance). Detection of these states can be used to provide qualitative feedback regarding where the user should focus their attention and / or invest practice time to improve their playing. This can include marking areas where the user did not play, where pitch transpositions occurred, where temporal (rhythmic) inaccuracies occurred, or where individual oversights or incorrect notes were made.

[0019]

[0029] 1A and 1B, an exemplary system 100 supporting various embodiments described herein is shown. Embodiments of system 100 use predictive techniques to estimate a maximum likelihood composite-state sequence from a user's performance (e.g., by playing a piano or other instrument) of a sequence of events notated on musical notation. Some embodiments of system 100 are adapted to generate a posteriori feedback by waiting until the user's performance is complete and then processing the entire sequence of performed notes and deriving qualitative feedback from the sequence. Other embodiments of system 100 are adapted to generate dynamic feedback by processing and evaluating the performed notes substantially in real time to provide qualitative feedback simultaneously with the user's performance. System 100 may be referred to as an automatic performance evaluation system, and such a system is understood to also perform automatic score performance alignment and other functions described herein.

[0020]

[0030] The version of system 100a shown in FIG. 1A performs processing of the user's performance audio stream within device I / O subsystem 105 and sends the performed notes to evaluator subsystem 150 for evaluation. The version of system 100b shown in FIG. 1B performs processing of the user's performance audio stream within evaluator subsystem 150. Similar components in the two versions of system 100 are labeled with similar reference labels and operate in the same manner unless otherwise noted. Thus, features of the two versions of the environment can be combined in any appropriate manner. Alternative embodiments can include both versions of the environment, selectable automatically and / or by the user as needed. For example, such a system can provide a particular type of qualitative feedback while the user is playing (e.g., to help the user make dynamic corrections), and the system can provide additional and / or different types of qualitative feedback after the performance is complete. In some embodiments, system 100a implements note processor 135 for post-feedback to generate a sequence of performed notes representing the user's entire performance before sending the sequence of performed notes to evaluation engine 155. In some embodiments, the system 100b implements a note processor 135 for dynamic feedback to dynamically process the user's performance and send one or more performed notes at a time to the evaluation engine 155 to support the dynamic presentation of qualitative feedback.

[0021]

[0031] System 100 includes a device input / output (I / O) subsystem 105, an evaluator subsystem 150, and a musical score (MS) data store 140. An embodiment of device I / O subsystem 105 includes a display interface 120 having a display processor 125 and an audio interface 130 having a note processor 135. An embodiment of system 100 is implemented on a computing device. In some embodiments, the computing device is a portable electronic device such as a laptop computer, a tablet computer, or a smartphone. In other embodiments, the computing device is an appliance such as a desktop computer. In other embodiments, the computing device is integrated with a musical instrument such as an electronic piano.

[0022]

[0032] Details of the musical score, including the sequence of score notes, are stored by MS data store 140. As shown, MS data store 140 stores at least MS visual data 145 and MS logical data 143. In some embodiments, MS visual data 145 includes all information used by display processor 125 to generate a graphical score representation, including a graphical representation of the score notes, for visual output by display interface 120. For example, MS visual data 145 includes a stored graphical representation of the score, including notes and rests, as well as any additional graphical score elements (e.g., staff lines, bar lines, clefs, accidentals, time signatures, key signatures, dynamic markings, tempo markings, expressive markings, titles, lyrics, etc.). In other embodiments, MS visual data 145 includes only a portion of the information used by display processor 125 to generate a visual output of the musical score by display interface 120. For example, the MS visual data 145 includes score elements other than notes and rests, and the display processor 125 generates visual representations of the notes and rests based on the MS logical data 143 of the score notes.

[0023]

[0033] MS logical data 143 logically defines a sequence of score notes, indicating at least the associated pitch, start time and offset time (and / or duration, note type, etc.) for each score note. In some implementations, MS logical data 143 includes other information about particular score notes, such as the note's pitch, whether the score note is part of a triplet or other tuplet, whether the score note is associated with a particular hand or finger, etc. In some embodiments, MS logical data 143 is score-compliant, such that the start time and / or offset time (and / or duration) are defined with respect to the underlying musical notation. The relationship to the score can be based on the type of note or rest (e.g., "quarter note"), the consumption fraction of one or more measures (e.g., half a measure), the number of metronome beats (e.g., three beats), etc. In the exemplary score notes, the start time is indicated as occurring on the second beat of the fourth measure, and the offset time is indicated as occurring on the fourth beat of the fourth measure. If necessary, the offset time can be converted to a duration using additional MS logical data 143. For example, MS logical data 143 may indicate that the musical notation is in "4 / 4" time, from which it may be derived that the exemplary score note is a half note. Additionally or alternatively, the offset time can be indicated as a duration in score-referenced MS logical data 143. Referring to the same exemplary score note, MS logical data 143 may indicate that the score note has a duration of two beats or that the score note is a half note. In such cases, the SNR offset timing can be derived, if desired, by adding the explicitly indicated duration to the start time. In either case, tempo can be used to convert score-referenced MS logical data 143 to real-time-referenced timing information. For example, additional MS logical data 143 may indicate a default tempo of 120 beats per minute, with a score-referenced half note consuming approximately one second of real time at the default tempo. The score-referenced definition of MS logical data 143 provides various features. For example, the real-time duration of all score notes can automatically adjust with changes in tempo.

[0024]

[0034] Alternatively, MS logic data 143 can be real-time based. In one implementation, each score note is defined as beginning a certain number of milliseconds from the beginning of the music (i.e., start time) and ending a certain number of milliseconds from the beginning of the music (i.e., offset time). In another implementation, each score note is defined as beginning a certain number of milliseconds from the beginning of the music (i.e., start time) and ending a certain number of milliseconds after the start time (i.e., offset time). Whether MS logic data 143 is score-based or real-time based, offset time is defined as the score note's associated end time or as a duration following the score note's associated start time.

[0025]

[0035] In some embodiments, display processor 125 generates the graphical representation of the musical notation and score notes based solely on MS visual data 145, without any processing of MS logical data 143. In other such embodiments, display processor 125 generates the graphical representation of the score notes partially based on MS logical data 143 (e.g., determining the horizontal placement of notes and rests corresponding to score notes based on start times, determining the type of note or rest based on offset times, etc.), while the remaining graphical representation of the musical notation and / or score note elements is generated from MS visual data 145. Some embodiments of display processor 125 include a notation rule engine (not shown) that can derive additional information from MS visual data 145 and / or MS logical data 143 that can be used to generate the graphical representation of the musical notation and / or score note elements. For example, the notation rule engine can determine whether a note has a flag, whether a note is tied across a bar boundary, whether a note's stem is pointing up or down, whether an accidental is indicated, etc.

[0026]

[0036] Embodiments of display interface 120 and display processor 125 are implemented in any suitable manner that supports the visual portion of the audiovisual output functionality described herein. The terms “visual” and “graphical” are used interchangeably herein. Display interface 120 may be implemented by any suitable display component, such as a tablet computer display, a smartphone display, a computer monitor, or the like. Display processor 125 may be implemented as any processor, portion of a processor, or group of processors capable of generating graphical output via display interface 120, as described herein. For example, display processor 125 may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), or the like. As described above, display processor 125 may use MS visual data 145 (e.g., and MS logic data 143) to generate a graphical representation of the musical score, including a graphical representation of the sequence of score notes played by the user.

[0027]

[0037] Embodiments of display processor 125 provide additional features. Some such features involve generating display output to graphically represent real-time musical score progression, such as by graphically indicating the current playback position. Other such features involve generating display output that graphically represents evaluation feedback. These and other features of display interface 120 and display processor 125 are described further below.

[0028]

[0038] Embodiments of audio interface 130 and note processor 135 may be implemented in any suitable manner that supports the aural portion of the audiovisual output functionality described herein. The terms “aural” and “audio” are used interchangeably herein. Audio interface 130 may be implemented by any suitable audio components, such as one or more audio transducers, speakers, headphones, etc. Note processor 135 may be implemented as any processor, portion of a processor, or group of processors capable of generating audio output via audio interface 130 as described herein. In some embodiments, audio representations of score notes are dynamically generated from MS logic data 143. In some such embodiments, the audio representation of a sequence of score notes indicates only the start times, such as by buzzing, vibrating, or clicking at each start time. In other such embodiments, the audio representation further indicates the offset times of the score notes (e.g., audio interface 130 buzzes or vibrates for the duration between the start and offset of each score note). In other such embodiments where MS logical data 143 includes additional audio information, the audio representation of the score notes indicates further audio information such as pitch, dynamics, etc. In other embodiments, the audio representation of the score notes is generated from recorded audio, for example, a sequence of score notes is played by playing a digital audio file from MS data store 140 that stores a recording of the musical piece.

[0029]

[0039] Embodiments of the note processor 135 provide additional features. Some such features involve generating additional audio during a user's performance, such as outputting a metronome and / or backing track to indicate tempo. When a user plays a sequence of score notes in tempo, the score notes and the notes played by the user can be easily registered to a common time reference (e.g., a score-based time reference). Other such features involve generating audio output to play back portions of a user's performance, such as for comparison with playback of an ideal version of a piece of music, a recorded version of a piece of music, a past performance by the user, etc.

[0030]

[0040] As a user plays an entire musical piece (a sequence of score notes represented on a musical staff), an embodiment receives audio of the user's performance via audio interface 130 and processes the audio via note processor 135. For example, a microphone may be used to convert the user's performance into an audio signal. Additionally or alternatively, the user may play on an instrument configured to generate output in Musical Instrument Digital Interface (MIDI) format, and device I / O subsystem 105 may include a MIDI input for receiving the played notes as digital event data. In such cases, audio interface 130 may include a MIDI interface, and note processor 135 may include MIDI processing capabilities. In cases such as MIDI, the digital data itself may indicate a sequence of musical events (e.g., note events) played by the user. In the absence of MIDI, the raw audio information is processed to extract a sequence of musical events corresponding to the user's performance. This extraction may occur in note processor 135 of the device I / O subsystem or in evaluation engine 155 as part of evaluator subsystem 150. Thus, the information sent from the note processor 135 to the evaluation engine 155 is a representation of the user's performance either in the form of raw / processed audio or in the form of a sequence of discrete musical events.

[0031]

[0041] As described herein, embodiments attempt to find the most likely set of played notes from a musical score and an audio signal representation of a user's performance, based on a sequence of known score notes. The sequence of score notes is received from MS data store 140, and the sequence of played notes is received from device I / O subsystem 105 (e.g., from note processor 135). A score-based time reference is used to define a series of time steps. For example, the user may play along with a metronome or rhythmic backing track to establish a shared time reference between the sequence of score notes and the user's performance.

[0032]

[0042] Embodiments can generate a sequence of played notes in one or more ways, depending on the type of input received via audio interface 130. The played notes may be received as a real-time audio stream, which may be a stream of raw analog audio information (e.g., received via a microphone or audio cable), a stream of raw digital audio information (e.g., an analog stream converted to a digital data stream), or a MIDI data stream received via a MIDI interface. References herein to “MIDI,” “MIDI stream,” “MIDI interface,” etc., may be extended to generally include any type of specialized audio encoding format that encodes at least the pitch, onset time, and offset time or duration of each note event. In contrast, references to a “stream of raw analog audio information” or a “stream of raw digital audio information” generally include any stream of analog or digital information that represents a user's audio performance but does not encode individual note events.

[0033]

[0043] In the case of a MIDI stream, the MIDI processor 137 of the note processor 135 takes the MIDI stream as input and generates an output representing a stream of played notes (each having at least a corresponding pitch, start time, and offset time or duration). In the case of a raw analog or digital audio stream, the received stream can be segmented by the note processor 135 for real-time processing. For example, the note processor 135 can divide the received audio stream into overlapping windows of fixed duration (e.g., 0.3 seconds), and each overlapping window can be fed to the transcription engine 139. For each overlapping window, the output of the transcription engine 139 represents the set of notes that started and the set of notes that ended during that window, if any. The note processor 135 can assemble the sets of notes from the transcription engine 139 into a stream of played notes (each having at least a corresponding pitch, start time, and offset time or duration). Thus, regardless of the form of the received real-time audio stream, the components of the note processor 135 are able to generate an output that represents the stream of played notes.

[0034]

[0044] In post-feedback embodiments, performance evaluation occurs after a performance is completed. In such cases, an embodiment of the note processor 135 may assemble the stream of played notes into a sequence of played notes that represents the entire performance. The sequence of played notes may be evaluated by the evaluation engine 155 on a segment-by-segment basis and / or in its entirety.

[0035]

[0045] In dynamic feedback embodiments, the performance evaluation is dynamic, based on real-time processing and evaluation of the performed notes. As used herein, "real-time" or "dynamic" is intended to mean concurrent with the user's performance, rather than after the user's performance is completed. In some implementations, the note processor 135 generates the performed notes and sends them to the evaluation engine 155 substantially continuously. For example, as each set of one or a small number of performed notes is generated, the set is sent to the evaluation engine 155. In some such implementations, the performed notes are sent based on a moving window, such as the same moving window used by the transcription engine 139 to generate the performed notes. In other implementations, the note processor 135 sends a sequence of played notes to the evaluation engine 155 based on one or more send trigger events, such as periodically (e.g., according to a fixed frequency such as once per second), whenever a certain number of played notes is reached, whenever a certain portion of the score elapses (e.g., once per measure), whenever a buffer of played note data is full, or whenever the evaluation engine 155 finishes processing the previous batch of played notes.

[0036]

[0046] An embodiment of the evaluation engine 155 can process the performed notes as they are received from the note processor 135. In some implementations, the evaluation engine 155 evaluates the performed notes based on a moving evaluation window. The evaluation window can be longer, shorter, or the same duration as a moving window used for transcription and / or transmission of the performed notes. For example, performed notes may be sent to the evaluation engine 155 by the note processor 135 multiple times per second, but the evaluation engine 155 may evaluate received performed notes only once every five seconds. In other implementations, the evaluation engine 155 performs evaluation of a sequence of performed notes periodically, based on one or more evaluation trigger events, such as after a certain number of performed notes are received, after a certain portion of the score has elapsed, or after a buffer of performed note data has filled. The evaluation trigger events used by the evaluation engine 155 may be the same as or different from those used by the note processor 135.

[0037]

[0047] Additionally, different moving window durations and / or evaluation triggers can be used for different types of evaluations. As one example, evaluating a moving window of three consecutive played notes can be useful for determining whether a user appears to be playing in the wrong octave, while evaluating a moving window of ten played notes can be useful for determining whether a user tends to play ahead or behind the beat. As another example, a particular evaluation may only occur after a certain number of occurrences of a particular occasion, such as a certain number of times the score indicates an accidental or a certain number of times the score indicates a mordent. Also, certain types of evaluations may only be performed at certain times and / or under certain conditions. As one example, an embodiment may evaluate whether a user completely fails to begin playing, but such evaluation is only relevant at the beginning of a performance. As another example, a particular evaluation may only be relevant when a user is playing with both hands. As another example, a particular evaluation may only be relevant when the score indicates a tempo change, a key signature change, a dynamic change, an accidental, etc.

[0038]

[0048] An embodiment of the evaluation engine 155 generates a maximum likelihood sequence of composite states. Each composite state in the sequence is generated based on a set of hypotheses. At each time step, an embodiment generates a playing hypothesis indicating whether the user is playing. If the user is assumed to be playing, an embodiment further generates a delta hypothesis indicating a hypothesis of the delta between what is being played and the corresponding score note. The delta hypothesis may indicate a timing delta, a pitch delta, and / or some other musically related delta (e.g., a dynamic delta). For example, a non-zero pitch delta indicates the hypothesis that the user is playing an incorrect pitch, and a non-zero timing delta indicates the hypothesis that the user is playing behind or ahead of a score-based time reference (e.g., corresponding to a metronome). One or more performance errors are determined based on the delta hypotheses, and one or more qualitative evaluations are made based on the performance errors. For example, determining a consistent pitch offset may qualitatively indicate that the user played a passage in the wrong octave, that the user consistently played a particular accidental incorrectly, etc.

[0039]

[0049] In some cases, a performance involves multiple simultaneous streams of auditory information, such as playing with both hands. In some such cases, a separate hypothesis can be generated for each track. For example, a maximum likelihood sequence of combined states is generated for the sequence of time steps for each hand. In other cases, a joint hypothesis is calculated for the sequence of combined states for both hands. While calculating the maximum likelihood sequence of combined hand states involves more computation, it is likely to give better results by avoiding inconsistent hypotheses between hands (e.g., both hands playing the same note).

[0040]

[0050] Embodiments operate in the context of a user performing with temporal guidance. For example, in a performance assessment application, playback of a musical piece is accompanied by a metronome and / or a display indicating the current score-based timing (e.g., a moving vertical line superimposed on the score). In this manner, embodiments can formulate hypotheses based on the assumption that the timeline of the user's performance is shared with the score-based timeline (i.e., the user is playing the music "in time"). Of course, features of the performance assessment system can allow the user to play the piece at different tempo settings (i.e., faster and slower speeds), with the score-based timeline adjusted accordingly.

[0041]

[0051] At each point in the performance, embodiments hypothesize whether a note is being played. In some embodiments, the hypotheses are whether a note is being played by each hand. Once a note is hypothesized to be played, embodiments further hypothesize a particular pitch delta and time delta between the hypothesized note and the corresponding score note. A non-zero pitch delta indicates that the user is playing an incorrect pitch, and a non-zero time delta indicates that the user is behind or ahead of the score-based timeline (e.g., ahead or behind the metronome). Further identification of note-level errors in the performance is determined based on the played / not played and pitch / time delta hypotheses, as described herein.

[0042]

[0052] Implementations of embodiments herein involve finding an automated technical approach to address the technical problem of finding the most likely state sequence given two sequences of notes: a sequence of score notes and a sequence of performed notes, where the state represents whether a user is playing a part / hand at a given time, and if so, the pitch / time delta the user is playing. It can be reasonably assumed that the observation probability at a given time step (i.e., the performed notes observed before and after this time step) is independent of the states at other time steps, given the state at the current time step (i.e., whether / how the user is playing at this time step). Thus, embodiments can formulate the technical problem as a Hidden Markov Model (HMM), where the state is unknown, the sequence of score notes is given as background information, and the performed notes are the raw observations for which it is desired to infer the most likely sequence of states.

[0043]

[0053] The approach implemented by the embodiments herein can provide several novel features. One novel feature is that the embodiments combine raw observations (played notes) with background information (score notes) to form observations (sets of time and pitch deltas) that are described by states. Another novel feature is that the state space is not constrained to be fully discrete or fully real-valued, but rather can be a hybrid between the two. Another novel feature is that the state space can be dynamically discretized at each time step to allow tractable inference using the Viterbi algorithm.

[0044]

[0054] As described herein, it may be desirable to provide qualitative feedback to the user during and / or after a performance, and embodiments herein provide such feedback based on aligning a performed note sequence (e.g., as performed by the user) with a score note sequence (as represented on the musical score). Alignment is performed by formulating a technical model (e.g., an HMM) that represents the problem, which may begin with formulating a technical definition of the notes and discretizing a time axis shared by the score and the performed note sequence.

[0045]

[0055] For the purposes of such alignment, both score notes and performed notes can be considered as onset-pitch pairs. For example, onset times can be expressed in seconds, and pitches can be expressed as a range of pitch identifiers from A0 to C8, as MIDI note numbers from 21 to 108, and / or in any other suitable manner. Implementations can ignore other (e.g., score-specific) attributes such as note duration and / or enharmonic spelling or articulation marking. We therefore define a score S and a performance P as sequences of (onset, pitch) pairs as follows: S=(s1,...,s K )(1) P=(p1,...,p L )(2) In the formula, s i ,p i ∈R×[21,108].

[0046]

[0056] Given a score S, we consider the Markov process {H i}, so that the HMM describes the sequence of unique note onset times of the score notes as follows: {H i}:t i ∈{onset(s)|s∈S} and ti <t i+1 (3) Therefore, t i is the H at time step i i is the time associated with

[0047]

[0057] Rather than treating played notes as observations directly, embodiments calculate the delta of a played note relative to a score note in terms of start time and pitch. More specifically, for a score note s and a played note p, the time and pitch (π) deltas are defined as follows, respectively:

number

[0048]

[0058] At this stage it is not known which score notes (if any) are associated with each played note, so time and pitch deltas are calculated between the score note and all played notes that fall within a given time and pitch delta range. This can be expressed as:

number

[0049]

[0059] For example, the time delta range is

number

number

number

[0050]

[0060] FIG. 2 shows an example set of score notes and performed notes in a "piano roll" layout 210, along with a table 220 of associated time and pitch delta values. In the piano roll layout, the rows represent time and the columns represent pitch. Pitch and time deltas can be calculated between each score note and any performed note that falls within a predetermined time and pitch delta range, as described above. As shown, for a particular score note (s_0), three performed notes (p_1, p_2, and p_3) fall within a predetermined time and pitch delta range. As one illustrative example, between score note s_0 and performed note p_1, there is a calculated pitch delta (dp_01) and a calculated time delta (dt_01). As another illustrative example, between score note s_0 and performed note p_3, there is a calculated pitch delta (dp_03) and a calculated time delta (dt_03).

[0051]

[0061] The calculated pitches and time deltas can be plotted on a two-dimensional plane. For example, Figure 3 shows a plot 300 of representative pitches and time deltas of several performed notes plotted on a two-dimensional plane against the pitches and timings of score notes. The pitches and timings of score notes correspond to position (0,0) on plot 300. Such a plot can be used to automatically determine regions of performed notes that fall within predetermined time and pitch delta ranges as candidate performed notes for alignment with the score notes.

[0052]

[0062] Embodiments generate a variable number of deltas for each score note, possibly zero (e.g., if no nearby notes were played). Even though the Markov process itself allows for state changes at each time step, embodiments account for the practical nature of the activity being modeled. In particular, it is recognized that the state of the hand being played evolves at a slower rate than the time steps of the Markov process. Thus, for example, it is useful to view an otherwise correctly played individual missed note as a mistake during a "played" state, rather than as a rapid transition from a "played" state to a "not played" state. For this reason, t i Instead of defining the observation at time step i as only the delta calculated from the note starting at t i More specifically, we can define a parameter ρ (an integer value) such that for each time step i and hand ∈ L,R, the deltas from the surrounding notes are aggregated. i This subset of score notes is then

number

number

[0053]

[0063] Observation at time step i

number

number

[0054]

[0064] Note that the weight associated with a pair of time / pitch delta values ​​is the weight of the score note from which the delta was calculated (e.g., as shown in table 220 of FIG. 2). For example, table 220 of FIG. 2 shows the observations as an array, while plot 300 of FIG. 3 shows the observations projected onto a 2D plane. The complete observations at step i are the hand-specific pair of observations:

number

[0055]

[0065] We can now define states. A state H represents the state of both hands at a given time. A hand can be either in a "not playing" state (represented by the empty set φ), or in a "playing" state with time and pitch deltas that fall within specified ranges. Thus, separate state spaces H for the left and right hands can be defined as follows: H (L) ,H (R )∈H:{φ}∪Δ One-handed state (12)

[0056]

[0066] Pitch delta value δ (p) can only take integer values, but to facilitate uniform treatment of time and pitch deltas, the possible pitch deltas Δ (p) We can define the space of H as a subset of the real numbers, which is convenient in the inference process. We can then represent the combined left and right hand states H as elements of a Cartesian product of the single hand state spaces: H=(H (L) ,H (R) )∈H×H Combined hand state (13)

[0057]

[0067] Toward the goal of achieving alignment between score note sequences and performance note sequences, embodiments recognize that such alignment corresponds to the best explanation of the observations represented by a sequence of hand states. In other words, embodiments recognize that the sequence of hand states H that best explains observation O. * Under the HMM formulation, this sequence can be characterized as follows:

number

[0058]

[0068] Since the state space is not a finite, discrete set, we use the Viterbi algorithm to find H * It is not possible to find H * Finding H over time is generally intractable, even though particle filter methods are typically defined over a purely real-valued state space. * One approach is to use so-called "particle filtering" to track the set of candidates. Alternatively, to support the application of the Viterbi algorithm, some embodiments dynamically discretize the state space at each time step. Such dynamic discretization preserves the advantages of a continuous state space (where delta values ​​can be accurately modeled at each time step without quantization) while allowing efficient inference using the Viterbi algorithm. The terms "Viterbi path" and "Viterbi alignment" are used generally herein to refer to H * Refers to...

[0059]

[0069] The state space defined in equations (12, 13) contains discrete terms, i.e., "unplayed" values ​​φ, and a continuous time / pitch delta space Δ. A simple approach to discretizing the time / pitch delta space is to use a predefined fixed quantization grid. The drawback of this approach is that there is an inherent trade-off between the precision of the state representation (the coarser the quantization, the larger the quantization error) and the computational burden of inference (finer quantization results in a larger state space).

[0060]

[0070] Since for any given time step i and hand, there are typically only a small number of possible candidate states in time / pitch delta space, a viable approach is to identify the most likely candidate states and classify them, along with "non-playing" states, as individual hand states.

number

number

[0061]

[0071] Observed value O i Use weighted kernel density estimation (KDE) on σ to identify candidate states from the time / pitch delta space at step i. i The KDE of O at point H i A non-negative function that returns the data density of

number

number

[0062]

[0072] Because deltas are calculated for many pairs of nearby score notes and played notes, delta values ​​are generally scattered around the space. However, when a user plays at a particular time and pitch delta relative to the score, the deltas between the score notes and their corresponding played notes accumulate around that point in delta space, forming a peak / mode in the KDE.

[0063]

[0073] A complicating factor is that repetitive patterns in the score (e.g., sequences of notes with the same duration and alternating pitches) tend to generate modes caused by deltas between non-corresponding score notes and performed notes, potentially causing misalignment between the score and the performance. To combat this phenomenon, embodiments may create a set of "negative" observations as follows:

number

number

[0064]

[0074] density

number

number

number

number

[0065]

[0075] There is no analytical method for finding the mode of a KDE (even when using a Gaussian kernel), but there are iterative methods such as the so-called "mean shift algorithm". i For computational efficiency, we use a mean shift algorithm seeded with density f i The gradient of is used as the shift vector for the candidate, rather than calculating the shift vector from the candidate's neighbors (which can be calculated analytically because we use a Gaussian kernel). Furthermore, because embodiments are only interested in the largest modes (rather than all modes), embodiments can discard any candidates that fall below a threshold (e.g., 0.5 times the density value of the current best candidate). For typical kernel bandwidth values, such an approach typically results in 1-5 candidates in fewer than 10 mean shift iterations.

[0066]

[0076] Figures 4A and 4B show exemplary mode selection / discretization for example data. An exemplary set of eight iterations (labeled "Iteration 1" through "Iteration 8") demonstrates iterative mode discovery in Gaussian kernel density estimation (KDE) of observed pitch / time deltas. Figure 4A shows Iterations 1-4, and Figure 4B shows Iterations 5-8. Each iteration is shown as a plot of pitch delta (in semitones) versus time delta (in seconds). In each plot, one or more "X" indicators represent the candidate pitch / time deltas for the iteration, and a heat map around those pitch / time deltas represents the KDE.

[0067]

[0077] Before iteration 1, all observed pitch / time deltas are potential candidates. In iteration 1, low density candidates are filtered out. In subsequent iterations, candidates are alternately moved along the gradient of the KDE towards local maxima and clustered. The iterations shown in Figures 4A and 4B are performed by generating a state H = (δ t ,δ p ) can be an example of a Gaussian KDE P(O|H). Embodiments use the modes (or a subset of the modes) of this distribution P(O|H) as the most likely hypotheses, typically resulting in a small number of hypotheses, for each of which a likelihood P(O|H) is calculated and used in the equations provided below (Equations (18) and (19)).

[0068]

[0078] As described herein, embodiments hypothesize each composite state based on at least one continuous delta variable. Thus, hypothesis ranking, including transition probability estimation, can include local smoothing and / or normalization. For example, suppose a score note corresponds to a D-flat in octave 4 on a piano. A first hypothesis is that the user played a D-natural in octave 4, and a second hypothesis is that the user played a D-natural in octave 5. Furthermore, suppose several preceding notes are all in octave 4, and the user played them all in octave 5. Without local smoothing, the first hypothesis would have a calculated pitch offset of +1, and the second hypothesis would have a calculated pitch offset of +13. Due to the continuous nature of the pitch delta variable, local smoothing can effectively treat the second hypothesis as having a pitch offset of +1 for a local transposition of +12 (i.e., an octave).

[0069]

[0079] Some embodiments may use the observation likelihood P(O|H)=P(O (L) ,O (R) |H (L) ,H (R)) can be defined to represent the likelihood of left and right hand observations given the left and right hand states. By simplifying assumptions, P(O (L) |H (L) ,H (R) )=P(O (L) |H (L) ) and P(O (R) |H (L) ,H (R) )=P(O (R) |H (R) ), where: P(O│H)=P(O (L) │H (L) )P(O (R) │H (R) )(18)

[0070]

[0080] The term P(O (L) |H (L) ) and P(O (R) |H (R) ) is defined as follows:

number

number

[0071]

[0081] The other components of the HMM defined to compute the Viterbi path are the transition probabilities P(H i |H i-1 ), where we can simplify the independence assumption to the observation likelihood as well: P(H i │H( i-1 ))=P(H (L) i │H (L) i-1 )P(H (R) i │H (R )i-1 )(20)

[0072]

[0082] The transition probabilities for each move are given by:

number

number

[0073]

[0083] The parameters used by embodiments herein can be optimized using Bayesian black-box optimization based on an expected improvement criterion. The evaluation function evaluates a set of parameter values ​​by computing an alignment for a set of score / performance pairs and comparing the computed alignment to an annotated ground truth alignment. The quantity to be minimized is the number of time steps the predicted hand states deviate from the ground truth.

[0074]

[0084] The result of the Viterbi algorithm is a sequence of hand states H that describes the time / pitch deltas for each hand and whether the hand is playing or not. * This result can be used as the basis for note-by-note alignment, which constructs a mapping between individual played notes and score notes. To construct the note-by-note alignment, embodiments apply a greedy search for the closest transformed score note for each played note. For example, embodiments may use a Viterbi path H, as described by the following pseudocode: * We can implement an algorithm for note-by-note alignment from P to S based on: 1: Procedure NOTEWISEALIGNMENT(S,P,H * ) 2:P'←P 3:S'←APPLYINFERREDDELTAS(S,H * ) 4:R←φ 5: While P'≠φ, do the following: 6:d←min s’∈S’ ||p-s'|| 7:s←argmin s’∈S’ ||p-s'|| 8:If d<η, do the following 9: R←R∪{(s,p)} 10:S'←S' / s 11: Otherwise, do the following: 12: R←R∪{(φ,p)} 13: End the if statement 14:P'←P' / p 15: End the while statement 16: Return R 17: End the procedure

[0075]

[0085] The above algorithm generates the estimated time / pitch delta H * to the score notes S. For example, the embedding algorithm can be described by the following pseudocode: 1: Procedure ApplyInferredDeltas(S,H * ) 2:S'←φ 3: For all hands ∈ L, R, do the following: 4: All alls∈S (hand ), do the following: 5:H i ∈H *(hand) , but t i = onset(s) 6:H i If ≠φ, do the following 7:s'←s+H i 8: S'←S'∪{s'} 9: End the if statement 10: Exit 11: Exit 12:Return S' 13: End the procedure

[0076]

[0086] H * While a greedy note-by-note alignment can in principle be performed without it, i.e. using the untransformed score notes S rather than S', if the time / pitch delta is too large the correct mapping cannot be established.

[0077]

[0087] 5A and 5B show an exemplary process flow 500 for aligning performance notes to score notes according to embodiments described herein. The process is represented as six stages. The first three stages (510, 520, 530) are shown in portion 500a of the flow shown in FIG. 5A, and the remaining three stages (540, 550, 560) are shown in portion 500b of the flow shown in FIG. 5B. Flow 500 can be implemented by evaluation engine 155 of FIG. 1A or 1B.

[0078]

[0088] An embodiment of flow 500 begins in a first stage 510 by receiving performance notes and score notes that fall within an evaluation window. For example, all performance notes and score notes can be received for post-feedback, or a subset of performance and score notes corresponding to a range of timestamps can be received for dynamic processing. In a second stage 520, time-step discretization and windowing are performed to effectively create processing chunks of score notes and potential candidate performance notes for alignment therewith. In a third stage 530, pitches and time deltas are calculated, and a KDE (e.g., Gaussian KDE) is performed based on the processing chunks generated in the second stage 520.

[0079]

[0089] Referring to the remaining portion 500b of the flow in Figure 5B, in a fourth stage 540, the state space of time and pitch delta from the third stage 530 is discretized. In a fifth stage 550, a Viterbi path (indicated by the bold lines connecting hypotheses at successive time steps) is computed through the discretized state space from the fourth stage 540. In a sixth stage, the Viterbi path computed in the fifth stage 550 is used to identify the set of hypotheses that best explain the observations, thereby generating the most likely note-by-note alignment between the score and the performed notes.

[0080]

[0090] For further clarity, Figures 6, 7A, and 7B show exemplary state-space representations of a given time window of user activity. Figure 6 shows a piano-roll format representation 600 of a given time window, including a sequence of score notes (dark thin horizontal lines) and played notes (lighter thick horizontal lines) for two-handed music. Dimmed lines are not part of the given time window. Figures 7A and 7B show kernel density plots 700 on the timing-pitch delta hypothesis space for left-hand and right-hand events, respectively, corresponding to a time step (t = 5.094 seconds) in the given time window of Figure 6. Circles indicate modes selected by the state-space discretization, and squares indicate modes that lie on the Viterbi path. In Figure 7A, it can be seen that one of the modes was selected to be part of the Viterbi path. In Figure 7B, in the left plot, the absence of a square indicates that the Viterbi path included a "not played" hypothesis at this time step.

[0081]

[0091] After applying the Viterbi algorithm, embodiments can move from score performance alignment to performance evaluation. In practice, the alignment of the performed notes to the score notes can be used as a most likely representation of the sequence of performance errors, such as between the played notes and the score notes, and the sequence of performance errors can be used to generate evaluation feedback. At least some of the evaluation feedback can be qualitative. For example, the feedback can indicate that an entire passage appears to have been played incorrectly, that an entire passage appears to have been played in the wrong octave or key, that certain types of score notes appear to be consistently incorrect (e.g., the user may not understand how to play dotted quarter notes, accidentals, triplets, etc.; or that the user needs to work on releasing notes during rests), that the user is late or early in one or more sections, etc. More subtle issues, such as one-off missed notes, pitch errors, or early / late note starts, can also follow directly from note-by-note alignment. Thus, feedback can be provided on different bases: based on time, based on pitch or timing or other (e.g., dynamics), based on the type of score note, at the individual event level, etc. Furthermore, multiple types of feedback can be provided simultaneously. For example, in a single passage, feedback can indicate that the user played the entire passage an octave higher, and also that they rushed a set of eighth notes in the middle of the passage, and also that they missed one of the notes (i.e., even taking into account the octave shift).

[0082]

[0092] Returning to FIG. 1 , the evaluator subsystem 150 includes an evaluation engine 155, an evaluation store 156, and a feedback engine 157. To support the novel evaluation techniques described herein, embodiments of the evaluation engine 155 are implemented in two stages: an alignment-based evaluation sub-engine and a note-based evaluation sub-engine. The alignment-based evaluation sub-engine performs alignment between the performed notes and the score notes, such as according to flow 500 in FIGS. 5A and 5B . The output of the alignment-based evaluation sub-engine can be a list of regions in the score, where each region corresponds to a time when the user is not playing or when the user is playing with some time and pitch offset (e.g., which may be 0 if the user is playing with correct pitch and timing). The note-by-note evaluation sub-engine generates a mapping (e.g., a one-to-one mapping) of performed notes to score notes. In some implementations, the note-by-note evaluation sub-engine queries each region from the alignment-based evaluation sub-engine, which returns its own region feedback. For example, if the user plays nothing during a particular region, the feedback may be a single feedback indicating that nothing was played there, or if the user plays something in a particular region, the feedback may be multiple feedback indicating a list of note-specific errors for the particular score notes.

[0083]

[0093] Some embodiments store the most likely sequence of composite states (e.g., as a sequence of performance errors) in the rating store 156. The rating store 156 can also store rating models usable by the rating engine 155, e.g., to generate an automated rating of a performance. For example, the rating model can provide a definition that enables the rating engine 155 to recognize when a sequence of performance errors suggests that the user was playing the wrong key. In some cases, the rating model may be defined algorithmically. For example, the rating model may indicate that a particular rating should be generated if a sequence of performance errors exceeding a threshold matches certain criteria. In other cases, the rating model may be defined as a pattern, mask, etc., and the rating engine 155 uses more complex statistical, pattern matching, machine learning, and / or other algorithmic techniques to determine whether to generate a particular rating. In some cases, one or more rating models are generated using artificial intelligence / machine learning (AI / ML) techniques. For example, numerous examples of a particular type of performance deficiency are assembled as training data, and an AI / ML engine is used to generate an appropriate rating model for subsequent identification of the deficiency. Although the description herein refers to identifying sequences of performance errors and performance deficiencies, the same techniques can be used to recognize performance improvements, successes, etc. For example, one or more evaluation models can be used by the evaluation engine 155 to identify instances where a user did an excellent job (e.g., performed a passage better than expected, or performed a passage better than the user performed in a previous attempt).

[0084]

[0094] The rating engine 155 generates one or more types of performance ratings (including, for example, a qualitative rating) of the user's performance and sends corresponding information to the feedback engine 157. The rating data is used by the feedback engine 157 to generate feedback data 159. The feedback data 159 may include micro-level feedback and / or macro-level feedback. The feedback data 159 can be used by the display processor 125 to generate graphical performance feedback that is output to the user via the display interface 120. In some embodiments, the feedback engine 157 generates feedback data 159 that can be further used by the note processor 135 to generate audible performance feedback that is output to the user via the audio interface 130.

[0085]

[0095] In some embodiments, the feedback data 159 is micro-level feedback that indicates event-by-event accuracy as a graphical overlay on the musical score (e.g., at the score note and / or played note level). The term "graphical overlay" is used generally herein to include any type of simultaneous graphical presentation that provides a visual juxtaposition between a graphical feedback element and the graphical score element to which the feedback applies. For example, such graphical overlays may include semi-transparently displaying the graphical feedback element over the graphical score element, recoloring the graphical score element to suggest the graphical feedback element (e.g., changing line thickness, etc.), adding text or an image representing the graphical feedback element along with the graphical score element, etc.

[0086]

[0096] In some implementations, score notes are displayed on the music staff in a first color to indicate skipped events, a second one or more colors to indicate well-played (or improved) events, and a third one or more colors to indicate poorly-played (or worsened) events. In another implementation, a first color is used to highlight the portion of the music staff during which the score note should be played (e.g., from the start time of a particular score note to an offset time), and a second color is used to highlight the portion of the music staff during which the corresponding played note was played (and / or a color to indicate whether the played note was played well or poorly). In other implementations, the generated feedback may include one or more numerical scores, colors, or other graphical elements used to overlay (or provide access to) past performance data, such as from the user or other performers.

[0087]

[0097] In other embodiments, the feedback data 159 is macro-level feedback that indicates feedback on multiple events at a time. Some macro-level feedback applies to an entire performance. Some macro-level feedback applies to a section (e.g., a page, a measure, a line, a hand, etc.). Some macro-level feedback applies to a passage (e.g., a musically related phrase, a grouped set of sixteenth notes, etc.). Some macro-level feedback applies to a category (e.g., a particular type of note or rest, all rests, triplets, dotted notes, staccato notes, etc.). As described herein, at least some of the macro-level feedback may be qualitative, such as based on identified patterns that indicate areas for improvement, areas showing current improvement, etc.

[0088]

[0098] Macro-level feedback may be presented in a separate portion of the display or in any suitable manner, as a graphical overlay on the musical score. In some such embodiments, the macro-level feedback indicates an overall performance score (e.g., in textual form and / or any other suitable format). For example, feedback data 159 indicates that the performance received an overall numerical score of 82%. In other such embodiments, evaluation engine 155 and / or feedback engine 157 perform one or more statistical analyses on the evaluation data to look for patterns or trends in the most likely sequences of composite states (e.g., in the sequence of performance errors), and feedback data 159 indicates the results of those analyses. One such analysis indicates patterns of performance across passages and / or sections of the performance. For example, feedback data 159 indicates that the user performed well (i.e., had high rhythmic correspondence) in measures 1-10, performed very poorly in measures 11-13, and performed moderately well in the remainder of the performance. Another such analysis matches performance patterns to predetermined performance categories. As one example, the feedback data 159 indicates that there was insufficient playing of rests rather than notes (e.g., the analysis compares the rhythmic correspondence of rests with score notes and non-score notes). As another example, the feedback data 159 indicates that there was an overall tendency to play ahead of or behind the beat (e.g., the analysis finds a statistical trend of performance event start times being earlier than or later than start times). As another example, the feedback data 159 indicates that there was a misunderstanding of a particular musical notation (e.g., the analysis finds that the user played two notes where the notation of two score notes is joined by a tie, indicating a misunderstanding of the tie, the user played triplets of all score notes as having the duration and interval of an eighth note, indicating a misunderstanding of the triplet, etc.).As another example, feedback data 159 may indicate that the performer is not attempting a portion of the music (e.g., the analysis may find that played notes are being missed for all of one hand in a two-handed piece, or for an entire section of a piece). As another example, feedback data 159 may indicate that the performer is struggling with a particular type of passage (e.g., the analysis may find that rhythmic correspondences tend to be higher during fast or slow sections of a piece). In any of the above examples, an embodiment of feedback engine 157 may generate feedback data 159 as text and / or graphical overlays in the same or separate portions of the interface to provide macro-level feedback to the user.

[0089]

[0099] In embodiments that provide dynamic feedback, the rating engine 155 can generate one or more types of performance ratings (e.g., including a qualitative rating) of the user's performance and can send corresponding information to the feedback engine 157 contemporaneously with the user's performance. The rating data is used by the feedback engine 157 to generate feedback data 159 (micro-level feedback and / or macro-level feedback) contemporaneously with the user's performance. As described above, the feedback data 159 can be used by the display processor 125 to generate graphical performance feedback (e.g., and / or audible feedback) that is output to the user via the display interface 120. Some or all of the feedback data 159 can be used to generate graphical and / or audible performance feedback contemporaneously with the user's performance.

[0090]

[0100] As one example, the dynamic performance evaluation determines that the user has not started playing the piece of music after a while. In response, dynamic qualitative feedback may be displayed (e.g., in a pop-up window, etc.) to ask whether the user wants to cancel or continue playing and / or reproducing the piece of music. In some implementations, such feedback may also include pausing playback. Other feedback (e.g., an audible bell, a change in display color, etc.) may also be used to make the user aware of the detected condition. As another example, while the user is playing, the dynamic performance evaluation determines that the user is playing in the wrong octave for several consecutive notes. In response, dynamic qualitative feedback may be displayed (e.g., in a pop-up window, on the displayed score, etc.) to instruct the user to change octaves. In some implementations, the dynamic feedback may change based on whether this is the first time the user has made this error, the user's playing level, the duration for which the same error has been detected, or any other suitable factor.

[0091]

[0101] 1 is shown as separate subsystems, the environment may be implemented in any suitable computing environment according to any suitable architecture. In some embodiments, the device I / O subsystem 105, the evaluator subsystem 150, and the MS data store 140 are all implemented in a single computing environment. In other embodiments, the device I / O subsystem 105, the evaluator subsystem 150, and the MS data store 140 are implemented in multiple computing environments communicatively coupled to each other by any suitable wired and / or wireless communication.

[0092]

[0102] In some embodiments, system 100 is implemented in a cloud-based environment. For example, one or more user devices communicate with one or more remote servers over one or more communication networks. Each user device may implement a respective instance of device I / O subsystem 105, thin client, and network interface. One or more remote servers implement evaluator subsystem 150 and a remote storage subsystem. The remote storage subsystem may include MS data store 140 and / or evaluation data store 156. For example, the remote storage subsystem may be used to maintain recordings of past performance data. In operation, a user uses a thin client (e.g., an app) on user device 210 to access an application that provides the functionality of rhythm evaluator subsystem 150 by communicating over a network. Operation of the thin client may involve communication of various types of data. For example, MS visual data 145 and MS logic data 143 are received by the user device from the remote storage subsystem of the remote server, performance data is sent from the user device back to the rhythm evaluator subsystem 150 of the remote server, and feedback data 159 is received by the user device from the rhythm evaluator subsystem 150 of the remote server. The network may include any suitable wired or wireless communication links with any one or more public and / or private networks, local and / or remote networks, etc.

[0093]

[0103] An embodiment of an automatic musical performance alignment and performance evaluation system (e.g., the automatic musical performance evaluation system 100 of FIGS. 1A and / or 1B) or components thereof can be implemented on and / or incorporated into one or more computer systems, as shown in FIG. 8. FIG. 8 illustrates a schematic diagram of one embodiment of a computer system 800 that can implement various system components and / or perform various steps of methods provided by various embodiments. Note that FIG. 8 is only meant to provide a generalized description of various components, any or all of which may be utilized as appropriate. Thus, FIG. 8 broadly illustrates how individual system elements can be implemented in a relatively separate or relatively more integrated manner.

[0094]

[0104] Computer system 800 is shown including hardware elements that can be electrically coupled (or otherwise communicate as needed) via bus 805. The hardware elements can include one or more processors 810, including, but not limited to, one or more general-purpose processors and / or one or more special-purpose processors (digital signal processing chips, graphics acceleration processors, video decoders, etc.). As shown, some embodiments include device I / O subsystem 105, which can include one or more I / O devices 815 and one or more I / O processors 817. I / O devices 815 can include display interface 120, audio interface 130, and / or any other suitable interface. I / O processor 817 can include display processor 125, note processor 135, and / or any other suitable processor. Additionally, input devices can include, but are not limited to, buttons, knobs, switches, keypads, touchscreens, remote controls, microphones, MIDI devices, etc., and output devices can include, but are not limited to, displays, speakers, indicators, gauges, etc. Some embodiments of computer system 800 interface with additional computers, peripheral devices, etc., such that device I / O subsystem 105 can include various physical and / or logical interfaces (e.g., ports, etc.) to facilitate interaction and control between components.

[0095]

[0105] The computer system 800 may further include (and / or be in communication with) one or more non-transitory storage devices 825, which may include, but are not limited to, local and / or network-accessible storage, and / or may include, but are not limited to, disk drives, drive arrays, optical storage devices, solid-state storage devices such as random access memory (“RAM”), and / or read-only memory (“ROM”), which may be programmable, flash-updateable, etc. Such storage devices may be configured to implement any suitable data store, including, but not limited to, various file systems, database structures, etc. In some embodiments, the storage device 825 includes non-transitory memory. In some embodiments, the storage device 825 may include the MS data store 140, the evaluation data store 156, and / or any other suitable data storage. The storage device 825 may also include buffers and / or other temporary storage used by the transcription engine 139, the note processor 135, the evaluation engine 155, and / or any other components.

[0096]

[0106] Computer system 800 may also include a communications subsystem 830, which may include, without limitation, any suitable antenna, transceiver, modem, network card (wireless or wired), infrared communication device, wireless communication device, chipset (e.g., Bluetooth™ device, 802.11 device, WiFi device, WiMax device, cellular communication device, etc.), and / or other communications components. As shown, communications subsystem 830 may also include network interface 220 for facilitating communications between a user device and a remote server over a communications network. Communications subsystem 830 may further facilitate communications with other computing systems.

[0097]

[0107] In many embodiments, computer system 800 further includes working memory 835, which may include RAM or ROM devices, as described herein. Computer system 800 may also include software elements shown as currently residing in working memory 835, which may include operating system 840, device drivers, executable libraries, and / or other code, such as one or more application programs 845, as described herein, and which may include computer programs provided by various embodiments and / or may be designed to implement methods and / or configure systems provided by other embodiments. By way of example only, one or more procedures described with respect to the methods described herein may be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in one aspect, such code and / or instructions can be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, operating system 840 and working memory 835 are used in conjunction with one or more processors 810 to implement score performance alignment and / or automatic performance evaluation functionality, as described herein.

[0098]

[0108] A set of these instructions and / or code may be stored on a non-transitory computer-readable storage medium, such as non-transitory storage device 825 described above. In some cases, the storage medium may be incorporated within a computer system, such as computer system 800. In other embodiments, the storage medium may be separate from the computer system (e.g., removable media such as a compact disc) and / or may be provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the stored instructions / code. These instructions may be in the form of executable code that is executable by computer system 800 and / or may be in the form of source and / or installable code that takes the form of executable code upon compilation and / or installation on computer system 800 (e.g., using any of a variety of commonly available compilers, installation programs, compression / decompression utilities, etc.).

[0099]

[0109] It will be apparent to those skilled in the art that substantial modifications can be made according to particular requirements. For example, customized hardware could be used and / or particular elements could be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connection to other computing devices, such as network input / output devices, could be used.

[0100]

[0110] As noted above, in one aspect, some embodiments may employ a computer system (such as computer system 800) to perform methods according to various embodiments of the present invention. According to one set of embodiments, some or all of the steps of such methods are performed by computer system 800 in response to processor 810 executing one or more sequences of one or more instructions contained in working memory 835 (which may be embedded in other code, such as operating system 840 and / or application program 845). Such instructions may be read into working memory 835 from another computer-readable medium, such as one or more of non-transitory storage devices 825. By way of example only, execution of the sequences of instructions contained in working memory 835 may cause processor 810 to perform one or more steps of the methods described herein.

[0101]

[0111] As used herein, the terms “machine-readable medium,” “computer-readable storage medium,” and “computer-readable medium” refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These media may be non-transitory. In embodiments implemented using computer system 800, various computer-readable media may participate in providing instructions / code to processor 810 for execution and / or may be used to store and / or carry such instructions / code. In many implementations, computer-readable media are physical and / or tangible storage media. Such media may take the form of non-volatile or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as non-transitory storage device 825. Volatile media include dynamic memory, such as, but not limited to, working memory 835. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium having a pattern of marks, a RAM, a PROM, an EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code. Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor 810 for execution. By way of example only, the instructions may initially be carried on a magnetic and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 800. The communications subsystem 830 (and / or its components) typically receives the signals, and the bus 805 may then carry the signals (and / or the data, instructions, etc. carried by the signals) to the working memory 835, from which the processor 810 retrieves and executes the instructions.The instructions received by the working memory 835 may optionally be stored on a non-transitory storage device 825 either before or after execution by the processor 810 .

[0102]

[0112] It should be further understood that components of computer system 800 may be distributed across a network. For example, some processing may be performed in one location using a first processor, while other processing may be performed by another processor remote from the first processor. Other components of computer system 800 may be distributed as well. Thus, computer system 800 may be interpreted as a distributed computing system that performs processing at multiple locations. In some cases, computer system 800 may be interpreted as a single computing device, such as separate laptops, desktop computers, etc., depending on the context.

[0103]

[0113] Figures 9A-9C show several plots representing an example performance that is generally correct, except for a silence between approximately 5 and 7 seconds. In the example performance, the user is playing with only the right hand, and the left-hand accompaniment is an automatic backing track. A pair of horizontal dashed lines is shown in all of Figures 9A-9C to represent the silence period. Figure 9A shows a piano roll representation 900 of the relevant portion of the performance. Similar to representation 600 in Figure 6, representation 900 shows score notes as thin lines and played notes as thicker shaded areas (overlaid on score notes where there is overlap). Figure 9B shows a plot 910 of the time and pitch delta of the right hand, respectively, as estimated in the Viterbi path. The black line represents the average time / pitch delta of the region, and the underlying gray line represents the change in delta over time. The breaks in the line represent the H i =φ, where the user is either not playing anything at all or is unable to meaningfully align the notes to the score. Figure 9C shows a plot 920 representing the estimated qualitative feedback, localized on the time line.

[0104]

[0114] 10A-10C show several plots representing an example performance with several issues, including the user playing an octave higher starting at 27 seconds, and the user playing several incorrect notes throughout the song. In the example performance, the user is playing only with the right hand, with the left-hand accompaniment being an automatic backing track. FIG. 10A shows a piano roll representation 1000 of the relevant portion of the performance. Similar to the representations in FIGS. 6 and 9A, representation 1000 shows score notes as thin lines and played notes as thicker shaded areas (overlaid on score notes where there is overlap). FIG. 10B shows a plot 1010 of the right hand's time and pitch delta, respectively, as estimated in the Viterbi path. The black line represents the average time / pitch delta of the region, and the underlying gray line represents the change in delta over time. The breaks in the line represent the H i =φ, where the user either played nothing at all or was unable to meaningfully align the notes to the score. Figure 10C shows a plot 1020 representing the estimated qualitative feedback, localized on the time line. In Figure 10C, it can be seen that the estimated qualitative feedback includes some incorrect notes that the user played, areas where the user did not play the score in a recognizable way, accidentals that the user overlooked, and intervals that the user missed.

[0105]

[0115] 11 shows a flow diagram of an exemplary method 1100 for score performance alignment for automatic performance evaluation, according to embodiments described herein. The embodiment begins, at step 1104, by receiving (e.g., by a processor-based evaluation engine) performed note data defining a sequence of performed notes representing a user's performance of a musical score according to a score time reference. Each of the performed note sequences is defined by at least a respective performed pitch and a respective performed start time. In some implementations, each performed note is further defined by a respective performed offset time and / or a performed duration.

[0106]

[0116] In some embodiments, receiving the performed note data in step 1104 includes receiving a raw audio stream via an audio interface and processing the raw audio stream by a transcription engine of a note processor to generate the performed note data. In other embodiments, receiving the performed note data in step 1104 includes receiving a MIDI stream via a Musical Instrument Digital Interface (MIDI) interface and processing the MIDI stream by a note processor to generate the performed note data. The note processor can be coupled to a processor-based evaluation engine.

[0107]

[0117] At stage 1108, an embodiment may receive (e.g., by a processor-based evaluation engine) score note data defining a sequence of score notes representing a musical score. Each of the sequence of score notes is defined by at least a respective score pitch and a score start time according to a score time base. In some implementations, each score note is further defined by a respective score offset time and / or score duration.

[0108]

[0118] In stage 1112, embodiments may calculate (e.g., by a processor-based evaluation engine) a note-by-note alignment between the performed note data and the score note data by calculating a most likely sequence of composite hand states for a sequence of time steps of a score time reference, given a sequence of score notes and a sequence of performed notes. Each time step in the sequence of time steps is defined as corresponding to the score start time of a respective one of the sequence of score notes (each time step corresponds to the start timing of a respective score note). In some embodiments, each composite hand state in the sequence of composite hand states is a combined left and right hand state calculated as an element of a Cartesian product of the left hand state space calculated for the time step and the right hand state space calculated for the time step.

[0109]

[0119] In some embodiments, calculating the note-by-note alignment in stage 1112 may include applying time-step discretization and windowing to the performance note data and the score note data to generate a sequence of regions, each associated with a respective time step in the sequence of time steps in the score time reference. Some such embodiments may include calculating region data for each region of the sequence of regions based on calculating a maximum likelihood sequence of composite hand states, the respective region data for each region indicating either a non-played observation or a played observation, the played observation further indicating associated pitch delta information and associated timing delta information for the region. In some embodiments, calculating the region data includes calculating the associated pitch delta information, associated timing delta information, and associated weighted kernel density estimates for the region, discretizing a hand state space for the region based on the associated pitch delta information, associated timing delta information, and associated kernel density estimates, and calculating a Viterbi path through the discretized hand state space. In some embodiments, calculating the note-by-note alignment in stage 1112 further includes generating, for each region of the sequence of regions, a one-to-one mapping between the performed notes and the score notes by querying the respective region data to obtain respective region feedback, each region feedback comprising either a single feedback indication of a non-played observation or one or more feedback indications representing the played observation as one or more note-specific errors for one or more particular score notes.

[0110]

[0120] Some such embodiments further include, for each region, calculating associated pitch delta information and associated timing delta information by: defining positive observations for the region as either non-played observations or playable observations; computing kernel density estimates (KDEs) for the positive observations, where the KDEs for the positive observations have modes corresponding to alignment between the score notes associated with the region and any played notes in the region; defining corresponding negative observations for the region; computing the KDEs for the negative observations, where the KDEs for the negative observations have modes corresponding to misalignment between the score notes associated with the region and any played notes in the region; and computing a non-normalized density for the region by subtracting the KDEs for the negative observations from the KDEs for the positive observations, where the non-normalized density has modes corresponding to most likely candidates for the associated pitch delta information and associated timing delta information for the region.

[0111]

[0121] In some embodiments, each hand composite-state of the maximum likelihood sequence of hand composite-states represents a hypothesis of whether the user is playing a predicted note at the time step and, if so, a hypothesis of a delta between the predicted note and one of the sequence of score-based note events. The delta may include a pitch delta and a time delta, where either or both of the pitch delta and the time delta may be represented by a continuous variable. In some such embodiments, calculating the note-by-note alignment in step 1112 includes, at each time step, applying a modified Hidden Markov Model (HMM) to generate a respective plurality of hypothesized hand composite-states for the time step and ranking each of the plurality of hypothesized hand composite-states to determine a respective most likely hand composite-state for the time step, where the most likely sequence of hand composite-states is the respective most likely hand composite-state for each of the sequence of time steps. Some such embodiments further include applying a Viterbi algorithm to calculate transition probabilities between each of the respective plurality of hypothesized composite-states at each time step, where the ranking is based on the transition probabilities.

[0112]

[0122] At step 1116, embodiments may automatically generate qualitative evaluation feedback by pattern matching the note-by-note alignment against a library of evaluation models. In some embodiments, the calculations at step 1112 and the automatic generation at step 1116 are performed at least in part for at least a portion of the sequence of time steps during the corresponding portion of the user's performance (i.e., simultaneously as the user plays the song). In such embodiments, the method may continue at step 1120 by displaying at least a portion of the qualitative evaluation feedback during the corresponding portion of the user's performance. In other embodiments, the automatic generation at step 1116 is performed at least in part upon completion of the user's performance. In such embodiments, the method may continue at step 1120 by displaying at least a portion of the qualitative evaluation feedback after completion of the user's performance.

[0113]

[0123] The methods, systems, and devices described above are examples. Various configurations may omit, substitute, or add various procedures or components, as appropriate. For example, in alternative configurations, methods may be performed in an order different from that described, and / or various steps may be added, omitted, and / or combined. Also, features described with respect to particular configurations may be combined in various other configurations. Different aspects and elements of the configurations may be combined in a similar manner. Also, technology evolves, and therefore, many of the elements are examples and do not limit the scope of the disclosure or the claims.

[0114]

[0124] Specific details are given in the description to provide a thorough understanding of example configurations (including implementations). However, configurations can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail to avoid obscuring the configurations. This description provides only example configurations and does not limit the scope, applicability, or configurations of the claims. Rather, the foregoing description of the configurations provides one skilled in the art with an enabling description for implementing the described technology. Various changes can be made in the function and arrangement of elements without departing from the spirit or scope of the present disclosure.

[0115]

[0125] Configurations may also be described as processes that are shown as flow diagrams or block diagrams. While each operation may be described as a sequential process, many of the operations may be performed in parallel or simultaneously. Additionally, the order of operations may be rearranged. A process may have additional steps not included in the diagrams. Furthermore, example methods may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. If implemented by software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored in a non-transitory computer-readable medium, such as a storage medium. A processor may perform the described tasks.

[0116]

[0126] While several example configurations have been described, various modifications, alternative configurations, and equivalents may be used without departing from the spirit of this disclosure. For example, the above elements may be components of larger systems, and other rules may take precedence over or otherwise modify the application of the present technology. Also, several steps may be taken before, during, or after the consideration of the above elements.

Claims

1. 1. A method for score performance alignment for automatic performance evaluation, comprising: receiving, by a processor-based evaluation engine, performed note data defining sequences of performed notes representing a user's performance of a musical score according to a score time standard, each of said sequences of performed notes defined by at least a respective played pitch and played start time; receiving, by said processor-based evaluation engine, score note data defining sequences of score notes representing said musical score, each of said sequences of score notes defined by at least a respective score pitch and a score start time that conforms to said score time standard; calculating, by said processor-based evaluation engine, a note-by-note alignment between said performance note data and said score note data by calculating a most likely sequence of composite hand states for a sequence of time steps in said score time base given said sequence of score notes and said sequence of performance notes; each time step of the sequence of time steps being defined as corresponding to the score start time of a respective one of the sequence of score notes; automatically generating qualitative evaluation feedback by pattern matching the note-by-note alignment with a library of evaluation models; A method comprising:

2. the steps of calculating the note-by-note alignment and automatically generating the qualitative evaluation feedback are performed at least in part for at least a portion of the sequence of time steps during a corresponding portion of the performance by the user; and displaying at least a portion of the qualitative evaluation feedback during the corresponding portion of the performance by the user. The method of claim 1.

3. the step of automatically generating the qualitative evaluation feedback is performed at least in part upon completion of the performance by the user; and and further comprising displaying at least a portion of the qualitative evaluation feedback after the user completes the performance. The method of claim 1.

4. 2. The method of claim 1 , wherein each composite hand state in the sequence of composite hand states is a combined left and right hand state computed as an element of a Cartesian product of a left hand state space computed for the time step and a right hand state space computed for the time step.

5. said step of calculating said note-by-note alignment further comprising: applying time step discretization and windowing to the performance note data and the score note data to generate a sequence of regions each associated with a respective time step of the sequence of time steps in the score time base; The method of claim 1 , comprising:

6. said step of calculating said note-by-note alignment further comprising: calculating region data for each region of the sequence of regions based on the step of calculating a maximum likelihood sequence of the composite hand states, the respective region data for each region indicating either a non-played observation or a played observation, the played observation further indicating associated pitch delta information and associated timing delta information for the region. The method of claim 5 further comprising:

7. For each region, the associated pitch delta information and the associated timing delta information defining a positive observation for said region as either said non-shot observation or said shot observation; calculating a kernel density estimate (KDE) of the positive observations, the KDE of the positive observations comprising modes corresponding to alignments between the score notes associated with the region and any played notes within the region; defining corresponding negative observations of said region; calculating a KDE for the negative observation, the KDE for the negative observation including a mode corresponding to a misalignment between the score notes associated with the region and any played notes within the region; and calculating a non-normalized density for the region by subtracting the KDE of the negative observations from the KDE of the positive observations, the non-normalized density having a mode corresponding to a most likely candidate for the associated pitch delta information and the associated timing delta information for the region; The method of claim 6 further comprising the step of calculating by:

8. For each region, the step of calculating the region data comprises: calculating the associated pitch delta information, the associated timing delta information, and an associated weighted kernel density estimate of the region; discretizing the hand state space for the region based on the associated pitch delta information, the associated timing delta information, and the associated kernel density estimate; computing a Viterbi path through the discretized hand state space; The method of claim 6, comprising:

9. said step of calculating said note-by-note alignment further comprising:

7. The method of claim 6, further comprising: generating, for each region of the sequence of regions, a one-to-one mapping between the performed notes and the score notes by querying the respective region data to obtain respective region feedback, the respective region feedback comprising either a single feedback indication of the non-played observation or one or more feedback indications expressing the played observation as one or more note-specific errors for one or more particular score notes.

10. each hand composite state of the most likely sequence of hand composite states represents a hypothesis of whether the user is playing a predicted note at that time step, and if so, a hypothesis of the delta between the predicted note and said one of the score-based sequences of note events; the deltas include a pitch delta and a time delta, and at least one of the pitch delta or the time delta is represented by a continuous variable; The method of claim 1.

11. The step of calculating comprises, at each time step: applying a modified Hidden Markov Model (HMM) to generate a respective plurality of hypothesized hand composite states for the time step; ranking each of the plurality of hypothesized composite hand states to determine a most likely composite hand state for each of the plurality of hypothesized composite hand states for the time step; wherein the most likely sequence of composite hand states is the most likely composite hand state for each of the sequence of time steps. The method of claim 10.

12. applying a Viterbi algorithm to calculate transition probabilities between each of the plurality of hypothesized composite states at each time step; further comprising the ranking is determined by combining the observation likelihood of each hypothesis with the transition probabilities between successive hypotheses; The method of claim 11.

13. The step of receiving musical note data includes: receiving a raw audio stream via an audio interface; processing the raw audio stream by a transcription engine of a note processor to generate the performed note data, the note processor being coupled to the processor-based evaluation engine; The method of claim 1 , comprising:

14. The step of receiving musical note data includes: receiving a MIDI stream via a Musical Instrument Digital Interface (MIDI) interface; processing the MIDI stream by a note processor to generate the performance note data, the note processor being coupled to the processor-based evaluation engine; The method of claim 1 , comprising:

15. The method of claim 1 , wherein each of the sequences of played notes is further defined by a respective played offset time and / or played duration.

16. The method of claim 1 , wherein each of the sequences of score notes is further defined by a respective score offset time and / or score duration.

17. a musical score data store storing score note data defining sequences of score notes representing a musical score, each of said sequences of score notes defined by at least a respective score pitch and a score start time conforming to a score time standard; an audio interface for receiving an audio stream during a user's performance of the score according to the score time reference; a note processor coupled to said audio interface for generating from said audio stream performed note data defining a sequence of performed notes representing said performance of said musical score by said user, each of said sequence of performed notes defined by at least a respective played pitch and played start time; a processor-based evaluation engine coupled to the note processor and the music score data store, receiving the played note data from the note processor; receiving the score note data from the music score data store; calculating, by said processor-based evaluation engine, a note-by-note alignment between said performance note data and said score note data by calculating a most likely sequence of composite hand states for a sequence of time steps in said score time base given said sequence of score notes and said sequence of performance notes; each time step of the sequence of time steps is defined as corresponding to the score start time of a respective one of the sequence of score notes; and automatically generating qualitative evaluation feedback by pattern-matching the note-by-note alignment with a library of evaluation models; A processor-based evaluation engine consisting of An automatic music performance evaluation system comprising:

18. displaying the score data on a display device as a visual representation of the score during the performance by the user; and Displaying the qualitative evaluation feedback on the display device. Display Processor 20. The system of claim 17, further comprising:

19. the processor-based evaluation engine is configured to calculate the note-by-note alignment for at least a portion of the sequence of time steps during a corresponding portion of the performance by the user and to automatically generate the qualitative evaluation feedback; and the display processor is configured to cause display of at least a portion of the qualitative evaluation feedback during the corresponding portion of the performance by the user.

20. The system of claim 18.

20. one or more processors coupled to the audio interface; a non-transitory processor-readable memory storing instructions that, when executed, cause the one or more processors to implement the musical note processor and the processor-based evaluation engine; 20. The system of claim 17, further comprising: