Music teaching system based on voice recognition

By using a speech recognition-based music teaching system, combined with technical feature quantification and emotion theory mapping models, personalized empathic dialogues and body movement instructions are generated, solving the problem that existing systems cannot provide emotional and artistic guidance, and realizing in-depth teaching and efficient learning.

CN121998802APending Publication Date: 2026-05-08SICHUAN YUNSHUFUZHI EDUCATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN YUNSHUFUZHI EDUCATION TECH CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing music teaching systems fail to provide guidance with emotional warmth and artistic inspiration, making it difficult to assess the emotional elements in music, such as emotional expression, sonic narrative, and cultural connotations. This results in one-sided teaching objectives and a lack of in-depth integration and intuitive experience.

Method used

A speech recognition-based music teaching system is adopted. By acquiring learners' music performance data through a technical feature quantification module, an emotion theory mapping model is constructed to generate personalized empathic dialogues and Dalcroze body rhythm teaching instructions, thereby realizing multi-dimensional assessment and in-depth teaching of music performance.

Benefits of technology

It enables a deeper understanding of musical performance and concrete teaching guidance. By integrating multi-dimensional objective indicators with the connotations of art philosophy, it generates personalized empathetic dialogues and physical movement instructions, thereby improving learners' perception and internalization efficiency of musical concepts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998802A_ABST
    Figure CN121998802A_ABST
Patent Text Reader

Abstract

The invention discloses a music teaching system based on voice recognition, and relates to the technical field of music feature fusion, a technical feature quantification module is used for calculating a rhythm time value deviation ratio, pitch accuracy, a pitch tension scalar, a mode type scalar, rhythm boundary detection intensity, a linearity scalar, a discontinuity scalar and a semantic definition scalar; the emotional reasoning module is used for constructing an emotional theory mapping model and generating an emotional expression scalar, an art connotation scalar and a sound narrative continuity scalar based on the emotional theory mapping model; and the teaching application module is used for generating a personalized emotion-sharing dialogue instruction and a Dalkerz posture rhythm teaching instruction aiming at the current technical characteristic deviation. The system is a complete system on deep perception, philosophy cognition and concrete teaching guidance of music performance, deep fusion of technical features and application features is realized, originally isolated objective scoring and subjective art guidance are unified, and the core problems of technology and art disjunction and rigid feedback are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of music feature fusion, and more particularly to a music teaching system based on speech recognition. Background Technology

[0002] Music education is a crucial pathway to cultivating aesthetic literacy and creativity. Traditional music teaching relies heavily on teachers' personal experience and subjective judgment, making it difficult to achieve large-scale, personalized, and standardized instruction. With the development of artificial intelligence technology, applications have emerged in music teaching that utilize speech recognition for pitch and rhythm scoring.

[0003] Currently, Chinese invention patent application number CN202510554921.0 discloses an artificial intelligence-based music teaching system, including a neurocognitive adaptation module, a music DNA map construction module, a cognitive load monitoring module, an anti-AI dependency adjustment module, and a multi-source data fusion unit. This invention's artificial intelligence-based music teaching system constructs a dynamically evolving personalized teaching model through AI-driven multimodal data fusion and real-time analysis. The neurocognitive adaptation module combines EEG feature analysis and physiological signal monitoring to accurately capture learners' cognitive preferences and ability bottlenecks; the music DNA map, based on quantum augmentation modeling technology, continuously updates and predicts skill development trajectories; and the AI ​​teaching strategy can be dynamically adjusted in milliseconds based on learners' neural feedback and behavioral data, significantly improving skill mastery efficiency and knowledge retention rate, breaking through the limitations of static and singular traditional teaching systems. Existing systems are mostly mechanical scoring systems with rigid feedback, unable to provide the emotional warmth and artistic inspiration that human teachers can offer. They struggle to build a deep understanding of music. Current systems can only assess quantifiable technical indicators and cannot understand and evaluate the emotional elements in music, such as emotional expression, vocal narrative, and cultural connotations. This leads to one-sided teaching objectives. Existing systems mostly remain at the level of listening, singing, imitation, and scoring feedback, lacking deep integration with mature music teaching methods and failing to transform abstract musical concepts into intuitive physical experiences. Summary of the Invention

[0004] The technical problem solved by this invention is that existing systems are mostly mechanical scoring systems with rigid feedback, unable to provide guidance with emotional warmth and artistic inspiration like human teachers, making it difficult to establish a deep understanding of music. Existing systems can only evaluate quantifiable technical indicators and cannot understand and evaluate the emotional elements in music, such as emotional expression, vocal narrative, and cultural connotations, leading to one-sided teaching objectives. Existing systems mostly stay at the level of listening, singing, imitation, and scoring feedback, lacking deep integration with mature music teaching methods and failing to transform abstract musical concepts into intuitive physical experiences.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a music teaching system based on speech recognition, comprising a technical feature quantification module, an empathy reasoning module, and a teaching application module: The technical feature quantization module is used to acquire the learner's music performance data through speech recognition and the Musical Instrument Digital Interface (MIDI). After aligning the music performance data with a time series based on dynamic time warping, it calculates the rhythmic value deviation rate, pitch accuracy, interval tension scalar, mode type scalar, prosodic boundary detection intensity, linearity scalar, discontinuity scalar, and semantic clarity scalar. The empathy reasoning module is used to construct an emotion theory mapping model, and based on the emotion theory mapping model, it generates scalars of emotion expression, artistic connotation, and audio narrative coherence. The teaching application module is used to take the emotional expression scalar, the voice narrative coherence scalar, and the artistic connotation scalar as conditions, input conditions to drive the dialogue generation model, and generate personalized empathic dialogue instructions and Dalcroze body rhythm teaching instructions that are tailored to the current technical feature deviations.

[0006] Preferably, the technical feature quantification module includes: The music performance data includes audio data of the learner's singing and playing, digital interface data of the instrument, and lyrics text corresponding to the audio data; The audio data of learners singing and playing includes the raw sound wave signal collected by the microphone when learners sing and play non-musical digital interface devices. Non-musical digital interface devices include human voices and acoustic musical instruments. The raw sound wave signal includes the amplitude sampling point sequence, fundamental frequency trajectory, loudness energy trajectory and Mel frequency cepstral coefficients of the sound signal stored in FLAC format on the time axis. The digital interface data for musical instruments includes a sequence of digital instructions generated when a learner plays an electronic musical instrument. The sequence of digital instructions includes note on events, note off events, pitch data, velocity values, and control change events. The note activation event includes the timestamp of the key being pressed and the velocity value of the note; The note-off event includes the timestamp of when the key was released; Pitch data includes discrete pitch values ​​represented by natural numbers from 1 to 127 recorded in the musical score; The force value is a numerical value attached to the note activation event, recording the speed and pressure at which the key is pressed. Control change events include changes in the usage status of the pedal, vibrato wheel, and pitch bend wheel controllers during performance.

[0007] Preferably, after time-series alignment of the audio data of the learner's singing and playing and the digital interface data of the instrument, music technical features are extracted. The music technical features include rhythmic value deviation rate, pitch accuracy rate, interval tension scalar and mode type scalar. Time series alignment includes: The standard sheet music uploaded by learners is converted into a standard note event set. The mathematical expression for the standard note event set is: ; ; Where S is the set of musical note events. For the first A standard musical note event, This is the standard timestamp for the note, which is the midpoint between the note's close event and its open event. Standard pitch data for musical notes. The total number of notes in a standard musical score; The audio data of the learner’s actual singing and playing is converted into actual note event sets. The actual note event sets include all actual note events of the learner’s actual singing and playing. The structure of the actual note event sets is the same as that of the standard note event sets. Each actual note event set includes the actual timestamp and actual pitch data of the notes. Construct a cost matrix, initializing it to have dimensions m×n, where m is the number of actual note events included in the actual note event set. The cost matrix represents the cost of the notes in the performance sequence. notes in the standard sequence The local matching cost incurred during the matching process; The mathematical expression for calculating the cost of local matching is: ; in, For local matching cost, For the actual pitch data of the i-th note in the actual note event set, Let i be the actual timestamp of the i-th note in the actual note event set. and These are the preset weighting coefficients.

[0008] Preferably, the m×n cumulative cost matrix D is initialized, and the initialization includes setting the starting point of the cumulative cost matrix D to... The first line is set to Following the rule that it can only move horizontally from the predecessor node on the left, the first column is set to... It follows the rule that it can only move from the predecessor node above; Using dynamic programming, we iteratively calculate each element in the cumulative cost matrix D to find the path with the minimum cumulative cost from the starting point to the ending point. The iterative calculation includes: Elements in the cumulative cost matrix D The value of is equal to the current local matching cost plus the minimum cumulative cost selected from the three neighboring predecessor nodes, and its mathematical expression is: ; After the iterative calculation is completed, the lower right element of the cumulative cost matrix D is... That is, the minimum total matching cost. Backtracking backwards along the path that leads to the minimum cumulative cost back to the starting point The backtracking path is the optimal alignment path. The path that results in the minimum cumulative cost is the path consisting of the predecessor nodes selected when the minimum cumulative cost is chosen. The optimal alignment path includes index pairs. This refers to the actual note event and the standard note event that are matched in time.

[0009] Preferably, the mathematical expression for the rhythmic time value deviation rate is: ; Where N is the total number of musical notes. This represents the duration of the i-th note in the actual note event set, where the duration is the difference between the timestamps of the note closing event and the note opening event. Let j be the duration of the standard note in the set of events that corresponds to the duration of the i-th note in the actual note event set. Let i be the weight of the i-th note. The calculation process includes: Extract key local features of the notes, including the dynamic value of the note, the relative value of the note duration, the signal-to-noise ratio of the note, boundary notes, and accented notes; The relative duration of a note is the proportion of the note's duration to the total duration of the musical segment; The signal-to-noise ratio of a note is the local signal-to-noise ratio of note i during the note's duration; Boundary notes are notes on potential prosodic boundaries; The accented note is the first note of each measure that matches the rhythm of the score. The set of all key local features corresponding to a random note is formed into a one-dimensional local note feature vector. The weight of the current note is calculated based on the local note feature vector, and its mathematical expression is as follows: ; ; in, For musical notes The local note feature vector, These are preset feature sensitivity parameters; The mathematical expression for pitch accuracy is: ; ; in, For the first The cent deviation of each note The preset maximum allowable pitch deviation threshold, For pitch accuracy; The mathematical expression for the interval tension scalar is: ; in, For interval tension scalar, The total number of consecutive note pairs or chords. This refers to the interval between the k-th note pairs or chords, expressed in semitones. It is a discrete interval tension function; The calculation process of mode type scalar includes: A pitch frequency histogram is collected. This histogram represents the frequency distribution of each of the 12 semitones in the corresponding musical segment from the audio data and instrument digital interface data. The pitch frequency histogram is normalized to obtain the actual pitch distribution. The cosine similarity between the actual pitch distribution and the target mode template is calculated to obtain the mode matching score. All cosine similarities are aggregated into a one-dimensional vector to obtain the mode type scalar. The expression for the mode type scalar is: ; in, For mode type scalar, For the target debugging template, For the actual pitch distribution and target mode template Cosine similarity between them; The highest-scoring element of the mode type scalar represents the main mode type of the current musical segment.

[0010] Preferably, prosodic structure analysis is performed on the audio data to extract prosodic boundary detection intensity and prosodic processing scalar. The prosodic structure analysis includes: Traverse the loudness energy trajectory of the audio data, identify regions where the loudness energy is lower than the preset loudness energy threshold and the duration exceeds the minimum inter-sentence pause threshold, and mark them as potential prosodic boundaries; The actual duration of the note preceding the potential prosodic boundary is detected to be greater than the standard duration. If the actual duration is greater than the standard duration, the note is marked as a phrase-end extended note, and the difference between the actual duration and the standard duration is calculated and recorded as the duration extension. The proportion of the duration extension to the total duration of the music segment corresponding to the current audio data is calculated and recorded as the duration extension ratio. The standard duration is the average of the actual durations of all notes in the current audio data. Detect whether the fundamental frequency of the note at the potential rhythm boundary drops. If the fundamental frequency drops, mark the note as the phrase end marker note. Calculate the average difference between the fundamental frequency of the current note and the fundamental frequency of the previous note, and record it as the fundamental frequency drop magnitude. The intensity feature value of the note on the potential prosodic boundary is obtained by multiplying the duration, duration extension ratio and fundamental frequency drop of the note on the potential prosodic boundary. The average value of the intensity feature value of all notes on the potential prosodic boundary is calculated to obtain the prosodic boundary detection intensity scalar. The prosodic processing scalar is obtained by multiplying the prosodic boundary detection intensity scalar with the rhythmic time value deviation rate.

[0011] Preferably, music structure analysis is performed on the audio data and instrument digital interface data to extract linearity scalars and discontinuous scalars. The music structure analysis process includes: A pattern matching algorithm is executed on the standard note event set to identify and obtain the motif set. The process of identifying and obtaining the motif set includes: Encode a relative feature vector for all corresponding notes in the standard note event set. The relative feature vector includes a relative pitch feature and a relative rhythm feature. The relative pitch feature is the number of semitones between the current note and the previous note. The relative rhythm feature is the ratio of the duration of the current note to the duration of the previous note. The minimum and maximum motive lengths are preset, with motive length expressed in units of notes. Traverse the standard note event set to obtain all complete melody fragments with lengths between the minimum and maximum motif lengths. The complete melody fragments are in units of measures. Calculate the note similarity between any two complete melody fragments using the Longest Common Subsequence (LCS) algorithm. Obtain the note similarity calculated between each complete melody fragment in the standard note event set and any other complete melody fragment. Count the number of note similarities corresponding to each complete melody fragment that exceed a preset note similarity threshold. When the number is greater than a preset minimum repetition threshold, output the current complete melody fragment as a valid motif fragment. Iteratively check whether all valid motivational fragments include shorter valid motivational fragments. If shorter valid motivational fragments are included, output the shorter valid motivational fragments as basic motivational fragments. If shorter valid motivational fragments are not included, directly output the valid motivational fragments as basic motivational fragments. Combine all basic motivational fragments into a motivational set. Record the pitch profile, rhythmic pattern, and digital instruction sequence of each basic motif fragment, and structure each element in the motif set into a feature vector including pitch profile, rhythmic pattern, and internal pitch tension to obtain a motif feature set; Traverse the standard note event set to obtain all occurrence positions of each basic motif fragment in the musical phrase, and identify the developmental variation type at each occurrence position. The developmental variation type includes repetition, reflection, expansion, and fragment development. Use big data to obtain the pitch change features and rhythm change features of each developmental variation type, and encode the developmental variation type and its pitch change features and rhythm change features in a non-repetitive manner to form a variation coding vector. Extract the actual motivic melody fragments at the corresponding occurrence positions of the actual note event set, obtain the variation encoding vector of the actual motivic melody fragments, and check whether the variation encoding vector of the actual motivic melody fragments matches the variation encoding vector of the basic motivic fragment at the corresponding occurrence position. The matching logic is that the elements at the corresponding positions of the variation encoding vectors are the same non-repeating encoding mathematical operation logic. Based on the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions, the linearity is calculated. The mathematical expression for linearity is: ; in, For linearity, The total number of variations being tracked. The Euclidean distance between the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions is given. Let k be the variation encoding vector of the basic motif segment. The variation encoding vector for the actual motif melody segment of the kth segment; Identify interval tension scalars that exceed the dissonance threshold, monitor the switching frequency of mode type scalars within a unit time window, identify sudden increases in rhythmic time value deviation rate at non-prosodic boundaries, sudden increases indicate discontinuous sound wave numerical growth, and calculate discontinuity. ; in, For discontinuity, The average value of the interval tension scalar that exceeds the dissonance threshold. The number of times the scalar switch to the mode type is performed. This represents the average amplitude of sudden changes in the rhythmic time value deviation rate.

[0012] Preferably, the standard tone of each Chinese character in the lyrics text is queried. The standard tone includes high level tone, rising tone, falling-rising tone and falling tone. The pitch direction is obtained according to the standard tone. The pitch direction corresponding to the standard tone is: high level tone and rising tone correspond to flat, falling-rising tone corresponds to rising, and falling tone corresponds to falling. Extract the actual pitch direction of the fundamental frequency trajectory of each Chinese character in the corresponding original sound wave signal segment and the standard pitch direction of the standard musical score of the corresponding segment. Calculate the tone melody violation rate and semantic clarity scalar based on the actual pitch direction and the standard pitch direction. The process of calculating the tone melody violation rate includes: comparing whether the actual pitch direction of the Chinese character is the same as the standard pitch direction of the corresponding position. If they are not the same, the Chinese character is marked as a tone melody violation. The ratio of the number of Chinese characters marked as tone melody violations to the total number of Chinese characters corresponding to the original sound wave signal is calculated to obtain the tone melody violation rate. Subtracting the tone-melody violation rate from 1 yields the semantic clarity scalar.

[0013] Preferably, the empathy reasoning module includes an emotion theory mapping model construction unit and an empathy vector extraction unit: The emotion theory mapping model building unit is used to construct an emotion theory mapping model, which includes: The rhythmic value deviation rate and mode type scalar are used as input features; The elements in the mode type scalar are multiplied by a weight matrix and then summed in a weighted manner. The weight matrix includes the sentiment tendencies of all mode types in the mode type scalar. The weighted sum is then normalized to obtain the sentiment polarity features. When the emotional polarity metric is greater than 0.5 and less than 1, it indicates a positive emotion. When the emotional polarity metric is less than 0.5 and greater than 0, it indicates negative emotion; The rhythmic timing deviation rate, the difference between 1 and the rhythmic timing deviation rate, and the emotional polarity feature are concatenated into a fusion feature vector. This vector is then input into the first fully connected layer and output through the nonlinear activation function ReLU. The output is then input into the second fully connected layer and output through the nonlinear activation function ReLU. Finally, the output is input into a linear layer and output through the Softmax activation function to obtain an emotional expression scalar. Each element in the emotional expression scalar represents the probability of matching the performer's musical performance with each emotion. The empathy vector extraction unit includes: By mapping the discontinuous scalar and the interval tension scalar to artistic connotations according to Nietzsche's philosophical interpretation rules, we obtain the artistic connotation scalar. These Nietzschean philosophical interpretation rules include: The discontinuous scalar, linearity scalar, and interval tension scalar are input into the emotion theory mapping model at the input positions of rhythmic time deviation rate, the difference between 1 and rhythmic time deviation rate, and emotional polarity features, and the output Dionysian spirit component features are then generated. The difference between 1 and the rhythmic time deviation rate, the rhythmic time deviation rate, and the interval tension scalar are input into the emotion theory mapping model at the input positions of the rhythmic time deviation rate, the difference between 1 and the rhythmic time deviation rate, and the emotional polarity feature, and the power will component feature is output. By combining the characteristics of the Dionysian spirit and the characteristics of the will to power, we obtain the scalar of artistic connotation; Input the semantic clarity scalar and the prosodic boundary detection intensity scalar into the emotion theory mapping model, and output the voice narrative coherence scalar.

[0014] Preferably, the teaching application module includes: Using the emotional expression scalar, the voice narrative coherence scalar, and the artistic connotation scalar as conditions, the input condition drives the dialogue generation model LLM to generate personalized empathetic dialogues that address the current technical feature deviations. The types of technical feature deviations include: when any one of the feature values ​​of rhythmic value deviation rate, pitch accuracy rate, interval tension scalar, and mode type scalar exceeds the preset corresponding threshold range, a technical feature deviation occurs. Based on the types of technical feature deviations included in the personalized empathic dialogue, index matching is performed in the preset Dalcroze body rhythm teaching instruction library. Based on the matched instruction template and specific deviation values, body movement instructions are generated. The body movement instructions include stride change instructions for correcting rhythm, body extension instructions for perceiving intervals, and imitation action instructions for reflecting mode types.

[0015] The beneficial effects of this invention are as follows: This invention achieves precise alignment of multimodal music performance data through dynamic time warping and innovatively generates multi-dimensional advanced technical indicators, including interval tension scalars, discontinuity scalars, and semantic clarity scalars, laying the foundation for accurate analysis. It transforms objective technical indicators into artistic cognition with profound philosophical connotations, identifies the emotional polarity of performance through an emotion theory mapping model, and quantifies high interval tension scalars and high discontinuity scalars into Dionysian spiritual components through an empathy vector extraction unit. Simultaneously, it transforms the high control power under complexity into a power will component, thereby achieving a deep assessment of the driving force of performing arts. This philosophically empowered analysis enables the system to understand the artistic motivation behind technical errors, generating personalized empathic dialogues with causal chains, solving the rigidity and one-sidedness of traditional feedback. In terms of teaching applications, this system deeply integrates abstract musical concepts with the Dalcroze body movement teaching method, directly matching technical deviations such as high rhythmic value deviation rates to concrete stride change instructions and body extension instructions, realizing a embodied cognitive teaching model, greatly improving learners' perception and internalization efficiency of musical concepts. Attached Figure Description

[0016] Figure 1 A basic flowchart of a speech recognition-based music teaching system is provided as an embodiment of the present invention. Figure 2 This is a schematic representation of a contrasting interval provided in one embodiment of the present invention; Figure 3 This is a schematic diagram of a discrete interval tension function provided in one embodiment of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Reference Figure 1 As an embodiment of the present invention, a music teaching system based on speech recognition is provided, including a technical feature quantification module, an empathy reasoning module, and a teaching application module: The technical feature quantization module is used to acquire learners' music performance data through speech recognition and the MIDI (Musical Instrument Digital Interface). After aligning the music performance data with a time series based on dynamic time warping, it calculates rhythmic value deviation rate, pitch accuracy, interval tension scalar, mode type scalar, prosodic boundary detection intensity, linearity scalar, discontinuity scalar, and semantic clarity scalar. The empathy reasoning module is used to construct an emotion theory mapping model, and based on the emotion theory mapping model, it generates scalars of emotion expression, artistic connotation, and audio narrative coherence. The teaching application module uses scalars of emotional expression, vocal narrative coherence, and artistic connotation as input conditions to drive the dialogue generation model, generating personalized empathic dialogue instructions and Dalcroze body rhythm teaching instructions that address current technical feature biases.

[0019] This invention provides a complete system for in-depth perception, philosophical understanding, and concrete teaching guidance in music performance. It achieves a deep integration of technical and application features, unifying previously isolated objective scoring and subjective artistic guidance. Through a technical feature quantification module, multi-dimensional objective data is acquired; an empathic reasoning module imbues it with artistic and philosophical connotations; and finally, a teaching application module generates personalized empathic dialogues and Dahlcroze body movement instructions. This completely solves the core problems of the disconnect between technology and art and rigid feedback in traditional music teaching, achieving a major breakthrough in intelligent and personalized teaching.

[0020] The technical feature quantification module includes: Music performance data includes audio data of learners singing and playing instruments, digital interface data of musical instruments, and lyrics text corresponding to the audio data; The audio data of learners singing and playing includes the raw sound wave signal collected by the microphone when learners sing and play non-musical digital interface devices. Non-musical digital interface devices include human voices and acoustic musical instruments. The raw sound wave signal includes the amplitude sampling point sequence, fundamental frequency trajectory, loudness energy trajectory and Mel frequency cepstral coefficients of the sound signal stored in FLAC format on the time axis. The digital interface data for musical instruments includes a sequence of digital instructions generated when a learner plays an electronic musical instrument. The sequence of digital instructions includes note on events, note off events, pitch data, velocity values, and control change events. The note activation event includes the timestamp of the key being pressed and the velocity value of the note; The note-off event includes the timestamp of when the key was released; Pitch data includes discrete pitch values ​​represented by natural numbers from 1 to 127 recorded in the score, representing different semitones, and is used to calculate pitch accuracy and interval tension scalars; The dynamic value is a numerical value attached to the note activation event, recording the speed and pressure of the key being pressed. It directly quantifies the performer's intention in loudness changes and is a key input for generating emotional expression scalars and analyzing dynamic performance details. Control events include changes in the use of pedals, vibrato wheel, and pitch bend wheel controllers during performance (in use and not in use), which are used to analyze the continuity and emotional expression of the performance, and are especially important when analyzing pedal use in piano playing.

[0021] The system achieves comprehensive acquisition and standardization of multimodal data. By simultaneously acquiring audio data, instrument digital interface data, and lyrics, it ensures that the system can analyze any type of musical performance. In particular, the acquisition of dynamic values ​​and control change events directly quantifies the performer's intentions in loudness changes and emotional embellishments, providing precise dynamic performance details for subsequent analysis of emotional expression and artistic connotation.

[0022] After time-series alignment of the audio data of learners’ singing and playing and the digital interface data of instruments, music technical features are extracted. The music technical features include rhythmic value deviation rate, pitch accuracy rate, interval tension scalar and mode type scalar. Time series alignment includes: The standard sheet music uploaded by learners is converted into a standard note event set. The mathematical expression for the standard note event set is: ; ; Where S is the set of musical note events. For the first A standard musical note event, This is the standard timestamp for the note, which is the midpoint between the note's close event and its open event. Standard pitch data for musical notes. The total number of notes in a standard musical score; The audio data of the learner’s actual singing and playing is converted into actual note event sets. The actual note event sets include all actual note events of the learner’s actual singing and playing. The structure of the actual note event sets is the same as that of the standard note event sets. Each actual note event set includes the actual timestamp and actual pitch data of the notes. Construct a cost matrix, initializing it to have dimensions m×n, where m is the number of actual note events included in the actual note event set. The cost matrix represents the cost of the notes in the performance sequence. notes in the standard sequence The local matching cost incurred during the matching process; The mathematical expression for calculating the cost of local matching is: ; in, For local matching cost, For the actual pitch data of the i-th note in the actual note event set, Let i be the actual timestamp of the i-th note in the actual note event set. and These are the preset weighting coefficients.

[0023] The preset weighting coefficients are manually set.

[0024] In this embodiment, the number of note events is equal to the number of notes, and note events include specific notes; Based on these matching pairs It allows for the safe calculation of rhythmic timing deviation rates, i.e., comparison. and .

[0025] Initialize the m×n cumulative cost matrix D. Initialization includes setting the starting point of the cumulative cost matrix D to... The first line is set to Following the rule that it can only move horizontally from the predecessor node on the left, the first column is set to... It follows the rule that it can only move from the predecessor node above; Using dynamic programming, we iteratively calculate each element in the cumulative cost matrix D to find the path with the minimum cumulative cost from the starting point to the ending point. The iterative calculation includes: Elements in the cumulative cost matrix D The value of is equal to the current local matching cost plus the minimum cumulative cost selected from the three neighboring predecessor nodes, and its mathematical expression is: ; After the iterative calculation is completed, the lower right element of the cumulative cost matrix D is... That is, the minimum total matching cost. Backtracking backwards along the path that leads to the minimum cumulative cost back to the starting point The backtracking path is the optimal alignment path. The path that results in the minimum cumulative cost is the path consisting of the predecessor nodes selected when the minimum cumulative cost is chosen. The optimal alignment path includes index pairs. This refers to the actual note event and the standard note event that are matched in time.

[0026] By aligning the time series, the learner's performance was successfully stretched or compressed in the time dimension, thus achieving precise quantification of rhythmic deviations. Even when there are changes in the performance speed, it can accurately identify which note is wrong. This solves the problem of the learner's performance not being perfectly synchronized with the standard score in time, ensuring that the actual note i can accurately match the standard note j.

[0027] This system solves the fundamental time synchronization problem in music teaching systems, achieving precise positioning and quantification of performance sequences. Utilizing a dynamic time warping algorithm, the system can calculate the minimum total matching cost path between the performance sequence and the standard score. This means that even if the learner's playing speed varies (e.g., fluctuating between fast and slow), the system can accurately identify which actual note corresponds to which standard note, thus achieving precise quantification of rhythmic deviations and laying the accurate foundation for all subsequent technical indicator calculations. This alignment method ensures that errors can be accurately attributed to specific notes during performance evaluation.

[0028] The mathematical expression for the rhythm time value deviation rate is: ; Where N is the total number of musical notes. This represents the duration of the i-th note in the actual note event set, where the duration is the difference between the timestamps of the note closing event and the note opening event. Let j be the duration of the standard note in the set of events that corresponds to the duration of the i-th note in the actual note event set. Let i be the weight of the i-th note. The calculation process includes: Extract key local features of the notes, including the dynamic value of the note, the relative value of the note duration, the signal-to-noise ratio of the note, boundary notes, and accented notes; The relative duration of a note is the proportion of the note's duration to the total duration of the musical segment; The signal-to-noise ratio of a note is the local signal-to-noise ratio of note i during the note's duration; Boundary notes are notes on potential prosodic boundaries; The accented note is the first note of each measure that matches the rhythm of the score. The set of all key local features corresponding to a random note is represented as a one-dimensional local note feature vector. The weight of the current note is calculated based on the local note feature vector, and its mathematical expression is as follows: ; ; in, For musical notes The local note feature vector, The preset feature sensitivity parameter is used to control the scaling and distribution of the input vector by the Softmax function. This dynamic weighting method ensures that the evaluation of rhythmic accuracy prioritizes the most structurally and expressively critical notes in the music.

[0029] The rhythmic duration deviation rate is calculated as a weighted average of the relative duration deviations of all notes, with the denominator being... This ensures error normalization, allowing for a fair comparison of errors between long and short notes. The closer it is to 0, the more accurate the rhythm. It is used to measure the relative error between the duration of a note in an actual performance and the standard timestamp in the score.

[0030] The mathematical expression for pitch accuracy is: ; ; in, For the first The cent deviation of each note The preset maximum allowable pitch deviation threshold, For pitch accuracy; In this embodiment, the preset maximum allowable pitch deviation threshold is 50 cents, which is a quarter of a semitone. In the process of calculating pitch accuracy, the absolute frequency deviation is converted into pitch cent deviation. Pitch cent deviation is the standard unit for measuring pitch difference in musical acoustics. The average value of the pitch cent deviation of all notes is calculated and normalized relative to the maximum allowable pitch cent deviation threshold. The closer the pitch accuracy is to 1, the more accurate the pitch is. It is used to measure the degree of deviation between the actual performance pitch and the standard pitch of the score. The mathematical expression for the interval tension scalar is: ; in, For interval tension scalar, The total number of consecutive note pairs or chords. This refers to the interval between the k-th note pairs or chords, expressed in semitones. It is a discrete interval tension function; Reference Figure 2 and Figure 3 These are the interval table and the discrete interval tension function, respectively. The discrete interval tension function is used to convert intervals... The tension value is mapped to the interval [0, 1], where 0 represents extreme consonance and 1 represents extreme dissonance. When calculating the interval tension scalar, it is directly queried. Figure 3 Get the corresponding semitone number Tension weight The higher the frequency of high-weight intervals in a musical passage, the greater the final interval tension scalar value. The higher the value, the more it is used in the Empathic Reasoning and Fusion module to enhance the artistic connotation of the Dionysian spirit in Nietzsche's musical philosophy.

[0031] The interval tension scalar is used to measure the rational manifestation of the emotional psychological tension generated by the interval relationship between consecutive note pairs and chords in music. It directly reflects the inherent harmony and conflict in music. In the empathy reasoning and fusion module, the high interval tension scalar is input into the empathy vector extraction unit to generate the artistic connotation scalar of the Dionysian spirit (conflict, chaos), which is used to measure the rational manifestation of the emotional psychological tension generated by the interval relationship between consecutive note pairs and chords in music.

[0032] The calculation process of mode type scalar includes: A pitch frequency histogram is collected. This histogram represents the frequency distribution of each of the 12 semitones in the corresponding music segment from the audio data and instrument digital interface data. The pitch frequency histogram is normalized to obtain the actual pitch distribution. The cosine similarity between the actual pitch distribution and the target mode template is calculated to obtain the mode matching score. All cosine similarities are aggregated into a one-dimensional vector to obtain the mode type scalar. The expression for the mode type scalar is: ; in, For mode type scalar, For the target debugging template, For the actual pitch distribution and target mode template Cosine similarity between them; The highest-scoring element of the mode type scalar represents the main mode type of the current musical segment.

[0033] Modal type scalars are used to identify and quantify the scale system upon which a musical fragment is based, particularly its cultural and emotional attributes. A modal type scalar is a multi-dimensional vector where each element represents the degree of matching between the fragment and a specific modal template. A higher modal matching score indicates that the musical fragment conforms more closely to the structure of the target modal template. The target modal template is obtained through big data retrieval, including the frequency of each chromatic scale in the 12 semitones corresponding to a unit measure of each standard key, and is then normalized.

[0034] This approach achieves musicological optimization and artistic relevance of objective indicators. The rhythmic value deviation rate employs dynamic weighting, assigning weights based on key local features such as note dynamics and accents. This ensures that the assessment prioritizes the most crucial notes in the musical structure, avoiding excessive punishment of non-emphasis notes and improving the musicological relevance of the score. The interval tension scalar, for the first time, transforms objective interval relationships into a rational representation of emotional psychological tension, directly reflecting the inherent harmony and conflict in music. Its high value is explicitly designated to enhance the artistic connotation of the Dionysian spirit in Nietzsche's musical philosophy, achieving a direct link between technical indicators and philosophical understanding.

[0035] Prosodic structure analysis is performed on audio data to extract prosodic boundary detection intensity and prosodic processing scalar. The prosodic structure analysis includes: Traverse the loudness energy trajectory of the audio data, identify regions where the loudness energy is lower than the preset loudness energy threshold and the duration exceeds the minimum inter-sentence pause threshold, and mark them as potential prosodic boundaries; The actual duration of the note preceding the potential prosodic boundary is detected to be greater than the standard duration. If the actual duration is greater than the standard duration, the note is marked as a phrase-end extended note, and the difference between the actual duration and the standard duration is calculated and recorded as the duration extension. The proportion of the duration extension to the total duration of the music segment corresponding to the current audio data is calculated and recorded as the duration extension ratio. The standard duration is the average of the actual durations of all notes in the current audio data. Detect whether the fundamental frequency of the notes at the potential rhythmic boundary drops. If the fundamental frequency drops, mark the note as the phrase end marker note. Calculate the average difference between the fundamental frequency of the current note and the fundamental frequency of the previous note, and record it as the fundamental frequency drop magnitude. The intensity feature value of the note on the potential prosodic boundary is obtained by multiplying the duration, duration extension ratio and fundamental frequency drop of the note on the potential prosodic boundary. The average value of the intensity feature value of all notes on the potential prosodic boundary is calculated to obtain the prosodic boundary detection intensity scalar. The prosodic processing scalar is obtained by multiplying the prosodic boundary detection intensity scalar with the rhythmic time value deviation rate. The higher the prosodic boundary detection intensity scalar, the clearer the performer's phrasing and the stronger the sense of structure; The prosodic processing scalar is used to measure a learner’s sensitivity to phrase segmentation. The prosodic processing scalar is a composite scalar that combines the prosodic boundary detection strength with the local performance of rhythmic time value deviation rate at the boundary. If the performer’s time value control at the boundary is accurate and intentionally prolonged, the prosodic processing scalar value is high, indicating strong prosodic processing ability.

[0036] Prosodic structure analysis aims to quantify the breathiness and phrasing clarity of musical performance, which are key to measuring performance fluency and structural comprehension. By analyzing loudness energy trajectories, duration, and fundamental frequency drop, prosodic boundary detection strength and prosodic processing scalars are extracted to objectively assess learners' understanding of phrase structure and emotional pauses. The prosodic processing scalar combines phrasing clarity with rhythmic control; a high scalar value indicates strong prosodic processing ability if the performer's control of time values ​​at boundaries is accurate and intentionally prolonged.

[0037] Music structure analysis is performed on audio data and instrument digital interface data to extract linearity scalars and discontinuous scalars. The music structure analysis process includes: A pattern matching algorithm is executed on the standard note event set to identify and obtain the motif set. The process of identifying and obtaining the motif set includes: Encode a relative feature vector for all corresponding notes in the standard note event set. The relative feature vector includes a relative pitch feature and a relative rhythm feature. The relative pitch feature is the number of semitones between the current note and the previous note, and the relative rhythm feature is the ratio of the duration of the current note to the duration of the previous note. The minimum and maximum motive lengths are preset, with motive length expressed in units of notes. Traverse the standard note event set to obtain all complete melody fragments with lengths between the minimum and maximum motif lengths. The complete melody fragments are in units of measures. Calculate the note similarity between any two complete melody fragments using the Longest Common Subsequence (LCS) algorithm. Obtain the note similarity of each complete melody fragment in the standard note event set with any other complete melody fragment. Count the number of note similarities for each complete melody fragment that exceed a preset note similarity threshold. When the number is greater than a preset minimum repetition threshold, output the current complete melody fragment as a valid motif fragment. Iteratively check whether all valid motivational fragments include shorter valid motivational fragments. If they do, output the shorter valid motivational fragment as a basic motivational fragment. If they do not include shorter valid motivational fragments, directly output the valid motivational fragment as a basic motivational fragment. Combine all basic motivational fragments into a motivational set. Record the pitch profile, rhythmic pattern, and digital instruction sequence of each basic motif fragment, and structure each element in the motif set into a feature vector including pitch profile, rhythmic pattern, and internal interval tension to obtain the motif feature set; Traverse the standard note event set to obtain all occurrence positions of each basic motif fragment in the musical phrase, and identify the developmental variation type at each occurrence position. The developmental variation types include repetition, reflection, expansion, and fragment development. Use big data to obtain the pitch change features and rhythm change features of each developmental variation type, and encode the developmental variation type and its pitch change features and rhythm change features in a non-repetitive manner, combining them into a variation coding vector. Extract the actual motivic melody fragments at the corresponding occurrence positions of the actual note event set, obtain the variation encoding vector of the actual motivic melody fragments, and check whether the variation encoding vector of the actual motivic melody fragments matches the variation encoding vector of the basic motivic fragment at the corresponding occurrence position. The matching logic is that the elements at the corresponding positions of the variation encoding vectors are the same non-repeating encoded mathematical operation logic, such as addition and proportion. Based on the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions, the linearity is calculated. The mathematical expression for linearity is: ; in, For linearity, The total number of variations being tracked. The Euclidean distance between the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions is given. Let k be the variation encoding vector of the basic motif segment. The variation encoding vector for the actual motif melody segment of the kth segment; Identify interval tension scalars that exceed the dissonance threshold, monitor the switching frequency of mode type scalars within a unit time window, such as a sudden switch from C major to F# major, reflecting polytonal overlap or tonal ambiguity, identify sudden increases in rhythmic value deviation rate at non-prosodic boundaries, sudden increases indicate discontinuous sound wave value growth, and calculate discontinuity. ; in, For discontinuity, The average value of the interval tension scalar that exceeds the dissonance threshold. The number of times the scalar switch to the mode type is performed. This represents the average amplitude of sudden changes in the rhythmic time value deviation rate.

[0038] The pitch profile is represented by a relative pitch sequence, which takes the first note as the original 0 value and records the interval value difference between all subsequent notes and the first note. The rhythm pattern is represented by a relative duration sequence, which takes the smallest note duration unit of the basic motif segment as the smallest duration unit (e.g., an eighth note) and records the ratio of the duration between all subsequent notes and the first note. The preset note similarity threshold is set between 75% and 95%. If the threshold is set too low, irrelevant segments will be incorrectly identified as motifs; if the threshold is set too high, normal developmental variations will be missed.

[0039] Big data was used to find fixed pitch variation data and fixed rhythm variation data for various types of expansion variation. Based on recurrent neural network (RNN), the fixed pitch variation data and fixed rhythm variation data were converted into pitch variation features and rhythm variation features. Linearity is used to measure the accuracy of all traced variations in the reproduction, and is designed to quantify the performer’s expression of the causal logic (linear thinking) of the melody.

[0040] High discontinuity (high conflict and chaos) is fed into the empathy vector extraction unit to enhance the artistic connotation scalar of the Dionysian spirit.

[0041] The system provides core technical inputs for Nietzsche's philosophical mapping, enabling it to understand the philosophical connotations of music at a structural level. The linearity scalar quantifies the logic and accuracy of melodic motif development through expansive variations (repetition, reflection, expansion, and fragmented development), representing the orderly and rational aspects of music. The discontinuity scalar, by integrating interval tension peaks, modal switching frequencies, and rhythmic abrupt changes, quantifies structural and emotional abrupt changes, conflicts, and a sense of chaos, directly representing the irrational and conflicting aspects of music (the Dionysian spirit). The generation of these two indicators allows the system to conduct advanced artistic analysis at the level of the inherent logic of musical structure. This musical structural analysis process is fundamental to understanding the inherent logic and artistic conflicts of melody.

[0042] The standard tone of each Chinese character in the lyrics text is retrieved. The standard tone includes high level tone, rising tone, falling-rising tone and falling tone. The pitch direction is obtained according to the standard tone. The pitch direction corresponding to the standard tone is as follows: high level tone and rising tone correspond to flat, falling-rising tone corresponds to rising, and falling tone corresponds to falling. Extract the actual pitch direction of the fundamental frequency trajectory of each Chinese character in the corresponding original sound wave signal segment and the standard pitch direction of the standard musical score of the corresponding segment. Calculate the tone melody violation rate and semantic clarity scalar based on the actual pitch direction and the standard pitch direction. The process of calculating the tone melody violation rate includes: comparing whether the actual pitch direction of the Chinese character is the same as the standard pitch direction of the corresponding position. If they are not the same, the Chinese character is marked as a tone melody violation. The ratio of the number of Chinese characters marked as tone melody violations to the total number of Chinese characters is calculated to obtain the tone melody violation rate. Subtracting the tone-melody violation rate from 1 yields the semantic clarity scalar.

[0043] The high semantic clarity scalar indicates that the lyrics are accurately conveyed, supporting clear audio narrative. This scalar is then used in the empathy reasoning and fusion module to assist in generating an audio narrative coherence scalar.

[0044] By analyzing acoustic signals, the brain's perception of the rhythmic boundaries of musical phrases is simulated, and a rhythmic processing scalar is output. Linearity scalar and discontinuous scalar representations quantify the linear thinking of melody by analyzing how melodic motifs develop through developmental variations. The semantic clarity scalar is obtained by calculating the tonal melody violation rate of the lyrics fundamental frequency trajectory and the melody pitch direction.

[0045] This method quantifies the degree of matching between language and music, significantly improving the quality assessment of vocal narrative in singing instruction. By comparing the standard tonal trends of Chinese characters with the actual melodic trends, the tonal-melody violation rate is calculated, and a semantic clarity scalar is obtained. This solves the musicological problem of clear pronunciation in singing instruction, ensuring that the semantics of lyrics can be accurately conveyed, and is a key foundation for generating a vocal narrative coherence scalar.

[0046] The empathy reasoning module includes an emotion theory mapping model construction unit and an empathy vector extraction unit: The emotion theory mapping model building unit is used to construct the emotion theory mapping model, which includes: The rhythmic value deviation rate and mode type scalar are used as input features; The elements in the mode type scalar are multiplied by a weight matrix and then summed in a weighted manner. The weight matrix includes the sentiment tendencies of all mode types in the mode type scalar. The weighted sum is then normalized to obtain the sentiment polarity features. When the emotional polarity metric is greater than 0.5 and less than 1, it indicates a positive emotion. When the emotional polarity metric is less than 0.5 and greater than 0, it indicates negative emotion; The rhythmic timing deviation rate, the difference between 1 and the rhythmic timing deviation rate, and the emotional polarity feature are concatenated into a fused feature vector. This vector is then input into the first fully connected layer and output through the ReLU nonlinear activation function. The output is then input into the second fully connected layer and output through the ReLU nonlinear activation function. Finally, the output is input into a linear layer and output through the Softmax activation function to obtain an emotional expression scalar. Each element in the emotional expression scalar represents the probability of matching the performer's musical performance with each emotion. The dimensions of the emotional expression scalar correspond to eight major emotional clusters, which include happiness, solemnity, calmness, sadness, vitality, fantasy, surprise, and humor. The weight matrix is ​​preset with initial values ​​manually based on all mode types in the mode type scalar. Through the above structure, the model realizes the transformation of the learner's objective rhythmic accuracy and tonality selection in the performance into an artistically meaningful scalar of emotional expression that can be used for empathetic dialogue.

[0047] The empathy vector extraction unit includes: By mapping the discontinuous scalar and the interval tension scalar to artistic connotations according to Nietzsche's philosophical interpretation rules, we obtain the artistic connotation scalar. Nietzsche's philosophical interpretation rules include: The discontinuous scalar, linearity scalar, and interval tension scalar are input into the emotion theory mapping model at the input positions of rhythmic time deviation rate, the difference between 1 and rhythmic time deviation rate, and emotional polarity features, and the output Dionysian spirit component features are then used to output the model. The Dionysian spirit represents irrationality, conflict, chaos, and fanaticism, primarily driven by the inherent tension and structural instability of music. A high level of Dionysian spirit indicates that the performance contains strong artistic expressions of irrationality, conflict, or fanaticism. The difference between 1 and the rhythmic time deviation rate, the rhythmic time deviation rate, and the interval tension scalar are input into the emotion theory mapping model at the input positions of the rhythmic time deviation rate, the difference between 1 and the rhythmic time deviation rate, and the emotional polarity feature, and the power will component feature is output. The will to power represents the artistic power to transcend obstacles, conquer and control conflicts. It is not simply accuracy, but accuracy under complexity. A high weight of the will to power indicates that the performer demonstrates a powerful and transcendent control in artistic expression. By combining the characteristics of the Dionysian spirit and the characteristics of the will to power, we obtain the scalar of artistic connotation; Nietzsche's philosophical interpretation rules are used to transform the structural conflicts and tensions manifested in performance into quantitative assessments of the Dionysian spirit and the will to power in Nietzsche's philosophy, thereby generating a scalar of artistic connotation.

[0048] This unit aims to map the structural conflicts and tensions in musical performances into deeper artistic and philosophical connotations.

[0049] Input the semantic clarity scalar and the prosodic boundary detection intensity scalar into the emotion theory mapping model, and output the voice narrative coherence scalar.

[0050] The empathic reasoning module is the intelligent hub of the system, responsible for transforming objective technical features into application feature scalars with emotional and artistic dimensions, and incorporating learners' personalized preferences.

[0051] The core creative aspect of this invention lies in its transformation of technological characteristics into artistic cognition and personalized preferences. Inputs of high discontinuity scalars (structural chaos) and high interval tension scalars (intrinsic conflict) are mapped, through a Nietzschean philosophical interpretive model, into a high-intensity Dionysian spiritual component. This demonstrates that the system objectively identifies the intense irrationality, conflict, and fervent artistic expression present in the performance. By quantifying the will to power through assessing the performer's control over complexity, it evaluates whether the performer can maintain an extremely low rhythmic value deviation rate (high control) even when the inherent difficulty of the music (high interval tension scalar) is high. This ability to manage conflict and overcome obstacles is directly quantified as a high-intensity will to power component.

[0052] This mapping enables the system to understand the artistic motivations behind a performer's technical errors, elevating technical data to a philosophical level and thus achieving a deep assessment of the artistic content.

[0053] By combining Hefner's theory to map rhythm (stability) and mode (polarity) into quantifiable emotional clusters, precise identification of the emotional nuances of a performance is achieved. Semantic clarity and prosodic boundary strength are combined to assess the fluency of a performer's vocal storytelling.

[0054] The teaching application module includes: Using scalars of emotional expression, scalars of vocal narrative coherence, and scalars of artistic connotation as conditions, the input condition drives the dialogue generation model LLM to generate personalized empathetic dialogues that address current technical feature deviations. Types of technical feature deviations include: when any one of the following features—rhythm value deviation rate, pitch accuracy, interval tension scalar, and mode type scalar—exceeds the preset threshold range, a technical feature deviation occurs. The corresponding threshold ranges are manually set. Based on the types of technical feature deviations included in the personalized empathic dialogue, index matching is performed in the preset Dalcroze body rhythm teaching instruction library. Based on the matched instruction template and specific deviation values, body movement instructions are generated. The body movement instructions include stride change instructions for correcting rhythm, body extension instructions for perceiving intervals, and imitation movement instructions for reflecting mode types.

[0055] The teaching application module is used to transform the emotional analysis results output by the empathic reasoning module into specific, actionable, and personalized teaching instructions, including verbal feedback and physical movement guidance.

[0056] In this embodiment, the language feedback includes: First, affirm the artistic effect of the performance, for example, "I felt the strong Dionysian spirit and sense of conflict in your performance"; Clearly point out the corresponding technical feature deviation, such as "This sense of conflict stems from your high rhythmic value deviation rate in section 3"; Explain the impact of technical biases on application features, such as "rhythmic instability weakens the coherence of your vocal narrative, making it difficult for listeners to follow the linear development of the melody."

[0057] The pre-defined Dalcroze body movement instruction library structure and matching logic include: Rhythm value deviation rate deviation type matching stride change instruction: When the personalized empathic dialogue shows a fast and unstable pace, the instruction is "Slow down your pace and use larger, more even steps to measure the value of the quarter note". Interval tension scalar deviation type matching body extension instructions: When the personalized empathic dialogue shows high tension in large leaps, the instruction is: "When playing large leaps, please open your arms at the same time and feel the stretch in your body to experience the tension of the interval." Tuning type deviation type matching mimics action instructions: When a matching ethnic mode is displayed in the personalized empathy dialogue, the instruction is "Please imitate the typical dance movements of this mode's cultural background to experience its unique cultural tone." This step aims to transform abstract musical concepts into concrete, perceptible, and executable physical movement instructions, thereby achieving a deep integration of the Dalcroze teaching method.

[0058] This approach transforms the results of empathic reasoning into concrete and actionable teaching actions. Utilizing the Large Language Model (LLM), which considers artistic connotation and emotional expression, it generates personalized empathic dialogues with causal chains. The dialogue content is no longer limited to mechanical pitch errors, but rather links technical deviations (such as rhythmic inaccuracies) with artistic consequences (such as weakened vocal narrative coherence). This overcomes the rigidity and one-sidedness of traditional feedback, enabling learners to understand how technical deficiencies directly affect their artistic expression. It directly matches abstract musical concepts with concrete bodily movement instructions, achieving an embodied cognitive teaching model. Learners internalize musical concepts through bodily perception, significantly improving learning efficiency and musical perception. This deep integration transforms abstract musical concepts into intuitive bodily experiences.

[0059] This invention achieves precise alignment of multimodal music performance data through dynamic time warping and innovatively generates multi-dimensional advanced technical indicators, including interval tension scalars, discontinuity scalars, and semantic clarity scalars, laying the foundation for accurate analysis. It transforms objective technical indicators into artistic cognition with profound philosophical connotations, identifies the emotional polarity of performance through an emotion theory mapping model, and quantifies high interval tension scalars and high discontinuity scalars into Dionysian spiritual components through an empathy vector extraction unit. Simultaneously, it transforms the high control power under complexity into a power will component, thereby achieving a deep assessment of the driving force of performing arts. This philosophically empowered analysis enables the system to understand the artistic motivations behind technical errors, generating personalized empathic dialogues with causal chains, overcoming the rigidity and one-sidedness of traditional feedback. In terms of pedagogical applications, this system deeply integrates abstract musical concepts with the Dalcroze body movement teaching method, directly matching technical deviations such as high rhythmic value deviation rates to concrete stride change instructions and body extension instructions, realizing a embodied cognitive teaching model, greatly improving learners' perception and internalization efficiency of musical concepts.

[0060] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A music teaching system based on speech recognition, characterized in that, It includes a module for quantifying technical characteristics, a module for empathic reasoning, and a module for teaching applications: The technical feature quantization module is used to acquire the learner's music performance data through speech recognition and the Musical Instrument Digital Interface (MIDI). After aligning the music performance data with a time series based on dynamic time warping, it calculates the rhythmic value deviation rate, pitch accuracy, interval tension scalar, mode type scalar, prosodic boundary detection intensity, linearity scalar, discontinuity scalar, and semantic clarity scalar. The empathy reasoning module is used to construct an emotion theory mapping model, and based on the emotion theory mapping model, it generates scalars of emotion expression, artistic connotation, and audio narrative coherence. The teaching application module is used to take the emotional expression scalar, the voice narrative coherence scalar, and the artistic connotation scalar as conditions, input conditions to drive the dialogue generation model, and generate personalized empathic dialogue instructions and Dalcroze body rhythm teaching instructions that are tailored to the current technical feature deviations.

2. The music teaching system based on speech recognition as described in claim 1, characterized in that, The technical feature quantification module includes: The music performance data includes audio data of the learner's singing and playing, digital interface data of the instrument, and lyrics text corresponding to the audio data; The audio data of learners singing and playing includes the raw sound wave signal collected by the microphone when learners sing and play non-musical digital interface devices. Non-musical digital interface devices include human voices and acoustic musical instruments. The raw sound wave signal includes the amplitude sampling point sequence, fundamental frequency trajectory, loudness energy trajectory and Mel frequency cepstral coefficients of the sound signal stored in FLAC format on the time axis. The digital interface data for musical instruments includes a sequence of digital instructions generated when a learner plays an electronic musical instrument. The sequence of digital instructions includes note on events, note off events, pitch data, velocity values, and control change events. The note activation event includes the timestamp of the key being pressed and the velocity value of the note; The note-off event includes the timestamp of when the key was released; Pitch data includes discrete pitch values ​​represented by natural numbers from 1 to 127 recorded in the musical score; The force value is a numerical value attached to the note activation event, recording the speed and pressure at which the key is pressed. Control change events include changes in the usage status of the pedal, vibrato wheel, and pitch bend wheel controllers during performance.

3. The music teaching system based on speech recognition as described in claim 2, characterized in that: After time-series alignment of the audio data of learners’ singing and playing and the digital interface data of instruments, music technical features are extracted. The music technical features include rhythmic value deviation rate, pitch accuracy rate, interval tension scalar and mode type scalar. Time series alignment includes: The standard sheet music uploaded by learners is converted into a standard note event set. The mathematical expression for the standard note event set is: ; ; Where S is the set of musical note events. For the first A standard musical note event, This is the standard timestamp for the note, which is the midpoint between the note's close event and its open event. Standard pitch data for musical notes. The total number of notes in a standard musical score; The audio data of the learner’s actual singing and playing is converted into actual note event sets. The actual note event sets include all actual note events of the learner’s actual singing and playing. The structure of the actual note event sets is the same as that of the standard note event sets. Each actual note event set includes the actual timestamp and actual pitch data of the notes. Construct a cost matrix, initializing it to have dimensions m×n, where m is the number of actual note events included in the actual note event set. The cost matrix represents the cost of the notes in the performance sequence. notes in the standard sequence The local matching cost incurred during the matching process; The mathematical expression for calculating the cost of local matching is: ; in, For local matching cost, For the actual pitch data of the i-th note in the actual note event set, Let i be the actual timestamp of the i-th note in the actual note event set. and These are the preset weighting coefficients.

4. The music teaching system based on speech recognition as described in claim 3, characterized in that: Initialize the m×n cumulative cost matrix D. Initialization includes setting the starting point of the cumulative cost matrix D to... The first line is set to Following the rule that it can only move horizontally from the predecessor node on the left, the first column is set to... It follows the rule that it can only move from the predecessor node above; Using dynamic programming, we iteratively calculate each element in the cumulative cost matrix D to find the path with the minimum cumulative cost from the starting point to the ending point. The iterative calculation includes: Elements in the cumulative cost matrix D The value of is equal to the current local matching cost plus the minimum cumulative cost selected from the three neighboring predecessor nodes, and its mathematical expression is: ; After the iterative calculation is completed, the lower right element of the cumulative cost matrix D is... That is, the minimum total matching cost. Backtracking backwards along the path that leads to the minimum cumulative cost back to the starting point The backtracking path is the optimal alignment path. The path that results in the minimum cumulative cost is the path consisting of the predecessor nodes selected when the minimum cumulative cost is chosen. The optimal alignment path includes index pairs. This refers to the actual note event and the standard note event that are matched in time.

5. The music teaching system based on speech recognition as described in claim 4, characterized in that: The mathematical expression for the rhythm time value deviation rate is: ; Where N is the total number of musical notes. This represents the duration of the i-th note in the actual note event set, where the duration is the difference between the timestamps of the note closing event and the note opening event. Let j be the duration of the standard note in the set of events that corresponds to the duration of the i-th note in the actual note event set. Let i be the weight of the i-th note. The calculation process includes: Extract key local features of the notes, including the dynamic value of the note, the relative value of the note duration, the signal-to-noise ratio of the note, boundary notes, and accented notes; The relative duration of a note is the proportion of the note's duration to the total duration of the musical segment; The signal-to-noise ratio of a note is the local signal-to-noise ratio of note i during the note's duration; Boundary notes are notes on potential prosodic boundaries; The accented note is the first note of each measure that matches the rhythm of the score. The set of all key local features corresponding to a random note is formed into a one-dimensional local note feature vector. The weight of the current note is calculated based on the local note feature vector, and its mathematical expression is as follows: ; ; in, For musical notes The local note feature vector, These are preset feature sensitivity parameters; The mathematical expression for pitch accuracy is: ; ; in, For the first The cent deviation of each note The preset maximum allowable pitch deviation threshold, For pitch accuracy; The mathematical expression for the interval tension scalar is: ; in, For interval tension scalar, The total number of consecutive note pairs or chords. This refers to the interval between the k-th note pairs or chords, expressed in semitones. It is a discrete interval tension function; The calculation process of mode type scalar includes: A pitch frequency histogram is collected. This histogram represents the frequency distribution of each of the 12 semitones in the corresponding musical segment from the audio data and instrument digital interface data. The pitch frequency histogram is normalized to obtain the actual pitch distribution. The cosine similarity between the actual pitch distribution and the target mode template is calculated to obtain the mode matching score. All cosine similarities are aggregated into a one-dimensional vector to obtain the mode type scalar. The expression for the mode type scalar is: ; in, For mode type scalar, For the target debugging template, For the actual pitch distribution and target mode template Cosine similarity between them; The highest-scoring element of the mode type scalar represents the main mode type of the current musical segment.

6. The music teaching system based on speech recognition as described in claim 5, characterized in that: Prosodic structure analysis is performed on audio data to extract prosodic boundary detection intensity and prosodic processing scalar. The prosodic structure analysis includes: Traverse the loudness energy trajectory of the audio data, identify regions where the loudness energy is lower than the preset loudness energy threshold and the duration exceeds the minimum inter-sentence pause threshold, and mark them as potential prosodic boundaries; The actual duration of the note preceding the potential prosodic boundary is detected to be greater than the standard duration. If the actual duration is greater than the standard duration, the note is marked as a phrase-end extended note, and the difference between the actual duration and the standard duration is calculated and recorded as the duration extension. The proportion of the duration extension to the total duration of the music segment corresponding to the current audio data is calculated and recorded as the duration extension ratio. The standard duration is the average of the actual durations of all notes in the current audio data. Detect whether the fundamental frequency of the note at the potential rhythm boundary drops. If the fundamental frequency drops, mark the note as the phrase end marker note. Calculate the average difference between the fundamental frequency of the current note and the fundamental frequency of the previous note, and record it as the fundamental frequency drop magnitude. The intensity feature value of the note on the potential prosodic boundary is obtained by multiplying the duration, duration extension ratio and fundamental frequency drop of the note on the potential prosodic boundary. The average value of the intensity feature value of all notes on the potential prosodic boundary is calculated to obtain the prosodic boundary detection intensity scalar. The prosodic processing scalar is obtained by multiplying the prosodic boundary detection intensity scalar with the rhythmic time value deviation rate.

7. The music teaching system based on speech recognition as described in claim 6, characterized in that: Music structure analysis is performed on audio data and instrument digital interface data to extract linearity scalars and discontinuous scalars. The music structure analysis process includes: A pattern matching algorithm is executed on the standard note event set to identify and obtain the motif set. The process of identifying and obtaining the motif set includes: Encode a relative feature vector for all corresponding notes in the standard note event set. The relative feature vector includes a relative pitch feature and a relative rhythm feature. The relative pitch feature is the number of semitones between the current note and the previous note. The relative rhythm feature is the ratio of the duration of the current note to the duration of the previous note. The minimum and maximum motive lengths are preset, with motive length expressed in units of notes. Traverse the standard note event set to obtain all complete melody fragments with lengths between the minimum and maximum motif lengths. The complete melody fragments are in units of measures. Calculate the note similarity between any two complete melody fragments using the Longest Common Subsequence (LCS) algorithm. Obtain the note similarity calculated between each complete melody fragment in the standard note event set and any other complete melody fragment. Count the number of note similarities corresponding to each complete melody fragment that exceed a preset note similarity threshold. When the number is greater than a preset minimum repetition threshold, output the current complete melody fragment as a valid motif fragment. Iteratively check whether all valid motivational fragments include shorter valid motivational fragments. If shorter valid motivational fragments are included, output the shorter valid motivational fragments as basic motivational fragments. If shorter valid motivational fragments are not included, directly output the valid motivational fragments as basic motivational fragments. Combine all basic motivational fragments into a motivational set. Record the pitch profile, rhythmic pattern, and digital instruction sequence of each basic motif fragment, and structure each element in the motif set into a feature vector including pitch profile, rhythmic pattern, and internal pitch tension to obtain a motif feature set; Traverse the standard note event set to obtain all occurrence positions of each basic motif fragment in the musical phrase, and identify the developmental variation type at each occurrence position. The developmental variation type includes repetition, reflection, expansion, and fragment development. Use big data to obtain the pitch change features and rhythm change features of each developmental variation type, and encode the developmental variation type and its pitch change features and rhythm change features in a non-repetitive manner to form a variation coding vector. Extract the actual motivic melody fragments at the corresponding occurrence positions of the actual note event set, obtain the variation encoding vector of the actual motivic melody fragments, and check whether the variation encoding vector of the actual motivic melody fragments matches the variation encoding vector of the basic motivic fragment at the corresponding occurrence position. The matching logic is that the elements at the corresponding positions of the variation encoding vectors are the same non-repeating encoding mathematical operation logic. Based on the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions, the linearity is calculated. The mathematical expression for linearity is: ; in, For linearity, The total number of variations being tracked. The Euclidean distance between the variation encoding vectors of the actual motif melodic fragment and the basic motif fragment at corresponding positions is given. Let k be the variation encoding vector of the basic motif segment. The variation encoding vector for the actual motif melody segment of the kth segment; Identify interval tension scalars that exceed the dissonance threshold, monitor the switching frequency of mode type scalars within a unit time window, identify sudden increases in rhythmic time value deviation rate at non-prosodic boundaries, sudden increases indicate discontinuous sound wave numerical growth, and calculate discontinuity. ; in, For discontinuity, The average value of the interval tension scalar that exceeds the dissonance threshold. The number of times the scalar switch to the mode type is performed. This represents the average amplitude of sudden changes in the rhythmic time value deviation rate.

8. The music teaching system based on speech recognition as described in claim 7, characterized in that: The standard tone of each Chinese character in the lyrics text is retrieved. The standard tone includes high level tone, rising tone, falling-rising tone and falling tone. The pitch direction is obtained according to the standard tone. The pitch direction corresponding to the standard tone is as follows: high level tone and rising tone correspond to flat, falling-rising tone corresponds to rising, and falling tone corresponds to falling. Extract the actual pitch direction of the fundamental frequency trajectory of each Chinese character in the corresponding original sound wave signal segment and the standard pitch direction of the standard musical score of the corresponding segment. Calculate the tone melody violation rate and semantic clarity scalar based on the actual pitch direction and the standard pitch direction. The process of calculating the tone melody violation rate includes: comparing whether the actual pitch direction of the Chinese character is the same as the standard pitch direction of the corresponding position. If they are not the same, the Chinese character is marked as a tone melody violation. The ratio of the number of Chinese characters marked as tone melody violations to the total number of Chinese characters corresponding to the original sound wave signal is calculated to obtain the tone melody violation rate. Subtracting the tone-melody violation rate from 1 yields the semantic clarity scalar.

9. The music teaching system based on speech recognition as described in claim 8, characterized in that, The empathy reasoning module includes an emotion theory mapping model construction unit and an empathy vector extraction unit: The emotion theory mapping model building unit is used to construct an emotion theory mapping model, which includes: The rhythmic value deviation rate and mode type scalar are used as input features; The elements in the mode type scalar are multiplied by a weight matrix and then summed in a weighted manner. The weight matrix includes the sentiment tendencies of all mode types in the mode type scalar. The weighted sum is then normalized to obtain the sentiment polarity features. When the emotional polarity metric is greater than 0.5 and less than 1, it indicates a positive emotion. When the emotional polarity metric is less than 0.5 and greater than 0, it indicates negative emotion; The rhythmic timing deviation rate, the difference between 1 and the rhythmic timing deviation rate, and the emotional polarity feature are concatenated into a fusion feature vector. This vector is then input into the first fully connected layer and output through the nonlinear activation function ReLU. The output is then input into the second fully connected layer and output through the nonlinear activation function ReLU. Finally, the output is input into a linear layer and output through the Softmax activation function to obtain an emotional expression scalar. Each element in the emotional expression scalar represents the probability of matching the performer's musical performance with each emotion. The empathy vector extraction unit includes: By mapping the discontinuous scalar and the interval tension scalar to artistic connotations according to Nietzsche's philosophical interpretation rules, we obtain the artistic connotation scalar. These Nietzschean philosophical interpretation rules include: The discontinuous scalar, linearity scalar, and interval tension scalar are input into the emotion theory mapping model at the input positions of rhythmic time deviation rate, the difference between 1 and rhythmic time deviation rate, and emotional polarity features, and the output Dionysian spirit component features are then generated. The difference between 1 and the rhythmic time deviation rate, the rhythmic time deviation rate, and the interval tension scalar are input into the emotion theory mapping model at the input positions of the rhythmic time deviation rate, the difference between 1 and the rhythmic time deviation rate, and the emotional polarity feature, and the power will component feature is output. By combining the characteristics of the Dionysian spirit and the characteristics of the will to power, we obtain the scalar of artistic connotation; Input the semantic clarity scalar and the prosodic boundary detection intensity scalar into the emotion theory mapping model, and output the voice narrative coherence scalar.

10. The music teaching system based on speech recognition as described in claim 9, characterized in that, The teaching application module includes: Using the emotional expression scalar, the voice narrative coherence scalar, and the artistic connotation scalar as conditions, the input condition drives the dialogue generation model LLM to generate personalized empathetic dialogues that address the current technical feature deviations. The types of technical feature deviations include: when any one of the feature values ​​of rhythmic value deviation rate, pitch accuracy rate, interval tension scalar, and mode type scalar exceeds the preset corresponding threshold range, a technical feature deviation occurs. Based on the types of technical feature deviations included in the personalized empathic dialogue, index matching is performed in the preset Dalcroze body rhythm teaching instruction library. Based on the matched instruction template and specific deviation values, body movement instructions are generated. The body movement instructions include stride change instructions for correcting rhythm, body extension instructions for perceiving intervals, and imitation action instructions for reflecting mode types.

Citation Information

Patent Citations

  • Music teaching system based on artificial intelligence

    CN120318041A