AI intelligent music classroom multi-module interactive teaching system and method
By collecting multimodal data in real time in the intelligent music classroom to generate three-dimensional feedback heat maps and combining AR equipment and edge computing, dynamically adjusting the performance equipment parameters, solving the problem that the existing system cannot adapt to individual differences, realizing personalized and immersive closed-loop teaching, and improving teaching quality and efficiency.
Patent Information
- Application Number
- CN202510539352.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
The existing intelligent music interactive teaching system lacks the precise evaluation mechanism and low-latency feedback mechanism of multimodal data, which makes it difficult for teaching strategies to dynamically adapt to individual students' differences, and it is impossible to achieve real-time correction and personalized teaching.
Through the perception acquisition module, students' performance posture, instrument touch pressure signals and ambient sound field data are collected in real time, a three-dimensional feedback heat map is generated, and adaptive teaching strategies are formulated based on historical performance data, and closed-loop feedback mechanism is established using AR equipment and edge computing technology to dynamically adjust the key damping coefficient and sustain pedal response to achieve immediate correction prompts.
It realizes high-dimensional accurate assessment of students' music expression and personalized and immersive teaching, improves teaching quality and efficiency, builds a low-latency and visual real-time interactive platform, and enhances the integration of physics training and digital teaching.
Smart Images

Figure CN120452262A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of teaching interaction technology, and specifically to a multi-module interactive teaching system and method for an AI intelligent music classroom. Background Art
[0002] Against the backdrop of the rapid development of digital education, intelligent music interactive teaching, as an innovative teaching method that integrates artificial intelligence, sensing technology, and interactive media, is gradually being applied to various music education scenarios. Typical intelligent music interactive teaching systems usually include basic modules such as audio analysis, performance scoring, rhythm matching, and performance movement recognition, which can achieve basic performance correction and teaching feedback. However, existing intelligent music interactive teaching systems, on the one hand, rely only on single audio input or basic image recognition, lack comprehensive perception and accurate evaluation of the multi-dimensional behavioral characteristics of students during performance, and are difficult to comprehensively and objectively reflect students' performance and provide sufficient basis for the formulation of teaching strategies. On the other hand, the existing feedback mechanism mainly relies on post-class scoring or text prompts, lacking real-time and interactivity, making it difficult to dynamically adjust teaching strategies according to individual differences among students, making it difficult for students to correct errors immediately during performance, affecting teaching effectiveness. Summary of the Invention
[0003] This application provides an AI smart music classroom multi-module interactive teaching system and method, which solves the technical problem that the existing technology lacks an accurate evaluation mechanism and low-latency feedback mechanism based on multimodal data, resulting in difficulty in dynamically adapting teaching strategies to individual differences among students. It achieves the technical effect of realizing personalized, immersive interactive teaching through somatosensory interactive perception, three-dimensional heat map visualization and AR closed-loop prompts, thereby improving the quality and efficiency of music teaching.
[0004] In view of the above problems, on the one hand, the present application provides an AI smart music classroom multi-module interactive teaching system, which includes: a perception acquisition module, which is used to perceive and acquire music expression evaluation indicators including student playing posture, instrument touch pressure signals and environmental sound field data in the music classroom, and segment the instrument track and background noise, generate performance error feature vectors and map them to a three-dimensional feedback heat map; a strategy formulation module, which is used to upload the three-dimensional feedback heat map to the interactive interface corresponding to the teacher port, and formulate an adaptive teaching strategy based on the historical performance data marked under the student port and the music expression evaluation indicators. The interactive interface supports gesture operation / touch operation for area labeling; an information synchronization module, which is used to synchronize the area labeling information to the AR visualization device of the student port in real time through the edge computing gateway, and establish a closed-loop feedback mechanism according to the adaptive teaching strategy; an environment construction module, which is used to dynamically adjust the key damping coefficient and the sustain pedal response sensitivity according to the quality of the music performance, and build a collaborative training environment of the physical layer and the digital layer; a feedback prompt module, which is used to provide instant correction prompts based on the closed-loop feedback mechanism in the collaborative training environment using the adaptive teaching strategy.
[0005] On the other hand, the present application also provides a multi-module interactive teaching method for an AI smart music classroom, which includes: in a music classroom, perceiving and collecting music expression evaluation indicators including students' playing postures, instrument touch and pressure signals, and environmental sound field data, and segmenting instrument tracks and background noise, generating performance error feature vectors and mapping them to three-dimensional feedback heat maps; uploading the three-dimensional feedback heat maps to the interactive interface corresponding to the teacher's port, and formulating adaptive teaching strategies based on the historical performance data marked under the student's port and the music expression evaluation indicators, and the interactive interface supports gesture operations / touch operations for area labeling; synchronizing the area labeling information to the AR visualization device of the student port in real time through the edge computing gateway, and establishing a closed-loop feedback mechanism according to the adaptive teaching strategy; at the same time, dynamically adjusting the key damping coefficient and the sustain pedal response sensitivity according to the quality of the music performance, and building a collaborative training environment for the physical layer and the digital layer; based on the closed-loop feedback mechanism, in the collaborative training environment, the adaptive teaching strategy is used for instant correction prompts.
[0006] One or more technical solutions provided in this application have at least the following beneficial effects:
[0007] The Perception Acquisition Module collects real-time data on student playing posture, instrument touch and pressure signals, and ambient sound field data to generate multi-dimensional musical performance assessment metrics. By separating audio tracks from noise, it achieves high-precision performance analysis, constructs performance error feature vectors, and generates a three-dimensional feedback heatmap, providing data support for subsequent teaching strategy development and visual feedback. The Strategy Formulation Module leverages the three-dimensional feedback heatmaps generated by the Perception Module and historical student performance data to automatically generate adaptive teaching strategies. It also supports teacher annotation through an interactive interface, enhancing the personalization and flexibility of strategies. The Information Synchronization Module, leveraging an edge computing gateway, transmits teacher annotation information and strategy content to the student's AR device with low latency, establishing a closed-loop feedback loop and enabling real-time strategy synchronization between the teacher and student. The Environment Construction Module extends the impact of teaching strategies to the physical performance equipment, dynamically adjusting the key damping coefficient and sustain pedal response based on the student's performance quality. This creates a human-computer collaborative training environment, deeply integrating physical training with digital teaching, and enhancing teaching immersion and responsiveness. In a collaborative training environment, the feedback prompt module provides students with timely error correction prompts or optimization suggestions based on teaching strategies, forming a fully closed-loop system of "perception-strategy-presentation-execution-feedback" to improve teaching efficiency and interactive effects.
[0008] In summary, this application has built an AI intelligent music interactive teaching system that integrates perception analysis, strategy formulation, real-time synchronization, environmental control and feedback prompts, and has opened up a complete teaching chain from multimodal performance behavior perception to personalized strategy execution. It not only achieves high-dimensional and accurate evaluation of students' musical expression, but also intuitively presents performance problems through three-dimensional heat maps, realizes dynamic teaching strategies collaboratively generated by teachers and AI, and uses AR devices and edge computing technology to build a low-latency, visual real-time interactive platform. Combined with the physical parameter control function of the performance equipment, it effectively integrates digital teaching and the real performance environment, and ultimately realizes a personalized, immersive, closed-loop music teaching experience, effectively improving the quality and efficiency of music teaching, and promoting intelligent music interactive teaching to develop in a more efficient, accurate and personalized direction.
[0009] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a structural diagram of the multi-module interactive teaching system of the AI smart music classroom provided in the embodiment of the present application.
[0011] Figure 2A flowchart of the multi-module interactive teaching method for the AI smart music classroom provided in an embodiment of the present application.
[0012] Description of the accompanying drawings: perception collection module 10, strategy formulation module 20, information synchronization module 30, environment construction module 40, feedback prompt module 50. DETAILED DESCRIPTION
[0013] The embodiments of the present application provide an AI smart music classroom multi-module interactive teaching system and method, which solves the technical problem in the existing technology that the teaching strategy is difficult to dynamically adapt to the individual differences of students due to the lack of an accurate evaluation mechanism and a low-latency feedback mechanism based on multimodal data. It achieves the technical effect of realizing personalized and immersive interactive teaching through somatosensory interactive perception, three-dimensional heat map visualization and AR closed-loop prompts, thereby improving the quality and efficiency of music teaching.
[0014] Example 1, as Figure 1 As shown, the embodiment of the present application provides an AI intelligent music classroom multi-module interactive teaching system, the system comprising:
[0015] The perception acquisition module 10 is used in the music classroom to perceive and acquire music expression evaluation indicators including students' playing postures, instrument touch and pressure signals, and environmental sound field data, and to segment instrument tracks and background noise, generate performance error feature vectors, and map them to three-dimensional feedback heat maps.
[0016] Specifically, playing posture refers to the dynamic position of a student's body, arms, fingers, and other parts while playing an instrument. This can be captured using posture recognition algorithms and IMU inertial sensors. Instrument touch pressure signals, which represent parameters such as the force and duration of a finger's pressure on the instrument, can be acquired using pressure sensors or capacitor arrays. Ambient sound field data, including classroom reverberation, background noise, and vocal interference, is collected using a microphone array.
[0017] The perception acquisition module 10 obtains students' playing postures, instrument touch and pressure signals, and environmental sound field data through interactive deployment of multiple types of sensor devices in the music classroom environment, including: IMU inertial sensors for detecting students' hand and fingertip movements, a touch and pressure sensor array installed on the keyboard surface, and a microphone array for collecting environmental sound field data. It also processes the audio data using spectrum analysis and noise suppression algorithms to separate pure instrument tracks. Subsequently, based on the pre-trained convolutional neural network model, error feature parameters are extracted during the performance to form a performance error feature vector containing multi-dimensional indicators such as pitch offset, rhythm drift, and force anomalies. The above vector is mapped to a three-dimensional feedback heat map, in which the three coordinate axes represent the time axis, pitch dimension, and error type dimension, respectively, to intuitively display the problem areas in students' performances.
[0018] This module realizes multi-dimensional perception of music performance, constructs a refined performance analysis model, provides a data basis for subsequent teaching strategies, and significantly improves the accuracy and comprehensiveness of teaching evaluation.
[0019] The strategy formulation module 20 is used to upload the three-dimensional feedback heat map to the interactive interface corresponding to the teacher port, formulate an adaptive teaching strategy based on the historical performance data marked under the student port and the music expression evaluation index. The interactive interface supports gesture operation / touch operation for area marking.
[0020] Specifically, historical performance data refers to students' performance data from previous lessons, such as error frequency and learning progress. Adaptive teaching strategies dynamically adjust teaching content, methods, and progress based on students' current performance and historical data to meet the personalized learning needs of different students. Regional annotation involves teachers manually selecting notes, rhythmic passages, and other areas that require attention or review on a heat map.
[0021] The strategy formulation module 20 receives the three-dimensional feedback heat map generated by the perception acquisition module 10 and uploads it to the interactive interface of the teacher's port. The interactive interface supports teachers to mark key areas in the heat map through gesture recognition or touch operations. After receiving the annotation information, the module combines the student's historical performance data (including but not limited to the frequency of error occurrence, practice completion time, progress assessment level, etc.) and music expression evaluation indicators, calls the built-in teaching strategy generation engine, and automatically formulates personalized adaptive teaching strategies. The teaching strategy may include repeated exercises of rhythm sections, special guidance on weak fingering, intensive training of playing style, etc., for teachers to further adjust and confirm.
[0022] This module uses a graphical strategy formulation method combined with individual student data to improve the personalization and targeting of teaching plans and enhance the synergy between teachers' subjective guidance and AI system's intelligent recommendations.
[0023] The information synchronization module 30 is used to synchronize the area annotation information to the AR visualization device of the student port in real time through the edge computing gateway, and establish a closed-loop feedback mechanism according to the adaptive teaching strategy.
[0024] Specifically, the information synchronization module 30 is designed based on the edge computing architecture, and synchronizes the teacher-side annotation information and teaching strategy content to the student-side AR visualization device in real time through the edge computing gateway deployed on the classroom side. After the synchronization is completed, when students wear AR devices to play, they can receive guidance information superimposed on layers in the real instrument environment, such as key highlighting, rhythm markings, error playback paths, etc., thereby realizing personalized immersive tutoring at the visual level. The module also supports recording students' performance feedback under prompts, and returns the data to the teacher side for strategy optimization, realizing a dynamic closed loop of the teaching process.
[0025] This module realizes real-time information interaction between the teacher and the student, so that the student can receive guidance and feedback from the teacher in a timely manner.
[0026] The environment building module 40 is used to dynamically adjust the key damping coefficient and the sustain pedal response sensitivity according to the quality of the music performance, and to build a collaborative training environment of the physical layer and the digital layer.
[0027] Specifically, the key damping coefficient is a parameter that affects the resistance to pressing and rebounding on instruments like pianos. Adjusting the key damping coefficient can alter the feel of the keys and the quality of the playing experience. The damper pedal response sensitivity controls the sensitivity of the instrument's sustain effect when the damper pedal is pressed, impacting the consistency and expressiveness of the performance.
[0028] The environment building module 40 is used to dynamically adjust the practice environment according to the quality of the student's music performance. The module changes the damping coefficient of the keys through the electronically controlled damping adjustment unit to provide students with different tactile feedback at different stages of practice. At the same time, the response sensitivity of the sustain pedal is controlled to train students' pedal coordination ability. The above-mentioned physical adjustment process is carried out synchronously with the digital teaching content, and the real-time rhythm and fingering trajectory are displayed in the AR overlay interface to guide students to establish corresponding cognition between physical performance and visual cues, realize a training environment that coordinates the physical and digital layers, effectively link physical control and virtual teaching, and realize "what you play is what you get" training perception feedback, which significantly improves students' practice immersion and technical mastery efficiency.
[0029] The feedback prompt module 50 is used to provide instant correction prompts based on the closed-loop feedback mechanism and the adaptive teaching strategy in the collaborative training environment.
[0030] Specifically, within the collaborative training environment, the feedback prompt module 50 monitors students' performance in real time based on the aforementioned closed-loop feedback mechanism. Upon detecting errors such as pitch errors, rhythmic drift, or unbalanced touch, immediate corrections are provided through visual prompts (e.g., note highlighting, dynamic arrow guidance) and auditory prompts (e.g., rhythmic sounds, correction tones). Furthermore, the feedback prompt module 50 features an adaptive adjustment function, automatically adjusting the intensity, frequency, and form of prompts based on the student's number of errors and feedback response speed, to better suit individual students with different learning rhythms and cognitive habits.
[0031] Furthermore, the strategy formulation module 20 also includes:
[0032] The audio analysis module is used to extract the harmonic distortion and rhythm offset of the performance audio through joint analysis in the time and frequency domains.
[0033] The mapping module is used to capture the angular velocity of the wrist joint through the IMU unit and establish a nonlinear mapping relationship between touch force and timbre response.
[0034] The spatiotemporal correlation analysis module is used to identify the focus of the music score and determine the spatiotemporal correlation between the distribution of visual attention and the characteristic vector of performance errors.
[0035] The strategy configuration module is used to configure the adaptive teaching strategy according to the harmonic distortion, rhythm offset, nonlinear mapping relationship between touch force and timbre response, and the spatiotemporal correlation between visual attention distribution and performance error feature vectors.
[0036] Specifically, to enhance the targetedness and individual adaptability of teaching strategies, the strategy formulation module 20 is further divided into the following submodules: an audio analysis module, a mapping module, a spatiotemporal correlation analysis module, and a strategy configuration module. These submodules work together to automatically generate personalized and adaptive teaching strategies.
[0037] The audio analysis module performs multi-level deconstruction of the student's performance audio based on a joint analysis method in the time and frequency domains. Specifically, algorithms such as STFT (Short-Time Fourier Transform) or CQT (Constant Q Transform) are used to extract harmonic structures from the audio. By comparing with the standard audio track, the harmonic distortion (overtone offset ratio) and rhythm offset (the deviation between the note trigger time and the beat reference) are calculated to quantify the performance stability and rhythm accuracy. For example, when it is detected that a student has continuous overtone spurs in the G-key performance, and the beat point always lags behind the standard beat by more than 20ms, it can be inferred that their pitch control and sense of rhythm are weak.
[0038] The mapping module collects the angular velocity data of students' wrists and finger joints through the IMU (inertial measurement unit), and pairs it with the touch pressure collected when typing and the generated timbre parameters (such as volume and brightness) for analysis. Based on these data, a nonlinear mapping relationship model between force and timbre is constructed to evaluate the finger control accuracy and consistency of expression. For example, when a student produces an uncontrolled sharp timbre when typing with medium to strong force, the module can mark the deviation of the "force-timbre" mapping, suggesting that the teaching strategy should strengthen control stability training.
[0039] The spatiotemporal correlation analysis module, combined with a visual tracking device (such as an eye tracker), captures the student's gaze focus during performance. By analyzing the spatial location and duration of the gaze points, a visual attention distribution map is constructed. This map is then spatiotemporally aligned with the performance error feature vector to identify whether distraction is a key factor in the error. If the student's gaze area frequently shifts before an error occurs, it can be determined that "inadequate visual attention" is highly correlated with the error, leading to an adjustment to the "visual guidance and paragraph decomposition" strategy.
[0040] The strategy configuration module integrates the data extracted from the three submodules above, performs matching operations based on the built-in teaching strategy knowledge base, and configures the final adaptive teaching strategy. This strategy may include adjustments in dimensions such as content rearrangement, skill reinforcement, and error type focus. For example, if a student exhibits a combination of characteristics such as "unstable rhythm, inconsistent force control, and easily distracted attention," the strategy configuration module may formulate a combination of intervention measures such as "rhythm and beat synchronization training, combined with dynamic touch feedback and visual guidance highlights," and push them to the teacher and student interfaces for execution.
[0041] Through the collaborative processing of the four sub-modules of the above-mentioned strategy formulation module 20, not only can fine-grained performance features be extracted from multimodal data, but also a causal logic chain between high-dimensional behaviors and performances can be established, and finally a highly personalized teaching strategy can be automatically generated, which greatly improves the efficiency and adaptability of music teaching.
[0042] Furthermore, a distributed pressure sensor array is embedded in the fingerboard of the musical instrument to quantify the discrete gradient of finger pressure. The system also includes:
[0043] The priority scoring module is used to align the standard score MIDI signal with the actual performance timing deviation, determine the pitch deviation and rhythm jitter coefficient, and generate a multi-dimensional error correction priority score in combination with the discrete gradient of the finger pressure intensity.
[0044] Specifically, in order to further refine the accuracy of error recognition and the scientific nature of intervention decisions, this system embeds a distributed pressure sensor array on the instrument's fingerboard to achieve high-resolution dynamic perception of the student's finger pressure, and combines standard performance benchmarks to build a priority scoring module to perform multi-dimensional error correction strategy sorting.
[0045] A distributed pressure sensor array consists of multiple micro-thin-film pressure sensors evenly distributed across the keyboard or fingerboard surface. Each sensor node independently captures the touch pressure value at a specific point in time and measures the pressure gradient with millimeter-level resolution, creating a three-dimensional force distribution map.
[0046] The priority scoring module first aligns and analyzes the MIDI data generated by the actual performance with the standard MIDI score, and extracts two key deviation parameters: pitch deviation and rhythm jitter coefficient. Among them, pitch deviation is the difference between the actual performance pitch and the pitch specified in the standard musical score. It is determined by judging whether there are wrong notes or glissando, and monitoring the number of semitones that deviate from the standard pitch, reflecting the pitch problem of the performance. The rhythm jitter coefficient is an indicator to measure the stability of the actual performance rhythm. It is determined based on the standard deviation between the performance timing point and the standard beat point, reflecting the rhythm stability. Combining the above two deviation parameters, and introducing the discrete gradient indicator of finger pressure, a three-dimensional error correction priority scoring model is constructed. This scoring model assigns an importance score to each error to guide teachers or systems to prioritize feedback or training strategy configuration for high-weight issues.
[0047] By embedding a distributed pressure sensor array in the fingerboard of an instrument, a high-precision force perception model can be constructed in real time. Combined with standard score alignment and performance deviation analysis, a multi-dimensional error correction priority score can be generated. This not only significantly improves the accuracy of error assessment, but also provides a clear and quantitative priority basis for subsequent teaching interventions, effectively optimizing feedback prompts and training path configuration strategies.
[0048] Furthermore, the perception collection module 10 is further configured to:
[0049] Step P11: Deploy edge computing nodes to perform real-time spectrum analysis and separate the fundamental frequency harmonic components of the instrument.
[0050] Step P12: Generate a noise reduction mask that matches the current ambient sound field through gated loop analysis.
[0051] Step P13: According to the fundamental frequency harmonic components of the instrument and the noise reduction mask, the directional sound source of the target instrument is enhanced, and a signal-to-noise ratio threshold is configured. The signal-to-noise ratio threshold is used to trigger the parameter update of the adaptive filter.
[0052] Specifically, the edge computing node refers to a lightweight computing device (such as an embedded GPU device, an industrial edge gateway, etc.) deployed locally at the teaching site, which can independently complete computing-intensive tasks. The perception acquisition module 10 performs spectral analysis on the real-time audio signal through this node, extracts the fundamental frequency and harmonic components of the instrument audio, and is used to identify the energy characteristics of the instrument in a complex sound field. Among them, the sound emitted by the instrument consists of a fundamental frequency and multiple harmonics. The fundamental frequency is the lowest frequency component of the sound, which determines the pitch; the harmonics are integer multiples of the fundamental frequency and determine the timbre.
[0053] Building on spectrum extraction, gated recurrent analysis is introduced. This method uses a gated neural network structure (such as a GRU) to model background noise patterns in a time series, generating a noise reduction mask that matches the current sound field state. The mask value represents the probability that each frequency component in the signal belongs to the target signal, and is used to filter out frequency components in the spectrogram that are unrelated to the target sound source.
[0054] The extracted fundamental harmonic components are combined with the noise reduction mask, and beamforming or sound source enhancement techniques are used to enhance the clarity of the target instrument's direction. The signal-to-noise ratio (SNR) is then calculated based on the enhanced signal. A pre-set SNR threshold is set in the module. When the calculated SNR falls below the threshold, an adaptive filter (such as a Wiener filter or a Kalman filter) is triggered to update its parameters, achieving dynamic noise suppression.
[0055] The aforementioned steps are performed locally and in real time by edge computing nodes, including spectrum extraction, harmonic modeling, and noise analysis. This significantly improves the recognizability of target instruments in complex sound fields. A gated loop structure is combined to achieve highly robust adaptive noise reduction, and filter updates are dynamically driven by the signal-to-noise ratio. This effectively enhances the system's discernment and perception of performance, providing a reliable audio foundation for subsequent teaching strategy development and feedback mechanism construction. Overall, this not only improves data quality but also reduces the system's dependence on central servers.
[0056] Furthermore, the environment building module 40 is also used to:
[0057] Step P41: Integrating a magnetorheological fluid damper, wherein the magnetorheological fluid damper is used to adjust the magnetic field strength according to the playing force error rate.
[0058] Step P42: Use a strain sensor to monitor the string striking velocity and establish a dynamic correlation sequence between the key rebound force and the note duration.
[0059] Step P43: Based on the dynamic association sequence, when continuous erroneous key touches are detected, activating the progressive damping enhancement feedback mode of the magnetorheological fluid damper.
[0060] Specifically, the environment building module 40 constructs a collaborative training environment that couples the physical layer and the perception layer through a magnetorheological fluid damper and a string-strike strain monitoring mechanism, thereby achieving an organic fusion of realistic tactile response and dynamic correction prompts. The magnetorheological fluid damper is an intelligent structural component based on magnetorheological fluid. When an external magnetic field is applied, the viscosity of the liquid inside it changes rapidly, thereby affecting its mechanical damping effect. The system embeds the device in a digital piano or keyboard device, and uses the performance force error rate in the teaching record (such as the mean square error between the expected keystroke force and the actual detection value) as an adjustment signal to dynamically control the magnetic field strength generated by the electromagnetic coil, thereby achieving an adaptive linkage between performance difficulty and feedback intensity.
[0061] Strain sensors are devices used to monitor minute physical deformations (such as tension or compression). Installed at the base of a piano key, they measure the velocity of the string and determine the strength and speed of each keystroke. Furthermore, the module dynamically correlates the monitored values with the note duration, quantifying the causal relationship between action and output during performance.
[0062] When the module identifies consecutive incorrect keystroke patterns (e.g., jerky rhythms, insufficient keystroke amplitude, etc.) through the rebound-duration sequence, it triggers the magnetorheological fluid damper to enter a progressive damping enhancement feedback mode, which increases resistance feedback based on the frequency of the error. For example, if a student makes the error of premature finger lift in four rhythmic passages, the module will significantly increase the damping in the fifth input, forcing the student to extend the keystroke duration, thus achieving physical intervention.
[0063] Furthermore, the system further comprises:
[0064] The second strategy module is used to determine a visual-tactile collaborative control strategy based on the physical resistance changes between the AR visualization device and the magnetorheological fluid damper; the visual-tactile collaborative control strategy is used to adjust the magnetic field gradient parameters of the magnetorheological fluid damper according to the spatiotemporal distribution characteristics of continuous erroneous key touches, so that the key damping coefficient increases exponentially with the error frequency. At the same time, the attention guidance plug-in embedded in the AR visualization device uses pulsed highlighting and three-dimensional arrows to dynamically point to the correct fingering area.
[0065] Specifically, in order to further improve the accuracy of intervention in erroneous behaviors and the sensory immersion experience, this system introduces a second strategy module based on the original teaching feedback. By integrating AR visualization equipment and magnetorheological fluid dampers, it realizes the dynamic adaptation and execution of visual-tactile collaborative control strategies, effectively dealing with students' high-frequency erroneous touch problems in specific areas and specific finger positions.
[0066] The second strategy module determines a control strategy for linking physical feedback and visual guidance based on the performance data from the perception acquisition module 10, particularly the spatiotemporal distribution characteristics of incorrect keystrokes. This strategy integrates two core mechanisms. The first is the tactile control dimension. Based on the error frequency detected within a specific key area, the magnetic field gradient parameters of the magnetorheological fluid damper are adjusted to exponentially increase the damping coefficient of the corresponding key position, guiding the performer to reduce keystroke speed or increase keystroke attention. The second is the visual guidance dimension. This dimension calls on the attention guidance plug-in embedded in the AR visualization device to indicate the correct playing area or finger placement through pulsed highlighting and three-dimensional arrow dynamic pointing functions.
[0067] Furthermore, the feedback prompt module 50 further includes:
[0068] The trajectory superposition module is used to superimpose fingering trajectory prediction lines on the music score projection interface, and the music score projection interface is used to compensate for visual display delay.
[0069] The beat guidance module is used to generate a phase-adjustable light pulse sequence on the music score projection interface, and the light pulse sequence is used to guide beat synchronization.
[0070] Specifically, the trajectory overlay module compares historical performance data with the target score's MIDI information, combines it with the performer's current finger positions and movement trends, and constructs a predicted fingering trajectory line based on Bezier curves or Kalman filters. This trajectory information is then overlaid on the score projection interface in real time. This predicted line features a slight transparency gradient and color change prompts to guide the student's next finger placement direction and path.
[0071] The beat guidance module generates a phase-adjustable light pulse sequence within the same music score projection interface. This light pulse sequence is based on the metronome rhythm and dynamically adjusts the pulse emission rhythm through time domain interpolation, allowing students to play the beat synchronously and perceive the accent position.
[0072] The feedback module 50, through the coordinated operation of the two modules mentioned above, combines fingering path prediction with beat guidance to simultaneously optimize students' performance control ability from both temporal and spatial dimensions. Trajectory superposition significantly improves movement foresight and reduces unconscious repetition errors. Beat guidance clearly indicates the rhythmic structure of the performance through dynamic changes, helping to develop a good sense of timing and fluency. Ultimately, this achieves coordinated error-correction training across the entire chain of rhythm, movement, and perception.
[0073] Furthermore, the beat guidance module is also used to:
[0074] Step P521: Based on the student's playing speed deviation rate, dynamically adjust the phase offset of the optical pulse sequence through time domain interpolation.
[0075] Step P522: Generate a tactile prompt waveform synchronized with the light pulse sequence in the fingertip area of the tactile feedback glove, wherein the tactile prompt waveform includes a pressure gradient pulse corresponding to the beat accent and a micro-vibration code of a sixteenth note interval.
[0076] Step P523: Based on the optical pulse sequence and the tactile prompt waveform, the propagation delay of the audio devices at the teacher port and the student port is compensated by the sound field phase controller.
[0077] Specifically, a time domain phase adjustment mechanism, a tactile prompt waveform generation mechanism, and a sound field delay compensation mechanism are introduced into the beat guidance module to build a more immersive and precise rhythm guidance feedback system, so as to further improve students' beat control accuracy and multimodal feedback consistency in complex music.
[0078] The beat guidance module first adjusts the phase shift of the currently generated light pulse sequence based on the real-time deviation rate of the student's playing speed through time-domain interpolation methods (such as spline interpolation, Lagrange interpolation, or dynamic time warping). If the student's tempo is detected to be slow, the light pulse is automatically shifted forward by a certain time window; if the tempo is fast, the pulse is delayed, achieving an adjustable guide beat line that dynamically matches the performance behavior.
[0079] To further enhance students' perception of rhythm, haptic feedback gloves (e.g., equipped with vibration motors and piezoelectric actuators) generate tactile cue waveforms at the fingertips that are strictly synchronized with light pulses. This waveform is embedded with pressure gradient pulses at the beat's accent positions to simulate the sensation of tapping, as well as low-amplitude micro-vibrations encoding sixteenth-note intervals to provide detailed rhythmic density cues.
[0080] In order to ensure strict synchronization of light prompts, tactile feedback and sound broadcasting between the teacher port and the student port, a sound field phase controller is introduced. Based on the current network delay parameters and device response time, it accurately compensates for the sound wave propagation phase difference of the audio devices at both ends, thereby achieving three-dimensional feedback alignment of hearing, vision and touch.
[0081] Furthermore, the system further comprises:
[0082] The model definition module is used to build a Markov decision process model of skill mastery and define note coherence and expressive richness as state transition rewards.
[0083] The task optimization module is used to select decomposition training tasks based on the Markov decision process model and adopt a dual deep Q network strategy to prioritize optimizing the direction of the maximum cumulative reward in the current performance stage.
[0084] The repertoire synthesis module is used to simultaneously synthesize progressively more difficult practice repertoires through an adversarial generative network and dynamically synchronize them with the student's muscle fatigue data.
[0085] Specifically, the model definition module is used to build a Markov decision process model with the goal of skill mastery. The modeling method is as follows: first, define the state space and action space. Among them, the state space is composed of note continuity (such as continuous pitch hit rate) and expressive richness (such as dynamic changes and emotional labeling accuracy), while the action space includes the decomposition selection of performance tasks, such as selected bar paragraphs, variation processing methods or rhythm reconstruction strategies. Then, the reward function is defined in combination with the performance improvement in the state transfer process to quantify the contribution of each performance task to skill growth. The Markov decision process model supports real-time updates to reflect the changing trend of students' current skill levels.
[0086] The task optimization module is based on the constructed Markov decision process model, which decomposes the performance training task into multiple subtasks, each of which corresponds to different performance skill points or music fragments. The dual-depth Q network algorithm is used to learn and optimize each subtask. The dual-depth Q network achieves the optimal scheduling selection for the performance training task through the collaborative work of two neural networks (evaluation network and target network). The current state of the student is input into the Q network, and the evaluation network outputs the corresponding Q value based on the current state and the executed actions; the target network updates the target Q value based on the output and reward signal of the evaluation network. The evaluation network is trained through the back-propagation algorithm, and the network parameters are continuously optimized to select the subtask that maximizes the long-term cumulative reward, and the output is the optimal subtask for the current stage.
[0087] In addition, to avoid the problem that traditional static teaching materials are difficult to adapt to students' current abilities and status, the repertoire synthesis module introduces a generative adversarial network to automatically synthesize practice repertoires with progressive difficulty. The generator of the generative adversarial network generates variation practice repertoires with new rhythmic patterns or harmonic structures based on a large amount of music repertoire data and predefined conditions such as music style and type, according to the difficulty level of the current performance task and the data of the skills already mastered; at the same time, the discriminator judges the rationality and effectiveness of the generated repertoire through the teacher's standard music score and the student's performance characteristics, ensuring that the generated repertoire meets the music norms and practice requirements and the difficulty is adapted to the student's current level. At the same time, a student muscle fatigue monitoring mechanism (such as finger pressure fluctuations and IMU frequency intensity analysis) is introduced. During the student's performance, muscle fatigue data is collected in real time, and the degree of muscle fatigue is analyzed using signal processing and machine learning algorithms. The analysis results are fed back to the generative adversarial network to dynamically adjust the difficulty parameters of the generated repertoire. Dynamic alignment of the synthesized repertoire with the current fatigue level is achieved to avoid muscle overload leading to decreased learning efficiency.
[0088] Through the above modules, Markov decision modeling, deep reinforcement learning training scheduling mechanism and dynamic generation mechanism of practice repertoire based on adversarial generative network are introduced into the teaching system to achieve dynamic closed-loop matching between teaching strategies and students' learning status.
[0089] In summary, the AI intelligent music classroom multi-module interactive teaching system provided by the embodiment of the present application has the following beneficial effects:
[0090] The embodiment of the present application builds an AI intelligent music interactive teaching system that integrates perception analysis, strategy formulation, real-time synchronization, environmental control and feedback prompts, and opens up a complete teaching chain from multimodal performance behavior perception to personalized strategy execution. It not only achieves high-dimensional and accurate evaluation of students' musical expression, but also intuitively presents performance problems through three-dimensional heat maps, realizes dynamic teaching strategies collaboratively generated by teachers and AI, and builds a low-latency, visual real-time interactive platform with the help of AR devices and edge computing technology. Combined with the physical parameter control function of the performance equipment, it effectively integrates digital teaching with the real performance environment, and ultimately realizes a personalized, immersive, closed-loop music teaching experience, effectively improving the quality and efficiency of music teaching, and promoting the development of intelligent music interactive teaching in a more efficient, accurate and personalized direction.
[0091] Example 2, as Figure 2 As shown, based on the same inventive concept as the aforementioned embodiment 1, the embodiment of the present application provides a multi-module interactive teaching method for an AI smart music classroom, the method comprising:
[0092] Step S1: In the music classroom, perception is performed to collect music performance evaluation indicators including student playing posture, instrument touch and pressure signals, and ambient sound field data, and the instrument track and background noise are separated to generate a performance error feature vector and map it to a three-dimensional feedback heat map.
[0093] Step S2: Upload the three-dimensional feedback heat map to the interactive interface corresponding to the teacher port, formulate an adaptive teaching strategy based on the historical performance data marked under the student port and the music expression evaluation index, and the interactive interface supports gesture operation / touch operation for area labeling.
[0094] Step S3: Synchronize the area annotation information to the AR visualization device of the student port in real time through the edge computing gateway, and establish a closed-loop feedback mechanism according to the adaptive teaching strategy.
[0095] Step S4: At the same time, according to the quality of the music performance, the key damping coefficient and the sustain pedal response sensitivity are dynamically adjusted to build a collaborative training environment for the physical layer and the digital layer.
[0096] Step S5: Based on the closed-loop feedback mechanism, in the collaborative training environment, immediate correction prompts are provided using the adaptive teaching strategy.
[0097] Furthermore, according to the historical performance data under the student port mark and in combination with the music expression evaluation index, an adaptive teaching strategy is formulated, which also includes:
[0098] Through joint analysis in the time and frequency domains, the harmonic distortion and rhythm offset of the performance audio are extracted; the angular velocity of the wrist joint is captured by the IMU unit, and a nonlinear mapping relationship between the touch force and the timbre response is established; the focus of the musical score is identified, and the spatiotemporal correlation between the visual attention distribution and the performance error feature vector is determined; and the adaptive teaching strategy is configured based on the harmonic distortion, rhythm offset, the nonlinear mapping relationship between the touch force and the timbre response, and the spatiotemporal correlation between the visual attention distribution and the performance error feature vector.
[0099] Furthermore, a distributed pressure sensor array is embedded in the instrument's fingerboard to quantify the discrete gradient of finger pressure; the standard music score MIDI signal is aligned with the actual performance timing deviation to determine the pitch deviation and rhythm jitter coefficient, and combined with the discrete gradient of finger pressure, a multi-dimensional error correction priority score is generated.
[0100] Furthermore, the instrument track and background noise are segmented, and a performance error feature vector is generated and mapped to a three-dimensional feedback heat map, which also includes:
[0101] Deploy edge computing nodes to perform real-time spectrum analysis and separate the fundamental frequency harmonic components of musical instruments. Through gated loop analysis, a noise reduction mask that matches the current ambient sound field is generated. Based on the fundamental frequency harmonic components of the musical instrument and the noise reduction mask, the sound source in the direction of the target instrument is enhanced, and a signal-to-noise ratio threshold is configured. The signal-to-noise ratio threshold is used to trigger the parameter update of the adaptive filter.
[0102] Furthermore, the key damping coefficient and sustain pedal response sensitivity are dynamically adjusted according to the quality of the musical performance, building a collaborative training environment between the physical and digital layers. This also includes:
[0103] An integrated magnetorheological fluid damper is used to adjust the magnetic field strength according to the error rate of playing force; a strain sensor is used to monitor the string striking speed and establish a dynamic correlation sequence between the key rebound force and the note duration; based on the dynamic correlation sequence, the magnetorheological fluid damper's progressive damping enhancement feedback mode is activated when continuous erroneous key touches are detected.
[0104] Furthermore, based on the changes in physical resistance between the AR visualization device and the magnetorheological fluid damper, a visual-tactile collaborative control strategy is determined; the visual-tactile collaborative control strategy is used to adjust the magnetic field gradient parameters of the magnetorheological fluid damper according to the spatiotemporal distribution characteristics of continuous erroneous key touches, so that the key damping coefficient increases exponentially with the error frequency, and at the same time, the attention guidance plug-in embedded in the AR visualization device uses pulsed highlighting and three-dimensional arrows to dynamically point to the correct fingering area.
[0105] Furthermore, in the collaborative training environment, providing instant correction prompts using the adaptive teaching strategy also includes:
[0106] Fingering trajectory prediction lines are superimposed on a music score projection interface, which is used to compensate for visual display delay; and a phase-adjustable light pulse sequence is generated on the music score projection interface, which is used to guide beat synchronization.
[0107] Furthermore, the optical pulse sequence is used to guide beat synchronization, and further includes:
[0108] Based on the student's playing speed deviation rate, the phase offset of the light pulse sequence is dynamically adjusted through time domain interpolation; a tactile prompt waveform synchronized with the light pulse sequence is generated in the fingertip area of the tactile feedback glove, wherein the tactile prompt waveform includes pressure gradient pulses corresponding to beat accents and micro-vibration codes of sixteenth note intervals; based on the light pulse sequence and the tactile prompt waveform, the propagation delay of the audio devices at the teacher port and the student port is compensated by a sound field phase controller.
[0109] Furthermore, the method described in the embodiment of the present application also includes:
[0110] A Markov decision process model of skill mastery is constructed, and note coherence and expressive richness are defined as state transition rewards. Based on the Markov decision process model, a dual-depth Q-network strategy is adopted to select decomposed training tasks, prioritizing the optimization of the direction of maximum cumulative reward in the current performance stage. At the same time, a generative adversarial network is used to synthesize progressively more difficult practice repertoires, which are dynamically synchronized with the student's muscle fatigue data.
[0111] Through the above detailed description of the AI smart music classroom multi-module interactive teaching system in this specification, those skilled in the art can clearly understand the AI smart music classroom multi-module interactive teaching method in this embodiment. As for the method disclosed in Example 2, since it corresponds to the system disclosed in Example 1 and has corresponding execution steps and beneficial effects, the relevant parts can be referred to the system part description.
[0112] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. AI intelligent music classroom multi-module interactive teaching system, characterized by: include: The sensory acquisition module is used in music classrooms to sense and collect musical performance evaluation indicators including student playing posture, instrument touch and pressure signals, and ambient sound field data. It also separates instrument tracks from background noise, generates performance error feature vectors, and maps them to a three-dimensional feedback heat map. A strategy formulation module is used to upload the three-dimensional feedback heat map to the interactive interface corresponding to the teacher's port, formulate an adaptive teaching strategy based on the historical performance data marked by the student port and the music expression evaluation index, and the interactive interface supports gesture operation / touch operation for area labeling; An information synchronization module is used to synchronize the regional annotation information to the AR visualization device of the student port in real time through the edge computing gateway, and establish a closed-loop feedback mechanism according to the adaptive teaching strategy; The environment building module is used to dynamically adjust the key damping coefficient and sustain pedal response sensitivity according to the quality of the music performance, and to build a collaborative training environment for the physical and digital layers; A feedback prompt module is used to provide instant correction prompts using the adaptive teaching strategy in the collaborative training environment based on the closed-loop feedback mechanism.
2. The AI intelligent music classroom multi-module interactive teaching system according to claim 1, characterized in that: The strategy formulation module also includes: An audio analysis module is used to extract the harmonic distortion and rhythm offset of the performance audio through joint analysis in the time and frequency domains; A mapping module is used to capture the wrist joint angular velocity through the IMU unit and establish a nonlinear mapping relationship between touch force and timbre response; A spatiotemporal correlation analysis module is used to identify the focus of attention on the score and determine the spatiotemporal correlation between the distribution of visual attention and the characteristic vector of performance errors; The strategy configuration module is used to configure the adaptive teaching strategy according to the harmonic distortion, rhythm offset, nonlinear mapping relationship between touch force and timbre response, and the spatiotemporal correlation between visual attention distribution and performance error feature vectors.
3. The AI intelligent music classroom multi-module interactive teaching system according to claim 2, characterized in that: A distributed pressure sensor array is embedded in the fingerboard of the musical instrument to quantify the discrete gradient of finger pressure. The system also includes: The priority scoring module is used to align the standard score MIDI signal with the actual performance timing deviation, determine the pitch deviation and rhythm jitter coefficient, and generate a multi-dimensional error correction priority score in combination with the discrete gradient of the finger pressure intensity.
4. The AI intelligent music classroom multi-module interactive teaching system according to claim 1, characterized in that: The perception collection module is also used for: Deploy edge computing nodes to perform real-time spectrum analysis and separate the fundamental frequency harmonic components of musical instruments; Generate a noise reduction mask that matches the current ambient sound field through gated loop analysis; According to the fundamental frequency harmonic components of the instrument and the noise reduction mask, the sound source in the direction of the target instrument is enhanced, and a signal-to-noise ratio threshold is configured. The signal-to-noise ratio threshold is used to trigger parameter update of the adaptive filter.
5. The AI intelligent music classroom multi-module interactive teaching system according to claim 1, characterized in that: The environment building module is also used to: An integrated magnetorheological fluid damper, which is used to adjust the magnetic field strength according to the playing force error rate; Use strain sensors to monitor string striking velocity and establish a dynamic correlation sequence between key rebound force and note duration; Based on the dynamic association sequence, when continuous erroneous key touches are detected, a progressive damping enhancement feedback mode of the magnetorheological fluid damper is activated.
6. The AI intelligent music classroom multi-module interactive teaching system according to claim 5, characterized in that: The system further comprises: a second strategy module, configured to determine a visual-tactile collaborative control strategy based on a change in physical resistance of the AR visualization device and the magnetorheological fluid damper; The visual-tactile collaborative control strategy is used to adjust the magnetic field gradient parameters of the magnetorheological fluid damper according to the spatiotemporal distribution characteristics of continuous incorrect key touches, so that the key damping coefficient increases exponentially with the error frequency. At the same time, the attention guidance plug-in embedded in the AR visualization device uses pulsed highlighting and three-dimensional arrows to dynamically point to the correct fingering area.
7. The AI intelligent music classroom multi-module interactive teaching system according to claim 6, characterized in that: The feedback prompt module also includes: a trajectory superposition module for superimposing fingering trajectory prediction lines on a music score projection interface, wherein the music score projection interface is used to compensate for visual display delay; The beat guidance module is used to generate a phase-adjustable light pulse sequence on the music score projection interface, and the light pulse sequence is used to guide beat synchronization.
8. The AI intelligent music classroom multi-module interactive teaching system according to claim 7, characterized in that: The beat guidance module is also used for: Dynamically adjusting the phase offset of the optical pulse sequence by time domain interpolation based on the student's playing speed deviation rate; Generating a tactile cue waveform synchronized with the light pulse sequence in the fingertip area of the tactile feedback glove, wherein the tactile cue waveform includes a pressure gradient pulse corresponding to the beat accent and a micro-vibration encoding of a sixteenth note interval; According to the optical pulse sequence and the tactile prompt waveform, the propagation delay of the audio devices at the teacher port and the student port is compensated by the sound field phase controller.
9. The AI intelligent music classroom multi-module interactive teaching system according to claim 8, characterized in that: The system further comprises: The model definition module is used to build a Markov decision process model of skill mastery and define note coherence and expressiveness as state transition rewards; A task optimization module, configured to select decomposed training tasks based on the Markov decision process model using a dual deep Q-network strategy, and prioritize optimizing the direction of maximum cumulative reward in the current performance phase; The repertoire synthesis module is used to simultaneously synthesize progressively more difficult practice repertoires through an adversarial generative network and dynamically synchronize them with the student's muscle fatigue data.
10. The multi-module interactive teaching method of AI intelligent music classroom is characterized by: The method is performed by the AI intelligent music classroom multi-module interactive teaching system according to any one of claims 1 to 9, comprising: In music classrooms, perceptual data is collected to assess musical performance, including student playing posture, instrument touch and pressure signals, and ambient sound field data. The instrument track is separated from background noise, and performance error feature vectors are generated and mapped to a three-dimensional feedback heat map. Uploading the three-dimensional feedback heat map to the interactive interface corresponding to the teacher's port, formulating an adaptive teaching strategy based on the historical performance data marked by the student's port and the music expression evaluation index, the interactive interface supports gesture operation / touch operation for area labeling; The area annotation information is synchronized in real time to the AR visualization device of the student port through the edge computing gateway, and a closed-loop feedback mechanism is established according to the adaptive teaching strategy; At the same time, according to the quality of the music performance, the key damping coefficient and sustain pedal response sensitivity are dynamically adjusted to build a collaborative training environment of physical and digital layers; Based on the closed-loop feedback mechanism, in the collaborative training environment, the adaptive teaching strategy is used to provide immediate correction prompts.