Intelligent music learning method
Through multimodal sensors, a personalized deviation feature library is constructed by combining hidden Markov model and transfer learning to generate targeted practice strategies and multimodal feedback signals, which solves the problems of low deviation recognition accuracy and single feedback in the existing music intelligent learning system, realizes personalized practice guidance and multi-dimensional feedback, and improves learning efficiency and performance level.
Patent Information
- Application Number
- CN202510422838.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing music intelligent learning systems lack the accuracy of performance deviation recognition, especially the ability to capture subtle deviations such as pitch and rhythm is limited, lack personalized deviation correction strategies, and the feedback signal is single, making it difficult to provide a multi-dimensional immersive learning experience.
Multimodal sensors are used to collect performance data, generate fidelity audio signals and tactile pressure timing data through anti-interference processing, and perform multi-dimensional deviation feature extraction. A personalized deviation feature library is constructed by combining hidden Markov models and transfer learning to generate targeted practice strategies and provide multimodal deviation correction feedback signals.
It realizes accurate identification and personalized correction of learners' pitch, rhythm and tone deviation, provides dynamic adjustment practice guidance and multi-dimensional feedback, and improves learning efficiency and performance level.
Smart Images

Figure CN120337139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent education, and in particular to a music intelligent learning method. Background Art
[0002] In the field of music education, the innovation of intelligent learning methods has always been committed to improving learning efficiency and personalized experience through technical means. With the integration of artificial intelligence and sensor technology, music intelligent learning systems have gradually become an important tool to assist instrument training. This type of method usually collects learners' performance data and generates feedback through algorithm analysis to help learners identify performance deviations and optimize practice paths. Its core goal is to reduce the dependence on professional teachers in traditional teaching through automated and data-driven methods, while providing accurate practice guidance for learners of different levels.
[0003] However, existing technologies still face significant bottlenecks in practical applications. First, the recognition accuracy of performance deviations is insufficient, especially the ability to capture subtle deviations such as pitch and rhythm is limited. Traditional audio analysis algorithms often have difficulty distinguishing between environmental noise and the performer's own technical errors, and are prone to misjudgment in complex performance scenarios (such as fast arpeggios and ornamentation processing). Secondly, existing systems lack the ability to deeply mine learners' long-term deviation patterns and are unable to build a personalized deviation feature library, resulting in serious homogeneity in practice suggestions and difficulty in providing dynamic adjustment strategies for the technical shortcomings of different learners. In addition, the singleness of the feedback signal (such as relying only on auditory or visual cues) further weakens the correction effect, making it difficult to form a multi-dimensional immersive learning experience. These problems jointly restrict the practicality and popularity of music intelligent learning systems, and have become technical problems that need to be solved urgently. Summary of the invention
[0004] Based on this, the purpose of the present invention is to provide a music intelligent learning method that can accurately identify multi-dimensional performance deviations and dynamically generate personalized correction strategies.
[0005] The purpose of the present invention is achieved by the following scheme:
[0006] In a first aspect, the present application provides a method for intelligent music learning, comprising the following steps:
[0007] S1: Collect the learner's performance data on the musical instrument based on the multimodal sensor, perform anti-interference processing on the performance data, and generate music data to be tested; the music data to be tested includes a fidelity audio signal and tactile pressure time series data;
[0008] S2: Based on the fidelity audio signal and tactile pressure time series data in the music data to be tested, multi-dimensional deviation feature extraction processing is performed to generate a deviation analysis parameter set, the deviation analysis parameter set includes a fundamental frequency deviation quantization parameter, a rhythm deviation matrix and a timbre spectrum feature;
[0009] S3: Based on the deviation analysis parameter set, construct a personalized deviation feature library through the Hidden Markov Model and transfer learning;
[0010] S4: Generate a targeted practice strategy according to the personalized deviation feature library, and perform real-time analysis on the learner's performance data based on the targeted practice strategy to generate a multi-modal correction feedback signal; the targeted practice strategy is used to dynamically adjust the speed, interval span, and repetition times of the practice repertoire; the multi-modal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
[0011] In one embodiment, S1 includes:
[0012] S11: Collect the original audio signal based on the microphone of the multi-modal sensor, and obtain the tactile pressure data based on the pressure sensor in the multi-modal sensor;
[0013] S12: Use the following formula to perform comb filter noise reduction processing on the original audio signal to generate a fidelity audio signal:
[0014]
[0015] where S'(f) is the spectrum after noise reduction, that is, the fidelity audio signal, S(f) is the original audio spectrum, N(f / k) is the energy of the environmental noise at the k-th harmonic, and α is the noise suppression factor;
[0016] S13: Perform Kalman filter smoothing processing on the tactile pressure data to generate standardized tactile pressure time series data;
[0017] S14: Organize the fidelity audio signal and the tactile pressure time series data to generate the music data to be inspected.
[0018] In one embodiment, S2 includes:
[0019] S21: Based on the wavelet packet transform technology, perform separation processing on the fidelity audio signal to obtain an isolated audio signal that separates the fundamental frequency component and the overtone frequency band;
[0020] S22: Use the following formula to correct the fundamental frequency offset of the isolated audio signal based on the tactile pressure time series data to generate a fundamental frequency deviation quantization parameter:
[0021] Δf(t) = |f user (t) - f std |
[0022] where Δf(t) is the fundamental frequency deviation quantization parameter, f user (t) is the real-time fundamental frequency of the isolated audio signal, f stdis the theoretical fundamental frequency value of the corresponding note in the standard musical score;
[0023] S23: Process the performance timing signal of the tactile pressure timing data in the music data to be inspected based on the dynamic time warping algorithm and the standard musical score time axis alignment technology, and generate a rhythm deviation matrix;
[0024] S24: Use the following formula to perform Mel spectrum conversion processing on the wide audio band of the isolated audio signal to generate timbre spectrum features:
[0025]
[0026] where D timbre is the timbre spectrum feature, and P std (m) is the energy distribution of the standard timbre in the m-th Mel band, and P user (m) is the corresponding energy distribution of the learner's performance timbre;
[0027] S25: Organize the fundamental frequency deviation quantization parameter, the rhythm deviation matrix, and the timbre spectrum features to generate a deviation analysis parameter set.
[0028] In one embodiment, S23 includes:
[0029] S231: Perform frame segmentation processing on the performance timing signal of the tactile pressure timing data based on the time series segmentation algorithm to generate a discrete timestamp sequence;
[0030] S232: Calculate the optimal alignment path between the discrete timestamp sequence and the standard musical score time axis through the dynamic time warping algorithm to generate a time deviation mapping table;
[0031] S233: Based on the alignment error in the time deviation mapping table, perform weighted average processing on the start and end times of each note to generate a rhythm deviation matrix.
[0032] In one embodiment, S3 includes:
[0033] S31: Construct a mixture Gaussian - hidden Markov model based on the deviation analysis parameter set;
[0034] S32: Perform spatio - temporal correlation analysis on the deviation analysis parameter set through the mixture Gaussian - hidden Markov model to generate an initial deviation feature library;
[0035] S33: Based on the deviation analysis parameter sets stored in the database for the historical learner group, perform classification processing on the initial deviation feature library through the spectral clustering algorithm to generate a group deviation feature library;
[0036] S34: Based on the transfer learning framework, inject the population deviation feature library as prior knowledge into the Gaussian mixture - Hidden Markov model for processing to generate a personalized deviation feature library.
[0037] In one embodiment, S4 includes:
[0038] S41: Process the pitch, rhythm, and timbre deviation patterns of the personalized deviation feature library based on the reinforcement learning strategy to generate a targeted practice strategy, where the targeted practice strategy includes track parameters with dynamically adjusted difficulty;
[0039] S42: Process the fundamental frequency deviation quantization parameter to generate a tactile feedback signal including vibration intensity;
[0040] S43: Process the timbre spectrum feature to generate a visual heatmap feedback signal, where the visual heatmap feedback signal maps the timbre deviation degree through color gradients;
[0041] S44: Integrate the tactile feedback signal and the visual heatmap feedback signal according to the targeted practice strategy to generate a multi - modal correction feedback signal.
[0042] In one embodiment, S41 includes:
[0043] S411: Based on the pitch deviation pattern in the personalized deviation feature library, construct the state space of reinforcement learning, where the state space includes the current pitch deviation value, the historical correction success rate, and the difficulty level of the practice track;
[0044] S412: Iteratively optimize the action strategy in the state space through the proximal policy optimization algorithm to generate dynamically adjusted track parameters, where the track parameters include the target speed and the interval span;
[0045] S413: Process the track parameters based on the reward function to generate a set of scoring results, and the calculation formula for each scoring result in the set of scoring results is:
[0046]
[0047] where, γ t is each scoring result in the set of scoring results, Δf t is the current pitch deviation value of the fundamental frequency deviation quantization parameter, is the average error of the rhythm deviation matrix, and β1, β2 are preset weight coefficients;
[0048] S414: Screen the set of scoring results for high scores to generate a set of candidate strategies;
[0049] S415: Based on the greedy algorithm, prioritize the candidate policy set, screen out the policy parameter combination with the largest deviation decrease rate, and process the policy parameter combination based on dynamic priority allocation to generate a targeted practice strategy.
[0050] In a second aspect, the present application provides a music intelligent learning system, which includes:
[0051] A data acquisition and processing module, which is used to collect the performance data of the learner on the musical instrument based on multi-modal sensors, perform anti-interference processing on the performance data, and generate the music data to be inspected; the music data to be inspected includes a fidelity audio signal and tactile pressure time series data;
[0052] A feature extraction module, which is used to perform multi-dimensional deviation feature extraction processing based on the fidelity audio signal and tactile pressure time series data in the music data to be inspected, and generate a deviation analysis parameter set, where the deviation analysis parameter set includes fundamental frequency deviation quantization parameters, rhythm deviation matrices, and timbre spectrum features;
[0053] A feature library construction module, which is used to construct a personalized deviation feature library based on the deviation analysis parameter set through a hidden Markov model and transfer learning;
[0054] A policy generation and feedback module, which is used to generate a targeted practice strategy according to the personalized deviation feature library, and perform real-time analysis on the performance data of the learner based on the targeted practice strategy to generate a multi-modal correction feedback signal; the targeted practice strategy is used to dynamically adjust the speed, interval span, and repetition times of the practice piece; the multi-modal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
[0055] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements any one of the above music intelligent learning methods.
[0056] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above music intelligent learning methods.
[0057] In summary, a music intelligent learning method provided by the present invention collects the performance data of learners through multi-modal sensors, and performs anti-interference processing on the collected research data to ensure the accuracy and integrity of the data; through multi-dimensional deviation feature extraction, it can comprehensively and meticulously analyze the deviation situations of learners in aspects such as pitch, rhythm, and timbre. And a personalized deviation feature library constructed by using the hidden Markov model and transfer learning can thus accurately capture the performance characteristics and problems of learners. Based on this, the generated targeted practice strategy and multi-modal correction feedback signal can provide personalized practice guidance and real-time feedback for learners. Through the above multi-dimensional deviation feature extraction method, the present invention can comprehensively and deeply analyze various problems of learners in the process of musical instrument performance, providing a solid foundation for personalized training and real-time feedback, helping learners quickly improve their performance level, overcome technical obstacles, and improve practice efficiency and learning effect.
[0058] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Brief Description of the Drawings
[0059] Figure 1 It is a schematic flowchart of a music intelligent learning method provided by an embodiment of the present application;
[0060] Figure 2 It is a schematic flowchart of generating a rhythm deviation matrix provided by an embodiment of the present application;
[0061] Figure 3 It is a schematic flowchart of generating a multi-modal correction feedback signal provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic structural diagram of a music intelligent learning system provided by another embodiment of the present application. Detailed Embodiments
[0063] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant accompanying drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the invention more thorough and comprehensive.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0065] In one embodiment, asFigure 1 As shown, a music intelligent learning method is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0066] S1: Based on the multi-modal sensors, collect the performance data of the learner on the musical instrument, perform anti-interference processing on the performance data, and generate the music data to be inspected; the music data to be inspected includes a fidelity audio signal and tactile pressure time series data.
[0067] Specifically, the music intelligent learning system comprehensively collects the performance data of the learner on the musical instrument by using multi-modal sensors. The sensors include high-precision audio sensors, pressure sensors, acceleration sensors, etc., which are respectively installed at key parts of the musical instrument, such as the strings, frets, and bridge of stringed instruments, and the keys and hammering devices of keyboard instruments, to capture multi-dimensional information such as audio signals, tactile pressure changes, movement amplitude, and speed during the performance process.
[0068] Preferably, the performance data collected by the music intelligent learning system is inevitably affected by external interferences, such as environmental noise, electromagnetic interference, and sensor self-noise. To ensure the accuracy and reliability of the data, advanced anti-interference processing algorithms can be used to purify the data. For example, for audio signals, adaptive filtering technology can be used. By analyzing the characteristics and variation rules of the noise, the filter parameters can be dynamically adjusted to effectively remove environmental noise and electromagnetic interference components. For tactile pressure time series data, a method combining median filtering and wavelet transform can be used. First, median filtering is used to remove salt-and-pepper noise in the data, and then wavelet transform is used to perform multi-scale decomposition and reconstruction on the signal. Threshold processing is performed in different frequency bands to remove high-frequency noise and retain the detailed characteristics of pressure changes, so as to obtain stable and real tactile pressure data.
[0069] After completing the anti-interference processing, the data after anti-interference processing is integrated into the music data to be inspected. Among them, the fidelity audio signal truly reflects the sound quality of the musical instrument, including information such as pitch, timbre, and volume; the tactile pressure time series data records the pressure change situation applied by the learner to the musical instrument during the performance process, such as the magnitude and change trend of the pressing force on the strings when playing stringed instruments, and the force and timing relationship of pressing keys on keyboard instruments, providing a comprehensive and accurate data basis for subsequent deviation analysis.
[0070] S2: Based on the fidelity audio signal and tactile pressure time series data in the music data to be inspected, perform multi-dimensional deviation feature extraction processing to generate a deviation analysis parameter set, and the deviation analysis parameter set includes fundamental frequency deviation quantization parameters, rhythm deviation matrices, and timbre spectrum features.
[0071] Specifically, the music intelligent learning system performs multi-dimensional deviation feature extraction processing based on the high-fidelity audio signal and tactile pressure time-series data in the music data to be inspected. The processing process is as follows:
[0072] For the high-fidelity audio signal, the fast Fourier transform (FFT) algorithm can be used to convert it from the time domain to the frequency domain to obtain the frequency spectrum of the signal. By analyzing the frequency spectrum, the fundamental frequency is accurately located, that is, the main frequency component of the musical instrument performance sound, and then compared with the standard fundamental frequency to calculate the fundamental frequency deviation value, and the distribution of the deviation is statistically analyzed, such as the mean, variance, maximum value, minimum value, etc. of the deviation, to form the fundamental frequency deviation quantization parameter, which intuitively reflects the deviation degree and stability of the learner in pitch control.
[0073] For the rhythm deviation analysis, the tactile pressure time-series data can be compared with a preset standard rhythm template. The standard rhythm template is preset according to the beat and rhythm pattern of the music, and contains information such as the start time, duration, and interval time of each note. Through the dynamic time warping (DTW) algorithm, the tactile pressure time-series data and the standard rhythm template are non-linearly aligned on the time axis, and the time difference at the corresponding positions is calculated to construct a rhythm deviation matrix. This matrix details the advance or lag of each note in the performance time, as well as the compliance of the rhythm pattern, providing an accurate quantitative basis for rhythm training.
[0074] In terms of timbre analysis, time-frequency analysis is performed on the high-fidelity audio signal to extract its spectral features, including parameters such as spectral centroid, spectral bandwidth, harmonic components, and noise components. These parameters comprehensively reflect the timbre characteristics of the audio signal. By comparing with the standard timbre template, the difference between the performance timbre and the ideal timbre is determined, and the timbre spectral features are generated to provide guidance for the learner in timbre control and optimization.
[0075] S3: Based on the deviation analysis parameter set, a personalized deviation feature library is constructed through the hidden Markov model and transfer learning.
[0076] First, the music intelligent learning system uses the fundamental frequency deviation quantization parameter, rhythm deviation matrix, and timbre spectral features in the deviation analysis parameter set as the observation sequences, and uses the HMM to model these observation sequences. Specifically, the HMM can capture the implicit relationships and change rules between deviation features, and determine the association patterns between different deviation features by learning the state transition probability and output probability in the training data, so as to construct a probability model that can reflect the internal structure of the learner's performance deviation features.
[0077] On this basis, transfer learning technology is introduced to transfer the deviation feature knowledge and experience of other similar learners to the model of the current learner. Specifically, transfer learning broadens the knowledge scope of the current learner's deviation feature library by selectively drawing on the excellent practice results and common deviation patterns of other learners in aspects such as intonation, rhythm, and timbre. This enables it to not only include the deviation features reflected in its own performance data but also absorb the beneficial experience of other learners, accelerating the training convergence speed of the model, enhancing the accuracy and generalization ability of the model, and laying a foundation for generating more targeted and effective practice strategies subsequently.
[0078] S4: Generate a targeted practice strategy based on the personalized deviation feature library and conduct real-time analysis of the learner's performance data based on the targeted practice strategy to generate a multimodal correction feedback signal; the targeted practice strategy is used to dynamically adjust the speed, interval span, and repetition times of the practice piece; the multimodal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
[0079] Specifically, the targeted practice strategy can accurately locate the problems and weak links existing in the learner's performance according to the data in the deviation feature library, such as specific pitch regions with unstable intonation, complex rhythm patterns prone to rhythm mistakes, and pitch segments with improper timbre control. In response to these problems, the music intelligent learning system can dynamically adjust the speed of the practice piece. In the initial stage, the speed is set slightly lower than the standard speed, allowing the learner to familiarize and correct the performance actions at a relatively relaxed rhythm. As the practice progresses and the mastery level improves, the speed is gradually increased until the standard speed or higher requirements are reached; and the interval span is reasonably planned. For pitch regions with prominent intonation problems, the practice proportion of the corresponding intervals is increased, and special interval practice pieces are designed to strengthen the learner's auditory perception and finger control ability for these intervals. At the same time, the music intelligent learning system can also scientifically set the repetition times, determining the appropriate number of repeated practice times for each practice link and piece segment according to the learner's mastery situation and deviation degree, ensuring that the learner forms correct performance habits and muscle memory through sufficient practice.
[0080] In summary, a music intelligent learning method provided by the present invention collects the performance data of learners through multi-modal sensors, and performs anti-interference processing on the collected research data to ensure the accuracy and integrity of the data; through multi-dimensional deviation feature extraction, it can comprehensively and meticulously analyze the deviation situations of learners in aspects such as pitch, rhythm, and timbre. And a personalized deviation feature library constructed by using the hidden Markov model and transfer learning can accurately capture the performance characteristics and problems of learners. Based on this, the targeted practice strategy and multi-modal correction feedback signal generated can provide personalized practice guidance and real-time feedback for learners. Through the above multi-dimensional deviation feature extraction method, the present invention can comprehensively and deeply analyze various problems of learners in the process of musical instrument performance, provide a solid foundation for personalized training and real-time feedback, help learners quickly improve their performance level, overcome technical obstacles, and improve practice efficiency and learning effect.
[0081] In one embodiment, step S1 of a music intelligent learning method in the embodiments of the present application specifically includes:
[0082] S11: Collect the original audio signal based on the microphone of the multi-modal sensor, and obtain the tactile pressure data based on the pressure sensor in the multi-modal sensor.
[0083] Specifically, the music intelligent learning system comprehensively collects the performance data of learners on musical instruments by using multi-modal sensors. Among them, a high-precision microphone is installed at the sound outlet of the musical instrument or near the sound-generating part to capture the original audio signal emitted by the musical instrument during the performance. This signal contains rich information such as the pitch, timbre, and volume of the musical instrument performance, reflecting the performance situation of the learner in aspects such as intonation and timbre control. At the same time, pressure sensors are arranged at the key contact points of the musical instrument, such as the strings, frets, and bridges of stringed instruments, and the keys of keyboard instruments, etc., to obtain the tactile pressure data exerted by the learner on the musical instrument during the performance. These data can reflect the operation details of the learner's fingering, string-pressing force, key-pressing depth, etc., and provide an important basis for analyzing performance skills and problems.
[0084] S12: Perform comb filter noise reduction processing on the original audio signal to generate a fidelity audio signal.
[0085] Specifically, the calculation formula of the fidelity audio signal is as follows:
[0086]
[0087] Among them, S'(f) is the spectrum after noise reduction, that is, the high-fidelity audio signal, S(f) is the original audio spectrum, N(f / k) is the energy of the environmental noise at the k-th harmonic, and α is the noise suppression factor, whose value range is usually between 0 and 1, and is used to control the intensity of noise reduction. Through this formula, the system can subtract the energy component of the environmental noise from the original audio spectrum, thereby effectively removing the interference of external factors such as background noise and electromagnetic interference on the audio signal, retaining the true sound characteristics of the instrument performance, and ensuring the accuracy of subsequent analysis.
[0088] S13: Perform Kalman filter smoothing on the tactile pressure data to generate standardized tactile pressure time series data.
[0089] Specifically, the music intelligent learning system performs Kalman filter smoothing on the collected tactile pressure data to remove random noise and interference in the data and generate standardized tactile pressure time series data. The Kalman filter uses a recursive algorithm to perform real-time estimation and correction of the pressure data through two steps: prediction and update. In the prediction step, based on the pressure data at the previous moment and the system model, the pressure value at the current moment is predicted; in the update step, the actually measured pressure data is compared with the predicted value, the Kalman gain is calculated, and the predicted value is corrected to obtain a more accurate estimated value. The tactile pressure time series data after Kalman filter processing is smoother and more stable, and can truly reflect the changes in the pressure exerted by the learner on the instrument during the performance, providing reliable data support for subsequent deviation analysis.
[0090] S14: Organize the high-fidelity audio signal and the tactile pressure time series data to generate the music data to be inspected.
[0091] Specifically, the high-fidelity audio signal ensures the high quality and authenticity of the audio information, and the tactile pressure time series data provides detailed information about the performance actions. The combination of these two types of data comprehensively records the performance of the learner during the instrument performance, providing a complete and accurate data basis for subsequent steps such as multi-dimensional deviation feature extraction, deviation analysis parameter set generation, and targeted practice strategy formulation, and is a key data input link in the entire instrument performance training method.
[0092] In one of the embodiments, as Figure 2 shown, step S2 of a music intelligent learning method in the embodiment of the present application specifically includes:
[0093] S21: Perform separation processing on the high-fidelity audio signal based on the wavelet packet transform technology to obtain an isolated audio signal that separates the fundamental frequency component and the overtone frequency band.
[0094] Specifically, the wavelet packet transform is a multi-resolution analysis tool that can decompose a signal into sub-band signals of different frequency bands. Specifically, the high-fidelity audio signal is decomposed into multiple frequency bands through the wavelet packet transform, and the fundamental frequency component and the overtone frequency band are separated. The fundamental frequency component corresponds to the main frequency of the musical instrument performance and reflects the pitch information; the overtone frequency band contains rich harmonic components and is closely related to the timbre characteristics. Through this separation process, the fundamental frequency and overtones can be accurately analyzed respectively, providing a basis for the subsequent extraction of deviation features.
[0095] S22: Based on the tactile pressure time series data, correct the fundamental frequency offset of the isolated audio signal to generate a fundamental frequency deviation quantization parameter.
[0096] Specifically, the calculation formula for the fundamental frequency deviation quantization parameter is as follows:
[0097] Δf(t) = |f user (t) - f std |
[0098] where Δf(t) is the fundamental frequency deviation quantization parameter, f user (t) is the real-time fundamental frequency of the isolated audio signal, and f std is the theoretical fundamental frequency value of the corresponding note in the standard music score. Through the above formula, the average deviation between the fundamental frequency during the learner's performance and the standard fundamental frequency is calculated, quantifying the error degree in pitch and providing an accurate parameter basis for pitch training.
[0099] S23: Based on the dynamic time warping algorithm and the standard music score time axis alignment technology, process the performance timing signal of the tactile pressure time series data in the music data to be tested to generate a rhythm deviation matrix.
[0100] Preferably, the rhythm deviation matrix is generated through the following steps:
[0101] S231: Based on the time series segmentation algorithm, perform frame segmentation on the performance timing signal of the tactile pressure time series data to generate a discrete time stamp sequence.
[0102] Specifically, the music intelligent learning system segments the continuous performance timing signal according to a preset time window and step size to generate a discrete time stamp sequence. Each time stamp corresponds to the time point of a performance action, such as the moment of pressing a key or a string. Through frame segmentation, the long-time performance signal is divided into multiple short-time segments, facilitating subsequent detailed analysis and processing.
[0103] S232: Calculate the optimal alignment path between the discrete time stamp sequence and the standard music score time axis through the dynamic time warping algorithm to generate a time deviation mapping table.
[0104] Specifically, the DTW algorithm can handle the non-linear changes in time series and find the best temporal match between two sequences. In this step, the music intelligent learning system aligns the performance timing signals of the learner with the time axis of the standard musical score, determines the corresponding positions of each performance action on the standard time axis, generates a time deviation mapping table, which records the deviation relationship between the actual performance time and the standard time.
[0105] S233: Based on the alignment errors in the time deviation mapping table, perform weighted average processing on the start and end times of each note to generate a rhythm deviation matrix.
[0106] Specifically, for each note, the system calculates the weighted average deviation value according to the start and end time deviations in the time deviation mapping table to form an element in the rhythm deviation matrix. This matrix details the early or late situation of each note in the performance time and the compliance of the rhythm pattern, providing a comprehensive quantitative analysis for rhythm training.
[0107] S24: Perform Mel-frequency spectrum conversion processing on the general audio segment of the isolated audio signal to generate timbre spectrum features.
[0108] Specifically, the timbre spectrum features are generated through the following formula:
[0109]
[0110] where D timbre is the timbre spectrum feature, P std (m) is the energy distribution of the standard timbre in the m-th Mel frequency band, and P user (m) is the corresponding energy distribution of the learner's performance timbre. Through Mel-frequency spectrum conversion, the energy distribution of the general audio segment is converted to the Mel frequency domain and compared with the energy distribution of the standard timbre to generate timbre spectrum features, which reflect the difference between the learner's performance timbre and the ideal timbre and provide guidance for timbre optimization.
[0111] S25: Organize the fundamental frequency deviation quantization parameters, the rhythm deviation matrix, and the timbre spectrum features to generate a set of deviation analysis parameters.
[0112] Specifically, the set of deviation analysis parameters comprehensively covers the performance deviation information of the learner in aspects such as pitch accuracy, rhythm, and timbre, providing rich data support for subsequent construction of a personalized deviation feature library, generation of targeted practice strategies, and multi-modal corrective feedback signals, and is a key link in realizing precise and efficient musical instrument performance training.
[0113] The above music intelligent learning method can effectively extract the fundamental frequency and overtone information by separating the audio signal using wavelet packet transform, providing a clear and accurate signal basis for subsequent deviation analysis. Secondly, by aligning the performance timing with the standard music score timeline through the dynamic time warping algorithm, the rhythm deviation can be accurately calculated, helping the practitioner better master the sense of rhythm; and through the Mel spectrum conversion process, the timbre spectrum characteristics can be analyzed from the perspective of auditory perception, enabling the practitioner to understand their problems in timbre control. Finally, the fundamental frequency deviation, rhythm deviation, and timbre spectrum characteristics are integrated into a deviation analysis parameter set, providing a comprehensive and systematic evaluation basis for music teachers or intelligent systems, helping to formulate a more scientific and personalized practice plan, and improving the efficiency and quality of music learning.
[0114] In one embodiment, step S3 of a music intelligent learning method in an embodiment of the present application specifically includes:
[0115] S31: Based on the deviation analysis parameter set, construct a Gaussian mixture-hidden Markov model.
[0116] Specifically, the Gaussian mixture-hidden Markov model (GMM-HMM) combines the ability of the Gaussian mixture model (GMM) to model the probability distribution of observed data and the advantage of the hidden Markov model (HMM) to model the state transition of time series data. Specifically, first, the fundamental frequency deviation quantization parameter, rhythm deviation matrix, and timbre spectrum characteristics in the deviation analysis parameter set are used as multi-dimensional observation vectors, and statistical analysis is performed on these observation vectors to determine their probability distribution. Then, using the state transition matrix and output probability matrix of the HMM, the time series relationship of the observation vectors is modeled to capture the change law and internal connection of the deviation characteristics during the performance process, thereby constructing a Gaussian mixture-hidden Markov model (GMM-HMM) that can comprehensively describe the performance deviation characteristics of the learner.
[0117] S32: Perform spatio-temporal correlation analysis on the deviation analysis parameter set through the Gaussian mixture-hidden Markov model to generate an initial deviation feature library.
[0118] Specifically, the system performs spatio-temporal correlation analysis on the set of deviation analysis parameters through a Gaussian Mixture Model - Hidden Markov Model (GMM - HMM). In the time dimension, it analyzes the changing trends and temporal relationships of deviation features during the performance process. For example, the fluctuation of the fundamental frequency deviation in different time periods, the distribution pattern of the rhythm deviation in different musical phrases, etc.; in the space dimension, it explores the correlations between different deviation features, such as the association between the fundamental frequency deviation and the timbre spectrum features, the coupling relationship between the rhythm deviation and the timbre change, etc. Through this spatio-temporal correlation analysis, the deep - level structure and laws of deviation features are excavated, generating an initial deviation feature library containing rich spatio-temporal information, providing a basis for subsequent classification and personalized processing.
[0119] S33: Based on the set of deviation analysis parameters stored in the database for the historical learner group, the initial deviation feature library is classified through a spectral clustering algorithm to generate a group deviation feature library.
[0120] Specifically, the spectral clustering algorithm can utilize the similarity matrix of the data to map data points into a low - dimensional space for clustering, and can effectively process non - linearly separable data. In this step, the system first calculates the similarity between different deviation feature vectors in the initial deviation feature library to construct a similarity matrix; then, through eigenvalue decomposition and clustering steps, the deviation feature vectors are divided into multiple categories, and each category represents a type of performance problem with similar deviation features. Finally, a group deviation feature library is generated, which contains the common deviation feature patterns and classification information of different learner groups during the performance process, providing an important reference basis for subsequent transfer learning and the construction of a personalized deviation feature library.
[0121] S34: Based on the transfer learning framework, the group deviation feature library is injected into the Gaussian Mixture Model - Hidden Markov Model as prior knowledge for processing to generate a personalized deviation feature library.
[0122] Specifically, transfer learning transfers useful information in the source domain (group deviation feature library) to the target domain (deviation feature analysis of the current learner) through a knowledge transfer mechanism, accelerating the training convergence speed of the model and improving the accuracy and generalization ability of the model. Specifically, the music intelligent learning system initializes and adjusts the parameters of the Gaussian Mixture Model - Hidden Markov Model (GMM - HMM) using the classification information and deviation patterns in the group deviation feature library, enabling it to adapt and converge faster when learning the deviation features of the current learner.
[0123] The model processed through transfer learning can more accurately capture the personalized differences and details of the current learner's performance deviation features, generating a personalized deviation feature library that truly conforms to their individual performance characteristics, providing a solid foundation for subsequent generation of targeted practice strategies and multi - modal correction feedback signals.
[0124] In one embodiment, as Figure 3 shown, step S4 of a music intelligent learning method in an embodiment of the present application specifically includes:
[0125] S41: Process the pitch, rhythm, and timbre deviation patterns in the personalized deviation feature library based on the reinforcement learning strategy to generate a targeted practice strategy, where the targeted practice strategy includes track parameters with dynamically adjusted difficulty.
[0126] Specifically, system reinforcement learning learns the optimal behavior strategy through the interaction between the agent and the environment to maximize the cumulative reward. In this step, the system can use the performance deviation features of the learner as the environmental state and the adjustment of the practice strategy as the action of the agent to construct a reinforcement learning model. The model evaluates the effects of different practice strategies on deviation correction based on the data in the deviation feature library, dynamically adjusts the difficulty parameters of the practice tracks, such as speed, interval span, and repetition times, etc., to generate an optimized targeted practice strategy, ensuring that the learner can efficiently train for their own problems during the practice process and gradually improve their performance level. Preferably, the targeted practice strategy is generated through the following steps:
[0127] S411: Based on the pitch deviation pattern in the personalized deviation feature library, construct the state space of reinforcement learning, where the state space includes the current pitch deviation value, historical correction success rate, and practice track difficulty level.
[0128] Specifically, the state space of reinforcement learning includes three main dimensions, namely the current pitch deviation value, historical correction success rate, and practice track difficulty level. Among them, the current pitch deviation value directly reflects the degree of deviation of the learner's intonation in the recent performance and is a real-time error indicator; the historical correction success rate counts the success probability of the learner correcting the intonation deviation in the past practice, reflecting the learner's progress trend and correction ability; the practice track difficulty level is divided according to factors such as the speed, interval span, and rhythm complexity of the track, and is used to evaluate the challenge of the current practice content. These three dimensions together constitute the state space of the reinforcement learning model, comprehensively describing the learning status of the learner in pitch control.
[0129] S412: Iteratively optimize the action strategy in the state space through the proximal policy optimization algorithm to generate dynamically adjusted track parameters, where the track parameters include the target speed and interval span.
[0130] Specifically, the Proximal Policy Optimization (PPO) algorithm, as an advanced reinforcement learning algorithm, can effectively improve the performance of the policy while ensuring the stability of policy updates. In this step, the PPO algorithm is based on the data in the learner's personalized deviation feature library and continuously tries different action policies, that is, different ways of adjusting the parameters of the practice repertoire, and evaluates its effect on correcting pitch deviation.
[0131] After multiple rounds of iterative optimization, dynamically adjusted repertoire parameters are finally generated, mainly including the target speed and the pitch interval. Among them, the target speed is set according to the learner's current ability. In the initial stage, it may be slightly lower than the standard speed. As the practice progresses and the deviation correction effect improves, the speed is gradually increased until it reaches or exceeds the standard requirements; the pitch interval is reasonably planned according to the deviation of the learner in different pitch ranges, and the corresponding pitch interval practice is increased for the pitch range with a large deviation to help the learner comprehensively improve the pitch control ability.
[0132] S413: Process the repertoire parameters based on the reward function to generate a set of scoring results.
[0133] Specifically, the calculation formula for each scoring result in the set of scoring results is as follows:
[0134]
[0135] where γ t is each scoring result in the set of scoring results, Δf t is the current pitch deviation value of the fundamental frequency deviation quantization parameter, is the average error of the rhythm deviation matrix, and β1, β2 are preset weight coefficients. This reward function aims to comprehensively evaluate the deviation correction effect of the practice strategy on both pitch and rhythm. The generated set of scoring results contains the comprehensive performance scores under different combinations of repertoire parameters, providing a quantitative basis for subsequent policy screening.
[0136] S414: Screen the set of scoring results for high scores to generate a set of candidate policies.
[0137] Specifically, the system screens out the policy parameter combinations with scores higher than the threshold by setting a certain scoring threshold and incorporates them into the set of candidate policies. These candidate policies represent the practice parameter combinations with better performance in correcting pitch and rhythm deviations, providing high-quality alternative solutions for further optimizing the practice strategy and ensuring the efficiency and pertinence of the targeted practice strategy.
[0138] S415: Based on the greedy algorithm, perform priority ranking on the set of candidate policies, screen out the policy parameter combination with the largest deviation decrease rate, and process the policy parameter combination based on dynamic priority allocation to generate a targeted practice strategy.
[0139] Specifically, the greedy algorithm uses the deviation reduction rate as the main indicator to evaluate and sort each combination of policy parameters in the candidate policy set. The deviation reduction rate reflects the proportion by which the learner's playing deviation is expected to decrease after using the policy. The larger the reduction rate, the more significant the policy's corrective effect on the deviation.
[0140] Specifically, the system filters out the combination of policy parameters with the largest deviation reduction rate according to the sorting result. To further improve the adaptability and flexibility of the policy, the selected combination of policy parameters is processed based on dynamic priority allocation. According to the learner's real-time practice progress and deviation changes, the priority and parameter details of the policy are dynamically adjusted, and finally a targeted practice policy is generated. This policy can accurately target the learner's current playing problems, dynamically adjust parameters such as the speed, interval span, and repetition times of the practice piece, provide the learner with optimized practice guidance, help them efficiently correct deviations, and improve their playing level.
[0141] S42: Process the fundamental frequency deviation quantization parameter to generate a tactile feedback signal containing vibration intensity.
[0142] Preferably, the system can design corresponding vibration patterns according to the degree and change trend of the fundamental frequency deviation. For example, when the fundamental frequency deviation is large, increase the vibration intensity to clearly indicate to the learner that there is a significant intonation problem; when the fundamental frequency deviation gradually decreases, correspondingly reduce the vibration intensity to encourage the learner to maintain the correct playing method. Specifically, the vibration signal is transmitted to the learner in real time through a tactile feedback device installed on the instrument. For example, a micro vibration motor is set at the fingerboard position of a string instrument or under the keys of a keyboard instrument to accurately transmit the tactile feedback to the learner's fingers, helping them immediately perceive and adjust the intonation during playing and form correct muscle memory.
[0143] S43: Process the timbre spectrum characteristics to generate a visual heat map feedback signal, and the visual heat map feedback signal maps the timbre deviation degree through color gradients.
[0144] Specifically, the music intelligent learning system maps the timbre deviation degree through color gradients and converts different timbre spectrum characteristics into intuitive visual information. For example, green is used to represent that the timbre is close to the standard timbre and the deviation is small; yellow indicates that there is a certain deviation and adjustment is needed; red indicates that the deviation is large and key practice is required. The heat map is displayed in real time on a display screen connected to the instrument or a smart device. The learner can clearly understand their timbre control situation during playing by observing the color changes of the heat map, and intuitively see which notes or pitch ranges have timbre problems, so as to conduct targeted timbre optimization practice.
[0145] S44: Integrate the tactile feedback signal and the visual heat map feedback signal according to the targeted practice strategy to generate a multi-modal correction feedback signal.
[0146] Specifically, the integration process of the tactile feedback signal and the visual heat map feedback signal needs to consider the complementarity and synergy of the two feedback signals to ensure that they can work together to provide comprehensive and accurate feedback guidance for the learner. For example, in the targeted practice strategy, when the learner is required to focus on practicing the intonation of a certain pitch range, the tactile feedback signal will correspondingly enhance the vibration prompt for that pitch range, while the visual heat map feedback signal will show a more obvious color change at the position of that pitch range. The dual feedback strengthens the learner's attention and adjustment to the problematic pitch range. The multi-modal correction feedback signal acts on the learner through multiple sensory channels simultaneously, improving the reception efficiency and correction effect of the feedback information, helping the learner to improve the playing skills more quickly and effectively, and enhancing the overall playing quality.
[0147] A music intelligent learning method provided by the above embodiment can dynamically generate a targeted practice strategy according to the actual deviation of the practitioner by using the reinforcement learning strategy, ensuring that the practice content highly matches the needs of the practitioner, improving the pertinence and effect of the practice, converting the fundamental frequency deviation into a tactile feedback signal, and converting the timbre spectrum characteristics into a visual heat map feedback signal, enriching the form and dimension of the feedback, enabling the practitioner to obtain feedback information from multiple sensory channels, and more comprehensively understanding their own playing or singing state. Finally, the generation of the multi-modal correction feedback signal realizes the synergy of tactile and visual feedback, further enhancing the accuracy and effectiveness of the feedback, helping the learner to more quickly and accurately discover and correct problems, and improving the music performance level.
[0148] Preferably, as Figure 4 shown, the present invention also provides a music intelligent learning system 500, which includes:
[0149] A data acquisition and processing module 510, configured to collect the playing data of the learner on the instrument based on the multi-modal sensor, perform anti-interference processing on the playing data, and generate the music data to be inspected; the music data to be inspected includes a fidelity audio signal and a tactile pressure time series data;
[0150] A feature extraction module 520, configured to perform multi-dimensional deviation feature extraction processing based on the fidelity audio signal and the tactile pressure time series data in the music data to be inspected, and generate a set of deviation analysis parameters, where the set of deviation analysis parameters includes a fundamental frequency deviation quantization parameter, a rhythm deviation matrix, and a timbre spectrum characteristic;
[0151] A feature library construction module 530, configured to construct a personalized deviation feature library based on the set of deviation analysis parameters through a hidden Markov model and transfer learning;
[0152] A strategy generation and feedback module 540 is configured to generate a targeted practice strategy based on a personalized deviation feature library, and perform real-time analysis on the performance data of a learner based on the targeted practice strategy to generate a multi-modal correction feedback signal; the targeted practice strategy is used to dynamically adjust the speed, interval span, and repetition times of a practice piece; the multi-modal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
[0153] In summary, a music intelligent learning system provided by the present invention uses a multi-modal sensor by a data acquisition and processing module 510 to collect the performance data of a learner, and performs anti-interference processing on the collected research data to ensure the accuracy and integrity of the data; a feature extraction module 520 can comprehensively and meticulously analyze the deviation conditions of a learner in aspects such as pitch, rhythm, and timbre through multi-dimensional deviation feature extraction. And a personalized deviation feature library constructed by a feature library construction module 530 using a hidden Markov model and transfer learning can accurately capture the performance characteristics and problems of a learner; a targeted practice strategy and a multi-modal correction feedback signal generated by a strategy generation and feedback module 540 based on this can provide personalized practice guidance and real-time feedback for a learner. Through the above multi-dimensional deviation feature extraction method, the music intelligent learning system provided by the present invention can comprehensively and deeply analyze various problems of a learner during the process of musical instrument performance, provide a solid foundation for personalized training and real-time feedback, help the learner quickly improve the performance level, overcome technical obstacles, and improve the practice efficiency and learning effect.
[0154] Preferably, the data acquisition and processing module 510 is configured with the following units:
[0155] A raw data acquisition unit 511 is configured to collect a raw audio signal based on a microphone of the multi-modal sensor, and obtain tactile pressure data based on a pressure sensor in the multi-modal sensor;
[0156] An audio noise reduction processing unit 512 is configured to perform comb filter noise reduction processing on the raw audio signal to generate a fidelity audio signal;
[0157] A data smoothing processing unit 513 is configured to perform Kalman filter smoothing processing on the tactile pressure data to generate standardized tactile pressure time series data;
[0158] A music data generation unit 514 is configured to organize the fidelity audio signal and the tactile pressure time series data to generate music data to be inspected.
[0159] Preferably, the feature extraction module 520 is configured with the following units:
[0160] A signal separation unit 521 is configured to perform separation processing on the fidelity audio signal based on wavelet packet transform technology to obtain an isolated audio signal separating the fundamental frequency component and the overtone frequency band;
[0161] An offset correction unit 522 for correcting the fundamental frequency offset of the isolated audio signal based on the tactile pressure time series data to generate a fundamental frequency deviation quantization parameter;
[0162] A timing processing unit 523 for processing the performance timing signal of the tactile pressure time series data in the music data to be detected based on the dynamic time warping algorithm and the standard music score timeline alignment technology to generate a rhythm deviation matrix;
[0163] Preferably, the timing processing unit 523 is configured with the following sub-units:
[0164] A timing framing sub-unit 5231 for framing the performance timing signal of the tactile pressure time series data based on the time series segmentation algorithm to generate a discrete time stamp sequence;
[0165] A path alignment sub-unit 5232 for calculating the optimal alignment path between the discrete time stamp sequence and the standard music score timeline through the dynamic time warping algorithm to generate a time deviation mapping table;
[0166] A rhythm deviation generation sub-unit 5233 for performing weighted average processing on the start and end times of each note based on the alignment error in the time deviation mapping table to generate a rhythm deviation matrix;
[0167] A feature generation unit 524 for performing Mel spectrum conversion processing on the wide audio segment of the isolated audio signal to generate timbre spectrum features.
[0168] Preferably, the feature library construction module 530 is configured with the following units:
[0169] A model construction unit 531 for constructing a mixture Gaussian - hidden Markov model based on the deviation analysis parameter set;
[0170] An initial library generation unit 532 for performing spatio-temporal correlation analysis on the deviation analysis parameter set through the mixture Gaussian - hidden Markov model to generate an initial deviation feature library;
[0171] A population library generation unit 533 for classifying the initial deviation feature library through a spectral clustering algorithm based on the deviation analysis parameter set stored in the database by the historical learner population to generate a population deviation feature library;
[0172] A personalized library generation unit 534 for injecting the population deviation feature library as prior knowledge into the mixture Gaussian - hidden Markov model for processing based on the transfer learning framework to generate a personalized deviation feature library.
[0173] Preferably, the policy generation feedback module 540 is configured with the following units:
[0174] The practice strategy unit 541 is used to process the pitch, rhythm, and timbre deviation patterns of the personalized deviation feature library based on the reinforcement learning strategy, and generate a targeted practice strategy. The targeted practice strategy includes track parameters with dynamically adjusted difficulty;
[0175] Preferably, the practice strategy unit 541 is configured with the following sub-units:
[0176] The state space construction sub-unit 5411 is used to construct the state space of reinforcement learning based on the pitch deviation pattern in the personalized deviation feature library. The state space includes the current pitch deviation value, the historical correction success rate, and the difficulty level of the practice track;
[0177] The strategy optimization sub-unit 5412 is used to iteratively optimize the action strategy in the state space through the proximal policy optimization algorithm, and generate dynamically adjusted track parameters. The track parameters include the target speed and the interval span;
[0178] The scoring result generation sub-unit 5413 is used to process the track parameters based on the reward function, and generate a set of scoring results;
[0179] The candidate strategy screening sub-unit 5414 is used to perform high-score screening on the set of scoring results, and generate a set of candidate strategies;
[0180] The final strategy determination sub-unit 5415 is used to perform priority sorting on the set of candidate strategies based on the greedy algorithm, screen out the strategy parameter combination with the largest deviation decrease rate, and process the strategy parameter combination based on dynamic priority allocation to generate a targeted practice strategy;
[0181] The tactile feedback unit 542 is used to process the fundamental frequency deviation quantization parameter, and generate a tactile feedback signal including the vibration intensity;
[0182] The visual feedback generation unit 543 is used to process the timbre spectrum feature, and generate a visual heat map feedback signal. The visual heat map feedback signal maps the timbre deviation degree through a color gradient;
[0183] The multi-modal feedback integration unit 544 is used to integrate the tactile feedback signal and the visual heat map feedback signal according to the targeted practice strategy, and generate a multi-modal correction feedback signal.
[0184] In one embodiment, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned music intelligent learning method.
[0185] In one embodiment, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-described music intelligent learning method is implemented.
[0186] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0187] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0188] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A music intelligent learning method, characterized in that, It includes the following steps: S1: Based on the multi-modal sensors to collect the performance data of the learner on the musical instrument, perform anti-interference processing on the performance data to generate the music data to be inspected; the music data to be inspected includes a fidelity audio signal and tactile pressure time series data; S2: Based on the fidelity audio signal and tactile pressure time series data in the music data to be inspected, perform multi-dimensional deviation feature extraction processing to generate a set of deviation analysis parameters, and the set of deviation analysis parameters includes a fundamental frequency deviation quantization parameter, a rhythm deviation matrix, and a timbre spectrum feature; S3: Based on the set of deviation analysis parameters, construct a personalized deviation feature library through a hidden Markov model and transfer learning; S4: Generate a targeted practice strategy according to the personalized deviation feature library, and perform real-time analysis on the performance data of the learner based on the targeted practice strategy to generate a multi-modal correction feedback signal; the targeted practice strategy is used to dynamically adjust the speed, interval span, and repetition times of the practice repertoire; the multi-modal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
2. The music intelligent learning method according to claim 1, characterized in that, The S1 includes: S11: Collect the original audio signal based on the microphone of the multi-modal sensor, and obtain the tactile pressure data based on the pressure sensor in the multi-modal sensor; S12: Use the following formula to perform comb filter noise reduction processing on the original audio signal to generate a fidelity audio signal: Where S'(f) is the spectrum after noise reduction, that is, the fidelity audio signal, S(f) is the original audio spectrum, N(f / k) is the energy of the environmental noise at the kth harmonic, and α is the noise suppression factor; S13: Perform Kalman filter smoothing processing on the tactile pressure data to generate standardized tactile pressure time series data; S14: Organize the fidelity audio signal and the tactile pressure time series data to generate the music data to be inspected.
3. The music intelligent learning method according to claim 2, wherein The S2 includes: S21: Based on the wavelet packet transform technology, perform separation processing on the fidelity audio signal to obtain an isolated audio signal that separates the fundamental frequency component and the overtone frequency band; S22: Use the following formula to correct the fundamental frequency offset of the isolated audio signal based on the tactile pressure time series data to generate a fundamental frequency deviation quantization parameter: Δf(t) = |f user (t) - f std | where Δf(t) is the fundamental frequency deviation quantization parameter, and f user (t) is the real-time fundamental frequency of the isolated audio signal, and f std is the theoretical fundamental frequency value of the corresponding note in the standard musical score; S23: Based on the dynamic time warping algorithm and the standard score time axis alignment technology, process the performance time series signal of the tactile pressure time series data in the music data to be inspected to generate a rhythm deviation matrix; S24: Use the following formula to perform Mel spectrum conversion processing on the overtone frequency band of the isolated audio signal to generate a timbre spectrum feature: Among them, D timbre is the color spectrum feature, and P std (m) is the energy distribution of the standard timbre in the m-th Mel band, and P user (m) is the corresponding energy distribution of the timbre played by the learner; S25: Organize the fundamental frequency deviation quantization parameter, the rhythm deviation matrix, and the timbre spectrum feature to generate a set of deviation analysis parameters.
4. The music intelligent learning method according to claim 3, wherein The S23 includes: S231: Based on the time series segmentation algorithm, perform frame segmentation processing on the performance time series signal of the tactile pressure time series data to generate a discrete time stamp sequence; S232: Calculate the optimal alignment path between the discrete time stamp sequence and the standard score time axis through the dynamic time warping algorithm to generate a time deviation mapping table; S233: Based on the alignment errors in the time deviation mapping table, perform weighted average processing on the start and end times of each note to generate a rhythm deviation matrix.
5. The music intelligent learning method according to claim 1, wherein The S3 includes: S31: Based on the deviation analysis parameter set, construct a mixture Gaussian - Hidden Markov model; S32: Through the mixture Gaussian - Hidden Markov model, perform spatio - temporal correlation analysis on the deviation analysis parameter set to generate an initial deviation feature library; S33: Based on the deviation analysis parameter set stored in the database by the historical learner group, perform classification processing on the initial deviation feature library through a spectral clustering algorithm to generate a group deviation feature library; S34: Based on a transfer learning framework, inject the group deviation feature library as prior knowledge into the mixture Gaussian - Hidden Markov model for processing to generate a personalized deviation feature library.
6. The music intelligent learning method according to claim 1, wherein The S4 includes: S41: Based on a reinforcement learning strategy, process the pitch, rhythm, and timbre deviation patterns in the personalized deviation feature library to generate a targeted practice strategy, where the targeted practice strategy includes track parameters with dynamically adjusted difficulty; S42: Process the fundamental frequency deviation quantization parameter to generate a tactile feedback signal containing vibration intensity; S43: Process the timbre spectrum feature to generate a visual heat map feedback signal, where the visual heat map feedback signal maps the timbre deviation degree through a color gradient; S44: Integrate the tactile feedback signal and the visual heat map feedback signal according to the targeted practice strategy to generate a multi - modal correction feedback signal.
7. The music intelligent learning method according to claim 6, wherein The S41 includes: S411: Based on the pitch deviation pattern in the personalized deviation feature library, construct a state space for reinforcement learning, where the state space includes the current pitch deviation value, historical correction success rate, and practice track difficulty level; S412: Through the proximal policy optimization algorithm, iteratively optimize the action policy in the state space to generate dynamically adjusted track parameters, where the track parameters include the target speed and interval span; S413: Based on a reward function, process the track parameters to generate a set of scoring results, and the calculation formula for each scoring result in the set of scoring results is: Among them, γ t is each scoring result in the scoring result set, and Δf t is the current pitch deviation value of the fundamental frequency deviation quantization parameter, is the average error of the rhythm deviation matrix, and β1 and β2 are preset weight coefficients; S414: Perform high - scoring screening on the set of scoring results to generate a candidate policy set; S415: Based on a greedy algorithm, perform priority sorting on the candidate policy set, screen out the policy parameter combination with the largest deviation reduction rate, and process the policy parameter combination based on dynamic priority allocation to generate a targeted practice strategy.
8. A music intelligent learning system, characterized in that, The system includes: A data acquisition and processing module, configured to collect the performance data of the learner on the musical instrument based on multi - modal sensors, perform anti - interference processing on the performance data to generate music data to be inspected; the music data to be inspected includes a fidelity audio signal and tactile pressure time - series data; A feature extraction module, configured to perform multi - dimensional deviation feature extraction processing based on the fidelity audio signal and tactile pressure time - series data in the music data to be inspected to generate a deviation analysis parameter set, where the deviation analysis parameter set includes a fundamental frequency deviation quantization parameter, a rhythm deviation matrix, and a timbre spectrum feature; A feature library construction module, configured to construct a personalized deviation feature library based on the set of deviation analysis parameters through a hidden Markov model and transfer learning; A policy generation feedback module, configured to generate a targeted practice policy according to the personalized deviation feature library, and perform real-time analysis on the performance data of the learner based on the targeted practice policy to generate a multimodal correction feedback signal; the targeted practice policy is used to dynamically adjust the speed, interval span, and repetition times of the practice piece; the multimodal correction feedback signal includes a tactile feedback signal and a visual heat map feedback signal.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.