Piano practice assisting method, system and equipment based on AI (Artificial Intelligence) and medium
By accurately synchronously processing and multi-dimensional analysis of the sound and action data in piano practice, combined with emotion analysis, personalized improvement suggestions are provided, and the problem that existing tools cannot fully reflect the performance status is solved, improving the efficiency and pertinence of the practice.
Patent Information
- Application Number
- CN202510421839.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-08
AI Technical Summary
The existing piano practice aid tools cannot achieve accurate synchronization of sound and movements, cannot fully reflect the learner's playing status, lack real-time performance analysis and personalized improvement suggestions, and lack of intelligent support for the practice plan, resulting in inefficiency in practice.
By obtaining the sound and action data of learners when playing piano, using signal processing technology and action recognition algorithms for denoising and segmentation, using a dual-modal Transformer evaluation network for multi-dimensional analysis, combined with sentiment analysis model, it provides personalized improvement directions and priority sequences.
A comprehensive performance evaluation and sentiment analysis are achieved, which improves the efficiency and pertinence of the exercises, provides accurate feedback and personalized suggestions, helping learners quickly discover and improve weak links.
Smart Images

Figure CN120452288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI-based piano practice assisting method, system, device and medium. Background Art
[0002] Music education, especially piano learning, is a key area for cultural heritage and personal development. It has attracted a large number of participants worldwide and continues to show strong demand. However, effective instruction and scientific learning management have always been major challenges for music educators and learners during piano practice.
[0003] Traditional piano practice relies primarily on face-to-face instruction from teachers or spontaneous practice by learners. While face-to-face instruction can provide immediate feedback and personalized advice, it's limited by time and location and can't always meet learners' needs. Spontaneous practice, on the other hand, often lacks a systematic feedback mechanism, resulting in inefficient practice and making it difficult for learners to quickly identify and improve their own shortcomings.
[0004] To aid piano practice, a number of tools, such as metronomes or simple recording devices, have emerged. These tools can, to a certain extent, record the performance and help learners master rhythm and pitch. However, they lack real-time performance analysis and personalized improvement suggestions. This makes it difficult for learners to obtain comprehensive feedback during practice, preventing them from accurately understanding their performance level and weaknesses.
[0005] In recent years, with the continuous development of technology, a number of AI-based piano practice assistance methods have gradually emerged. However, these methods still face many technical challenges. First, the technology for synchronously capturing sound signals and movement expressions is still immature. In the piano performance process, sound and movement are two inseparable parts, together constituting the performer's complete performance. However, existing acquisition technology has difficulty achieving precise synchronization of sound and movement, resulting in an inability to fully reflect the learner's performance.
[0006] Secondly, existing analysis methods have limitations when processing complex musical information. Piano playing involves multiple technical dimensions, and existing analysis methods often only assess certain dimensions, failing to balance technical accuracy with emotional expression. This results in learners not receiving comprehensive feedback during practice, making it difficult to fully improve their performance.
[0007] Finally, the generation and recording of practice plans lacks intelligent support. Developing a sound practice plan is crucial for improving piano learning efficiency. However, existing practice plans are often developed by teachers or learners based on their own experience, lacking scientific and targeted approaches. Furthermore, recording practice progress lacks intelligent management tools, resulting in a lack of continuity and targeted learning. Summary of the Invention
[0008] The purpose of the present invention is to provide an AI-based piano practice assistance method, system, device and medium to provide learners with comprehensive performance evaluation and emotional analysis, effectively improving the efficiency and pertinence of piano practice, so as to solve at least one of the above-mentioned existing technical problems.
[0009] In a first aspect, the present invention provides an AI-based piano practice assistance method, the method specifically comprising:
[0010] Acquire the learner's sound data and movement data when playing the piano, and align and match the sound data and movement data;
[0011] De-noising and segmenting the sound data using signal processing technology, and extracting finger keystroke timing features of the action data using an action recognition algorithm to obtain a sound feature set and an action feature set;
[0012] Based on the sound feature set and the action feature set, a bimodal Transformer evaluation network is used to perform multidimensional analysis to obtain the first evaluation score;
[0013] Acquiring emotional expression data of the learner during piano playing, analyzing and processing the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result;
[0014] Based on the first sentiment analysis result and the first evaluation score, a deep learning algorithm is used to classify the learner's playing style and weaknesses, and to determine personalized improvement directions and priority sequences.
[0015] In a second aspect, the present invention provides an AI-based piano practice auxiliary system, the system specifically comprising:
[0016] The first auxiliary module is used to obtain the sound data and movement data of the learner when playing the piano, and align and match the sound data and movement data;
[0017] The second auxiliary module is configured to perform denoising and segmentation on the sound data using a signal processing technique, and extract finger keystroke timing features of the action data using an action recognition algorithm to obtain a sound feature set and an action feature set;
[0018] The third auxiliary module is used to perform multi-dimensional analysis based on the sound feature set and the action feature set using a bimodal Transformer evaluation network to obtain a first evaluation score;
[0019] A fourth auxiliary module is used to obtain emotional expression data of the learner during piano playing, and analyze and process the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result;
[0020] The fifth auxiliary module is used to classify the learner's playing style and weak links based on the first sentiment analysis result and the first evaluation score using a deep learning algorithm to determine personalized improvement directions and priority sequences.
[0021] In a third aspect, the present invention provides a computer device comprising: a memory and a processor and a computer program stored in the memory, wherein when the computer program is executed on the processor, the AI-based piano practice assistance method as described in any one of the above methods is implemented.
[0022] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an AI-based piano practice assistance method as described in any one of the above methods.
[0023] Compared with the prior art, the present invention has at least one of the following technical effects:
[0024] 1. The present invention provides learners with comprehensive performance evaluation and emotional analysis, effectively improving the efficiency and pertinence of piano practice.
[0025] 2. The present invention uses spectral subtraction and wavelet threshold method to jointly reduce the noise of sound data, thereby improving the clarity of sound data; and uses short-time Fourier transform to extract spectral features to obtain an accurate sound feature set.
[0026] 3. The present invention uses the HRNet network to detect the key points of the hand in the motion data, and extracts the motion features based on the key action judgment model, thereby achieving accurate capture and feature extraction of the motion data, and providing a reliable data basis for subsequent evaluation and analysis.
[0027] 4. The present invention realizes multidimensional analysis of sound feature sets and action feature sets through a bimodal Transformer evaluation network. It can calculate the correlation between action and audio and the response of audio to action to form fusion features, and perform multidimensional analysis on the fusion features to obtain explanations and contribution weights of multiple evaluation dimensions, which helps to comprehensively and accurately evaluate the learner's performance level and discover existing problems and weak links.
[0028] 5. The present invention projects the emotion analysis results onto a 2D plane through an emotion visualization layer, providing learners with intuitive emotional feedback, which helps to improve the expressiveness and appeal of the performance.
[0029] 6. The present invention provides customized improvement suggestions based on the learner's actual situation and needs, which helps learners to quickly improve their performance level.
[0030] 7. The present invention can monitor the learner's learning progress and stress status in real time, ensure the rationality and feasibility of the practice plan, and help maintain the learner's learning motivation and interest.
[0031] 8. The present invention can provide personalized music recommendations based on the learner's actual environment and skill level, which helps to enrich the practice content and improve the practice effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0033] Figure 1 1 is a flowchart of an AI-based piano practice assistance method provided by one embodiment of the present invention;
[0034] Figure 2 1 is a schematic structural diagram of an AI-based piano practice assistance system provided by one embodiment of the present invention;
[0035] Figure 3 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0036] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0037] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0038] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0039] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0040] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0041] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0042] In the embodiments of the present application, the execution subject of the process includes a terminal device, which includes but is not limited to: a server, a computer, a smart phone, a tablet computer, and other devices capable of executing the method disclosed in the present application. Figure 1 A flowchart of the AI-based piano practice assisting method disclosed in the first embodiment of the present invention is shown, and is described in detail as follows:
[0043] S101, obtaining the sound data and movement data of the learner when playing the piano, and aligning and matching the sound data and movement data.
[0044] In this embodiment, high-precision sensors are installed on the electronic piano to detect the keyboard pressure and force changes. These sensors can accurately record the touch time and force of each key when the learner plays. A high-definition camera is deployed on the piano or in the practice area to capture the learner's hand movements and facial expressions. The high-definition camera ensures the clarity and accuracy of the image. A professional audio acquisition device (such as a microphone) is used to record the sound data of the learner's performance. The learner plays the piano according to the teaching requirements. At the same time, the sensors, camera and audio acquisition device work synchronously to collect movement data and sound data respectively. The collected sound data is subjected to noise reduction processing to remove background noise and interference during the recording process, thereby improving the quality of the sound data. The collected movement data is smoothed to remove small fluctuations caused by sensor jitter or camera shaking, ensuring the accuracy of the movement. Acoustic features such as fundamental frequency (F0), short-time energy (Short-time energy), Mel-frequency cepstral coefficients (MFCC) are extracted from the sound data. These features can reflect the essential attributes and distinguishing properties of the sound. Key movement features are extracted from the movement data, such as finger number, placement, movement trajectory, and dynamics. These features can reflect the learner's playing technique and fingering accuracy. The sound and movement data are aligned using the Dynamic Time Warping (DTW) algorithm. The DTW algorithm elastically scales the time axis to better align two time series data. This allows for the calculation of similarities between sounds and movements of varying lengths or varying degrees of variation along the time axis. During the alignment process, multiple characteristic dimensions of sound and movement, such as pitch, rhythm, and dynamics, can be comprehensively considered to improve the accuracy of the alignment. The aligned data is analyzed to assess the learner's playing technique, sense of rhythm, and emotional expression. The analysis results are intuitively displayed in charts, animations, and other formats to help learners and teachers better understand the learner's performance. Based on the analysis results, learners are provided with personalized teaching suggestions and guidance to help them improve their playing techniques and enhance their musical expression.
[0045] S102, using signal processing technology to denoise and segment the sound data, and using a motion recognition algorithm to extract finger keystroke timing features of the motion data to obtain a sound feature set and a motion feature set.
[0046] In this embodiment, a sound signal is subjected to wavelet decomposition to obtain wavelet coefficients of different scales and frequencies. Based on the distribution characteristics of the wavelet coefficients of noise and signal at different scales, an appropriate threshold is selected for threshold processing to remove the noise coefficient. The retained signal coefficients are then used for wavelet reconstruction to obtain a denoised sound signal. Features such as short-time energy and zero-crossing rate are extracted from the denoised sound signal. Based on the extracted features, endpoint detection of the sound signal is performed using methods such as a double threshold method or an energy ratio method to determine the start and end positions of each note. The sound signal is segmented into multiple independent note segments, each corresponding to the performance of a note. Acoustic features such as fundamental frequency, spectral centroid, and Mel-Frequency Cepstral Coefficients (MFCC) are extracted from each segmented note segment to construct a sound feature set. The preprocessed action data is analyzed using an action recognition algorithm to extract the temporal features of finger keystrokes, including the time of keystrokes, the order of keystrokes, and the force of keystrokes. The extracted temporal features of finger keystrokes are sorted and organized to construct a motion feature set.
[0047] In this embodiment, a wavelet transform-based denoising algorithm effectively removes noise interference from the sound data, improving the clarity and accuracy of the sound signal. The segmentation of the sound data allows for independent analysis of the performance of each note, facilitating subsequent feature extraction and pattern recognition. A deep learning-based motion recognition algorithm accurately identifies the learner's finger keystrokes and extracts detailed timing features. These timing features not only reflect the learner's keystroke timing and sequence, but also include dynamic information such as the force of the keystrokes, providing strong support for analyzing the learner's playing technique.
[0048] S103: Based on the sound feature set and the action feature set, a bimodal Transformer evaluation network is used to perform multidimensional analysis to obtain a first evaluation score.
[0049] In this embodiment, a bimodal Transformer evaluation network is constructed, comprising two independent encoders (one for processing the sound feature set and the other for processing the motion feature set) and a shared decoder. The sound encoder utilizes a multi-layer Transformer encoder structure, with each layer incorporating a self-attention mechanism and a feedforward neural network. This encoder is capable of capturing temporal dependencies and interactions between features within the sound feature set. The motion encoder also utilizes a multi-layer Transformer encoder structure, but is optimized for the characteristics of the motion feature set, such as adding positional encoding to capture the timing information of keystrokes. The decoder utilizes a multi-layer Transformer decoder structure, responsible for fusing the outputs of the sound encoder and motion encoder to generate the final evaluation result. The sound feature set and motion feature set are input into their respective encoders. The encoder processes the input feature sets and generates their respective encoded representations. The decoder fuses the outputs of the sound encoder and motion encoder, capturing the intrinsic connections between the two modalities through a self-attention mechanism, to generate the final evaluation result. A cross-entropy loss function or a mean squared error loss function is used, depending on the requirements of the specific evaluation task. The bimodal Transformer evaluation network is trained using labeled performance data to optimize network parameters. The sound feature set and action feature set to be evaluated are input into the trained network, and the first evaluation score is obtained through network calculation.
[0050] In this embodiment, the bimodal Transformer evaluation network can process the sound feature set and the action feature set simultaneously to achieve a comprehensive evaluation of the learner's performance. This multidimensional analysis method is more accurate and comprehensive than the traditional unimodal evaluation method. Through the Transformer's self-attention mechanism, the network can capture the intrinsic connection between the sound feature set and the action feature set, such as the correspondence between the key action and the sound generation. This ability makes the evaluation results more accurate and reliable. The bimodal Transformer evaluation network can process the sound feature set and the action feature set in parallel, improving the evaluation efficiency. At the same time, the structural design of the network enables it to quickly adapt to different evaluation tasks and data sets.
[0051] S104, acquiring emotional expression data of the learner during piano playing, analyzing and processing the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result.
[0052] In this embodiment, the emotional expression data includes facial micro-expressions and upper body posture, physiological signals, and key force-velocity joint features. An appropriate emotional analysis model is selected based on the application scenario and data characteristics. For example, a neural network model based on deep learning, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer, can be used. The model is trained using labeled emotional expression data. This data should include performance samples of different emotional types and corresponding emotional labels. Through training, the model can learn the mapping relationship between emotional features and emotional types. The model calculates the emotional probability distribution based on the input features, that is, the probability of each emotional type. Based on the probability distribution, the main emotional types expressed by the learner during the performance are determined, and an emotional analysis result is generated. The emotional analysis results are presented to the learner or teacher in a visual manner, such as through charts, animations, etc. to show the emotional type, emotional intensity, and emotional change trends. Based on the emotional analysis results, personalized emotional expression suggestions and guidance are provided to learners to help them improve their playing skills and enhance their emotional expression ability.
[0053] In this embodiment, based on the emotion analysis results, the system can provide learners with personalized emotion expression suggestions and guidance. These suggestions and guidance are based on the learners' actual performance and are therefore more targeted and practical.
[0054] S105 , based on the first sentiment analysis result and the first evaluation score, a deep learning algorithm is used to classify the learner's playing style and weaknesses, and to determine personalized improvement directions and priority sequences.
[0055] In this embodiment, the first sentiment analysis result (such as sentiment type, sentiment intensity, etc.) and the first evaluation score (such as skill score, sense of rhythm score, etc.) are integrated to form a performance data set of the learner. The data is standardized or normalized to ensure that data of different dimensions are comparable. A classification model in a deep learning algorithm, such as a deep learning version of a convolutional neural network (CNN) or a support vector machine (SVM) (through a kernel method or deep feature extraction combined with SVM), is used to classify the performance data set by style. The classification categories can be set according to the needs of the teaching system, such as romanticism, classicalism, modernism, etc., or more fine-grained style classification. By training the classification model, the system can automatically identify the learner's performance style and label it accordingly. The clustering or anomaly detection model in the deep learning algorithm, such as a self-organizing map (SOM) or an anomaly detection network based on deep learning, is used to identify weak links in the performance data set. Weak links include deficiencies in technique, lack of sense of rhythm, inadequate emotional expression, etc. Through model analysis, the system can identify and categorize learners' specific weaknesses, such as technical weaknesses and emotional expression weaknesses. Based on the performance style classification results and the identification of weaknesses, the system uses a rule engine or a deep learning-based decision-making model to determine personalized improvement directions for the learner. These may include strengthening specific techniques, improving rhythmic control, and enhancing emotional expression. The system also prioritizes these improvement directions based on the severity of the weaknesses and their overall impact on learning and playing. These personalized improvement directions and priority sequences are presented to the learner in a visual format. Based on the system's suggestions, the learner can develop a targeted practice plan to gradually improve their performance skills.
[0056] In this embodiment, by using a deep learning algorithm to classify playing styles and weaknesses, the system can provide learners with personalized improvement suggestions to meet the needs of different learners. This personalized teaching method helps stimulate learners' interest and enthusiasm in learning, improving teaching effectiveness. The system can accurately identify and classify learners' weaknesses, helping learners to clearly understand their shortcomings, helping them to develop targeted practice plans, and improving practice efficiency. Based on the severity of the weaknesses and the impact on the overall performance, the system determines the priority sequence of improvement directions, helps learners to reasonably arrange practice time, and helps learners achieve the greatest progress within a limited time.
[0057] In some embodiments, in step S102, the denoising and segmenting of the sound data using signal processing technology, and the extraction of finger keystroke timing features of the action data using a motion recognition algorithm to obtain a sound feature set and a motion feature set specifically include:
[0058] Performing a joint noise reduction process on the sound data using a spectral subtraction method and a wavelet threshold method to obtain first sound data;
[0059] Performing spectrum analysis on the first sound data using short-time Fourier transform to extract spectrum features and obtain a sound feature set;
[0060] Using the HRNet network to detect the hand key points of the motion data, marking the coordinates of the hand key points, and obtaining the first motion data;
[0061] Based on the first action data, a key action determination model is used to detect the key press interval, duration and key force value to obtain an action feature set.
[0062] In this embodiment, spectral subtraction is used to perform noise reduction processing on sound data to obtain first noise reduction data. The first noise reduction data is further processed by a wavelet threshold method to obtain first sound data. Short-time Fourier transform is used to perform spectral analysis on the first sound data to extract spectral features to obtain a preliminary feature set. For the preliminary feature set, if the feature value exceeds a preset threshold, mean filtering and smoothing are performed to obtain a smoothed feature set. Based on the smoothed feature set, principal component analysis is used to reduce the feature dimension to obtain a reduced-dimensional feature set. Using the reduced-dimensional feature set, a clustering algorithm is used to group the sound features to obtain a grouped feature set. For the grouped feature set, the central feature of each group is obtained to obtain a sound feature set.
[0063] The HRNet network is used to perform detection on the motion data, mark the coordinates of the key points of the hand, and obtain the first motion data. Based on the first motion data, the key action judgment model is used to calculate the key interval, duration, and force value to obtain the motion feature set. According to the motion feature set, the key interval features are extracted to determine whether there is an abnormal interval and obtain the interval judgment result. The duration distribution is analyzed through the motion feature set. If the duration exceeds the preset threshold, the abnormal time period is determined. The force value in the motion feature set is used to obtain the force value change trend, and the motion stability is judged through trend analysis. Based on the interval judgment result and the abnormal time period, the force value change trend is integrated, and the motion consistency index is obtained through logical operation. Based on the motion consistency index, the preset classification rules are used to judge the integrity of the motion feature set to obtain the motion feature set.
[0064] For example, by performing STFT on a 10-second speech, a spectrum graph that changes over time can be obtained, and then sound features such as pitch and formant can be extracted. In terms of motion data processing, the HRNet network is used to detect hand key points, such as locating the coordinates of key points such as finger joints and fingertips. For example, when analyzing hand movements in a video, HRNet can output the coordinates of 21 hand key points in each frame to form a hand posture sequence. Based on the detected hand key points, the key action judgment model is used to analyze the key pressing behavior. Specifically, the key pressing interval can be determined by calculating the displacement of key points between adjacent frames, the key pressing duration can be obtained by tracking the duration of the pressed state, and the key force value can be estimated based on the degree of finger bending. It can be understood that this method can accurately capture the detailed features of the key pressing action, which is helpful for further gesture recognition or human-computer interaction applications. Through the joint analysis of the above-mentioned sound and motion data, more comprehensive and accurate information acquisition can be achieved, providing strong support for subsequent recognition or interaction tasks.
[0065] In some embodiments, in step S103 above, the bimodal Transformer evaluation network includes a feature embedding layer, a multi-head cross attention layer, a feature fusion layer, and a multidimensional evaluation layer; the bimodal Transformer evaluation network is used to perform multidimensional analysis based on the sound feature set and the action feature set to obtain the first evaluation score, specifically including:
[0066] The MFCC feature vector of the sound feature set and the action timing matrix and hand key point coordinates of the action feature set are input into the feature embedding layer. A multi-head cross attention layer is used to calculate the association between the action and the audio and the response of the audio to the action, thereby obtaining the first sound feature set and the first action feature set.
[0067] Using a feature fusion layer to perform layer normalization and time dimension pooling on the first sound feature set and the first action feature set to form a fused feature;
[0068] A multidimensional evaluation layer is used to perform a multidimensional analysis on the fusion features to obtain explanations and contribution weights of multiple evaluation dimensions, wherein the evaluation dimensions include rhythm accuracy, force control, note accuracy, hand shape standardization and expressiveness.
[0069] In this embodiment, the MFCC vector, time series matrix and key point coordinates are processed by the feature embedding layer to obtain preliminary embedded features. The preliminary embedded features are calculated using a multi-head cross attention layer to obtain the correlation between the action and the audio and the responsiveness of the audio to the action. Significant features are extracted from the correlation and responsiveness to obtain a first sound feature set and a first action feature set. The first sound feature set and the first action feature set are layer-normalized using a feature fusion layer to generate a first intermediate feature set. Time dimension pooling is performed on the first intermediate feature set to obtain fused features. The fused features are multidimensionally analyzed by the multidimensional evaluation layer to obtain preliminary weight values for each evaluation dimension. If the preliminary weight value exceeds a preset threshold, the weight value is classified and adjusted using a support vector machine algorithm to determine the final weight value. The fused features are weighted according to the final weight value to obtain evaluation results for rhythm accuracy, dynamics control, note accuracy and expressiveness. The evaluation results are dimensionally interpreted using preset mapping rules to obtain a contribution interpretation for each evaluation dimension.
[0070] For example, for a 5-second audio clip of a piano performance, its MFCC feature vectors can be extracted to reflect the spectral characteristics of the sound. For the corresponding performance video, a motion timing matrix can be generated to record the hand movement trajectory. This is then combined with the coordinates of the 21 hand key points detected by HRNet to form a motion feature set. In one possible implementation, the feature embedding layer maps this heterogeneous data to a common dimension through linear transformation, for example, expanding the MFCC vectors from 13 to 64 dimensions and adjusting the motion timing matrix to a 64-dimensional vector for ease of subsequent processing. This approach preserves the details of the original features and lays the foundation for cross-modal analysis. It should be noted that the multi-head cross-attention layer is the core component of the network, capturing the interplay between sound and movement. Specifically, imagine a pianist playing a fast scale. The rapid finger tapping of the keys is highly correlated with the rhythm of the notes. The multi-head cross-attention layer, through parallel computation using eight attention heads, analyzes how the displacement speed of the finger key points affects the rhythmic accuracy of the notes, while also detecting whether the pitch changes in the audio are consistent with the hand gestures. For example, when a performer strikes five notes in succession, the attention mechanism can output a score for the correlation between the action and the audio, such as 0.85, indicating high synchronization; the score for the response of the audio to the action might be 0.78, reflecting the tendency of the sound to drive the hand. This bidirectional analysis helps to reveal the deep connection between the modalities. Preferably, the feature fusion layer integrates the first sound feature set and the first action feature set into a unified representation. The layer normalization operation can eliminate scale differences between the data, such as normalizing the numerical range of the sound features from -1 to 1 and mapping the action features from 0 to 100 to the same range. Subsequently, through time dimension pooling, the feature sequence of the 5-second performance is compressed into a single vector, for example, from a 64x50 matrix to a 64-dimensional vector. This fusion method can smooth out noise interference and highlight the overall features. It can be understood that the design of the multidimensional evaluation layer is intended to analyze the expressiveness of the fused features from multiple perspectives. For example, in analyzing a piano performance video, rhythmic accuracy can be assessed by comparing the time intervals between keypoint displacements and the degree of alignment with the audio beat. Assuming an ideal beat of 0.5 seconds per note, the actual detected beat averaged 0.52 seconds, with a deviation of only 4%. Dynamic control can be inferred from changes in finger bending angle; for example, a change from 30 degrees to 45 degrees corresponds to a 10dB increase in volume. Note accuracy relies on the matching of MFCC features with standard pitch. Hand shape standardization is determined based on keypoint coordinates, such as whether the distance between the thumb and index finger remains within 2 cm. Expressiveness can be comprehensively scored by combining the amplitude of rhythm and dynamic fluctuations. The contribution weight of each dimension can be dynamically adjusted, for example, 30% for rhythmic accuracy and 25% for expressiveness, reflecting the evaluation focus. This multidimensional analysis provides detailed feedback to performers and data support for automated teaching systems, demonstrating significant application value.
[0071] In some embodiments, in step S104, the sentiment analysis model includes a local temporal coding layer, a global attention pooling layer, a context-aware fusion layer, a sentiment evaluation layer, and a sentiment visualization layer; and analyzing and processing the sentiment expression data according to the preset sentiment analysis model to obtain the sentiment analysis results specifically includes:
[0072] A bidirectional LSTM network of a local temporal layer is used to extract short-term dependency features of the emotion expression data to obtain a first emotion feature;
[0073] Using a multi-head self-attention mechanism of a global attention pooling layer to calculate each feature weight of the first emotion feature, to obtain a second emotion feature;
[0074] globally weighting the second emotion feature using a context-aware fusion layer to form a third emotion feature;
[0075] Using the emotion evaluation layer to perform emotion dimension mapping and emotion intensity preset on the third emotion feature to obtain an emotion analysis result;
[0076] The emotion analysis results are projected onto a 2D plane using an emotion visualization layer, and a dynamic emotion intensity layer is superimposed on the performance video screen.
[0077] In this embodiment, a bidirectional LSTM network is used to process emotion expression data and extract short-term dependency feature sequences. Based on the extracted short-term dependency feature sequences, the hidden state vector is calculated for each time step. A multi-head self-attention mechanism is used to weight the hidden state vectors to obtain a weighted feature representation. A global pooling operation is performed on the weighted feature representation to generate a fixed-dimensional feature vector. A context-aware module is used to further semantically enhance and refine the feature vector. An emotion assessment layer is used to map the emotion feature data to obtain emotion dimension data. The intensity of the emotion feature data is calculated using preset parameters to obtain emotion intensity data. If both the emotion dimension data and the emotion intensity data exceed a preset threshold, the emotion dimension data and the emotion intensity data are integrated using a weighted fusion algorithm to obtain emotion analysis result data. A visualization layer is used to perform a projection transformation on the emotion analysis result data to obtain 2D plane coordinate data. The 2D plane coordinate data is superimposed with the video image data using image processing tools to obtain superimposed image data. The dynamic intensity in the superimposed image data is used to obtain the time series change trend and obtain dynamic intensity change data. The dynamic intensity change data is processed by a smoothing algorithm to obtain the smoothed dynamic emotion intensity layer data.
[0078] For example, in a sentiment analysis model, the local temporal encoding layer uses a bidirectional LSTM network to capture the short-term dependency characteristics of emotional expression data in a performance video. The advantage of a bidirectional LSTM is that it can simultaneously consider information from previous and subsequent time steps. For example, in a piano performance, the performer plays a slow melody for three seconds, with gentle finger movements in the first second and gradually accelerating in the last two seconds. This short-term dependency can be reflected as a trend in finger keystroke speed increasing from once per second to twice per second. Combined with the change in audio volume from low to high, this creates a primary emotional signature, reflecting the transition from calm to passionate emotion. The global attention pooling layer uses a multi-head self-attention mechanism to assign weights to the primary emotional signature. Assuming there are five key moments in a performance video, the multi-head mechanism might analyze the performance video in parallel using four attention heads, determining that the rapid finger movements in the third second contribute more to the overall emotion, with a weight of 0.35, while the slower movements in the first second only have a weight of 0.15. This allows the secondary emotional signature to highlight the emotional expression at these key moments, facilitating subsequent analysis. It's important to note that the context-aware fusion layer globally weights the second emotional feature to form a more unified third emotional feature. Specifically, during a climax, a performer's finger movements might increase from 5 cm to 10 cm, and the audio tempo might accelerate from 2 beats per second to 3 beats per second. The fusion layer integrates this information and uses global weighting to emphasize the emotional intensity of the climax, ensuring feature consistency. The emotion assessment layer maps the third emotional feature to an emotional dimension and pre-sets its intensity. For example, when playing a sad piece, finger movements are slow, with an average of one key press per second and the volume maintained at a low level, such as 30 decibels. The emotional dimension might be mapped to "sadness" with a preset intensity of 0.7. On the other hand, in an upbeat piece, with a key press frequency of 3 per second and a volume of 60 decibels, the emotional dimension becomes "joyful" with an intensity of 0.9. This mapping provides the basis for emotion quantification. As you can understand, the emotion visualization layer projects the analysis results onto a 2D plane and overlays a dynamic emotion intensity layer. In one embodiment, on the performance video screen, sad passages may be displayed with a blue gradient bar, whose width increases from 2 pixels to 5 pixels with intensity and extends dynamically over time; happy passages use an orange bar with a width of up to 8 pixels. This visualization intuitively shows emotional changes, helping performers understand the dynamic trends of their own expressions, while providing an immersive experience for the audience. For example, when analyzing a 5-second performance clip, the local temporal coding layer may capture the emotional turning point of the finger suddenly pausing for 0.5 seconds in the 2nd second. The global attention pooling layer increases the weight of this moment to 0.4. The context-aware fusion layer integrates it to highlight its impact. The emotional evaluation layer determines it as a "nervous" emotion with an intensity of 0.6. The visualization layer superimposes a briefly flashing purple mark on the screen. This multi-layer collaboration ensures the comprehensiveness and intuitiveness of emotional analysis, providing a detailed reference for performance improvement.
[0079] In some embodiments, in step S105, the process of using a deep learning algorithm to classify the learner's playing style and weaknesses based on the first sentiment analysis result and the first evaluation score, and determining personalized improvement directions and priority sequences, specifically includes:
[0080] Correlating the first sentiment analysis result and the first evaluation score to form a sentiment and skill correlation matrix;
[0081] Based on the emotion and skill association matrix, a sampling graph network algorithm constructs a dynamic relationship graph, and calculates the attention coefficient between every two nodes of the dynamic relationship graph through a neighbor aggregation formula;
[0082] Based on the dynamic relationship graph, a gated recurrent unit is used to analyze the learner's weak links and determine a weight coefficient for each weak link;
[0083] Based on the dynamic relationship graph, a variational autoencoder is used to construct a style latent space and determine a style latent vector;
[0084] The learner's weak links and the corresponding weight coefficients and style latent vectors are taken as the state space, the improvement methods for the learners are taken as the action space, the skill improvement speed, emotional expression richness and practice time efficiency are taken as the reward function, and the reinforcement learning algorithm is used to determine the personalized improvement direction and priority sequence.
[0085] In this embodiment, correlation information is extracted through sentiment analysis and evaluation scores to obtain preliminary mapping data of sentiment and skills. The preliminary mapping data is processed by constructing a matrix method to generate a sentiment and skill correlation matrix. Structural information is obtained from the sentiment and skill correlation matrix, and a dynamic relationship graph is generated using a graph network algorithm. For the connection between nodes in the dynamic relationship graph, the attention coefficient is calculated using the neighbor aggregation formula to determine the relationship strength between the nodes. If the attention coefficient exceeds a preset threshold, the dynamic relationship between the corresponding nodes is retained to obtain an optimized relationship graph. Based on the optimized relationship graph, the key node features are obtained by extracting information to determine the deep correlation pattern between sentiment and skills. Based on the deep correlation pattern, the dynamic relationship graph is updated using a graph network algorithm to obtain the final sentiment and skill mapping result.
[0086] A gated recurrent unit processes the dynamic relationship graph to produce a relationship map. Weak links are identified within the relationship map, and analysis yields the link identification results. A computational process processes the identified link results to obtain preliminary weights for the weak links. The gated recurrent unit adjusts the preliminary weights to determine the output weights. If the output weight exceeds a preset threshold, the computational process updates the analysis results. Based on the analysis results, dynamic relationships are extracted to produce an updated relationship map. The updated relationship map is used to determine the changing trends of the weak links and determine the final weight coefficients.
[0087] A relationship graph is extracted from the dynamic relationships to obtain graph data. A variational autoencoder is used to encode the graph data to obtain a vector representation. A latent space is constructed from the vector representation to determine the style latent space. A spatial mapping is mapped based on the style latent space to obtain the style latent. If the style latent space is consistent with the latent vector, the spatial mapping is adjusted using the vector representation. A clustering algorithm is used to determine the distribution characteristics of the dynamic relationships after the adjustment to determine the change trend. The graph data is updated based on the change trend to determine the final style latent vector.
[0088] The learner's weaknesses, weight coefficients, and style vectors are obtained and combined to form a state space. Through state space analysis, improvement methods are identified and an action space is constructed. For the improvement methods in the action space, the speed of skill improvement, the richness of emotional expression, and the time efficiency of practice are calculated to obtain a reward function. A reinforcement learning algorithm is used to process the state and action spaces. Through iterative optimization of the reward function, personalized improvement directions are determined. A priority sequence is extracted from the improvement directions and a ranking result is determined. Based on the ranking result, the action space is adjusted to obtain an optimized improvement method. The optimized improvement method is then used to update the state space, obtaining a new weak link and style vector.
[0089] For example, in constructing an emotion-skill correlation matrix, the first emotion analysis result might be derived from the emotional expressions in a learner's performance video, such as a five-second continuous bowing variation in a piano solo. The first assessment score is based on the proficiency of playing techniques, such as intonation and rhythmic stability. Imagine that the emotion analysis identifies that the performer displays "softness" emotion in the first two seconds, scoring 0.8, and then shifts to "passion" in the next three seconds, scoring 0.7. The skill assessment then indicates 90% intonation and 85% rhythmic stability. Through correlation construction, the matrix associates "softness" with high intonation scores and "passion" with medium-to-high rhythmic stability, intuitively reflecting the coupling relationship between emotion and skill. This matrix provides basic data support for subsequent analysis. In one possible implementation, based on the emotion-skill correlation matrix, when generating a dynamic relationship graph using a sampling graph network algorithm, emotions and skills can be treated as nodes, for example, "softness" and "pitch" as a node pair. When calculating the attention coefficient using the neighbor aggregation formula, it is assumed that the connection strength between "softness" and "pitch" is high due to frequent co-occurrence, resulting in an attention coefficient of 0.6; while the coefficient between "passion" and "rhythmic stability" is 0.4 due to some errors. Specifically, the performer's tempo occasionally tends to be too fast during passionate passages. By aggregating neighbor information, the graph network amplifies the significance of this weakness, helping to precisely pinpoint the problem. It is important to note that the gated recurrent unit focuses on skill fluctuations over time when analyzing weak links. For example, a learner's pitch drops from 90% to 80% during a performance, possibly due to insufficient finger velocity control. The gated recurrent unit captures this trend and identifies "velocity control" as a weak link, assigning a weight of 0.5, while rhythmic stability receives a weight of 0.3. This weighting reflects the severity of the problem, allowing for prioritization of critical deficiencies during subsequent improvements. When constructing the style latent space, the variational autoencoder can extract latent features from performance data. Suppose a learner prefers slow bowing in gentle passages and fast vibrato in passionate passages. A stylistic latent space generates a vector, such as [0.7, 0.3], indicating a dominant gentle style, and [0.4, 0.6], indicating a predominant passionate style. This vector quantifies stylistic preferences and provides a basis for personalized recommendations. The reinforcement learning algorithm defines weak areas, such as dynamic control, with a weight of 0.5, and the style vector [0.7, 0.3] as the state space, and improvement methods, such as "increasing slow practice" or "strengthening technical training," as the action space. The reward function integrates the speed of skill improvement, the richness of emotional expression, and the time efficiency of practice. For example, after slow practice, a learner's pitch accuracy improves to 95%, their emotional expression shifts from monotonous to richly layered, and their practice efficiency increases by 20%. The algorithm iterates based on this information and prioritizes dynamic control first, then optimizes rhythm. This approach ensures that the direction of improvement aligns with individual needs. Clearly, this process, through the deep connection between emotion and skill, fosters comprehensive improvement for learners.For example, a pianist's pitch was stable during soft passages but lacked emotional depth. Reinforcement learning recommended extending practice time and incorporating expression training, ultimately increasing the richness of their emotional expression by 30%, achieving simultaneous improvement in both skill and emotion. This multi-dimensional analysis and optimization significantly enhances learning outcomes.
[0090] In some embodiments, in steps S101 to S105 above, the method further includes:
[0091] Set the progress index formula based on the learner's playing accuracy, progress speed and emotional expression;
[0092] Obtaining historical progress data of the learner, inputting the historical progress data into a hidden Markov model, and performing modeling and prediction in combination with the progress index formula to obtain predicted progress data of the learner;
[0093] Detecting the learner's cognitive load estimate and physiological fatigue value according to a preset learning stress assessment model, comparing the cognitive load estimate and the physiological fatigue value with their respective baseline values to obtain learning stress assessment data;
[0094] The difficulty of the piano practice tutoring program for the learner is dynamically adjusted according to the predicted progress data and the learning pressure assessment data.
[0095] In this embodiment, a progress index is calculated based on performance accuracy, progress speed, and emotional expression, and an initial progress index value is obtained using a formula setting method. Learner historical data is obtained, and features of performance accuracy, progress speed, and emotional expression are extracted from the historical data to obtain a structured historical data set. The structured historical data set is processed using a hidden Markov model, and combined with the initial progress index value, a state sequence after model training is obtained. Based on the state sequence and the formula setting method, a dynamic progress index is calculated to obtain an updated progress index value. If the updated progress index value exceeds a preset threshold, the future state sequence is predicted through model processing to obtain predicted progress data.
[0096] Using a preset detection model, cognitive load estimates and physiological fatigue values are obtained from learner data to generate initial detection results. The initial detection results are compared against preset baseline values to determine learning stress assessment data. The predicted progress data is processed using a random forest algorithm to determine progress trends. The adjustment range for practice difficulty is determined based on the learning stress assessment data and progress trends. If the adjustment range exceeds the preset threshold, the tutoring plan is optimized using a linear regression algorithm to obtain adjusted plan parameters. The piano practice tutoring plan is updated based on the adjusted plan parameters to obtain personalized practice content for the learner. The applicability of the adjustment plan is verified using real-time learner feedback data to obtain an optimized tutoring plan.
[0097] For example, when setting a progress index formula, a learner's performance accuracy, progress rate, and emotional expression can be combined to reflect their overall ability. Assume that performance accuracy refers to the percentage of correct notes, progress rate refers to the amount of practice completed per unit time, and emotional expression refers to the clarity of emotional transmission during performance. A formula can be envisioned that combines these three factors according to weights, such as 40% for accuracy, 30% for speed, and 30% for emotional expression. In one possible implementation, if a learner's accuracy is 85%, their progress rate is three pieces per week, and their emotional expression score is 0.7, then the progress index can reflect their current level and provide a basis for subsequent predictions. Specifically, when historical progress data is input into a hidden Markov model, the model predicts future performance by analyzing state transitions in the time series. Suppose that over the past six weeks, a learner's accuracy has gradually increased from 80% to 85%, their speed has increased from two pieces per week to three pieces per week, and their emotional expression has increased from 0.5 to 0.7. The model identified a trend of improvement, predicting that over the next four weeks, the accuracy could reach 88%, the tempo could stabilize at three pieces, and the emotional expressiveness could rise to 0.75. This prediction helps plan practice content ahead of time and avoid blind adjustments. The learning stress assessment model measures cognitive load and physiological fatigue by estimating these values based on the learner's reaction time and heart rate during practice. Assuming a baseline reaction time of 0.5 seconds and a heart rate of 80 beats per minute, while the learner's actual values are 0.7 seconds and 90 beats per minute, these deviations indicate slightly higher cognitive load and increased physiological fatigue. Importantly, this assessment data can reveal whether a learner's efficiency is declining due to excessive practice intensity, thus providing a basis for adjustment plans. Dynamic adjustments to piano tutoring plans combine predicted progress data with stress assessment data to determine difficulty. For example, if the prediction indicates that a learner's ability is improving but stress is high, the difficulty of the new piece could be reduced from intermediate to elementary, while increasing repetition of familiar pieces. Understandably, this adjustment ensures skill consolidation while preventing regression due to excessive workload. For example, a learner originally planned to learn a complex piece, but due to excessive fatigue, switched to reviewing basic scales. The following week, their accuracy rate increased from 85% to 90%, and their emotional expression became more stable. In one possible implementation, if predicted data indicates rapid improvement and manageable stress, the difficulty level can be appropriately increased, such as by introducing pieces requiring more fingering variations. This approach ensures that challenge aligns with ability and unleashes potential. For example, after basic practice, a learner attempted a more difficult piece, maintaining their accuracy at 83% and improving their emotional expression to 0.8, demonstrating the effectiveness of the adjustment. This multi-dimensional, dynamic adjustment significantly enhances the learning experience and outcomes.
[0098] In some embodiments, in steps S101 to S105 above, the method further includes:
[0099] Acquiring acoustic environment data, and extracting noise energy and music spectrum energy based on the acoustic environment data to form acoustic features;
[0100] Acquiring visual environment data, and obtaining a visual emotion vector by performing lighting condition detection and color emotion mapping on the visual environment data;
[0101] Using a fine-grained classification model based on ResNet-50 to perform spatial style classification on the visual emotion vector to obtain visual features;
[0102] Based on the acoustic features and the visual features, combined with the user's current skill vector and the required skill level of the repertoire, suitable piano repertoires are dynamically recommended to the learner who is playing the piano.
[0103] In this embodiment, raw data from an acoustic environment is acquired through sensor collection to obtain environmental data. The noise portion is separated from the environmental data, and the noise energy is extracted using short-time Fourier transform. Spectral energy is obtained for the music portion of the environmental data through spectral analysis. If the noise energy exceeds a preset threshold, the environmental data is subjected to noise separation processing to obtain purified data. Acoustic features are formed using principal component analysis based on the purified data and spectral energy. The changing trend in the acoustic features is acquired to determine whether the feature formation is complete. Based on the changing trend, cluster analysis is used to determine the classification result of the acoustic features.
[0104] Visual environment data is acquired by collecting raw images and lighting information through sensors to obtain an initial environment dataset. The initial environment dataset is then subjected to lighting condition detection and brightness analysis tools to separate lighting parameters and obtain lighting adjustment data. A color-emotion mapping method is used to calculate color distribution and emotion weights for the lighting adjustment data to obtain a visual emotion vector. A ResNet-50 model is used to perform spatial style classification on the visual emotion vector, extracting deep features and obtaining a visual style feature set. If the visual style feature set contains multiple styles, a clustering algorithm is used to group the features and determine the primary style category. Based on the primary style category and combined with environmental analysis attributes, spatial layout information is obtained to obtain a final style description. Feature matching techniques are used to determine the similarity between the final style description and the preset style library to obtain the classification result.
[0105] A preset algorithm is used to fuse acoustic and visual features to determine a comprehensive feature set. The user's skill vector is matched against this comprehensive feature set to determine the current skill status. The correspondence between skill levels and repertoire requirements is obtained from the piano repertoire database to determine a candidate repertoire set. If the current skill status matches the repertoire requirements, suitable repertoires are selected from the candidate repertoire set to generate recommendations. Based on the recommendation results, the user's skill vector is updated to generate a dynamically adjusted repertoire list.
[0106] For example, when acquiring acoustic environment data, a microphone can be used to capture ambient audio signals and analyze the distribution of noise energy and the spectral energy of the music. For example, suppose a learner is practicing in a living room with a faint background sound of a television and a running fan. The noise energy is likely concentrated in the low-frequency band, while the spectral energy of the piano music is more distributed in the mid- and high-frequency bands. By separating these two energy components, an acoustic signature can be formed, reflecting the degree to which the environment interferes with the performance. For example, during a particular practice session, the noise energy may account for 20% of the total sound energy, the music spectral energy may account for 70%, and the remainder may be spurious signals. This indicates a relatively quiet environment suitable for focused practice. In one possible implementation, visual environment data can be acquired by recording the lighting and color information of the practice area through a camera. Lighting condition detection can determine whether the brightness is uniform, while color emotion mapping converts the ambient color tone into an emotional orientation. For example, suppose a learner is practicing at dusk with warm yellow light inside the room and bluish-gray natural light outside the window. The lighting detection indicates moderate but slightly dim brightness, and the color emotion mapping generates an emotional vector indicating warmth and relaxation. This visual emotional vector can indicate whether the emotional atmosphere of the current environment is conducive to emotional expression. When a fine-grained classification model based on ResNet-50 classifies visual emotion vectors into spatial styles, it can categorize environments into simple, warm, or cluttered categories. If a learner's practice area has neatly arranged bookshelves and soft lighting, the model might classify it as "warm" and output corresponding visual features. This feature not only describes the spatial style but also indirectly reflects the learner's psychological comfort. For example, a warm environment may make a learner more engaged when playing lyrical music. It should be noted that when recommending music based on acoustic and visual features combined with the user's skill vector, the learner's actual ability and the difficulty of the music can be considered. Suppose a learner's current skill vector indicates a rhythm mastery of 0.8 and a fingering flexibility of 0.6. A piece requires a rhythm of 0.7 and a fingering of 0.5. If the acoustic features indicate a quiet environment and the visual features suggest a warm atmosphere, then this piece is a suitable recommendation. Conversely, if the noise energy is too high, a piece with a lower rhythmic requirement may be recommended to reduce the impact of interference. In one embodiment, dynamic music recommendations can also take the learner's practice goals into consideration. If the goal is to improve fingering flexibility, the system may prioritize recommending pieces with more fast scales, and when the visual features indicate a cluttered environment, it may suggest adjusting the space to improve concentration. For example, after the learner cleans up the desk, the visual features become simpler, the recommended pieces become slightly more difficult, and the practice effect is better. Preferably, the recommendation system can also adjust its strategy based on long-term data. If the acoustic features show that the noise level is always high during a certain period of time, it may be recommended to switch to a quiet period for practice; if the visual features continue to be warm, then pieces that focus more on emotional expression are recommended. This approach ensures that the recommendations are highly consistent with the learner's actual situation. It is understandable that this combination of multi-dimensional features can more accurately match the piece of music with the learner's state.For example, a learner was recommended an etude with a simple melody in a quiet and softly lit environment. After playing it, the learner's rhythm mastery improved to 0.85 and his fingering flexibility became smoother, proving that the recommendation met the needs and was effective.
[0107] Reference Figure 2 An embodiment of the present invention provides an AI-based piano practice auxiliary system 2, wherein the system 2 specifically includes:
[0108] The first auxiliary module 201 is used to obtain the sound data and movement data of the learner when playing the piano, and align and match the sound data and movement data;
[0109] The second auxiliary module 202 is configured to perform denoising and segmentation on the sound data using signal processing technology, and extract finger keystroke timing features of the action data using an action recognition algorithm to obtain a sound feature set and an action feature set;
[0110] A third auxiliary module 203 is configured to perform a multi-dimensional analysis based on the sound feature set and the action feature set using a bimodal Transformer evaluation network to obtain a first evaluation score;
[0111] The fourth auxiliary module 204 is used to obtain the learner's emotional expression data when playing the piano, and analyze and process the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result;
[0112] The fifth auxiliary module 205 is used to classify the learner's playing style and weaknesses using a deep learning algorithm based on the first sentiment analysis result and the first evaluation score, and determine personalized improvement directions and priority sequences.
[0113] It is understandable that if Figure 1 The contents of the embodiment of the AI-based piano practice auxiliary method shown in the figure are applicable to the embodiment of the present AI-based piano practice auxiliary system. The functions specifically implemented by the embodiment of the AI-based piano practice auxiliary system are similar to those in the embodiment of the present AI-based piano practice auxiliary system. Figure 1 The embodiment of the piano practice auxiliary method based on AI shown in FIG. Figure 1 The beneficial effects achieved by the AI-based piano practice assistance method embodiment shown are also the same.
[0114] It should be noted that the information interaction, execution process and other contents between the above-mentioned systems are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0115] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0116] Reference Figure 3 An embodiment of the present invention further provides a computer device 3, comprising: a memory 302 and a processor 301 and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, the AI-based piano practice auxiliary method as described in any one of the above methods is implemented.
[0117] The computer device 3 may be a desktop computer, a notebook computer, a PDA, a cloud server or other computing devices. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 This is merely an example of the computer device 3 and does not constitute a limitation on the computer device 3 . The computer device 3 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device 3 may also include input and output devices, network access devices, etc.
[0118] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0119] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 3. Furthermore, the memory 302 may include both an internal storage unit of the computer device 3 and an external storage device. The memory 302 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 302 may also be used to temporarily store data that has been output or is about to be output.
[0120] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the AI-based piano practice auxiliary method as described in any one of the above methods.
[0121] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process of the above-mentioned method embodiment by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, mobile hard drive, magnetic disk, or optical disk. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0122] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0123] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0124] In the embodiments disclosed in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0125] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. An AI-based piano practice assisting method, characterized in that: The method specifically includes: Acquire the learner's sound data and movement data when playing the piano, and align and match the sound data and movement data; De-noising and segmenting the sound data using signal processing technology, and extracting finger keystroke timing features of the action data using an action recognition algorithm to obtain a sound feature set and an action feature set; Based on the sound feature set and the action feature set, a bimodal Transformer evaluation network is used to perform multidimensional analysis to obtain the first evaluation score; Acquiring emotional expression data of the learner during piano playing, analyzing and processing the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result; Based on the first sentiment analysis result and the first evaluation score, a deep learning algorithm is used to classify the learner's playing style and weaknesses, and to determine personalized improvement directions and priority sequences.
2. The method according to claim 1, characterized in that The method of using signal processing technology to denoise and segment the sound data, and using an action recognition algorithm to extract the finger keystroke timing features of the action data to obtain a sound feature set and an action feature set specifically includes: Performing a joint noise reduction process on the sound data using a spectral subtraction method and a wavelet threshold method to obtain first sound data; Performing spectrum analysis on the first sound data using short-time Fourier transform to extract spectrum features and obtain a sound feature set; Using the HRNet network to detect the hand key points of the motion data, marking the coordinates of the hand key points, and obtaining the first motion data; Based on the first action data, a key action determination model is used to detect the key press interval, duration and key force value to obtain an action feature set.
3. The method according to claim 1, characterized in that The bimodal Transformer evaluation network includes a feature embedding layer, a multi-head cross attention layer, a feature fusion layer, and a multidimensional evaluation layer. The bimodal Transformer evaluation network is used to perform multidimensional analysis based on the sound feature set and the action feature set to obtain a first evaluation score, specifically including: The MFCC feature vector of the sound feature set and the action timing matrix and hand key point coordinates of the action feature set are input into the feature embedding layer. A multi-head cross attention layer is used to calculate the association between the action and the audio and the response of the audio to the action, thereby obtaining the first sound feature set and the first action feature set. Using a feature fusion layer to perform layer normalization and time dimension pooling on the first sound feature set and the first action feature set to form a fused feature; A multidimensional evaluation layer is used to perform a multidimensional analysis on the fusion features to obtain explanations and contribution weights of multiple evaluation dimensions, wherein the evaluation dimensions include rhythm accuracy, force control, note accuracy, hand shape standardization and expressiveness.
4. The method according to claim 1, wherein The sentiment analysis model includes a local temporal coding layer, a global attention pooling layer, a context-aware fusion layer, a sentiment evaluation layer, and a sentiment visualization layer; the sentiment expression data is analyzed and processed according to the preset sentiment analysis model to obtain the sentiment analysis results, specifically including: A bidirectional LSTM network of a local temporal layer is used to extract short-term dependency features of the emotion expression data to obtain a first emotion feature; Using a multi-head self-attention mechanism of a global attention pooling layer to calculate each feature weight of the first emotion feature, to obtain a second emotion feature; globally weighting the second emotion feature using a context-aware fusion layer to form a third emotion feature; Using the emotion evaluation layer to perform emotion dimension mapping and emotion intensity preset on the third emotion feature to obtain an emotion analysis result; The emotion analysis results are projected onto a 2D plane using an emotion visualization layer, and a dynamic emotion intensity layer is superimposed on the performance video screen.
5. The method according to claim 1, wherein The method of using a deep learning algorithm to classify the learner's playing style and weaknesses based on the first sentiment analysis result and the first evaluation score, and determining personalized improvement directions and priority sequences, specifically includes: Correlating the first sentiment analysis result and the first evaluation score to form a sentiment and skill correlation matrix; Based on the emotion and skill association matrix, a sampling graph network algorithm constructs a dynamic relationship graph, and calculates the attention coefficient between every two nodes of the dynamic relationship graph through a neighbor aggregation formula; Based on the dynamic relationship graph, a gated recurrent unit is used to analyze the learner's weak links and determine a weight coefficient for each weak link; Based on the dynamic relationship graph, a variational autoencoder is used to construct a style latent space and determine a style latent vector; The learner's weak links and the corresponding weight coefficients and style latent vectors are taken as the state space, the improvement methods for the learners are taken as the action space, the skill improvement speed, emotional expression richness and practice time efficiency are taken as the reward function, and the reinforcement learning algorithm is used to determine the personalized improvement direction and priority sequence.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Set the progress index formula based on the learner's playing accuracy, progress speed and emotional expression; Obtaining historical progress data of the learner, inputting the historical progress data into a hidden Markov model, and performing modeling and prediction in combination with the progress index formula to obtain predicted progress data of the learner; Detecting the learner's cognitive load estimate and physiological fatigue value according to a preset learning stress assessment model, comparing the cognitive load estimate and the physiological fatigue value with their respective baseline values to obtain learning stress assessment data; The difficulty of the piano practice tutoring program for the learner is dynamically adjusted according to the predicted progress data and the learning pressure assessment data.
7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Acquiring acoustic environment data, and extracting noise energy and music spectrum energy based on the acoustic environment data to form acoustic features; Acquiring visual environment data, and obtaining a visual emotion vector by performing lighting condition detection and color emotion mapping on the visual environment data; Using a fine-grained classification model based on ResNet-50 to perform spatial style classification on the visual emotion vector to obtain visual features; Based on the acoustic features and the visual features, combined with the user's current skill vector and the required skill level of the repertoire, suitable piano repertoires are dynamically recommended to the learner who is playing the piano.
8. An AI-based piano practice auxiliary system, characterized in that: The system specifically includes: The first auxiliary module is used to obtain the sound data and movement data of the learner when playing the piano, and align and match the sound data and movement data; The second auxiliary module is configured to perform denoising and segmentation on the sound data using a signal processing technique, and extract finger keystroke timing features of the action data using an action recognition algorithm to obtain a sound feature set and an action feature set; The third auxiliary module is used to perform multi-dimensional analysis based on the sound feature set and the action feature set using a bimodal Transformer evaluation network to obtain a first evaluation score; A fourth auxiliary module is used to obtain emotional expression data of the learner during piano playing, and analyze and process the emotional expression data according to a preset emotional analysis model to obtain an emotional analysis result; The fifth auxiliary module is used to classify the learner's playing style and weak links based on the first sentiment analysis result and the first evaluation score using a deep learning algorithm to determine personalized improvement directions and priority sequences.
9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, which, when executed on the processor, implements the AI-based piano practice assisting method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the AI-based piano practice assisting method as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Guqin playing action and audio synchronous analysis method based on multi-modal fusion
CN120635654A
A guqin playing action and audio synchronization analysis method based on multi-modal fusion
CN120635654B