Intelligent music teaching auxiliary system and method based on motion perception
Through the intelligent music teaching assistance system, students' performance videos are analyzed using action perception technology and OpenPose model, providing comparative feedback with standard actions, solving the problem of lack of instant feedback in traditional teaching, and improving teaching efficiency and learning experience.
Patent Information
- Application Number
- CN202510248382.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional music teaching lacks a real-time feedback mechanism, which makes it difficult for students to quickly correct performance errors, and teachers have strong subjective feedback and lacks objective standards.
Using an intelligent music teaching assistance system based on action perception, students' performance videos are captured through cameras, joint movements are analyzed using OpenPose model, compared with standard action databases, fine-grained coding vectors are generated, and action coherence feedback is provided.
It realizes automated analysis and instant feedback on students' playing movements, improves teaching efficiency and learning experience, and helps students master the correct performance skills faster.
Smart Images

Figure CN120183041A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of music teaching assistance, and more specifically, to an intelligent music teaching assistance system and method based on motion perception. Background Art
[0002] With the popularization of music education globally, how to improve teaching quality and learning efficiency has become an important research topic. Especially in the learning process of musical instrument performance, especially for beginners, correct postures, finger positions, and the fluency of movements are all key factors affecting performance effects. However, due to the lack of an instant feedback mechanism, students often take a long time to realize and correct their mistakes. In addition, different students have differences in physical conditions and comprehension abilities, which requires teaching methods to be highly personalized and adaptable.
[0003] Traditional music teaching mainly relies on teachers' experience and direct observation to guide students' practice. Although this method is effective, it has certain limitations. Specifically, in traditional music teaching, teachers' feedback is often based on personal experience and subjective judgment, which may lead to different views on the same problem by different teachers and lack objective criteria. In addition, without technical assistance, it is difficult for teachers to give students instant and detailed feedback every time they practice, especially for some subtle movement mistakes, which may need to be repeated many times to be discovered and corrected.
[0004] Therefore, an intelligent music teaching assistance system is desired. Summary of the Invention
[0005] This application provides an intelligent music teaching assistance system and method based on motion perception, which can not only improve teaching efficiency, but also help students master correct performance techniques faster and enhance the learning experience.
[0006] According to one aspect of this application, an intelligent music teaching assistance system based on motion perception is provided, including:
[0007] A motion data acquisition module for using a camera to capture a monitoring video of the performance process of a target student object;
[0008] A performance motion analysis module for using the OpenPose model to process the monitoring video of the performance process to obtain a time series of performance joint images;
[0009] A standard performance motion extraction module for extracting a time series of standard music performance joint images from a standard action database;
[0010] A synchronization module, configured to align the time series of the performance joint images and the time series of the standard music performance joint images to obtain the aligned time series of the performance joint images and the time series of the standard music performance joint images;
[0011] A performance action feature extraction module, configured to extract performance action joint features from the aligned time series of the performance joint images and the time series of the standard music performance joint images to obtain the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors;
[0012] A performance action matching and analysis module, configured to calculate the position-by-position differential feature vectors between each pair of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors in the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors to obtain the time queue of the performance action difference fine-grained encoding vectors;
[0013] An evaluation and feedback module, configured to generate and display feedback opinions on the screen based on the time queue of the performance action difference fine-grained encoding vectors, where the feedback opinions are used to indicate whether the action coherence meets a preset standard.
[0014] According to another aspect of the present application, there is provided an intelligent music teaching assistance method based on action perception, including:
[0015] Using a camera to capture a performance process monitoring video of a target student object;
[0016] Using the OpenPose model to process the performance process monitoring video to obtain the time series of the performance joint images;
[0017] Extracting the time series of the standard music performance joint images from a standard action database;
[0018] Aligning the time series of the performance joint images and the time series of the standard music performance joint images to obtain the aligned time series of the performance joint images and the time series of the standard music performance joint images;
[0019] Extracting performance action joint features from the aligned time series of the performance joint images and the time series of the standard music performance joint images to obtain the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors;
[0020] Calculate the time series of the performance joint image feature encoding vectors and the differential feature vectors by position between each corresponding pair of performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors in the time series of the standard music performance joint image feature encoding vectors to obtain the time queue of the performance action difference fine-grained encoding vectors;
[0021] Generate and display feedback on the screen based on the time queue of the performance action difference fine-grained encoding vectors, where the feedback is used to indicate whether the action coherence meets the preset standard.
[0022] An intelligent music teaching assistance system and method based on action perception provided by the present application can automatically analyze students' performance actions using deep learning algorithms based on artificial intelligence and image analysis, compare the action differences with the standard music performance actions of musical instruments, and perform time series analysis to identify problems in students' actions, and provide targeted feedback and improvement suggestions on the coherence of students' performance actions, thereby assisting in intelligent music teaching. This can not only improve teaching efficiency but also help students master correct performance techniques faster and enhance the learning experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings of the embodiments of the present application will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present application and do not limit the present application.
[0024] Figure 1 It is a schematic block diagram of the intelligent music teaching assistance system based on action perception according to an embodiment of the present application.
[0025] Figure 2 It is a schematic diagram of the data flow of the intelligent music teaching assistance system based on action perception according to an embodiment of the present application.
[0026] Figure 3 It is a schematic block diagram of the evaluation and feedback module in the intelligent music teaching assistance system based on action perception according to an embodiment of the present application.
[0027] Figure 4 It is a schematic block diagram of the performance action context association analysis unit in the intelligent music teaching assistance system based on action perception according to an embodiment of the present application.
[0028] Figure 5 It is a schematic block diagram of the causal association topology feature construction subunit in the intelligent music teaching assistance system based on action perception according to an embodiment of the present application.
[0029] Figure 6 It is a schematic flowchart of the intelligent music teaching assistance method based on action perception according to an embodiment of the present application. Detailed implementation manners
[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts also belong to the scope of protection of the present application.
[0031] With the rapid development of artificial intelligence and machine learning technologies, their applications in the education field have become increasingly widespread. In recent years, the development of motion capture and analysis technologies has brought new possibilities to music teaching. Based on this, in the technical solution of the present application, an intelligent music teaching assistance system based on motion perception is proposed, which can use deep learning algorithms based on artificial intelligence and image analysis to automatically analyze students' performance actions, and compare the action differences and perform timing analysis with the standard music performance actions of musical instruments to identify problems existing in students' actions, and provide targeted feedback and improvement suggestions on the coherence of students' performance actions, thereby assisting in intelligent music teaching. This can not only improve teaching efficiency, but also help students master correct performance techniques faster and enhance the learning experience.
[0032] Specifically, in the technical solution of the present application, as Figure 1 and Figure 2As shown, the action perception-based intelligent music teaching assistance system includes: an action data acquisition module 10 for using a camera to capture a monitoring video of the performance process of a target student object; a performance action analysis module 20 for using the OpenPose model to process the monitoring video of the performance process to obtain a time series of performance joint images; a standard performance action extraction module 30 for extracting a time series of standard music performance joint images from a standard action database; a synchronization module 40 for performing alignment processing on the time series of the performance joint images and the time series of the standard music performance joint images to obtain an aligned time series of performance joint images and a time series of standard music performance joint images; a performance action feature extraction module 50 for extracting performance action joint features from the aligned time series of performance joint images and the time series of standard music performance joint images to obtain a time series of performance joint image feature encoding vectors and a time series of standard music performance joint image feature encoding vectors; a performance action matching and analysis module 60 for calculating the position-differential feature vectors between each pair of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors in the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors to obtain a time queue of performance action difference fine-grained encoding vectors; an evaluation and feedback module 70 for generating and displaying feedback opinions on the screen based on the time queue of the performance action difference fine-grained encoding vectors, where the feedback opinions are used to indicate whether the action coherence meets a preset standard.
[0033] Exemplarily, in the action data acquisition module 10, a camera is used to capture a monitoring video of the performance process of a target student object. It should be understood that the video captured by the camera can comprehensively display details such as the body posture, finger positions, and action fluency of the student when playing the instrument. This is particularly important for music teaching because the correct posture and finger positions directly affect the performance effect and learning efficiency. Traditional teaching methods rely on the experience and direct observation of teachers. Although this method is effective, it has subjectivity and limitations. By using a camera to capture videos combined with deep learning algorithms, the performance actions of students can be objectively and quantitatively analyzed, compared with standard actions, and subtle action deviations can be identified and targeted improvement suggestions can be given.
[0034] In one embodiment, first, a suitable camera needs to be set up to ensure that it can cover all the key parts of the student's performance, such as the hand, arm, and body posture changes. The camera should have sufficient resolution and frame rate to ensure that the captured actions are clear and smooth enough. Then, connect the camera to a computer or smart device and record the video through the corresponding software program. During this process, to ensure the validity and consistency of the data, students are usually required to perform according to certain specifications, such as a fixed performance piece or segment.
[0035] Exemplarily, in the performance action analysis module 20, the OpenPose model is used to process the performance process monitoring video to obtain a time series of performance joint images. It should be understood that in the actual process of performance action evaluation and music teaching assistance, the physical conditions and learning abilities of each target student object are different. In order to analyze the unique performance postures and actions of each student and provide more personalized feedback suggestions based on individual differences, it is necessary to capture the human performance action characteristics and posture changes during the performance more accurately and effectively. Based on this, in the technical solution of this application, the OpenPose model is used to process the performance process monitoring video to obtain a time series of performance joint images. In particular, OpenPose is a deep learning-based human pose estimation algorithm that can detect the positions of human key points (such as joints, fingers, etc.) in real time. By processing the monitoring video during the performance, the OpenPose model can extract the position information of each joint of the performer in each frame, and then form a sequence of joint positions that changes over time (i.e., a time series). These data are crucial for understanding the dynamic characteristics of performance actions.
[0036] Exemplarily, in the standard performance action extraction module 30, a time series of standard music performance joint images is extracted from the standard action database. In particular, considering that the time series of the standard music performance joint images contains timing data on standard music performance actions, which are extracted from the performance process of professional performers and represent the correct performance postures and action patterns. Subsequently, by comparing the student's actions with the standard actions, the system can objectively evaluate the student's performance quality and avoid the limitations of relying on the subjective judgment of teachers in traditional teaching.
[0037] Exemplarily, in the synchronization module 40, the time series of the performance joint images and the time series of the standard music performance joint images are aligned to obtain the aligned time series of the performance joint images and the time series of the standard music performance joint images. It should be understood that the music performance actions are also considered to have high temporality and spatiality. The standard action time series not only includes the spatial position information of the joints, but also includes the time evolution process of the actions. In order to compare the performance actions during the performance process with the standard music performance actions in terms of features, it is necessary to ensure the consistency of the student actions and the standard actions in the time dimension, so as to more accurately analyze the coherence and rhythm of the actions. Based on this, in the technical solution of the present application, the time series of the performance joint images and the time series of the standard music performance joint images are further aligned to obtain the aligned time series of the performance joint images and the time series of the standard music performance joint images. It is worth mentioning that the alignment process can eliminate the offset on the time axis (such as the difference in performance speed), enabling the system to focus on the differences in the actions themselves rather than the temporal asynchrony. This is crucial for identifying subtle action deviations (such as finger position deviations, joint angle deviations, etc.).
[0038] In one embodiment, in order to enable the two to be compared on the same time axis, alignment processing is required. During this process, the system considers two key factors: one is the difference in performance speed, and the other is the difference in the action start points. For example, a student may perform a certain segment at a slower speed, while a professional performer may complete the same segment at a faster speed. In addition, a student may start performing a certain action at a specific time point, while the action start point of the professional performer may be slightly different. Specifically, this goal can be achieved through the Dynamic Time Warping (DTW) algorithm. DTW is a method commonly used for time series matching, which can find the optimal alignment path between two time series, even if the time lengths or speeds of the two series are different. First, the joint position time series of the student and the professional performer are input into the DTW algorithm. The algorithm will calculate a path to maximize the similarity between the two. This path describes how to "stretch" or "compress" one of the time series to align it with the other time series. For a specific example, if a student starts playing a note at the 5th second, while a professional performer starts playing the same note at the 4th second, and the student takes 10 seconds to complete this passage, while the professional performer only takes 8 seconds. Then, through the DTW algorithm, the time series of the student can be adjusted to align it with the time series of the professional performer. In this way, when checking the actions of the student and the professional performer at the 6th second, it can be ensured that they are indeed performing the corresponding parts in the same music segment, thus effectively analyzing the action differences.
[0039] Exemplarily, in the performance action feature extraction module 50, performance action joint features are extracted from the time series of the aligned performance joint images and the time series of the standard music performance joint images to obtain the time series of the performance joint image feature coding vectors and the time series of the standard music performance joint image feature coding vectors. It should be understood that since the time series of the aligned performance joint images and the time series of the standard music performance joint images usually contain a large amount of high-dimensional data (such as 3D coordinates, speed, acceleration, etc. of each joint). Directly using these raw data for calculation will result in high computational complexity, large storage requirements, and being easily affected by noise interference. In addition, music performance actions have specific semantic information, such as the bending angle of the fingers, the rotation angle of the wrist, the movement trajectory of the elbow, etc., and these semantic information are the key to evaluating the performance quality. Therefore, in the technical solution of the present application, performance action joint features are further extracted from the time series of the aligned performance joint images and the time series of the standard music performance joint images to obtain the time series of the performance joint image feature coding vectors and the time series of the standard music performance joint image feature coding vectors. Through feature extraction technology, more valuable information for evaluating performance actions can be refined from these data, such as joint angles, limb directions, or the position, speed, and acceleration of the fingers, etc., so as to more intuitively reflect the semantics of performance actions. Converting these features into feature coding vectors can more effectively capture the essential characteristics of the actions.
[0040] In one embodiment, the performance action feature extraction module is configured to: respectively input the time series of the aligned performance joint images and the time series of the standard music performance joint images into a performance action joint feature extractor based on a convolutional neural network model to obtain the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors. Specifically, input the time series of the aligned performance joint images into the CNN model. At the same time, also input the time series of the standard music performance joint images into the same model. The model will process these image data layer by layer, and each layer will extract different levels of features, such as low-level features like edges and textures, as well as high-level semantic features like joint angles and limb directions. Through this multi-level feature extraction process, a series of feature encoding vectors can be finally obtained, and each vector represents the joint position and its related characteristics at a specific time point. For the time series of the aligned performance joint images, the CNN model outputs a time series of encoding vectors containing the joint positions and their change features during the performance. Similarly, for the time series of the standard music performance joint images, the model will also generate the corresponding time series of encoding vectors, but these vectors represent the joint positions and action features that should be in the ideal state. With these two sets of encoding vectors, the difference between the two can be further calculated, so as to provide targeted feedback suggestions.
[0041] Exemplarily, in the performance action matching analysis module 60, a time series of the performance joint image feature encoding vectors and a time series of the standard music performance joint image feature encoding vectors are calculated to obtain a time queue of performance action difference fine-grained encoding vectors by calculating the position-wise differential feature vectors between each pair of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors. It should be understood that after performing performance action joint feature extraction on each performance joint image in the time series of the performance joint images and the time series of the standard music performance joint images, further comparing the performance action with the standard action helps to more accurately identify the action deviation during the student's performance, that is, clearly point out in which specific joint actions or postures the student deviates from the standard, so as to provide targeted improvement suggestions. Based on this, in the technical solution of this application, a time series of the performance joint image feature encoding vectors and a time series of the standard music performance joint image feature encoding vectors are further calculated to obtain a time queue of performance action difference fine-grained encoding vectors by calculating the position-wise differential feature vectors between each pair of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors. In particular, by calculating the position-wise difference between the time series of the performance joint image feature encoding vectors of the standard action containing the student's performance joint image features and the standard music performance joint image features at each time point and the time series feature encoding vectors of the standard music performance joint image feature encoding vectors, the spatial difference between the two features at each time point can be intuitively quantified. For example, the differential vector can represent the deviation of the student's finger position from the standard position, the deviation of the wrist angle, etc. In addition, each performance action difference fine-grained encoding vector in the time queue of the performance action difference fine-grained encoding vectors can not only represent the magnitude of the performance action difference, but also locate the specific position of the difference (such as a certain joint or a certain time point). This enables the system to generate targeted feedback opinions, such as "At the 3rd second, the position of the left little finger is 2 cm higher". That is to say, by comparing the student's performance action with the standard action, the system can identify the deviations of the student in terms of posture, finger position, and action fluency. In another specific embodiment, the Euclidean distance between each pair of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors can also be calculated to quantify the difference between the student's action and the standard action. This comparison includes not only the spatial difference but also the temporal difference, ensuring that the timing of the action is fully considered.
[0042] Exemplarily, in the evaluation and feedback module 70, feedback opinions are generated based on the time queue of the performance action difference fine-grained encoding vectors and displayed on the screen, and the feedback opinions are used to indicate whether the action coherence meets the preset standard. In one embodiment, as Figure 3As shown, the evaluation and feedback module 70 includes: a performance action context correlation analysis unit 71, configured to perform dynamic walk correlation encoding based on the performance action difference time series on the time queue of the performance action difference fine-grained encoding vectors to obtain performance action difference time series context correlation encoding vectors; and a feedback opinion generation unit 72, configured to generate the feedback opinion based on the performance action difference time series context correlation encoding vectors.
[0043] Exemplarily, in the performance action context correlation analysis unit 71, it is configured to perform dynamic walk correlation encoding based on the performance action difference time series on the time queue of the performance action difference fine-grained encoding vectors to obtain performance action difference time series context correlation encoding vectors. It should be understood that the time queue of the performance action difference fine-grained encoding vectors contains the action state difference information between the student's actions and the standard actions at each time point in space. There are temporal causal correlations and context dependencies between these fine-grained features of the performance action differences at different time points. However, this difference information is usually high-dimensional and complex. Directly using this information for feedback generation will cause some problems. For example, the time queue of the performance action difference fine-grained encoding vectors may contain irrelevant noise information, thus affecting the accuracy of the analysis and judgment of action coherence. Moreover, music performance actions have strong temporal dependencies, and it is difficult for traditional methods to model such complex temporal relationships and context dependencies. Based on this, in the technical solution of this application, the time queue of the performance action difference fine-grained encoding vectors is further subjected to dynamic walk correlation encoding based on the performance action difference time series to obtain performance action difference time series context correlation encoding vectors.
[0044] In one embodiment, as Figure 4 shown, the performance action context correlation analysis unit 71 includes: a performance action implicit feature mining subunit 711, configured to perform implicit feature mining on each performance action difference fine-grained encoding vector in the time queue of the performance action difference fine-grained encoding vectors to obtain a set of performance action difference fine-grained deep implicit feature encoding vectors; a causal association topological feature construction subunit 712, configured to construct the causal association topological features of the set of performance action difference fine-grained deep implicit feature encoding vectors to obtain a performance action difference semantic causal association topological feature matrix; and a context dynamic walk fusion subunit 713, configured to use the performance action difference semantic causal association topological feature matrix as the semantic causal association structure information, and perform context dynamic walk fusion on the time queue of the performance action difference fine-grained encoding vectors and the set of performance action difference fine-grained deep implicit feature encoding vectors to obtain the performance action difference time series context correlation encoding vectors.
[0045] Exemplarily, in the performance action implicit feature mining subunit 711, implicit feature mining is performed on each vector in the time queue of the performance action difference fine-grained encoding vectors. This process aims to reveal the deep features and patterns hidden beneath the surface differences, helping to identify potential problem sources and improvement directions. Specifically, during the implicit feature mining process, the system applies point convolution encoding and a convolution encoding activation function to further process these fine-grained encoding vectors. These algorithms can automatically learn and extract complex patterns and features, revealing the implicit structure in the data. Through implicit feature mining, the system can obtain a set containing richer information - a set of performance action difference fine-grained deep implicit feature encoding vectors. Each vector in this set represents the deep features at a specific time point, including not only the specific positions of the joints but also the overall posture of the limbs, the smoothness of the actions, and the coordination between different parts. Specifically, the processing process of the performance action implicit feature mining subunit 711 can be expressed by the formula as follows:
[0046] O = {x1, x2,..., x i ,..., x n}
[0047] v i = Sigmoid[Conv 1×1 (x i )]
[0048] D = {v1, v2,..., v i ,..., v n}
[0049] where O is the time queue of the performance action difference fine-grained encoding vectors, x1, x2, x i and x n are respectively the 1st, 2nd, ith, and nth performance action difference fine-grained encoding vectors in the time queue of the performance action difference fine-grained encoding vectors, Conv 1×1 is point convolution encoding, Sigmoid is the convolution encoding activation function, v1, v2, v i , v j and v n are respectively the 1st, 2nd, ith, jth, and nth performance action difference fine-grained deep implicit feature encoding vectors in the set of the performance action difference fine-grained deep implicit feature encoding vectors, and D is the set of the performance action difference fine-grained deep implicit feature encoding vectors.
[0050] In one embodiment, as Figure 5As shown, the causal association topological feature construction subunit 712 includes: a semantic causal association factor calculation secondary subunit 7121, configured to calculate a semantic causal association factor between any two performance action difference fine-grained depth implicit feature encoding vectors in the set of performance action difference fine-grained depth implicit feature encoding vectors to obtain a performance action difference semantic causal association topological matrix composed of multiple performance action difference semantic causal association factors; a causal trigger processing secondary subunit 7122, configured to input the performance action difference semantic causal association topological matrix into a causal trigger network based on a gated activation function to obtain the performance action difference semantic causal association topological feature matrix.
[0051] In one embodiment, the semantic causal association factor calculation secondary subunit 7121 is configured to: calculate an association matrix between any two performance action difference fine-grained depth implicit feature encoding vectors in the set of performance action difference fine-grained depth implicit feature encoding vectors to obtain a set of performance action difference association matrices. Specifically, this process is represented by the formula:
[0052]
[0053] Where is matrix multiplication, v j T is the transposed vector of v j and M i-j is the performance action difference association matrix between v i and v j .
[0054] Calculate the semantic causal association factors of each performance action difference association matrix in the set of performance action difference association matrices to obtain the performance action difference semantic causal association topological matrix composed of multiple performance action difference semantic causal association factors. The performance action difference semantic causal association factor is calculated from the mean, variance, maximum value, and causal association bias value of the performance action difference association matrix. Specifically, this process is represented by the formula:
[0055]
[0056] Where σ 2 (M i-j ) is the variance of M i-j , max(M i-j ) is to take the maximum value in M i-j , μ(M i-j ) is the mean of M i-j , λ is the causal association bias value, t i-j is the performance action difference semantic causal association factor corresponding to M i-j .1-1 , t 1-n , t n-1 and t n-n are the semantic causal association factors of the performance action differences at various positions in the semantic causal association topology matrix of performance action differences respectively, and T is the semantic causal association topology matrix of the performance action differences.
[0057] In one embodiment, the semantic causal association factors of each performance action difference association matrix in the set of the performance action difference association matrices are calculated to obtain the semantic causal association topology matrix of performance action differences composed of multiple semantic causal association factors of performance action differences. The semantic causal association factors of performance action differences are calculated from the mean, variance, maximum value, and causal association bias value of the performance action difference association matrix, including: in response to the variance of the performance action difference association matrix being greater than or equal to a predetermined threshold, taking the weighted average of the distances between any two performance action fine-grained depth implicit feature encoding vectors in the set of the performance action fine-grained depth implicit feature encoding vectors as the causal association bias value; in response to the variance of the performance action difference association matrix being less than the predetermined threshold, taking the weighted value of the variance of the performance action difference association matrix as the causal association bias value. Specifically, this process is represented by the formula:
[0058]
[0059] where d(v i , v j ) is the distance between v i and v j , L is the number of vectors in D, ε is the predetermined threshold, and α and β are weighted hyperparameters.
[0060] Specifically, the processing process of the causal trigger processing secondary subunit 7122 can be represented by the formula:
[0061]
[0062] where softmax is a non-linear activation function, τ is a normalization threshold, f trigger (T) is the gated activation processing of T, and M is the semantic causal association topology feature matrix of performance action differences.
[0063] It should be understood that, first, the semantic causal association factor between the fine-grained depth implicit feature encoding vectors of the performance action differences is calculated to reveal the causal relationship between the action differences at different time points. Specifically, the system calculates the association matrix between any two fine-grained depth implicit feature encoding vectors of the performance action differences in the set, and obtains the performance action difference semantic causal association topology matrix composed of multiple performance action difference semantic causal association factors. In this process, the system not only focuses on the action differences at each time point, but also tries to understand the causal relationship of these differences in the time dimension. For example, when analyzing a student's performance of a certain segment, the system finds that there is a certain connection between the finger position deviation and the wrist angle change of the student. This connection may be due to the change of finger position causing the adjustment of wrist angle, or vice versa. To capture these complex relationships, the system calculates the association matrix between each pair of encoding vectors and further extracts the semantic causal association factors. These factors can quantify the degree of mutual influence between different action differences and form a topology matrix. Subsequently, the system inputs this performance action difference semantic causal association topology matrix into the causal trigger network based on the gated activation function. The causal trigger network is a neural network structure specially designed to process causal relationships. It can selectively activate or inhibit certain paths through the gating mechanism, so as to better model complex causal relationships. In the network, the system further processes the topology matrix to generate the performance action difference semantic causal association topology feature matrix. Through this series of steps, the system can more deeply understand the problems existing in the student's performance actions and the reasons behind them. For example, the system can identify that the finger position deviation of the student at a certain specific time point is actually caused by the slight change of the wrist angle in the previous few seconds. This deep understanding makes the feedback more accurate and targeted.
[0064] In one embodiment, the context dynamic walk fusion subunit 713 is configured to: input the performance action difference semantic causal association topology feature matrix and the time queue of the performance action difference fine-grained encoding vectors into the feature sequence dynamic walk encoder based on the graph convolutional neural network model to obtain the performance action difference surface context dynamic walk semantic encoding vector. Specifically, this process is represented by the formula:
[0065]
[0066] where GCN is the graph convolutional processing, represents any given, H surface is the performance action difference surface context dynamic walk semantic encoding vector.
[0067] Input the semantic causal association topological feature matrix of the performance action differences and the set of fine-grained depth implicit feature encoding vectors of the performance action differences into the feature sequence dynamic walk encoder based on the graph convolutional neural network model to obtain the hidden layer context dynamic walk semantic encoding vector of the performance action differences. Specifically, this process is represented by the formula:
[0068]
[0069] where, H hidden is the hidden layer context dynamic walk semantic encoding vector of the performance action differences.
[0070] Fuse the hidden layer context dynamic walk semantic encoding vector of the performance action differences and the surface layer context dynamic walk semantic encoding vector of the performance action differences to obtain the temporal context association encoding vector of the performance action differences. Specifically, this process is represented by the formula:
[0071] H final = γ · H surface + (1 - γ) · H hidden
[0072] where, γ is the fusion weighting parameter, and H final is the temporal context association encoding vector of the performance action differences.
[0073] It should be understood that, first, the system inputs the time queue of the semantic causal association topological feature matrix of performance action differences and the fine-grained encoded vector of performance action differences into the feature sequence dynamic walk encoder based on the graph convolutional neural network (GCN) model to obtain the surface context dynamic walk semantic encoded vector of performance action differences. This process aims to capture explicit action differences and their evolution laws over time. For example, when analyzing the deviation between the student's finger position and the standard action, the system not only focuses on the specific differences at each time point but also considers how these differences change and develop over time. This surface semantic encoding can intuitively reflect the student's action state at different time points and its comparison with the standard action. At the same time, the system also inputs the set of the semantic causal association topological feature matrix of performance action differences and the fine-grained deep implicit feature encoded vector of performance action differences into the same feature sequence dynamic walk encoder based on the graph convolutional neural network model to obtain the hidden layer context dynamic walk semantic encoded vector of performance action differences. The hidden layer semantic encoding aims to capture implicit action patterns and deep relationships. For example, the system can identify how the change in wrist angle affects the finger position, or how the elbow posture affects the coordination of the entire arm. These implicit causal relationships are crucial for understanding the fundamental problems existing in the student's performance. Through the hidden layer semantic encoding, the system can reveal subtle but important action differences that are difficult to observe on the surface, thus providing more targeted improvement suggestions. After these two steps are completed, the system will obtain two types of semantic encoded vectors: the hidden layer context dynamic walk semantic encoded vector of performance action differences and the surface context dynamic walk semantic encoded vector of performance action differences. To integrate these two types of information, the system will fuse these two vectors to generate the final temporal context association encoded vector of performance action differences. This process enables the system to comprehensively describe the overall characteristics of the student's performance action and its differences from the standard action. For example, the system can not only point out the specific deviation of the student's finger position at a specific time point but also explain how this deviation is caused by the previous changes in wrist or elbow posture and predict its impact on subsequent actions.
[0074] Specifically, by constructing a causal association topology matrix, the causal relationship of performance action differences in time series can be explicitly modeled. For example, the deviation of the wrist angle may cause the deviation of the finger position, and this causal relationship can be captured by the causal trigger network. Causal modeling can reveal the internal driving relationship of action differences, helping the system to more accurately identify the root cause of problems, rather than just the surface differences. Moreover, by simulating the propagation of features in the topological structure through the dynamic walk mechanism, the evolution law of action differences over time can be captured. For example, the deviation of the finger position may start at a certain time point and gradually increase, and this time series dependence can be modeled by the dynamic walk mechanism. The enhancement of time series context dependence enables the system to more comprehensively analyze the evolution process of action differences, thus providing more accurate feedback. The encoder generates context semantic vectors for the surface layer and the hidden layer through hierarchical dynamic walk encoding. The surface layer semantics captures explicit action differences (such as the deviation of finger position), while the hidden layer semantics captures implicit action differences (such as the impact of wrist angle deviation on the overall action). The multi-level semantic representation can more comprehensively describe action differences, avoiding the semantic loss problem caused by single representation in traditional methods. At the same time, compressing the high-dimensional difference vector into a low-dimensional semantic feature representation can also remove redundant information and noise.
[0075] Exemplarily, in the feedback opinion generation unit 72, the feedback opinion is generated based on the performance action difference time series context association encoding vector. In one embodiment, the feedback opinion generation unit is configured to: input the performance action difference time series context association encoding vector into an intelligent feedback module based on a large language model to obtain the feedback opinion. That is, the performance action difference time series context association encoding features are used to judge the action coherence, so as to identify the problems existing in the student's actions and provide targeted feedback and improvement suggestions on the student's performance action coherence to assist in intelligent music teaching. This can not only improve the teaching efficiency, but also help students master the correct performance skills faster and enhance the learning experience.
[0076] In a specific embodiment, the encoded vector of the sequential context correlation of performance action differences is input into a pre-trained large language model. For example, the information received by the model may indicate that during the 3rd to 5th seconds, the position of the little finger of the left hand was 2 centimeters higher, resulting in insufficient strength for some notes; and during the 10th to 12th seconds, the angle of the right hand wrist was too low, affecting the coherence and fluency of the notes. Based on these details and combined with its rich language and music knowledge base, the large language model will generate the following feedback: "Between the 3rd and 5th seconds of your performance, it is noticed that the position of your left hand little finger was slightly raised by about 2 centimeters, which may cause insufficient strength when playing some notes. It is recommended that you try to lower the height of the little finger to ensure it is on the same level as the other fingers, so that the force can be distributed more evenly and each note can sound clearly and powerfully. Additionally, during the 10th to 12th seconds, the angle of your right hand wrist appears to be too low, which may affect the coherence and fluency between the notes. Try to keep the wrist slightly raised to make the movement of the arm more natural and smooth, which helps to improve the overall quality of the performance." Furthermore, the intelligent feedback module can also provide some visual aids or animation demonstrations to show the correct postures and movements, helping students to more intuitively understand and correct their mistakes. For example, the system can display a virtual hand model on the screen to simulate the correct finger positions and wrist angle changes, allowing students to compare their movements in real time and make adjustments. In this way, the system can not only identify the subtle problems in the students' performances, but also provide specific and actionable improvement suggestions to help students master the correct performance techniques faster and enhance the learning efficiency and experience. This method combines advanced machine learning techniques and meticulous language expression capabilities, making the feedback more accurate, personalized, and easy to understand.
[0077] In summary, the intelligent music teaching assistance system and method based on action perception according to the embodiments of the present application are elucidated. It can use deep learning algorithms based on artificial intelligence and image analysis to automatically analyze the performance actions of students, and conduct action difference comparison and sequential analysis with the standard music performance actions of musical instruments to identify the problems existing in the students' actions, and provide targeted feedback and improvement suggestions regarding the coherence of the students' performance actions, thereby assisting in intelligent music teaching. This can not only improve the teaching efficiency, but also help students master the correct performance techniques faster and enhance the learning experience.
[0078] Figure 6 It is a schematic flowchart of the intelligent music teaching assistance method based on action perception according to the embodiments of the present application. As Figure 6As shown, the action perception-based intelligent music teaching assistance method includes: S1, using a camera to capture a monitoring video of the performance process of a target student object; S2, using the OpenPose model to process the monitoring video of the performance process to obtain a time series of performance joint images; S3, extracting a time series of standard music performance joint images from a standard action database; S4, performing alignment processing on the time series of the performance joint images and the time series of the standard music performance joint images to obtain the aligned time series of the performance joint images and the time series of the standard music performance joint images; S5, extracting performance action joint features from the aligned time series of the performance joint images and the time series of the standard music performance joint images to obtain a time series of performance joint image feature encoding vectors and a time series of standard music performance joint image feature encoding vectors; S6, calculating the position-by-position differential feature vectors between each group of corresponding performance joint image feature encoding vectors and standard music performance joint image feature encoding vectors in the time series of the performance joint image feature encoding vectors and the time series of the standard music performance joint image feature encoding vectors to obtain a time queue of performance action difference fine-grained encoding vectors; S7, generating and displaying feedback opinions on the screen based on the time queue of the performance action difference fine-grained encoding vectors, where the feedback opinions are used to indicate whether the action coherence meets a preset standard.
[0079] Here, those skilled in the art can understand that the specific operations of each step in the above action perception-based intelligent music teaching assistance method have been introduced in detail in the description of the Figures 1 to 5 action perception-based intelligent music teaching assistance system above, and therefore, the repeated description thereof will be omitted.
[0080] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0081] It should be understood that the specific examples in this article are only to help those skilled in the art better understand the embodiments of this application, rather than limiting the scope of the embodiments of this application.
[0082] It should also be understood that in various embodiments of this application, the magnitude of the serial numbers of each process does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0083] It should also be understood that the various embodiments described in this specification can be implemented either individually or in combination, and the embodiments of the present application do not limit this.
[0084] Unless otherwise specified, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field of the present application. The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the scope of the present application. The term "and / or" used in the present application includes any and all combinations of one or more of the related listed items. The singular forms "a", "above-mentioned", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0085] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0086] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0087] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0088] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0089] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An intelligent music teaching auxiliary system based on motion perception, characterized in that: include: A motion data acquisition module is used to capture a monitoring video of the performance process of a target student object using a camera; A performance action analysis module, used to process the performance process monitoring video using an OpenPose model to obtain a time series of performance joint images; A standard performance action extraction module is used to extract a time series of standard music performance joint images from a standard action database; A synchronization module, used for aligning the time sequence of the performance joint images and the time sequence of the standard music performance joint images to obtain an aligned time sequence of the performance joint images and a time sequence of the standard music performance joint images; A performance action feature extraction module, used to extract performance action joint features from the aligned performance joint image time series and the standard music performance joint image time series to obtain a performance joint image feature encoding vector time series and a standard music performance joint image feature encoding vector time series; A performance action matching analysis module is used to calculate the position difference feature vectors between each group of corresponding performance joint image feature coding vectors and the standard music performance joint image feature coding vectors in the time series of the performance joint image feature coding vectors and the time series of the standard music performance joint image feature coding vectors to obtain a time queue of the performance action difference fine-grained coding vectors; The evaluation and feedback module is used to generate and display feedback on the screen based on the time queue of the fine-grained encoding vector of the performance action difference, and the feedback is used to indicate whether the action coherence meets the preset standard.
2. The intelligent music teaching auxiliary system based on motion perception according to claim 1 is characterized in that: The performance action feature extraction module is used to: The aligned time series of playing joint images and the time series of standard music playing joint images are respectively input into a playing action joint feature extractor based on a convolutional neural network model to obtain a time series of playing joint image feature coding vectors and a time series of standard music playing joint image feature coding vectors.
3. The intelligent music teaching auxiliary system based on motion perception according to claim 2 is characterized in that: The evaluation and feedback module includes: A performance action context association analysis unit, used for performing dynamic walking association encoding based on the performance action difference timing sequence on the time queue of the performance action difference fine-grained encoding vector to obtain a performance action difference timing sequence context association encoding vector; The feedback generating unit is used to generate the feedback based on the performance action difference timing context associated coding vector.
4. The intelligent music teaching auxiliary system based on motion perception according to claim 3 is characterized in that: The performance action context association analysis unit comprises: A playing action implicit feature mining subunit, used for performing implicit feature mining on each playing action difference fine-grained coding vector in the time queue of the playing action difference fine-grained coding vector to obtain a set of playing action difference fine-grained deep implicit feature coding vectors; A causal correlation topological feature construction subunit is used to construct the causal correlation topological features of the set of fine-grained deep implicit feature encoding vectors of the playing action differences to obtain a playing action difference semantic causal correlation topological feature matrix; The context dynamic walking fusion sub-unit is used to use the semantic causal association topological feature matrix of the performance action difference as the semantic causal association structure information, and perform context dynamic walking fusion on the time queue of the performance action difference fine-grained encoding vector and the set of the performance action difference fine-grained deep implicit feature encoding vector to obtain the performance action difference temporal context association encoding vector.
5. The intelligent music teaching auxiliary system based on motion perception according to claim 4 is characterized in that: The causal association topological feature construction subunit includes: A semantic causal association factor calculation secondary subunit is used to calculate the semantic causal association factor between any two playing action difference fine-grained deep implicit feature coding vectors in the set of playing action difference fine-grained deep implicit feature coding vectors to obtain a playing action difference semantic causal association topological matrix composed of multiple playing action difference semantic causal association factors; The causal trigger processing secondary sub-unit is used to input the performance action difference semantic causal association topological matrix into the causal trigger network based on the gated activation function to obtain the performance action difference semantic causal association topological feature matrix.
6. The motion-aware intelligent music teaching auxiliary system according to claim 5 is characterized in that: The semantic causal association factor calculation secondary subunit is used to: Calculating the association matrix between any two playing action difference fine-grained deep implicit feature coding vectors in the set of playing action difference fine-grained deep implicit feature coding vectors to obtain a set of playing action difference association matrices; The semantic causal association factors of each playing action difference association matrix in the set of the playing action difference association matrices are calculated to obtain the playing action difference semantic causal association topological matrix composed of multiple playing action difference semantic causal association factors, and the playing action difference semantic causal association factors are calculated from the mean, variance, maximum value and causal association bias value of the playing action difference association matrix.
7. The motion-aware intelligent music teaching auxiliary system according to claim 6 is characterized in that: Calculating the semantic causal association factor of each playing action difference association matrix in the set of the playing action difference association matrices to obtain the playing action difference semantic causal association topological matrix composed of a plurality of playing action difference semantic causal association factors, wherein the playing action difference semantic causal association factor is calculated from the mean, variance, maximum value and causal association bias value of the playing action difference association matrix, including: In response to the variance of the performance action difference association matrix being greater than or equal to a predetermined threshold, taking a weighted average of the distances between any two performance action difference fine-grained deep implicit feature coding vectors in the set of the performance action difference fine-grained deep implicit feature coding vectors as the causal association bias value; In response to the variance of the performance action difference association matrix being smaller than the predetermined threshold, a weighted value of the variance of the performance action difference association matrix is used as the causal association bias value.
8. The motion-aware intelligent music teaching auxiliary system according to claim 7 is characterized in that: The context dynamic walk fusion subunit is used to: Input the performance action difference semantic causal association topological feature matrix and the time queue of the performance action difference fine-grained encoding vector into a feature sequence dynamic walk encoder based on a graph convolutional neural network model to obtain a performance action difference surface context dynamic walk semantic encoding vector; Inputting the performance action difference semantic causal association topological feature matrix and the performance action difference fine-grained deep implicit feature encoding vector set into the feature sequence dynamic walking encoder based on the graph convolutional neural network model to obtain the performance action difference hidden layer context dynamic walking semantic encoding vector; The playing action difference hidden context dynamic wandering semantic coding vector and the playing action difference surface context dynamic wandering semantic coding vector are fused to obtain the playing action difference temporal context association coding vector.
9. The motion-aware intelligent music teaching auxiliary system according to claim 8, characterized in that: The feedback generating unit is used to input the performance action difference time sequence context association coding vector into the intelligent feedback module based on the large language model to obtain the feedback.
10. An intelligent music teaching auxiliary method based on motion perception, characterized in that: include: Use a camera to capture surveillance video of the target student's performance; Using the OpenPose model to process the monitoring video of the performance process to obtain a time series of performance joint images; Extracting time series of standard music performance joint images from a standard action database; Aligning the time series of the playing joint images and the time series of the standard music playing joint images to obtain aligned time series of the playing joint images and time series of the standard music playing joint images; Extracting performance action joint features from the aligned performance joint image time series and the standard music performance joint image time series to obtain a performance joint image feature encoding vector time series and a standard music performance joint image feature encoding vector time series; Calculate the position difference feature vectors between each group of corresponding performance joint image feature coding vectors and the standard music performance joint image feature coding vectors in the time series of the performance joint image feature coding vectors and the time series of the standard music performance joint image feature coding vectors to obtain a time queue of the performance action difference fine-grained coding vectors; Feedback is generated based on the time queue of the fine-grained encoding vector of the performance action difference and displayed on the screen, wherein the feedback is used to indicate whether the action continuity meets the preset standard.
Citation Information
Cited By
Buffer bin dehumidification system
CN120010572A
Fingerprint identification method, fingering identification system and fingering identification equipment applied to musical instrument system, and medium
CN121459429A
A fingering recognition method, system, device and medium applied to a musical instrument system
CN121459429B
Event identification method and system applied to musical instrument system
CN121884452A