Dancing motion real-time analysis and evaluation system based on computer vision

By constructing a multi-dimensional assessment system and combining multi-scale feature extraction and spatiotemporal attention fusion, multi-dimensional assessment and real-time feedback of dance movements are realized, solving the problems of single assessment dimensions and lack of real-time feedback in existing technologies, and improving the assessment accuracy and training efficiency of dance teaching.

CN121768075APending Publication Date: 2026-03-31JIANGXI NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing dance movement assessment systems have limited evaluation dimensions, lack real-time feedback, have fixed weight allocations, limited feature extraction capabilities, and lack historical data analysis, making it difficult to provide comprehensive and objective evaluations and personalized suggestions.

Method used

A multi-dimensional evaluation system based on computer vision is adopted, including an image acquisition module, a multi-scale feature extraction module, a spatiotemporal attention fusion module, a multi-dimensional evaluation engine, an adaptive weight adjustment module, and a historical data management module, to achieve a comprehensive evaluation of posture accuracy, rhythm matching, movement smoothness, and expressiveness, and to provide real-time feedback and historical data analysis.

Benefits of technology

It enables multi-dimensional comprehensive assessment, improves assessment accuracy and real-time performance, enhances feature extraction capabilities, provides personalized training suggestions and development plans, and improves the efficiency and effectiveness of dance teaching and training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768075A_ABST
    Figure CN121768075A_ABST
Patent Text Reader

Abstract

The invention provides a dance movement real-time analysis and evaluation system based on computer vision, and belongs to the technical field of computer vision and human body movement analysis. The system comprises an image acquisition module, a multi-scale feature extraction module, a space-time attention fusion module, a multi-dimensional evaluation engine, a self-adaptive weight adjustment module, a real-time feedback generation module and a historical data management module, the system carries out feature extraction through Vision Transform and multi-scale convolution, a space-time attention mechanism is adopted to capture the coherence and coordination of actions, and the time-space attention fusion is realized. A multi-dimensional evaluation system including posture accuracy, rhythm matching degree, motion fluency and expressive force is constructed, evaluation weight is adaptively adjusted according to dance types, real-time visual feedback and historical data analysis are provided, and practical application shows that the system evaluation accuracy reaches 92% of the professional teacher level, the training efficiency is improved by 38%, and the training effect is good. And an efficient technical support tool is provided for dance teaching training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and human motion analysis technology, and in particular to a real-time analysis and evaluation system for dance movements based on computer vision, which is specifically applied to scenarios such as dance teaching, training, competition scoring, and self-practice. Background Technology

[0002] As a performing art that blends artistry and technique, dance's accuracy, fluidity, and expressiveness directly impact the overall performance. Traditional dance teaching and assessment primarily rely on manual observation and subjective judgment, a method with several limitations. First, manual assessment is influenced by the assessor's experience, level of concentration, and subjective preferences, making it difficult to standardize assessment criteria and potentially leading to different evaluations for the same dance movement. Second, manual assessment lacks quantifiable objective data support, making it difficult to precisely pinpoint the location and extent of movement deviations. Third, traditional assessment methods are limited by time and space, hindering dancers from continuous self-practice and feedback after class. Furthermore, manual assessment cannot achieve systematic tracking and data analysis of a dancer's long-term training process, making it difficult to scientifically evaluate training effectiveness and develop personalized improvement plans.

[0003] To address the aforementioned issues, automatic dance movement evaluation systems based on computer vision technology have emerged in recent years. For example, CN113516005A discloses a dance movement evaluation system based on deep learning and pose estimation. This system acquires dance videos using an RGB camera, extracts the coordinates of the human skeleton's joints using VideoPose3D, then encodes the movement sequence using an LSTM network and an Attention layer, and finally completes the evaluation by calculating the cosine similarity of the encoded data. This system uses the MMD_NCA loss function as a constraint during the training of the motion analysis neural network, clustering similar dances in the encoding space and separating different dances, thereby optimizing the encoding results.

[0004] However, the existing technologies still have the following shortcomings: First, the evaluation dimension is singular, relying solely on the similarity of posture encoding for evaluation. This fails to comprehensively reflect the multi-dimensional characteristics of dance movements, neglecting important evaluation indicators such as rhythm, fluency, and expressiveness, making it difficult to provide comprehensive and objective evaluation results. Second, there is a lack of real-time feedback mechanisms. The system only provides a final score after the dance movement is completed, failing to offer immediate movement correction suggestions during training, thus reducing training efficiency. Third, the weight allocation is fixed, failing to consider the different levels of emphasis placed on each evaluation dimension by different dance types. For example, classical ballet emphasizes posture accuracy, while modern dance emphasizes expressiveness; fixed weights cannot adapt to different application scenarios. Fourth, the feature extraction capability is limited. LSTM-based sequence encoding methods have limitations in capturing long-distance spatiotemporal dependencies and do not fully utilize multi-scale features and attention mechanisms to enhance the expression of key features. Fifth, there is a lack of historical data analysis capabilities, making it impossible to track and predict the long-term training process of dancers, and thus difficult to provide personalized training suggestions and development plans.

[0005] Therefore, there is an urgent need to develop a dance movement analysis and evaluation system that can perform multi-dimensional comprehensive evaluation, support real-time feedback, have adaptive weight adjustment capabilities, adopt advanced feature extraction technology, and provide historical data analysis functions, in order to solve the above-mentioned defects of existing technologies and meet the actual needs of dance teaching and training. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a real-time analysis and evaluation system for dance movements based on computer vision. It aims to solve problems such as the single evaluation dimension, lack of real-time feedback, fixed weight allocation, limited feature extraction capabilities, and lack of historical data analysis in existing dance evaluation systems.

[0007] The technical solution of this invention is: a real-time analysis and evaluation system for dance movements based on computer vision, including an image acquisition module, a multi-scale feature extraction module, a spatiotemporal attention fusion module, a multi-dimensional evaluation engine, an adaptive weight adjustment module, a real-time feedback generation module, and a historical data management module. The image acquisition module captures video streams of dance performance scenes in real time using a high-precision camera. The multi-scale feature extraction module extracts multi-scale features from preprocessed video frames, identifies the coordinates of key points on the three-dimensional human skeleton, and generates human posture sequence data. The spatiotemporal attention fusion module performs spatiotemporal correlation analysis on the human posture sequence data, capturing the continuity and coordination features of movements through attention mechanisms in the time and space dimensions, respectively. The multi-dimensional evaluation engine receives the spatiotemporal fusion feature vector output by the spatiotemporal attention fusion module, including a posture accuracy evaluation unit, a rhythm matching degree evaluation unit, a movement fluency evaluation unit, and an expressiveness evaluation unit. Each evaluation unit extracts feature data of the corresponding dimension from the spatiotemporal fusion feature vector, and quantitatively evaluates the dance movements from different dimensions. The adaptive weight adjustment module dynamically determines the weight coefficients for four evaluation dimensions—accuracy score, rhythm matching score, movement fluency score, and expressiveness score—based on the dance type and evaluation scenario. The real-time feedback generation module generates visual feedback information based on the overall score and scores for each dimension. Joints with posture deviations exceeding preset deviation thresholds are marked as areas requiring improvement, and movement frames with rhythm deviations exceeding preset time thresholds are marked as time points requiring improvement. The preset deviation thresholds are joint position errors greater than 50mm or angle errors greater than 15°, and the preset time thresholds are time differences from the music beat exceeding 100ms. The historical data management module stores the dancers' historical evaluation data and performs longitudinal comparative analysis.

[0008] Compared with the prior art, the present invention has the following advantages:

[0009] First, it achieves multi-dimensional comprehensive evaluation. This invention constructs a multi-dimensional evaluation system encompassing four dimensions: posture accuracy, rhythm matching, movement fluency, and expressiveness. Compared to existing technologies that rely solely on single-dimensional evaluation based on posture similarity, this system provides a more comprehensive and objective reflection of the overall quality of dance movements. The multi-dimensional evaluation system simulates the evaluation mindset of professional dance teachers, conducting quantitative analysis from multiple perspectives, including technical standardization, musical coordination, movement continuity, and artistic expression. The evaluation accuracy reaches 92% of that of professional teachers, representing a 15 percentage point improvement over existing technologies.

[0010] Secondly, it provides a real-time feedback mechanism. The real-time feedback generation module of this invention can analyze movement deviations and generate visual feedback in real time during dance performances. Compared to existing technologies that only provide a final score after the performance, this significantly shortens the feedback delay time, allowing dancers to adjust and correct movement errors promptly. The real-time feedback mechanism improves training efficiency by 38%, enabling dancers to obtain more effective improvement information in a single training session.

[0011] Third, adaptive weight adjustment is achieved. This invention dynamically adjusts the weight coefficients of each evaluation dimension according to the dance type and evaluation scenario. For example, it increases the weight of posture accuracy for classical ballet, increases the weight of expressiveness for modern dance, and increases the weight of rhythm matching for rhythmic dance. This adaptive mechanism enables the system to adapt to the characteristics and evaluation requirements of different dance styles. The correlation coefficient between the evaluation results and professional judges reaches 0.89, which is 0.12 higher than the fixed weight method.

[0012] Fourth, advanced feature extraction technology is employed. The multi-scale feature extraction module of this invention combines the VisionTransformer architecture and multi-scale convolution, which, compared with the existing LSTM encoding method, can more effectively capture long-distance spatiotemporal dependencies and multi-scale features. The spatiotemporal attention fusion module enhances the expressive power of key action features through a dual attention mechanism in the time and space dimensions, improving feature representation accuracy by 21%.

[0013] Fifth, it provides historical data analysis capabilities. The historical data management module of this invention can systematically store dancers' training data, generate progress trajectory curves and development trend predictions through longitudinal comparative analysis, and provide dancers and teachers with data-driven training effect evaluations and personalized improvement suggestions. This long-term tracking mechanism makes dancers' training plans more scientific and reasonable, and accelerates the improvement of training effects by 25%.

[0014] In summary, this invention comprehensively improves the accuracy, real-time performance, and practicality of dance movement assessment by constructing a multi-dimensional evaluation system, a real-time feedback mechanism, adaptive weight adjustment, advanced feature extraction technology, and historical data analysis functions, providing a powerful technical support tool for dance teaching and training. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the overall architecture of the system of the present invention;

[0016] Figure 2 This is a schematic diagram of the multi-scale feature extraction module;

[0017] Figure 3 This is a schematic diagram of the workflow of the spatiotemporal attention fusion module;

[0018] Figure 4 This is a schematic diagram of the weight calculation process of the adaptive weight adjustment module;

[0019] Figure 5 This is a schematic diagram of the data analysis process in the historical data management module. Detailed Implementation

[0020] Please refer to the attached document. Figures 1-5 The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be noted that the following description is merely illustrative and should not be construed as limiting the scope of the invention.

[0021] Reference Figure 1 This invention provides a real-time analysis and evaluation system for dance movements based on computer vision, including an image acquisition module 1, a multi-scale feature extraction module 2, a spatiotemporal attention fusion module 3, a multi-dimensional evaluation engine 4, an adaptive weight adjustment module 5, a real-time feedback generation module 6, and a historical data management module 7.

[0022] Image acquisition module 1 is used to capture video streams of dance performance scenes in real time using a high-precision camera device, and to extract and preprocess the video stream frames.

[0023] In one embodiment of the present invention, the high-precision camera device can be an RGB camera or an RGB-D depth camera. An RGB camera acquires the dancer's appearance and posture information by capturing color images of the red, green, and blue channels, with a preferred frame rate of 60fps to 120fps, enabling it to capture subtle changes in dance movements. An RGB-D depth camera adds a depth information channel to the RGB image, enabling it to more accurately acquire distance information between the dancer and the camera, improving the accuracy of three-dimensional pose estimation. In a preferred embodiment, an RGB camera with a frame rate of 90fps is used, ensuring both temporal resolution for motion capture and controlling the computational burden of data processing.

[0024] The installation position and angle of the camera device have a significant impact on the quality of data acquisition. Preferably, the camera is installed directly in front of the dancer, 3m to 5m away, with the camera's optical axis at an angle of 10° to 15° to the ground, ensuring that the dancer's full-body movements can be captured completely. In some embodiments, a multi-camera shooting scheme can be used, with cameras installed in front of, to the left of, and to the right of the dancer, respectively. Multi-view fusion technology can be used to improve the robustness and accuracy of pose estimation.

[0025] The video stream preprocessing includes image denoising, contrast enhancement, and frame extraction. Image denoising employs Gaussian filtering with a 5×5 pixel kernel size and a standard deviation σ=1.2, effectively removing image noise while preserving edge information. Contrast enhancement uses histogram equalization to map the image's grayscale value distribution to the full range of 0 to 255, improving the image's visual quality and subsequent processing effectiveness. Frame extraction extracts image frames from the video stream at fixed time intervals, preferably 10ms to 20ms, corresponding to a frame rate of 50fps to 100fps.

[0026] In addition, image acquisition module 1 also performs human detection and region cropping. A deep learning-based target detection algorithm is used to identify human targets in video frames and generate human detection boxes. Preferably, the YOLO series detection algorithm is used, which has fast detection speed and high accuracy, meeting the requirements of real-time processing. After detecting a human target, the detection box region is cropped and its size normalized, uniformly scaling the human image to 256×256 pixels or 512×512 pixels, which serves as the standard input for subsequent feature extraction modules.

[0027] Reference Figure 2 The multi-scale feature extraction module 2 is connected to the image acquisition module 1 and is used to extract multi-scale features from the preprocessed video frames, identify the coordinates of the dancer's three-dimensional human skeleton key points, and generate human posture sequence data.

[0028] In one embodiment of the present invention, the multi-scale feature extraction module 2 employs a Vision Transformer architecture as the backbone network to extract image features. The Vision Transformer divides the input image into 16×16 pixel image blocks. Each image block is flattened into a one-dimensional vector and then mapped to a 768-dimensional feature space through a linear projection layer. Positional encoding information is added to the feature space, enabling the model to learn the spatial relationships between image blocks. Then, the feature vector sequence is input into a multi-layer Transformer encoder for self-attention computation and feature extraction. The Transformer encoder contains 12 layers, each including a multi-head self-attention mechanism and a feedforward neural network, effectively capturing both global and local features of the image.

[0029] To enhance feature extraction capabilities for different body parts, the multi-scale feature extraction module 2 combines multi-scale convolutional kernels with the VisionTransformer for differentiated feature extraction. Specifically, multi-scale convolution operations are performed on the feature map output by the VisionTransformer, using 3×3, 5×5, and 7×7 convolutional kernels to extract fine-grained, medium-grained, and coarse-grained features, respectively. The 3×3 convolutional kernel is suitable for capturing fine features of small joints such as fingers and ankles, the 5×5 convolutional kernel is suitable for capturing medium-scale limb features such as arms and calves, and the 7×7 convolutional kernel is suitable for capturing large-scale body parts such as the torso and thighs. The features from the three scales are fused through channel concatenation to generate a multi-scale fused feature map.

[0030] Multi-scale fused feature maps are input into a keypoint detection network to predict the locations of keypoints in the human skeleton. In a preferred embodiment, the keypoint detection network employs a heatmap regression method to generate a heatmap for each keypoint. The value of each pixel in the heatmap represents the probability that the location contains a keypoint. The peak position of the heatmap corresponds to the two-dimensional image coordinates of the keypoint. This invention predicts 25 human keypoints, including 1 keypoint for the head, 5 keypoints for the torso, 8 keypoints for the upper limbs, 10 keypoints for the lower limbs, and 1 keypoint each for the hands and feet, comprehensively describing the posture information of the human body.

[0031] After obtaining the two-dimensional keypoint coordinates, the multi-scale feature extraction module 2 further estimates the three-dimensional spatial coordinates of the keypoints. Three-dimensional pose estimation employs a depth regression network, using the two-dimensional keypoint coordinates and image features as input to predict the depth value of each keypoint. The depth regression network contains four fully connected layers with 2048, 1024, 512, and 25 neurons respectively, with the last layer outputting the depth values ​​of 25 keypoints. By combining the two-dimensional coordinates and depth values, the three-dimensional spatial coordinates of the keypoints are reconstructed. In a preferred embodiment, the origin of the three-dimensional coordinates is set at the center of the dancer's pelvis, with the x-axis pointing to the dancer's right side, the y-axis pointing upwards, and the z-axis pointing forwards, forming a right-handed coordinate system.

[0032] The multi-scale feature extraction module 2 repeatedly performs the above feature extraction and keypoint detection process on consecutive video frames to generate human pose sequence data. The pose sequence data is organized in time series form, with each time step containing the three-dimensional coordinates of 25 keypoints, represented as a 25×3 matrix. Preferably, the length of the pose sequence is set to 90 to 120 frames, corresponding to a dance movement segment of 1.5 to 2.0 seconds, which can cover the basic movement units of dance while maintaining reasonable computational efficiency.

[0033] Reference Figure 3The spatiotemporal attention fusion module 3 is connected to the multi-scale feature extraction module 2 and is used to perform spatiotemporal correlation analysis on human posture sequence data. It captures the continuity features of the action through the time dimension attention mechanism and captures the limb coordination features through the spatial dimension attention mechanism to generate a spatiotemporal fusion feature vector.

[0034] The spatiotemporal attention fusion module 3 adopts a spatiotemporally decoupled processing architecture, applying attention mechanisms in both the temporal and spatial dimensions, and then fusing features from both dimensions. Compared to direct 3D spatiotemporal convolution or 3D attention, this decoupled design offers higher computational efficiency and stronger feature representation capabilities.

[0035] In the temporal attention mechanism, for each joint in the pose sequence, its motion trajectory along the time axis is analyzed. Specifically, the coordinate sequence of a joint across all time steps is organized into a time-series vector and input into the temporal attention network. The temporal attention network employs a Transformer encoder structure, including multi-head self-attention layers and feedforward neural network layers. The multi-head self-attention layers calculate the correlation between different moments in the time series, automatically learning which poses are most important for the current action. Preferably, the temporal attention network contains 6 Transformer encoder layers, each with 8 attention heads, enabling the learning of temporal dependencies from different perspectives. The output of the temporal attention mechanism is a 256-dimensional temporal feature vector for each joint, encoding the joint's motion pattern across the entire time series.

[0036] In the spatial attention mechanism, for each time step in the pose sequence, the spatial relationships between all relevant nodes at that moment are analyzed. Specifically, the coordinates of all joints at a given time step are organized into a spatial feature matrix and input into the spatial attention network. The spatial attention network also employs a Transformer encoder structure to calculate the correlations between joints. The spatial attention mechanism can learn human kinematic constraints, such as the coordinated movements of the arms and shoulders, and the balance relationships between the torso and legs. The spatial attention network contains four Transformer encoder layers, each with four attention heads. The output of the spatial attention mechanism is a 256-dimensional spatial feature vector for each time step, encoding the coordination state of various parts of the human body at that moment.

[0037] Temporal and spatial feature vectors are fused using a feature fusion network. The fusion network employs a gated fusion mechanism, dynamically learning the fusion weights for temporal and spatial features. The formula for gated fusion is:

[0038] ,

[0039] in, The fused feature vector For time feature vectors, For spatial feature vectors, For the gated weight matrix, This represents the concatenation of eigenvectors. It is the Sigmoid activation function. This is the element-wise multiplication operator. The formula adaptively adjusts the fusion ratio of temporal and spatial features through a gating mechanism, enabling the model to flexibly adjust the fusion strategy according to the characteristics of different dance movements. For dance segments emphasizing movement continuity, the gating mechanism increases the weight of temporal features; for keyframes emphasizing pose accuracy, the gating mechanism increases the weight of spatial features.

[0040] The fused feature vectors are processed through fully connected layers and normalization layers to generate the final spatiotemporal fusion feature vector. The spatiotemporal fusion feature vector has a dimension of 512, comprehensively encoding the temporal coherence and spatial coordination information of dance movements, and serves as the input to the multidimensional evaluation engine.

[0041] The multidimensional evaluation engine 4 is connected to the spatiotemporal attention fusion module 3, receiving the spatiotemporal fusion feature vector output by the module as the data basis for evaluation calculation. The spatiotemporal fusion feature vector has 512 dimensions, with the first 256 dimensions encoding the temporal dimension of motion continuity features and the latter 256 dimensions encoding the spatial dimension of limb coordination features. The multidimensional evaluation engine 4 includes a posture accuracy evaluation unit, a rhythm matching degree evaluation unit, a motion fluency evaluation unit, and an expressiveness evaluation unit. Each unit decodes the kinematic parameters of the corresponding dimension from the spatiotemporal fusion feature vector, thus quantitatively evaluating dance movements from different dimensions.

[0042] The posture accuracy assessment unit is used to calculate the spatial distance between the dancer's current posture and the standard posture at multiple joints and generate an accuracy score.

[0043] The posture accuracy assessment unit first retrieves the standard posture sequence corresponding to the current movement segment from the standard movement database. The standard movement database pre-stores standard movements for different dance types, performed by professional dancers and verified by experts. Using a movement recognition algorithm, the system automatically matches the dance type and specific movement name of the current movement segment and extracts the corresponding standard posture sequence from the database.

[0044] Posture accuracy is calculated based on multi-joint Euclidean distance and angle deviation. For each time step, the differences in joint position and joint angle between the dancer's posture and the standard posture are calculated. The joint position difference is measured using a three-dimensional Euclidean distance, calculated using the following formula:

[0045] ,

[0046] in, The positional difference of the i-th joint. Let i be the three-dimensional coordinates of the i-th joint of the dancer. Let represent the three-dimensional coordinates of the i-th joint in the standard pose. This formula calculates the spatial distance between the dancer's joint and the standard joint; the smaller the distance, the more accurate the pose.

[0047] Joint angle discrepancies are calculated based on the angles between bone vectors. For each joint, a joint angle is defined by adjacent bones; for example, the elbow angle is determined by the angle between the upper arm and forearm bones. The formula for calculating joint angle deviation is:

[0048] ,

[0049] in, Let the angle deviation of the j-th joint be , For the angle of the dancer's j-th joint, Let be the angle of the j-th joint in the standard posture. Joint angles reflect the degree of limb flexion and postural shape, and are an important indicator for assessing postural accuracy.

[0050] The attitude accuracy score is calculated by combining positional differences and angular deviations. The score uses an exponential decay function, mapping distance and angular deviations to a score range of 0 to 100.

[0051] ,

[0052] in, For attitude accuracy scoring, K is the total number of joints (K=25). The position sensitivity coefficient, This is the angle weighting coefficient. In a preferred embodiment, A value of 0.5 allows the score to respond significantly to positional deviations of 5cm to 10cm. The value is set to 0.3, making the weight of the angle deviation on the score approximately 30% of that of the position deviation. This formula achieves a non-linear mapping through an exponential decay function, so that small deviations have little impact on the score, while large deviations will significantly reduce the score, which is consistent with the human perception of posture accuracy.

[0053] The rhythm matching evaluation unit is used to extract music beat features and calculate the time alignment between motion keyframes and music beats.

[0054] The rhythm matching evaluation unit extracts beat timestamps from the input audio signal. The audio signal is converted to a time-frequency domain representation through a short-time Fourier transform, and then a neural network-based beat detection algorithm is used to identify the strong and weak beats in the music. The beat detection algorithm outputs a timestamp sequence, with each timestamp marking the occurrence time of a beat. Preferably, a beat tracking model based on a combination of convolutional neural networks and recurrent neural networks is used, which can accurately identify rhythmic patterns in different styles of music.

[0055] Keyframes in dance movements refer to the moments when the amplitude of the movement reaches its peak or when the direction changes. Keyframe detection is based on changes in joint velocity and acceleration. For each joint, its velocity across consecutive frames is calculated, and the moment corresponding to the peak velocity is the keyframe for that joint. Additionally, moments when acceleration changes from positive to negative or vice versa, representing turning points in the direction of movement, are also considered keyframes. By combining keyframe information from all relevant nodes, a clustering algorithm is used to merge keyframes that are temporally close, resulting in a keyframe sequence for the overall movement.

[0056] Rhythm matching is calculated based on the alignment between keyframe timestamps and music beat timestamps. For each action keyframe, the music beat closest in time is found, and the time difference is calculated. The smaller the time difference, the higher the rhythm matching between the action and the music. The formula for calculating the rhythm matching score is:

[0057] ,

[0058] in, The rhythm matching score is given, where M is the total number of keyframes. The time difference between the m-th keyframe and the most recent beat. This is the time sensitivity coefficient. In a preferred embodiment, A value of 10.0 ensures that the score responds significantly to time deviations ranging from 50ms to 100ms. This formula also employs an exponential decay function, ensuring that keyframes with smaller time deviations contribute more to the score, while keyframes with larger time deviations have a smaller impact.

[0059] The motion smoothness evaluation unit is used to calculate the smoothness index of motion based on the changes in joint velocity and acceleration between adjacent frames.

[0060] Fluency reflects the continuity and smoothness of dance movements. Fluent dance movements are characterized by smooth changes in the velocity and acceleration of joint movements, while non-fluent movements are characterized by abrupt changes in velocity and acceleration. The fluidity assessment unit quantifies fluidity by analyzing the time derivatives of joint movements.

[0061] Joint velocity is calculated by dividing the position difference between adjacent frames by the time interval:

[0062] ,

[0063] in, Let be the velocity vector of the i-th joint at time t. Let be the position vector of the key point at time t. This represents the inter-frame time interval. The velocity vector reflects the direction and speed of motion of the joints.

[0064] Joint acceleration is calculated by the change in velocity between adjacent time points:

[0065] ,

[0066] in, Let be the acceleration vector of the i-th joint at time t. Acceleration reflects the rate of change of motion velocity; abrupt changes in acceleration indicate a lack of smoothness in the motion.

[0067] The smoothness of movement is assessed based on the rate of change of velocity and the root mean square deviation of acceleration. The rate of change of velocity represents the degree of fluctuation in velocity and is calculated as the standard deviation of the velocity vector magnitude. The root mean square deviation of acceleration represents the degree of dispersion of acceleration and is calculated as the root mean square of the acceleration vector magnitude. The formula for calculating the smoothness score is:

[0068] ,

[0069] in, Rate the smoothness of the movement. The standard deviation of the magnitude of the velocity vector. The root mean square of the magnitude of the acceleration vector. Sensitivity coefficient This is the acceleration weighting coefficient. In a preferred embodiment, The value is 0.1. A value of 2.0 is used to ensure the score reflects both velocity fluctuations and acceleration changes on smoothness. Smaller standard deviations and root mean squares indicate smoother, more fluid motion and higher scores.

[0070] The expressiveness assessment unit is used to evaluate the expressiveness of a dance by analyzing the range of limb movements and changes in the body's center of gravity.

[0071] The expressiveness of dance is reflected in the tension, range of motion, and control of the center of gravity. Expressive dance movements are characterized by full extension of the limbs, rich variations in the range of motion, and flexible control of the body's center of gravity. The expressiveness assessment unit extracts quantitative indicators related to expressiveness from kinematic characteristics.

[0072] Limb movement amplitude is calculated using statistical characteristics of joint range of motion. For each joint, its positional distribution throughout the entire movement sequence is calculated, including the range of motion in the x, y, and z axes. A larger range of motion indicates greater range of motion and stronger expressiveness. Simultaneously, the standard deviation of the joint position is calculated; a larger standard deviation indicates richer movement variation. The movement amplitude score is calculated by combining the range of motion and standard deviation.

[0073] ,

[0074] in, As an indicator of the range of motion, , , These represent the range of motion of the k-th joint along the x, y, and z axes, respectively. Let be the standard deviation of the position of the k-th joint. This is a weighting coefficient. In a preferred embodiment, The value is set to 0.5 to balance the contribution of activity range and variability.

[0075] Shifts in the body's center of gravity reflect a dancer's balance control and dynamic expressiveness. The center of gravity is calculated as a weighted average of all relevant points on the body, with weights determined based on the mass distribution of different body parts. The torso and head account for a larger proportion of mass, while the limbs account for a smaller proportion. Changes in the trajectory of the center of gravity are comprehensively assessed using the standard deviation of the center of gravity position and the length of the trajectory. Dramatic shifts in the center of gravity indicate strong dynamism and rich expressiveness in the dance movements.

[0076] The performance score incorporates indicators of range of motion and shift in center of gravity.

[0077] ,

[0078] in, To score performance, This is an indicator of the change in the center of gravity during the current action. and These are the baseline values ​​for the range of motion and center of gravity shift of the reference movement, determined through statistical analysis of a large amount of dance movement data. The formula uses a normalization method to map expressiveness to a scoring range of 0 to 100, with higher scores indicating stronger expressiveness.

[0079] Reference Figure 4 The adaptive weight adjustment module 5 is connected to the multi-dimensional evaluation engine 4. It is used to dynamically determine the weight coefficients of each evaluation dimension according to the dance type and evaluation scenario, and to generate a comprehensive score by weighting and fusing the scores of each dimension based on the weight coefficients.

[0080] Different dance styles place varying degrees of emphasis on each evaluation dimension. For example, classical ballet emphasizes standardized postures and precise movements, so posture accuracy should be given a higher weight; modern dance focuses on individual expression and emotional delivery, so expressiveness should be given a higher weight; rhythmic dances such as street dance and jazz dance have strict requirements for mastering musical rhythm, so rhythm matching should be given a higher weight. The adaptive weight adjustment module 5 can automatically adjust the weight allocation scheme according to the dance style.

[0081] In one embodiment of the present invention, the system presets weight allocation schemes for various dance types, which are stored in a weight configuration database. The weight configuration database includes weight schemes for common dance types such as classical ballet, modern dance, folk dance, street dance, jazz dance, and Latin dance. Each dance type corresponds to four weight coefficients, representing the weights of posture accuracy, rhythm matching, movement fluency, and expressiveness, respectively. The weight coefficients range from 0 to 1, and the sum of the four weight coefficients is 1, satisfying the normalization constraint.

[0082] For example, for classical ballet, the weighting scheme is: posture accuracy 0.45, rhythm matching 0.20, movement fluidity 0.25, and expressiveness 0.10. This weighting reflects the high requirements of classical ballet for posture precision. For modern dance, the weighting scheme is: posture accuracy 0.20, rhythm matching 0.20, movement fluidity 0.25, and expressiveness 0.35. This weighting highlights modern dance's emphasis on expressiveness. For street dance, the weighting scheme is: posture accuracy 0.15, rhythm matching 0.45, movement fluidity 0.20, and expressiveness 0.20. This weighting reflects street dance's emphasis on rhythm.

[0083] After the user selects the current dance type in the system, the adaptive weight adjustment module 5 loads the corresponding weight coefficients from the weight configuration database. In some embodiments, the system also supports user-defined weight coefficients, allowing teachers or dancers to adjust the weight allocation according to specific training goals, thereby achieving personalized assessment.

[0084] The overall score is calculated based on a weighted fusion method:

[0085] ,

[0086] in, For comprehensive scoring, , , , These are the weighting coefficients for posture accuracy, rhythm matching, movement fluency, and expressiveness, respectively. For accuracy scoring, Rate the rhythm matching Scoring for the smoothness of movement, For performance scoring, the four scores correspond one-to-one with the four evaluation units of the multi-dimensional evaluation engine. The weighting coefficients satisfy... This formula uses a linear weighted fusion of scores from various dimensions to generate a comprehensive score that reflects the overall quality of the dance. The comprehensive score ranges from 0 to 100, with a higher score indicating better dance quality.

[0087] In a preferred embodiment of the present invention, the adaptive weight adjustment module 5 further includes a weight optimization function based on historical data. The system records scoring data from professional judges on a large number of dance performances, learns the scoring patterns of the professional judges through regression analysis, and optimizes the weight coefficients to maximize the correlation between the system scores and the professional judge scores. The weight optimization employs a gradient descent algorithm to minimize the mean squared error between the system scores and the professional judge scores.

[0088] ,

[0089] Where L is the loss function and N is the number of samples. The overall score for the nth sample. The system assigns scores to the professional judges for the nth sample. By minimizing the loss function and adjusting the weight coefficients, the system scores approximate the professional judges' scores, thereby improving the accuracy and reliability of the evaluation.

[0090] The real-time feedback generation module 6 is connected to the adaptive weight adjustment module 5. It is used to generate visual feedback information based on the comprehensive score and the scores of each dimension, and to mark the action details and time nodes that need to be improved.

[0091] The core function of Real-Time Feedback Generation Module 6 is to transform quantitative evaluation results into intuitive and easy-to-understand visual feedback, helping dancers quickly understand their strengths and weaknesses. Visual feedback includes various forms such as score display, movement annotation, and improvement suggestions.

[0092] On the display interface, the overall score is prominently displayed in large numbers in the center of the screen, allowing dancers to quickly understand their overall performance level. Next to the overall score, the score level is displayed; for example, 90 points or above is excellent, 80-89 points is good, 70-79 points is average, 60-69 points is passable, and below 60 points is failable. Simultaneously, the scores for each of the four dimensions are displayed in radar chart format, allowing dancers to intuitively compare their performance levels in each dimension. The four vertices of the radar chart correspond to posture accuracy, rhythm matching, movement fluidity, and expressiveness, respectively. The current score is marked in red on the radar chart, while the standard reference score is marked in blue. The degree of overlap between the two reflects how close the dancer is to the standard level.

[0093] For joints with low posture accuracy, the real-time feedback generation module 6 marks them with different colors in the video playback. Specifically, when the joint position error is less than 30mm and the angle error is less than 10°, it is marked in green to indicate accurate posture; when the joint position error is between 30mm and 50mm or the angle error is between 10° and 15°, it is marked in yellow to indicate slight posture deviation; when the joint position error is greater than 50mm or the angle error is greater than 15°, it is marked in red to indicate significant posture deviation, and this joint is judged as a movement detail that needs improvement. The markings are superimposed on the joint position in the form of circles or highlighted boxes, allowing dancers to clearly see which parts of the posture need improvement. At the same time, time nodes with significant posture deviations are marked on the timeline, and dancers can click on the time nodes to jump to the corresponding video frames to view the specific posture problems.

[0094] For keyframes with low rhythm matching, the real-time feedback generation module 6 marks the music beat position and the motion keyframe position on the timeline, with the deviation between the two indicated by arrows or lines. If the motion keyframe is earlier than the music beat, the arrow points backward, prompting the dancer to delay the movement; if the motion keyframe is later than the music beat, the arrow points forward, prompting the dancer to advance the movement. The time difference between the beat and the keyframe is displayed in milliseconds, allowing the dancer to accurately understand the magnitude of the rhythm deviation.

[0095] For segments with low fluidity, the real-time feedback generation module 6 displays the change in joint movement speed over time using a speed curve graph. Fluent movements are characterized by a smooth and continuous speed curve, while non-fluid movements are characterized by abrupt changes or jagged fluctuations in the speed curve. The system marks the time periods with large fluctuations on the speed curve and prompts the dancers that the continuity of the movements in these periods is insufficient and requires more practice.

[0096] For assessments showing low expressiveness, the real-time feedback generation module 6 provides improvement suggestions in text or voice format. The suggestions are generated based on the specific details of the range of motion and changes in center of gravity. For example, if the range of motion is insufficient, the system prompts, "Please increase the extension of your arms, fully extending your elbows and wrists"; if the changes in center of gravity are insufficient, the system prompts, "Please increase the forward and backward movement of your body, controlling the dynamic changes in your center of gravity." The improvement suggestions are expressed in concise and clear language, avoiding the use of technical jargon, ensuring that dancers can understand and implement them.

[0097] In a preferred embodiment, the real-time feedback generation module 6 also supports a comparison playback function. The system displays the dancer's performance video and the standard movement video side by side, allowing the dancer to intuitively compare the differences between their movements and the standard movements through synchronized playback. During the comparison playback, the two videos maintain the same playback speed and time alignment, enabling the dancer to observe the differences in posture frame by frame. In addition, the system also provides a slow-motion playback function, reducing the video playback speed to 25% to 50%, allowing the dancer to carefully observe the details of the movements and changes in posture.

[0098] Reference Figure 5 The historical data management module 7 is connected to the real-time feedback generation module 6. It is used to store the dancer's historical evaluation data and perform longitudinal comparative analysis to generate progress trajectory curves and development trend predictions.

[0099] The historical data management module 7 uses a time-series database to store dancers' historical scoring data. After each dance evaluation is completed, the system automatically saves the evaluation results, including the evaluation date, dance type, overall score, scores for each dimension, keyframe pose data, and other information. The time-series database organizes data in chronological order, supporting efficient time-range queries and trend analysis.

[0100] The historical data management module 7 provides a longitudinal comparative analysis function, displaying the dancer's score trends over a period of time. The progress trajectory curve is presented in the form of a line graph, with time on the horizontal axis and scores on the vertical axis. Different colored curves represent the overall score and the scores for each dimension, respectively. By observing the trend of the curve, dancers can understand in which areas they have made progress and in which areas still need improvement. For example, if the posture accuracy curve continues to rise, it indicates that the dancer's posture standardization is constantly improving; if the fluidity curve fluctuates greatly, it indicates that the dancer's movement stability needs to be improved.

[0101] The historical data management module 7 uses regression analysis to predict dancers' future skill development trends. Regression analysis fits a trend model based on historical scoring data; commonly used models include linear regression, multinomial regression, and exponential regression. Linear regression is suitable for scores showing a linear growth trend, multinomial regression is suitable for scores that show acceleration or deceleration, and exponential regression is suitable for scores that have grown rapidly and then stabilized. The system automatically selects the model with the best fit for trend prediction, forecasting the score trend over the next week to month. The prediction results are displayed as a dashed line on the progress trajectory curve, allowing dancers to anticipate their progress rate and the time it will take to reach their target level.

[0102] In a preferred embodiment, the historical data management module 7 also provides a personalized training suggestion function. The system analyzes the dancer's historical data, identifies dimensions where scores are growing slowly or stagnating, and recommends targeted training programs. For example, if posture accuracy is consistently lower than other dimensions, the system recommends strengthening basic skills practice, focusing on correcting posture accuracy; if rhythm matching is low, the system recommends rhythm training, listening to more music and practicing hitting the beat; if fluency is insufficient, the system recommends continuity training to reduce pauses and stiffness in movements. The training suggestions are based on a large number of dancers' training experience and are highly targeted and practical.

[0103] The historical data management module 7 also supports multi-dancer comparative analysis. In dance training institutions or groups, teachers can view the scoring data of multiple dancers for horizontal comparison. The system generates an overall scoring distribution chart for the class or team, showcasing the scoring levels and differences among dancers. Through comparative analysis, teachers can understand the overall training effect, identify dancers requiring focused attention, and develop differentiated teaching strategies. Simultaneously, dancers can also view their ranking and relative level within the team, stimulating training motivation and a sense of competition.

[0104] The application scenarios and actual effects of the system of the present invention are illustrated below through a specific embodiment.

[0105] In a dance training institution, teachers use the system of this invention to assess students' classical ballet training. The system is equipped with an RGB camera with a shooting frame rate of 90fps, installed 4 meters away from the dancer in front of the training room, with the camera's optical axis at a 12° angle to the ground. Students perform a 2-minute classical ballet segment to music, and the system captures and analyzes the video stream in real time.

[0106] Image acquisition module 1 captures video at a frame rate of 90fps, obtaining a total of 10,800 frames. After human detection and region cropping, the trainee's human image sequence is extracted, with each frame measuring 512×512 pixels.

[0107] The multi-scale feature extraction module 2 employs the Vision Transformer architecture to extract image features, combining 3×3, 5×5, and 7×7 multi-scale convolutional kernels for differential feature extraction. The system successfully identified 25 key points of the trainee's skeletal structure, including key locations such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. The 3D pose estimation accuracy reached an average positional error of 12mm, meeting the requirements for high-precision evaluation.

[0108] The spatiotemporal attention fusion module 3 performs spatiotemporal correlation analysis on the posture sequence data. The temporal dimension attention mechanism detects insufficient continuity in the trainees' turning movements, while the spatial dimension attention mechanism reveals poor upper and lower limb coordination in the trainees' leg-raising movements. The spatiotemporal fusion feature vector comprehensively encodes these spatiotemporal features, providing rich information for subsequent evaluation.

[0109] The multi-dimensional evaluation engine 4 assesses the trainee's performance from multiple dimensions. The posture accuracy evaluation unit calculates the differences in joint positions and angle deviations between the trainee's posture and the standard posture, resulting in a posture accuracy score of 82, indicating that the trainee's posture is basically correct but still has room for improvement. The rhythm matching evaluation unit analyzes the timing alignment between the trainee's keyframes and the music beat, finding a 0.15s to 0.20s beat delay in some movements, resulting in a rhythm matching score of 75. The movement fluidity evaluation unit calculates the rate of change of velocity and the root mean square error of acceleration at joint movements, finding sudden changes in velocity when connecting movements, resulting in a fluidity score of 78. The expressiveness evaluation unit analyzes the trainee's range of motion and center of gravity changes, finding that the trainee's arm extension is good but body center of gravity control is slightly stiff, resulting in an expressiveness score of 70.

[0110] The adaptive weight adjustment module 5, based on the weighting scheme of classical ballet, sets the weight for posture accuracy to 0.45, rhythm matching to 0.20, movement fluency to 0.25, and expressiveness to 0.10. Based on these weight coefficients, the system calculates a comprehensive score of 78.5 points, which is rated as medium.

[0111] The real-time feedback generation module 6 generates visual feedback information. On the display interface, the overall score of 78.5 is shown in large numbers, with the rating marked "Medium". The radar chart shows the sub-scores across four dimensions, allowing the learner to visually see that while the posture accuracy is relatively high, the expressiveness is relatively low. In video playback, the system highlights the left wrist and right knee joints in yellow, indicating deviations in these postures. On the timeline, the system marks the segment from 45s to 50s, noting that the rhythm of this segment is lagging by approximately 0.18s, and suggesting that the learner initiate the torso rotation earlier during this period. The system also generates text suggestions: "Please increase the extension range of your right arm, ensuring your arm is in a straight line with your shoulder; initiate the torso rotation 0.2s earlier to precisely align the movement with the music beat."

[0112] The historical data management module 7 saves the assessment data and compares it with the trainee's assessment records from the past month. The progress trajectory curve shows that the trainee's posture accuracy improved from 73 points one month ago to 82 points, an increase of 9 points, demonstrating significant progress. The fluency score improved from 76 points to 78 points, a relatively slow improvement. The expressiveness score fluctuated around 70 points, without significant improvement. The system predicts through regression analysis that if the trainee maintains the current training intensity, the overall score is expected to reach 85 points within the next two weeks, reaching a good level. At the same time, the system generates personalized training suggestions: "The current expressiveness dimension is progressing slowly. It is recommended to increase body control training, practice center of gravity transfer and dynamic balance, and conduct at least 3 specific training sessions per week."

[0113] After three months of continuous training and systematic evaluation, the trainees' overall score improved from an initial 65 points to 88 points, with significant improvements across all dimensions. Teachers reported that the objective quantitative assessment and real-time feedback provided by this system greatly improved teaching efficiency, enabling trainees to identify and correct movement problems more quickly, resulting in training effects significantly superior to traditional manual instruction methods.

[0114] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A real-time analysis and evaluation system for dance movements based on computer vision, characterized in that, include: The image acquisition module is used to capture the video stream of the dance performance scene in real time through a high-precision camera device, and to extract and preprocess the video stream frames. A multi-scale feature extraction module, connected to the image acquisition module, is used to extract multi-scale features from the preprocessed video frames, identify the coordinates of key points of the dancer's three-dimensional human skeleton, and generate human posture sequence data. The spatiotemporal attention fusion module, connected to the multi-scale feature extraction module, is used to perform spatiotemporal correlation analysis on human posture sequence data. It captures the continuity features of actions through a time-dimensional attention mechanism and captures the limb coordination features through a spatial-dimensional attention mechanism, generating a spatiotemporal fusion feature vector. A multi-dimensional evaluation engine, connected to the spatiotemporal attention fusion module, is used to perform multi-dimensional quantitative evaluation of dance movements based on the spatiotemporal fusion feature vector. This includes a posture accuracy evaluation unit, a rhythm matching evaluation unit, a movement fluency evaluation unit, and an expressiveness evaluation unit. The posture accuracy evaluation unit extracts joint spatial coordinates from the spatiotemporal fusion feature vector, calculates the spatial distance between the dancer's current posture and a standard posture, and generates an accuracy score. The rhythm matching evaluation unit extracts music beat features and calculates the time alignment between movement keyframes and music beats. The movement fluency evaluation unit calculates the smoothness index of the movement based on joint speed and acceleration changes between adjacent frames. The expressiveness evaluation unit evaluates the expressiveness of the dance by analyzing the amplitude of limb movements and changes in the body's center of gravity. An adaptive weight adjustment module, connected to the multi-dimensional evaluation engine, is used to dynamically determine the weight coefficients of each evaluation dimension according to the dance type and evaluation scenario, and to generate a comprehensive score by weighting and fusing the scores of each dimension based on the weight coefficients. The real-time feedback generation module, connected to the adaptive weight adjustment module, is used to generate visual feedback information based on the comprehensive score and the scores of each dimension, and to mark the positions of joints where the posture deviation exceeds the preset deviation threshold and the time nodes where the rhythm deviation exceeds the preset time threshold. The historical data management module, connected to the real-time feedback generation module, is used to store the dancer's historical evaluation data and perform longitudinal comparative analysis to generate progress trajectory curves and development trend predictions.

2. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The high-precision camera device includes an RGB camera or an RGB-D depth camera, with a shooting frame rate of 60fps to 120fps.

3. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The multi-scale feature extraction module uses the Vision Transformer architecture to extract image features and combines multi-scale convolution kernels to extract differentiated features for different body parts.

4. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The temporal dimension attention mechanism of the spatiotemporal attention fusion module is used to analyze the motion trajectory of joints between consecutive frames, while the spatial dimension attention mechanism is used to analyze the relative positional relationship between joints within a single frame.

5. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The posture accuracy evaluation unit generates an accuracy score by calculating the Euclidean distance and angular deviation between the dancer's keyframe posture and the standard posture at multiple joints.

6. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The rhythm matching evaluation unit extracts the beat timestamp from the input audio signal and calculates the time difference between the moment when the dancer's key movements occur and the timetamp of the music beat.

7. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The motion smoothness evaluation unit assesses the continuity and smoothness of motion by calculating the rate of change of velocity vectors and the root mean square error of acceleration at joints between adjacent frames.

8. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The adaptive weight adjustment module presets different weight allocation schemes according to the dance type. For classical ballet, it emphasizes the weight of posture accuracy; for modern dance, it emphasizes the weight of expressiveness; and for rhythm dance, it emphasizes the weight of rhythm matching.

9. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The real-time feedback generation module marks joints with large posture deviations in different colors on the display interface and provides improvement suggestions in text or voice form.

10. The real-time analysis and evaluation system for dance movements based on computer vision according to claim 1, characterized in that, The historical data management module uses a time-series database to store dancers' historical rating data and uses regression analysis to predict dancers' future skill development trends.

Citation Information

Patent Citations

  • Dancing motion evaluation system based on deep learning and attitude estimation

    CN113516005A