Dance movement image analysis method, system and equipment based on image time sequence recognition
Through the method of image segmentation and multimodal feature fusion, the displacement trajectory and high-dimensional features of dance movements were extracted, and the timing model was used to use long and short-term memory networks to construct a dance movement representation model, which solved the problems of displacement trajectory capture and motion recognition in dance movement image analysis, and achieved high-accuracy dance movement analysis and recognition.
Patent Information
- Application Number
- CN202510152916.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In dance action image analysis, how to accurately capture and quantify the displacement trajectors at different time points, especially when facing the complexity and diversity of dance moves, differences in clothing and body shape, and changes in the stage environment.
By acquiring dance action image sequences, an image segmentation algorithm such as MaskR-CNN is used to determine the dancer's pixel position information, calculate the initial displacement trajectory, and extract high-dimensional features. Combining joint angle and limb posture data, a multimodal feature matrix was constructed, long and short-term memory network was used for timing modeling, key displacement patterns were extracted through cluster analysis, dance movement representation model was constructed, and classification training was performed using support vector machines.
It realizes accurate identification and classification of dance movements, can effectively capture the timing characteristics and semantic information of the movements, build an accurate dance movement representation model, and improves the accuracy of analysis and recognition.
Smart Images

Figure CN120088854A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to the analysis of dance movement images for image time series recognition. Background Art
[0002] In the analysis of dance movement images, how to accurately capture and quantify the displacement trajectories of dancers at different time points is a key technical problem. Due to the complexity and diversity of dance movements, it is difficult to comprehensively characterize the temporal features of movements solely relying on the pixel position changes of dancers in the images. In addition, the differences in dancers' costumes, body types, and the changes in the stage environment will all affect the accuracy of displacement estimation.
[0003] At the same time, considering the continuity and smoothness of dance movements, methods of time series analysis can be introduced to perform smoothing filtering and interpolation on the displacement trajectories to reduce the displacement estimation errors caused by image quality and frame rate limitations. Another issue that needs attention is how to extract displacement information related to the semantics of dance movements from high-dimensional image features, and fuse it with other temporal features such as the joint angles and limb postures of dancers to construct a more comprehensive and accurate dance movement representation model. Summary of the Invention
[0004] The present invention provides an analysis of dance movement images for image time series recognition, mainly including:
[0005] Obtain a sequence of dance movement images, and determine the pixel position information of the dancer through image segmentation; calculate the initial displacement trajectory of the dancer according to the change trend of the pixel position information of the dancer; extract the high-dimensional features of the dancer in the image sequence, and combine with the initial displacement trajectory to extract displacement information related to the semantics of dance movements; obtain the joint angle and limb posture data of the dancer, perform normalization processing on the joint angle and limb posture data and the semantic displacement information, and construct a multi-modal feature matrix; use a long short-term memory network to perform temporal modeling processing on the multi-modal feature matrix to obtain a temporal modeling result; extract key displacement patterns through clustering analysis according to the temporal modeling result, and construct a dance movement representation model; use a support vector machine to perform classification training on the dance movement representation model to obtain a classification training result; optimize the parameters of the dance movement representation model according to the classification training result.
[0006] Further, the obtaining of the dance movement image sequence and the determination of the dancer pixel position information through image segmentation include: obtaining a continuous image sequence containing dance movements, and using the MaskR-CNN algorithm based on deep learning to perform pixel-level segmentation on the dancer region in the image; obtaining a binary mask image of the dancer region in each frame of the image, where the dancer pixel positions are marked with a specific value and the background region pixels are marked with another specific value; performing a bitwise AND operation on the binary mask image and the original image frame to extract the pixel information of the dancer region; performing a morphological closing operation on the pixel information of the dancer region to eliminate holes and break point noises in the image; and extracting the pixel coordinates of the dancer contour through a contour extraction algorithm to determine the pixel coordinate range of the dancer region in each frame of the image.
[0007] Further, the calculation of the initial displacement trajectory of the dancer according to the change trend of the dancer pixel position information includes: for each frame of the image, using a human pose estimation model to detect and locate the key body part coordinates of the dancer; using the Hungarian algorithm to match the body part coordinates detected in two adjacent frames of the image to establish the correspondence of the same dancer between different frames; according to the correspondence, calculating the displacement vector of each dancer between two adjacent frames and connecting them in chronological order to form an initial displacement trajectory sequence; applying the Kalman filter algorithm to the initial displacement trajectory sequence for smoothing processing to estimate the true displacement trajectory of the dancer; and according to the true displacement trajectory, calculating the instantaneous velocity vector and acceleration vector of the dancer at each moment through numerical differentiation.
[0008] Further, the extraction of the high-dimensional features of the dancer in the image sequence and the extraction of the displacement information related to the dance movement semantics in combination with the initial displacement trajectory include: for each frame of the image in the video sequence, extracting its high-dimensional feature information such as color histogram, local binary pattern texture feature, and edge feature; performing normalization processing on the high-dimensional feature information to scale the feature values into a specific interval; inputting the normalized high-dimensional feature information into a pre-trained convolutional neural network, and realizing the reduction of the feature dimension through multiple iterations of the convolutional layer and the pooling layer; in the first frame of the video sequence, using an object detection algorithm to obtain the initial position coordinates of the dancer; for each frame of the image in the video sequence, splicing the low-dimensional features output by the convolutional neural network and the reference point coordinates into a vector; and calculating the Euclidean distance between the vector and the reference point coordinates to obtain the displacement vector of the current frame dancer position relative to the initial position.
[0009] Further, the method of obtaining the joint angles and limb posture data of the dancer, normalizing the joint angles and limb posture data with the semantic displacement information, and constructing a multi-modal feature matrix includes: obtaining the joint angle data and limb posture data of the dancer through sensors and preprocessing them to remove outliers and noise; using a word embedding model to extract semantic information from the text description of the dance movement and mapping each word to a real number vector with a fixed dimension; performing min-max normalization on the joint angle data, limb posture data, and semantic displacement vector respectively to uniformly map the numerical range to a specific interval; aligning the three normalized data according to the time stamp; and concatenating the aligned joint angle data, limb posture data, and semantic displacement vector according to the feature dimension to construct a multi-modal feature matrix.
[0010] Further, the method of performing temporal modeling on the multi-modal feature matrix using a long short-term memory network to obtain a temporal modeling result includes: performing temporal modeling on the multi-modal feature matrix using a long short-term memory network, controlling the information flow through the input gate, forget gate, and output gate of the network, and capturing the continuity and smoothness characteristics of the dance movement; at each time step, updating the long-term dependence information in the memory unit according to the input feature at the current time and the hidden state at the previous time step, and generating the hidden feature representation at the current time step through the output gate; repeating the above process until all time steps of the feature matrix are processed to obtain a complete temporal modeling result; and inputting the temporal modeling result into a fully connected layer to classify the dance movement through an activation function.
[0011] Further, the method of extracting key displacement patterns through cluster analysis based on the temporal modeling result and constructing a dance movement representation model includes: modeling and analyzing the action displacement characteristics using a time series analysis method according to the obtained dance movement data to obtain an action displacement feature sequence; performing feature clustering on the action displacement feature sequence through a clustering algorithm, and if the distance between features is less than a preset threshold, classifying them into the same category to obtain multiple key displacement patterns; constructing a dance movement representation model according to the key displacement patterns; modeling the key displacement patterns using a hidden Markov model, taking the key displacement patterns as the observation states of the hidden Markov model, defining the hidden state as the basic unit of the dance movement, and estimating the transition probability and emission probability parameters of the hidden Markov model through a specific algorithm to obtain a dance movement hidden Markov model.
[0012] Further, optimizing the parameters of the dance movement representation model according to the classification training result includes: obtaining the classification training result of the dance movement representation model on the training set; according to the classification training result, using a hyperparameter optimization method to optimize and adjust the model parameters of the dance movement representation model; using the optimized parameters to retrain the dance movement representation model on the training set to obtain an optimized model; evaluating the optimized dance movement representation model on the test set to obtain the classification accuracy of the model on the test set; if the classification accuracy is lower than a preset threshold, triggering a fusion weight adjustment mechanism; the fusion weight adjustment mechanism optimizes the weight coefficients of the modal features through a gradient descent algorithm based on the contribution degree of different modal features to the classification result; applying the optimized modal fusion weight to multi-modal feature fusion to generate an optimized fusion feature matrix.
[0013] A dance movement image analysis system for image time series recognition, based on the above-mentioned dance movement image analysis method for image time series recognition, includes: an image acquisition and segmentation module, a displacement trajectory calculation module, a high-dimensional feature extraction module, a posture data acquisition and normalization module, a processing module, and an evaluation and adjustment module;
[0014] The image acquisition and segmentation module is used to acquire a dance movement image sequence and determine the pixel position information of the dancer through image segmentation;
[0015] The displacement trajectory calculation module is used to calculate the initial displacement trajectory of the dancer according to the change trend of the pixel position information of the dancer;
[0016] The high-dimensional feature extraction module is used to extract the high-dimensional features of the dancer in the image sequence, and combine the initial displacement trajectory to extract the displacement information related to the semantics of the dance movement;
[0017] The posture data acquisition and normalization module is used to acquire the joint angle and limb posture data of the dancer, normalize the joint angle and limb posture data and the semantic displacement information, and construct a multi-modal feature matrix;
[0018] The processing module is used to propose and establish a model for the collected information, and perform classification training and optimize data parameters on the data;
[0019] The evaluation and adjustment module is used to evaluate the model performance on the test set and trigger a fusion weight adjustment mechanism to optimize the weight coefficients of different modal features.
[0020] A computer device includes: a memory and a processor; the memory stores a computer program, and is characterized in that: when the processor executes the computer program, the steps of the above-mentioned dance movement image analysis method for image time series recognition are implemented.
[0021] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0022] The present invention discloses a method for analyzing dance motion images by image time series recognition. The method first performs segmentation processing on the dance image sequence to extract the initial displacement trajectory and high-dimensional features of the dancer. Through convolutional neural network dimensionality reduction and semantic association, displacement information related to the meaning of dance movements is obtained. Combining joint angle and limb posture data, a multi-modal feature matrix is constructed. A long short-term memory network is used for time series modeling to capture the continuity and smoothness of movements. Key displacement patterns are extracted through clustering analysis to construct a dance motion representation model. Finally, a support vector machine is used for classification training, and the model performance is iteratively optimized by adjusting the fusion weights. The present invention can effectively identify and classify the displacement characteristics of different dance movements, providing a new technical solution for dance motion analysis and recognition. Brief Description of the Drawings
[0023] Figure 1 It is a flowchart of a method, system, and device for analyzing dance motion images by image time series recognition according to the present invention. Detailed Embodiments
[0024] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] Such as Figure 1 , a method for analyzing dance motion images by image time series recognition in this embodiment may specifically include:
[0026] S101. Obtain a dance motion image sequence, perform segmentation processing on the dancer area in the image sequence to obtain the pixel position information of the segmented dancer.
[0027] Obtain a continuous image sequence containing dance movements to form the original image dataset to be processed. For each frame image in the image sequence, use the MaskR-CNN algorithm based on deep learning to perform pixel-level segmentation on the dancer region in the image. After the segmentation process, a binary mask image of the dancer region in each frame image is obtained, where the pixel positions of the dancer are marked as 1 and the pixel positions of the background region are marked as 0. Perform a bitwise AND operation on the segmented binary mask image and the original image frame to extract the pixel information of the dancer region and obtain the segmented image of the dancer region. Perform a morphological closing operation on the segmented dancer image, use a structuring element to perform dilation and erosion operations on the image, eliminate noises such as holes and breakpoints in the image, and obtain a complete and smooth dancer contour. Through a contour extraction algorithm, such as the findContours function, extract the pixel coordinates of the dancer contour, determine the pixel coordinate range of the dancer region in each frame image, and obtain the position information of the dancer in the image. Organize the dancer position information of each frame image in chronological order to obtain a complete dance movement pixel position sequence for subsequent dance movement analysis and recognition.
[0028] Exemplarily, dance motion analysis is an important application in the field of computer vision. By processing a continuous image sequence, the recognition and analysis of dance motions can be achieved. First, a continuous image sequence containing dance motions is obtained, which can be realized by using a high-speed camera to shoot a dance performance at a speed of 30 frames per second. For example, by shooting a 3-minute ballet performance, approximately 5400 frames of images can be obtained. Next, the MaskR-CNN algorithm is used to perform pixel-level segmentation on each frame of the image. MaskR-CNN is a deep learning-based instance segmentation algorithm that can not only detect objects in the image but also accurately segment the contours of the objects. In this process, the algorithm generates a binary mask image, where the pixel values of the dancer region are 1 and the background region is 0. For example, for an image with a resolution of 1920×1080, approximately 200,000 pixels in the mask image are marked as the dancer region. Performing a bitwise AND operation on the binary mask image and the original image can extract the pixel information of the dancer region. This step can effectively remove background interference and highlight the morphological features of the dancer. To further optimize the segmentation result, a morphological closing operation is performed on the extracted dancer image. This process includes dilation and erosion operations, which can fill small holes in the image and connect disconnected regions, thereby obtaining a more complete and smooth dancer contour. Using a contour extraction algorithm, such as the findContours function in the OpenCV library, the pixel coordinates of the dancer contour can be extracted. These coordinate information describes the exact position and pose of the dancer in the image. For example, for a standard ballet dance pose, the coordinates of approximately 2000 contour points may be obtained. Organizing the dancer position information of each frame of the image in chronological order yields a complete dance motion pixel position sequence. This sequence contains the spatio-temporal information of the dance motion and can be used for subsequent motion analysis and recognition. For example, by analyzing the position changes between consecutive frames, the movement speed and acceleration of the dancer can be calculated, thereby identifying specific actions such as rapid rotation or jumping. The advantage of this method is that it can accurately capture the details of dance motions, and can not only be used for automatic scoring and analysis of dance motions, but also be applied to fields such as dance teaching, motion capture, and virtual reality. By processing and analyzing a large amount of dance video data, a dance motion database can also be established to provide technical support for the digital preservation and inheritance of dance art.
[0029] S102. Calculate the initial displacement trajectory of the dancer in each frame of the image according to the change trend of the dancer pixel position information.
[0030] Obtain a sequence of consecutive frame images containing dancers. For each frame image, use a human pose estimation model such as OpenPose to detect and locate the coordinates of the key body parts of the dancers. Match the coordinates of the body parts detected in two adjacent frame images through the Hungarian algorithm to establish the correspondence of the same dancer between different frames. According to the matching results, calculate the displacement vector of each dancer between two adjacent frames, and connect them in chronological order to form an initial displacement trajectory sequence. Apply the Kalman filter algorithm to the initial displacement trajectory sequence for smoothing processing to estimate the true displacement trajectory of the dancers and remove the trajectory jitter noise caused by detection errors. According to the smoothed displacement trajectory, calculate the instantaneous velocity vector and acceleration vector of the dancers at each moment through numerical differentiation. Conduct K-means clustering analysis on the velocity vector sequence and acceleration vector sequence, and identify the time points where the cluster centers are located as key action change points. Take the trajectory segment between two adjacent key action change points as a basic dance action unit. Extract features such as velocity, acceleration, and duration for each dance action unit to construct a dance action feature dataset. Manually perform semantic annotation on some data and define dance action categories such as "spin", "jump", etc. Use a support vector machine (SVM) classifier to predict the categories of unannotated dance action units. To evaluate the performance of dance action recognition, randomly select a part of the samples as the test set and calculate evaluation metrics such as accuracy, recall rate, and F1 value of the recognition results. Optimize the hyperparameters of the SVM classifier through cross-validation to improve the recognition performance. Finally, obtain the complete dance action recognition results, and each action unit corresponds to a semantic label and a category label.
[0031] Exemplarily, dance motion recognition is an important application in the field of computer vision. By processing and analyzing a sequence of consecutive images, automatic recognition and classification of dance motions can be achieved. First, a high-speed camera is used to capture a dance performance at a speed of 30 frames per second, obtaining a sequence of consecutive frame images containing dancers. For example, by shooting a 3-minute modern dance performance, approximately 5400 frames of images can be obtained. Next, for each frame image, a human pose estimation model such as OpenPose is used to detect and locate the coordinates of the key body parts of the dancers. OpenPose is a real-time multi-person pose estimation algorithm based on deep learning, which can simultaneously detect the body key points of multiple people, including the head, shoulders, elbows, wrists, hips, knees, and ankles, etc. For a standard dance pose, OpenPose can usually detect the coordinates of 18 key points. To establish the correspondence between the same dancer in different frames, the Hungarian algorithm is used to match the coordinates of the body parts detected in two adjacent frame images. The Hungarian algorithm is a classic algorithm for solving the assignment problem, which can correctly correspond the key point coordinates of each dancer to their positions in the next frame when multiple dancers appear in the picture simultaneously. This step is crucial for accurately tracking the movement trajectory of the dancers. According to the matching results, the displacement vector of each dancer between two adjacent frames can be calculated. For example, if the coordinates of a dancer's right hand in the first frame are (100, 200) and in the second frame are (110, 210), then the displacement vector of this key point is (10, 10). Connecting these displacement vectors in chronological order forms the initial displacement trajectory sequence. However, due to the influence of detection errors and noise, the initial trajectory may jitter. To solve this problem, the Kalman filter algorithm is applied to the initial displacement trajectory sequence for smoothing. The Kalman filter is a recursive state estimation algorithm, which can perform an optimal estimation of the movement trajectory considering measurement noise and system dynamics. Through this step, a smoother and more accurate displacement trajectory of the dancers can be obtained. Based on the smoothed displacement trajectory, the instantaneous velocity vector and acceleration vector of the dancers at each moment can be calculated through numerical differentiation. These vectors contain rich motion information. For example, a high speed value may correspond to fast rotation or jumping motions, while a high acceleration value may indicate sudden changes or pauses in the motion. To identify the key motion change points, K-means clustering analysis is performed on the velocity vector sequence and acceleration vector sequence. K-means is a commonly used unsupervised learning algorithm, which can cluster similar data points together. Here, the time points where the cluster centers are located can be identified as the key motion change points. For example, for a 30-second dance segment, 10 - 15 key change points may be identified.Taking the trajectory segment between two adjacent key action change points as a basic dance action unit, the features of each unit can be extracted, such as average speed, maximum acceleration, duration, etc. These features constitute the dance action feature dataset, laying the foundation for subsequent action classification. By manually performing semantic annotation on some data, dance action categories can be defined, such as "spin", "jump", "slide step", etc. These annotated data will be used to train a support vector machine (SVM) classifier. SVM is a powerful supervised learning algorithm, especially suitable for dealing with classification problems in high-dimensional feature spaces. Through training, SVM can learn the decision boundaries between different dance action categories, so as to accurately predict the category of unannotated dance action units. To evaluate the performance of dance action recognition, about 20% of the samples can be randomly selected as the test set, and evaluation metrics such as accuracy, recall rate, and F1 value of the recognition results can be calculated. By optimizing the hyperparameters of the SVM classifier, such as kernel function type, regularization parameter, etc., through cross-validation, the recognition performance can be further improved. Finally, a complete dance action recognition system is obtained, which can automatically decompose continuous dance video sequences into a series of basic action units with semantic labels. This technology can not only be applied to dance teaching and scoring, but also provide strong support for fields such as dance choreography, motion capture, and virtual reality.
[0032] S103. Extract the high-dimensional features of the dancer from the image sequence. The high-dimensional features include color, texture, and edge information. Use a convolutional neural network to perform dimensionality reduction processing on the high-dimensional features, and combine the initial displacement trajectory to extract the displacement information related to the semantics of the dance action. The displacement information related to semantics refers to the spatial position change data related to the meaning of the dance action.
[0033] Obtain a video sequence containing dance movements. For each frame image in the video sequence, extract high-dimensional feature information such as its color histogram, Local Binary Pattern (LBP) texture features, and Canny edge features. Normalize the extracted high-dimensional feature information, scale the feature values to the interval [0, 1], and eliminate the dimensional differences between different features. Input the normalized high-dimensional feature information into a pre-trained convolutional neural network (such as AlexNet or VGGNet). Through multiple iterative calculations of the convolutional layer and pooling layer, reduce the feature dimension to obtain a more compact and semantically rich feature representation. In the first frame of the video sequence, use an object detection algorithm (such as YOLO or FasterR-CNN) to obtain the initial position coordinates of the dancer (such as the center point coordinates or the coordinates of the bounding rectangle) as a reference point for subsequent displacement trajectory calculation. For each frame image in the video sequence, concatenate the low-dimensional features output by the convolutional neural network with the reference point coordinates into a vector as the representation of the dancer's position in the current frame. Calculate the Euclidean distance between this vector and the reference point coordinates to obtain the displacement vector of the dancer's position in the current frame relative to the initial position. Connect the displacement vectors of multiple consecutive frames (such as 10 frames) in chronological order to form a displacement trajectory. Smooth this trajectory to remove the influence of noise and jitter. Then, use the Dynamic Time Warping (DTW) algorithm to perform similarity matching between this trajectory and a predefined typical dance movement trajectory template, and calculate the minimum matching distance between the two trajectories. If the minimum matching distance is less than a preset threshold (such as 2), mark this trajectory segment with the corresponding dance movement semantic label. To achieve the synchronization of the semantic label and the trajectory, corresponding label information can be added at the start frame and end frame of each trajectory segment. Finally, output a dictionary structure with the video frame number as the key and a tuple containing the displacement vector and the action semantic label as the value. This can facilitate subsequent tasks such as analyzing, recognizing, and generating dance movements in the video.
[0034] Exemplarily, the core of the dance movement recognition system lies in extracting effective features from the video sequence and performing intelligent analysis. First, color histograms, LBP texture features, and Canny edge features are extracted from each frame of the image. The color histogram reflects the overall color distribution of the image and can capture the color information of the dancer's clothing and the background. For example, in a modern dance performance, if the dancer is wearing a red top and black pants, and the background is white, then the histogram will show obvious peaks in the red, black, and white color regions. The LBP texture feature can describe the local texture patterns of the image and is very effective for identifying subtle changes in dance movements. For instance, when a ballet dancer makes an elegant arm movement, the texture contrast between the arm and the background will form a unique LBP pattern. The Canny edge feature can highlight the dancer's contour and movement lines, helping to capture the key shapes of dance movements. After feature extraction, normalization is performed to eliminate the dimensional differences between different features, enabling the subsequent neural network to learn better. After normalization, these high-dimensional features are input into a pre-trained convolutional neural network. Taking AlexNet as an example, it contains 5 convolutional layers and 3 fully connected layers, and can extract the abstract features of the image layer by layer. When processing dance videos, the lower-level convolutional kernels may identify simple edges and textures, while the higher-level convolutional kernels may identify complex pose and movement patterns. Object detection algorithms such as YOLO are used to locate the initial position of the dancer. YOLO divides the image into grids, and each grid predicts multiple bounding boxes and class probabilities, enabling it to quickly and accurately locate the dancer. After obtaining the initial position of the dancer, combined with the low-dimensional features output by the convolutional neural network, a vector representing the current position of the dancer can be constructed. By calculating the Euclidean distance between this vector and the initial position, the displacement information of the dancer can be obtained. The smoothing of the displacement trajectory can use methods such as moving average or Kalman filtering. This step can remove the noise caused by detection errors or the dancer's slight jitters, making the trajectory more coherent. The smoothed trajectory is matched with a predefined action template using the dynamic time warping (DTW) algorithm. DTW can handle sequences of different lengths and speeds and is very suitable for the comparison of dance movements. For example, a spinning movement may have different durations in different performances, but DTW can find the best alignment method. The finally output dictionary structure facilitates subsequent analysis. For example, the occurrence frequency and duration of different actions can be counted to analyze the structure of the dance choreography; the performance of the dancer can also be evaluated, such as the smoothness and accuracy of the movements. Such a system can not only be used for dance teaching and scoring, but also be applied to dance creation assistance, motion capture, and virtual reality and other fields, providing technical support for the development of dance art.
[0035] S104. Obtain the joint angle and limb posture data of the dancer, and perform normalization processing on the joint angle and limb posture data together with the semantic displacement information. The normalization processing unifies data of different scales into the same numerical range. Combine the normalized data to construct a multi-modal feature matrix.
[0036] Obtain the joint angle data and limb posture data of the dancer through sensors, and preprocess these data to remove outliers and noise. Use a word embedding model such as Word2Vec in natural language processing technology to extract semantic information from the text description of the dance movement. Map each word to a real-valued vector of a fixed dimension, and then obtain a numerical vector representing the entire movement displacement by averaging or weighted averaging the vectors. Perform min-max normalization processing on the joint angle data, limb posture data, and semantic displacement vector respectively, and uniformly map their numerical ranges to the interval [0, 1]. Align the three normalized data according to the time stamp to ensure that they correspond one by one at each moment. Specifically, the data can be resampled to the same time interval, or an interpolation algorithm can be used for alignment. Concatenate the aligned joint angle data, limb posture data, and semantic displacement vector according to the feature dimension to construct a multi-modal feature matrix. Each row represents a time point, and each column represents a data feature. For example, if there are 10 features for joint angles, 15 features for limb postures, and 5 features for semantic displacements, then each row of the concatenated matrix has 30 elements. Use a convolutional neural network to extract and fuse features from the multi-modal feature matrix. The input of the network is the aligned multi-modal feature matrix. First, use a 1D convolutional layer to extract local features, then perform downsampling through a max pooling layer, and then alternate through multiple convolutional layers and pooling layers to gradually extract high-level semantic features. The size of the convolutional kernel can be adjusted according to the actual situation, such as 3x3 or 5x5. Use a fully connected layer at the top of the convolutional neural network to flatten the features and connect a Softmax classifier to map the features to predefined dance movement categories. The network can be trained using a labeled dance dataset, adopting a cross-entropy loss function and an Adam optimizer, and updating the network parameters through the backpropagation algorithm. In the test phase, input the dance data to be recognized into the trained network to obtain the prediction result of the action category. To evaluate the performance of the model, the hold-out method can be used to divide the dataset into a training set, a validation set, and a test set. Adjust the hyperparameters according to the performance on the validation set during the training process, and finally evaluate the generalization ability of the model on the test set. Common evaluation metrics include accuracy, precision, recall, and F1 value, etc. At the same time, the intermediate features of the model can also be analyzed through visualization tools to better understand the internal mechanism of the model.
[0037] Exemplarily, the core of the dance movement recognition system lies in the fusion and processing of multi-modal data. First, the joint angles and body posture data of the dancer are collected through sensors. For example, in ballet, inertial measurement unit (IMU) sensors can be used to capture the angle changes of the dancer's toes, knees, and hips. These data can accurately reflect the unique movements in ballet, such as Pointe and Pirouette. In the preprocessing stage, median filtering is used to remove outliers, such as instantaneous spikes caused by sensor jitter. At the same time, natural language processing techniques are used to extract the semantic information of dance movements. Taking modern dance as an example, the description "smooth wave arms" can be transformed into a numerical vector through the Word2Vec model. Assuming the word embedding dimension is 100, then "smooth", "wave", and "arms" will each be mapped to a 100-dimensional vector. By performing weighted averaging on these vectors, a vector that can represent the semantics of the entire movement can be obtained. The weights can be set according to the importance of the words in the dance field. For example, "wave" may have a higher weight than "smooth". Data normalization is a crucial step to ensure the comparability of data from different sources. For joint angle data, the possible range is 0° to 360°; body posture data may be represented as three-dimensional space coordinates; while the value ranges of the dimensions of the semantic vectors vary. Through min-max normalization, these heterogeneous data are uniformly mapped to the [0,1] interval. For example, an original angle value of 180° may be normalized to 0.5. Data alignment is the basis for multi-modal fusion. Suppose the sampling rate of joint angle data is 100Hz, the body posture data is 50Hz, and the semantic vectors are generated once per second. All data can be resampled to 50Hz, or aligned to 100Hz using linear interpolation. This ensures that at each time point, there are corresponding values for the three types of data. The multi-modal matrix formed by feature concatenation provides rich information for subsequent analysis. Taking a street dance movement as an example, suppose there are 10 joint angle features, 15 body posture features, and 5 semantic features. Then each row of the matrix will contain 30 elements, representing the comprehensive state of the dancer at that moment. The application of convolutional neural networks enables the system to automatically learn the spatio-temporal features of dance movements. Taking the "Freeze" movement in hip-hop as an example, the characteristic of this movement is sudden stillness. The shallow layer of the network may learn the short-term changes in limb positions, while the deep layer may capture the overall structure of the movement. 1D convolution is used because the data is essentially time series. For example, using a convolutional kernel of size 5, local features of approximately 0.1 seconds (assuming a sampling rate of 50Hz) can be captured. The Softmax classifier maps the extracted features to predefined dance movement categories. For example, for hip-hop, the possible categories include "Toprock", "Freeze", "PowerMove", etc. The network is trained through the cross-entropy loss function to enable it to distinguish these categories.The adaptive learning rate feature of the Adam optimizer helps to process features of different scales. The hold-out method is used for model evaluation, and the dataset can be divided into a training set, a validation set, and a test set in a ratio of 7:2:1. Hyperparameters such as the number of convolutional layers and the size of convolutional kernels are adjusted on the validation set. Finally, the model is evaluated on the test set to obtain metrics such as accuracy and precision. For example, if the precision of the model on the "Freeze" action reaches 95%, it indicates that it can well recognize this specific action. By visualizing intermediate features, such as using the t-SNE dimensionality reduction technique, the distribution of different dance actions in the feature space can be observed, which helps to understand the decision boundary of the model.
[0038] S105. Perform temporal modeling processing on the multi-modal feature matrix using a long short-term memory network to obtain a temporal modeling result that captures the continuity and smoothness of dance movements.
[0039] Based on the dance movement data collected by multi-modal sensors, construct a multi-modal feature matrix, including features such as joint positions, velocities, and accelerations. Use a long short-term memory network to perform temporal modeling processing on the multi-modal feature matrix, and control the information flow through the input gate, forget gate, and output gate of the network to capture the continuity and smoothness features of dance movements. At each time step, update the long-term dependence information in the memory unit according to the input features at the current moment and the hidden state at the previous moment, and generate the hidden feature representation at the current moment through the output gate. Repeat the above process until all time steps of the feature matrix are processed to obtain a complete temporal modeling result. Input the temporal modeling result into the fully connected layer, and perform dance movement classification through the Softmax activation function. Use the cross-entropy loss function to calculate the difference between the prediction result and the true label, and update the network parameters through the backpropagation algorithm to implement the training process of dance movement recognition. In the test stage, input the dance movement data to be recognized into the trained model, obtain the predicted probability distribution of the action category, and select the category with the highest probability as the final recognition result.
[0040] Exemplarily, the acquisition of dance motion data by multi-modal sensors is the basis for dance motion recognition. Taking ballet as an example, an inertial measurement unit (IMU) sensor can be used to collect data on the joint positions, velocities, and accelerations of dancers. For example, in an elegant Arabesque pose, the sensor can capture the stability of the supporting leg, the height and angle of the free leg, and the inclination of the upper body. These data together constitute a multi-modal feature matrix, providing rich information for subsequent temporal modeling. The application of long short-term memory networks (LSTMs) enables the system to effectively capture the continuity and smoothness of dance motions. In a sequence of smooth modern dance movements, the input gate of the LSTM can determine which new information needs to be recorded at the current moment, such as sudden changes in direction or speed. The forget gate is responsible for filtering information in long-term memory, for example, retaining the rhythm characteristics of the entire dance sequence while downplaying certain transient unstable movements. The output gate controls which information will be used to predict the motion state at the current moment. Updating the memory cell at each time step is the core operation of the LSTM. Suppose in a hip-hop performance, the dancer is performing a complex floor move. The input features at the current moment may include the dancer's body posture and movement speed, while the hidden state at the previous moment contains the context information of the previous movements. The LSTM will update its internal state based on this information to capture the coherence of the movements, such as the entire process from squatting to spinning and then standing up. Feeding the temporal modeling results into the fully connected layer is for the final classification decision. At this stage, the network needs to map the temporal features extracted by the LSTM to predefined dance motion categories. For example, in Latin dance, possible categories include Cha-cha, Rumba, Samba, etc. The application of the Softmax activation function ensures that the output is a probability distribution, with each category corresponding to a probability value. The use of the cross-entropy loss function helps to measure the difference between the prediction result and the true label. During training, if the model misclassifies a Cha-cha step as a Rumba, the cross-entropy loss will give a large penalty value. The backpropagation algorithm then uses this loss value to adjust the network parameters, gradually improving the recognition accuracy of the model. In practical applications, the performance evaluation of a dance motion recognition system usually adopts multiple metrics. In addition to the overall accuracy, the precision and recall of each dance motion also need to be considered. For example, for a difficult move like the Grand Jeté in ballet, special attention may be paid to its recognition accuracy. The confusion matrix can visually display the misclassification situations between different types of movements, enabling targeted improvement of the model. In addition, the generalization ability of the model is also an important consideration. An excellent dance motion recognition system should be able to adapt to the performance styles of different dancers. For example, for the figure-eight step in tango, different dancers may have subtle differences. By including diverse samples in the training data, the model's adaptability to these variations can be improved.
[0041] S106. According to the timing modeling results, extract key displacement patterns through cluster analysis and construct a dance movement representation model. The key displacement pattern refers to the representative displacement characteristics in dance movements.
[0042] Based on the obtained dance movement data, use time series analysis method to conduct modeling analysis on the movement displacement characteristics to obtain the movement displacement characteristic sequence. For the movement displacement characteristic sequence, perform feature clustering through the K-means clustering algorithm. If the Euclidean distance between features is less than the preset threshold, then classify them into the same category to obtain multiple representative key displacement patterns. According to the key displacement patterns, construct a dance movement representation model. Use the Hidden Markov Model (HMM) to model the key displacement patterns, take the key displacement patterns as the observation states of the HMM, define the hidden state as the basic unit of the dance movement, and estimate parameters such as the transition probability and emission probability of the HMM through the Baum-Welch algorithm to obtain the dance movement HMM model. For the newly input dance movement data, by extracting its movement displacement characteristics, perform forward algorithm calculation using the constructed dance movement HMM model to obtain the probability of the movement sequence under each HMM model. If the probability is greater than the preset threshold, then classify the movement sequence into the corresponding key displacement pattern. According to the classification result, obtain the key displacement pattern that best matches the input dance movement, and obtain the displacement feature representation vector corresponding to the key displacement pattern by looking up the dance movement representation model. Use the Support Vector Machine (SVM) classifier to classify the displacement feature representation vector of the dance movement, and obtain the SVM classification model through the training sample data. Input the displacement feature representation vector of the newly input dance movement into the SVM classification model, and calculate through the classification decision function to obtain the category label to which the movement belongs, realizing the automatic classification and recognition of dance movements. Finally, complete the representation and recognition of dance movements based on key displacement patterns.
[0043] Exemplarily, the time series analysis of dance movement data is a key step in understanding dance movement characteristics. Taking ballet as an example, an elegant arabesque movement involves coordinated changes in the supporting leg, free leg, and upper body. By analyzing the displacement changes of these parts in the time dimension, the coherence and smoothness of the movement can be captured. The application of the K-means clustering algorithm enables the extraction of representative key displacement patterns from complex movement sequences. In Latin dance, the basic steps of the cha-cha may be identified as a key displacement pattern. By setting an appropriate Euclidean distance threshold, similar step movements can be classified, thus simplifying the subsequent analysis process. The introduction of the Hidden Markov Model (HMM) provides a powerful tool for the probabilistic modeling of dance movements. In hip-hop dance, a complex floor movement sequence can be regarded as a combination of multiple basic movement units. The hidden states of the HMM can represent these basic units, while the observed states correspond to the key displacement patterns obtained through clustering. The application of the Baum-Welch algorithm enables the learning of the transition probabilities and emission probabilities between these states from the training data. The use of the forward algorithm enables the evaluation of the matching degree between a new dance movement sequence and the established HMM model. For example, when identifying a modern dance performance, the probability of this movement sequence under different HMM models can be calculated. If the probability given by a certain model is significantly higher than the preset threshold, this movement sequence can be classified as the corresponding dance style or specific movement combination. The introduction of Support Vector Machine (SVM) provides strong support for the final classification of dance movements. By converting the key displacement patterns into feature vectors, an SVM classifier can be trained to distinguish different types of dance movements. For example, when distinguishing between jumping movements in ballet, the SVM can learn the subtle differences between different jumping types (such as grand jeté, petit jeté), thus achieving accurate classification. This method of dance movement representation and recognition based on key displacement patterns has multiple advantages. First, it can effectively handle the temporal characteristics of dance movements, capturing the continuity and changing trends of movements. Second, through clustering and HMM modeling, the most representative movement patterns can be extracted from massive data, greatly improving the efficiency and generalization ability of the algorithm. Finally, the application of the SVM classifier enables the system to handle complex non-linear classification problems and adapt to the diversity of different dance genres and performance styles. In practical applications, this method can be used in dance teaching assistance systems, dance movement scoring systems, or intelligent choreography systems. For example, in a dance teaching scenario, the system can analyze the movements of students in real time, identify the deviations from the standard movements, and give targeted improvement suggestions. This can not only improve teaching efficiency but also provide students with a personalized learning experience. In addition, in dance competition scoring, this method can provide objective technical references for judges, reduce the influence of subjective factors, and improve the fairness and accuracy of scoring.
[0044] S107. Use a support vector machine to perform classification training on the dance movement representation model to obtain a classification training result for identifying displacement features of different dance movements.
[0045] According to the pre-established dance movement representation model, obtain the dance movement sequence data to be recognized. For the obtained dance movement sequence data, use human pose estimation tools such as OpenPose to extract the key point coordinates therein, and calculate the displacement of the key points between adjacent frames to obtain a displacement feature sequence. Divide the displacement feature sequence into segments of a fixed length, with each segment corresponding to a dance movement. Normalize the divided displacement feature segments and input them into a pre-trained support vector machine classifier for classification judgment. The support vector machine classifier is trained using pre-annotated dance movement segments, with each category corresponding to a dance movement. If the classification result indicates that the current displacement feature segment belongs to a certain category of dance movement, then output that category as the recognition result. According to the recognition result, obtain the label text corresponding to the current dance movement category from the pre-constructed action label dictionary. Align the recognized action label with the original dance movement sequence data on the time axis to obtain a dance movement sequence with action annotations. Use the dance movement sequence with annotations as training data to fine-tune the dance movement representation model through deep learning frameworks such as TensorFlow to improve the action recognition accuracy of the model. During the fine-tuning training process, strategies such as data augmentation and L2 regularization can be used to control overfitting.
[0046] Exemplarily, the core of the dance movement recognition system lies in accurately capturing and analyzing human motion characteristics. Taking ballet as an example, an elegant Arabesque movement involves the coordination of multiple key points. Tools such as OpenPose can precisely locate the key points of the dancer's shoulders, elbows, wrists, hips, knees, ankles, etc. By calculating the displacements of these points between adjacent frames, a displacement feature sequence reflecting the smoothness and strength of the movement can be obtained. Dividing the continuous dance movement sequence into fixed-length segments is the key to identifying individual movements. For example, a complete ballet spin may last about 2 seconds and contain 60 frames of images. By setting an appropriate window size, this movement can be accurately segmented. Normalizing each segment to a unified scale, such as the interval [-1, 1], can eliminate the influence of differences in the body shapes of different dancers and improve the generalization ability of recognition. The support vector machine (SVM) classifier performs well in dance movement recognition. By using a large amount of pre-annotated training data, such as segments containing various basic ballet movements (such as jumps, spins, backbends, etc.), the SVM can learn the decision boundaries between different movement categories. In the recognition stage, when a new displacement feature segment is input into the SVM, it can quickly determine which dance movement category this segment most likely belongs to. To make the recognition results more intuitive, an action label dictionary can be constructed. For example, map the digital category "1" to "Arabesque", "2" to "Grand Jeté", etc. In this way, the system output is no longer an abstract number, but a dance term with practical meaning, which is convenient for dancers and coaches to understand and use. Timeline alignment is a crucial step to ensure the practicality of the recognition results. By corresponding the recognized action labels with the original video sequence in the time dimension, the start and end times of each action can be accurately located. This is particularly important for dance teaching and scoring systems because it can help analyze the duration, rhythm, and coherence of movements. Fine-tuning the model using annotated dance movement sequences can significantly improve the recognition accuracy. For example, using TensorFlow to build a deep neural network containing convolutional layers and recurrent layers can better capture the spatio-temporal characteristics of dance movements. During the training process, data augmentation techniques (such as adding noise, rotation transformation, etc.) can be used to improve the robustness of the model. At the same time, techniques such as L2 regularization can effectively prevent overfitting and ensure that the model can also maintain good performance on new dance videos. The application prospects of this dance movement recognition system are broad. In professional dance teaching, it can provide instant feedback to students, pointing out the subtle deviations in movements. In dance competition scoring, it can provide objective references for judges and improve the fairness of scoring. In addition, in intelligent choreography systems, this technology can help creators analyze and combine different dance elements and inspire creative inspiration. Through continuous optimization and iteration, this system is expected to become an important tool for promoting the development of dance art.
[0047] S108. Optimize the parameters of the dance motion representation model according to the classification training results. If the classification accuracy of the dance motion representation model on the test set is lower than the preset threshold, readjust the fusion weights of the multi-modal feature matrix. The fusion weights refer to the importance of different modal features in the matrix. By adjusting the fusion weights, iteratively optimize the performance of the dance motion representation model.
[0048] Obtain the classification training results of the dance motion representation model on the training set. According to the classification training results, use hyperparameter optimization methods such as grid search to optimize and adjust the model parameters of the dance motion representation model, such as learning rate, batch size, regularization coefficient, etc. Use the optimized parameters to retrain the dance motion representation model on the training set to obtain an optimized model. Evaluate the optimized dance motion representation model on the test set to obtain the classification accuracy of the model on the test set. Compare the classification accuracy of the model on the test set with the preset performance threshold. If the classification accuracy is lower than the preset threshold, trigger the fusion weight adjustment mechanism. The fusion weight adjustment mechanism optimizes the weight coefficients of the modal features through the gradient descent algorithm based on the contribution degree of different modal features to the classification results, so as to improve the classification accuracy. Specifically, first extract the features of different modalities based on the current model, such as bone key point coordinates, RGB images, optical flow, etc., and construct a multi-modal feature matrix. Initialize the weight coefficients of each modal feature and generate fused features in a weighted fusion manner. Use the fused features as input and the bone key point action category as the label to construct a classification loss function. The loss function performs backpropagation on the fused features, calculates the gradients of the weight coefficients of each modal feature, and uses the gradient descent algorithm to update the weight coefficients to minimize the classification loss. Continuously iterate and optimize until the weight coefficients converge or reach the preset number of iterations. Apply the optimized modal fusion weights to the multi-modal feature fusion to generate an optimized fused feature matrix. Use this fused feature matrix as input to retrain the dance motion representation model to obtain an optimized model. Repeat the above test evaluation and fusion weight adjustment process until the classification accuracy of the model on the test set reaches the preset threshold. Finally, obtain a dance motion representation model with better generalization performance.
[0049] Exemplarily, the optimization of the dance movement representation model is a cyclic iterative process aimed at improving the model's classification accuracy and generalization ability. First, through hyperparameter optimization methods such as grid search, the impact of different parameter combinations on the model performance can be systematically explored. For example, in ballet movement recognition, it may be found that when the learning rate is set to 0.001, the batch size is 64, and the L2 regularization coefficient is 0.0001, the model performs best on the training set. The optimized model needs to be evaluated on the test set to verify its generalization ability. Suppose the set performance threshold is a classification accuracy of 95%, but the optimized model only reaches 92% on the test set. At this time, the fusion weight adjustment mechanism needs to be triggered to further improve the model performance. The core idea of the fusion weight adjustment mechanism is to make full use of multimodal information. In dance movement recognition, in addition to the coordinates of skeletal key points, RGB images and optical flow information can also be considered. Each modality can provide unique movement features: skeletal key points reflect the spatial position relationship of the limbs, RGB images contain rich visual details, and optical flow captures the continuity and speed changes of the movement. Initially, equal weights may be assigned to these three modalities. However, through the gradient descent algorithm, these weights can be gradually adjusted to make the model pay more attention to the features that contribute more to the classification result. For example, when recognizing the "grand jeté" movement in ballet, it may be found that the skeletal key points and optical flow information are more important than RGB images, so the algorithm will correspondingly increase the weights of these two modalities. The weight adjustment process is iterative. The classification loss of the current fused features is calculated in each iteration, and then the weights are updated through backpropagation. This process can not only improve the accuracy of the model but also help understand the importance of different modalities in dance movement recognition. For example, it may be found that when recognizing subtle hand movements, the weight of RGB images will increase significantly because it can capture finger postures that are difficult to accurately represent by skeletal key points. After multiple rounds of iterative optimization, suppose the accuracy of the model on the test set has increased to 96%, exceeding the preset threshold of 95%. At this time, a dance movement representation model with better performance and stronger generalization ability is obtained. This optimized model can not only more accurately recognize known dance movements but also better handle new and unseen movement variants. The reason why this method of multimodal fusion and dynamic weight adjustment is effective is that it can adaptively utilize different types of information. In practical applications, this means that the system can better handle various complex dance scenarios. For example, in an environment with insufficient light, the model may rely more on skeletal key point information; while in a fast continuous movement sequence, optical flow information may be given a higher weight. This flexibility enables the model to maintain high accuracy in various different dance styles and performance environments, thus providing more reliable technical support for dance teaching, scoring, and creation.
[0050] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made thereto based on the present invention, which will be obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection claimed by the present invention.
Claims
1. A dance action image analysis method for image time sequence recognition, characterized in that: include: Obtain a dance action image sequence and determine the dancer's pixel position information through image segmentation; Calculating the dancer's initial displacement trajectory according to the change trend of the dancer's pixel position information; Extracting high-dimensional features of the dancer in the image sequence, and combining the initial displacement trajectory to extract displacement information related to the semantics of dance movements; Acquiring the dancer's joint angle and limb posture data, normalizing the joint angle and limb posture data with the semantic displacement information, and constructing a multimodal feature matrix; Using a long short-term memory network to perform time series modeling processing on the multimodal feature matrix to obtain a time series modeling result; According to the time series modeling results, key displacement patterns are extracted through cluster analysis to construct a dance movement representation model; Using a support vector machine to perform classification training on the dance movement representation model to obtain a classification training result; According to the classification training results, the parameters of the dance movement representation model are optimized.
2. The method according to claim 1, characterized in that The step of acquiring a dance action image sequence and determining the pixel position information of the dancer by image segmentation includes: Obtain a continuous image sequence containing dance movements, and use the MaskR-CNN algorithm based on deep learning to perform pixel-level segmentation on the dancer area in the image; A binary mask image of the dancer area in each frame of the image is obtained, wherein the dancer's pixel position is marked with a specific value, and the background area pixel is marked with another specific value; Performing a bitwise AND operation on the binary mask image and the original image frame to extract pixel information of the dancer area; Performing a morphological closing operation on pixel information of the dancer region to eliminate holes and breakpoint noise in the image; The pixel coordinates of the dancer's outline are extracted through the contour extraction algorithm, and the pixel coordinate range of the dancer's area in each frame of the image is determined.
3. The method according to claim 1, characterized in that The step of calculating the dancer's initial displacement trajectory according to the change trend of the dancer's pixel position information includes: For each frame of the image, the human pose estimation model is used to detect and locate the coordinates of the dancer’s key body parts; The coordinates of body parts detected in two adjacent frames are matched by the Hungarian algorithm to establish the correspondence between different frames of the same dancer. According to the corresponding relationship, the displacement vector of each dancer between two adjacent frames is calculated, and the displacement vectors are connected in time sequence to form an initial displacement trajectory sequence; Applying a Kalman filter algorithm to smooth the initial displacement trajectory sequence to estimate the dancer's true displacement trajectory; According to the real displacement trajectory, the dancer's instantaneous velocity vector and acceleration vector at each moment are obtained through numerical differentiation calculation.
4. The method according to claim 1, characterized in that The step of extracting high-dimensional features of the dancer in the image sequence and combining the initial displacement trajectory to extract displacement information related to the semantics of the dance movement includes: For each frame of the video sequence, extract its color histogram, local binary pattern texture features and edge feature high-dimensional feature information; Normalizing the high-dimensional feature information to scale the feature values to a specific range; The normalized high-dimensional feature information is input into a pre-trained convolutional neural network, and the feature dimension is reduced through multiple iterative calculations of the convolution layer and the pooling layer; In the first frame of the video sequence, the object detection algorithm is used to obtain the dancer’s initial position coordinates; For each frame in the video sequence, the low-dimensional features output by the convolutional neural network and the coordinates of the reference point are concatenated into a vector; The Euclidean distance between the vector and the reference point coordinates is calculated to obtain a displacement vector of the dancer's position in the current frame relative to the initial position.
5. The method according to claim 1, characterized in that The step of obtaining the dancer's joint angle and limb posture data, normalizing the joint angle and limb posture data and the semantic displacement information, and constructing a multimodal feature matrix includes: The dancer’s joint angle data and limb posture data are acquired through sensors, and pre-processed to remove outliers and noise; Using the word embedding model, semantic information is extracted from the text description of the dance movements, and each word is mapped into a real number vector of fixed dimension; The joint angle data, limb posture data and semantic displacement vector are respectively normalized to the maximum and minimum values, and the numerical range is uniformly mapped to a specific interval; Align the three normalized data according to timestamps; The aligned joint angle data, limb posture data and semantic displacement vectors are concatenated according to the feature dimension to construct a multimodal feature matrix.
6. The method according to claim 1, characterized in that The method of using a long short-term memory network to perform time series modeling processing on the multimodal feature matrix to obtain a time series modeling result includes: According to the multimodal feature matrix, a long short-term memory network is used to perform time series modeling processing, and the flow of information is controlled through an input gate, a forget gate, and an output gate of the network to capture the continuity and smoothness characteristics of the dance movements; At each time step, based on the input features at the current moment and the hidden state at the previous moment, the long-term dependency information in the memory unit is updated, and the hidden feature representation of the current moment is generated through the output gate; Repeat the above process until the feature matrices of all time steps are processed and the complete time series modeling results are obtained; The time series modeling results are input into the fully connected layer, and the dance movements are classified through the activation function.
7. The method according to claim 1, characterized in that According to the time series modeling results, the key displacement patterns are extracted through cluster analysis to construct a dance movement representation model, including: According to the acquired dance movement data, the time series analysis method is used to model and analyze the movement displacement characteristics, and the movement displacement feature sequence is obtained; For the action displacement feature sequence, feature clustering is performed using a clustering algorithm. If the distance between features is less than a preset threshold, they are classified into the same category to obtain multiple key displacement patterns. constructing a dance movement representation model according to the key displacement patterns; A hidden Markov model is used to model the key displacement pattern, and the key displacement pattern is used as the observation state of the hidden Markov model. The hidden state is defined as the basic unit of dance movements. The transition probability and emission probability parameters of the hidden Markov model are estimated through a specific algorithm to obtain a hidden Markov model of dance movements.
8. The method according to claim 1, characterized in that Optimizing the parameters of the dance movement representation model according to the classification training results includes: Obtain the classification training results of the dance movement representation model on the training set; According to the classification training results, using a hyperparameter optimization method, optimizing and adjusting the model parameters of the dance movement representation model; Retraining the dance movement representation model on the training set using the optimized parameters to obtain an optimized model; Evaluating the optimized dance movement representation model on a test set to obtain the classification accuracy of the model on the test set; If the classification accuracy is lower than a preset threshold, a fusion weight adjustment mechanism is triggered; The fusion weight adjustment mechanism optimizes the weight coefficient of the modal features through a gradient descent algorithm based on the contribution of different modal features to the classification results; The optimized modality fusion weights are applied to multimodal feature fusion to generate an optimized fusion feature matrix.
9. A dance movement image analysis system for image time sequence recognition, based on the dance movement image analysis method for image time sequence recognition according to any one of claims 1 to 8, characterized in that: Image acquisition and segmentation module, displacement trajectory calculation module, high-dimensional feature extraction module, posture data acquisition and normalization module, processing module and evaluation and adjustment module; An image acquisition and segmentation module is used to acquire a dance action image sequence and determine the pixel position information of the dancer through image segmentation; A displacement trajectory calculation module, used to calculate the dancer's initial displacement trajectory according to the change trend of the dancer's pixel position information; A high-dimensional feature extraction module, used to extract high-dimensional features of the dancer in the image sequence, and extract displacement information related to the semantics of dance movements in combination with the initial displacement trajectory; A posture data acquisition and normalization module is used to obtain the dancer's joint angle and limb posture data, normalize the joint angle and limb posture data with the semantic displacement information, and construct a multimodal feature matrix; The processing module is used to propose and build models for the collected information, and to classify and train the data and optimize the data parameters; The evaluation and adjustment module is used to evaluate the model performance on the test set and trigger the fusion weight adjustment mechanism to optimize the weight coefficients of different modal features.
10. A computer device comprising: A memory and a processor; the memory stores a computer program, wherein the processor implements the steps of the dance movement image analysis method for image timing recognition according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Cited By
Dance motion auxiliary generation method and system based on three-dimensional modeling
CN120765860A
Dance motion auxiliary generation method and system based on three-dimensional modeling
CN120765860B
Coprocessing-based mobile terminal 3D human body posture estimation method and system
CN121640581A
A mobile terminal 3D human posture estimation method and system based on cooperative processing
CN121640581B
Degradation robustness-oriented quality perception adaptive fusion continuous sign language recognition method
CN122157314A