Boxing action recognition method and system based on priori knowledge and multivariate time sequence classification

Through data preprocessing, filtering and feature derivation combined with multivariate timing classifiers, the problems of high computational cost and low accuracy in boxing action recognition are solved, and efficient and real-time boxing action recognition is achieved.

CN120296603APending Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510445526.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems in boxing movement recognition with high calculation costs, poor feedback real-time performance, and difficulty in quantifying athlete performance, and identification based on bone point movement timing data is affected by occlusion and complexity, and has low accuracy.

Method used

Using a method based on prior knowledge and multivariate timing classification, smooth and accurate identification of bone point motion timing data containing noise and outliers is achieved through data preprocessing, traceless Kalman filtering, bone length constraint filtering, boxing scene-specific feature derivation, shapelet mining and dual-channel Transformer classifiers.

Benefits of technology

It improves the accuracy and real-timeness of boxing action recognition, reduces calculation costs, and can effectively identify complex boxing actions and provide real-time feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296603A_ABST
    Figure CN120296603A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data mining and artificial intelligence, and discloses a boxing action recognition method and system based on priori knowledge and multivariate time sequence classification, a non-contact action capture system is adopted to obtain skeleton point motion time sequence data, and a skeleton structure is stabilized through coarse-grained filtering based on unscented Kalman filtering, so that the recognition accuracy of the boxing action is improved. Fine-grained filtering based on kinematics priori knowledge inhibits noise and abnormity on a time sequence coordinate value level, comprehensive analysis is carried out on skeleton point motion time sequence data, and attack and defense actions in a boxing scene are effectively identified and classified by utilizing the kinematics priori knowledge and a multivariable time sequence classifier. According to the method, the problem that the action recognition accuracy is reduced due to low data quality caused by shielding is effectively solved, and a real-time feedback result is provided for a boxing scene based on the multivariable time sequence classifier of the dual-channel Transform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of data mining and artificial intelligence, and specifically relates to a boxing action recognition method and system based on prior knowledge and multivariate time series classification. Background Art

[0002] In recent years, with the development of deep learning, human action recognition, including action recognition in sports, has become a popular research field. The action recognition obtained through advanced computer vision algorithms has significant effects and practical application value. However, there are still some problems in its application in the sports field. The vision-based model has a high computational cost and poor real-time feedback, making it difficult to meet the real-time requirements of real sports scenarios and unable to quantify the performance of athletes. For example, the model can recognize a boxer's straight punch, but does not know the punching speed and acceleration.

[0003] Many sports practitioners hope to extract possible points for improving techniques and tactics from the actions of athletes themselves. Therefore, in the sports field, the temporal data of the skeletal point movements of athletes is usually extracted through channels such as motion capture, and the corresponding actions are recognized through temporal information. However, the temporal data generated by motion capture often has outliers and jitters, and this phenomenon is exacerbated due to factors such as the occlusion of the boxing ring. At the same time, the complexity of boxing actions also makes it more difficult to recognize boxing actions based on the temporal data of skeletal point movements. These challenges have made boxing action recognition a thorny and urgent problem to be solved. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a boxing action recognition method and system based on prior knowledge and multivariate time series classification, which can accurately recognize actions for the noisy and outlier-containing temporal data of skeletal point movements in a boxing scenario.

[0005] The boxing action recognition method based on prior knowledge and multivariate time series classification according to the present invention includes the following steps:

[0006] S1. Based on the collected temporal data of skeletal point movements, perform data preprocessing on the temporal data to obtain preprocessed data;

[0007] S2. Use a coarse-grained filter based on the unscented Kalman filter to smooth the temporal data of skeletal point movements;

[0008] S3. Apply kinematic prior knowledge to design a fine-grained filter based on bone length constraints to correct the positional relationship of skeletal points;

[0009] S4. Derive boxing scenario-specific features based on the prior knowledge of the boxing scenario;

[0010] S5. Mine and screen shapelets of the skeletal point motion time series data, and obtain the local matching scores between the shapelets and the sample data;

[0011] S6. Design a multivariate time series classifier based on a dual-channel Transformer, and train it using cross-entropy loss to obtain the recognition result.

[0012] Furthermore, in S1, the skeletal point motion time series data is collected by a contactless motion capture device, and includes the coordinate value sequences of multiple skeletal points changing over time; preprocess the collected data, specifically:

[0013] S11. Set the length value of the action segment time series, and align the sequences to the same length to ensure the consistency of the time series length;

[0014] S12. Use improved linear interpolation to fill in the missing values, and fill the "0" coordinate values with the real coordinate values;

[0015] S13. Use the min-max normalization method to scale the three-dimensional coordinate values of each skeletal node to the interval [0, 1].

[0016] Furthermore, S2 is specifically as follows:

[0017] S21. Set the initial value of the state vector, the state covariance matrix, the process noise covariance matrix, and the observation noise covariance matrix;

[0018] S22. Based on the state vector, calculate a set of sigma points at each time step;

[0019] S23. Based on the sigma points and the state transition function, iteratively update the state vector and the state covariance matrix, and output a new set of sigma point sets;

[0020] S24. Calculate the Kalman gain according to the cross-covariance between the state vector and the observation value and the observation covariance matrix, and construct a new state prediction and state covariance matrix.

[0021] Furthermore, S3 is specifically as follows:

[0022] S31. According to the human skeleton structure, select appropriate skeletal point pairs, and calculate the skeletal point pair length sequences according to the skeletal point motion time series data information;

[0023] S32. Sort the skeletal point pair length sequences and delimit the intervals, and take the average to obtain the approximate real bone length;

[0024] S33. By calculating the bone midpoint and the bone direction vector, correct the skeleton structure depicted by the skeletal point motion time series data.

[0025] Further, S4 is specifically as follows:

[0026] S41. Through principal component analysis, screen the key joints with distinguishing features in offensive and defensive actions, and derive their speed and acceleration features;

[0027] S42. Incorporate the trajectory information of key joint points, add the trajectory curvature information of the joints, and capture the difficult-to-classify features such as swinging punches and hook punches;

[0028] S43. Derive limb coordination features including the shoulder-wrist-hip triangle angle to describe the coordination relationship between limbs in boxing actions;

[0029] S44. Define the key stages of the action, and obtain the data segments corresponding to the key stages of the action through manual annotation or peak detection, which are the multivariate time series data.

[0030] Further, S5 is specifically as follows:

[0031] S51. Obtain shapelets by sliding window cropping;

[0032] S52. Use FastDTW to measure the distance between shapelets, and screen high-discrimination shapelets through information gain, and then obtain the optimal shapelet set;

[0033] S53. Measure the matching degree between the bone point motion time series data samples and shapelets through gait convolution to obtain local matching scores.

[0034] Further, in S6, the construction process of the multivariate time series classifier based on the dual-channel Transformer is specifically as follows:

[0035] S61. Design a dual-channel Transformer mechanism, one is a class-specific channel, and the other is a global feature capture channel;

[0036] S62. The class-specific channel obtains a query matrix based on the local matching score, maps the key stage sequence of each action sample to obtain the key and value matrices, and further obtains class-specific features through multi-head self-attention;

[0037] S63. The global feature capture channel obtains long-range dependence based on the multivariate time series output by S4;

[0038] S64. Design a gating mechanism to dynamically weight and fuse the output features of the two channels to achieve the complementarity of different types of features;

[0039] S65. Through minimizing the cross-entropy loss, the loss is propagated to the entire network for end-to-end training;

[0040] S66. The classifier outputs the predicted distribution of each possible action in the time series, and the action label is obtained to complete the boxing action recognition task.

[0041] The present invention also provides a boxing action recognition system based on prior knowledge and multivariate time series classification, including:

[0042] Data preprocessing module: Based on the collected kinematic sequence data of skeleton points, preprocess the sequence data to obtain the preprocessed data;

[0043] Coarse-grained filtering module: Apply coarse-grained filtering based on the unscented Kalman filter to the preprocessed data to smooth the kinematic sequence data of skeleton points;

[0044] Fine-grained filtering module: Utilize the prior knowledge of kinematics to design a fine-grained filter based on bone length constraints to suppress noise and anomalies at the coordinate value level;

[0045] Boxing-specific feature derivation module: Derive boxing-scene-specific features based on the prior knowledge of the boxing scene;

[0046] Shapelet feature mining module: Mine and screen the shapelets of the kinematic sequence data of skeleton points to obtain the local matching scores between the shapelets and the sample data;

[0047] Multivariate time series classifier: Design a multivariate time series classifier based on a two-channel Transformer and train it using cross-entropy loss to obtain the recognition result.

[0048] The beneficial effects of the present invention are as follows: The present invention utilizes the prior knowledge of kinematics, introduces the speed change under coarse-grained filtering, and uses the unscented Kalman filter to iteratively update the stable skeleton structure of the skeleton state change; designs a fine-grained filter based on bone length constraints to suppress noise and anomalies at the coordinate value level, effectively solving the problem of the decrease in action recognition accuracy caused by the low data quality due to occlusion. Utilize the prior knowledge of the boxing scene to derive action features such as punching; mine the shapelets in the key stages of action execution as highly discriminative subsequences for action recognition. In addition, the present invention designs a two-channel Transformer, in which the class-specific channel relies on the local matching scores obtained from the highly discriminative shapelet features to mine the local features of the action sequence, and the global feature capture channel captures the global features of the key stages of the action sequence, and fuses these two features through a gating mechanism to improve the action recognition accuracy. Description of the Drawings

[0049] Figure 1 It is a flowchart of the method described in the present invention;

[0050] Figure 2Schematic diagram of a multivariate time series classifier provided by an embodiment of the present invention;

[0051] Figure 3 Architecture diagram of a boxing action recognition system of the present invention;

[0052] Figure 4 Structure diagram of shapelet feature mining provided by an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of a dual-channel Transformer structure provided by an embodiment of the present invention;

[0054] Figure 6 Schematic diagram of collecting boxing actions;

[0055] Figure 7 Schematic diagram of boxing action recognition performance and loss;

[0056] Figure 8 Schematic diagram of bone point distribution;

[0057] Figure 9 Schematic diagram of the confusion matrix of the boxing action dataset;

[0058] Figure 10 Schematic diagram showing the anomalies of the original skeleton structure, abnormal coordinate values, and the data repair capabilities of coarse-grained filtering and fine-grained filtering;

[0059] Figure 11 Visualization diagram of noisy time series data and filtered data of partial bone point coordinates;

[0060] Figure 12 Visualization diagram of the length constraint of bone point pairs under fine-grained filtering;

[0061] Figure 13 Schematic diagram of the accuracy and loss curves of only class-specific channels;

[0062] Figure 14 Schematic diagram of the accuracy and loss curves of only the global feature capture channel. Detailed implementation manners

[0063] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below according to specific embodiments and in conjunction with the accompanying drawings.

[0064] The present invention provides a new idea for complex boxing action recognition. This method and system can output the corresponding action category labels only by inputting noisy original data, and propose a solution that can recognize complex boxing actions, provide real-time feedback, and has a low computational cost.

[0065] Such as Figure 1As shown in the figure, a boxing action recognition method based on prior knowledge and multivariate time series classification according to the present invention includes the following steps:

[0066] S1. Based on the collected kinematic time series data of skeletal points, perform data preprocessing on the time series data to obtain the preprocessed data;

[0067] S2. Perform coarse-grained filtering based on the unscented Kalman filter on the preprocessed data to smooth the kinematic time series data of skeletal points;

[0068] S3. Utilize kinematic prior knowledge to design fine-grained filtering based on bone length constraints to suppress noise and anomalies at the coordinate value level;

[0069] S4. Derive boxing scene-specific features based on prior knowledge of the boxing scene;

[0070] S5. Mine and screen the shapelets of the kinematic time series data of skeletal points to obtain the local matching scores between the shapelets and the sample data;

[0071] S6. Design a multivariate time series classifier based on a dual-channel Transformer and train it using cross-entropy loss to obtain the recognition result.

[0072] Among them, the kinematic time series data of skeletal points described in S1 specifically refers to the three-dimensional spatial coordinates of key skeletal nodes during the boxing process, which change dynamically over time. The key skeletal nodes include but are not limited to the head, shoulders, elbows, hands, hips, knees, and ankles, and their x, y, z coordinate values constitute the complete time series data for detailed characterization of the details of boxing actions. The collection of the kinematic time series data of skeletal points can adopt a non-contact motion capture system, and the collection frequency is maintained above 30Hz.

[0073] Further, in order to ensure the consistency of the sequence length of the action period, set the sequence length value of the action segment. For sequences with a length less than the set value, fill 0 before and after to extend to the set value length; for sequences with a length greater than the set value, mark the middle frame of the action and select the consecutive frames before and after the middle frame to ensure that the length of the final sequence meets the set value.

[0074] Further, due to the clinching, leaning actions of boxers, the boxing ring fence, and even the bodies of opponents blocking the key skeletal points of the athletes, resulting in missing values in the collected kinematic time series data of skeletal points, an improved linear filling is designed.

[0075] Specifically, for the missing values in time series data, the present invention adopts an improved linear interpolation strategy. First, if there are missing values at the beginning of the sequence, traverse the subsequent data and fill them with the first valid value. Second, for the missing values in the middle of the sequence, use linear interpolation method to fill them in combination with the front and back valid values. Finally, if there are missing values at the end of the sequence, perform predictive filling using linear regression according to the change trend of the valid data before the end, so as to ensure the continuity of the action.

[0076] Further, in the present invention, the coordinate value of a certain dimension of a certain bone point is defined as a variable. For example, the x coordinate of the hand joint is a variable. Therefore, the motion time series data of bone points is a multi-variable time series from the perspective of time series classification. Each variable is scaled by applying min-max normalization:

[0077]

[0078] where, X origin is the noisy original time series coordinate value, X min is the minimum value under the original time series coordinate value, X max is the maximum value, X norm is the normalized noisy original time series coordinate value, and the coordinate value range is [0, 1].

[0079] In S2, considering the instability of the motion time series data of bone points, accompanied by jitter and outliers, which has a negative impact on the accuracy of action recognition, the motion time series data of bone points is smoothed by coarse-grained filtering. The coarse-grained filtering uses the unscented Kalman filter adapted to non-linear motion data to smooth the x, y, and z coordinate time series data of each bone point.

[0080] Further, the method defines the spatial coordinates of the bone point in the earth coordinate system as the state vector to describe the motion state at the current moment. In order to accurately describe the motion state, the coarse-grained filtering introduces the velocity change as part of the state vector. The present invention iteratively calculates the state vector according to the state transition function and the observation function, and processes the non-linearity in the state transition function and the observation function through unscented transformation.

[0081] Specifically, the initial value of the state vector is determined as the information contained in the first timestamp of the motion time series data of the bone point. In order to record the change of velocity, a zero vector of the same dimension is concatenated. The coarse-grained filtering recursively calculates the state vector s and the observation vector o in time. The state transition function F is used to iterate the state vector of the previous timestamp to the current timestamp, and introduces the process noise p with a mean of 0 and a process noise covariance matrix of Q noi, to simulate the uncertainty of the state transition process. The observation function H maps the state vector at the current timestamp to the observation space and introduces an observation noise O with a mean of 0 and an observation noise covariance matrix of R noi , to simulate the uncertainty of the observation process;

[0082] s i = F(s i-1 ) + p noi

[0083] o i = H(s i ) + o noi

[0084] where i represents the i-th timestamp, F(·) is the state transition function, and F(s i-1 ) is the state vector s i-1 at the previous timestamp being transferred to the state at the current timestamp; H(·) is the observation function, and H(s i ) maps the state vector s i at the current timestamp to the observation space, and o i represents the observation value in the observation space at the i-th timestamp.

[0085] During iterative update, the unscented Kalman filter calculates a set of sigma points at each time step. These points are used to approximate the non-linear probability distribution, with a mean of and a state covariance matrix of P. These points have weights W for mean reconstruction and covariance reconstruction. The following formula can generate a set of 2m + 1 sigma points, where m is the dimension of the skeletal point motion time series data:

[0086]

[0087] where col represents the col-th column of the matrix; λ is the scaling factor, λ = α 2 (m + κ) - m; α determines the distribution range of the sigma points and is usually taken as a small value, and κ is an auxiliary parameter.

[0088] For mean reconstruction and covariance reconstruction, the sigma points have different weights. The mean reconstruction weight W (mean) and the covariance reconstruction weight W (cov) are:

[0089]

[0090] where is the mean reconstruction weight for the current state estimate; is the covariance reconstruction weight for the current state estimate, and k represents the index value of the sigma point.

[0091] These sigma points are input into the non-linear state transition function, i.e., and the state vector and state covariance matrix are iteratively updated, where χ k is the sigma point of the k-th point, represents the k-th sigma point after state transition.

[0092]

[0093] Among them, for the i-th timestamp, represents the predicted value of the updated state vector; is the predicted sigma point obtained by the sigma point through the state transition function; is the predicted value of the updated state covariance matrix; Q i is the process noise matrix.

[0094] Furthermore, the sigma points output by the state transition function are input into the observation function H to reconstruct the observed values, i.e., and the observed value covariance matrix and the cross-covariance between the state vector and the observed values are calculated:

[0095]

[0096] Among them, represents the estimation of the observed value; represents the observed value of the k-th sigma point at the i-th timestamp; is the observed covariance matrix; is the cross-covariance between the state vector and the observed values; R i is the observed noise covariance matrix.

[0097] The calculation formula for the Kalman gain is:

[0098]

[0099] Furthermore, combining the state vector and observed value information and the Kalman gain, a new state prediction and state covariance are constructed:

[0100]

[0101] The above steps are carried out cyclically, and the number of cycles is determined by the length of the time series.

[0102] In S3, the fact that the length of the bone remains unchanged during movement is a basic physical property, which is designed as the "bone length constraint" in human pose estimation. This invention adopts this constraint and improves it to fine-grained filtering, combining the naturalness of movement with the bone length consistency in time to correct the positional relationship of bone points.

[0103] In this invention, the selection of bone point pairs usually needs to be determined according to the specific characteristics of the dataset, and different datasets have different bone point information. For example, in the boxing action dataset, key joints such as the head, shoulders, elbows, hands, hips, knees, and ankles are usually included to support the information required for action recognition. The selected bone point pairs are usually the two ends of important bones, and there is a direct bone connection between the points, such as (knee, ankle) and (shoulder, elbow); however, (head, shoulder) usually cannot be used as a bone point pair because there is no direct bone connection.

[0104] Further, receive the data obtained by coarse-grained filtering. For the b-th bone or bone point pair, set the start end of the bone point pair and the end end to include the temporal coordinate values of x, y, and z of this bone point.

[0105]

[0106] where l represents the length of the time series, and x, y, and z correspond to the three-dimensional coordinates of this bone point.

[0107] Consider the two bone points at both ends of a bone and which represent the head bone point and the tail bone point respectively; sort and remove the data in the first l / 4 and the last l / 4 to avoid the influence of outliers.

[0108] The approximate true bone length is:

[0109]

[0110] where BL b is the length of the b-th bone, mean(·) is the operation of taking the mean, sort(·) is the sorting operation, which can be in ascending or descending order, and finally take the value within the interval.

[0111] Calculate the bone point position error vector, that is, the difference that can be repaired by fine-grained filtering:

[0112]

[0113] This invention is based on the bone midpoint M p and the bone direction vector v p, correct the positions of the bone points. The calculation formulas for the middle bone point and the bone direction vector are as follows:

[0114]

[0115] The chronological coordinate values of the corrected head bone point and tail bone point are:

[0116]

[0117] Perform fine-grained filtering on the selected bone point pairs, adjust the positions of the bone points for each timestamp from the perspective of the coordinate values, and finally output the chronological data of the bone point movements after filtering.

[0118] In S4, for the data output by the fine-grained filtering, select key joints to reduce the dimension of the data itself. The present invention uses the principal component analysis (PCA) method to capture key joints with more information content, and derives features applicable to the boxing scenario and having discrimination for downstream tasks based on the key joints.

[0119] First, extract the chronological 3D coordinates of the key joint points of the data output by the fine-grained filtering. Assume that the fine-grained filtering X is obtained, with a dimension of l×m, where l is the length of the time series and m is the dimension of X, that is, the dimension of the chronological data of the bone point movements.

[0120] Calculate the covariance matrix applicable to the PCA algorithm:

[0121]

[0122] Among them, is the mean vector of the data. The covariance matrix S reflects the characteristics of each joint, that is, the linear correlation degree between the joint coordinate dimensions at different timestamps. The elements on the diagonal are the variances of each feature, and the elements off the diagonal are the covariances between different features.

[0123] Perform eigenvalue decomposition on S, solve the characteristic equation |S - μI| = 0, and obtain the eigenvalues μ1 ≥ μ2 ≥ … ≥ μ s ≥ … ≥ μ m and the corresponding eigenvectors e1, e2, …, e s , …, e m . The eigenvalue μ s represents the variance size of the s-th principal component. The larger the variance, the more information the principal component contains. e s represents the projection direction of the original feature on the s-th principal component.

[0124] Arrange them in descending order of eigenvalues, and select the eigenvectors corresponding to the first a largest eigenvalues to form the principal component matrix E a ; The selection of a is based on the cumulative contribution rate CR aOK, the formula is:

[0125]

[0126] Among them, m i is the index value of the variable, ranging from 1 to m. Select CR a Reach 80% to 95% of the principal components.

[0127] Observe the selected principal component matrix E a The coefficients corresponding to the original joint features in the equation are as follows. Since each joint has three coordinate dimensions, the coefficients corresponding to the three dimensions are comprehensively considered, and the square sum of the corresponding coefficients of each joint is calculated and then squared to obtain the comprehensive coefficient of each joint. Joints with larger absolute values ​​of comprehensive coefficients have higher weights in the principal components. These joints play a key role in distinguishing actions (such as offensive and defensive actions). They are identified as key joints, and the three-dimensional coordinate time series data corresponding to the key joints are used as part of the multivariate time series output in this step S4.

[0128] Secondly, derive features from key joints; specifically, derive velocity v t With acceleration a t :

[0129]

[0130] Among them, P t is the 3D coordinate of a key joint point at time t, T is the time length between two timestamps, capturing the explosiveness characteristics of the action execution.

[0131] Derived joint trajectory curvature κ t , used to distinguish the trajectory differences of difficult-to-distinguish categories such as swing punches and uppercuts:

[0132]

[0133] Derived shoulder-wrist-hip triangle angle:

[0134]

[0135] in They are the coordinates of the shoulder, wrist, and hip joints, respectively, reflecting the coordination pattern between the upper limb force and the trunk.

[0136] Furthermore, data other than the time-series 3D coordinates and derived features of key joints are removed. Through manual annotation or automatic peak detection, the key stages of the action are determined, including the intention to start the action, the execution of the action, and the return to the forward / reverse frame state at the end of the action.

[0137] Finally, the data segments of the key stages of the action are obtained as multivariate time series data.

[0138] In S5, a shapelet is a feature representation method for time series classification, which is a short and representative time series segment. Mine the shapelets of the skeletal point motion time series data and screen out the optimal set of shapelets.

[0139] As Figure 4 shown, it is the shapelet feature mining structure diagram provided by the embodiment of the present invention.

[0140] Furthermore, use a sliding window with a window length of τ to crop the multivariate time series data and generate a candidate Shapelet set. Each shapelet needs to satisfy:

[0141]

[0142] Among them, τ v is the speed threshold of the joint point with the most intense movement, and the default value is 2m / s. The value range of t is 1 to τ, and the average speed is calculated by traversing the entire window; since v t contains the speeds of multiple joint points, so if the highest speed in this vector, that is, the maximum value, reaches 2m / s, the condition can be satisfied.

[0143] Furthermore, use FastDTW to calculate the distance metric between shapelets. Given a candidate shapelet S and a time series X K , this time series X K is taken from the multivariate time series data, that is, the data segment of the key action phase, and the distance calculation formula is obtained as:

[0144] Distance(S,X K ) = FastDTW(S,X K ).

[0145] Furthermore, evaluate the discrimination ability of each candidate shapelet. The higher the information gain of the shapelet, the more effectively it can distinguish different categories of time series. Specifically, calculate the distances {d i} between each candidate shapelet and all training samples in the data set, select the median as the threshold d split , and divide the data set into two parts; then, calculate the difference in information entropy before and after the division, that is, the information gain, which is defined as:

[0146]

[0147] Among them, H(D) is the information entropy of the data set D, and the data set D is the set of all multivariate time series data output by S4; D + and D -are distances d respectively i ≤d split and d i >d split of the sample subsets

[0148] Retain the high-discrimination shapelets with IG(S)≥ε ig =0.15 to obtain the optimal shapelet set. Among them, ε ig is used as the threshold for screening shapelets. If the information gain of S is greater than this value, then this S is classified into the optimal shapelet set

[0149] Furthermore, through deep learning frameworks such as pytorch, the feature structure of the shapelet is encoded into the weight matrix of the convolution kernel, that is, the numerical values in the time dimension are mapped to the weight values of the convolution kernel in the corresponding dimension, with the dimension of τ×m

[0150] Furthermore, define the gait convolution formula to calculate the local matching score

[0151] D conv (X K ,S q )=X K *W q

[0152] where S q is the q-th sub-shapelet of the optimal shapelet set, 1≤q≤s num ; W q is the convolution kernel weight matrix; finally, the local matching score is obtained where S num is the number of shapelets, w is the number of matching values generated by the convolution; if the step size is 1, then w=l k -τ+1, l k is the sequence length of the input X K

[0153] In S6, the boxing action recognition task is performed by a multivariate time series classifier, and the preprocessed time series data of boxers during training and competition scenarios in the ring is used as the input. This scenario involves two major categories of actions, namely offensive actions and defensive actions; among them, the offensive actions include: straight punches, swinging punches, hook punches, two-punch combinations; the defensive actions include: covering the head, blocking, leaning, dodging, parrying, arm swinging, sliding steps, cross steps

[0154] ​The present invention designs a multivariate time series classifier dedicated to complex boxing motion recognition, and assigns corresponding action labels to the motion time series data of each skeletal point. In the training stage, the training set and the test set are divided in the ratio of 8:2, and each sample corresponds to a true action label, which is annotated by professional sports science practitioners and computer-related practitioners.

[0155] Such as Figure 5 is a schematic diagram of a multivariate time series classifier provided by an embodiment of the present invention.

[0156] Design a dual-channel Transformer mechanism. One is a class-specific channel, and the input is the local matching score matrix and the action key stage sequence of each action sample, which captures class-related local features. The other is a global feature capture channel, and the input is the multivariate time series data output by S4, which captures long-range dependencies.

[0157] Further, for the class-specific channel, it includes a Transformer encoder. The Transformer encoder includes a linear projection module that converts the input into a dimension adapted to the model, a multi-head attention mechanism, and a feed-forward neural network FFN. The matrix S is linearly projected by the linear projection module scores into a query Q matrix, and the action key stage sequence of each action sample is linearly projected into a key K and a value V matrix. For each head, based on the multi-head self-attention layer of this channel, the dot product similarity between the query and the key is calculated respectively:

[0158]

[0159] wherein, d model is the dimension of the vector. These similarities are normalized by the softmax function to obtain the attention weight matrix Attn; the value matrix of each head is weighted and summed based on this attention weight to obtain the output of each head; then the outputs of each head are concatenated and linearly transformed again, and the linearly transformed features are input into the feed-forward neural network to obtain the class-specific feature h class .

[0160] Furthermore, for the global feature capture channel, it includes a Transformer encoder. The Transformer encoder includes a linear projection module that converts the input into a dimension adapted to the model, a multi-head attention mechanism, and a feed-forward neural network FFN. The multivariate time series data output by S4 is used to capture long-range class relationships. This data is linearly mapped to query Q, key K, and value V matrices through the linear projection module. For each head, the dot product similarity between the query and the key is calculated as the attention weight, and the value matrices of each head are weighted and summed to obtain the output of each head. Then, the outputs of multiple heads are concatenated, linearly transformed, and the linearly transformed features are input into the feed-forward neural network FFN to obtain the global feature h global .

[0161] Furthermore, a gating mechanism is designed for dynamic weighted fusion:

[0162] g = σ(W g [h class ; h global )

[0163] h = g ⊙ h class + (1 - g) ⊙ h global

[0164] where h is the weighted feature after the gating mechanism fusion, and g is the weight of the gating mechanism, generated by the softmax function σ(·). When g is close to 1, it indicates that class-specific features are more important, and the fusion result will be more biased towards the features of the class-specific pathway. When g is close to 0, the features of the global feature capture pathway dominate in the fusion result. Through this dynamic weighting method, the complementarity of two different types of features is achieved, enhancing the model's representation ability for complex patterns.

[0165] Furthermore, the fused features are passed through an average pooling layer to obtain the probability distribution of each action category.

[0166] Furthermore, cross-entropy loss is used as the loss function for classifier training. The cross-entropy loss is expressed as:

[0167] Loss CE = -∑ c p c log(q c )

[0168] where q c represents the probability of the true label in the c-th class, and p i represents the probability predicted by the classifier for the c-th class. By minimizing the cross-entropy loss, the model is trained to improve the consistency between its predicted probability distribution and the true distribution. During the training process, the parameters of the model are adjusted through the backpropagation algorithm and the gradient descent optimizer to reduce the cross-entropy loss.

[0169] As Figure 3 , on the other hand, an embodiment of the present invention provides a boxing motion recognition system, including a data preprocessing module, a coarse-grained filtering module, a fine-grained filtering module, a boxing-specific feature derivation module, a Shapelet feature mining module, and a multivariate time series classifier.

[0170] Data preprocessing module: Based on the collected bone point motion time series data, perform data preprocessing on the time series data to obtain preprocessed data.

[0171] The input of the data preprocessing module is the original bone point motion time series data, that is, the sequence of three-dimensional coordinates of human key bone nodes changing with time, and it contains noise and anomalies. Considering that the time series lengths of each sample are different, the present invention unifies all sequences to the same length. For the missing time series data caused by occlusion, an improved linear filling is used to fill the missing values. In order to eliminate the dimensionality effect between different dimensional coordinate values, a min-max normalization method is embedded. The output is preprocessed data, with a range of [0,1].

[0172] Coarse-grained filtering module: Perform coarse-grained filtering based on the unscented Kalman filter on the preprocessed data to smooth the bone point motion time series data.

[0173] The input of the coarse-grained filtering module is the preprocessed data, which still contains noise and anomalies, and the time series data is smoothed through the coarse-grained filtering module. The coarse-grained filtering module starts from the skeleton level and models it as a motion chain system. In order to accurately describe the motion state of the system, the change in speed is introduced as part of the state vector, and the nonlinearity of the state transition function and the observation function is processed through unscented transformation. The initial value of the system state vector is determined as the information contained in the first timestamp of the bone point motion time series data. In order to record the change in speed, a zero vector of the same dimension is concatenated. Determine the covariance matrix, state noise, and observation noise matrix. The covariance matrix represents the system uncertainty; the state noise matrix describes the random noise in the system dynamic model; the observation noise matrix represents the random noise in the observation model. For the determination of the matrices, the characteristics of the data set need to be considered. For example, if the coordinate value of a certain variable has a greater degree of jitter and anomaly due to factors such as occlusion, its covariance matrix should be set larger. The coarse-grained filtering module calculates a set of sigma points at each time step to approximate the probability distribution of the nonlinear system. These points have weights W (mean) and weight W (cov), a total of 2m+1 sigma point sets are generated; these sigma points are input into the nonlinear state transfer function, and the system and covariance matrices are iteratively updated. The sigma points output by the state transfer function are input into the observation function H to reconstruct the observations, and the observation covariance matrix and the cross covariance between the system and the observations are calculated. The Kalman gain is calculated, and the new state prediction and state covariance are constructed by combining the state transfer model and the observation model information. The above steps are repeated and the output is coarse-grained filtered data.

[0174] Fine-grained filtering module: Using kinematic prior knowledge, we design fine-grained filtering based on bone length constraints to suppress noise and anomalies at the coordinate value level.

[0175] The input of the fine-grained filtering module is the output of the coarse-grained filtering module. The principle of bone length invariance is used to fine-tune the position of the bone points after coarse-grained filtering. In combination with the characteristics of the data set, appropriate bone point pairs are selected. The bone point pairs should contain important joint information for boxing action execution to support boxing action recognition, and should comply with the bone point pair selection rules, that is, there should be direct bone links between the bone points. The distance of the bone point pairs with time as the dimension is calculated to obtain the bone point pair length sequence. In order to avoid the influence of noise and anomalies, the sequence is sorted and intervals are delineated, and the average is taken to obtain the approximate true bone length. Combined with the approximate true bone length and the bone point pair length sequence, the position error vector is calculated to obtain the difference that can be repaired by fine-grained filtering. The bone point motion timing data is updated through the bone middle point and the bone direction vector, and the skeleton structure depicted by the bone point motion timing data is corrected.

[0176] Boxing-specific feature derivation module: Based on the prior knowledge of boxing scenes, boxing scene-specific features are derived.

[0177] Specifically, PCA is used to capture key joints with greater information content. The covariance matrix of the PCA algorithm is calculated, and eigendecomposition is performed to obtain the eigenvectors corresponding to the first a largest eigenvalues ​​to form a principal component matrix; the comprehensive coefficient of each joint is obtained based on the principal component matrix, and the key joints are obtained by sorting; the velocity and acceleration, joint trajectory curvature, and shoulder-wrist-hip triangle angle are derived for the key joint points; data other than the time series 3D coordinates and derived features of the key joint points are eliminated to obtain multivariate time series data.

[0178] Shapelet feature mining module: mines and filters shapelets of skeleton point motion time series data to obtain the local matching score between shapelet and sample data.

[0179] Specifically, for the multivariate time series obtained by the boxing-specific feature derivation module, the key stages of the action are determined through manual annotation or automatic peak detection; the key stages of the action are cropped by a sliding window; the distance metric between shapelets is calculated using FastDTW; high-discrimination shapelets are selected through the information gain metric, and then the optimal shapelet set is obtained; through deep learning frameworks such as pytorch, the feature structure of the optimal shapelet set is encoded into the weight matrix of the convolutional kernel, and the local matching score is calculated through the gait convolution formula for class-specific feature capture of the multivariate time series classifier.

[0180] Multivariate time series classifier: Design a multivariate time series classifier based on a two-channel Transformer, and train it using cross-entropy loss to obtain the recognition result.

[0181] The two-channel Transformer mechanism includes a class-specific channel and a global feature capture channel. For the class-specific channel, the local matching scores output by the Shapelet feature mining module are linearly mapped to the query Q matrix, and the key stage sequence of each action sample is linearly mapped to the key K and value V matrices; class-specific features are obtained through the multi-head self-attention mechanism and linear transformation. For the global feature capture channel, the multivariate time series obtained by the boxing-specific feature derivation module is linearly mapped to the query Q, key K, and value V matrices; global features are obtained through the multi-head self-attention mechanism and linear transformation. The class-specific features and global features are dynamically fused through a gating mechanism, and finally the fused features pass through an average pooling layer to obtain the probability distribution of each action category.

[0182] The effects of the present invention are illustrated based on the following experiments.

[0183] In this embodiment, four Z CAM E2 high-frame-rate cameras are arranged around the four corners outside the boxing ring and used to synchronously record the actions of boxers during training at a resolution of 1080P and a frame rate of 60fps. The dataset contains a total of 175 confrontation videos of red and blue boxers, covering 13 action categories, with a total duration of 14 minutes and 17 seconds. These 13 action categories specifically include 4 offensive actions: straight punch, swinging punch, hook punch, and two-punch combination; and 9 defensive or footwork actions: dodge, block, sway, slap, slide step, lean, cover head, retreat, and circumferential step. Subsequently, the Fastmove motion capture software is used to process each video to calculate the three-dimensional spatial coordinate sequences (x, y, z) of 21 key joints of the red and blue athletes respectively. The schematic diagram of data acquisition is as Figure 6 shown, and the bone point distribution is as Figure 8 shown.

[0184] In this embodiment, the 350 action sequences generated from 175 videos are divided into a training set and a test set at a ratio of 8:2.

[0185] The boxing action recognition task of the present invention is performed by a multivariate time series classifier, which can capture time and variable dependencies. Regarding the 3D coordinates of the joint points as variables can also achieve higher accuracy in the action recognition task.

[0186] The multivariate time series classifier is constructed based on a dual-channel Transformer model. One is a class-specific channel, with the input being the local matching score matrix and the action key phase sequence of each action sample, which captures class-related local features. The other is a global feature capture channel, with the input being the multivariate time series data output by S4, which captures long-range dependencies. The boxing action recognition performance and loss are as Figure 7 shown, and the highest accuracy is 0.8827. Figure 7 (a) in is the image of the boxing action recognition accuracy, Figure 7 (b) in is the loss image. The confusion matrix corresponding to the highest accuracy is as Figure 9 shown. Because the class-specific channel captures the local features of the action sequence, and the local features incorporate the prior knowledge of the boxing scene, including specific features such as measuring punches and upper limb exertion in the boxing scene, and high-discriminative shapelets are extracted as queries to help the Transformer in the class-specific channel capture local features. In addition, the global feature capture channel uses the Transformer to capture the dependencies of long sequences, thereby mining global features. The class-specific features and global features are fused through a gating mechanism, giving an adaptive weight fusion, so that the fused features have high discriminability for complex actions in the boxing scene.

[0187] In the present invention, coarse-grained filtering uses the unscented Kalman filter adapted to non-linear motion data to smooth the time series data of the x, y, and z coordinates of each bone point. Among them, a prediction-update mechanism is designed to predict and update the current state based on the previous sequence.

[0188] Using the kinematic prior knowledge that the length of the bone remains unchanged during movement, fine-grained filtering is designed to combine the naturalness of movement with the consistency of bone length in time, and correct outliers at the coordinate value level.

[0189] Since the influence degrees of noise and outliers on each sample are different, the basis for selecting cases is the situation where the influence of noise is relatively large; therefore, the second sample of 70 action sequences in the test data set is selected, corresponding to the red side making a swinging punch and the blue side making a sliding step movement to avoid the swinging punch in the video. Due to the boxing fence blocking the ankle joint, heel, and toes, abnormal skeletons and coordinate values occur.

[0190] Figure 10Show the anomalies of the original skeleton structure, abnormal coordinate values, and the data repair capabilities of coarse-grained filtering and fine-grained filtering. For the skeleton structure, the original 3D skeletons (blue) and the filtered skeleton structures (red) of the 125th, 129th, and 138th frames are intercepted. It can be seen that the coarse-grained filtering has anti-jitter ability. Because when designing this invention, not only displacement changes but also speed changes are considered, and it has better noise suppression ability in complex boxing scenarios. For some abnormal coordinate value situations in the time series, fine-grained filtering is adopted. The idea of this prior knowledge of bone length constraint has an observable effect on the anomaly suppression of real data.

[0191] Figure 11 Show the visualization of the noisy time series data and the filtered data of some bone point coordinates, where Figure 11 (a) shows the time series curve of the z-axis of the right wrist; Figure 11 (b) shows the time series curve of the z-axis of the right elbow. Figure 12 Show the visualization of the bone point pair length constraint under fine-grained filtering. It can be noted that there are more noises and mutation values during the action execution from the 100th frame to the 200th frame, and after coarse-grained filtering and fine-grained filtering, there are better smoothing and anomaly suppression effects.

[0192] Figure 12 Show the visualization of the bone point pair length constraint under fine-grained filtering, where Figure 12 (a) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the neck and left shoulder; Figure 12 (b) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the left elbow and left wrist; Figure 12 (c) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the right knee and right ankle; Figure 12 (d) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the right shoulder and right elbow; Figure 12 (e) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the left knee and left ankle; Figure 12 (f) shows the original bone length sequence and the bone length sequence after fine-grained filtering of the right ankle and right heel. The original sequence is affected by outliers and noises, and the skeleton is severely deformed so that the bone length is not within a reasonable range. Due to the boxing ring fence blocking the leg and foot related joints, the mutation values are more serious, as shown in Figure 12 (e) and (f). The fine-grained filtering proposed by this invention has an effective suppression of outliers when the change range of the bone length is within 30 millimeters.

[0193] The present invention also evaluates the effectiveness of dual-channel attention of a multivariate time series classifier in boxing action recognition, and explores the changes in the accuracy of boxing action recognition under two strategies: only class-specific channels and only global feature capture channels, and analyzes the discriminative ability of these two features in action recognition. When the gating fusion mechanism is removed from the present invention, the accuracy and loss curves when there are only class-specific channels are as shown in Figure 13 in (a) of Figure 13 and (b) of Figure 13 where (a) of Figure 14 is an accuracy image, and the highest accuracy is 0.8657; the accuracy and loss curves when there are only global feature capture channels are as shown in

[0194] and the highest accuracy is 0.8228. It can be seen that when there are only class-specific features, the classifier loses the ability to model long sequences. Compared with the classifier based on dual-channel attention, the best accuracy is 0.8827, and there is a certain loss in accuracy. When there are only global features, the classifier loses the high discriminative ability for complex actions in the boxing scenario, and only relies on the advantage of the Transformer itself to capture long-range dependencies, resulting in a large loss in the accuracy of boxing action recognition. Therefore, the multivariate time series classifier based on dual-channel attention is effective.

[0195] The above is only the preferred solution of the present invention, and is not intended to further limit the present invention. All equivalent changes made by using the content of the specification and drawings of the present invention are within the protection scope of the present invention.

Claims

1. A boxing action recognition method based on prior knowledge and multivariate time series classification, characterized in that It includes the following steps: S1. Based on the collected kinematic point motion time-series data, preprocess the time-series data to obtain the preprocessed data; S2. Apply coarse-grained filtering based on unscented Kalman filter to the preprocessed data to smooth the kinematic point motion time-series data; S3. Utilize kinematic prior knowledge to design fine-grained filtering based on bone length constraint to correct the position relationship of kinematic points; S4. Based on the prior knowledge of boxing scenarios, derive multivariate time series with boxing-scenario-specific features; S5. Mine and screen the time-series segments (shapelets) of kinematic point motion time-series data, and obtain the local matching scores between the shapelets and the sample data; S6. Design a multivariate time series classifier based on dual-channel Transformer, and train it using cross-entropy loss to obtain the recognition result.

2. The boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 1, wherein In S1, the kinematic point motion time-series data is collected by a contactless motion capture device, which contains the coordinate value sequences of multiple kinematic points changing over time; the preprocessing of the collected data is specifically as follows: S11. Set the time-series length value of the action segment, and align the sequences to the same length to ensure the consistency of time-series length; S12. Use improved linear interpolation to fill in the missing values, and fill the "0" coordinate values with real coordinate values; S13. Use the min-max normalization method to scale the three-dimensional coordinate values of each kinematic node to the interval [0, 1].

3. The boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 1, characterized in that S2 is specifically as follows: S21. Set the initial value of the state vector, the state covariance matrix, the process noise covariance matrix, and the observation noise covariance matrix; S22. Based on the state vector, calculate a set of sigma points at each time step; S23. Based on the sigma points and the state transition function, iteratively update the state vector and the state covariance matrix, and output a new set of sigma point sets; S24. Calculate the Kalman gain according to the cross-covariance between the state vector and the observation value and the observation covariance matrix, and construct a new state prediction and state covariance matrix.

4. The boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 1, characterized in that S3 is specifically as follows: S31. According to the human skeleton structure, select appropriate kinematic point pairs, and calculate the bone point pair length sequences based on the kinematic point motion time-series data information; S32. Sort the bone point pair length sequences and delimit intervals, and take the average to obtain an approximate real bone length; S33. Correct the skeleton structure depicted by the kinematic point motion time-series data by calculating the bone midpoint and the bone direction vector.

5. A boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 1, characterized in that, S4 is specifically as follows: S41. Through principal component analysis, screen the key joints with distinguishing features in offensive and defensive actions, and derive their velocity and acceleration features; S42. Incorporate the trajectory information of the key joint points, add the trajectory curvature information of the joints, and capture the difficult-to-classify category features; S43. Derive limb coordination features including the shoulder-wrist-hip triangle angle to describe the coordination relationship between the limbs in boxing actions; S44. Define the key stages of the action, and obtain the data segments corresponding to the key stages of the action through manual annotation or peak detection, which are the multivariate time series data.

6. The boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 1, characterized in that S5 is specifically as follows: S51. Obtain the shapelets by cropping through a sliding window; S52. The distance between shapelets is measured using FastDTW, and high-discriminative shapelets are screened through information gain to obtain the optimal shapelet set; S53. The matching degree between the kinematic time-series data samples of skeleton points and shapelets is measured through gait convolution to obtain local matching scores.

7. A boxing action recognition method based on prior knowledge and multivariate time series classification according to claim 6, characterized in that S6 is specifically as follows: S61. A dual-channel Transformer mechanism is designed, one is a class-specific channel, and the other is a global feature capture channel; S62. The class-specific channel obtains a query matrix based on the local matching scores, maps the key-phase sequence of each action sample to obtain the key and value matrices, and obtains class-specific features through multi-head self-attention; S63. The global feature capture channel obtains long-range dependencies based on the multivariate time series output by S4; S64. A gating mechanism is designed to dynamically weight and fuse the output features of the dual channels to achieve the complementarity of different types of features; S65. By minimizing the cross-entropy loss, the loss is propagated throughout the network for end-to-end training; S66. The classifier outputs the prediction distribution of each possible action of the time series, and the action label is obtained to complete the boxing action recognition task.

8. A boxing action recognition system based on prior knowledge and multivariate time series classification, characterized in that, For implementing the method according to any one of claims 1-7, the system includes: Data preprocessing module: Based on the collected kinematic time-series data of skeleton points, preprocess the time-series data to obtain preprocessed data; Coarse-grained filtering module: Perform coarse-grained filtering based on unscented Kalman filtering on the preprocessed data to smooth the kinematic time-series data of skeleton points; Fine-grained filtering module: Utilize kinematic prior knowledge to design fine-grained filtering based on bone length constraints to suppress noise and anomalies at the coordinate value level; Boxing-specific feature derivation module: Derive boxing-scene-specific features based on the prior knowledge of the boxing scene; Shapelet feature mining module: Mine and screen the shapelets of the kinematic time-series data of skeleton points to obtain the local matching scores between shapelets and sample data; Multivariate time series classifier: Design a multivariate time series classifier based on a dual-channel Transformer, and train it using cross-entropy loss to obtain the recognition result.

Citation Information

Cited By

  • Motion control method, boxing control method and system of humanoid robot

    CN121989238A

  • Vehicle longitudinal interaction intention estimation method, electronic equipment and storage medium

    CN122347870A