Boxing scene-oriented data refinement and boxing action recognition method and system
Through multi-bone point time series data acquisition and recognition model based on LSTM and TCN networks, the problems of data distortion and missing in boxing movement recognition are solved, and the precise identification and classification of boxing movements are achieved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202510154566.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art has problems of data distortion and missing in boxing movement recognition, and traditional methods cannot accurately identify boxing movements, especially ignoring the important role of other joints such as legs, torso, etc. in boxing.
Using multi-bone point time series data collection, through abnormal and missing detection, processing and filling, a boxing action recognition model based on LSTM and TCN network is constructed to achieve accurate identification and classification of boxing actions.
It improves the accuracy of boxing movement recognition, can more comprehensively identify complex movements in boxing, reduces the impact of data distortion, and provides real-time feedback.
Smart Images

Figure CN119992662A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and in particular relates to a boxing scene-oriented data refinement and boxing action recognition method and system. Background Art
[0002] As a high-intensity, fast-response combat sport, boxing has extremely high requirements on the athlete's technique, strength and strategy. In the field of boxing training, the coach's guidance of the athlete's movements is crucial. Accurate movement recognition can help the coach to promptly detect and correct the athlete's technical errors and improve the training effect. However, boxing movements are complex and varied, including a variety of offensive punches, as well as various defensive movements, footwork and feints. Traditional manual observation and evaluation methods are inefficient and subjective, and cannot provide real-time feedback.
[0003] Traditional IMU or other sensing methods use sensors attached to the boxer's wrist; some marking methods require sticking markers on the athlete. In boxing matches, these devices or methods will not only affect the athlete's normal performance, but will also cause the markers to fall off or data to be distorted due to collisions.
[0004] Traditional boxing action recognition methods often only focus on the movement of the hands or a few key joints, ignoring the important role of other joints in boxing, such as legs and torso. This method has certain limitations in recognizing complex boxing actions, such as footwork and body rotation.
[0005] In summary, the existing algorithms have the following shortcomings:
[0006] 1. Traditional IMU or other sensing methods have limitations in boxing match scenarios and cannot accurately identify movements;
[0007] 2. Traditional manual observation and evaluation methods are easily affected by personal bias, resulting in subjective evaluation results.
[0008] 3. Traditional boxing motion recognition methods only focus on the movement of the hands or a few key joints, ignoring the important role of other joints in boxing such as legs and torso.
[0009] 4. Traditional IMU or other sensing methods will affect the normal performance of athletes in boxing competition scenarios and may cause markers to fall off or data to be distorted.
[0010] Therefore, how to solve the problems of data distortion and data missing in action recognition and improve the accuracy of action recognition. Summary of the invention
[0011] The purpose of the present invention is to provide a data refinement and boxing action recognition system and method for boxing scenes to solve the problems raised in the above background technology.
[0012] The object of the present invention is achieved by: a data refinement and boxing action recognition method for boxing scenes, characterized in that the method comprises the following steps:
[0013] Step S1: Collecting time series data of multiple skeleton points of a boxer;
[0014] Step S2: perform anomaly and missing detection on time series data;
[0015] Step S3: Processing the detected abnormal and missing time series data;
[0016] Step S4: construct a boxing action recognition model based on LSTM and TCN network, and perform training;
[0017] Step S5: input the processed time series data into the boxing action recognition model to realize the boxing action recognition task.
[0018] Preferably, the step S1 collects the time series data of multiple skeleton points of the boxer, and the specific operations are:
[0019] Step S1-1: four cameras are arranged around the boxing ring, and real-time video streams are acquired by the cameras;
[0020] Step S1-2: Use the OpenPose deep learning algorithm to obtain the corresponding two-dimensional coordinate sequence of multiple skeleton points in the video stream;
[0021] Step S1-3: Using the pinhole imaging principle, the projection of the two-dimensional coordinates on the imaging plane to the three-dimensional coordinates in space is realized to obtain the three-dimensional coordinates of multiple skeleton points;
[0022] Step S1-4: Convert the three-dimensional coordinates of multiple skeleton points into time series data.
[0023] Preferably, in step S2, the time series data is subjected to abnormality and missing detection, specifically:
[0024] Step S2-1: Represent the multi-skeleton point time series data in matrix form, as follows:
[0025] F=[f1,f2,…,f T ];
[0026] Among them, f i =[x i,1 ,y i,1 ,z i,1 ,...,x i,K ,yi,K ,z i,K ] T ,i=1,...,T, t is the total number of frames in the time series, K is the number of skeleton points;
[0027] Step S2-2: define anomalies and missingness of time series data;
[0028] Missing data means that in a multivariate time series, the data of certain timestamps are not recorded or lost, which is manifested as null values for some timestamps in the matrix F;
[0029] Data anomaly refers to the situation in which the data of some variables at certain timestamps in a multivariate time series deviate significantly from the normal pattern, which is manifested in the oscillation of some data in the matrix F and the outliers of data at a certain time length;
[0030] Step S2-3: label samples of abnormal and missing time series data;
[0031] If there are missing or abnormal data, the multivariate time series samples are labeled as [1,1];
[0032] If there is no missing or abnormal data, the multivariate time series samples are labeled as [0,0].
[0033] Preferably, the detected abnormal and missing time series data are processed in step S3, specifically:
[0034] Step S3-1: For long time sequences of different action combinations, the fuzzy C-means clustering method is used to divide the long sequence, so that the same or similar motion sequences can be classified into the same segment:
[0035] For the matrix F = [f1,f2,…,f T ], according to the membership of each data point to each class w slices f i (i=1,...,w) is divided into c fuzzy groups, and the membership matrix S ranges from [0,1]. After normalization, the sum of the membership of a data set is equal to 1, that is:
[0036]
[0037] Among them, s ij Represents the membership of slice j to cluster i. The closer the value is to 1, the more the slice belongs to cluster i.
[0038] The cost function of fuzzy C-means is:
[0039]
[0040] Among them, 0 <sij <1, c i is the cluster center of fuzzy group i, d ij =||c i -f j || is the Euclidean distance between the i-th cluster center and the j-th slice, m is the fuzzy index, m>1;
[0041] Construct a new objective function to ensure that the value function can be optimized along with the constraints. The constructed new objective function is shown as follows:
[0042]
[0043] The necessary condition for minimizing the value function by taking the derivative of the input is:
[0044]
[0045] Among them, x j is the time series data; the membership matrix S is initialized by taking values in the range [0,1] and satisfies Condition, by calculating the cluster center c i And calculate the value function, iteratively update S, and finally get the cluster center c when the value function is minimum i and the membership matrix S;
[0046] Step S3-2: reconstruct all elements in the recovery matrix according to the low rank of the data matrix, and convert the filling problem into a convex optimization problem;
[0047] The convex optimization problem is:
[0048]
[0049] Among them, X is the restored matrix, M is the missing data matrix, Ω is the index set of known elements, and P Ω is the linear projection operator;
[0050]
[0051] Step S3-3: solving the convex optimization problem by projected approximate point algorithm;
[0052] Step S3-4: According to the growing Lagrange multiplier algorithm, the filled complete multi-skeleton point temporal motion data is obtained.
[0053] Preferably, in step S3-3, the convex optimization problem is solved by a projected approximate point algorithm, specifically:
[0054] For each sub-segment the missing data matrix where t wis the number of frames in the wth sequence after cluster segmentation, d is the data dimension, d = 3K;
[0055] Step S3-3-1: Set initialization parameters: X0=0, V0=0, n=1,
[0056] Step S3-3-2: Calculate the projection matrix Y of the missing data matrix M n , the iteration formula is
[0057] Step S3-3-3: Project the matrix Y n Perform singular value decomposition and solve the approximate point matrix X n ,Right now:
[0058] [U n ,S n ,V n ]=SVD(Y n );
[0059]
[0060] in, is the contraction operator, and its operation rules are:
[0061]
[0062] pass For the matrix S n Perform a contraction operation;
[0063] Step S3-3-4: During the solution process of the projection approximate point algorithm, the following convergence conditions are set:
[0064]
[0065] The convergence condition is and ε is the convergence threshold; through the first convergence condition, the low-rank matrix X is measured n and the reconstruction error of the original matrix M to determine the low-rank matrix X of the current iteration n Is it close enough to the observation matrix M? The second convergence condition is used to measure the difference between the results of two iterations, indicating the degree of convergence of the matrix update.
[0066] In step S3-4, the filled complete multi-skeleton point temporal motion data is obtained according to the growing Lagrange multiplier algorithm, specifically:
[0067] According to the growing Lagrange multiplier algorithm, V n The iterative update formula is
[0068] When the convergence condition is met, or n exceeds the threshold value n of the number of operations max , the solution X n As the optimal solution, the complete multi-skeleton point temporal motion data X after filling is obtained.
[0069] Preferably, the action recognition model includes a long short-term memory module LSTM, a multi-head attention mechanism module and a time domain convolution module TCN, the long short-term memory module LSTM is used to maintain the hidden vector h and the memory vector m, respectively control the state update and input of each timestamp, and capture the time dependency of the multivariate time series;
[0070] The long short-term memory module LSTM includes a forget gate, an input gate, and an output gate. The forget gate determines how much of the cell state at the previous moment is retained to the current moment c. t , the input gate determines the input x of the network at the current moment t How much is saved to the cell state c t , the output gate controls the control unit state c t How many outputs are there to the current output value h of the LSTM t ;
[0071] The input gate includes three inputs, which are the input value x of the network at the current moment. t , the output value h of LSTM at the previous moment t-1 And the cell state c at the previous moment t-1 ;
[0072] The output gate includes two outputs, the output value h of the LSTM at the current moment. t And the cell state c at the previous moment t-1 ;
[0073] The multi-head attention mechanism module includes a linear transformation module, a self-attention calculation module and a splicing module. The linear transformation module performs linear transformation to obtain the query, key and value of multiple heads:
[0074] Q i =QW i q , K i =KW i k , V i =VW i v ;
[0075] Among them, Q i is the query matrix of the i-th head, K i is the key matrix of the ith head, V i is the value matrix of the i-th head; l is the sequence length, n is the feature dimension of TCN output; Q is the query, K is the key, and V is the value;
[0076] The self-attention calculation module performs self-attention calculation on each head:
[0077]
[0078] Among them, n h is the number of heads;
[0079] The concatenation concatenates the outputs of all heads and restores them back to the input sequence dimension through linear transformation:
[0080] MultiHead(Q,K,V)=Concat(head1,...,head h )W O ;
[0081] in, is a trainable linear transformation weight matrix;
[0082] The self-attention calculation module performs independent self-attention operations on the input query q, key k, and value v through multiple heads, concatenates the results of all heads and obtains the final output through linear transformation, and pays attention to different parts of the input sequence to different degrees through multiple heads;
[0083] The temporal convolution module is composed of multiple layers of causal convolution kernels and dilated convolutions, and the output of each layer is transmitted through residual connections;
[0084] The time domain convolutional network (TCN) uses causal convolution kernels to model time series data, and models long time series data through residual connections and extended convolutions to capture local temporal dependencies of time series.
[0085] For the input signal x t And the causal convolution kernel W, the output of the causal convolution kernel is:
[0086]
[0087] Among them, w i is the weight of the convolution kernel; x t-i is the time series data input to TCN;
[0088] The extended convolution increases the receptive field of the time domain convolutional network for time series data and captures longer-term dependencies. The output of the extended convolution is:
[0089]
[0090] Among them, r is the expansion factor. As the number of layers increases, the receptive field gradually expands.
[0091] Preferably, the training of the action recognition model is performed by constructing a loss function, specifically:
[0092] Step S4-1: construct a smooth loss function to ensure the smoothness of the motion sequence;
[0093] X is the matrix obtained after data refinement; repeat the boundary elements of X to obtain X′:X1′ ,i =X2′ ,i =X 1,i , X′ 3K+1,i =X′ 3K,i =X 3K,i , Among them, 1≤i≤T;
[0094] Define O to be a symmetric tridiagonal matrix that satisfies the following conditions:
[0095]
[0096]
[0097] Among them, 2≤j≤t-1, Since the intervals between frames are equal, when 1≤j≤T, h j =1, then
[0098]
[0099] Step S4-2: Define a smooth loss function, which is:
[0100]
[0101] Where T is the number of frames in the sequence;
[0102] Step S4-3: According to the prior knowledge of the kinematic model, a bone length loss function is introduced. The bone length loss function is:
[0103] Set {l b |1≤b≤J-1} is the bone length sequence, where b is the bone index and the bone length loss is:
[0104]
[0105] in, and is the three-dimensional position of the joints at both ends of the skeleton in the i-th frame;
[0106] Step S4-4: Introduce cross entropy loss to evaluate and optimize the action recognition model;
[0107] The cross entropy loss function is used to calculate the difference between the probability distribution of each action category output by the classifier and the probability distribution corresponding to the true action label. The cross entropy loss is expressed as:
[0108] Loss c =-∑ i q i log(p i );
[0109] Among them, q i Represents the probability of the true label in the i-th category, p i Represents the probability that the classifier predicts the i-th category;
[0110] The total loss function of the action recognition model is:
[0111] Loss = λ1Loss s +λ2Loss b +λ3Loss c ;
[0112] Among them, λ1, λ2, and λ3 are the weight coefficients of smoothing loss, bone length loss, and cross entropy loss, respectively. The weight coefficients are hyperparameters in the training process and are fine-tuned according to the training situation and the action recognition performance of the test set.
[0113] Preferably, in step S5, the processed time series data is input into a boxing action recognition model to implement a boxing action recognition task, specifically:
[0114] Step S5-1: Split the processed time series data into a training set and a test set in a ratio of 8:2, normalize the time series data, and input the training set into the action recognition model;
[0115] Step S5-2: When the time series length T is greater than the number of variables 3K, it is input into the long short-term memory module LSTM, and the long short-term memory module LSTM provides the data of the variable at all T timestamps;
[0116] The dataset dimension is R N×T×3K , where N is the number of samples in the data set, T is the time series length of each sample, and K is the number of bone points;
[0117] Step S5-3: Obtain the local features of the time series data through the one-dimensional convolution block of the time domain convolutional network TCN, and then adjust the mean and standard deviation of each batch of data through batch normalization to make the distribution of the time series data more stable;
[0118] Step S5-4: The action recognition model learns complex patterns and relationships in the data by introducing nonlinearity through the ReLU activation function; the local features of the output sequence are used through the multi-head attention mechanism to obtain the long-range dependencies between features and capture the global context;
[0119] Step S5-5: The features of the dual-branch outputs are aggregated through concat, and the input vector is converted into the probability distribution of each class through softmax to achieve the effect of sequence classification;
[0120] Step S5-6: Adopt Adam optimizer and cross entropy loss function to back propagate the whole network.
[0121] A data refinement and boxing action recognition system for boxing scenes, characterized by:
[0122] The motion recognition system comprises an unmarked motion capture module, a data missing and anomaly detection module, a multi-skeletal point missing data filling module, a multi-skeletal point data repair module and a boxing motion recognition module. The unmarked motion capture module is used for collecting the time series data of multiple skeleton points of a boxer in actual boxing training and competition scenes to obtain the time series motion data of multiple skeleton points of the boxer; the data missing and anomaly detection module is used for detecting missing or anomaly of the output time series data, and marking the data with type labels by detecting whether there is missing or anomaly; the multi-skeletal point missing data filling module processes incomplete skeleton time series data, and reconstructs the skeleton motion data close to the real human body through the fuzzy C-means clustering algorithm and the projection approximate point algorithm;
[0123] The multi-skeleton point data repair module realizes the abnormal elimination and smoothing processing of the multi-skeleton point time series motion data, and the boxing action recognition module uses the boxing action recognition model to realize the boxing action recognition task.
[0124] Compared with the prior art, the present invention has the following improvements and advantages:
[0125] 1. By combining the fuzzy C-means clustering algorithm and the projection approximate point algorithm, we can reconstruct the human skeleton motion data close to the real one. By comprehensively analyzing the motion information of the entire human skeleton including multiple bone points, we can achieve accurate recognition and classification of boxing movements and improve the accuracy of movement recognition.
[0126] 2. Further improve the accuracy of action recognition by smoothing the motion sequence using different loss functions and fine-tuning different data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Figure 1 Flow chart of the method of the present invention.
[0128] Figure 2It is a structural diagram of the system of the present invention.
[0129] Figure 3 This is the structural diagram of the action recognition model. DETAILED DESCRIPTION
[0130] The present invention is further summarized below with reference to the accompanying drawings.
[0131] like Figure 1 As shown, a data refinement and boxing action recognition method for boxing scenes includes the following steps:
[0132] Step S1: Collecting time series data of multiple skeleton points of a boxer;
[0133] Step S1-1: four cameras are arranged around the boxing ring, and real-time video streams are acquired by the cameras;
[0134] Four cameras are connected to the markerless motion capture module through a synchronization interface, and the motion capture module supports one-click synchronous shooting. To ensure the stability of the time series output by the markerless motion capture module, a master camera is selected from the four cameras and connected to the motion capture module. The master camera and other slave cameras are connected through a dedicated synchronization line to achieve signal synchronization. The user sends a trigger signal to the master camera by pressing a button. After receiving the signal, the master camera immediately transmits the signal to all slave cameras through the synchronization line, so that the four cameras can start and end shooting at the same time, ensuring that the start and end times of the shooting are completely synchronized.
[0135] Step S1-2: Use the OpenPose deep learning algorithm to obtain the corresponding two-dimensional coordinate sequence of multiple skeleton points in the video stream;
[0136] Openpose can identify and locate the athlete's body posture and movement trajectory in the video, and extract key joints, such as the head, shoulders, elbows, wrists, hips, knees and ankles. For the i-th camera (i=1, 2, 3, 4), in the video shot by the i-th camera, the algorithm can obtain the two-dimensional coordinates of multiple bone points on the imaging plane. Openpose supports real-time video stream input and accurately locates the key joints of each frame in the video.
[0137] Step S1-3: Using the pinhole imaging principle, the projection of the two-dimensional coordinates on the imaging plane to the three-dimensional coordinates in space is realized to obtain the three-dimensional coordinates of multiple skeleton points;
[0138] For a joint point P in a certain space, (u i ,v i ) is the two-dimensional coordinate of P on the imaging plane of the i-th camera. The imaging on the camera can be represented by the pinhole imaging model, which is expressed as follows:
[0139]
[0140]
[0141] Among them, f x and f y are the horizontal focal length and vertical focal length of the camera respectively; u0, v0 are the pixel coordinates of the intersection of the camera optical axis and the imaging plane, f x 、f y , u0, v0 are the internal parameters of the camera, M i is the external parameter matrix from the world coordinate system to the camera i coordinate system, x w ,y w , z w is the three-dimensional coordinate of the joint point P in space, z c is the proportional coefficient, the default value is 1;
[0142] For the i-th camera, the three-dimensional coordinate point P is derived from the above formula: i (X w ,Y w ,Z w ) and the two-dimensional plane coordinates of the point imaged in a certain camera:
[0143]
[0144] in,
[0145] Step S1-4: converting the three-dimensional coordinates of multiple skeleton points into time series data;
[0146] The above mapping relationship is applicable to the four cameras. The rays emitted by the four cameras can intersect at one point, namely P(x w ,y w ,z w ). However, due to the inevitable noise in the data, the ray l i Deviation occurs, resulting in multiple rays not intersecting at one point, so it is necessary to estimate the spatial position of the joint point P in space; considering the distance between the rays, set a spherical area, the radius of the sphere is 10 times the shortest distance between the rays, and the moving point P' moves in the spherical area; combined with the OpenPose confidence parameter, the spatial point with the minimum distance to multiple rays in space is taken as the intersection of multiple rays, and this spatial point is taken as the spatial position of the joint point P.
[0147] The OpenPose deep learning algorithm outputs a two-dimensional coordinate sequence. Through the above process, we can further obtain a three-dimensional multivariate time series covering the motion information of multiple joints of the human body, that is, time series data. The variable of the multivariate time series is a certain dimension of a certain joint point.
[0148] Step S2: Perform anomaly and missing detection on time series data, specifically:
[0149] Step S2-1: Represent the multi-skeleton point time series data in matrix form, as follows:
[0150] F=[f1,f2,…,f T ];
[0151] Among them, f i =[x i,1 ,y i,1 ,z i,1 ,...,x i,K ,y i,K ,z i,K ] T ,i=1,...,T, t is the total number of frames in the time series, K is the number of skeleton points;
[0152] Step S2-2: define anomalies and missingness of time series data;
[0153] Missing data means that in a multivariate time series, the data of certain timestamps are not recorded or lost, which is manifested as null values for some timestamps in the matrix F;
[0154] Data anomaly refers to the situation in which the data of some variables at certain timestamps in a multivariate time series deviate significantly from the normal pattern, which is manifested in the oscillation of some data in the matrix F and the outliers of data at a certain time length;
[0155] Step S2-3: label samples of abnormal and missing time series data;
[0156] If there are missing or abnormal data, the multivariate time series samples are labeled as [1,1];
[0157] If there is no missing or abnormal data, the multivariate time series samples are labeled as [0,0].
[0158] Step S3: Process the detected abnormal and missing time series data, specifically:
[0159] Step S3-1: For long time sequences of different action combinations, the fuzzy C-means clustering method is used to divide the long sequence, so that the same or similar motion sequences can be classified into the same segment:
[0160] For the matrix F = [f1,f2,…,f T ], according to the membership of each data point to each class w slices f i(i=1,...,w) is divided into c fuzzy groups, and the membership matrix S ranges from [0,1]. After normalization, the sum of the membership of a data set is equal to 1, that is:
[0161]
[0162] Among them, s ij Represents the membership of slice j to cluster i. The closer the value is to 1, the more the slice belongs to cluster i.
[0163] The cost function of fuzzy C-means is:
[0164]
[0165] Among them, 0 ij <1, c i is the cluster center of fuzzy group i, d ij =||c i -f j || is the Euclidean distance between the i-th cluster center and the j-th slice, m is the fuzzy index, m>1;
[0166] Construct a new objective function to ensure that the value function can be optimized along with the constraints. The constructed new objective function is shown as follows:
[0167]
[0168] The necessary condition for minimizing the value function by taking the derivative of the input is:
[0169]
[0170] Among them, x j is the time series data; the membership matrix S is initialized by taking values in the range [0,1] and satisfies Condition, by calculating the cluster center c i And calculate the value function, iteratively update S, and finally get the cluster center c when the value function is minimum i and the membership matrix S;
[0171] Step S3-2: reconstruct all elements in the recovery matrix according to the low rank of the data matrix, and convert the filling problem into a convex optimization problem;
[0172] The convex optimization problem is:
[0173]
[0174] Among them, X is the restored matrix, M is the missing data matrix, Ω is the index set of known elements, and P Ω is the linear projection operator;
[0175]
[0176] Step S3-3: solving the convex optimization problem by projected approximate point algorithm;
[0177] For each sub-segment the missing data matrix where t w is the number of frames in the wth sequence after cluster segmentation, d is the data dimension, d = 3K;
[0178] Step S3-3-1: Set initialization parameters: X0=0, V0=0, n=1,
[0179] Step S3-3-2: Calculate the projection matrix Y of the missing data matrix M n , the iteration formula is
[0180] Step S3-3-3: Project the matrix Y n Perform singular value decomposition and solve the approximate point matrix X n ,Right now:
[0181] [U n ,S n ,V n ]=SVD(Y n );
[0182]
[0183] in, is the contraction operator, and its operation rules are:
[0184]
[0185] pass For the matrix S n Perform a contraction operation;
[0186] Step S3-3-4: During the solution process of the projection approximate point algorithm, the following convergence conditions are set:
[0187]
[0188] The convergence condition is and ε is the convergence threshold; through the first convergence condition, the low-rank matrix X is measured n and the reconstruction error of the original matrix M to determine the low-rank matrix X of the current iteration n Is it close enough to the observation matrix M? The second convergence condition is used to measure the difference between the results of two iterations, indicating the degree of convergence of the matrix update.
[0189] Step S3-4: According to the growing Lagrange multiplier algorithm, the filled complete multi-skeleton point temporal motion data is obtained.
[0190] According to the growing Lagrange multiplier algorithm, V n The iterative update formula is
[0191] When the convergence condition is met, or n exceeds the threshold value n of the number of operations max , the solution X n As the optimal solution, the complete multi-skeleton point temporal motion data X after filling is obtained.
[0192] Step S4: construct a boxing action recognition model based on LSTM and TCN network, and perform training;
[0193] The action recognition model includes a long short-term memory module LSTM, a multi-head attention mechanism module, and a time domain convolution module TCN. The long short-term memory module LSTM is used to maintain the hidden vector h and the memory vector m, respectively controlling the state update and input of each timestamp, and capturing the time dependency of multivariate time series.
[0194] The long short-term memory module LSTM includes a forget gate, an input gate, and an output gate. The forget gate determines how much of the cell state at the previous moment is retained to the current moment c t , the input gate determines the input x of the network at the current moment t How much is saved to the cell state c t , the output gate controls the control unit state c t How many outputs are there to the current output value h of the LSTM t ;
[0195] The input gate includes three inputs, which are the input value x of the network at the current moment t , the output value h of LSTM at the previous moment t-1 And the cell state c at the previous moment t-1 ;
[0196] The output gate includes two outputs, the output value h of LSTM at the current moment t And the cell state c at the previous moment t-1 ;
[0197] The multi-head attention mechanism module includes a linear transformation module, a self-attention calculation module, and a splicing module. The linear transformation module performs linear transformation to obtain the query, key, and value of multiple heads:
[0198] Q i =QW i q , K i =KW ik , V i =VW i v ;
[0199] Among them, Q i is the query matrix of the i-th head, K i is the key matrix of the ith head, V i is the value matrix of the i-th head; l is the sequence length, n is the feature dimension of TCN output; Q is the query, K is the key, and V is the value;
[0200] The self-attention calculation module performs self-attention calculation on each head:
[0201]
[0202] Among them, n h is the number of heads;
[0203] The concatenation concatenates the outputs of all heads and restores them back to the input sequence dimension through linear transformation:
[0204] MultiHead(Q,K,V)=Concat(head1,...,head h )W O ;
[0205] in, is a trainable linear transformation weight matrix;
[0206] The self-attention calculation module performs independent self-attention operations on the input query q, key k, and value v through multiple heads, concatenates the results of all heads and obtains the final output through linear transformation, and pays attention to different parts of the input sequence to different degrees through multiple heads;
[0207] The temporal convolution module is composed of multiple layers of causal convolution kernels and dilated convolutions, and the output of each layer is transmitted through residual connections;
[0208] The time domain convolutional network (TCN) uses causal convolution kernels to model time series data, and models long time series data through residual connections and extended convolutions to capture local temporal dependencies of time series.
[0209] For the input signal x t And the causal convolution kernel W, the output of the causal convolution kernel is:
[0210]
[0211] Among them, w i is the weight of the convolution kernel; x t-i is the time series data input to TCN;
[0212] The extended convolution increases the receptive field of the time domain convolutional network for time series data and captures longer-term dependencies. The output of the extended convolution is:
[0213]
[0214] Among them, r is the expansion factor. As the number of layers increases, the receptive field gradually expands.
[0215] The training of the action recognition model is carried out by constructing a loss function, specifically:
[0216] Step S4-1: construct a smooth loss function to ensure the smoothness of the motion sequence;
[0217] X is the matrix obtained after data refinement; repeat the boundary elements of X to obtain X′:X1′ ,i =X2′ ,i =X 1,i , X′ 3K+1,i =X′ 3K,i =X 3K,i , Among them, 1≤i≤T;
[0218] Define O to be a symmetric tridiagonal matrix that satisfies the following conditions:
[0219]
[0220] Among them, 2≤j≤t-1, Since the intervals between frames are equal, when 1≤j≤T, h j =1, then
[0221]
[0222] Step S4-2: Define a smooth loss function, which is:
[0223]
[0224] Where T is the number of frames in the sequence;
[0225] Step S4-3: According to the prior knowledge of the kinematic model, a bone length loss function is introduced. The bone length loss function is:
[0226] Set {l b |1≤b≤J-1} is the bone length sequence, where b is the bone index and the bone length loss is:
[0227]
[0228] in, and is the three-dimensional position of the joints at both ends of the skeleton in the i-th frame;
[0229] Step S4-4: Introduce cross entropy loss to evaluate and optimize the action recognition model;
[0230] The cross entropy loss function is used to calculate the difference between the probability distribution of each action category output by the classifier and the probability distribution corresponding to the true action label. The cross entropy loss is expressed as:
[0231] Loss c =-∑ i q i log(p i );
[0232] Among them, q i Represents the probability of the true label in the i-th category, p i Represents the probability of the classifier predicting the i-th category;
[0233] The total loss function of the action recognition model is:
[0234] Loss = λ1Loss s +λ2Loss b +λ3Loss c ;
[0235] Among them, λ1, λ2, and λ3 are the weight coefficients of smoothing loss, bone length loss, and cross entropy loss, respectively. The weight coefficients are hyperparameters in the training process and are fine-tuned according to the training situation and the action recognition performance of the test set.
[0236] Step S5: Input the processed time series data into the boxing action recognition model to implement the boxing action recognition task, specifically:
[0237] Step S5-1: Split the processed time series data into a training set and a test set in a ratio of 8:2, normalize the time series data, and input the training set into the action recognition model;
[0238] Step S5-2: When the time series length T is greater than the number of variables 3K, it is input into the long short-term memory module LSTM, and the long short-term memory module LSTM provides the data of the variable at all T timestamps;
[0239] The dataset dimension is R N×T×3K , where N is the number of samples in the data set, T is the time series length of each sample, and K is the number of bone points;
[0240] Step S5-3: Obtain the local features of the time series data through the one-dimensional convolution block of the time domain convolutional network TCN, and then adjust the mean and standard deviation of each batch of data through batch normalization to make the distribution of the time series data more stable;
[0241] Step S5-4: The action recognition model learns complex patterns and relationships in the data by introducing nonlinearity through the ReLU activation function; the local features of the output sequence are used through the multi-head attention mechanism to obtain the long-range dependencies between features and capture the global context;
[0242] Step S5-5: The features of the dual-branch outputs are aggregated through concat, and the input vector is converted into the probability distribution of each class through softmax to achieve the effect of sequence classification;
[0243] Step S5-6: Adopt Adam optimizer and cross entropy loss function to back propagate the whole network.
[0244] A motion recognition system for boxing scene data refinement and boxing motion recognition method generation, the motion recognition system includes a markerless motion capture module, a data missing and anomaly detection module, a multi-skeletal point missing data filling module, a multi-skeletal point data repair module and a boxing motion recognition module, the markerless motion capture module is used for collecting multi-skeletal point time series data of boxers in actual boxing training and competition scenes, and obtaining multi-skeletal point time series motion data of boxers; the data missing and anomaly detection module is used for detecting missing or anomaly of output time series data, and annotating type labels for data by detecting whether there are missing or anomalies; the multi-skeletal point missing data filling module processes incomplete bone time series data, and reconstructs approximate real human bone motion data through fuzzy C-means clustering algorithm and projection approximate point algorithm;
[0245] The multi-skeleton point data repair module realizes the anomaly elimination and smoothing of the multi-skeleton point temporal motion data. The boxing action recognition module uses the boxing action recognition model to realize the boxing action recognition task.
[0246] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A method for data refinement and boxing action recognition for boxing scenes, characterized by: The method comprises the following steps: Step S1: Collecting time series data of multiple skeleton points of a boxer; Step S2: perform anomaly and missing detection on time series data; Step S3: Processing the detected abnormal and missing time series data; Step S4: construct a boxing action recognition model based on LSTM and TCN network, and perform training; Step S5: input the processed time series data into the boxing action recognition model to realize the boxing action recognition task.
2. The method for data refinement and boxing action recognition for boxing scenes according to claim 1, characterized in that: In step S1, the time series data of multiple skeleton points of the boxer are collected, and the specific operations are as follows: Step S1-1: four cameras are arranged around the boxing ring, and real-time video streams are acquired by the cameras; Step S1-2: Use the OpenPose deep learning algorithm to obtain the corresponding two-dimensional coordinate sequence of multiple skeleton points in the video stream; Step S1-3: Using the pinhole imaging principle, the projection of the two-dimensional coordinates on the imaging plane to the three-dimensional coordinates in space is realized to obtain the three-dimensional coordinates of multiple skeleton points; Step S1-4: Convert the three-dimensional coordinates of multiple skeleton points into time series data.
3. The method for data refinement and boxing action recognition for boxing scenes according to claim 1, characterized in that: In step S2, the time series data is detected for anomalies and missing data, specifically: Step S2-1: Represent the multi-skeleton point time series data in matrix form, as follows: F=[f1,f2,…,f T ]; Among them, f i =[x i,1 ,y i,1 ,z i,1 ,...,x i,K ,y i,K ,z i,K ] T ,i=1,...,T, t is the total number of frames in the time series, K is the number of skeleton points; Step S2-2: define anomalies and missingness of time series data; Missing data means that in a multivariate time series, the data of certain timestamps are not recorded or lost, which is manifested as null values for some timestamps in the matrix F; Data anomaly refers to the situation in which the data of some variables at certain timestamps in a multivariate time series deviate significantly from the normal pattern, which is manifested in the oscillation of some data in the matrix F and the outliers of data at a certain time length; Step S2-3: label samples of abnormal and missing time series data; If there are missing or abnormal data, the multivariate time series samples are labeled as [1,1]; If there is no missing or abnormal data, the multivariate time series samples are labeled as [0,0].
4. The method for data refinement and boxing action recognition for boxing scenes according to claim 1, characterized in that: In step S3, the detected abnormal and missing time series data are processed, specifically: Step S3-1: For long time sequences of different action combinations, the fuzzy C-means clustering method is used to divide the long sequence, so that the same or similar motion sequences can be classified into the same segment: For the matrix F = [f1,f2,…,f T ], according to the membership of each data point to each class w slices f i (i=1,...,w) is divided into c fuzzy groups, and the membership matrix S ranges from [0,1]. After normalization, the sum of the membership of a data set is equal to 1, that is: Among them, s ij Represents the membership of slice j to cluster i. The closer the value is to 1, the more the slice belongs to cluster i. The cost function of fuzzy C-means is: Among them, 0 ij <1, c i is the cluster center of fuzzy group i, d ij =||c i -f j || is the Euclidean distance between the i-th cluster center and the j-th slice, m is the fuzzy index, m>1; Construct a new objective function to ensure that the value function can be optimized along with the constraints. The constructed new objective function is shown as follows: The necessary condition for minimizing the value function by taking the derivative of the input is: Among them, x j is the time series data; the membership matrix S is initialized by taking values in the range [0,1] and satisfies Condition, by calculating the cluster center c i And calculate the value function, iteratively update S, and finally get the cluster center c when the value function is minimum i and the membership matrix S; Step S3-2: reconstruct all elements in the recovery matrix according to the low rank of the data matrix, and transform the filling problem into a convex optimization problem; The convex optimization problem is: Among them, X is the restored matrix, M is the missing data matrix, Ω is the index set of known elements, and P Ω is the linear projection operator; Step S3-3: solving the convex optimization problem by projected approximate point algorithm; Step S3-4: According to the growing Lagrangian multiplier algorithm, the filled complete multi-skeleton point temporal motion data is obtained.
5. The method for data refinement and boxing action recognition for boxing scenes according to claim 4, characterized in that: In step S3-3, the convex optimization problem is solved by using the projected approximate point algorithm, specifically: For each sub-segment the missing data matrix where t w is the number of frames in the wth sequence after cluster segmentation, d is the data dimension, d = 3K; Step S3-3-1: Set initialization parameters: X0=0, V0=0, n=1, Step S3-3-2: Calculate the projection matrix Y of the missing data matrix M n , the iteration formula is Step S3-3-3: Project the matrix Y n Perform singular value decomposition and solve the approximate point matrix X n ,Right now: [U n ,S n ,V n ]=SVD(Y n ); Among them, σ λn-1 is the contraction operator, and its operation rules are: pass For the matrix S n Perform a contraction operation; Step S3-3-4: During the solution process of the projection approximate point algorithm, the following convergence conditions are set: The convergence condition is and ε is the convergence threshold; through the first convergence condition, the low-rank matrix X is measured n and the reconstruction error of the original matrix M, to determine the low-rank matrix X of the current iteration n Is it close enough to the observation matrix M? The second convergence condition is used to measure the difference between the two iteration results, indicating the degree of convergence of the matrix update. In step S3-4, the filled complete multi-skeleton point temporal motion data is obtained according to the growing Lagrange multiplier algorithm, specifically: According to the growing Lagrange multiplier algorithm, V n The iterative update formula is When the convergence condition is met, or n exceeds the threshold value n of the number of operations max , the solution X n As the optimal solution, the complete multi-skeleton point temporal motion data X after filling is obtained.
6. The method for data refinement and boxing action recognition for boxing scenes according to claim 1, characterized in that: The action recognition model includes a long short-term memory module LSTM, a multi-head attention mechanism module and a time domain convolution module TCN. The long short-term memory module LSTM is used to maintain the hidden vector h and the memory vector m, respectively control the state update and input of each timestamp, and capture the time dependency of the multivariate time series; The long short-term memory module LSTM includes a forget gate, an input gate, and an output gate. The forget gate determines how much of the cell state at the previous moment is retained to the current moment c. t , the input gate determines the input x of the network at the current moment t How much is saved to the cell state c t , the output gate controls the control unit state c t How many outputs are there to the current output value h of the LSTM t ; The input gate includes three inputs, which are the input value x of the network at the current moment. t , the output value h of LSTM at the previous moment t-1 And the cell state c at the previous moment t-1 ; The output gate includes two outputs, the output value h of the LSTM at the current moment. t And the cell state c at the previous moment t-1 ; The multi-head attention mechanism module includes a linear transformation module, a self-attention calculation module and a splicing module. The linear transformation module performs linear transformation to obtain the query, key and value of multiple heads: Q i =QW i q ,K i =KW i k ,V i =VW i v ; Among them, Q i is the query matrix of the i-th head, K i is the key matrix of the ith head, V i is the value matrix of the i-th head; l is the sequence length, n is the feature dimension of TCN output; Q is the query, K is the key, and V is the value; The self-attention calculation module performs self-attention calculation on each head: Among them, n h is the number of heads; The concatenation concatenates the outputs of all heads and restores them back to the input sequence dimension through linear transformation: MultiHead(Q,K,V)=Concat(head1,...,head h )W O ; in, is a trainable linear transformation weight matrix; The self-attention calculation module performs independent self-attention operations on the input query q, key k, and value v through multiple heads, concatenates the results of all heads and obtains the final output through linear transformation, and pays attention to different parts of the input sequence to different degrees through multiple heads; The temporal convolution module is composed of multiple layers of causal convolution kernels and dilated convolutions, and the output of each layer is transmitted through residual connections; The time domain convolutional network (TCN) uses causal convolution kernels to model time series data, and models long time series data through residual connections and extended convolutions to capture local temporal dependencies of time series. For the input signal x t And the causal convolution kernel W, the output of the causal convolution kernel is: Among them, w i is the weight of the convolution kernel; x t-i is the time series data input to TCN; The extended convolution increases the receptive field of the time domain convolutional network for time series data and captures longer-term dependencies. The output of the extended convolution is: Among them, r is the expansion factor. As the number of layers increases, the receptive field gradually expands.
7. The method for data refinement and boxing action recognition for boxing scenes according to claim 6, characterized in that: The training of the action recognition model is performed by constructing a loss function, specifically: Step S4-1: construct a smooth loss function to ensure the smoothness of the motion sequence; X is the matrix obtained after data refinement; repeat the boundary elements of X to obtain X′:X1′ ,i =X2′ ,i =X 1,i , X3′ K+1,i =X3′ K,i =X 3K,i , Among them, 1≤i≤T; Define O to be a symmetric tridiagonal matrix that satisfies the following conditions: Among them, 2≤j≤t-1, Since the intervals between frames are equal, when 1≤j≤T, h j =1, then Step S4-2: Define a smooth loss function, which is: Where T is the number of frames in the sequence; Step S4-3: According to the prior knowledge of the kinematic model, a bone length loss function is introduced. The bone length loss function is: Set {l b |1≤b≤J-1} is the bone length sequence, where b is the bone index and the bone length loss is: in, and is the three-dimensional position of the joints at both ends of the skeleton in the i-th frame; Step S4-4: Introduce cross entropy loss to evaluate and optimize the action recognition model; The cross entropy loss function is used to calculate the difference between the probability distribution of each action category output by the classifier and the probability distribution corresponding to the true action label. The cross entropy loss is expressed as: Loss c =-∑ i q i log(p i ); Among them, q i Represents the probability of the true label in the i-th category, p i Represents the probability that the classifier predicts the i-th category; The total loss function of the action recognition model is: Loss=λ1Loss s +λ2Loss b +λ3Loss c ; Among them, λ1, λ2, and λ3 are the weight coefficients of smoothing loss, bone length loss, and cross entropy loss, respectively. The weight coefficients are hyperparameters in the training process and are fine-tuned according to the training situation and the action recognition performance of the test set.
8. The method for data refinement and boxing action recognition for boxing scenes according to claim 7, characterized in that: In step S5, the processed time series data is input into the boxing action recognition model to realize the boxing action recognition task, specifically: Step S5-1: Split the processed time series data into a training set and a test set in a ratio of 8:2, normalize the time series data, and input the training set into the action recognition model; Step S5-2: When the time series length T is greater than the number of variables 3K, it is input into the long short-term memory module LSTM, and the long short-term memory module LSTM provides the data of the variable at all T timestamps; The dataset dimension is R N×T×3K , where N is the number of samples in the data set, T is the time series length of each sample, and K is the number of bone points; Step S5-3: Obtain the local features of the time series data through the one-dimensional convolution block of the time domain convolutional network TCN, and then adjust the mean and standard deviation of each batch of data through batch normalization to make the distribution of the time series data more stable; Step S5-4: The action recognition model learns complex patterns and relationships in the data by introducing nonlinearity through the ReLU activation function; the local features of the output sequence are used through the multi-head attention mechanism to obtain the long-range dependencies between features and capture the global context; Step S5-5: The features of the dual-branch outputs are aggregated through concat, and the input vector is converted into the probability distribution of each class through softmax to achieve the effect of sequence classification; Step S5-6: Adopt Adam optimizer and cross entropy loss function to back propagate the whole network.
9. The action recognition system according to any one of claims 1 to 8, characterized in that: The motion recognition system comprises an unmarked motion capture module, a data missing and anomaly detection module, a multi-skeletal point missing data filling module, a multi-skeletal point data repair module and a boxing motion recognition module. The unmarked motion capture module is used for collecting the time series data of multiple skeleton points of a boxer in actual boxing training and competition scenes to obtain the time series motion data of multiple skeleton points of the boxer; the data missing and anomaly detection module is used for detecting missing or anomaly of the output time series data, and marking the data with type labels by detecting whether there is missing or anomaly; the multi-skeletal point missing data filling module processes incomplete skeleton time series data, and reconstructs the skeleton motion data close to the real human body through the fuzzy C-means clustering algorithm and the projection approximate point algorithm; The multi-skeleton point data repair module realizes the abnormal elimination and smoothing processing of the multi-skeleton point time series motion data, and the boxing action recognition module uses the boxing action recognition model to realize the boxing action recognition task.
Citation Information
Patent Citations
Motion data denoising method and system
CN110232672A
Bone movement data enhancement method and system based on Kinect
CN111507920A
Multi-person three-dimensional motion capture method, storage medium and electronic equipment
CN112379773A
Multi-modal dynamic gesture recognition method based on lightweight 3D residual network and TCN
CN112507898A
Video action segmentation by mixed temporal domain adaption
US20210174093A1
Cited By
Complex posture behavior analysis method based on machine learning
CN121167252A