Classified sensing end-to-end action recognition method and system

By using traceless Kalman filtering and multivariate time series classifiers in action recognition, combined with multi-scale feature extraction network and masking mechanism, the problems of high computational costs and complex feature extraction in action recognition are solved, and efficient and accurate action recognition and the effect of reducing computational costs are achieved.

CN119964251APending Publication Date: 2025-05-09NANJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445527.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art relies on visual data in action recognition, which has problems such as high computational cost, complex feature extraction, and susceptible to environmental influences, and the strong time dependence of multivariate time series data is difficult to effectively handle.

Method used

Coarse-grained filtering based on traceless Kalman filtering and a robust multivariate time series classifier are used to extract multi-scale feature extraction network and masking mechanism to achieve action recognition.

Benefits of technology

Effectively suppress noise and outliers, improve the accuracy of action recognition, reduce calculation costs, facilitate deployment, and improve the robustness of the model and multi-scale feature extraction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964251A_ABST
    Figure CN119964251A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of time sequence data mining and action recognition, discloses a classification perception end-to-end action recognition method and system, and aims to design a classification perception action recognition method and system by mining a time sequence mode of obtained time sequence data and adopting a multivariable time sequence classification thought. An accurate classification result can be recognized only by inputting original noisy data, meanwhile, the calculation cost is reduced, and deployment is facilitated. According to the invention, noise suppression of a data level is realized through coarse-grained filtering; extracting a network mining time sequence mode through a mask mechanism and multi-scale features; the mask mechanism captures attention scores of filtering data by using a Transform encoder, perceives and classifies important timestamps, and masks the important timestamps; in combination with a time slicing mechanism of a multi-scale feature extraction network, input of different scales and a destruction-recovery mechanism are constructed to provide robustness for the feature extraction network, the multi-scale feature extraction network captures multi-scale time features, and then an action label is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series data mining and action recognition, and specifically relates to a classification-aware end-to-end action recognition method and system. Background Art

[0002] At present, the field of action recognition mainly relies on the analysis of visual data such as videos and images, which are usually used to train computer vision models to achieve the recognition and classification of human actions. However, this type of method has significant limitations. First, the computational cost of models based on visual data is high, especially when processing high-resolution videos, which requires powerful hardware support and a large amount of computing resources, making the deployment of the system complicated and expensive. Second, due to the high dimensionality of video data and the complex feature extraction process, the model training time is long and the efficiency is low. In addition, visual data is easily affected by environmental factors such as lighting changes and occlusion, which further increases the difficulty of model design.

[0003] In contrast, wearing sensors (such as accelerometers, gyroscopes, etc.) to capture the multivariate time series of daily human movements, and then completing the action recognition task through a multivariate time series classifier, provides a more efficient and economical solution. Sensors can accurately record parameters such as coordinates and speed during movement to generate multivariate time series data. For athlete movements in the sports field, more advanced technologies such as OpenPose, motion capture systems or other time series data extraction methods can also be used to obtain more accurate motion data.

[0004] However, using multivariate time series classifiers to handle action recognition tasks faces a series of challenges: first, the features captured by the classifier are not directly related to the downstream tasks of action recognition, which will compromise the classifier's performance; second, in the field of action recognition, multivariate time series data has significant strong time dependence, and the relevant information of athletes' actions is usually reflected in continuous time intervals. Therefore, it is crucial to mine its multi-scale time features. Summary of the invention

[0005] To solve the above technical problems, the present invention provides a classification-aware end-to-end action recognition method and system. At the data level, a coarse-grained filter based on unscented Kalman filtering is designed to suppress jitter noise and outliers in time series data; at the model level, a robust multivariate time series classifier is designed to extract multi-scale temporal features for multivariate time series to support action recognition tasks.

[0006] The classification-aware end-to-end action recognition method of the present invention comprises the following steps: S1, collect time series data of human body movements and perform data preprocessing; S2, performing coarse-grained filtering on the preprocessed time series data based on an unscented Kalman filter to obtain filtered time series data; S3, using a masking mechanism to capture the classification key timestamps and mask the filtered time series data; S4. Utilize a multi-scale feature extraction network, build inputs of different scales based on the time slicing mechanism, and perform position encoding to extract multi-scale robust temporal features from multivariate time series to complete action recognition.

[0007] Furthermore, in S1, data preprocessing is specifically as follows: S11. Design an improved linear interpolation strategy to complete the missing value filling of time series data; S12, z-score standardization is performed on multivariate time series, and the value range of all multivariate time series is mapped to the interval [0,1]; S13. Align each multivariate time series to the same length.

[0008] Furthermore, in S11, the improved linear interpolation strategy is specifically as follows: S111. For missing values ​​at the beginning of the sequence, fill them in by traversing subsequent data and using the first valid value; S112. For missing data in the middle of the sequence, linear interpolation method is used to fill in the missing data by combining the previous and next valid data points; S113. When there are missing values ​​at the end of the sequence, linear regression prediction is used to fill in the missing values ​​according to the changing trend of the valid data before the end, so that the action sequence is continuous and complete.

[0009] Furthermore, S2 is specifically: S21, setting the initial state value, covariance matrix, state noise and observation noise matrix of each multivariate time series at the first moment; S22, according to the currently estimated state mean and covariance matrix, generate a set of Sigma points to represent the state space distribution and perform state prediction; S23, mapping the Sigma point to the observation space through coarse-grained filtering to obtain the observation value of the state space distribution; S24, calculating the weighted average of all Sigma points to obtain the updated state prediction value and covariance matrix; S25. Based on the difference between the observed value and the predicted value, the state estimate and covariance are updated through the Kalman gain to correct the state estimate.

[0010] Furthermore, in S3, the masking mechanism includes a key timestamp perception network and a masking module, the input of which is the filtered time series data output by S2, and the attention score vector consistent with the time length dimension is calculated. , the mask module relies on the attention score vector Masking of key timestamps for a fixed ratio; The backbone network of the key timestamp perception network is a Transformer encoder network, including a time domain projection module and a multi-head attention mechanism; the filtered time series data is mapped into query q, key k, and value v matrices, and the variable dimensions are converted into dimensions that are adapted to the model; the time domain projection module reduces the time dimension of the key k and value v matrices to R times the original dimension; the multi-head attention mechanism receives the query q matrix and the projected k and v matrices to capture the attention score consistent with the time dimension , input Attn into the mask module; l is the length of the time series; The mask module sums the first dimension of the received Attn matrix to obtain , combined with sampling randomization, select the attention score vector Before Timestamps from Randomly select from timestamps timestamp and mask, where .

[0011] Furthermore, the multi-scale feature extraction network is a three-layer architecture, where the output of the previous layer is the input of the next layer, and each layer includes a time slicing mechanism and a Transformer encoder network; In the first layer of the multi-scale feature extraction network, the input of the time slicing mechanism is mask data. The time slicing mechanism constructs inputs of different scales through its built-in time domain convolutional network, so that the mask timestamp compensates its own information according to the context in the slice; The input of the time slicing mechanism of the second and third layers of the multi-scale feature extraction network is the output of the previous layer of Transformer encoder network; The Transformer encoder network includes a time domain projection module, a multi-head attention mechanism, and a feedforward neural network FFN; for the ath layer, 1≤a≤3, the input of the Transformer encoder network is the shallow encoding of the slice output by the time slicing mechanism, and the shallow encoding is mapped to the query q, key k, and value v matrix; the time domain projection module reduces the time dimension of the key k and value v matrix to the original dimension times; the multi-head attention mechanism receives the query q matrix and the projected k and v matrices to capture the long-range dependencies of the slices; the multi-head attention mechanism and FFN respectively use residual connections to prevent the gradient from disappearing, the output of the multi-head attention mechanism and the shallow encoding of the slice output by the time slicing mechanism are fused and input to FFN, and the output of FFN and the feature fusion of the input of FFN are used as the output of the Transformer encoder network; the features of the FFN output of the last layer are passed through the average pooling layer to obtain the probability distribution of the action category; the cross entropy loss is used as the loss function for training the Transformer encoder network.

[0012] Furthermore, S4 is specifically: S41, constructing the shallow encoding of the current layer based on the time slicing mechanism to complete the time slicing; S42, a one-dimensional zero-filled convolution operation is performed on the slice shallow code to give the slice shallow code a context-aware position attribute, i.e., position coding, and the slice shallow code is fused with the slice shallow code as the input of the Transformer encoder; S43, mapping the shallow encoding of the current layer to the query q, key k, value v matrix of the Transformer encoder; S44. Design a time domain projection mechanism to compress the time dimension of key k and value v to save computational cost; S45, the multi-head attention mechanism receives the query q matrix and the projected key k and value v matrices to capture the long-range dependencies of the slices; S46, concatenate the attention scores calculated by each head and reshape them to the dimension of the input sequence, and send the reshaped attention features to the feed-forward neural network FFN; S47. The output of the feedforward neural network of the last layer of Transformer encoder is passed through the average pooling layer and softmax to obtain the probability distribution of each action category.

[0013] The present invention also provides a classification-aware end-to-end action recognition system for implementing the above method, the system comprising: Data preprocessing module: obtain time series data and perform data preprocessing; Coarse-grained filtering module: Based on unscented Kalman filtering, it realizes smoothing of nonlinear motion time series data through unscented transformation and prediction-update mechanism; Masking mechanism: Capture and mask the key timestamps of classification for filtered time series data, including masking a fixed ratio of key timestamps based on the attention scores calculated by the Transformer encoder network. Multi-scale feature extraction network: Relying on the time slicing mechanism to construct inputs of different scales, it extracts multi-scale robust temporal features from multivariate time series and performs position encoding to complete action recognition.

[0014] The beneficial effects described in the present invention are as follows: the present invention mines time series patterns from skeleton point time series data obtained based on sensor data, motion capture and other methods, adopts the concept of multivariate time series classification, and designs classification-aware action recognition methods and systems; only the original noisy data needs to be input to identify accurate classification results, and the corresponding action category labels are output, which effectively improves the accuracy of action recognition, and greatly reduces the computational cost, making it easy to deploy. The present invention achieves noise suppression at the data level through coarse-grained filtering; mines time series patterns through masking mechanisms and multi-scale feature extraction networks; the masking mechanism uses the attention scores captured by the Transformer encoder on the filtered data, perceives the key timestamps of classification and masks them, and combines the time slicing mechanism of the multi-scale feature extraction network to construct inputs of different scales and the destruction-recovery mechanism, restores the mask timestamp information, establishes the correlation between features and downstream tasks, and provides robustness for the feature extraction network. The multi-scale feature extraction network captures multi-scale time features and then outputs action labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A flow chart of a classification-aware end-to-end action recognition method provided by an embodiment of the present invention; Figure 2 A classification-aware end-to-end action recognition model architecture diagram provided by an embodiment of the present invention; Figure 3 An architecture diagram of a classification-aware end-to-end action recognition system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments in conjunction with the accompanying drawings.

[0017] In this embodiment, the classification perception end-to-end action recognition model architecture is constructed as follows Figure 2 The classification-aware end-to-end action recognition method described in the present invention is as follows: Figure 1 As shown, the following steps are included: S1, collect time series data of human body movements and perform data preprocessing; S2, performing coarse-grained filtering on the preprocessed time series data based on an unscented Kalman filter to obtain filtered time series data; S3, using a masking mechanism to capture the classification key timestamps and mask the filtered time series data; S4. Utilize a multi-scale feature extraction network, build inputs of different scales based on the time slicing mechanism, and perform position encoding to extract multi-scale robust temporal features from multivariate time series to complete action recognition.

[0018] In S1, time series data is obtained based on sensor data, motion capture and other methods, and data preprocessing is performed. The preprocessing process includes filling missing values, performing z-score standardization, and aligning time series lengths.

[0019] Specifically, the time series data refers to multivariate time series data formed by the changes in the x, y, and z coordinate values ​​of important skeletal points of the human body over time. These data can be located by sensors at specific skeletal points or extracted using motion capture technology.

[0020] In the actual collected time series data, missing values ​​are inevitable, which are caused by sensor failure, data collection problems, or communication interruptions. The time series data collected by the motion capture system also usually contains missing values. The missing values ​​are generated because the skeleton points of the athletes in the video are blocked and the light problems lead to the failure of skeleton point recognition. If the time period corresponding to the missing value exceeds one-third of the total sample length, consider removing the sample, because the filled data is not the real data collected and cannot reflect the real movement.

[0021] Furthermore, the present invention adopts an improved linear filling strategy to interpolate missing values: first, for missing values ​​at the beginning of the sequence, they are filled by traversing the subsequent data and using the first valid value; second, for missing data in the middle of the sequence, a linear interpolation method is used to fill in the missing values ​​by combining the previous and next valid data points; finally, when there are missing values ​​at the end of the sequence, linear regression prediction is used to fill in the missing values ​​according to the changing trend of the valid data before the end, so as to ensure the continuity and integrity of the action sequence.

[0022] In order to eliminate the dimensional differences between different variables, the data is mapped to the [0,1] interval through the z-score normalization method. In the present invention, the coordinate value of a certain dimension of a certain bone point is defined as a variable, for example, the x coordinate of the wrist joint is a variable. Therefore, the time series data is a multivariate time series from the perspective of time series classification. The following formula is used to process multivariate time series: , Among them, x is the time series value of a variable in the multivariate time series, and z is the standardized result of the time series value of the variable. is the mean, is the standard deviation.

[0023] Since the time series length of each action sample is different, the time series length alignment mechanism is used to align the length of each multivariate time series. First, a fixed time series length value is set. For sequences with a length less than the set value, zeros are filled in the front and back of the sequence to extend it to the set length. For sequences with a length greater than the set value, the middle frame of the action is calibrated and the front and back continuous frames are selected to ensure that the final sequence length is consistent with the set value.

[0024] After the time series lengths are aligned, z is concatenated in variable dimensions: , Among them, l is the time series length of the aligned multivariate time series, m is the number of variables, is the multivariate time series obtained after splicing, which serves as the pre-filtered data of S2.

[0025] In S2, the preprocessed time series data is coarse-grained filtered based on the unscented Kalman filter, and the smoothing of the nonlinear motion time series data is achieved through the unscented transformation and prediction-update mechanism.

[0026] Specifically, the multivariate time series collected by the sensor data and the motion capture system usually have noise. For example, the sensor may be affected by environmental factors and mechanical vibrations, resulting in fluctuations in the collected time series data. In addition, the motion capture system may generate noise due to calibration errors, occlusions, or rapid changes in motion trajectories. Therefore, the pre-processed time series data is smoothed by coarse-grained filtering. In this embodiment, the coarse-grained filtering uses an unscented Kalman filter that adapts to nonlinear data to smooth the time series data.

[0027] Furthermore, the state estimate is initialized to measure the initial state , state estimation Initialize the covariance matrix to describe the uncertainty of the initial state estimate. ; Initialize the process noise covariance matrix Q to describe the statistical characteristics of the noise introduced during the state transfer process; initialize the observation noise covariance matrix R to reflect the statistical characteristics of the noise during the observation process.

[0028] Furthermore, for the kth moment, an m-dimensional state vector , it is necessary to calculate 2m+1 sigma points, which can represent the probability distribution of the state vector to a certain extent; the specific calculation method is as follows: Central sigma point: ,in is the state estimate obtained at the k-1th moment, The subscript indicates the k-1th moment, and the superscript indicates the central sigma point; Positive offset sigma point: ,in Representation Matrix The positive offset sigma point is obtained by adding an offset related to the covariance matrix to the current state estimate to cover the possible range of values ​​in the state space; the negative offset sigma point: ; The negative offset sigma point subtracts the corresponding offset from the current state estimate, and together with the positive offset sigma point describes the distribution of the state variable; in, is the scaling factor, ; Determines the distribution range of the sigma point, usually taking a smaller value. is an auxiliary parameter; According to the state transfer function , predict the state and covariance at the next moment.

[0029] Furthermore, the calculated sigma point is passed through the state equation , get the predicted sigma point: ,in, is the control input at time k, is the process noise.

[0030] Furthermore, based on the predicted sigma points, the mean of the predicted state is calculated by weighted summation. : , in, It is the weighting coefficient used to calculate the mean, and different sigma points have different weights; is the predicted sigma point.

[0031] Furthermore, the covariance matrix of the predicted state is calculated: , in, is the weighting coefficient used to calculate the covariance; is the process noise covariance matrix, which reflects the uncertainty in the system state transfer process.

[0032] Weighting coefficient and The value of is: , , , Among them, β is used to incorporate prior knowledge of state variables.

[0033] The observation update state uses the actual observation value to correct the predicted state and improve the accuracy of the state estimation.

[0034] Furthermore, the predicted sigma point is passed through the observation equation , get the observed sigma point ;in, represents the state transition function, is the actual observed value at time k; is the observation noise; is the mean of the predicted states.

[0035] Furthermore, the mean of the observations is calculated by weighted summation based on the sigma points of the observations: ; Furthermore, the covariance matrix of the observations is calculated : , in, is the observation noise covariance matrix; Furthermore, the cross-covariance matrix between states and observations is calculated: ; Furthermore, the Kalman gain is calculated based on the cross-covariance matrix and the observation covariance matrix: : , in, is the cross-covariance matrix between the state and observation at the kth moment. The Kalman gain determines the degree to which the observation corrects the state estimate.

[0036] Furthermore, the Kalman gain and the difference between the observed value and the predicted value are used to update the state estimate: ;

[0037] Furthermore, the covariance matrix of the state estimate is updated: .

[0038] The unscented Kalman filter selects sigma points through unscented transformation to approximate the probability distribution of state variables, avoiding the complex process of linearizing nonlinear functions, so that the mean and covariance of state variables can be estimated more accurately. By continuously updating time and observations, the unscented Kalman filter can achieve real-time tracking and prediction of system states.

[0039] In S3, given that multivariate time series data usually has information sparsity characteristics, and there are significant differences in the contribution of different timestamps to classification tasks, the present invention designs a classification-aware masking mechanism, including a key timestamp-aware network and a masking module, which masks the key timestamps and extracts the multi-scale features of the masked sequence to achieve robust feature extraction. The backbone network of the key timestamp-aware network is a Transformer encoder network. The present invention determines the important timestamps in the process of learning multivariate time series based on the attention weights output by the Transformer encoder, and shields them in the masking stage.

[0040] Suppose someone is doing an action, and obtain the action sample data through coarse-grained filtering , can be regarded as a multivariate time series sample, containing l length and m variables; for the tth timestamp, ; Pass X through the multi-head self-attention layer to obtain The attention feature weight of .

[0041] Furthermore, the first dimension is accumulated to obtain the attention score vector consistent with the time length l ; Timestamps corresponding to positions with higher weights have higher importance for classification. Arrange from high to low, the higher the score, the more important the corresponding timestamp is for classification, and get the index corresponding to the timestamp with a higher score.

[0042] Furthermore, set the mask ratio , take the attention score vector Before The index corresponding to the value is masked, where .

[0043] However, during the feature extraction network training process, the key timestamp for a given sample does not change with the number of training times, which makes the network only focus on certain specific timestamps and learn the patterns of specific timestamps in the training data, resulting in overfitting; and each training time only masks the same set of timestamps, which is also not conducive to improving robustness and extraction robustness.

[0044] Furthermore, setting the ratio , , select the attention score vector Before timestamp, and ,from Randomly select from timestamps timestamp and mask.

[0045] This masking mechanism combines attention scores with sampling randomization, so that the masking mechanism does not always mask the same set of timestamps. The time slicing mechanism can focus on different timestamps at each training epoch and learn a wider range of patterns, and has a certain perception ability for classification tasks.

[0046] In S4, the multi-scale feature extraction network is a layered architecture with three layers. The output of the previous layer is the input of the next layer. Each layer contains a time slicing mechanism and a Transformer encoder network. Time slices are used to construct time scale inputs of different layers, and the dependencies between slices are captured through the Transformer encoder. The output of the previous layer of the multi-scale feature extraction mechanism is the input of the next layer to mine time features at different scales.

[0047] Specifically, a time slicing mechanism is designed to perform time slicing on the masked multivariate time series. Since the feature extraction network is multi-scale, the window slice length of each layer is different. Suppose the slice length of the ath layer is The time slicing mechanism consists of a trainable time-domain convolutional network. In the ath layer, each The consecutive points are grouped and fed into a trainable temporal convolutional network, which consists of 1D convolutional modules to transform the feature dimension into .

[0048] Furthermore, for the first layer, the input data is The data is passed through the convolutional layer to obtain the shallow encoding of the slice , to achieve time slicing; for the second and third layers, the input data is the output of the Transformer encoder network of the previous layer, with a dimension of , through the time slicing mechanism, the shallow encoding of the current layer is constructed to complete the time slicing. Considering that in a slice, the mask timestamp can perceive the context information in the slice, so that the information of the mask timestamp can still be restored or predicted based on the context information of the previous and next timestamps when it is damaged, and then the time slice of information aggregation is obtained. This destruction-recovery mechanism provides model robustness.

[0049] Considering that the Transformer architecture captures time invariance that is crucial to multivariate time series analysis, that is, the ability of the model to recognize multivariate time series patterns, while requiring learnable location information for each timestamp; therefore, a position encoding method is designed to ensure that context awareness and pattern recognition capabilities can be obtained during feature extraction.

[0050] Specifically, at the output of the time slice of the multi-scale feature extraction network, a one-dimensional zero-filled convolution operation is used as the context-aware position encoding for the shallow encoding of the slice, which only requires a convolution kernel size of , the contextual information of adjacent local sequences can be captured, and the absolute position information can be provided through zero padding operation.

[0051] Furthermore, the shallow encoding features are added to the position encoding, and the shallow encoding features and the position encoding have the same dimension, giving the slice shallow encoding a context-aware position attribute and using it as the input of the Transformer encoder.

[0052] Specifically, a Transformer encoder network is designed, which has the same structure as the Transformer encoder network in the mask mechanism. The input of the Transformer encoder network is the slice shallow encoding output by the time slicing mechanism; the Transformer encoder network captures the global context of inputs of different scales through an attention mechanism.

[0053] Specifically, for layer a: , in, is the output of the time slicing mechanism at layer a, which is the shallow encoding of the slice.

[0054] The calculation rules that define the time domain projection mechanism are: , in, represents the input sequence, ,pass Reshape the input sequence into ; h is the number of attention layer heads, ,Will The dimension is divided into h heads; As the projection layer matrix, the input sequence length is reduced to ,in To reduce the scale factor of the computational cost; the Norm layer uses layer normalization operation.

[0055] Considering that the multi-head attention mechanism uses inner product operation at each time point to calculate the attention scores of all other time stamps in the sequence, the computational complexity increases quadratically with the length of the sequence and occupies a large amount of memory, a method for reducing the computational cost of the time dimension is designed, relying on the linear attention idea to improve the multi-head attention. Traditionally, the query q, key k, and value v are inputs obtained through three different linear transformations and are of the same dimension. The present invention designs a time domain projection mechanism to compress the dimensions of the key k and value v matrices and reduce the computational cost.

[0056] Furthermore, the time domain projection mechanism is calculated as: , , While q is not compressed by the time domain projection mechanism, q is obtained through traditional linear transformation: ; Furthermore, the attention calculation mechanism is: .

[0057] The present invention applies the time domain projection mechanism to the above formula to obtain the attention score of each head: , in, , , is a trainable weight matrix, j is the jth head of the multi-head attention mechanism, 1≤j≤h.

[0058] The attention scores of each head are concatenated and restored to the dimension of the input sequence; in this way, the feature extraction method can process longer input sequences without requiring more resources, balancing classification accuracy and computational overhead.

[0059] The reshaped attention features are sent to the feed-forward neural network FFN. To avoid the gradient vanishing problem, the multi-head attention layer and the FFN layer are respectively connected with residual connections to improve the convergence ability of the model.

[0060] The probability distribution of each action category is obtained through the average pooling layer and softmax.

[0061] Furthermore, the cross entropy loss is used as the loss function, which is used to measure the difference between the predicted probability distribution of each action category output and the probability distribution of the true label. The cross entropy loss function is expressed as: , in, represents the probability of the true label in the cth category, represents the probability of the classifier predicting the cth class. By minimizing the cross entropy loss, the goal of the feature extraction network is to make its predicted probability distribution as close as possible to the true label distribution. During the training process, the parameters are updated through the back propagation algorithm and the gradient descent optimizer to continuously reduce the cross entropy loss, thereby improving the prediction accuracy of action recognition.

[0062] like Figure 3As shown, another aspect of an embodiment of the present invention provides a classification-aware end-to-end action recognition system, including a data preprocessing module, a coarse-grained filtering module, a mask mechanism, and a multi-scale feature extraction network.

[0063] The data preprocessing module performs data preprocessing based on time series data obtained by methods such as sensor data and motion capture.

[0064] Coarse-grained filtering module: Based on the unscented Kalman filter, the module realizes the smoothing of nonlinear motion time series data through unscented transformation and prediction-update mechanism. The input of this module is preprocessed data, and the output is filtered time series data. The data dimensions of input and output are consistent, and the noise contained in the preprocessed data is suppressed and smoothed.

[0065] Specifically, determine the initial value of the system state vector, covariance matrix, state noise and observation noise matrix, consider the characteristics of the data set and adjust the weight factor of the matrix. If the value of a variable fluctuates greatly, the scale of the corresponding covariance matrix should be increased; generate a set of sigma points at each time step to approximate the distribution of the state space, and obtain the predicted observation value by inputting these sigma points into the nonlinear state transfer function; input the sigma points output by the state transfer function into the observation function to reconstruct the observation value, and obtain the state observation value and covariance matrix. According to the difference between the observed value and the predicted value, the state estimate and covariance are updated through the Kalman gain to correct the state estimate.

[0066] Masking mechanism: Capture and mask the key timestamps for classification of filtered time series data, including taking the attention score calculated by the Transformer encoder network as the basis, and masking a fixed proportion of key timestamps through the masking module. The input of the masking mechanism is the output of the coarse-grained filtering module, that is, the filtered data, and the output is the masked data of the filtered data. The masking mechanism obtains the index corresponding to the timestamp with a higher score based on the attention score output by the Transformer encoder network; through sampling randomization, a fixed proportion of key timestamps are randomly selected from the selected index to mask, perceive the classification task while ensuring that the feature extraction network does not capture local patterns.

[0067] Multi-scale feature extraction network: Using the multi-scale feature extraction network and relying on the time slicing mechanism to construct inputs of different scales, we extract multi-scale robust time features from multivariate time series and perform position encoding to complete action recognition. The input of the multi-scale feature extraction network is the masked filtered data, and the output is passed through the average pooling layer and softmax to obtain the probability distribution of each action category, and then the action label corresponding to the sample can be obtained.

[0068] The present invention combines a mask mechanism and a multi-scale feature extraction network. The attention score output by the Transformer encoder network in the mask mechanism provides a reference for the downstream classification task for the mask operation. According to the attention score, it is determined which timestamps are more critical in the classification task, and the information of certain timestamps is destroyed to obtain mask data. The reference is to enable the Transformer encoder network to have the perception ability of the downstream classification task through each training number of training, and the mask mechanism and the Transformer encoder network of the multi-scale feature extraction network share the same architecture. For the time slicing mechanism in the multi-scale feature extraction network, the mask timestamp will perceive the context information in the slice, and restore and predict the information of the mask timestamp; this destruction-recovery mechanism can improve the robustness of the model. The slice lengths of the time slicing mechanism are different. After the output of the previous layer of the feature extraction network, the slices will continue to be sliced ​​through the time slicing mechanism. At this time, the time dimension is smaller, contains more global information and lower computational cost. Therefore, the system constructed by the present invention considers the deployment cost and provides a short-time, labor-saving and effective action recognition solution.

[0069] The present invention selects five advanced deep learning models mentioned in the prior art: TARNet, Todynet, FormerTime, ConvTran, and ShapeFormer. All models are run on the same experimental platform. Accuracy is selected as the performance indicator, and the evaluation indicator of TARNet in the prior art is referenced. The experiment of this embodiment takes the best performance of three experiments on each data set. The performance of the method proposed in the present invention and the model selected by the comparative experiment on some UEA data sets is shown in Table 1, where bold indicates that the model is the best than other models, and underline indicates that the model is the second best than other models.

[0070] Ours 1-to-1 Wins / Draws / Losses: The number of datasets where our method is better than / same as / worse than the corresponding baseline, which facilitates a one-to-one comparison between our method and the model selected in the comparative experiment.

[0071] Average rank: The average rank of a model across all datasets. The lower the average rank, the better the model classification performance.

[0072] Table 1 Performance of the proposed method and the models selected in the comparative experiment on some UEA datasets

[0073] In Table 1, bold represents the method or model with the best performance in this dataset, and underlined represents the method or model with the second best performance in this dataset.

[0074] To verify the effect of the present invention, the ablation experiment of the mask mechanism is conducted to compare two strategies: no mask mechanism and mask mechanism, as shown in Table 2. The ablation experiment of time slicing in the multi-scale feature extraction network is conducted to compare two strategies: no time slicing mechanism and time slicing mechanism, as shown in Table 3. At the same time, the present invention also verifies the effectiveness of the multi-scale feature extraction method. The present invention compares three strategies: the influence of using one layer (no layered architecture), two layers, and three layers on the accuracy of action recognition, as shown in Table 4. The present invention uses the action recognition accuracy to measure the effect of the present invention.

[0075] Table 2 Ablation experiment of mask mechanism

[0076] Table 3 Ablation experiment of time slices in multi-scale feature extraction network

[0077] Due to the interaction between the mask mechanism and time slices, the action recognition performance is greatly improved. The addition of this destruction-recovery mechanism designed by the method described in the present invention brings strong robustness to the method described in the present invention. The mask mechanism will timestamp the input sequence according to the criticality of the timestamp to the classification task. The time slice mechanism of the multi-scale feature extraction network captures useful classification information in the slice, and restores the information of the mask timestamp in the process of aggregating the slice information, which brings powerful sequence modeling capabilities and significant improvement of action recognition performance.

[0078] Table 4 Ablation experiments of multi-scale feature extraction methods

[0079] The present invention uses different levels from 1 to 3, while keeping the number of multi-head attention modules the same and other conditions unchanged. The results show that the performance is highest when the number of levels is equal to 3, and the performance at multiple scales is always better than that at single scale.

[0080] The above description is only a preferred embodiment of the present invention and is not intended to be a further limitation of the present invention. All equivalent changes made using the contents of the present specification and drawings are within the protection scope of the present invention.

Claims

1. A classification-aware end-to-end action recognition method, characterized in that: The following steps are involved: S1, collect time series data of human body movements and perform data preprocessing; S2, performing coarse-grained filtering on the preprocessed time series data based on an unscented Kalman filter to obtain filtered time series data; S3, using a masking mechanism to capture the classification key timestamps and mask the filtered time series data; S4. Utilize a multi-scale feature extraction network, build inputs of different scales based on the time slicing mechanism, and perform position encoding to extract multi-scale robust temporal features from multivariate time series to complete action recognition.

2. The classification-aware end-to-end action recognition method according to claim 1, characterized in that: In S1, data preprocessing is as follows: S11. Design an improved linear interpolation strategy to complete the missing value filling of time series data; S12, z-score standardization is performed on multivariate time series, and the value range of all multivariate time series is mapped to the interval [0,1]; S13. Align each multivariate time series to the same length.

3. The classification-aware end-to-end action recognition method according to claim 2, characterized in that: In S11, the improved linear interpolation strategy is as follows: S111. For missing values ​​at the beginning of the sequence, fill them in by traversing subsequent data and using the first valid value; S112. For missing data in the middle of the sequence, linear interpolation method is used to fill in the missing data by combining the previous and next valid data points; S113. When there are missing values ​​at the end of the sequence, linear regression prediction is used to fill in the missing values ​​according to the changing trend of the valid data before the end, so that the action sequence is continuous and complete.

4. The classification-aware end-to-end action recognition method according to claim 1, characterized in that: S2 is specifically: S21, setting the initial state value, covariance matrix, state noise and observation noise matrix of each multivariate time series at the first moment; S22, according to the currently estimated state mean and covariance matrix, generate a set of Sigma points to represent the state space distribution and perform state prediction; S23, mapping the Sigma point to the observation space through coarse-grained filtering to obtain the observation value of the state space distribution; S24, calculating the weighted average of all Sigma points to obtain the updated state prediction value and covariance matrix; S25. Based on the difference between the observed value and the predicted value, the state estimate and covariance are updated through the Kalman gain to correct the state estimate.

5. The classification-aware end-to-end action recognition method according to claim 1, characterized in that: In S3, the mask mechanism includes a key timestamp perception network and a mask module, which takes the filtered time series data output by S2 as input and calculates the attention score vector consistent with the time length dimension. , the mask module relies on the attention score vector Masking of key timestamps for a fixed ratio; The backbone network of the key timestamp perception network is a Transformer encoder network, including a time domain projection module and a multi-head attention mechanism; the filtered time series data is mapped into query q, key k, and value v matrices, and the variable dimensions are converted into dimensions that are adapted to the model; the time domain projection module reduces the time dimension of the key k and value v matrices to R times the original dimension; the multi-head attention mechanism receives the query q matrix and the projected k and v matrices to capture the attention score consistent with the time dimension , input Attn into the mask module; l is the length of the time series; The mask module sums the first dimension of the received Attn matrix to obtain , combined with sampling randomization, select the attention score vector Before Timestamps from Randomly select from timestamps timestamp and mask, where .

6. The classification-aware end-to-end action recognition method according to claim 5, characterized in that: The multi-scale feature extraction network is a three-layer architecture, where the output of the previous layer is the input of the next layer, and each layer includes a time slicing mechanism and a Transformer encoder network; In the first layer of the multi-scale feature extraction network, the input of the time slicing mechanism is mask data. The time slicing mechanism constructs inputs of different scales through its built-in time domain convolutional network, so that the mask timestamp compensates its own information according to the context in the slice; The input of the time slicing mechanism of the second and third layers of the multi-scale feature extraction network is the output of the previous layer of Transformer encoder network; The Transformer encoder network includes a time domain projection module, a multi-head attention mechanism, and a feedforward neural network FFN; for the ath layer, 1≤a≤3, the input of the Transformer encoder network is the shallow encoding of the slice output by the time slicing mechanism, and the shallow encoding is mapped to the query q, key k, and value v matrix; the time domain projection module reduces the time dimension of the key k and value v matrix to the original dimension times; the multi-head attention mechanism receives the query q matrix and the projected k and v matrices to capture the long-range dependencies of the slices; the multi-head attention mechanism and FFN respectively use residual connections to prevent the gradient from disappearing, the output of the multi-head attention mechanism and the shallow encoding of the slice output by the time slicing mechanism are fused and input to FFN, and the output of FFN and the feature fusion of the input of FFN are used as the output of the Transformer encoder network; the features of the FFN output of the last layer are passed through the average pooling layer to obtain the probability distribution of the action category; the cross entropy loss is used as the loss function for training the Transformer encoder network.

7. The classification-aware end-to-end action recognition method according to claim 6, characterized in that: S4 is specifically: S41, constructing the shallow encoding of the current layer based on the time slicing mechanism to complete the time slicing; S42, a one-dimensional zero-filled convolution operation is performed on the slice shallow code to give the slice shallow code a context-aware position attribute, i.e., position coding, and the slice shallow code is fused with the slice shallow code as the input of the Transformer encoder; S43, the shallow encoding of the current layer is mapped to the query q, key k, value v matrix of the Transformer encoder; S44. Design a time domain projection mechanism to compress the time dimension of key k and value v; S45, the multi-head attention mechanism receives the query q matrix and the projected key k and value v matrices to capture the long-range dependencies of the slices; S46, concatenate the attention scores calculated by each head and reshape them to the dimension of the input sequence, and send the reshaped attention features to the feed-forward neural network FFN; S47. The output of the feedforward neural network of the last layer of Transformer encoder is passed through the average pooling layer and softmax to obtain the probability distribution of each action category.

8. A classification-aware end-to-end action recognition system, characterized in that: For implementing the method according to any one of claims 1 to 7, the system comprises: Data preprocessing module: obtain time series data and perform data preprocessing; Coarse-grained filtering module: Based on unscented Kalman filtering, it realizes smoothing of nonlinear motion time series data through unscented transformation and prediction-update mechanism; Masking mechanism: Capture and mask the key timestamps of classification for filtered time series data, including masking a fixed ratio of key timestamps based on the attention scores calculated by the Transformer encoder network. Multi-scale feature extraction network: Relying on the time slicing mechanism to construct inputs of different scales, it extracts multi-scale robust temporal features from multivariate time series and performs position encoding to complete action recognition.